Compare commits

..

2 Commits

Author SHA1 Message Date
Dai Ha 5da912ae8e CB-633: sshAuthSock: block must win over an SSH_AUTH_SOCK allow: entry, and report the credential gap on the non-zsh fallback path
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 1m39s
Fixes two defects in PR #174:

1. applyEnvironmentAllowListPolicy unioned memberCredentials.allow directly into the
   derived allow-list, so an operator who wrote both `sshAuthSock: block` and
   SSH_AUTH_SOCK on `allow:` (the live fleetd.yaml shape) got the block silently
   defeated. SSH_AUTH_SOCK is now excluded from that union and governed only by
   sshAuthSock:, with a one-time WARN when the two controls conflict.

2. The non-zsh login-shell fallback branch dropped its logCredentialGap call, so an
   allow-list spawn on a non-zsh shell produced no gap report at all — exactly the
   weakest, overlay-only path that most needs one. The call is restored, and
   logCredentialGap now picks its WARN/INFO wording from whether a scrub-derived
   allow-list set was actually computed (effectiveAllowed != null) rather than from
   the policy alone, so this path correctly gets the WARN ("inherited UNBLOCKED")
   wording instead of the allow-list INFO wording.

3. credentialGapLogged was one shared AtomicBoolean guarding both report kinds;
   split into allowListGapLogged/unprotectedGapLogged so an early INFO on one spawn
   can no longer suppress a later WARN on a live-reloaded policy.

Each fix is proven by reverting it and observing the corresponding test fail, then
restoring it.
2026-08-31 08:44:30 +07:00
Dai Ha f6c150e99a CB-633: honor explicit member credential keeps
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m19s
2026-08-28 05:25:46 +07:00
55 changed files with 624 additions and 6197 deletions
+2 -13
View File
@@ -113,19 +113,8 @@ PY
**What this tier cannot see:** it proves facts only about the Mac daemon at `127.0.0.1:8765`.
It cannot show the fleet01 daemon, broker queue depth, or broker consumers. The fleet01 REST service
at `10.10.20.13:8765` is not reachable from the Mac. Say this in the report rather than omitting
fleet01.
**But fleet01 IS reachable over SSH — checked 2026-08-28.** An older version of this line said SSH
was denied. That is true only for the user `dai.ha`. The host alias `fleet01` maps to user `ltms`,
and `ssh fleet01` works with key auth:
```bash
ssh -o BatchMode=yes -o ConnectTimeout=6 fleet01 'echo $(id -un)@$(hostname)'
```
So fleet01's daemon PID, uptime, jar and `/healthz` **can** be reported — over SSH, not over REST.
Do that rather than writing `not reachable`. `ltms` also has passwordless sudo there.
at `10.10.20.13:8765` is not reachable from the Mac, and SSH as `dai.ha@10.10.20.13` is denied.
Say this in the report rather than omitting fleet01.
## 3. Tier 2 — the shared broker (run when management access exists)
-208
View File
@@ -1,208 +0,0 @@
# Wiki audit for #168
**Source checked:** `.wiki-snapshot/` at `68e32c6` (2026-08-31). I did not use
`wiki/`. Code references below are from the current `fleetd` source tree. A quoted
line is a concrete claim that needs correction, unless the table says `KEEP`.
| Page | Verdict | One-line reason |
|---|---|---|
| `Home.md` | REVISE | Good overview, but it still names the retired product. |
| `_Sidebar.md` | REVISE | The heading still says `claude-bridge`. |
| `1-Architecture.md` | REBUILD | Its component contract mixes current names with removed tools, routes, and planned backends. |
| `2-Message-Server.md` | REBUILD | The claimed MCP schema, mount command, REST/SSE surface, and fallback paths are pre-build design. |
| `3-Approaches.md` | REVISE | Useful research history, but it presents unbuilt AgentAPI as a selectable fallback. |
| `4-Setup.md` | RETIRE | It is an intentional stub that only redirects to chapter 13. |
| `5-Operations.md` | RETIRE | It is an intentional stub that only redirects to chapter 13. |
| `6-Team.md` | REBUILD | It teaches role-addressed sends and a Claude-only team model that the shipped API does not have. |
| `7-Use-Cases.md` | REBUILD | Its flagship flow depends on removed `ccs` profiles and removed send parameters. |
| `8-Roadmap.md` | REBUILD | It is a historical plan, but it presents old implementation choices and planned work as the current stack. |
| `9-Implementation.md` | REBUILD | Its package, class, endpoint, and outcome map has drifted from the source. |
| `10-Cross-Host-Messaging.md` | REVISE | It labels most federation work proposed, but misses the shipped `coordinator:` lead channel. |
| `11-Features.md` | REVISE | It is the right catalogue, but code-path names are old and it misses the second-herdr-daemon capability. |
| `12-Claude-to-OpenCode.md` | REVISE | The porting guide is mostly current, but calls the product and spawned-member path a bridge. |
| `13-User-Guide.md` | REVISE | It is the best operator page, but needs the product rename and the second-herdr-daemon setup. |
## Pages needing work
### `Home.md` — REVISE
- Quote: `# claude-bridge` (line 1) and `` `claude-bridge` keeps`` (line 11).
The product is `fleet` / `fleetd`. The MCP server identifies itself as `fleet` in
`fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:313-315`.
- Quote: `AgentAPI ... swappable fallback injector` (lines 73-76).
There is no AgentAPI implementation under `fleetd/src/main/java`; the actual
launchers are selected by `Profile.kind` in
`fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:265-270`.
### `_Sidebar.md` — REVISE
- Quote: `### 📖 claude-bridge` (line 1).
Rename it to `fleet`. `FleetMcp` registers the current product-facing tool set at
`fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:301-326`.
### `1-Architecture.md` — REBUILD
- Quote: `` `claude-bridge` lets`` (line 3). The product was renamed; the MCP
server name is `fleet` (`FleetMcp.java:313-315`).
- Quote: ``fleet_read`` in the tool list (line 102). No such tool is registered.
The complete registered list is `fleet_send` through `fleet_whoami` at
`FleetMcp.java:301-326`; `fleet_read` is absent.
- Quote: `SSE (GET /events)` (line 143). `FleetApp.build()` registers no `/events`
route; its routes are listed at `FleetApp.java:143-159`.
- Quote: `Redis Streams / NATS JetStream, or an embedded queue` (line 106).
The shipped durable inbox is AMQP, configured by `broker`, at
`FleetConfig.java:49-50` and `FleetConfig.java:655-714`.
- Quote: `AgentAPI (fallback)` (line 107). No AgentAPI adapter exists; shipped
launcher kinds are `claude-code` and `opencode` (`FleetConfig.java:265-270`).
### `2-Message-Server.md` — REBUILD
- Quote: `claude mcp add --transport http bridge http://127.0.0.1:8080/mcp`
(line 67). The daemon defaults to port `8765` in `FleetConfig.java:183-187`,
and identifies its server as `fleet` at `FleetMcp.java:313-315`.
- Quote: ``fleet_send(message, target?, {block, timeout_seconds, auto_spawn,
turn_id})`` (line 80). The real parameters are `sessionId`, `content`,
`timeoutMs`, `wait`, `turnId`, and `coordId` (`FleetMcp.java:1096-1108`).
- Quote: ``fleet_read(target, source)`` (line 85). It is not registered; see the
complete registration at `FleetMcp.java:301-326`.
- Quote: `docs/MCP-Contract.md ... normative` (lines 87-88). That is not a valid
reference: only §6 is current, as the current operator guide itself says at
`.wiki-snapshot/13-User-Guide.md:466`.
- Quote: `SSE (GET /events)` (line 45). No route exists in the built REST surface,
`FleetApp.java:143-159`.
### `3-Approaches.md` — REVISE
- Quote: `AgentAPI ... remains a swappable fallback injector` (lines 78-84).
It was never built. The shipped adapter selection is only `claude-code` or
`opencode` (`FleetConfig.java:265-270`). Keep it as discarded research, not an
operational fallback.
- Quote: `claude-bridge` (line 109). Rename the product to `fleet`; the runtime
package is `dev.ltms.fleet`, for example `FleetMcp.java:1`.
### `4-Setup.md` — RETIRE
It is a 25-line redirect and says its procedure was never written (lines 3-9).
Chapter 13 is the maintained install procedure. Keeping a second navigation page
adds no working documentation.
### `5-Operations.md` — RETIRE
It is a 35-line redirect and says its runbook was never written (lines 3-14).
Chapter 13 now owns run and recovery instructions.
### `6-Team.md` — REBUILD
- Quote: `fleet_send {role: w-claude, prompt: A}` (line 98). `fleet_send` accepts
`sessionId` and `content`, not `role` or `prompt` (`FleetMcp.java:1096-1108`).
- Quote: `some on Claude, some on the remote local LLM` (lines 3-5) and `Every
worker is ... Claude Code` (line 25). `opencode` is a first-class launcher kind,
not a Claude worker (`FleetConfig.java:265-270`).
- Quote: `fleetd's concurrency policy` (line 121). The configured capacity control
is per-profile `maxLoad` (`FleetConfig.java:251-264`), not the role routing model
described here.
### `7-Use-Cases.md` — REBUILD
- Quote: `ccs profile` (line 10), `ccs + herdr` (line 22), and `ccs-spawn`
(line 45). The configuration has `profiles` and `fleet`, not `ccs`:
`FleetConfig.java:34-58` and `FleetConfig.java:81-101`.
- Quote: `fleet_send({"to", "kind", "body", "block"})` (lines 55-62).
None of those are the shipped send parameters. The schema is
`FleetMcp.java:1096-1108`.
- Quote: `fleet_list() → { "profiles": ... }` (lines 74-80). `fleet_list` is a
roster view; `fleet_profiles` is the configured-backend view, as registered at
`FleetMcp.java:307-311` and described at `FleetMcp.java:1176-1182`.
### `8-Roadmap.md` — REBUILD
- Quote: `Java 21+` (line 43). The current project guidance and source use Java 25;
the `FleetConfig` source itself uses Java 25 unnamed lambda parameters, for
example `FleetConfig.java:102`.
- Quote: `herdr 0.7.0 / protocol 14` (line 46). The current REST health endpoint
reports the live protocol returned by herdr (`FleetApp.java:240-244`), while the
current operator guide records protocol 19 at
`.wiki-snapshot/13-User-Guide.md:76-85`.
- Quote: `ccs <profile> claude` and `ccs env <profile>` (lines 47-48). Shipped
configuration uses `Profile` records and launcher `kind`,
`FleetConfig.java:313-330` and `FleetConfig.java:265-270`.
- Quote: `Redis Streams via Lettuce` (line 50). The actual durable inbox is AMQP
`broker`, `FleetConfig.java:655-714`.
### `9-Implementation.md` — REBUILD
- Quote: `rest.FleetdApp` and `mcp.BridgeMcp` (lines 29-30). The classes are
`rest.FleetApp` and `mcp.FleetMcp` (`FleetApp.java:46`; `FleetMcp.java:67`).
- Quote: `dev.ltms.fleetd` (line 67). The source package is `dev.ltms.fleet`
(`FleetMcp.java:1`).
- Quote: `WorkerPresence` (line 110). The current class is `MemberPresence`, as
imported and used by `FleetMcp` at `FleetMcp.java:12` and `465-469`.
- Quote: the outcome list ending in `STALE_TURN` (lines 128-131). The code also
has `BACKEND_EXHAUSTED` (`FleetMcp.java:550-554`) and async `ASKING` handling
(`FleetMcp.java:664-668`).
- Quote: `FleetdApp` (line 207) and `FleetdConfig` (line 211). These names do not
resolve; current classes are `FleetApp` and `FleetConfig`.
### `10-Cross-Host-Messaging.md` — REVISE
- Quote: the chapter says the cross-host fabric is proposed except for the
single-host inbox (lines 3-8). Cross-host **lead-to-lead** delivery shipped:
`fleet_send` accepts `coordId` (`FleetMcp.java:1094-1107`) and publishes it at
`FleetMcp.java:616-641`; configuration has `coordinator` at
`FleetConfig.java:74-78` and `99-101`.
- Quote: `bridge.dlx` (line 90). This product name is stale. The shipped lead path
uses `LeadChannel`, not the proposed exchange flow (`FleetMcp.java:95-96` and
`616-641`). Keep the proposed federation design, but add a clear shipped/proposed
boundary for CB-637.
### `11-Features.md` — REVISE
- Quote: `mcp/BridgeMcp` (line 22), `config/FleetdConfig` (lines 25-27), and other
index references. These paths no longer resolve; the source classes are
`mcp/FleetMcp` (`FleetMcp.java:67`) and `config/FleetConfig`
(`FleetConfig.java:81`).
- Quote: `fleet_whoami` returns only `primary` or `worker` (lines 99-100).
It also returns `architect` (`FleetMcp.java:1235-1244`).
- The page needs the missing separate member-herdr-daemon feature listed below.
### `12-Claude-to-OpenCode.md` — REVISE
- Quote: `same bridge mount` (line 5) and `a bridge-spawned worker` (line 94).
Rename the product path to `fleet`. The daemon exposes the MCP server as `fleet`
(`FleetMcp.java:313-315`), and profiles select OpenCode with `kind: opencode`
(`FleetConfig.java:332-335`).
- Quote: the sample mount name is `fleetd` (line 67). The server name is `fleet`;
update the sample to avoid teaching a second product name.
### `13-User-Guide.md` — REVISE
- Quote: `The bridge is the only channel` (line 63). The invariant is correct, but
the product term needs the `fleet` rename. The daemon's MCP server name is
`fleet` (`FleetMcp.java:313-315`).
- Quote: it describes one herdr socket (lines 72-85). It needs the optional
`memberHerdrSocket` setup and two-daemon health meaning. The config key is in
`FleetConfig.java:34-37`, and `/healthz` checks both daemons when configured at
`FleetApp.java:210-245`.
## MISSING
`11-Features.md` has a body section for **routing members through a separate herdr daemon**
(`## memberHerdrSocket`, line 2174), but **no row in the index table** at the top of the page
(lines 20-95). That table is how the page is meant to be read, so a capability absent from it is
effectively undiscoverable. Lead note: this is my own omission — I added the section on 2026-08-31
and did not add the matching row. Fixed in the wiki at `68e32c6`'s successor.
The original audit stated the feature had no entry at all. That was wrong: the section exists. The
gap is the index row. Recorded here rather than silently corrected, because the difference matters —
"undocumented" and "documented but unindexed" are different jobs.
Evidence for the feature itself: `FleetConfig.java:34-37` and `FleetApp.java:103-115`, `210-245`,
and `247-263`.
## Audit method and coverage
I checked all 15 pages. I checked concrete tool, route, config, class, file, and
product-name claims claim-by-claim on 11 pages: Home, Sidebar, 1, 2, 4, 5, 6, 7, 9,
11, and 13. I skimmed the remaining four long historical or research pages (3, 8, 10,
12), then checked their concrete claims that affect the verdict. This is an audit of
the supplied snapshot, not a wiki rewrite.
-30
View File
@@ -114,18 +114,6 @@ bind:
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
# Optional socket for member panes. Omit this to use herdrSocket for both leads and members.
# memberHerdrSocket: /Users/member/.config/herdr/herdr.sock
# fleetd #213: the login shell the member OS user (memberHerdrSocket above) actually runs. ONLY
# read when memberHerdrSocket is set — fleetd's own $SHELL says nothing about a pane running
# under a different OS user, and there is no channel to ask herdr for that user's shell, so this
# must be told rather than guessed. Absent, blank, or anything not ending in "zsh" is treated the
# same as "not zsh": the memberCredentials.policy: allow-list ZDOTDIR scrub (see worktreeGroup
# below) is skipped in favour of the weaker CB-596 sentinel overlay — a degraded control, never a
# refusal to spawn. When memberHerdrSocket is absent this key is never consulted at all.
# memberLoginShell: /bin/zsh
# How member sessions are spawned. Define one or more named profiles (backends) under
# `profiles`; each key is the profile name (also the ccs profile). A profile says only WHICH
# BACKEND — model, CLI adapter, credentials, cost. It says nothing about what a member spawned on
@@ -643,24 +631,6 @@ guard:
# to a sibling directory of the repo root.
# worktreeRoot: /Users/me/src/.bridged-worktrees
# Worktree group sharing (fleetd #185 stage 3). OPTIONAL, off by default. Names an OS group
# that a provisioned worktree's repo is made group-writable for (git config
# core.sharedRepository group, plus a one-time chgrp/chmod/setgid fix-up), so a member spawned
# under a DIFFERENT OS user (see memberHerdrSocket) can write its own worktree, its
# per-worktree git metadata, and its own commit objects — without it, every file GitWorktrees
# creates is owned by fleetd's own uid and unwritable by another user.
# CAUTION: this isolates credentials, not the repository — a member in the group can still
# write the operator's git objects and refs in the shared repo. The operator running fleetd
# must already be a member of the named group, or every provisioning spawn fails loudly.
#
# fleetd #213: this is also the ONE group the memberCredentials.policy: allow-list ZDOTDIR scrub
# reuses when memberHerdrSocket is set — deliberately not a second config key. Under
# memberHerdrSocket, the scrub directory is generated under worktreeRoot (never java.io.tmpdir,
# which the member OS user cannot reach) and shared read-only with this group. If worktreeGroup
# is unset while memberHerdrSocket is set, the scrub cannot be guaranteed reachable by the member,
# so fleetd falls back to the weaker CB-596 sentinel overlay instead (a WARN names the gap).
# worktreeGroup: fleet-workers
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
-18
View File
@@ -28,7 +28,6 @@
<testcontainers.version>1.20.4</testcontainers.version>
<commons-compress.version>1.27.1</commons-compress.version>
<commons-lang3.version>3.18.0</commons-lang3.version>
<sqlite-jdbc.version>3.53.4.0</sqlite-jdbc.version>
</properties>
<!--
@@ -45,12 +44,6 @@
3.0-rc5; bumping Jackson 3 to the patched 3.2.x breaks the SDK (annotation mismatch).
Only the loopback /mcp endpoint parses this JSON, from trusted local Claude clients.
The 11.0.23 -> 11.0.25 bump did clear jetty CVE-2024-8184 (5.9) and CVE-2024-6763.
fleetd #206: org.xerial:sqlite-jdbc 3.53.4.0 (added for OpenCodeSessionDiscovery) — the
only known advisory against this artifact is CVE-2023-32697 (RCE via an attacker-controlled
JDBC URL), fixed in 3.41.2.2; 3.53.4.0 is well past that fix and OSV.dev reports no open
advisory against it. Checked via the OSV.dev API (no Mend.io/JetBrains IDE MCP mount
available from this worktree) on 2026-08-31.
-->
<!-- Force the latest patched Jetty 11.x across all Javalin-pulled Jetty modules (no version
@@ -130,17 +123,6 @@
<version>${amqp.version}</version>
</dependency>
<!-- fleetd #206: opencode moved its session store from a JSON tree to SQLite
(opencode.db). This is the JDBC driver OpenCodeSessionDiscovery uses to read it
read-only. Ships bundled native libraries (linux/mac/windows, several archs), so it
is a heavier jar than most deps here — see the pom's dependency-security note below
for the size/CVE tradeoff actually measured. -->
<dependency>
<groupId>org.xerial</groupId>
<artifactId>sqlite-jdbc</artifactId>
<version>${sqlite-jdbc.version}</version>
</dependency>
<!-- Logging -->
<dependency>
<groupId>org.slf4j</groupId>
+19 -32
View File
@@ -7,7 +7,6 @@ import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.herdr.LeadTabScanner;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.herdr.PaneLocator;
@@ -150,12 +149,9 @@ public final class Fleetd {
: UnixSocketHerdrClient.defaultSocketPath();
UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect(socket, new com.fasterxml.jackson.databind.ObjectMapper());
UnixSocketHerdrClient memberHerdr = cfg.memberHerdrSocket() != null && !cfg.memberHerdrSocket().isBlank()
? UnixSocketHerdrClient.connect(Path.of(cfg.memberHerdrSocket()), new com.fasterxml.jackson.databind.ObjectMapper())
: herdr;
AtomicReference<Supplier<Map<String, String>>> leadsRef = new AtomicReference<>(Map::of);
HerdrRouter router = new HerdrRouter(herdr, memberHerdr,
target -> leadsRef.get().get().containsKey(target));
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
@@ -173,18 +169,18 @@ public final class Fleetd {
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(new ClaudeCodeLauncher(router.memberAgents(), router.memberSpaces(), guard,
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(),
() -> config.get().memberCredentials(), null, config::get));
() -> config.get().memberCredentials()));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(router.memberAgents(), router.memberSpaces(),
adapters.add(new OpenCodeLauncher(agents, spaces,
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(),
() -> config.get().memberCredentials(), config::get));
() -> config.get().memberCredentials()));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
@@ -224,7 +220,7 @@ public final class Fleetd {
contextCap = cfg.lifecycle().contextCap();
}
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot(), cfg.worktreeGroup()),
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()),
System::nanoTime, contextCap, clearAfterTurn);
liveCountRef.set(profileName -> (int) sessions.roster().stream()
.filter(s -> profileName.equals(s.profile()))
@@ -275,7 +271,6 @@ public final class Fleetd {
// operational cadence, not identity, so there is no correctness reason to give every
// lead its own scanner.
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
// This must use the lead daemon: scanning member tabs would demote the lead to a worker.
leads = new LeadTabScanner(herdr, tabToName, Set.of(),
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), System::nanoTime);
log.info("lead scan: tabs {} host a lead (rescan every {}s, shared fleet space)",
@@ -283,14 +278,13 @@ public final class Fleetd {
} else {
leads = () -> leadTerminals;
}
leadsRef.set(leads);
// CB-558: start any declared lead that is not already running. After the scanner is built,
// because both read the same tab labels and the ordering makes that dependency visible; and
// only when herdr answered, because the launcher's whole safety property is that it can
// count live leads first — it must never guess and risk a second orchestrator.
if (herdrUp && !leaders.isEmpty()) {
int launched = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg).ensureLeads();
int launched = new LeadLauncher(agents, spaces, cfg).ensureLeads();
if (launched > 0) {
log.info("lead auto-launch: {} lead(s) started", launched);
}
@@ -346,7 +340,6 @@ public final class Fleetd {
+ "BACKEND_EXHAUSTED): {}", credentialId,
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
});
AgentControl agents = router.memberAgents();
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns, exhaustionSink);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
@@ -387,10 +380,9 @@ public final class Fleetd {
sessions.onTurnFailed(target);
}
};
Predicate<String> deliverable = deliverableTo(presence, leads);
Injector injector = new Injector(router, turnListener, deliverable,
Injector injector = new Injector(agents, turnListener, deliverableTo(presence, leads),
presence::forget);
StatusPoller poller = new StatusPoller(router, injector, Injector.POLL_INTERVAL_MILLIS);
StatusPoller poller = new StatusPoller(agents, injector, Injector.POLL_INTERVAL_MILLIS);
poller.start();
// CB-307: reply inbox. A broker: block selects the AMQP-backed durable adapter; absent (or
@@ -427,7 +419,7 @@ public final class Fleetd {
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
// CB-512: the push loop takes it too, so nudge outcomes (delivered|exhausted) are counted.
Metrics metrics = FleetMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, router.leadAgents(), replyInbox,
var pushLoop = new ReplyPushLoop(primaryRegistry, agents, replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
@@ -437,7 +429,7 @@ public final class Fleetd {
Thread.ofVirtual().name("bridge-heartbeat-").unstarted(r));
if (cfg.leadHeartbeat() != null) {
var hb = cfg.leadHeartbeat();
heartbeat = new LeadHeartbeatLoop(primaryRegistry, router.leadAgents(), replyInbox, sessions::roster,
heartbeat = new LeadHeartbeatLoop(primaryRegistry, agents, replyInbox, sessions::roster,
pushLoop, heartbeatScheduler, System::nanoTime,
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
metrics);
@@ -446,7 +438,7 @@ public final class Fleetd {
heartbeat = null;
heartbeatScheduler.shutdownNow();
}
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
@@ -498,11 +490,8 @@ public final class Fleetd {
// MCP server face (CB-105): fleet_send/fleet_reply/fleet_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
// CB-185: a caller's pane can live on either daemon (a lead's on the lead daemon, a
// member's on the member daemon) — search both, lead first. Collapses to one scan when
// memberHerdrSocket is unset (herdr == memberHerdr).
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr, memberHerdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
new PaneLocator(herdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
@@ -548,7 +537,7 @@ public final class Fleetd {
if (leadMailbox != null) {
var leadCoordScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-leadcoord-").unstarted(r));
leadCoordLoop = new LeadCoordLoop(leadMailbox, router.leadAgents(), leads, leadCoordScheduler,
leadCoordLoop = new LeadCoordLoop(leadMailbox, agents, leads, leadCoordScheduler,
LEAD_COORD_INTERVAL_MS);
leadCoordLoop.start();
leadCoordSchedulerRef = leadCoordScheduler;
@@ -599,13 +588,11 @@ public final class Fleetd {
log.debug("lead mailbox close: {}", e.toString());
}
}
router.close();
herdr.close();
}));
// CB-185: give FleetApp both daemons — /healthz must require both to answer and
// GET /sessions must merge across both, or a down/unpolled member daemon is invisible.
Javalin app = new FleetApp(herdr, memberHerdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics, deliverable).build();
Javalin app = new FleetApp(herdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics).build();
app.start(cfg.bind().host(), cfg.bind().port());
log.info("fleetd listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
@@ -71,7 +71,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
/** Keys that cannot change under a running daemon — see the class doc. */
private static final Set<String> COLD_KEYS =
Set.of("bind", "herdrSocket", "memberHerdrSocket", "broker", "auth");
Set.of("bind", "herdrSocket", "broker", "auth");
private final Path path;
private final AtomicReference<FleetConfig> current;
@@ -190,9 +190,6 @@ public final class ConfigRef implements Supplier<FleetConfig> {
if (!Objects.equals(old.herdrSocket(), fresh.herdrSocket())) {
changed.add("herdrSocket");
}
if (!Objects.equals(old.memberHerdrSocket(), fresh.memberHerdrSocket())) {
changed.add("memberHerdrSocket");
}
if (!Objects.equals(old.broker(), fresh.broker())) {
changed.add("broker");
}
@@ -32,8 +32,7 @@ import java.util.Set;
* silently dropping a whole block is indistinguishable from honouring it.
*
* @param bind REST/MCP listen host:port
* @param herdrSocket path to the lead herdr Unix socket ({@code null} → client default)
* @param memberHerdrSocket optional member herdr Unix socket ({@code null}/blank → lead socket)
* @param herdrSocket path to herdr's Unix socket ({@code null} → client default)
* @param profiles named backend profiles, keyed by profile name (multi-backend fleet). A
* profile answers <em>which backend</em> — model, CLI adapter, credentials,
* cost. It says nothing about what the member spawned on it is for; that is
@@ -76,32 +75,11 @@ import java.util.Set;
* stay on {@code broker}'s vhost). {@code null} → no lead mailbox is opened.
* Config parsing + accessors only — nothing here wires it into a live
* {@code LeadMailbox}; that is a separate ticket. See {@link Coordinator}.
* @param worktreeGroup optional OS group name (fleetd #185 stage 3) that makes a provisioned
* worktree's repo group-shared, so a member running as a different OS user
* (see {@code memberHerdrSocket}) can write its own worktree, its per-worktree
* git metadata, and its own commit objects. {@code null}/blank/empty ⇒ off,
* today's behaviour unchanged (every file stays owned by fleetd's own uid).
* <strong>This isolates credentials, not the repository</strong>: a member in
* the group can still write the operator's git objects and refs in the shared
* repo. See {@link dev.ltms.fleet.session.Worktrees#shareWithGroup}.
* @param memberLoginShell fleetd #213: the login shell the member's OS user actually runs, ONLY
* meaningful (and only ever read) when {@code memberHerdrSocket} is
* configured — that mode spawns member panes under a different OS user than
* fleetd's own process, so fleetd's own {@code $SHELL} says nothing about what
* that pane runs. There is no channel to ask herdr for another user's shell, so
* this must be told, never guessed. {@code null}/blank (or a value not ending
* in {@code zsh}) is treated the same as "not zsh": the {@code
* memberCredentials.policy: allow-list} ZDOTDIR scrub is skipped in favour of
* the CB-596 sentinel overlay — a degraded control, never a refusal to spawn.
* When {@code memberHerdrSocket} is NOT configured this field is never
* consulted at all; fleetd keeps reading its own {@code $SHELL}, exactly as
* before this field existed.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record FleetConfig(
Bind bind,
String herdrSocket,
String memberHerdrSocket,
Map<String, Profile> profiles,
Guard guard,
String worktreeRoot,
@@ -118,33 +96,7 @@ public record FleetConfig(
ConfigReload configReload,
Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials,
Coordinator coordinator,
String worktreeGroup,
String memberLoginShell) {
/** Back-compat form before the {@code memberLoginShell} key was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup, null);
}
/** Back-compat form before the {@code worktreeGroup} key was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, null, null);
}
Coordinator coordinator) {
/** Back-compat form before the {@code coordinator:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
@@ -153,9 +105,9 @@ public record FleetConfig(
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials) {
this(bind, herdrSocket, null, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, null, null);
configReload, quarantineCooldownSeconds, memberCredentials, null);
}
/** Back-compat form before the CB-596 {@code memberCredentials:} block was added. */
@@ -164,9 +116,9 @@ public record FleetConfig(
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds) {
this(bind, herdrSocket, null, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, null, null);
configReload, quarantineCooldownSeconds, null);
}
/** Default cooldown (CB-578 stage B) when {@code quarantineCooldownSeconds} is absent/non-positive. */
@@ -177,8 +129,8 @@ public record FleetConfig(
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, String placement, Auth auth) {
this(bind, herdrSocket, null, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, null, null, null, null);
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, null, null);
}
/** Back-compat form before the optional {@code health:} block was added. */
@@ -186,8 +138,8 @@ public record FleetConfig(
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, String placement, Auth auth, ConfigReload configReload) {
this(bind, herdrSocket, null, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, configReload, null, null, null);
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, configReload, null);
}
/** Back-compat form before the CB-578 stage B {@code quarantineCooldownSeconds} field was added. */
@@ -196,9 +148,9 @@ public record FleetConfig(
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload) {
this(bind, herdrSocket, null, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, null, null, null);
configReload, null);
}
/**
@@ -1368,10 +1320,10 @@ public record FleetConfig(
* {@code fleetd.yaml} itself is gitignored.
*/
static final Set<String> KNOWN_TOP_LEVEL_KEYS = Set.of(
"bind", "herdrSocket", "memberHerdrSocket", "profiles", "guard", "worktreeRoot",
"bind", "herdrSocket", "profiles", "guard", "worktreeRoot",
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
"memberCredentials", "coordinator", "worktreeGroup", "memberLoginShell");
"memberCredentials", "coordinator");
/** Load and validate config from {@code path}. */
public static FleetConfig load(Path path) {
@@ -1989,14 +1941,9 @@ public record FleetConfig(
: new MemberCredentials(null, List.of(), List.of());
// coordinator is left as-is, like broker/primary above: null keeps no LeadMailbox opened,
// and this ticket's Coordinator is config-only anyway (nothing yet reads it at startup).
// worktreeGroup is left as-is (fleetd #185 stage 3): null/blank is "off", and there is no
// sane non-null default — an OS group name is operator-specific.
// memberLoginShell is left as-is (fleetd #213), like worktreeGroup: null/blank is "not
// configured", and there is no sane non-null default — a member's login shell is
// operator-specific and only meaningful when memberHerdrSocket is also set.
return new FleetConfig(b, herdrSocket, memberHerdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
return new FleetConfig(b, herdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
quarantineCooldown, mc, coordinator, worktreeGroup, memberLoginShell);
quarantineCooldown, mc, coordinator);
}
/**
@@ -38,11 +38,6 @@ public final class AgentControl {
this.herdr = herdr;
}
/** The herdr daemon this control object sends its agent calls to. */
public HerdrClient herdr() {
return herdr;
}
/** One agent-targeted call, translating a terminal id to its pane id (retrying once fresh). */
private JsonNode agentCall(String method, String target, Map<String, Object> extra) {
String resolved = resolveTarget(target);
@@ -1,40 +0,0 @@
package dev.ltms.fleet.herdr;
import java.util.Objects;
import java.util.function.Predicate;
/** Routes lead operations and member operations to their owning herdr daemon. */
public final class HerdrRouter implements AutoCloseable {
private final HerdrClient lead;
private final HerdrClient member;
private final AgentControl leadAgents;
private final AgentControl memberAgents;
private final WorkspaceControl leadSpaces;
private final WorkspaceControl memberSpaces;
private final Predicate<String> isLead;
public HerdrRouter(HerdrClient lead, HerdrClient member, Predicate<String> isLead) {
this.lead = Objects.requireNonNull(lead, "lead");
this.member = member != null ? member : lead;
this.isLead = Objects.requireNonNull(isLead, "isLead");
leadAgents = new AgentControl(this.lead);
memberAgents = this.member == this.lead ? leadAgents : new AgentControl(this.member);
leadSpaces = new WorkspaceControl(this.lead);
memberSpaces = this.member == this.lead ? leadSpaces : new WorkspaceControl(this.member);
}
public AgentControl leadAgents() { return leadAgents; }
public WorkspaceControl leadSpaces() { return leadSpaces; }
public AgentControl memberAgents() { return memberAgents; }
public WorkspaceControl memberSpaces() { return memberSpaces; }
public AgentControl agentsFor(String targetId) { return isLead.test(targetId) ? leadAgents : memberAgents; }
HerdrClient leadClient() { return lead; }
HerdrClient memberClient() { return member; }
@Override
public void close() {
lead.close();
if (member != lead) member.close();
}
}
@@ -2,11 +2,7 @@ package dev.ltms.fleet.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.OptionalLong;
import java.util.Set;
/**
* Resolves which herdr pane a process belongs to — the herdr half of connection-based MCP
@@ -15,130 +11,46 @@ import java.util.Set;
* is calling without the worker sending anything spoofable.
*
* <p>herdr owns the PID→pane truth: {@code pane.process_info} reports each pane's {@code shell_pid}
* and foreground process PIDs. A pid that is neither of those directly — e.g. a grandchild a
* worker spawned, such as a {@code python3} or {@code curl} helper that opens its own MCP
* connection — is resolved by walking its ancestry (via {@link ParentResolver}) up to the root and
* matching any ancestor against a pane's {@code shell_pid} or foreground pids (CB-161). Without
* this walk such a pid matches no pane, and the caller falls through to loopback-trust and is
* resolved as the primary — a worker→primary privilege escalation.
*
* <p>This scans agent panes; a spawn-time {@code pid→terminal} cache is the obvious optimization
* once wired into {@code ClaudeCodeLauncher}.
*
* <p>CB-185 split the fleet across two herdr daemons — lead operations on one, members on the
* other ({@code memberHerdrSocket}). A caller's pane can live on <em>either</em> daemon (a lead's
* MCP connection resolves against the lead daemon; a member's against the member daemon), so this
* must be able to search more than one client. {@link #PaneLocator(HerdrClient, HerdrClient)}
* searches the lead client first, then the member client, and collapses to a single scan when the
* two are the same object (the historical single-daemon deployment). The caller's ancestor set is
* computed once per {@link #terminalForPid} call and reused across every client searched — it
* does not depend on which daemon a pane happens to live on.
* and foreground process PIDs. This scans agent panes; a spawn-time {@code pid→terminal} cache is
* the obvious optimization once wired into {@code ClaudeCodeLauncher}.
*/
public final class PaneLocator {
/**
* Bound on how many ancestor generations {@link #ancestorsOf} walks. This runs on every MCP
* call, so a cycle or a pathologically deep process tree must not hang identity resolution;
* 32 generations is far more than any real worker→helper process tree needs.
*/
private static final int MAX_ANCESTRY_DEPTH = 32;
private final HerdrClient herdr;
private final List<HerdrClient> herdrs;
private final ParentResolver parentResolver;
/** Search only this client — the single-daemon deployment. */
public PaneLocator(HerdrClient herdr) {
this(herdr, ParentResolver.PROCESS_HANDLE);
}
/** Search only this client, resolving ancestry through {@code parentResolver} — for tests. */
public PaneLocator(HerdrClient herdr, ParentResolver parentResolver) {
this.herdrs = List.of(herdr);
this.parentResolver = parentResolver;
}
/**
* Search {@code lead} first, then {@code member} — the two-daemon deployment (CB-185). When
* the caller passes the same client for both (no {@code memberHerdrSocket} configured), this
* collapses to one client and one scan, exactly {@link #PaneLocator(HerdrClient)}'s behaviour.
*/
public PaneLocator(HerdrClient lead, HerdrClient member) {
this(lead, member, ParentResolver.PROCESS_HANDLE);
}
/** Two-daemon deployment, resolving ancestry through {@code parentResolver} — for tests. */
public PaneLocator(HerdrClient lead, HerdrClient member, ParentResolver parentResolver) {
this.herdrs = lead == member ? List.of(lead) : List.of(lead, member);
this.parentResolver = parentResolver;
this.herdr = herdr;
}
/**
* The {@code terminal_id} of the agent pane whose process tree contains {@code pid}, or
* {@code null} if no agent pane on any searched daemon owns it (e.g. the caller is the
* primary, or off-host).
* {@code null} if no agent pane owns it (e.g. the caller is the primary, or off-host).
*/
public String terminalForPid(long pid) {
if (pid <= 0) {
return null;
}
Set<Long> ancestry = ancestorsOf(pid);
for (HerdrClient herdr : herdrs) {
String terminal = terminalForPid(herdr, ancestry);
if (terminal != null) {
return terminal;
}
}
return null;
}
/**
* {@code pid} itself plus its ancestor chain, walked through {@link #parentResolver} up to
* {@link #MAX_ANCESTRY_DEPTH} generations or pid 1, whichever comes first. A vanished ancestor
* ({@link ParentResolver#parentOf} returning empty) ends the walk without error — it just means
* the chain is shorter than the bound. A cycle in a fake resolver is caught by the "already
* seen" check and also ends the walk, so this can never loop.
*/
private Set<Long> ancestorsOf(long pid) {
Set<Long> ancestry = new LinkedHashSet<>();
long current = pid;
for (int depth = 0; depth < MAX_ANCESTRY_DEPTH; depth++) {
if (current <= 0 || !ancestry.add(current)) {
break; // vanished/invalid pid, or a cycle back to a pid already recorded
}
if (current == 1) {
break; // reached the root of the process tree
}
OptionalLong parent = parentResolver.parentOf(current);
if (parent.isEmpty()) {
break; // vanished ancestor — not an error, just the end of the chain
}
current = parent.getAsLong();
}
return ancestry;
}
private static String terminalForPid(HerdrClient herdr, Set<Long> ancestry) {
for (JsonNode pane : herdr.call("pane.list", Map.of()).path("panes")) {
String paneId = pane.path("pane_id").asText(null);
if (paneId != null && paneOwnsAnyOf(herdr, paneId, ancestry)) {
if (paneId != null && paneOwnsPid(paneId, pid)) {
return pane.path("terminal_id").asText(null);
}
}
return null;
}
private static boolean paneOwnsAnyOf(HerdrClient herdr, String paneId, Set<Long> ancestry) {
private boolean paneOwnsPid(String paneId, long pid) {
JsonNode info;
try {
info = herdr.call("pane.process_info", Map.of("pane_id", paneId)).path("process_info");
} catch (HerdrException e) {
return false; // pane vanished mid-scan — just skip it
}
if (ancestry.contains(info.path("shell_pid").asLong(-1))) {
if (info.path("shell_pid").asLong(-1) == pid) {
return true;
}
for (JsonNode p : info.path("foreground_processes")) {
if (ancestry.contains(p.path("pid").asLong(-1))) {
if (p.path("pid").asLong(-1) == pid) {
return true;
}
}
@@ -1,21 +0,0 @@
package dev.ltms.fleet.herdr;
import java.util.OptionalLong;
/**
* Resolves a pid's parent pid — the seam {@link PaneLocator} walks a process's ancestry through,
* so its tests can drive the walk from a fake pid→parent map instead of spawning real processes.
*
* <p>{@link #PROCESS_HANDLE} is the production implementation, backed by {@link ProcessHandle}.
*/
public interface ParentResolver {
/** The parent pid of {@code pid}, or empty if {@code pid} is gone or has no known parent. */
OptionalLong parentOf(long pid);
/** Production resolver: asks the JVM's {@link ProcessHandle} view of the OS process tree. */
ParentResolver PROCESS_HANDLE = pid -> ProcessHandle.of(pid)
.flatMap(ProcessHandle::parent)
.map(parent -> OptionalLong.of(parent.pid()))
.orElse(OptionalLong.empty());
}
@@ -6,14 +6,12 @@ import dev.ltms.fleet.msg.TurnToken;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.time.Duration;
import java.util.List;
import java.util.Objects;
import java.util.Set;
import java.util.TreeSet;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.LongSupplier;
import java.util.regex.Pattern;
/**
@@ -63,25 +61,6 @@ public final class CompletionResolver implements TurnListener {
/** Cap the scraped tail so a long transcript can't return an unbounded blob. */
static final int MAX_SCRAPE_CHARS = 4000;
/**
* fleetd#164: the floor below which a {@code BUSY -> DONE} transition cannot be real work. A
* backend that rejects a turn outright (e.g. an HTTP 400 from the model, before the worker read
* a single file or produced a token) drives the exact same confirmed {@code working -> idle}
* transition a genuine completion does — just in about a second instead of the many seconds a
* real turn costs. {@link #onTurnComplete} cannot tell those two cases apart from the transition
* alone, so a turn that settles inside this floor is treated as a crash signature and resolved
* as a failure, never as a (possibly empty) success.
*/
public static final long MIN_TURN_NANOS = Duration.ofSeconds(2).toNanos();
/**
* fleetd#164 (part 2): one stable, explicit backend-failure marker seen on a Claude Code pane
* when the backend itself rejected the turn (e.g. {@code "API Error: 400 invalid request body"}).
* Kept deliberately narrow — a growing list of ad-hoc error strings rots as backends change their
* wording; broader backend-error surfacing is out of scope here (fleetd#164 point 3).
*/
private static final Pattern BACKEND_ERROR = Pattern.compile("(?i)\\bAPI Error\\s*:");
private static final String CLIPPED_PANE_TAIL_MARKER =
"[Pane tail clipped: member did not call fleet_reply.]";
@@ -89,7 +68,6 @@ public final class CompletionResolver implements TurnListener {
private final Rendezvous rendezvous;
private final ExhaustedPatternLookup exhaustedPatterns;
private final ExhaustionSink exhaustionSink;
private final LongSupplier nowNanos;
/**
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
@@ -101,21 +79,8 @@ public final class CompletionResolver implements TurnListener {
* reference: a completion scrape equal to it means the worker produced no new output (the previous
* turn's wind-down sampled as this boundary), so it is suppressed. Overwritten on each delivery;
* cleared when the turn resolves. Package-private so tests can capture and replay a specific turn.
*
* <p>{@code deliveredAtNanos} (fleetd#164) is the {@link #nowNanos} reading taken at delivery —
* the other half of the {@link #MIN_TURN_NANOS} floor check, compared against a fresh reading at
* resolution time.
*/
record InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline, long deliveredAtNanos) {
/**
* Convenience for tests exercising scrape/suppression logic that don't care about turn
* timing: back-dates the delivery far enough that {@link #MIN_TURN_NANOS} can never fire.
* Not used by production code — {@link #captureBaseline} always records a real reading.
*/
InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline) {
this(waiter, baseline, Long.MIN_VALUE / 2);
}
record InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline) {
}
private final ConcurrentHashMap<String, InFlight> inFlight = new ConcurrentHashMap<>();
@@ -132,25 +97,10 @@ public final class CompletionResolver implements TurnListener {
*/
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
ExhaustionSink exhaustionSink) {
this(agents, rendezvous, exhaustedPatterns, exhaustionSink, System::nanoTime);
}
/**
* Test constructor with an injectable clock (fleetd#164), matching the {@code LongSupplier}
* pattern {@link dev.ltms.fleet.session.SessionManager} and {@link dev.ltms.fleet.msg.MessageService}
* already use: lets a test place a turn's delivery and its resolution at an exact, controllable
* distance apart around the {@link #MIN_TURN_NANOS} floor, without a real sleep. Public (rather
* than package-private like those two) because callers that wire a full {@code MessageService}
* fixture — e.g. {@code MessageServiceTest} — construct this resolver directly from another
* package.
*/
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
ExhaustionSink exhaustionSink, LongSupplier nowNanos) {
this.agents = agents;
this.rendezvous = rendezvous;
this.exhaustedPatterns = Objects.requireNonNull(exhaustedPatterns, "exhaustedPatterns");
this.exhaustionSink = Objects.requireNonNull(exhaustionSink, "exhaustionSink");
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
}
@Override
@@ -180,7 +130,7 @@ public final class CompletionResolver implements TurnListener {
baseline = null; // fail open: no baseline ⇒ no suppression
log.debug("delivery baseline for {} failed: {}", target, e.getMessage());
}
inFlight.put(target, new InFlight(waiter, baseline, nowNanos.getAsLong()));
inFlight.put(target, new InFlight(waiter, baseline));
}
/** The turn currently baselined for {@code target}, or {@code null} — a test hook for the captureBaseline path. */
@@ -226,61 +176,31 @@ public final class CompletionResolver implements TurnListener {
inFlight.remove(target, turn);
return;
}
// fleetd#164: a BUSY -> DONE transition inside the floor cannot be real work — it's a crash
// signature (e.g. a backend HTTP 400 before the worker did anything), not a fast answer. Fail
// it before spending a scrape on the ordinary path; the reason still carries whatever is on
// screen, since that is usually the backend's own error.
long elapsedNanos = nowNanos.getAsLong() - turn.deliveredAtNanos();
if (elapsedNanos < MIN_TURN_NANOS) {
fail(target, turn, tooFastReason(target, elapsedNanos));
return;
}
String tail;
String assistantBlock = null;
String rawScrape = null;
int originalLength = 0;
boolean clipped = false;
boolean scrapeFailed = false;
try {
rawScrape = agents.read(target, SCRAPE_SOURCE);
assistantBlock = lastAssistantBlock(rawScrape);
assistantBlock = lastAssistantBlock(agents.read(target, SCRAPE_SOURCE));
originalLength = assistantBlock.strip().length();
clipped = originalLength > MAX_SCRAPE_CHARS;
tail = clip(assistantBlock);
} catch (RuntimeException e) {
log.warn("completion scrape for {} failed: {}", target, e.getMessage());
// The worker finished but we couldn't read its screen — still resolve the send so the
// caller unblocks; an empty tail beats hanging until the caller's timeout.
log.warn("completion scrape for {} failed; resolving with an empty tail: {}",
target, e.getMessage());
tail = "";
scrapeFailed = true;
}
// fleetd#164: a scrape nobody could read, and a scrape that read cleanly but produced nothing,
// both used to resolve the send as a SUCCESS carrying "" — indistinguishable from a worker that
// genuinely finished with nothing to say. That is the defect: fail loudly instead, naming the
// member, so a caller (including a lead deciding whether to delegate again) can tell a lost
// turn from a real empty answer.
if (scrapeFailed || tail.isEmpty()) {
// fleetd#211: lastAssistantBlock() found nothing usable — most often a pane with no ⏺
// marker at all, whose boundary scan then starts at the top of the raw screen and breaks
// immediately on the first line of TUI chrome (╭, │, ❯, …). Before giving up as a lost
// turn, run the same exhaustion/backend-error classification against the RAW scrape as a
// fallback, ONLY here. A pane that already yielded a usable block never reaches this
// branch, so the narrow (trimmed) match on the normal path below is completely unchanged
// — zero new false positives there. Every pane this fallback examines was already headed
// for the empty-scrape failure, so a wrong label here is strictly less bad than silently
// losing an exhaustion signal: the alternative outcome is already a failure, just one that
// never quarantines the credential. lastAssistantBlock stays the source of the reply
// TEXT everywhere else; only classification ever consults the raw scrape, and only here.
if (rawScrape != null && classifyRawScrapeFallback(target, turn, waiter, rawScrape)) {
return;
}
fail(target, turn, emptyScrapeReason(target, scrapeFailed));
return;
}
// Misattribution guard (CB-115): if the scrape is byte-identical to the pane content at
// delivery, this turn produced no new output — the boundary belongs to the previous turn's
// wind-down (common on rapid back-to-back sends). Suppress rather than resolve the send with
// a stale answer; the real fleet_reply (or a later genuine completion) resolves it instead.
// A scrape that failed to read is exempt — an empty tail there is "couldn't see", not "no change".
String baseline = turn.baseline();
if (baseline != null && baseline.equals(tail)) {
if (!scrapeFailed && baseline != null && baseline.equals(tail)) {
log.debug("suppressing misattributed completion for {} (no output change since delivery)",
target);
return; // keep the in-flight record: a later genuine completion still needs it
@@ -288,33 +208,21 @@ public final class CompletionResolver implements TurnListener {
// CB-578 stage A: a turn that ended with no fleet_reply AND whose scrape matches the
// backend's configured usage-limit pattern is a refusal, not an answer. Classify it as
// BACKEND_EXHAUSTED rather than handing the caller a scrape that reads like a real reply.
Pattern exhausted = exhaustedPatterns.patternFor(target);
String matchedLine = exhausted == null ? null : firstMatchingLine(assistantBlock, exhausted);
if (matchedLine != null) {
String reason = "backend exhausted (usage limit): " + matchedLine;
if (rendezvous.resolveExhausted(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("completion for {} classified BACKEND_EXHAUSTED (no fleet_reply; scrape "
+ "matched the profile's exhausted pattern): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
if (!scrapeFailed) {
Pattern exhausted = exhaustedPatterns.patternFor(target);
String matchedLine = exhausted == null ? null : firstMatchingLine(assistantBlock, exhausted);
if (matchedLine != null) {
String reason = "backend exhausted (usage limit): " + matchedLine;
if (rendezvous.resolveExhausted(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("completion for {} classified BACKEND_EXHAUSTED (no fleet_reply; scrape "
+ "matched the profile's exhausted pattern): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
}
return;
}
return;
}
// fleetd#164 (part 2): a scrape that read cleanly and produced content still isn't a real
// reply when that content is the backend's own rejection (e.g. an HTTP 400 before the worker
// did any work). Classify it as a failure naming the member, rather than handing the caller a
// scrape that reads like a completed answer.
String backendError = firstMatchingLine(assistantBlock, BACKEND_ERROR);
if (backendError != null) {
// Carry the whole scrape, not just the matched line. The pattern is a heuristic: a member
// that forgot fleet_reply while reporting *about* a backend error matches it too. Failing
// is still right — the caller must not read a scrape as an answer — but dropping the rest
// of the pane would destroy the report, which is the same defect fleetd#164 is about.
fail(target, turn, "member " + target + " ended on a backend error: " + backendError
+ "\n--- pane tail ---\n" + tail);
return;
}
String completion = clipped ? tail + "\n" + CLIPPED_PANE_TAIL_MARKER : tail;
if (rendezvous.resolveCompletion(waiter, completion)) {
@@ -329,43 +237,6 @@ public final class CompletionResolver implements TurnListener {
}
}
/**
* fleetd#211: the raw-scrape fallback classification, run only when {@link #lastAssistantBlock}
* found nothing usable (see the call site in {@link #resolve}). Mirrors the two classifications
* the normal path already applies to the trimmed assistant block — exhaustion first, then the
* narrow {@link #BACKEND_ERROR} pattern — against {@code raw} instead, and reports whether one of
* them handled the turn (resolved the waiter or failed it) so the caller skips the empty-scrape
* failure. Never runs on the normal (non-empty-block) path, and never touches the reply text.
*/
private boolean classifyRawScrapeFallback(String target, InFlight turn,
CompletableFuture<Rendezvous.Resolution> waiter, String raw) {
Pattern exhausted = exhaustedPatterns.patternFor(target);
String matchedLine = exhausted == null ? null : firstMatchingLine(raw, exhausted);
if (matchedLine != null) {
String reason = "backend exhausted (usage limit): " + matchedLine;
if (rendezvous.resolveExhausted(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("completion for {} classified BACKEND_EXHAUSTED from the raw scrape (no "
+ "usable assistant block; no fleet_reply): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
}
return true;
}
String backendError = firstMatchingLine(raw, BACKEND_ERROR);
if (backendError != null) {
// Carry the pane, not just the matched line — the same fleetd#164 rule the normal path
// above applies. Here it matters more, not less: the trimmed block was empty, so the raw
// scrape is the ONLY copy of whatever the member managed to say. Clipped to the same cap
// the normal path uses, since a raw screen has no boundary trimming to bound it.
fail(target, turn, "member " + target + " ended on a backend error: " + backendError
+ "\n--- pane tail ---\n" + clip(raw));
return true;
}
return false;
}
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
void fail(String target, InFlight turn) {
fail(target, turn, null);
@@ -401,36 +272,6 @@ public final class CompletionResolver implements TurnListener {
}
}
/**
* fleetd#164: the failure reason for a turn that settled inside {@link #MIN_TURN_NANOS} — names
* the member and both timings, and appends whatever the pane shows (usually the backend's own
* error) so the caller sees the cause, not just "it failed".
*/
private String tooFastReason(String target, long elapsedNanos) {
String scrape;
try {
scrape = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
scrape = "";
}
String reason = String.format(
"member %s went BUSY -> DONE in %dms (floor %dms) — too fast to be real work, most "
+ "likely a backend error before any work started",
target, elapsedNanos / 1_000_000, MIN_TURN_NANOS / 1_000_000);
return scrape.isBlank() ? reason : reason + ": " + scrape;
}
/**
* fleetd#164: the failure reason for a scrape that produced zero characters — names the member
* and says plainly that the turn produced nothing, so a caller (a lead deciding whether to
* delegate again included) never mistakes a lost turn for a genuinely empty reply.
*/
private static String emptyScrapeReason(String target, boolean scrapeFailed) {
return "member " + target + " turn completed with an empty scrape (0 chars) — "
+ (scrapeFailed ? "its pane could not be read; " : "")
+ "treating as a lost turn, not a real answer";
}
/**
* The first line of {@code text} matching {@code pattern}, stripped — the CB-578 stage A
* evidence carried in a {@code BACKEND_EXHAUSTED} reason so the operator sees the real refusal
@@ -2,7 +2,6 @@ package dev.ltms.fleet.inject;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.msg.TurnToken;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -90,7 +89,6 @@ public final class Injector {
public static final long POLL_INTERVAL_MILLIS = 250;
private final AgentControl agents;
private final HerdrRouter router;
private final TurnListener turnListener;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
private final Consumer<String> forget; // CB-114: clear a gone worker's readiness/presence
@@ -126,25 +124,11 @@ public final class Injector {
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
this.agents = agents;
this.router = null;
this.turnListener = turnListener;
this.ready = ready;
this.forget = forget;
}
public Injector(HerdrRouter router, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
this.agents = null;
this.router = router;
this.turnListener = turnListener;
this.ready = ready;
this.forget = forget;
}
private AgentControl agentsFor(String target) {
return router != null ? router.agentsFor(target) : agents;
}
/** A pending message and the future that completes when it has been delivered. */
private record Pending(String text, TurnToken token, CompletableFuture<Void> delivered) {
}
@@ -269,7 +253,7 @@ public final class Injector {
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
try {
agentsFor(target).send(target, p.text());
agents.send(target, p.text());
t.queue.poll();
t.awaitingPickup = true;
t.awaitingCompletion = true;
@@ -329,7 +313,7 @@ public final class Injector {
// thread while it holds the target lock.
if (resubmit) {
try {
agentsFor(target).submit(target); // nudge a raced Enter so the pending paste submits
agents.submit(target); // nudge a raced Enter so the pending paste submits
} catch (RuntimeException e) {
log.debug("resubmit to {} failed (will retry next poll): {}", target, e.getMessage());
}
@@ -3,7 +3,6 @@ package dev.ltms.fleet.inject;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.HerdrRouter;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -23,7 +22,6 @@ public final class StatusPoller {
private static final Logger log = LoggerFactory.getLogger(StatusPoller.class);
private final AgentControl agents;
private final HerdrRouter router;
private final Injector injector;
private final StatusRefiner refiner;
private final long intervalMillis;
@@ -37,24 +35,11 @@ public final class StatusPoller {
public StatusPoller(AgentControl agents, Injector injector, StatusRefiner refiner,
long intervalMillis) {
this.agents = agents;
this.router = null;
this.injector = injector;
this.refiner = refiner;
this.intervalMillis = intervalMillis;
}
public StatusPoller(HerdrRouter router, Injector injector, long intervalMillis) {
this.agents = null;
this.router = router;
this.injector = injector;
// CB-185: this refiner's own AgentControl (member) is only a default for the legacy 2-arg
// refine() overload — the loop below always calls the 3-arg refine(target, raw, control)
// with the per-target control from router.agentsFor(target), so a lead target is refined
// against the LEAD daemon even though this field points at the member one.
this.refiner = new StatusRefiner(router.memberAgents());
this.intervalMillis = intervalMillis;
}
/** Start the polling loop on a virtual thread. Idempotent. */
public synchronized void start() {
if (running) return;
@@ -71,11 +56,7 @@ public final class StatusPoller {
try {
// herdr's agent_status can misreport a settled worker as `unknown`; refine it
// against the pane content before it drives delivery/completion (CB-115).
// CB-185: refine THROUGH the same control the raw status came from — a router
// splits lead/member targets across two herdr daemons, and reading a lead's pane
// through the (fixed) member refiner never finds it, wedging that lead at UNKNOWN.
AgentControl control = router != null ? router.agentsFor(target) : agents;
AgentStatus status = refiner.refine(target, control.status(target), control);
AgentStatus status = refiner.refine(target, agents.status(target));
injector.onStatus(target, status);
} catch (HerdrException e) {
// The worker's agent is gone — stop trying and unblock its waiters.
@@ -41,32 +41,15 @@ public final class StatusRefiner {
}
/**
* Return a trustworthy status for {@code target}, reading its pane through this refiner's own
* {@link AgentControl}. Equivalent to {@link #refine(String, AgentStatus, AgentControl)} with
* that control — kept for callers that only ever talk to one herdr daemon.
* Return a trustworthy status for {@code target}. Any non-{@code UNKNOWN} {@code raw} is returned
* unchanged; an {@code UNKNOWN} triggers a pane read and content classification. A read failure
* leaves it {@code UNKNOWN} (the safe default: no delivery, and the stall path still applies).
*/
public AgentStatus refine(String target, AgentStatus raw) {
return refine(target, raw, agents);
}
/**
* Return a trustworthy status for {@code target}. Any non-{@code UNKNOWN} {@code raw} is returned
* unchanged; an {@code UNKNOWN} triggers a pane read (through {@code control}) and content
* classification. A read failure leaves it {@code UNKNOWN} (the safe default: no delivery, and
* the stall path still applies).
*
* <p>CB-185: {@code control} must be the {@link AgentControl} for the <em>same</em> daemon the
* raw status was sampled from — a router splits lead and member targets across two herdr
* daemons, and reading a lead's pane through the member client (or vice versa) fails to find
* the pane and leaves the target wedged at {@code UNKNOWN} forever. Callers that route per
* target (e.g. {@code StatusPoller}) must pass that target's control explicitly rather than
* relying on the control fixed at construction.
*/
public AgentStatus refine(String target, AgentStatus raw, AgentControl control) {
if (raw != AgentStatus.UNKNOWN) return raw;
String pane;
try {
pane = control.read(target, PROBE_SOURCE);
pane = agents.read(target, PROBE_SOURCE);
} catch (RuntimeException e) {
log.debug("status refine read for {} failed; leaving UNKNOWN: {}", target, e.getMessage());
return AgentStatus.UNKNOWN;
@@ -951,10 +951,7 @@ public final class FleetMcp {
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm))
.toList();
// fleetd #209: this is the caller-driven fleet_list read that actually reports
// agentSessionId (via memberCapacityView -> SessionManager.rosterView), so it uses the
// resolving roster; the heartbeat/health/metrics timers stay on the plain sessions.roster().
List<MemberSession> roster = sessions.rosterResolved();
List<MemberSession> roster = sessions.roster();
List<Map<String, Object>> out = roster.stream()
.map(s -> memberCapacityView(s, live.get(s.terminalId()), messages, capacity.clock().getAsLong()))
.toList();
@@ -94,26 +94,13 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials) {
this(agents, spaces, guard, profiles, defaultProfile, env, spawnReadyTimeoutMs, spawnReadyPollMs,
fleet, memberCredentials, null, null);
}
/** Production constructor, plus the live config for URI environment exclusions. */
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<Set<String>> hostEnvNames,
Supplier<FleetConfig> config) {
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
fleet, memberCredentials, hostEnvNames, config);
fleet, memberCredentials);
}
/**
@@ -166,22 +153,10 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials) {
this(agents, spaces, guard, profiles, defaultProfile, env, spawnReadyTimeoutMs, nowMillis, sleeper,
fleet, memberCredentials, null, null);
}
/** Full testability constructor, plus the live config for URI environment exclusions. */
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Function<String, String> env, long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper, Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<Set<String>> hostEnvNames,
Supplier<FleetConfig> config) {
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials, hostEnvNames, config);
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials);
this.guard = guard;
}
@@ -199,7 +174,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<Set<String>> hostEnvNames) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials, hostEnvNames, null);
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials, hostEnvNames);
this.guard = guard;
}
@@ -278,22 +253,18 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/**
* Add the Claude-specific session-identity flags to {@code argv} and return the peer's OWN
* session id — the resume handle. A resume passes the prior id via {@code -r} and returns
* that id; every other spawn mints a new UUID, passes it via {@code --session-id}, and
* returns the mint. The bridge's logical name rides along as {@code -n} when present.
*
* <p>fleetd #214: the mint is unconditional. A plain {@code fleet_spawn} passes neither
* sessionName nor resumeSessionId, yet the member must still be resumable, and this id is
* the only resume handle a claude-code member has — unlike opencode, nothing resolves it
* after the launch. Checked against the real binary (claude 2.1.252): the flag is safe on
* every spawn. The binary takes only a valid UUID — it refuses any other value at argument
* parsing ("Invalid session ID. Must be a valid UUID.") — so the {@code UUID.randomUUID()}
* mint is required, not incidental. The only flag interaction the binary documents is with
* {@code -r} (both claim the session id), and the resume branch above never combines the two.
* session id — the resume handle. A resume request passes the prior id via {@code -r} and
* returns that id; a fresh named session mints a new UUID, passes it via {@code --session-id},
* and returns the mint. The bridge's logical name rides along as {@code -n} when present. When
* <em>no</em> identity is requested (sessionName and resumeSessionId both blank) this adds
* nothing and returns {@code null}, keeping the legacy no-identity launch byte-identical.
*/
private static String applySessionIdentity(List<String> argv, String sessionName, String resumeSessionId) {
boolean resuming = resumeSessionId != null && !resumeSessionId.isBlank();
boolean named = sessionName != null && !sessionName.isBlank();
if (!resuming && !named) {
return null; // no identity requested — keep the legacy launch byte-identical
}
if (named) {
argv.add("-n");
argv.add(sessionName);
@@ -324,19 +295,11 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
*
* <p>CB-618: Claude Code refuses to start when BOTH {@code --append-system-prompt} and
* {@code --append-system-prompt-file} are on the command line ("Cannot use both ... Please use
* only one"), so the two charters can never travel on separate flags. They are concatenated
* into the one file, role charter first and reply charter last — last is where the reply rule
* must sit, because it is the rule that must survive.
*
* <p>fleetd #220: a lone reply charter used to ride inline on {@code --append-system-prompt},
* which put ~800 bytes of prose on the command line herdr types into the pane. That line is
* capped at {@value HerdrPeerLauncher#PANE_COMMAND_BYTE_LIMIT} bytes by the pty itself, and
* everything past the cap is dropped with no error from any layer. The charter alone left about
* 50 bytes of headroom, so adding one flag ({@code --session-id}, fleetd #214) truncated the
* LAST argument instead — {@code --autocompact 250000} arrived as {@code --autocompact 25},
* claude rejected it, and every claude-code spawn died as an unexplained readiness timeout.
* The charter now always travels as a file, which takes the prose off the command line for
* good; {@link HerdrPeerLauncher#checkPaneCommandFits} is the backstop for whatever grows next.
* only one"), so the two charters can never travel on separate flags. When both are present they
* are concatenated into the one file, role charter first and reply charter last — last is where
* the reply rule must sit, because it is the rule that must survive. When only the reply charter
* is present it keeps its proven inline {@code --append-system-prompt} delivery, which is also
* the only form that reaches a member with no repo checkout.
*/
private List<String> argvWithFleet(FleetConfig.Profile cfg, LaunchSpec spec) {
String roleCharter = nonBlank(spec.roleCharter());
@@ -353,13 +316,16 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
argv.add("--mcp-config");
argv.add(mcpConfigJson(cfg));
}
// Combine the charters in order role -> reply, dropping any that are absent. They ride one
// --append-system-prompt-file (CB-618 forbids the inline flag and the file flag together),
// always — fleetd #220: charter prose on the command line overruns the pane's byte cap.
// Combine the charters in order role -> reply, dropping any that are absent. When two
// or more survive they must ride one --append-system-prompt-file (CB-618 forbids the inline
// flag and the file flag together). A lone reply charter keeps its proven inline delivery.
List<String> charters = new java.util.ArrayList<>(2);
if (roleCharter != null) charters.add(roleCharter);
if (replyCharter != null) charters.add(replyCharter);
if (!charters.isEmpty()) {
if (charters.size() == 1 && replyCharter != null && roleCharter == null) {
argv.add("--append-system-prompt");
argv.add(replyCharter);
} else if (!charters.isEmpty()) {
argv.add("--append-system-prompt-file");
argv.add(writeCharterFile(String.join("\n\n", charters)).toString());
}
@@ -2,8 +2,6 @@ package dev.ltms.fleet.member;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerHandle;
@@ -23,7 +21,6 @@ import java.util.ArrayList;
import java.util.Collections;
import java.util.EnumSet;
import java.util.HashSet;
import java.util.IdentityHashMap;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
@@ -47,17 +44,12 @@ import java.util.stream.Collectors;
* the single adapter that declares it. Profiles partition cleanly across adapters: the
* constructor rejects a name claimed by two.</li>
* <li><strong>By pane id</strong> — {@link #stop} routes to the adapter that spawned that pane
* (recorded at spawn time). A pane the composite never spawned, or one whose record was lost
* to a daemon restart (CB-185 blocker 1 — {@link #spawnedBy} is in-memory only), can use the
* fallback route in a one-daemon fleet. With more than one herdr daemon, {@link #probeOwner}
* asks each configured daemon which one actually knows the pane: exactly one match routes
* (and caches); no match is treated as already-gone; more than one match is a genuine
* ambiguity (pane ids are per-daemon counters, so two daemons really can both hold, say,
* {@code w1:p1}) and stop refuses rather than closing a pane on an arbitrary herdr daemon.</li>
* (recorded at spawn time). A pane the composite never spawned (only real for a caller that
* hand-rolls an id) falls back to the first delegate; teardown is pane-id addressed and
* tab cleanup is single-occupant guarded, so it is safe either way.</li>
* <li><strong>Fleet-wide</strong> — {@link #reapOrphanWorkers} and {@link #capabilities} fan out
* and combine. {@link #list} is deduplicated by (owning daemon, pane id): delegates that share
* one herdr connection report the same global agent set, but two daemons can each hold a pane
* called {@code w1:p1}, so the daemon has to be part of the key.</li>
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
* shares one herdr connection and so reports the same global agent set.</li>
* </ul>
*
* <p>CB-518: an unqualified spawn is routed through a {@link PlacementPolicy}. The default
@@ -443,98 +435,12 @@ public final class CompositePeerLauncher implements PeerLauncher {
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.get(id);
HerdrPeerLauncher d = spawnedBy.remove(id);
if (d == null) {
if (herdrDaemonCount() == 1) {
log.debug("stop({}) — no recorded owner in a single-daemon fleet", id);
d = delegates.getFirst();
} else {
d = probeOwner(id);
if (d == null) {
// No configured herdr daemon has ever heard of this pane. CB-185 blocker 1: this
// is the normal case right after a daemon restart empties spawnedBy for a member
// that has ALREADY been torn down since — the caller retried a stop that already
// succeeded. Nothing to close and no owner to cache; matching the tolerance
// HerdrPeerLauncher#stop already gives an already-gone pane (agent.close swallows
// that as success), stop() here is a no-op rather than a refusal.
log.debug("stop({}) — no configured herdr daemon knows this pane; "
+ "treating as already stopped", id);
return;
}
}
log.debug("stop({}) — no recorded owner, routing to the first adapter (pane-addressed)", id);
d = delegates.getFirst();
}
// Drop the owner record only after the delegate accepted the stop. Removing it first meant a
// delegate that threw left the pane alive with its owner forgotten, so the retry fell into
// the ambiguous branch above and refused the id for good.
d.stop(id);
spawnedBy.remove(id);
}
/**
* CB-185 blocker 1: recover a spawnedBy cache miss by asking every distinct herdr daemon which
* one actually knows {@code id} — the fix for "after a restart, every surviving member becomes
* un-stoppable" (spawnedBy is in-memory only, so a restart empties it, and members intentionally
* outlive the daemon).
*
* <p>Grouped by daemon identity, not by delegate, for the same reason {@link #list()} groups
* that way: two adapters (claude-code, opencode) sharing one herdr connection would otherwise be
* probed twice, and a pane on their shared daemon would look owned by two adapters instead of
* one daemon.
*
* <p>A daemon that fails to answer {@code list()} (e.g. it is down) is treated as "does not know
* this pane" rather than aborting the whole probe — one unreachable daemon must never make a
* pane that a <em>different</em>, healthy daemon actually owns un-stoppable too, which would
* resurrect the exact bug this method exists to fix.
*
* @return the owning delegate — cached into {@link #spawnedBy} so the next call is free — or
* {@code null} when no daemon knows the pane
* @throws IllegalArgumentException when more than one daemon claims the pane: pane ids are
* per-daemon counters, so two daemons really can both hold, say, {@code w1:p1}, and there
* is no way to tell which one the caller means
*/
private HerdrPeerLauncher probeOwner(String id) {
Map<HerdrClient, HerdrPeerLauncher> byDaemon = new IdentityHashMap<>();
for (HerdrPeerLauncher delegate : delegates) {
byDaemon.putIfAbsent(delegate.herdr(), delegate);
}
List<HerdrPeerLauncher> owners = new ArrayList<>();
for (HerdrPeerLauncher representative : byDaemon.values()) {
List<Agent> agents;
try {
agents = representative.list();
} catch (HerdrException e) {
log.warn("stop({}) probe: a configured herdr daemon was unreachable ({}); "
+ "treating it as not knowing this pane", id, e.getClass().getSimpleName());
continue;
}
boolean knows = agents.stream().anyMatch(a -> id.equals(a.paneId()));
if (knows) {
owners.add(representative);
}
}
if (owners.size() > 1) {
throw new IllegalArgumentException("ambiguous paneId '" + id + "': "
+ owners.size() + " configured herdr daemons report this pane — "
+ "no way to tell which one the caller means");
}
if (owners.isEmpty()) {
return null;
}
HerdrPeerLauncher owner = owners.get(0);
spawnedBy.put(id, owner);
return owner;
}
/**
* Count actual herdr daemons, not peer adapter kinds. Identity is intentional: separate client
* objects may represent different daemons even if a client later implements value equality.
*/
private int herdrDaemonCount() {
Set<HerdrClient> daemons = Collections.newSetFromMap(new IdentityHashMap<>());
for (HerdrPeerLauncher delegate : delegates) {
daemons.add(delegate.herdr());
}
return daemons.size();
}
@Override
@@ -569,24 +475,14 @@ public final class CompositePeerLauncher implements PeerLauncher {
return route(profileName).capabilities();
}
/**
* Every herdr agent, deduplicated by (owning daemon, pane id).
*
* <p>Delegates that share one {@link HerdrClient} see the same global agent set, so listing them
* both would report every agent twice — that is what the dedupe is for. But pane ids are
* per-daemon counters, so two daemons really can both hold {@code w1:p1} on different panes.
* Keying on the pane id alone would silently drop one of them from {@code fleet_list} and from
* every status view built on it. The daemon is part of the key for exactly that reason.
*/
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
Map<HerdrClient, Integer> daemonIndex = new IdentityHashMap<>();
Map<String, Agent> byPane = new LinkedHashMap<>();
for (HerdrPeerLauncher d : delegates) {
int daemon = daemonIndex.computeIfAbsent(d.herdr(), _ -> daemonIndex.size());
for (Agent a : d.list()) {
if (a.paneId() != null) {
byPane.putIfAbsent(daemon + "\u0000" + a.paneId(), a);
byPane.putIfAbsent(a.paneId(), a);
}
}
}
@@ -7,9 +7,6 @@ import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.GroupPrincipal;
import java.nio.file.attribute.PosixFileAttributeView;
import java.nio.file.attribute.PosixFilePermissions;
import java.time.Duration;
import java.time.Instant;
import java.util.ArrayList;
@@ -119,77 +116,6 @@ public final class EnvAllowListScrub {
}
}
/**
* fleetd #213: as {@link #generate(Path, Set)}, plus share the generated directory with
* {@code group} — the member OS user's group (the operator's existing {@code worktreeGroup:}
* name, reused rather than inventing a second one) — so a member running under a different OS
* user than fleetd's own process can still read what it needs from a directory placed outside
* {@code java.io.tmpdir}. {@code group} null/blank ⇒ identical to {@link #generate(Path, Set)};
* this is the single-daemon (no {@code memberHerdrSocket}) shape, where the pane is fleetd's own
* uid and no group sharing is needed.
*
* @throws UncheckedIOException also when {@code group} does not resolve on this host, or a
* group-ownership/permission call is refused — the same "fail
* loudly rather than start unprotected" contract as above: a scrub
* the configured member user cannot even read is not a working
* control.
*/
public static Path generate(Path parentDir, Set<String> allowedNames, String group) {
Path dir = generate(parentDir, allowedNames);
if (group != null && !group.isBlank()) {
shareWithGroup(dir, group);
}
return dir;
}
/**
* chgrp/chmod-equivalent over a freshly generated directory and the flat files already written
* into it: owner keeps full access, {@code group} gets traverse+read on the directory ({@code
* rwxr-x---}, so a member process — a login shell reading it via {@code ZDOTDIR}, or another
* process simply opening a file under it — running under that group can find and read the
* files) and read-only on each file ({@code rw-r-----}) — deliberately no group WRITE anywhere,
* since a member never needs to add or change what fleetd generated. (For the ZDOTDIR scrub
* specifically, this also means the scrub script's own report write inside the pane fails
* closed rather than open — see {@code scrub.zsh}'s trailing {@code 2>/dev/null} — which {@link
* dev.ltms.fleet.member.HerdrPeerLauncher#releaseZdotdir} already treats as "cannot be
* confirmed to have run" rather than success.)
*
* <p>Package-private and named generically on purpose: fleetd #213 built this for the ZDOTDIR
* scrub directory, and fleetd #219 reuses it verbatim for {@link
* dev.ltms.fleet.member.OpenCodeLauncher}'s ephemeral {@code opencode.json} directory — both are
* "a fleetd-generated directory of flat files that a different-uid member process must read but
* never write," so the sharing mechanism is shared rather than copied a second time.
*/
static void shareWithGroup(Path dir, String group) {
try {
GroupPrincipal principal = dir.getFileSystem().getUserPrincipalLookupService()
.lookupPrincipalByGroupName(group);
setGroupAndPermissions(dir, principal, "rwxr-x---");
try (Stream<Path> entries = Files.list(dir)) {
for (Path file : entries.toList()) {
setGroupAndPermissions(file, principal, "rw-r-----");
}
}
} catch (IOException e) {
throw new UncheckedIOException("cannot share generated directory " + dir + " with group '"
+ group + "' — the group must exist, and the fleetd operator ("
+ System.getProperty("user.name") + ") must be a member of it", e);
} catch (UnsupportedOperationException e) {
throw new UncheckedIOException("cannot share generated directory " + dir + " with group '"
+ group + "' — this filesystem does not support POSIX group ownership",
new IOException(e));
}
}
private static void setGroupAndPermissions(Path path, GroupPrincipal group, String perms) throws IOException {
PosixFileAttributeView view = Files.getFileAttributeView(path, PosixFileAttributeView.class);
if (view == null) {
throw new IOException("POSIX file attributes are not supported for " + path);
}
view.setGroup(group);
Files.setPosixFilePermissions(path, PosixFilePermissions.fromString(perms));
}
/** One operator-sourcing startup file: source the {@code $HOME} counterpart, change nothing else. */
private static String homeSourcingFile(String name) {
return """
@@ -3,7 +3,6 @@ package dev.ltms.fleet.member;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
@@ -112,12 +111,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* daemon's own process is started the same way (a login shell sourcing the same secret store —
* see CB-592's investigation of {@code secrets.sh}), so on a single-host deployment its env
* mirrors what the pane's login shell is about to export.
*
* <p>fleetd #185 stage 2: that mirroring assumption holds only while the member pane runs under
* the SAME OS user as the daemon. When {@code memberHerdrSocket:} is configured, member panes
* run on a second herdr owned by a different user — different {@code $HOME}, different {@code
* secrets.sh}, different environment entirely — so this field's data no longer describes what a
* member pane inherits. See {@link #logCredentialGap} for how that mode is handled.
*/
private final Supplier<Set<String>> hostEnvNames;
@@ -176,8 +169,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
/** Guards {@link #warnNonZsh} to one WARN per launcher instance, not one per spawn. */
private final AtomicBoolean nonZshShellWarned = new AtomicBoolean();
/** Live config provides URI environment names that must never enter member panes. */
private final Supplier<FleetConfig> config;
/**
* @param namePrefix label prefix for this peer kind (drives naming and reap)
@@ -246,19 +237,8 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<Set<String>> hostEnvNames) {
this(namePrefix, agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs, nowMillis,
sleeper, fleet, memberCredentials, hostEnvNames, null);
}
/** As above, plus the live full config for secret-bearing URI environment names. */
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Function<String, String> env, long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper, Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<Set<String>> hostEnvNames, Supplier<FleetConfig> config) {
Supplier<Set<String>> hostEnvNames) {
this.fleet = fleet;
this.namePrefix = namePrefix;
this.agents = agents;
@@ -271,7 +251,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
this.sleeper = sleeper;
this.memberCredentials = memberCredentials;
this.hostEnvNames = hostEnvNames != null ? hostEnvNames : () -> System.getenv().keySet();
this.config = config;
}
// --- adapter seams -------------------------------------------------------------------------
@@ -541,11 +520,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
req.sessionName(), spawned.agentSessionId(), spawned.receipt());
}
/** The herdr daemon that owns this launcher's pane coordinates. */
public HerdrClient herdr() {
return agents.herdr();
}
@Override
public String effectiveCwd(SpawnRequest req) {
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
@@ -682,7 +656,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
// argv[0] — the configured executable — is dropped and only the extra args are passed.
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
checkPaneCommandFits(cfg, argv);
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
@@ -698,59 +671,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
throw last;
}
/**
* fleetd #220: herdr does not exec the launch command — it TYPES it into the pane as one line,
* and a pty line buffer holds only {@value #PANE_COMMAND_BYTE_LIMIT} bytes (BSD/macOS {@code
* MAX_CANON}). Everything past that byte is dropped. Nothing reports it: herdr answers "agent
* started", the backend exits on the mangled argument it was handed, the pane closes, and the
* only symptom is {@link #waitUntilInjectableOrThrow} timing out 20 seconds later with no
* reason. That is exactly how #214 broke every claude-code spawn — one 50-byte flag pushed a
* 978-byte command to 1028, and the tail that got cut was {@code --autocompact 250000}.
*
* <p>So measure it here and refuse, loudly and immediately, rather than spawn something that
* cannot work. The estimate is deliberately conservative: fleetd cannot see herdr's quoting, so
* every argument is charged its own bytes plus a separator and a quote pair. An over-estimate
* costs a clear error at a length that was already unsafe; an under-estimate would let the
* silent truncation back in.
*
* @throws PeerUnreachableException when the command cannot fit — the same failure the spawn
* would have hit anyway, named at the point it is still
* explainable
*/
private void checkPaneCommandFits(FleetConfig.Profile cfg, List<String> argv) {
int bytes = 0;
String longest = null;
int longestBytes = 0;
for (String arg : argv) {
int argBytes = arg == null ? 0 : arg.getBytes(java.nio.charset.StandardCharsets.UTF_8).length;
bytes += argBytes + QUOTING_OVERHEAD_PER_ARG;
if (argBytes > longestBytes) {
longestBytes = argBytes;
longest = arg;
}
}
if (bytes <= PANE_COMMAND_BYTE_LIMIT) {
return;
}
String culprit = longest == null ? "<none>"
: longest.substring(0, Math.min(longest.length(), 60)) + (longest.length() > 60 ? "…" : "");
throw new PeerUnreachableException(
"launch command for profile " + cfg.profile() + " is about " + bytes + " bytes, over the "
+ PANE_COMMAND_BYTE_LIMIT + "-byte limit of the pane line herdr types it into. "
+ "The pty would drop the tail silently and the backend would exit on a mangled "
+ "argument. Longest argument is " + longestBytes + " bytes: " + culprit
+ " — move it off the command line (a file flag) or shorten it.");
}
/**
* The pty line buffer herdr types a launch command into: BSD/macOS {@code MAX_CANON}. Not a
* fleetd choice and not configurable — see {@link #checkPaneCommandFits}.
*/
static final int PANE_COMMAND_BYTE_LIMIT = 1024;
/** Per-argument allowance for the separating space and a shell quote pair fleetd cannot see. */
private static final int QUOTING_OVERHEAD_PER_ARG = 3;
/** Start the agent into {@code paneId}, waiting out the seed shell's boot with the sleeper. */
private Agent startAwaitingShellPrompt(String name, List<String> args, String paneId) {
HerdrException busy = null;
@@ -914,51 +834,20 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
*/
private void waitUntilInjectableOrThrow(String paneId) {
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
Object lastStatus = null;
while (nowMillis.getAsLong() < deadline) {
var status = agents.status(paneId);
lastStatus = status;
if (status.injectable()) {
if (agents.status(paneId).injectable()) {
log.debug("peer pane={} reached injectable state", paneId);
return;
}
sleeper.run();
}
// fleetd #220: read the pane BEFORE stop() closes it. Without this the gate says only that
// it timed out, which is true of every cause — a backend that never launched, a binary that
// rejected an argument and exited, a trust prompt, a login shell that hung. The pane holds
// the one copy of that answer and it is destroyed a line later.
log.warn("peer pane={} did not become injectable within {}ms (last status {}) — closing. "
+ "Pane tail:\n{}",
paneId, spawnReadyTimeoutMs, lastStatus, readPaneQuietly(paneId));
log.warn("peer pane={} did not become injectable within {}ms — closing", paneId, spawnReadyTimeoutMs);
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " did not reach injectable state within "
+ spawnReadyTimeoutMs + "ms");
}
/**
* fleetd #220: the pane's recent output, clipped, for the readiness-gate timeout log — or a
* short note when it cannot be read. Best-effort by construction: this runs on a path that is
* already failing, so it must never replace the real error with one of its own.
*/
private String readPaneQuietly(String paneId) {
try {
String pane = agents.read(paneId, "recent");
if (pane == null || pane.isBlank()) {
return "<pane read returned nothing>";
}
return pane.length() <= SPAWN_FAILURE_PANE_CHARS
? pane
: pane.substring(pane.length() - SPAWN_FAILURE_PANE_CHARS);
} catch (RuntimeException e) {
return "<pane could not be read: " + e.getMessage() + ">";
}
}
/** How much of a failed spawn's pane the timeout log carries. */
private static final int SPAWN_FAILURE_PANE_CHARS = 4000;
/**
* A concrete {@link PeerHandle} wrapping herdr agent coordinates, the profile that spawned it,
* the session identity the launch resolved (CB-547a): the bridge's logical name and the peer's
@@ -1111,15 +1000,8 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
*
* <p>CB-633: under {@code policy: allow-list} this overlay is NOT the control anymore — it is
* applied before the login shell runs and a sourced file can (and did) undo it. The control is
* the ZDOTDIR scrub ({@link #applyEnvironmentAllowListPolicy}); {@code known}/{@code allow}
* remain as reporting only via {@link #logCredentialGap}.
*
* <p>CB-633 follow-up (#192): under {@code allow-list} this method does NOT call {@link
* #logCredentialGap} itself — at this point (called from {@link #baseEnv}, before {@link
* #applyEnvironmentAllowListPolicy} runs) we do not yet know whether the pane's shell is zsh, so
* we cannot yet pick correct wording. That decision, and the call, are deferred entirely to
* {@link #applyEnvironmentAllowListPolicy}, which knows by then whether the scrub will actually
* run.
* the ZDOTDIR scrub ({@link #applyEnvironmentAllowListPolicy}); {@code allow} adds explicit
* keeps to its derived base, while {@code known} remains reporting metadata.
*/
private void applyMemberCredentialPolicy(Map<String, String> workerEnv) {
FleetConfig.MemberCredentials creds = memberCredentials == null ? null : memberCredentials.get();
@@ -1133,11 +1015,9 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
/** Put {@link #BLOCKED_CREDENTIAL_SENTINEL} over every blocked name in the pane-creation env map. */
private void overlayBlockedCredentials(Map<String, String> workerEnv,
FleetConfig.MemberCredentials creds) {
Set<String> blocked = new java.util.TreeSet<>(creds.blockedSet());
blocked.addAll(brokerUriEnvNames());
for (String name : blocked) {
private static void overlayBlockedCredentials(Map<String, String> workerEnv,
FleetConfig.MemberCredentials creds) {
for (String name : creds.blockedSet()) {
workerEnv.put(name, BLOCKED_CREDENTIAL_SENTINEL);
}
}
@@ -1153,201 +1033,79 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* pass it through {@code tab.create}/{@code pane.split}. Returns the directory for teardown
* registration, or {@code null} when the policy does not apply.
*
* <p>The allow-list handed to the generator is the derived profile set UNIONed with the
* operator's own {@code memberCredentials.allow:} names ({@link MemberEnvAllowList#derive(
* Collection, Set)} — CB-633 follow-up) and with the exact keys of THIS launch's env map —
* names the daemon itself injects must survive its own control. {@code SSH_AUTH_SOCK} is added
* ONLY when the config explicitly allows it, EVEN IF the operator also listed it under
* {@code allow:}; by default it is absent, so the scrub blanks it like any other non-derived
* name. It stays a one-off decision because it is a live handle to the operator's ssh-agent, not
* a value — a member holding it can sign with every key the agent holds, so letting it ride in
* on the generic {@code allow:} list would hand that out for an unrelated reason.
*
* <p>fleetd #213: which shell decides the zsh gate, and where the generated directory lives,
* both depend on whether {@code memberHerdrSocket:} is configured — see {@link
* #memberHerdrSocketConfigured()}'s javadoc for why fleetd's own {@code $SHELL} and {@code
* java.io.tmpdir} describe the wrong process once member panes run under a different OS user.
* <ul>
* <li>{@code memberHerdrSocket} ABSENT (today's only mode): byte-identical to before this
* fix — fleetd's own {@code $SHELL} decides zsh, and the directory is generated under
* {@code java.io.tmpdir}.</li>
* <li>{@code memberHerdrSocket} PRESENT: the configured {@code memberLoginShell:} decides
* zsh instead — fleetd's own {@code $SHELL} is never consulted, since it names a
* different user's shell, not the member's. Absent/non-zsh falls back exactly like the
* non-zsh case below. When it IS zsh, the directory still cannot go under {@code
* java.io.tmpdir} (mode 0700, unreadable by another uid — the exact gap fleetd #213
* exists to close), so it is generated under {@code worktreeRoot} instead and shared
* read-only with {@code worktreeGroup} — the same group {@link
* dev.ltms.fleet.session.Worktrees#shareWithGroup} already uses, reused rather than
* inventing a second group key. Either one missing means the scrub cannot be guaranteed
* reachable by the member, which is the same "cannot guarantee the scrub runs" case as a
* non-zsh shell, so it gets the identical fallback.</li>
* </ul>
* In every branch: never refuse to spawn. A degraded credential control must not become an
* outage for an opt-in feature.
* <p>The allow-list handed to the generator is the derived profile set ({@link
* MemberEnvAllowList#derive}) UNIONed with {@code memberCredentials.allow} and the exact keys of
* THIS launch's env map. This lets the operator keep inherited variables by name without putting
* their values in config, and names the daemon itself injects survive its own control. {@code
* SSH_AUTH_SOCK} is EXCLUDED from that union and added ONLY when {@code
* memberCredentials.sshAuthSock: allow} — {@code sshAuthSock} is the sole control for that one
* name, so it cannot be widened back in by naming it on {@code allow:} too; a conflicting entry
* there gets a WARN ({@link #warnSshAuthSockConflict}) and loses. By default {@code
* SSH_AUTH_SOCK} is absent, so the scrub blanks it like any other non-derived name.
*/
private Path applyEnvironmentAllowListPolicy(FleetConfig.Profile cfg, Launch launch) {
FleetConfig.MemberCredentials creds = memberCredentials == null ? null : memberCredentials.get();
if (creds == null || !creds.isAllowList()) {
return null;
}
Set<String> allowed = derivedAllowedNames(creds, launch);
boolean memberHerdrSocket = memberHerdrSocketConfigured();
// fleetd #213 defect 1: under memberHerdrSocket the member pane runs as a DIFFERENT OS
// user, so fleetd's own $SHELL says nothing about what that pane runs — resolveEnv("SHELL")
// must not even be called on this path, only the explicit memberLoginShell: config can
// answer it. With memberHerdrSocket absent, nothing here changes: fleetd's own $SHELL is
// still the input, exactly as before this fix.
String loginShell = memberHerdrSocket ? configuredMemberLoginShell() : resolveEnv("SHELL");
boolean zsh = isZshShell(loginShell);
String loginShell = resolveEnv("SHELL");
boolean zsh = loginShell != null && (loginShell.endsWith("/zsh") || loginShell.equals("zsh"));
if (!zsh) {
// A non-zsh login shell ignores ZDOTDIR entirely: NO scrub would run, so pretending
// otherwise would be worse than saying so. Warn loudly and fall back to the CB-596
// sentinel overlay over the enumerated known: names — weaker (a sourced file can undo
// it), but strictly better than nothing. Deliberately no "allowed N of M" line here: the
// scrub this count describes does not run on this path, so printing it would tell an
// operator that a fraction of names were blocked when the real number blocked is zero.
// logCredentialGap's WARN (below) is the only signal for this path.
// it), but strictly better than nothing.
warnNonZsh(loginShell);
overlayBlockedCredentials(launch.env(), creds);
// Nothing is scrubbed on this path — it is the CB-596 overlay only, which a sourced
// startup file can undo. Report the gap with the WARN wording (an inherited-unblocked
// exposure), not the allow-list INFO wording, which would falsely claim a scrub blanks it.
logCredentialGap(creds, null);
return null;
}
Path parentDir;
String group = null;
if (memberHerdrSocket) {
// fleetd #213 defect 2: java.io.tmpdir is fleetd's own per-user temp dir (mode 0700 on
// macOS) — a member running as a different uid cannot even traverse it, let alone read
// the generated files. worktreeRoot is the only configured location a different-uid
// member can be given access to, and only WITH worktreeGroup to grant that access —
// absent either, the scrub cannot be guaranteed reachable, so this falls back exactly
// like the non-zsh case above rather than generating a directory nothing can read.
parentDir = memberScrubParentDir();
group = memberGroup();
if (parentDir == null || group == null) {
warnCannotShareScrubDirectory();
overlayBlockedCredentials(launch.env(), creds);
logCredentialGap(creds, null);
return null;
Set<String> allowed = new java.util.TreeSet<>(MemberEnvAllowList.derive(profiles.values()));
for (String name : creds.allow()) {
// SSH_AUTH_SOCK is governed ONLY by sshAuthSock: below, never by allow: — an operator
// who writes both `sshAuthSock: block` and `SSH_AUTH_SOCK` on `allow:` means the block
// to win, not to be silently widened away by the second control.
if (!SSH_AUTH_SOCK.equals(name)) {
allowed.add(name);
}
} else {
parentDir = Path.of(System.getProperty("java.io.tmpdir"));
}
// Only reached when the scrub is actually about to run — the count below describes that
// scrub, so it must not be logged before this gate (see the non-zsh branch above). Same
// reasoning gates logCredentialGap's wording: passing the derived `allowed` set (non-null)
// here, and ONLY here, is what tells it the scrub will really blank an unkept name — #192.
logAllowListCoverage(allowed);
if (creds.sshAuthSockAllowed()) {
allowed.add(SSH_AUTH_SOCK);
} else if (creds.allow().contains(SSH_AUTH_SOCK)) {
warnSshAuthSockConflict();
} // blocked by default: absent from the set ⇒ blanked by the scrub like any other name
allowed.addAll(launch.env().keySet());
logCredentialGap(creds, allowed);
Path dir = memberHerdrSocket
? EnvAllowListScrub.generate(parentDir, allowed, group)
: EnvAllowListScrub.generate(parentDir, allowed);
Path dir = EnvAllowListScrub.generate(Path.of(System.getProperty("java.io.tmpdir")), allowed);
launch.env().put("ZDOTDIR", dir.toAbsolutePath().toString());
log.info("memberCredentials policy=allow-list: profile={} generated ZDOTDIR {} — derived "
log.info("memberCredentials policy=allow-list: profile={} generated ZDOTDIR {} — effective "
+ "allow-list holds {} name(s); the pane reports allowed N of M at release",
cfg.profile(), dir.getFileName(), allowed.size());
return dir;
}
/** True when {@code shell} is a zsh login shell path or bare name — the ZDOTDIR gate. */
private static boolean isZshShell(String shell) {
return shell != null && (shell.endsWith("/zsh") || shell.equals("zsh"));
}
/**
* fleetd #213: the configured {@code memberLoginShell:}, or {@code null} when unconfigured.
* Called ONLY from the {@code memberHerdrSocket}-configured branch of {@link
* #applyEnvironmentAllowListPolicy} — fleetd's own {@code $SHELL} is never read on that path.
* {@link #config} being {@code null} (an older test call site, or a launcher that never
* threaded the full config through) is treated the same as "not configured".
*/
private String configuredMemberLoginShell() {
FleetConfig cfg = config == null ? null : config.get();
return cfg == null ? null : cfg.memberLoginShell();
}
/**
* fleetd #213: {@code worktreeRoot}, as the ZDOTDIR scrub's parent directory under {@code
* memberHerdrSocket}, or {@code null} when unconfigured — the same "cannot guarantee the scrub
* runs" gap as {@link #memberGroup()} being unset (see {@link
* #applyEnvironmentAllowListPolicy}). Deliberately no sibling-of-repo-root default here, unlike
* {@code GitWorktrees}' own {@code worktreeRoot} resolution: that default is a convenience for
* provisioning a worktree that will exist regardless, whereas an unconfigured value here means
* fleetd has no operator-endorsed location to put a credential-bearing directory a different OS
* user must reach, so falling back to the overlay is the honest answer, not a guess.
*
* <p>Package-private (fleetd #219) so {@link OpenCodeLauncher} can reuse the exact same
* "different OS user, put it under worktreeRoot instead of java.io.tmpdir" resolution for its
* own ephemeral {@code opencode.json} directory, rather than re-reading {@code config} a second
* time with a second copy of this null/blank handling.
*/
Path memberScrubParentDir() {
FleetConfig cfg = config == null ? null : config.get();
if (cfg == null || cfg.worktreeRoot() == null || cfg.worktreeRoot().isBlank()) {
return null;
}
return Path.of(cfg.worktreeRoot());
}
/**
* fleetd #213: the configured {@code worktreeGroup:}, or {@code null} when unset/blank. Reuses
* the group {@link dev.ltms.fleet.session.Worktrees#shareWithGroup} already establishes for
* provisioned worktrees, rather than a second group key — see {@link
* #applyEnvironmentAllowListPolicy}.
*
* <p>Package-private (fleetd #219) — reused by {@link OpenCodeLauncher} alongside {@link
* #memberScrubParentDir()}; see that method's javadoc.
*/
String memberGroup() {
FleetConfig cfg = config == null ? null : config.get();
if (cfg == null || cfg.worktreeGroup() == null || cfg.worktreeGroup().isBlank()) {
return null;
}
return cfg.worktreeGroup();
}
/**
* The full kept-name set for this spawn: the profile-derived names, unioned with {@code
* memberCredentials.allow:} (CB-633 follow-up — previously ignored by this whole policy), the
* ssh-agent handle when explicitly allowed, and the exact keys of THIS launch's own env map.
*/
private Set<String> derivedAllowedNames(FleetConfig.MemberCredentials creds, Launch launch) {
Set<String> brokerUriEnvNames = brokerUriEnvNames();
Set<String> allowed = new java.util.TreeSet<>(
MemberEnvAllowList.derive(profiles.values(), creds.allowSet(), brokerUriEnvNames));
if (creds.sshAuthSockAllowed()) {
allowed.add(SSH_AUTH_SOCK);
} // blocked by default: absent from the set ⇒ blanked by the scrub like any other name
allowed.addAll(launch.env().keySet());
allowed.removeAll(brokerUriEnvNames);
return allowed;
}
private Set<String> brokerUriEnvNames() {
return MemberEnvAllowList.brokerUriEnvNames(config == null ? null : config.get());
}
/**
* CB-633 follow-up: one INFO line per allow-list spawn WHOSE SCRUB ACTUALLY RUNS, so an operator
* can read a single log line and know the scrub ran and how much of the visible environment it
* will keep. Callable ONLY from the zsh branch of {@link #applyEnvironmentAllowListPolicy}, after
* the shell gate — logging it before that gate (or on the non-zsh fallback, where nothing is
* scrubbed) would tell an operator a fraction of names were blocked when the real number blocked
* is zero, which is worse than not logging at all. {@code M} is {@link #hostEnvNames}' size (the
* daemon's own environment — see that field's javadoc for why it stands in for the pane's, which
* the daemon has no channel to inspect at spawn time) and {@code N} is how many of those names
* survive {@code allowed} (including the {@code LC_*} prefix rule). Neither number is a constant:
* both come from the actual derived set and the actual environment this spawn sees. Never logs a
* variable NAME or VALUE — only the counts.
*/
private void logAllowListCoverage(Set<String> allowed) {
Set<String> hostNames = hostEnvNames.get();
long kept = hostNames.stream().filter(name -> MemberEnvAllowList.keeps(allowed, name)).count();
log.info("member credentials: allowed {} of {}", kept, hostNames.size());
}
/** The operator ssh-agent handle — kept ONLY by explicit config decision, never by default. */
private static final String SSH_AUTH_SOCK = MemberEnvAllowList.SSH_AUTH_SOCK;
private static final String SSH_AUTH_SOCK = "SSH_AUTH_SOCK";
/** Guards {@link #warnSshAuthSockConflict} to one WARN per launcher instance, not one per spawn. */
private final AtomicBoolean sshAuthSockConflictWarned = new AtomicBoolean();
/**
* Defect fix: {@code SSH_AUTH_SOCK} named on {@code memberCredentials.allow} while
* {@code sshAuthSock} is not {@code allow} is a conflict between the two controls — say which
* one wins, once per launcher instance, without printing any value.
*/
private void warnSshAuthSockConflict() {
if (sshAuthSockConflictWarned.compareAndSet(false, true)) {
log.warn("memberCredentials policy=allow-list: SSH_AUTH_SOCK is named on allow: but "
+ "sshAuthSock is not 'allow' — sshAuthSock: block wins and SSH_AUTH_SOCK "
+ "stays blocked. Remove it from allow: or set sshAuthSock: allow if the "
+ "member should keep it.");
}
}
/**
* CB-633: a non-zsh login shell means the allow-list control CANNOT run — say so once per
@@ -1363,30 +1121,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
}
/** Guards {@link #warnCannotShareScrubDirectory} to one WARN per launcher instance. */
private final AtomicBoolean cannotShareScrubDirWarned = new AtomicBoolean();
/**
* fleetd #213: {@code memberHerdrSocket} is configured and the member login shell IS zsh, but
* {@code worktreeRoot} and/or {@code worktreeGroup} is missing, so the generated ZDOTDIR cannot
* be placed anywhere the member's OS user can reach — {@code java.io.tmpdir} is fleetd's own
* 0700 temp dir, unreadable by another uid, which is the exact gap this ticket exists to close.
* Say so once per launcher instance, instead of either generating a directory nothing can read
* (protection theatre) or refusing to spawn (turning a degraded credential control into an
* outage for an opt-in feature).
*/
private void warnCannotShareScrubDirectory() {
if (cannotShareScrubDirWarned.compareAndSet(false, true)) {
log.warn("memberCredentials policy=allow-list: memberHerdrSocket is configured and the "
+ "member login shell is zsh, but worktreeRoot and/or worktreeGroup is not "
+ "configured — the generated ZDOTDIR cannot be placed where the member's OS "
+ "user can read it (java.io.tmpdir is fleetd's own, unreadable by another uid), "
+ "so the scrub cannot be guaranteed to run. Falling back to the CB-596 sentinel "
+ "overlay. Configure both worktreeRoot and worktreeGroup to enable the "
+ "allow-list scrub under memberHerdrSocket.");
}
}
/**
* CB-633 teardown half: read the pane's scrub report (the denominator report the generated
* scrub wrote) and delete the directory. Called from {@link #stop}, which is the one funnel
@@ -1433,176 +1167,58 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
Pattern.compile("(?i).*(TOKEN|SECRET|_KEY|APIKEY|PASSWORD|CREDENTIAL|AUTH).*");
/**
* Guards the {@code effectiveAllowed == null} branch of {@link #logCredentialGap} — the
* genuinely-unprotected report (deny-by-default, and the allow-list non-zsh fallback) — to one
* WARN per launcher instance, not one per spawn.
*
* <p>CB-633 follow-up (#192): kept SEPARATE from {@link #allowListGapLogged} on purpose.
* {@code memberCredentials} is a live, re-read-per-spawn supplier, so the policy can change
* between two spawns on the same launcher. A single shared flag would let a harmless allow-list
* INFO on spawn 1 permanently suppress the real deny-by-default WARN a later spawn deserves —
* the report that matters most getting hidden by the report that doesn't. Two flags mean each
* report kind fires exactly once, independent of what the other kind already logged.
*/
private final AtomicBoolean unprotectedGapLogged = new AtomicBoolean();
/**
* Guards the {@code effectiveAllowed != null} branch of {@link #logCredentialGap} — the
* allow-list-scrub-covered report — to one INFO per launcher instance. See {@link
* #unprotectedGapLogged}'s javadoc for why this is a separate flag rather than a shared one.
* Guards {@link #logCredentialGap}'s two report kinds separately — one flag per wording, not
* one shared flag — so an early INFO ("will be blanked by the scrub") on one spawn can never
* suppress a later, more serious WARN ("inherited UNBLOCKED") on another. {@code
* memberCredentials} is live-reloadable, so the policy really can change between spawns on the
* same launcher instance.
*/
private final AtomicBoolean allowListGapLogged = new AtomicBoolean();
/**
* fleetd #185 stage 2: guards {@link #warnUnknownMemberEnvironment} to one WARN per launcher
* instance, not one per spawn — the same one-per-instance shape as {@link #unprotectedGapLogged}
* and {@link #allowListGapLogged}, kept as its own flag for the same reason those two are split:
* this mode is orthogonal to which of the other two branches would otherwise have fired.
*/
private final AtomicBoolean unknownMemberEnvironmentWarned = new AtomicBoolean();
/** See {@link #allowListGapLogged} — the WARN-wording counterpart. */
private final AtomicBoolean unprotectedGapLogged = new AtomicBoolean();
/**
* fleetd #185 stage 2: whether {@code memberHerdrSocket:} is configured, i.e. member panes run
* on a second herdr owned by a different OS user than the daemon's own process. Re-read from the
* live config on every call (same hot-reload shape as {@link #memberCredentials}), never cached,
* so a config reload takes effect on the next spawn without a restart.
*
* <p>{@link #config} is {@code null} on any call site that never threaded the full config
* through (every production {@code HerdrPeerLauncher} does; a handful of older tests do not) —
* treated the same as "not configured", which is the correct, permissive default: it is exactly
* today's single-daemon behaviour.
*
* <p>Package-private (fleetd #219) — {@link OpenCodeLauncher} reuses this same gate to decide
* where its own ephemeral {@code opencode.json} directory (site 1) and its opencode session
* discovery (site 2) may run, rather than re-deriving "is this a multi-uid fleet" a second way.
*/
boolean memberHerdrSocketConfigured() {
if (config == null) {
return false;
}
FleetConfig cfg = config.get();
return cfg != null && cfg.memberHerdrSocket() != null && !cfg.memberHerdrSocket().isBlank();
}
/**
* fleetd #185 stage 2: the single replacement WARN for {@link #logCredentialGap}'s usual
* conclusions when {@code memberHerdrSocket:} is configured. {@link #hostEnvNames} (and
* everything derived from it — {@code known}/{@code allow} coverage, the allow-list scrub's
* derived set) describes the DAEMON's own environment; under this config key member panes run as
* a different OS user with a different environment entirely, so neither "every member pane
* inherits them UNBLOCKED" nor "the scrub blanks them" is evidence-backed here — both would be
* reporting on the wrong process. Logged once, names the config key, and states the honest
* conclusion: the gap for member panes is UNKNOWN, not clean, so {@code memberCredentials} cannot
* be verified from this daemon. The one count it does report is scoped explicitly to fleetd's own
* environment, never presented as if it said anything about the member's — see {@link
* #logCredentialGap}'s javadoc for why this branch exists.
*/
private void warnUnknownMemberEnvironment(FleetConfig.MemberCredentials creds) {
if (!unknownMemberEnvironmentWarned.compareAndSet(false, true)) {
return;
}
Set<String> covered = new HashSet<>(creds.known());
covered.addAll(creds.allow());
Set<String> hostNames = hostEnvNames.get();
long gapInFleetdsOwnEnv = hostNames.stream()
.filter(name -> CREDENTIAL_SHAPED_NAME.matcher(name).matches())
.filter(name -> !covered.contains(name))
.count();
log.warn("memberCredentials gap: memberHerdrSocket is configured, so member panes run under "
+ "a different OS user than fleetd's own process, with a different environment "
+ "entirely — fleetd has no channel to read that user's environment. {} of the "
+ "{} names in fleetd's OWN environment are credential-shaped and not on "
+ "known:/allow:, but that count describes fleetd's process, not the member "
+ "herdr's. The credential gap for member panes is UNKNOWN, not clean, and "
+ "memberCredentials cannot be verified from here.",
gapInFleetdsOwnEnv, hostNames.size());
}
/**
* CB-596 criterion 4: a credential-shaped host env var name on neither {@code known} nor
* {@code allow} is not silently allowed — it is reported. {@link #hostEnvNames} enumerates the
* CB-596 criterion 4: report credential-shaped host env var names that the active policy does
* not classify. Callers that pass a non-null {@code effectiveAllowed} are reporting against a
* kept-name set a scrub will actually enforce — an absent name gets an INFO stating that the
* scrub will blank it, because that is the control working rather than an exposure. Every
* other caller — deny-by-default, and an allow-list spawn whose scrub could not run (non-zsh
* login shell, falls back to the overlay) — passes {@code null} and gets the original WARN,
* because in both cases a gap name is inherited unblocked. {@link #hostEnvNames} enumerates the
* daemon's own environment (see that field's javadoc for why the daemon's env is read rather
* than the spawned pane's, which the daemon has no channel to inspect at spawn time); this logs
* every such NAME — never a value, a prefix of a value, or a hash of a value, so the log itself
* cannot leak anything.
*
* <p>CB-633 follow-up (#192): {@code effectiveAllowed} picks the wording, and it must NOT be
* picked from {@code creds.isAllowList()} — see {@link #applyEnvironmentAllowListPolicy}'s
* javadoc for the reasoning this mirrors. {@code null} means no scrub-derived allow-list was
* computed for this call — true on the deny-by-default path AND on the allow-list non-zsh
* fallback, where nothing is ever scrubbed — so the whole gap is real and gets the WARN,
* unchanged from before this fix. Non-null means this call came from the zsh branch of {@link
* #applyEnvironmentAllowListPolicy}, reachable ONLY after that method's own zsh gate — but
* {@code effectiveAllowed} is a SUPERSET of {@code known ∪ allow}: {@link MemberEnvAllowList#derive}
* also unions in every profile's {@code gitTokenEnv}/{@code gitHostEnv}/{@code tokenEnv}/
* {@code env:} keys, and {@link #derivedAllowedNames} further unions in this very spawn's own
* env keys — so a name can be in the gap (uncovered by {@code known}/{@code allow}) AND still be
* kept by the derived allow-list, in which case the scrub does NOT blank it and the member DOES
* inherit it. Lead review on #192 caught this: the first cut of this fix reported the WHOLE gap
* as scrub-blanked without checking that, which reported a real leak as safe — the exact
* inversion #192 exists to remove. So on this path the gap is split with {@link
* MemberEnvAllowList#keeps}, the SAME predicate the generated scrub itself evaluates, so this
* split cannot drift from what the scrub actually does: the names it says are kept get the WARN
* (same severity, and same guard, as the deny-by-default case — a name genuinely reaching a
* member unprotected is equally serious whichever path put it there), and the names it says are
* blanked keep the INFO.
*
* <p>fleetd #185 stage 2: everything above assumes the member pane runs under the same OS user
* as the daemon, so {@link #hostEnvNames} mirrors what the pane inherits — see that field's
* javadoc. When {@code memberHerdrSocket:} is configured that assumption is false: the member
* pane runs on a second herdr owned by a <em>different</em> user, and neither conclusion below
* ("inherits them UNBLOCKED" / "the scrub blanks them") is backed by evidence about that user's
* environment. So this method checks that first and, when configured, reports the honest
* "unknown, not clean" conclusion instead — see {@link #warnUnknownMemberEnvironment}. When
* {@code memberHerdrSocket:} is absent (the default, and the only mode this host runs) this
* branch is never taken and every line below is unchanged.
* than the spawned pane's). This logs names only, at most once per launcher instance per report
* kind — never a value, a prefix of a value, or a hash of a value, so the log itself cannot
* leak anything.
*/
private void logCredentialGap(FleetConfig.MemberCredentials creds, Set<String> effectiveAllowed) {
if (memberHerdrSocketConfigured()) {
warnUnknownMemberEnvironment(creds);
return;
Set<String> covered = effectiveAllowed == null
? new HashSet<>(creds.known()) : effectiveAllowed;
if (effectiveAllowed == null) {
covered.addAll(creds.allow());
}
Set<String> covered = new HashSet<>(creds.known());
covered.addAll(creds.allow());
List<String> gap = hostEnvNames.get().stream()
.filter(name -> CREDENTIAL_SHAPED_NAME.matcher(name).matches())
.filter(name -> !covered.contains(name))
.filter(name -> effectiveAllowed == null
? !covered.contains(name) : !MemberEnvAllowList.keeps(covered, name))
.sorted()
.toList();
if (gap.isEmpty()) {
return;
}
if (effectiveAllowed == null) {
warnGapUnprotected(gap);
return;
}
List<String> keptByDerivedList = gap.stream()
.filter(name -> MemberEnvAllowList.keeps(effectiveAllowed, name))
.toList();
List<String> blankedByScrub = gap.stream()
.filter(name -> !MemberEnvAllowList.keeps(effectiveAllowed, name))
.toList();
if (!keptByDerivedList.isEmpty() && unprotectedGapLogged.compareAndSet(false, true)) {
log.warn("memberCredentials gap: {} credential-shaped env var name(s) are on neither "
+ "known: nor allow: — the derived allow-list keeps them anyway (a profile's "
+ "gitTokenEnv/gitHostEnv/tokenEnv/env: names one, or this spawn injects it), "
+ "so every member pane inherits them UNBLOCKED — {}. Add each to "
+ "memberCredentials.known (or .allow if a member legitimately needs it), or "
+ "remove it from whatever profile setting derives it in.",
keptByDerivedList.size(), keptByDerivedList);
}
if (!blankedByScrub.isEmpty() && allowListGapLogged.compareAndSet(false, true)) {
log.info("memberCredentials gap: {} credential-shaped env var name(s) are on neither "
+ "known: nor allow: — {}. The allow-list scrub blanks them anyway (they "
+ "are not on the derived allow-list), so no member pane keeps them; add "
+ "each to memberCredentials.known or .allow to make that explicit.",
blankedByScrub.size(), blankedByScrub);
}
}
/** The deny-by-default (and allow-list non-zsh fallback) WARN — unchanged byte-for-byte by #192. */
private void warnGapUnprotected(List<String> gap) {
if (unprotectedGapLogged.compareAndSet(false, true)) {
// effectiveAllowed != null means a derived allow-list set was actually computed (the
// scrub will run against it) — that is the INFO case. A null effectiveAllowed means no
// scrub protects this spawn (deny-by-default, or an allow-list spawn that fell back to
// the overlay-only path because the login shell is not zsh) — that is the WARN case.
// See #allowListGapLogged for why the two kinds use separate guards.
if (effectiveAllowed != null) {
if (allowListGapLogged.compareAndSet(false, true)) {
log.info("memberCredentials allow-list: {} credential-shaped host env var name(s) "
+ "are not kept and will be blanked by the scrub — {}. Add any name "
+ "a member legitimately needs to memberCredentials.allow.",
gap.size(), gap);
}
} else if (unprotectedGapLogged.compareAndSet(false, true)) {
log.warn("memberCredentials gap: {} credential-shaped env var name(s) are on neither "
+ "known: nor allow: — every member pane inherits them UNBLOCKED — {}. "
+ "Add each to memberCredentials.known (blocked by default) or .allow "
@@ -7,14 +7,16 @@ import java.util.Set;
import java.util.TreeSet;
/**
* CB-633: the set of environment variable NAMES a spawned member is allowed to keep under
* {@code memberCredentials.policy: allow-list} — DERIVED from what the launcher itself injects,
* never hand-typed.
* CB-633: the base set of environment variable NAMES a spawned member is allowed to keep under
* {@code memberCredentials.policy: allow-list}. This class derives the base from what the launcher
* itself injects. The launcher then adds the operator's explicitly named {@code
* memberCredentials.allow} keeps and the exact keys of this launch's env map.
*
* <p>A hand-typed allow-list is the defect this class exists to prevent: a name an operator forgets
* to type is a credential that passes through to every member, and a profile added to config later
* would silently break spawns whose scrub did not know its names. Derivation closes both ends. The
* kept-name set is the union of:
* <p>A hand-typed list that <em>replaces</em> derivation is the defect this class exists to prevent:
* a profile added to config later would silently break spawns whose scrub did not know its names.
* An explicit list that only adds keeps is safe because adding a name can only widen the set; it
* cannot make another spawn lose a name that derivation already kept. The derived base is the union
* of:
*
* <ul>
* <li>every configured {@link FleetConfig.Profile profile}'s {@code gitTokenEnv},
@@ -26,37 +28,18 @@ import java.util.TreeSet;
* or the agent binary genuinely needs to function.</li>
* </ul>
*
* <p>Because the union spans EVERY profile (not just the one spawning), adding a new profile can
* only ever widen the list — it cannot break another spawn's scrub. And because the launcher also
* unions in the exact keys of each spawn's own env map at generation time (see {@code
* HerdrPeerLauncher}), anything the daemon deliberately injects for THIS spawn survives its own
* control.
* <p>Because the base spans EVERY profile (not just the one spawning), adding a new profile can only
* ever widen the list — it cannot break another spawn's scrub. The launcher then adds the operator's
* explicit keeps and the exact keys of each spawn's own env map at generation time (see {@code
* HerdrPeerLauncher}), so member binaries can keep named inherited variables without putting their
* secret values in config, and anything the daemon injects for THIS spawn survives its own control.
*
* <p>{@code SSH_AUTH_SOCK} is deliberately NOT here. It is a handle to the operator's ssh-agent — a
* member holding it can sign with the operator's keys — so keeping it is a config decision
* ({@code memberCredentials.sshAuthSock: allow}), not a derivation default.
*
* <p><b>CB-633 follow-up:</b> the union also includes {@code memberCredentials.allow:} — the
* operator's own explicit list. Before this, {@code policy: allow-list} silently ignored every name
* an operator wrote under {@code allow:} unless a profile happened to carry it too, which meant
* turning the policy on could blank credentials working members already depended on. {@code
* SSH_AUTH_SOCK} and configured broker URI environment names are exceptions: even when the operator
* lists them under {@code allow:}, they are excluded here. {@code SSH_AUTH_SOCK} is added back ONLY
* by the caller when {@code sshAuthSock: allow} is explicitly set
* (see {@link #SSH_AUTH_SOCK}'s javadoc) — it is a live handle to the operator's own ssh-agent, not
* a value, so treating it like any other allow-listed name would hand a member every key the
* operator's agent holds the moment they typed the name under {@code allow:} for an unrelated
* reason.
*/
public final class MemberEnvAllowList {
/**
* The operator's ssh-agent socket path. Deliberately excluded from {@link #derive}'s union of
* {@code memberCredentials.allow:} — see the class javadoc's CB-633 follow-up note. Governed
* ONLY by {@code memberCredentials.sshAuthSock}, never by appearing in {@code allow:}.
*/
public static final String SSH_AUTH_SOCK = "SSH_AUTH_SOCK";
/**
* Names that are not credentials and that a login shell or agent binary genuinely needs.
*
@@ -92,31 +75,9 @@ public final class MemberEnvAllowList {
/**
* Derive the allowed NAME set from the given profiles plus {@link #INFRASTRUCTURE_PASSTHROUGH}.
* Equivalent to {@link #derive(Collection, Set)} with no operator-configured names — kept for
* callers (and existing tests) that only care about the profile-derived half.
* Deterministic (sorted) so generated scrub files are diffable run-to-run.
*/
public static Set<String> derive(Collection<FleetConfig.Profile> profiles) {
return derive(profiles, Set.of());
}
/**
* Derive the allowed NAME set: the profile-derived union above, PLUS {@code configuredAllow} —
* the operator's own {@code memberCredentials.allow:} list (CB-633 follow-up). {@code
* SSH_AUTH_SOCK} is dropped from {@code configuredAllow} even if the operator listed it there;
* see the class javadoc for why. Deterministic (sorted) so generated scrub files are diffable
* run-to-run.
*/
public static Set<String> derive(Collection<FleetConfig.Profile> profiles, Set<String> configuredAllow) {
return derive(profiles, configuredAllow, Set.of());
}
/**
* As {@link #derive(Collection, Set)}, while excluding names that fleetd knows carry credentials.
* A configured broker URI contains its AMQP password inline, so it must never reach a member,
* even when an operator put its variable name in {@code memberCredentials.allow:}.
*/
public static Set<String> derive(Collection<FleetConfig.Profile> profiles, Set<String> configuredAllow,
Set<String> excludedNames) {
Set<String> derived = new TreeSet<>(INFRASTRUCTURE_PASSTHROUGH);
if (profiles != null) {
for (FleetConfig.Profile p : profiles) {
@@ -128,44 +89,9 @@ public final class MemberEnvAllowList {
}
}
}
if (configuredAllow != null) {
for (String name : configuredAllow) {
if (name != null && !name.isBlank() && !SSH_AUTH_SOCK.equals(name)) {
derived.add(name);
}
}
}
if (excludedNames != null) {
derived.removeAll(excludedNames);
}
return Set.copyOf(derived);
}
/**
* The host environment names whose values are AMQP URIs with inline passwords. Both broker
* connections belong to fleetd, never to a member pane. Blank and absent configuration changes
* nothing.
*
* <p><b>How strong this exclusion is depends on the policy, and the difference matters.</b>
* Under {@code policy: allow-list} it is enforced by the generated ZDOTDIR scrub, which runs
* AFTER the pane's shell has sourced the operator's chain — so a login shell that re-exports the
* name is still blanked. Under the deny-list policy there is no scrub: the name is only removed
* from the pre-shell env map, and a login shell that sources the operator's secret store
* re-exports it. That is the long-standing weakness of deny-list (a sourced file can undo it),
* not something this exclusion introduces, but it means deny-list deployments do NOT get this
* guarantee. The same caveat applies to the non-zsh path, which has no scrub at all — see
* {@code HerdrPeerLauncher#applyEnvironmentAllowListPolicy}.
*/
public static Set<String> brokerUriEnvNames(FleetConfig config) {
if (config == null) {
return Set.of();
}
Set<String> names = new TreeSet<>();
addUriEnvIfPresent(names, config.broker());
addUriEnvIfPresent(names, config.coordinator());
return Set.copyOf(names);
}
/**
* Whether {@code name} survives the scrub when {@code allowedNames} is the derived set: an exact
* match, or an infrastructure-prefixed name ({@code LC_*}). Prefix rules live ONLY here and in
@@ -185,16 +111,4 @@ public final class MemberEnvAllowList {
into.add(name);
}
}
private static void addUriEnvIfPresent(Set<String> into, FleetConfig.Broker broker) {
if (broker != null && broker.hasUriEnv()) {
into.add(broker.uriEnv());
}
}
private static void addUriEnvIfPresent(Set<String> into, FleetConfig.Coordinator coordinator) {
if (coordinator != null && coordinator.hasUriEnv()) {
into.add(coordinator.uriEnv());
}
}
}
@@ -20,8 +20,6 @@ import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.function.BooleanSupplier;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
@@ -118,23 +116,12 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs, spawnReadyPollMs,
fleet, memberCredentials, null);
}
/** Production constructor, plus the live config for URI environment exclusions. */
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Function<String, String> env, long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<FleetConfig> config) {
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot(), defaultDiscoveryRoot(), fleet, memberCredentials, config);
defaultConfigRoot(), defaultDiscoveryRoot(), fleet, memberCredentials);
}
/**
@@ -195,20 +182,8 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
Path configRoot, Path discoveryRoot,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs, nowMillis, sleeper,
configRoot, discoveryRoot, fleet, memberCredentials, null);
}
/** Full testability constructor, plus the live config for URI environment exclusions. */
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Function<String, String> env, long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper, Path configRoot, Path discoveryRoot,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<FleetConfig> config) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials, null, config);
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials);
this.configRoot = configRoot;
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
}
@@ -217,43 +192,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
return Path.of(System.getProperty("java.io.tmpdir"));
}
/**
* The default opencode storage root: {@code ~/.local/share/opencode} (the XDG data dir) —
* always FLEETD's OWN {@code user.home}, whichever OS user runs the daemon.
*
* <p><b>fleetd #219 site 2 — a decision, not a patch.</b> Under {@code memberHerdrSocket:} the
* member pane runs as a <em>different</em> OS user, and opencode writes {@code opencode.db}
* under <em>that</em> user's {@code $HOME}, not fleetd's. Scanning fleetd's own {@code
* user.home} is therefore looking in the wrong place — a wrong-LOCATION failure, not a
* wrong-PERMISSION one like site 1, and it fails quietly: {@link
* SessionAwareHandle#agentSessionId()} would keep returning {@code null} forever, which reads
* as "opencode does not support resume" rather than "fleetd looked in the wrong home." fleetd
* #209 is the reason that silence is unacceptable.
*
* <p>Three ways to close the gap were weighed:
* <ol>
* <li><b>Make the member's home configurable.</b> Correct in principle, but this ticket's
* scope is the two existing call sites, not a new config key — {@code memberHerdrSocket}
* already carries the second herdr's socket path, not its user's home, and inventing a
* parallel key here without also wiring it through discovery's actual callers is a
* half-shipped feature (the exact shape CB-596/CB-611 warn against).</li>
* <li><b>Derive it</b> (e.g. from {@code worktreeRoot}'s owner, or {@code getent passwd}).
* Rejected: nothing in this codebase resolves a Unix username to a home directory today,
* and guessing wrong would silently point discovery at a THIRD wrong location — worse
* than the current gap, because it would look like it should work.</li>
* <li><b>Declare discovery unavailable</b> under {@code memberHerdrSocket}, and say so once,
* loudly, instead of scanning a directory that structurally cannot hold the answer.</li>
* </ol>
*
* <p>Option 3 is taken — the one this ticket says to default to when unsure. {@link
* OpenCodeLauncher#spawn} routes {@link SessionAwareHandle#agentSessionId()} through {@link
* HerdrPeerLauncher#memberHerdrSocketConfigured()} before ever calling {@link
* OpenCodeSessionDiscovery#sessionIdForDirectory}, so under {@code memberHerdrSocket} the
* database at this root is never even opened, and one WARN per launcher instance names the gap
* instead of the {@code null} return reading as "unsupported." Capability advertising is
* unaffected: {@link #capabilities()} always includes {@code SESSION_RESUME}, since {@code
* memberHerdrSocket} absent (today's only live mode) is unchanged by this decision.
*/
/** The default opencode storage root: {@code ~/.local/share/opencode} (the XDG data dir). */
private static Path defaultDiscoveryRoot() {
return Path.of(System.getProperty("user.home"), ".local", "share", "opencode");
}
@@ -369,7 +308,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
*/
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd) {
try {
Path dir = Files.createTempDirectory(configParentDir(), "fleetd-opencode-");
Path dir = Files.createTempDirectory(configRoot, "fleetd-opencode-");
dir.toFile().deleteOnExit();
ObjectNode root = JSON.createObjectNode();
@@ -440,14 +379,6 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
// carries operator-supplied values (URL, model id, api key), so escaping must be real.
Files.writeString(cfgFile, JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root));
cfgFile.toFile().deleteOnExit();
if (memberHerdrSocketConfigured()) {
// fleetd #219: the same "different OS user" gap fleetd #213 closed for the ZDOTDIR
// scrub — share read-only with worktreeGroup rather than leaving the directory under
// fleetd's own 0700 java.io.tmpdir, where the member's OS user could not even
// traverse it. memberGroup() cannot be null here: configParentDir() above already
// refused this spawn if either worktreeRoot or worktreeGroup was missing.
EnvAllowListScrub.shareWithGroup(dir, memberGroup());
}
return cfgFile;
} catch (IOException e) {
throw new UncheckedIOException(
@@ -455,57 +386,6 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
}
}
/**
* fleetd #219 site 1: where {@link #writeConfig} creates its per-spawn directory.
*
* <ul>
* <li>{@code memberHerdrSocket} ABSENT (today's only mode): byte-identical to before this
* fix — always {@link #configRoot} (defaults to {@code java.io.tmpdir}, fleetd's own
* process).</li>
* <li>{@code memberHerdrSocket} PRESENT: {@code java.io.tmpdir} is fleetd's own per-user temp
* dir (mode {@code 0700} on macOS) — the member pane runs as a DIFFERENT OS user under
* this config key and cannot even traverse it, so the directory holding {@code
* opencode.json} (which tells the member where the bridge MCP is) and the member charter
* would be unreadable to the very process it is written for. The directory instead goes
* under {@code worktreeRoot}, shared read-only with {@code worktreeGroup} via {@link
* EnvAllowListScrub#shareWithGroup} — the SAME mechanism fleetd #213 built for the ZDOTDIR
* scrub, reused here rather than duplicated (see {@link
* HerdrPeerLauncher#memberScrubParentDir()}).</li>
* </ul>
*
* <p><b>Unlike the ZDOTDIR scrub, a missing {@code worktreeRoot}/{@code worktreeGroup} here
* REFUSES the spawn instead of degrading.</b> The ZDOTDIR scrub is a credential CONTROL: a
* degraded control (CB-596's sentinel overlay) is still worth having. This config file is not a
* control — it is the ONLY way the member learns where the bridge MCP lives. Writing it
* somewhere the member cannot read would not degrade anything; it would spawn a member that
* occupies a pane and never becomes deliverable, since {@code fleet_send} waits ~60s on the
* readiness gate and then fails with nothing pointing at a temp directory as the cause. Refusing
* up front, with a message that names the missing config key, is the honest failure — an
* undeliverable member is not a working spawn either way, so nothing is lost by refusing loudly
* instead of failing silently later.
*
* @throws IllegalStateException when {@code memberHerdrSocket} is configured but {@code
* worktreeRoot} and/or {@code worktreeGroup} is not
*/
private Path configParentDir() {
if (!memberHerdrSocketConfigured()) {
return configRoot;
}
Path root = memberScrubParentDir();
String group = memberGroup();
if (root == null || group == null) {
throw new IllegalStateException("memberHerdrSocket is configured, so opencode's config "
+ "directory (opencode.json + member charter) must be placed where the member's "
+ "OS user can read it — worktreeRoot, shared via worktreeGroup — but "
+ (root == null ? "worktreeRoot" : "worktreeGroup") + " is not configured. "
+ "Refusing to spawn rather than write a config the member cannot read: that "
+ "member would occupy a pane and never become deliverable, with nothing "
+ "pointing at the real cause. Configure both worktreeRoot and worktreeGroup to "
+ "enable opencode member spawns under memberHerdrSocket.");
}
return root;
}
/**
* Declare a custom OpenAI-compatible provider so the worker talks to a pinned endpoint (a local
* vLLM, say) instead of opencode's default gateway (CB-508).
@@ -614,16 +494,11 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
}
/** One WARN per launcher instance for the fleetd #219 site-2 discovery-unavailable gap. */
private final AtomicBoolean discoveryUnavailableWarned =
new AtomicBoolean();
/** Add lazy on-disk session discovery to the base handle. */
@Override
public PeerHandle spawn(SpawnRequest req) {
PeerHandle inner = super.spawn(req);
return new SessionAwareHandle(inner, discovery, effectiveCwd(req),
this::memberHerdrSocketConfigured, discoveryUnavailableWarned);
return new SessionAwareHandle(inner, discovery, effectiveCwd(req));
}
/**
@@ -638,17 +513,11 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
private final PeerHandle delegate;
private final OpenCodeSessionDiscovery discovery;
private final String cwd;
private final BooleanSupplier discoveryUnavailable;
private final AtomicBoolean discoveryUnavailableWarned;
SessionAwareHandle(PeerHandle delegate, OpenCodeSessionDiscovery discovery, String cwd,
BooleanSupplier discoveryUnavailable,
AtomicBoolean discoveryUnavailableWarned) {
SessionAwareHandle(PeerHandle delegate, OpenCodeSessionDiscovery discovery, String cwd) {
this.delegate = delegate;
this.discovery = discovery;
this.cwd = cwd;
this.discoveryUnavailable = discoveryUnavailable;
this.discoveryUnavailableWarned = discoveryUnavailableWarned;
}
@Override
@@ -673,22 +542,6 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
@Override
public String agentSessionId() {
// fleetd #219 site 2: under memberHerdrSocket the member pane runs as a different OS
// user, so opencode.db lives under THAT user's $HOME, not the one discoveryRoot was
// built from (see OpenCodeLauncher#defaultDiscoveryRoot's javadoc for the full
// reasoning). Scanning fleetd's own $HOME under that config would only ever find "no
// row" and read as "resume unsupported" — declare it unavailable instead, once, loudly.
if (discoveryUnavailable.getAsBoolean()) {
if (discoveryUnavailableWarned.compareAndSet(false, true)) {
log.warn("opencode session discovery unavailable: memberHerdrSocket is "
+ "configured, so opencode's on-disk session database lives under the "
+ "MEMBER's own $HOME, not fleetd's ({}) — agentSessionId will stay null "
+ "for every opencode member under this config, and SESSION_RESUME "
+ "cannot be honored (fleetd #209/#219).",
System.getProperty("user.home"));
}
return null;
}
// Lazy + retried, never a spawn-time blocker: opencode writes the session record only
// when the session is first persisted, so null here is the correct interim answer and
// the caller re-calls later (each call re-scans, picking up a record that has since
@@ -1,87 +1,58 @@
package dev.ltms.fleet.member;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import org.sqlite.SQLiteConfig;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.sql.Connection;
import java.sql.PreparedStatement;
import java.sql.ResultSet;
import java.sql.SQLException;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.stream.Stream;
/**
* Resolves the opencode session id for a fleetd worker from opencode's on-disk storage — the
* only place this adapter touches opencode's private layout, and deliberately the <em>only</em>
* class that does.
*
* <p><strong>Why this is isolated behind one seam.</strong> The layout is version-coupled and not
* a stable contract: opencode persists its session state in a SQLite database at
* {@code <storageRoot>/opencode.db} (a {@code session} table, one row per session, keyed by id and
* carrying a {@code directory} column). That schema can move between opencode releases exactly
* like the JSON-file layout it replaced did (opencode migrated off a one-JSON-file-per-session
* tree under {@code <storageRoot>/storage/session/<projectID>/ses_*.json} in January 2026 — that
* tree is now a frozen migration artefact nothing writes, which is why this class no longer reads
* it). opencode also ships a headless HTTP server that may supersede both of these entirely.
* Everything this adapter knows about that private storage — its shape and column names — lives
* here, so a layout change, or a switch to the HTTP server, changes exactly one class and nothing
* in {@link OpenCodeLauncher}.
* <p><strong>Why this is isolated behind one seam.</strong> The layout is version-coupled and not a
* stable contract: opencode writes one JSON file per session under
* {@code <storageRoot>/session/<projectID>/<ses_*.json>}, and each record carries a
* {@code "version"} field (e.g. {@code "1.1.31"}), so the exact directory shape, file naming, and
* field names can move between opencode releases. opencode also ships a headless HTTP server that
* may supersede file scanning entirely. Everything this adapter knows about that private storage —
* its shape, naming, and field names — lives here, so a layout change, or a switch to the HTTP
* server, changes exactly one class and nothing in {@link OpenCodeLauncher}.
*
* <p>The determinism that makes this useful is structural, not a guess: every fleetd worker runs
* in its own unique git worktree, so the row's {@code directory} (its project root) equals the
* in its own unique git worktree, so the record's {@code directory} (its project root) equals the
* worker's cwd identifies <em>its</em> session unambiguously. We match on {@code directory} rather
* than diffing {@code opencode session list} before/after — that races under concurrent spawns, and
* the CLI listing does not even show the directory.
*
* <p>All reads are best-effort and never throw: a missing or unreadable database, a query that
* fails, or a directory with no row yet all yield {@code null}, and the caller (the session
* handle) treats that as "identity not resolved yet" and retries later. The database is opened
* read-only and never written to: opencode itself may be running and writing it concurrently (WAL
* mode), and this class must never disturb that.
* <p>All reads are best-effort and never throw: a missing or unreadable storage root, a record that
* fails to parse, or a directory with no record yet all yield {@code null}, and the caller (the
* session handle) treats that as "identity not resolved yet" and retries later.
*/
final class OpenCodeSessionDiscovery {
private static final Logger log = LoggerFactory.getLogger(OpenCodeSessionDiscovery.class);
private final Path storageRoot; // e.g. ~/.local/share/opencode (injectable for tests)
private final Path databasePath;
private final AtomicBoolean warnedMissingDatabase = new AtomicBoolean(false);
private final ObjectMapper json;
OpenCodeSessionDiscovery(Path storageRoot) {
this.storageRoot = storageRoot;
this.databasePath = storageRoot.resolve("opencode.db");
this.json = new ObjectMapper();
}
/**
* A connection to {@link #databasePath} opened with SQLite's {@code SQLITE_OPEN_READONLY}
* flag: it never creates the file, never writes, and never touches WAL or journal mode.
* opencode may be running and writing this database concurrently, and this class must never
* disturb it.
*
* <p>Package-private so a test can hold the connection and prove it refuses a write. That is
* the only way to pin this property: making the file unwritable does <em>not</em> work,
* because SQLite silently downgrades a read-write open of an unwritable file to read-only, so
* such a test passes whether or not the flag is set.
*/
Connection openReadOnly() throws SQLException {
SQLiteConfig config = new SQLiteConfig();
config.setReadOnly(true);
return config.createConnection("jdbc:sqlite:" + databasePath);
}
/**
* The opencode session id whose row references {@code directory} (the worker's cwd), or
* {@code null} when no row matches yet. When several rows share the directory — e.g. repeated
* spawns into the same worktree — the row with the highest {@code time_updated} wins: it is
* The opencode session id whose record references {@code directory} (the worker's cwd), or
* {@code null} when no record matches yet. When several records share the directory — e.g.
* repeated spawns into the same worktree — the <em>most recently modified</em> one wins: it is
* the session the pane most likely corresponds to.
*
* <p>Never throws: a missing {@code opencode.db}, a locked/unreadable database, a query
* failure, or a directory that has not been persisted yet all resolve to {@code null} rather
* than failing a spawn. A fleetd worker's session row is written lazily (when the session is
* first persisted), so {@code null} here is the normal answer right after the pane is ready,
* and the caller retries later.
* <p>Never throws: a missing {@code storageRoot}, an unreadable/malformed record, or a
* directory that has not been persisted yet all resolve to {@code null} rather than failing a
* spawn. A fleetd worker's session record is written lazily (when the session is first
* persisted), so {@code null} here is the normal answer right after the pane is ready, and the
* caller retries later.
*
* @param directory the worker's cwd, as resolved for this spawn
* @return the matching session id, or {@code null} if none is known yet
@@ -90,31 +61,62 @@ final class OpenCodeSessionDiscovery {
if (directory == null || directory.isBlank()) {
return null;
}
if (!Files.isRegularFile(databasePath)) {
if (warnedMissingDatabase.compareAndSet(false, true)) {
log.warn("opencode session database not found at {} — opencode's on-disk layout "
+ "may have moved again; session discovery will keep returning null",
databasePath);
}
Path sessionRoot = storageRoot.resolve("session");
if (!Files.isDirectory(sessionRoot)) {
return null;
}
String sql = "SELECT id FROM session WHERE directory = ? ORDER BY time_updated DESC LIMIT 1";
try (Connection connection = openReadOnly();
PreparedStatement statement = connection.prepareStatement(sql)) {
statement.setString(1, directory);
try (ResultSet rows = statement.executeQuery()) {
if (rows.next()) {
return rows.getString("id");
String best = null;
long bestMtime = Long.MIN_VALUE;
try (Stream<Path> projectDirs = Files.list(sessionRoot)) {
for (Path projectDir : projectDirs.filter(Files::isDirectory).toList()) {
try (Stream<Path> records = Files.list(projectDir)) {
for (Path record : records.toList()) {
String id = matchId(record, directory);
if (id == null) {
continue;
}
long mtime = lastModifiedEpochMillis(record);
if (mtime > bestMtime) {
bestMtime = mtime;
best = id;
}
}
} catch (IOException ignored) {
// one project dir unreadable — skip it; another may still match
}
}
} catch (SQLException e) {
// Locked, corrupt, or otherwise unreadable — never fatal to a spawn. Not the
// "database moved" signal (the file exists), so this stays below WARN.
log.debug("opencode session database unreadable at {}: {}", databasePath, e.toString());
} catch (IOException ignored) {
// storage root vanished or became unreadable — "no session known yet"
return null;
}
log.debug("no opencode session row for directory (root={}, directory={})",
storageRoot, directory);
return null;
return best;
}
/**
* The record's session id when it references {@code directory}, else {@code null}. A record
* that is not JSON, lacks {@code id}/{@code directory}, or points at a different directory is
* simply not our session; a malformed one is skipped, never fatal.
*/
private String matchId(Path record, String directory) {
try {
JsonNode node = json.readTree(record.toFile());
JsonNode id = node == null ? null : node.get("id");
JsonNode dir = node == null ? null : node.get("directory");
if (id == null || dir == null || !directory.equals(dir.asText())) {
return null;
}
return id.asText();
} catch (IOException e) {
return null;
}
}
/** The record's last-modified epoch ms, or {@code Long.MIN_VALUE} if unreadable (never wins). */
private static long lastModifiedEpochMillis(Path record) {
try {
return Files.getLastModifiedTime(record).toMillis();
} catch (IOException e) {
return Long.MIN_VALUE;
}
}
}
@@ -29,9 +29,8 @@ import java.util.concurrent.TimeoutException;
*
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target has a durable
* queue {@code agent.<target>.inbox}. The gateway that owns the target starts a manual-ack consumer
* ({@link #own}) that pulls persistent messages, up to its prefetch window, off that queue into an
* in-memory <em>held</em> map (keyed by {@code msgId}) but does <em>not</em> ack them.
* {@link #peek} returns that snapshot;
* ({@link #own}) that pulls persistent messages off that queue into an in-memory <em>held</em> map
* (keyed by {@code msgId}) but does <em>not</em> ack them. {@link #peek} returns that snapshot;
* {@link #ack} acks the broker delivery-tag and drops the entry. Because messages stay unacked until
* the owning gateway actually drains them, a crash (or a {@code java -jar} bounce) before caller-ack
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
@@ -2,14 +2,12 @@ package dev.ltms.fleet.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.List;
import java.util.UUID;
import java.util.concurrent.CompletableFuture;
@@ -168,31 +166,15 @@ public final class MessageService {
private static final class Task {
private final String ticket;
private final String target;
/**
* When this task was created (#137 fix): the tiebreaker for which of several open tasks on
* one target gets a recovered reply in {@link #abandon} — the oldest, since it is the one
* that has been waiting longest.
*/
private final long createdNanos;
private final CompletableFuture<Reply> future = new CompletableFuture<>();
/**
* When {@link #future} resolved, or {@code null} while it is still pending — the clock
* {@link #pruneTerminalTickets} measures the TTL from (#197). Deliberately a boxed
* {@code Long} rather than a {@code long} with a sentinel: {@link System#nanoTime} may
* legitimately return any value, zero and negatives included, so no numeric sentinel can mean
* "not stamped yet". Stamped by a {@code whenComplete} hook registered in the constructor, so
* every completion path stamps it — a reply, the completion fallback, a timeout, a failure,
* or an abandon on teardown — without each of those having to remember to.
*/
private volatile Long completedNanos;
private final long createdNanos;
private volatile Reply question;
private volatile String turnId;
private Task(String ticket, String target, LongSupplier nowNanos) {
private Task(String ticket, String target, long createdNanos) {
this.ticket = ticket;
this.target = target;
this.createdNanos = nowNanos.getAsLong();
future.whenComplete((reply, ex) -> completedNanos = nowNanos.getAsLong());
this.createdNanos = createdNanos;
}
}
@@ -207,7 +189,6 @@ public final class MessageService {
}
private final AgentControl agents;
private final HerdrRouter router;
private final Injector injector;
private final Rendezvous rendezvous;
private final ReplyInbox inbox;
@@ -276,7 +257,6 @@ public final class MessageService {
MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox,
ReplyPushLoop pushLoop, Metrics metrics, LongSupplier nowNanos) {
this.agents = agents;
this.router = null;
this.injector = injector;
this.rendezvous = rendezvous;
this.inbox = inbox;
@@ -285,20 +265,6 @@ public final class MessageService {
this.nowNanos = nowNanos;
}
public MessageService(HerdrRouter router, Injector injector, Rendezvous rendezvous, ReplyInbox inbox,
ReplyPushLoop pushLoop, Metrics metrics) {
this.agents = null;
this.router = router;
this.injector = injector;
this.rendezvous = rendezvous;
this.inbox = inbox;
this.pushLoop = pushLoop;
this.metrics = metrics;
this.nowNanos = System::nanoTime;
}
private AgentControl agentsFor(String target) { return router != null ? router.agentsFor(target) : agents; }
/** Create with an explicit {@link ReplyInbox} and no push loop. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox) {
this(agents, injector, rendezvous, inbox, null);
@@ -311,7 +277,7 @@ public final class MessageService {
/** Current lifecycle status of a worker (the {@code GET /sessions/{id}/status} surface). */
public AgentStatus status(String target) {
return agentsFor(target).status(target);
return agents.status(target);
}
/** Read-only delegation fact for fleet views. */
@@ -374,61 +340,21 @@ public final class MessageService {
}
/**
* Route a worker's explicit {@code fleet_reply}: resolve an open send, complete an async ticket
* still parked waiting on this exact turn's answer, or — only once neither applies — queue it in
* the inbox. Unlike the bare {@link Rendezvous#resolve}, a no-waiter result is <em>not</em> a
* failure — the reply is held for later drain.
* Route a worker's explicit {@code fleet_reply}: resolve an open send, or queue it in the
* inbox if no send is currently open. Unlike the bare {@link Rendezvous#resolve}, a no-waiter
* result is <em>not</em> a failure — the reply is held for later drain.
*
* <p><strong>Do NOT use this for mid-turn questions.</strong> {@code fleet_ask} /
* {@link Rendezvous#resolveQuestion} must keep today's {@code NO_WAITER} behaviour — questions
* are interactive and must never be queued.
*
* <p><strong>Ambiguous match also falls to the inbox.</strong> {@link #askAnsweredAsyncTasks}
* cannot actually return more than one entry today (see its own javadoc for why — in short,
* {@link #hasAsyncQuestion} keeps a target BUSY, so no second task can reach this state, for as
* long as an earlier one's {@code turnId} is still stamped). That is an emergent guarantee from
* two other facts, not one this method enforces, so this branch stays in as defence in depth
* rather than being removed as dead code: if it ever weakens, returning whichever candidate a
* {@code ConcurrentHashMap} iteration reaches first would let a genuine reply complete the
* <em>wrong</em> ticket — silently handing the lead something that reads like a correct answer to
* a delegation the worker never touched, which is worse than a failure because the lead acts on
* it. When more than one candidate exists, guessing is not safe: fall back to the inbox exactly
* as the zero-candidate case does, and let {@link #abandon} apply the eventual recovery
* deterministically instead.
*
* @return always {@code true} — the reply resolved a live send, completed a parked ticket, or
* was queued
* @return always {@code true} — the reply either resolved a live send or was queued
*/
public boolean reply(String session, String content) {
if (rendezvous.resolve(session, content)) {
count(FleetMetrics.REPLIES, "path", "rendezvous");
return true; // a live send took it — unchanged fast path
}
// #137: no live rendezvous waiter, but this may be the worker's real fleet_reply resuming a
// turn that {@link #answer} already gave up waiting on. answer()'s own bounded wait (the
// primary's fleet_send{turnId} call, capped well under a minute) can time out and close its
// waiter long before the worker — now actually resuming real work — finishes and replies. That
// reply used to have nowhere to land but the session inbox, leaving the async ticket's future
// unresolved forever: fleet_poll{ticket} stayed PENDING until fleet_stop's abandon() forced it
// FAILED with a misleading "session released before it replied" reason, even though the reply
// had, in fact, arrived. Completing the matching ticket directly here means fleet_poll{ticket}
// sees the real reply instead.
List<Task> candidates = askAnsweredAsyncTasks(session);
if (candidates.size() == 1) {
Task orphan = candidates.get(0);
if (orphan.future.complete(new Reply(Outcome.REPLIED, content))) {
if (orphan.turnId != null) {
asyncTasksByTurn.remove(orphan.turnId, orphan);
}
count(FleetMetrics.REPLIES, "path", "async-recovered");
return true; // the ticket itself took it — no inbox stranding at all
}
} else if (candidates.size() > 1) {
List<String> tickets = candidates.stream().map(t -> t.ticket).toList();
log.warn("reply from {} matches {} open async tickets {} — cannot tell which one it "
+ "answers, queuing to the inbox instead of guessing", session, candidates.size(),
tickets);
}
inbox.publish(session, UUID.randomUUID().toString(), content);
// CB-640: record the stranding itself (not just the reply text) so fleet health can see a
// worker whose replies keep missing their waiter, not only the queue depth this leaves behind.
@@ -442,41 +368,6 @@ public final class MessageService {
return true; // held, not lost
}
/**
* Every still-open async task on {@code target} whose {@code fleet_ask} was already answered —
* its {@link Task#turnId} is stamped but its {@link Task#question} was cleared by {@link #answer}
* — yet whose future is not resolved yet (#137). Empty if no such task exists, including the
* common case where {@code target}'s worker never used {@code fleet_ask} at all (a task that was
* never asked has {@code turnId == null}, so it can never match here and only ever completes
* through the ordinary rendezvous fast path in {@link #reply}).
*
* <p><strong>Returns at most one entry today — verified, not assumed.</strong> {@link #send}
* refuses to open a waiter on {@code target} while {@link #hasAsyncQuestion} is true, and that
* check matches ANY task whose {@code turnId} is still stamped in {@code asyncTasksByTurn} —
* not only while its question is still open. {@link #answer} deliberately leaves that stamp in
* place ({@code clearAsyncQuestion(turnId, false)}) until the resumed turn's own future actually
* resolves, at which point {@link #finishAsyncTask} both removes the stamp AND completes that
* task's future in the same call. So a second task can never reach "{@code turnId} stamped, future
* still open" — the exact pair this method matches on — while a first one already holds it: by
* the time the stamp is gone, so is the eligibility. This is an emergent property of those two
* facts holding together, not something this method (or its callers) enforces on its own — flip
* {@code forgetTurn} to {@code true} in that one {@link #answer} call and it silently stops being
* true, with nothing left to fail loudly. The callers below still handle "more than one" as
* defence in depth against exactly that, not because they exercise it today: {@link #reply}
* treats it as unresolvable and falls back to the inbox; {@link #abandon} would pick the oldest
* deterministically (its own {@code matching} list has no such guarantee — see its javadoc).
*/
private List<Task> askAnsweredAsyncTasks(String target) {
List<Task> candidates = new ArrayList<>();
for (Task task : tasks.values()) {
if (target.equals(task.target) && task.question == null && task.turnId != null
&& !task.future.isDone()) {
candidates.add(task);
}
}
return candidates;
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
if (metrics != null) {
@@ -518,108 +409,19 @@ public final class MessageService {
* <p>Resolving the waiter as a failure — rather than letting it time out — also means the
* outcome is counted, so a torn-down delegation stops being invisible to {@code /metrics}.
*
* <p><strong>#137 defence in depth.</strong> {@link #reply} already hands a worker's real
* {@code fleet_reply} straight to the async ticket it belongs to whenever exactly one is still
* parked waiting for it (see {@link #askAnsweredAsyncTasks}), so by the time a session is
* released its tasks are normally already resolved — this loop's {@code complete} calls are then
* harmless no-ops (a {@link CompletableFuture} can only resolve once). But should some other path
* someday strand a reply in the inbox without completing its ticket, checking
* {@link #hasStrandedReply(String)} here — before ever writing a failure — means a torn-down
* session whose worker in fact replied is still reported {@code REPLIED} with that reply's own
* text, never the misleading "the worker session was released before it replied" (which also
* means the snapshot/worktree recovery hint that follows it never prints once a reply exists).
*
* <p><strong>At most one task gets the recovered reply — and here, unlike {@link #reply}'s
* {@link #askAnsweredAsyncTasks}, {@code matching.size() >= 2} alone is reachable today.</strong>
* This method's {@code matching} filter has no {@code turnId != null} requirement, so it matches
* any plain (never-asked) open task too — and {@link #sendAsync} does not limit a target to one
* of those: a second {@code fleet_send{wait:false}} at a target that is still busy returns its own
* ticket immediately and simply parks its {@link #send} behind the target's session lock for up
* to {@link #ASYNC_TIMEOUT_MS}, exactly as {@code abandonFailsEveryPendingAsyncTicketForTheReleasedTarget}
* already proves. Before this fix, the loop below drained the strand once and then reused that
* same {@code Reply} for <em>every</em> task it walked past — so two open tasks really did both
* complete {@code REPLIED} with the same text (see the pre-fix loop in commit 97f6c33's parent).
* A stranded reply is one worker answer, so it can settle at most one open task on this target —
* never every open task, and never a guess. When more than one task is still open here, the
* recovered reply goes to the <em>oldest</em> (lowest {@link Task#createdNanos}) — it has been
* waiting longest, so it is the one most likely to be what the reply actually answers. Every
* other open task keeps the ordinary {@code WORKER_FAILED} path it would take without a stranded
* reply at all.
*
* <p><strong>{@code matching.size() >= 2} together with {@code hadStrandedReply} is a different
* question, and today it is defence in depth rather than a path this codebase's public API can
* drive.</strong> This class has exactly two sites that ever acquire a target's entry in
* {@code sessionLocks} — {@link #send} and {@link #answer} — and both open a {@link Rendezvous}
* waiter for that same target as the very first thing they do after acquiring the lock, then hold
* lock and waiter together for the rest of their critical section ({@link #send} also clears
* {@link #strandedReplies} right there, the instant it opens its waiter — before it ever enqueues
* delivery). So "the session lock is held" and "a live waiter is open for it" are the same fact
* throughout this class, and {@link #reply}'s fast path always resolves a currently-open waiter
* directly rather than stranding. The two facts this method wants therefore cannot be produced
* side by side: while the lock is held, a real reply resolves the open waiter directly and never
* reaches {@link #strandedReplies}; the instant the lock is free, any parked matching task's own
* {@link #send} that is scheduled next wins it and, by opening its waiter, clears the strand again
* before this method ever runs. There is no way to hold that lock open-but-unaccepted from outside
* {@link #send}/{@link #answer} to freeze a window in between. Constructing both facts at once
* through {@code sendAsync}/{@code reply}/{@code ask}/{@code answer} would need a race against
* virtual-thread scheduling, not a deterministic sequence — so the oldest-wins code below stays as
* defence in depth against a regression to that mechanism (e.g. clearing {@link #strandedReplies}
* on a narrower condition than "any acceptance"), not because today's test suite exercises the
* conjunction. {@code matching.size() >= 2} alone, without a strand, is exactly what
* {@code abandonFailsEveryPendingAsyncTicketForTheReleasedTarget} already covers.
*
* <p><strong>Do not "fix" that gap with a test that reaches past this class.</strong> A test can
* build both facts by calling {@link Rendezvous#close} itself on the waiter an accepted
* {@link #send} is still blocked on: the send keeps the lock, no waiter is registered any more, a
* second parked task stays open, and the next {@link #reply} then strands. That was checked, and
* such a test does go red against the pre-fix loop. But it only goes red because it broke the
* lock-and-waiter invariant above from outside — no caller of this class ever does that — so it
* pins a state production cannot reach, and would read to the next person as if it could.
*
* @return true if a live waiter or an async task was failed (never true for one recovered as a
* reply — see the note above)
* @return true if a live waiter was failed
*/
public boolean abandon(String target, String reason) {
boolean hadStrandedReply = hasStrandedReply(target);
// CB-640: the session is gone — nothing will ever accept or deliver into it now.
strandedReplies.remove(target);
queuedDeliveries.remove(target);
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
boolean failed = waiter != null && !waiter.isDone() && rendezvous.resolveFailure(waiter, reason);
boolean asyncFailed = false;
List<Task> matching = new ArrayList<>();
for (Task task : tasks.values()) {
if (target.equals(task.target) && task.question == null && !task.future.isDone()) {
matching.add(task);
}
}
Task recoveryTask = null;
if (hadStrandedReply && !matching.isEmpty()) {
recoveryTask = matching.get(0);
for (Task candidate : matching) {
if (candidate.createdNanos < recoveryTask.createdNanos) {
recoveryTask = candidate;
}
}
}
Reply recovered = recoveryTask != null ? recoverStrandedReply(target) : null;
for (Task task : matching) {
boolean isRecovery = task == recoveryTask && recovered != null;
Reply outcome = isRecovery ? recovered : new Reply(Outcome.WORKER_FAILED, reason);
if (task.future.complete(outcome)) {
if (outcome.outcome() == Outcome.WORKER_FAILED) {
asyncFailed = true;
} else if (task.turnId != null) {
asyncTasksByTurn.remove(task.turnId, task);
}
} else if (isRecovery) {
// The recovered reply was already drained out of the inbox, but this task resolved
// through another path (e.g. a concurrent reply() or a second abandon() racing this
// one) between us choosing it and completing it here. Put the reply back rather than
// lose it silently — it may still belong to some other still-open task, or the next
// caller that drains this target's inbox.
inbox.publish(target, UUID.randomUUID().toString(), recovered.text());
if (target.equals(task.target) && task.question == null
&& task.future.complete(new Reply(Outcome.WORKER_FAILED, reason))) {
asyncFailed = true;
}
}
if (failed) {
@@ -628,25 +430,6 @@ public final class MessageService {
return failed || asyncFailed;
}
/**
* Drain {@code target}'s inbox and hand its content back as a {@link Outcome#REPLIED} result
* (#137 defence in depth for {@link #abandon}) — {@code null} if it turned out empty (the
* stranding fact raced away, e.g. a lead's own {@code fleet_poll} on the raw session already
* drained it first). When more than one message is queued, only the newest is the worker's actual
* final answer ({@link #drainReplies} returns them oldest-first).
*
* <p>This does drain (removes the messages from the inbox) before the caller knows whether the
* task it is recovering for will actually accept them — {@link #abandon} is the one that puts a
* reply back if its {@code complete} call turns out to lose the race.
*/
private Reply recoverStrandedReply(String target) {
var messages = drainReplies(target);
if (messages.isEmpty()) {
return null;
}
return new Reply(Outcome.REPLIED, messages.get(messages.size() - 1).content());
}
/**
* Acknowledge a specific reply by {@code msgId} for {@code target}. Removes it from the inbox
* so that a subsequent drain or peek no longer returns it.
@@ -921,7 +704,7 @@ public final class MessageService {
*/
public String sendAsync(String target, String content, Runnable onAccepted) {
String ticket = "task-" + ticketSeq.incrementAndGet();
Task task = new Task(ticket, target, nowNanos);
Task task = new Task(ticket, target, nowNanos.getAsLong());
tasks.put(ticket, task);
if (pushLoop != null) {
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
@@ -1005,7 +788,7 @@ public final class MessageService {
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
private String liveStatus(String target) {
try {
return agentsFor(target).status(target).name().toLowerCase();
return agents.status(target).name().toLowerCase();
} catch (RuntimeException e) {
return "unknown";
}
@@ -1021,25 +804,12 @@ public final class MessageService {
* (or one the reminder cap already gave up on) is pruned here but never collected there, so it
* lingers in {@code pendingTickets} forever and rides along on every later nudge to the same lead
* — naming a ticket {@code fleet_poll} can no longer find (CB-588 follow-up).
*
* <p>The TTL runs from **completion**, not from creation (#197). It used to compare against
* {@code createdNanos}, which made the real collection window {@code TTL minus however long the
* task ran}: a delegation that took longer than the TTL had its reply destroyed on the first
* sweep after it landed, every time. That is the normal case here — real work runs well past ten
* minutes — and the reply lives only in {@code future}, so pruning it discards the worker's whole
* report with nothing to fall back on. Measuring from completion gives every ticket the same full
* window whatever its runtime, and still bounds {@code tasks}.
*
* <p>A task whose future is done but whose {@code completedNanos} is not stamped yet is left
* alone. That window is the few instructions between {@code complete()} and the constructor's
* {@code whenComplete} hook running; the next sweep collects it.
*/
private void pruneTerminalTickets() {
long cutoff = nowNanos.getAsLong() - TICKET_TTL_NANOS;
tasks.entrySet().removeIf(e -> {
Task t = e.getValue();
Long completed = t.completedNanos;
boolean expired = t.future.isDone() && completed != null && completed < cutoff;
boolean expired = t.future.isDone() && t.createdNanos < cutoff;
if (expired && pushLoop != null) {
pushLoop.ticketCollected(e.getKey());
}
@@ -31,7 +31,6 @@ import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.stream.Collectors;
/**
@@ -55,12 +54,11 @@ public final class FleetApp {
/** Context attribute under which the resolved caller is stashed by the auth filter. */
private static final String CALLER = "fleetd.caller";
private final HerdrClient herdr; // lead daemon
private final HerdrClient memberHerdr; // CB-185: member daemon (same object when unconfigured)
private final HerdrClient herdr;
private final PeerLauncher workers;
private final SessionManager sessions; // CB-301: authoritative session registry
private final MessageService messages;
private final Predicate<String> deliverable;
private final MemberPresence presence; // CB-113: which workers are MCP-connected (available)
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
private final CallerResolver auth; // CB-501: null → authz not enforced (legacy behaviour)
private final Metrics metrics; // CB-502: null → /metrics not exposed
@@ -72,9 +70,9 @@ public final class FleetApp {
* behaviour without each needing an auth fixture.
*/
public FleetApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet) {
this(herdr, workers, sessions, messages, presence, mcpServlet, null, null, presence::isPresent);
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet) {
this(herdr, workers, sessions, messages, presence, mcpServlet, null, null);
}
/**
@@ -84,39 +82,13 @@ public final class FleetApp {
* the endpoint
*/
public FleetApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics) {
this(herdr, workers, sessions, messages, presence, mcpServlet, auth, metrics, presence::isPresent);
}
/**
* @param deliverable the injector's readiness gate, shared so status reports its real result
*/
public FleetApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
Predicate<String> deliverable) {
this(herdr, herdr, workers, sessions, messages, presence, mcpServlet, auth, metrics, deliverable);
}
/**
* @param herdr the lead daemon's client
* @param memberHerdr the member daemon's client (CB-185); pass the same instance as
* {@code herdr} for a single-daemon deployment — {@code healthz}/{@code
* sessions} then make exactly one herdr call each, unchanged from before
* the two-daemon router existed
* @param deliverable the injector's readiness gate, shared so status reports its real result
*/
public FleetApp(HerdrClient herdr, HerdrClient memberHerdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
Predicate<String> deliverable) {
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics) {
this.herdr = herdr;
this.memberHerdr = memberHerdr != null ? memberHerdr : herdr;
this.workers = workers;
this.sessions = sessions;
this.messages = messages;
this.deliverable = deliverable;
this.presence = presence;
this.mcpServlet = mcpServlet;
this.auth = auth;
this.metrics = metrics;
@@ -207,85 +179,30 @@ public final class FleetApp {
ctx.status(200).contentType("text/plain; version=0.0.4; charset=utf-8").result(metrics.render());
}
/**
* Liveness + herdr reachability. 200 only when BOTH daemons answer ping — 503 otherwise
* (CB-185). With no {@code memberHerdrSocket} configured {@code memberHerdr == herdr}, so this
* makes exactly the one {@code ping} call it always did and reports the same body; with a
* second daemon configured, a member daemon that is down must not be masked by a healthy lead
* daemon — every spawn goes through the member daemon and would otherwise fail silently behind
* a green {@code /healthz}.
*
* <p>CB-185 blocker 2: the {@code herdr} key always carries the <em>lead</em> daemon's
* version/protocol, unchanged, because two consumers — {@code scripts/redeploy-fleetd.sh} and
* {@code scripts/rename-checkout.sh} — read this endpoint already (both only check the HTTP
* status code and print the body verbatim; neither parses a specific field, so adding a key
* alongside {@code herdr} is safe). But it is the <em>member</em> daemon's protocol that decides
* whether a spawn works, so when a second daemon is configured its version/protocol is reported
* too, under a separate {@code member} key — never folded into {@code herdr}, which would make a
* mismatch invisible to whichever consumer only reads that key. If the two protocol numbers
* differ, {@code protocolMismatch: true} calls it out explicitly rather than leaving it to be
* spotted by comparing two numbers by eye.
*/
/** Liveness + herdr reachability. 200 when herdr answers ping, 503 otherwise. */
private void healthz(Context ctx) {
JsonNode pong;
try {
pong = herdr.call("ping");
JsonNode pong = herdr.call("ping");
ctx.status(200).json(Map.of(
"status", "ok",
"herdr", Map.of(
"version", pong.path("version").asText(""),
"protocol", pong.path("protocol").asInt())));
} catch (HerdrException e) {
ctx.status(503).json(Map.of(
"status", "degraded",
"herdr", "unreachable",
"detail", e.getMessage()));
return;
}
Map<String, Object> body = new LinkedHashMap<>();
body.put("status", "ok");
body.put("herdr", Map.of(
"version", pong.path("version").asText(""),
"protocol", pong.path("protocol").asInt()));
if (memberHerdr != herdr) {
JsonNode memberPong;
try {
memberPong = memberHerdr.call("ping");
} catch (HerdrException e) {
ctx.status(503).json(Map.of(
"status", "degraded",
"herdr", "member unreachable",
"detail", e.getMessage()));
return;
}
int leadProtocol = pong.path("protocol").asInt();
int memberProtocol = memberPong.path("protocol").asInt();
body.put("member", Map.of(
"version", memberPong.path("version").asText(""),
"protocol", memberProtocol));
if (leadProtocol != memberProtocol) {
body.put("protocolMismatch", true);
}
}
ctx.status(200).json(body);
}
/**
* Sessions view derived from herdr {@code workspace.list} (one workspace → one row), merged
* across both daemons (CB-185). With no {@code memberHerdrSocket} configured {@code
* memberHerdr == herdr}, so this calls {@code workspace.list} exactly once, same as before the
* router existed; with a second daemon configured, calling it twice would silently drop every
* member workspace (they live on the member daemon only).
*/
/** Sessions view derived from herdr {@code workspace.list} (one workspace → one row). */
private void sessions(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
JsonNode result = herdr.call("workspace.list");
List<Map<String, Object>> out = new ArrayList<>();
collectSessions(herdr, out);
if (memberHerdr != herdr) {
collectSessions(memberHerdr, out);
}
ctx.status(200).json(Map.of("sessions", out));
}
private static void collectSessions(HerdrClient client, List<Map<String, Object>> out) {
JsonNode result = client.call("workspace.list");
for (JsonNode w : result.path("workspaces")) {
out.add(Map.of(
"id", w.path("workspace_id").asText(""),
@@ -294,6 +211,7 @@ public final class FleetApp {
"paneCount", w.path("pane_count").asInt(),
"agentStatus", w.path("agent_status").asText("unknown")));
}
ctx.status(200).json(Map.of("sessions", out));
}
/** Discovery: every agent herdr tracks, keyed by its Claude session UUID. */
@@ -315,18 +233,10 @@ public final class FleetApp {
.map(Agent.class::cast)
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
// fleetd #209: this REST roster reports agentSessionId via SessionManager.rosterView, so it
// uses the resolving roster read (caller-driven, not a timer) rather than the plain one.
List<Map<String, Object>> out = sessions.rosterResolved().stream()
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
Map<String, Object> body = new LinkedHashMap<>();
// fleetd #199: the endpoint became /members in the CB-634 rename but the body key stayed
// "workers", so a caller that read "members" saw an empty fleet and reported no members at
// all. "members" is the canonical key; "workers" stays as a deprecated alias so an existing
// REST consumer keeps working — the out-of-band path a lead falls back to when its MCP mount
// drops reads this endpoint. Drop the alias once nothing reads it.
body.put("members", out);
body.put("workers", out);
// CB-586: operator visibility for the refs/wip snapshot store without shelling into the
// repo — how many snapshot refs exist and roughly what they cost. Present only once a
@@ -597,9 +507,9 @@ public final class FleetApp {
/**
* Live lifecycle status of a worker (MCP `fleet_status` wraps this in CB-105), plus its
* <em>readiness</em>: {@code ready} is true when the injector can deliver to the target. A
* spawned member must connect the bridge MCP first, while a known lead is ready without member
* presence. This differs from bare {@code idle}, which is also true during member boot.
* <em>readiness</em> (CB-113): {@code ready} is true once the worker's Claude has connected the
* bridge MCP — the reliable "available to receive a task" signal, unlike bare {@code idle}, which
* is also true during boot.
*/
private void sessionStatus(Context ctx) {
String id = ctx.pathParam("id");
@@ -610,7 +520,7 @@ public final class FleetApp {
Map<String, Object> body = new LinkedHashMap<>();
body.put("sessionId", id);
body.put("status", messages.status(id).name().toLowerCase());
body.put("ready", deliverable.test(id));
body.put("ready", presence.isPresent(id));
// CB-582: a worker paused mid-turn in an async fleet_ask is otherwise invisible to a
// status poll — surface the open question and how to answer it, same as fleet_poll's
// Phase.ASKING view.
@@ -7,8 +7,6 @@ import java.io.BufferedReader;
import java.io.IOException;
import java.io.InputStreamReader;
import java.io.UncheckedIOException;
import java.net.URI;
import java.net.URISyntaxException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
@@ -22,8 +20,6 @@ import java.util.Optional;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Consumer;
import java.util.function.Function;
import java.util.stream.Collectors;
/**
@@ -68,10 +64,6 @@ public final class GitWorktrees implements Worktrees {
/** What {@link #isolateToolSurface} writes for {@code .autoenv}: a valid, empty env file. */
private static final String NEUTRAL_AUTOENV_CONFIG = "";
/** A credential helper command which reads only an environment variable at Git call time. */
private static final String ENVIRONMENT_CREDENTIAL_HELPER = "!f() { if [ \"$1\" = get ]; then "
+ "printf 'username=%s\\npassword=%s\\n\\n' git \"$WORKER_GITEA_TOKEN\"; fi; }; f";
/**
* A tracked project config that is hostile in a provisioned worktree, and what to replace it
* with. {@link #file} is the repo-relative path; {@link #stub} is a neutral but VALID payload for
@@ -89,69 +81,21 @@ public final class GitWorktrees implements Worktrees {
);
private final String configuredRoot;
/** OS group name for {@link #shareWithGroup} (fleetd #185 stage 3); {@code null} ⇒ feature off. */
private final String group;
private final Consumer<String> afterWorktreeAdded;
/** How {@link #shareWithGroup}'s processes (git config / chgrp / chmod / find) actually run.
* Defaults to the real {@link #exec(String...)}. Package-private test seam so a unit test can
* prove "no group configured ⇒ zero processes spawned" and inspect exactly what a configured
* group runs, without a real second OS user or OS group on this host. */
private final Function<String[], String> shareGroupRunner;
private final SecureRandom random = new SecureRandom();
private final AtomicLong seq = new AtomicLong();
/** Default constructor: worktree root is derived per-repo as {@code <repoRoot>/../.bridged-worktrees}. */
public GitWorktrees() {
this(null, (String) null);
this(null);
}
/**
* @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of
* the repo root. No {@code worktreeGroup} configured — {@link #shareWithGroup}
* is a no-op.
*/
/** @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of the repo root. */
public GitWorktrees(String configuredRoot) {
this(configuredRoot, (String) null);
}
/**
* @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of
* the repo root.
* @param group optional OS group name (fleetd #185 stage 3, {@code worktreeGroup:} in
* config); null/blank ⇒ {@link #shareWithGroup} is a no-op.
*/
public GitWorktrees(String configuredRoot, String group) {
this(configuredRoot, group, _ -> {});
}
/** Test seam for changing a real worktree between its creation and its security check. */
GitWorktrees(String configuredRoot, Consumer<String> afterWorktreeAdded) {
this(configuredRoot, null, afterWorktreeAdded);
}
/** Test seam combining a configurable {@code group} with {@link #afterWorktreeAdded}. */
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded) {
this(configuredRoot, group, afterWorktreeAdded, null);
}
/**
* Full test seam: also overrides how {@link #shareWithGroup}'s processes run (fleetd #185
* stage 3), so a unit test can prove "no group configured ⇒ no process spawned" and inspect
* exactly what commands a configured group runs, without a real second OS user/group.
*
* @param shareGroupRunner {@code null} ⇒ the real {@link #exec(String...)}.
*/
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded,
Function<String[], String> shareGroupRunner) {
this.configuredRoot = configuredRoot;
this.group = (group == null || group.isBlank()) ? null : group;
this.afterWorktreeAdded = afterWorktreeAdded == null ? _ -> {} : afterWorktreeAdded;
this.shareGroupRunner = shareGroupRunner != null ? shareGroupRunner : this::exec;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
reportRemoteUrlsWithUserInfo(repoRoot);
String base = (baseRef == null || baseRef.isBlank()) ? "HEAD" : baseRef;
String nonce = nonce();
Path root = resolveRoot(repoRoot);
@@ -163,244 +107,11 @@ public final class GitWorktrees implements Worktrees {
}
String wt = path.toAbsolutePath().toString();
log.info("adding worktree branch={} path={} base={}", branch, wt, base);
removeUserInfoFromHttpsOrigin(repoRoot);
exec("git", "-C", repoRoot, "worktree", "add", wt, "-b", branch, base);
afterWorktreeAdded.accept(wt);
requireCredentialFreeHttpsOrigin(wt);
configureEnvironmentCredentialHelper(repoRoot, wt);
configureHttpsUrlRewriteForSshOrigin(repoRoot, wt);
isolateToolSurface(wt);
return wt;
}
/**
* A linked worktree shares its primary checkout's git config. Remove HTTPS user info before
* adding one, so a credential accidentally embedded in that config cannot reach the member.
*/
private void removeUserInfoFromHttpsOrigin(String repoRoot) {
if (exitCode("git", "-C", repoRoot, "config", "--get", "remote.origin.url") != 0) {
return;
}
String origin = execRedacted("git", "-C", repoRoot, "config", "--get", "remote.origin.url").trim();
URI uri;
try {
uri = new URI(origin);
} catch (URISyntaxException e) {
throw new WorktreeException("origin URL is invalid; cannot provision a safe worktree", e);
}
if (!"https".equalsIgnoreCase(uri.getScheme()) || uri.getUserInfo() == null) {
return;
}
int schemeEnd = origin.indexOf("://") + 3;
int userInfoEnd = origin.indexOf('@', schemeEnd);
if (userInfoEnd < schemeEnd) {
throw new WorktreeException("origin URL has invalid HTTPS user info; cannot provision safely");
}
String cleanOrigin = origin.substring(0, schemeEnd) + origin.substring(userInfoEnd + 1);
// Plain exec, not execRedacted, is correct here: this call WRITES cleanOrigin (already
// stripped of user-info above) rather than reading a URL back from stdout, so there is
// nothing secret left in either its argv or its stdout to redact.
exec("git", "-C", repoRoot, "remote", "set-url", "origin", cleanOrigin);
log.info("removed HTTPS user info from forge origin before provisioning worktree");
}
/** Refuse the worktree if Git still resolves any HTTPS origin URL with embedded credentials. */
private void requireCredentialFreeHttpsOrigin(String worktreePath) {
if (exitCode("git", "-C", worktreePath, "remote", "get-url", "--all", "origin") != 0) {
return;
}
String origins = execRedacted("git", "-C", worktreePath, "remote", "get-url", "--all", "origin");
for (String origin : origins.split("\\R")) {
try {
URI uri = new URI(origin);
if ("https".equalsIgnoreCase(uri.getScheme()) && uri.getUserInfo() != null) {
throw new WorktreeException("worktree origin contains HTTPS user info; refusing provision");
}
} catch (URISyntaxException e) {
throw new WorktreeException("worktree origin URL is invalid; refusing provision", e);
}
}
}
/**
* Report — never refuse — every remote whose fetch or push URL carries user-info outside SSH.
* A linked worktree shares its parent repository's git config, so a credential on ANY remote
* (not only {@code origin}) or in a {@code pushurl} is just as readable by a member as one on
* {@code origin}'s HTTPS fetch URL — the one case {@link #removeUserInfoFromHttpsOrigin} and
* {@link #requireCredentialFreeHttpsOrigin} already strip and refuse. This check is additive: it
* only logs a warning, it never mutates config and never refuses the provision.
*
* <p>A reporting-only check must never be able to abort a provision — PR #173 shipped one that
* ran unguarded at the top of {@link #add}, and every git call inside it can throw ({@link #exec}
* turns a non-zero exit or its 30-second timeout into a {@link WorktreeException}). Every git call
* here is therefore wrapped, and on failure only the exception's <em>class</em> is logged, never
* its message: the enumerating {@code git remote} call is not redacted, and its stderr is read
* from the very config that may hold the URL this check exists to find.
*/
private void reportRemoteUrlsWithUserInfo(String repoRoot) {
List<String> remotes;
try {
remotes = exec("git", "-C", repoRoot, "remote").lines()
.map(String::trim)
.filter(r -> !r.isBlank())
.toList();
} catch (RuntimeException e) {
log.warn("could not enumerate remotes to check for credentialed URLs in {}: {}",
Path.of(repoRoot).toAbsolutePath().normalize(), e.getClass().getName());
return;
}
for (String remote : remotes) {
reportOneRemoteUrlsWithUserInfo(repoRoot, remote);
}
}
private void reportOneRemoteUrlsWithUserInfo(String repoRoot, String remote) {
boolean leaks = remoteUrlsLeakUserInfo(repoRoot, remote, false)
|| remoteUrlsLeakUserInfo(repoRoot, remote, true);
if (leaks) {
log.warn("member worktree shares a remote URL containing user-info: remote={} repository={}; "
+ "remove credentials from the repository's git config",
remote, Path.of(repoRoot).toAbsolutePath().normalize());
}
}
/**
* True when any resolved fetch (or, if {@code push}, push) URL for {@code remote} carries
* non-empty user-info outside SSH. Never throws — a git failure here is caught, logged (its
* class only, per the javadoc above), and treated as "nothing found", so it cannot abort or
* otherwise affect provisioning. Uses {@link #execRedacted} because the command's stdout is
* itself the URL this check exists to find.
*/
private boolean remoteUrlsLeakUserInfo(String repoRoot, String remote, boolean push) {
try {
String out = push
? execRedacted("git", "-C", repoRoot, "remote", "get-url", "--push", "--all", remote)
: execRedacted("git", "-C", repoRoot, "remote", "get-url", "--all", remote);
return out.lines().anyMatch(url -> !url.isBlank() && urlLeaksUserInfo(url.trim()));
} catch (RuntimeException e) {
log.warn("could not read the {} URL for remote {} to check for credentials: {}",
push ? "push" : "fetch", remote, e.getClass().getName());
return false;
}
}
/**
* True when {@code rawUrl} parses as an absolute URI with a non-SSH-family scheme and non-empty
* user-info. An unparsable or scheme-less URL — including the ssh scp-like shorthand
* ({@code user@host:path}) — is not this check's concern and is treated as "no finding", the
* same way {@link #configureHttpsUrlRewriteForSshOrigin} leaves that shorthand untouched.
*/
private static boolean urlLeaksUserInfo(String rawUrl) {
URI uri;
try {
uri = new URI(rawUrl);
} catch (URISyntaxException e) {
return false;
}
String scheme = uri.getScheme();
if (scheme == null || isSshLikeScheme(scheme)) {
return false;
}
String userInfo = uri.getUserInfo();
return userInfo != null && !userInfo.isEmpty();
}
/**
* SSH-family schemes deliberately excluded from {@link #urlLeaksUserInfo}: there, the user part
* selects an account and authentication itself happens over SSH, so it is not a credential the
* way HTTPS/HTTP user-info is.
*/
private static boolean isSshLikeScheme(String scheme) {
return "ssh".equalsIgnoreCase(scheme) || "git+ssh".equalsIgnoreCase(scheme)
|| "ssh+git".equalsIgnoreCase(scheme);
}
/** Configure a per-worktree helper that supplies a token from the member environment at call time. */
private void configureEnvironmentCredentialHelper(String repoRoot, String worktreePath) {
exec("git", "-C", repoRoot, "config", "extensions.worktreeConfig", "true");
// An empty helper resets values inherited from the system or global config. Without it Git
// asks the next helper after this one, which can expose an operator-level credential.
exec("git", "-C", worktreePath, "config", "--worktree", "--replace-all", "credential.helper", "");
exec("git", "-C", worktreePath, "config", "--worktree", "--add", "credential.helper",
ENVIRONMENT_CREDENTIAL_HELPER);
}
/**
* {@link #configureEnvironmentCredentialHelper} only ever fires for an HTTPS origin — Git never
* consults a {@code credential.helper} for an SSH transport. This repo's own origin is
* {@code ssh://git@git.ltms.dev:2224/fleet/fleetd.git}, so a member sitting on that origin never
* reaches the helper and the repo-scoped {@code WORKER_GITEA_TOKEN} is simply not used.
*
* <p>An earlier version of this javadoc justified the rewrite by claiming a member <em>cannot</em>
* push once {@code memberCredentials.policy: allow-list} blocks {@code SSH_AUTH_SOCK}, because
* "there is no private key file on this host, only an ssh-agent socket". That premise is false
* (fleetd #184): {@code ssh -G} resolves a readable, passphrase-free {@code IdentityFile} outside
* {@code ~/.ssh}, and a member — same OS user — pushes over SSH with the socket blanked. The
* rewrite is still worth having, but for the reason below rather than that one: it routes the
* member through its own scoped token instead of the operator's ssh identity, which is what makes
* a member's pushes attributable and revocable.
*
* <p>The fix is a <em>worktree-scoped</em> URL rewrite: {@code url.<https-base>.insteadOf
* <ssh-base>}, set with {@code --worktree} so it lands only in
* {@code <worktree>/.git/worktrees/<name>/config.worktree} (enabled by
* {@code extensions.worktreeConfig}, already turned on above) and never touches the shared
* repo-level config the primary checkout also reads. {@code insteadOf} — not
* {@code pushInsteadOf} — because a member may also need to fetch or rebase, and both should go
* through the member's own token for the same reason.
*
* <p>The host (and, for the rewrite's SSH-side match, the port) come from parsing the origin
* itself — never a hardcoded forge host, which is exactly what #177 removed. An origin that is
* already {@code https://} is left alone; the credential helper already covers it. An origin
* that is neither {@code ssh://} nor {@code https://} — including the scp-like shorthand
* ({@code git@host:path}, no scheme) — is left untouched deliberately: that shorthand's
* {@code host:path} split is defined by the user's ssh_config aliases, not by URI syntax, so
* guessing at it risks rewriting to the wrong place. A repo provisioned from that form keeps
* today's (broken, if the policy blocks the agent) SSH-only behaviour rather than a wrong rewrite.
*/
/**
* Blank the user-info of a remote URL before it reaches a log. A remote URL is not obviously a
* credential channel, which is exactly why one has leaked here three times ({@code git remote -v}
* printing a token inline, and fleetd #157 / #182). An {@code ssh://} authority normally carries
* only {@code git@}, so this usually changes nothing — it is here so that the one origin that
* does carry a secret cannot print it. Matches every {@code ://…@} pair, not just the first.
*/
private static String redactUserInfo(String url) {
return url == null ? null : url.replaceAll("://[^@/]*@", "://<redacted>@");
}
private void configureHttpsUrlRewriteForSshOrigin(String repoRoot, String worktreePath) {
if (exitCode("git", "-C", repoRoot, "config", "--get", "remote.origin.url") != 0) {
return;
}
String origin = execRedacted("git", "-C", repoRoot, "config", "--get", "remote.origin.url").trim();
URI uri;
try {
uri = new URI(origin);
} catch (URISyntaxException e) {
log.warn("origin URL {} is not a valid URI; skipping worktree HTTPS rewrite", redactUserInfo(origin));
return;
}
String scheme = uri.getScheme();
if (!"ssh".equalsIgnoreCase(scheme)) {
// Already https:// (the credential helper covers it), or a scheme-less/scp-like origin
// left alone on purpose — see the javadoc above.
log.debug("origin scheme is not ssh ({}) — no worktree HTTPS rewrite needed", redactUserInfo(origin));
return;
}
String host = uri.getHost();
String authority = uri.getRawAuthority();
if (host == null || host.isBlank() || authority == null || authority.isBlank()) {
log.warn("ssh origin {} has no resolvable host; skipping worktree HTTPS rewrite", redactUserInfo(origin));
return;
}
String sshBase = "ssh://" + authority + "/";
String httpsBase = "https://" + host + "/";
exec("git", "-C", worktreePath, "config", "--worktree", "--replace-all",
"url." + httpsBase + ".insteadOf", sshBase);
log.info("worktree {} rewrites {} to {} (worktree-scoped; parent checkout untouched)",
worktreePath, redactUserInfo(sshBase), httpsBase);
}
/**
* Neutralize the worktree's worktree-hostile project configs so a worker inherits only the tools
* and environment its launcher mounts (the bridge via {@code --mcp-config}, the opencode config
@@ -648,126 +359,6 @@ public final class GitWorktrees implements Worktrees {
return deleted;
}
/**
* {@inheritDoc}
*
* <p>fleetd #185 stage 3. No-op — no process spawned, nothing logged — when {@link #group} is
* null/blank. Otherwise:
* <ol>
* <li>{@code git -C repoRoot config core.sharedRepository group} so every future write by
* either uid stays group-writable;</li>
* <li>a one-time {@code chgrp}/{@code chmod g+rwX} fix-up over the worktree directory and,
* under the repo's <em>common</em> git directory, {@code objects}, {@code refs},
* {@code logs}, {@code worktrees} and {@code packed-refs} — with setgid
* ({@code chmod g+s}) applied only to the directories among them, so files created later
* inherit the group;</li>
* <li>one INFO line naming the group and the paths touched.</li>
* </ol>
*
* <p><b>Every path is skipped when it does not exist.</b> {@code .git/logs} is absent in a repo
* with {@code core.logAllRefUpdates=false} or one that has had no ref update yet, and
* {@code packed-refs} is absent until refs are packed. Passing a missing path to {@code chgrp}
* exits non-zero, which would fail <em>every</em> provisioning spawn with a message blaming a
* group that is in fact fine.
*
* <p><b>The git directory is resolved, not assumed.</b> {@code <repoRoot>/.git} is a
* <em>file</em>, not a directory, when the checkout is itself a linked worktree — the very
* thing this class creates for every member. {@code git rev-parse --git-common-dir} gives the
* real shared store, and it may answer relatively, so it is resolved against {@code repoRoot}.
*
* <p><b>The fix-up re-runs on every spawn, by design.</b> {@code core.sharedRepository=group}
* governs only what git writes <em>after</em> it is set; the walk is what covers everything
* already on disk. It is not redundant work to optimise away — dropping it silently leaves
* pre-existing objects unreadable to the member. It costs three walks of the object store per
* spawn (about 3000 files in this repo, well under a second, but it grows with the repo).
*
* <p>This only fixes up file ownership/permissions on the operator's shared repo so a
* different-uid member can write to it — it isolates credentials, not the repository. A member
* in the group can still write the operator's git objects and refs.
*
* <p>Fails loudly: a missing group, or a {@code chgrp}/{@code chmod} refused because the
* operator is not a member of it, becomes a {@link WorktreeException} naming the group — never
* a silent skip that leaves a member unable to work with nothing in the log to explain why.
*/
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
if (group == null) {
return;
}
List<String> touched = new ArrayList<>();
try {
shareGroupRunner.apply(new String[]{"git", "-C", repoRoot, "config", "core.sharedRepository", "group"});
String commonDir = gitCommonDir(repoRoot);
shareGroupPathIfPresent(worktreePath, true, touched);
for (String name : List.of("objects", "refs", "logs", "worktrees")) {
shareGroupPathIfPresent(commonDir + "/" + name, true, touched);
}
shareGroupPathIfPresent(commonDir + "/packed-refs", false, touched);
} catch (WorktreeException e) {
throw new WorktreeException("cannot share worktree with group '" + group + "': "
+ e.getMessage() + " — the group must exist, and the fleetd operator ("
+ System.getProperty("user.name") + ") must be a member of it", e);
}
log.info("worktreeGroup={} shared repoRoot={} worktreePath={} paths={}",
group, repoRoot, worktreePath, touched);
}
/**
* The repo's <em>common</em> git directory as an absolute path — where {@code objects},
* {@code refs} and {@code worktrees} actually live. {@code git rev-parse --git-common-dir}
* answers relative to {@code repoRoot} in the ordinary case ({@code .git}) and absolutely for a
* linked worktree, so the answer is resolved against {@code repoRoot} either way. Never
* hardcode {@code repoRoot + "/.git"}: that is a FILE when the checkout is itself a linked
* worktree.
*/
private String gitCommonDir(String repoRoot) {
String answer = shareGroupRunner.apply(
new String[]{"git", "-C", repoRoot, "rev-parse", "--git-common-dir"});
String trimmed = answer == null ? "" : answer.trim();
if (trimmed.isEmpty()) {
trimmed = ".git";
}
return Path.of(repoRoot).resolve(trimmed).normalize().toString();
}
/**
* {@link #shareGroupPath} when {@code path} exists, recording it in {@code touched}; otherwise
* nothing at all. A missing path is normal, not an error — see {@link #shareWithGroup}'s
* javadoc for which ones are routinely absent and why passing them to {@code chgrp} would fail
* every spawn.
*/
private void shareGroupPathIfPresent(String path, boolean recursive, List<String> touched) {
if (!Files.exists(Path.of(path))) {
return;
}
shareGroupPath(path, recursive);
touched.add(path);
}
/**
* {@code chgrp}/{@code chmod g+rwX} {@code path} to {@link #group}. When {@code recursive},
* also walks the directories under {@code path} (including {@code path} itself, when it is a
* directory) and sets setgid on each — directories only, per the javadoc on
* {@link #shareWithGroup}.
*/
private void shareGroupPath(String path, boolean recursive) {
List<String> chgrp = new ArrayList<>(List.of("chgrp"));
if (recursive) chgrp.add("-R");
chgrp.add(group);
chgrp.add(path);
shareGroupRunner.apply(chgrp.toArray(new String[0]));
List<String> chmod = new ArrayList<>(List.of("chmod"));
if (recursive) chmod.add("-R");
chmod.add("g+rwX");
chmod.add(path);
shareGroupRunner.apply(chmod.toArray(new String[0]));
if (recursive) {
shareGroupRunner.apply(new String[]{"find", path, "-type", "d", "-exec", "chmod", "g+s", "{}", "+"});
}
}
/**
* Every {@code refs/wip/*} ref (see {@link WipRef}). The committer date is read as a unix
* count of seconds and converted to millis. {@code %00} (NUL) separates the fields because a
@@ -861,31 +452,6 @@ public final class GitWorktrees implements Worktrees {
/** Same as {@link #exec(String...)}, with extra environment variables set on the child process. */
private String exec(Map<String, String> extraEnv, String... command) {
return exec(extraEnv, false, command);
}
/**
* Same as {@link #exec(String...)}, for a command whose stdout may itself carry a credential
* (e.g. {@code git remote get-url}, whose output is a URL). Stdout is still returned normally on
* success — callers still get the URL to inspect — but it is suppressed from BOTH the timeout
* message and the non-zero-exit message, so a failing call here can never copy it into a
* {@link WorktreeException}, and from there into a caller's log.
*/
private String execRedacted(String... command) {
return exec(Map.of(), true, command);
}
/**
* Shared implementation for {@link #exec(Map, String...)} and {@link #execRedacted(String...)}.
* {@code redactOutput} suppresses captured stdout from both failure messages below.
*
* <p>Package-private, not {@code private}: also a test seam, the same way the
* {@code afterWorktreeAdded} constructor parameter is. It lets a test drive the redaction
* guarantee directly — a synthetic failing command whose stdout carries a test marker passed
* through {@code extraEnv} rather than argv — without depending on finding a real git failure
* mode that happens to echo a URL onto stdout before exiting non-zero.
*/
String exec(Map<String, String> extraEnv, boolean redactOutput, String... command) {
String out;
int code;
Process p;
@@ -907,8 +473,7 @@ public final class GitWorktrees implements Worktrees {
try {
if (!p.waitFor(30, TimeUnit.SECONDS)) {
p.destroyForcibly();
throw new WorktreeException("command timed out: " + String.join(" ", command)
+ (redactOutput || out.isBlank() ? "" : "\n" + out));
throw new WorktreeException("command timed out: " + String.join(" ", command) + "\n" + out);
}
code = p.exitValue();
} catch (InterruptedException e) {
@@ -918,7 +483,7 @@ public final class GitWorktrees implements Worktrees {
}
if (code != 0) {
throw new WorktreeException("exit " + code + " for: " + String.join(" ", command)
+ (redactOutput || out.isBlank() ? "" : "\n" + out));
+ (out.isBlank() ? "" : "\n" + out));
}
return out;
}
@@ -89,15 +89,4 @@ public record MemberSession(
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount + 1, state, worktree, branch, charterReceipt, agentSessionId);
}
/**
* Return a copy with {@code agentSessionId} resolved to a non-null value (fleetd #209). Some
* adapters (opencode) cannot answer {@link dev.ltms.fleet.peer.PeerHandle#agentSessionId()} at
* spawn time — the peer has not persisted its session record yet — so the id is discovered on
* a later poll and swapped into the otherwise-immutable session via this wither.
*/
public MemberSession withAgentSessionId(String agentSessionId) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId);
}
}
@@ -46,15 +46,6 @@ public final class SessionManager implements TurnListener {
private final PeerLauncher launcher;
private final Worktrees worktrees;
private final ConcurrentHashMap<String /*paneId*/, MemberSession> registry = new ConcurrentHashMap<>();
/**
* fleetd #209: the live {@link PeerHandle} for every registered pane, retained solely so
* {@link #resolveAgentSessionId} can re-poll {@link PeerHandle#agentSessionId()} after spawn.
* The handle used to go out of scope at the end of the spawn method, so a launcher that answers
* the id lazily (opencode — the on-disk session row is written after the pane is created) could
* never be re-asked, and {@code fleet_list}/{@code fleet_spawn resumeSessionId} never saw it.
* Populated on every spawn path, removed on {@link #release}.
*/
private final ConcurrentHashMap<String /*paneId*/, PeerHandle> handles = new ConcurrentHashMap<>();
private final MemberPresence presence;
private final SecureRandom nonceRandom = new SecureRandom();
private final AtomicLong nonceSeq = new AtomicLong();
@@ -214,7 +205,6 @@ public final class SessionManager implements TurnListener {
handle.charterReceipt(),
handle.agentSessionId());
registry.put(handle.id(), session);
handles.put(handle.id(), handle);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
@@ -267,10 +257,6 @@ public final class SessionManager implements TurnListener {
*/
private void release(String paneId, ReleaseCause cause) {
MemberSession removed = registry.remove(paneId);
// fleetd #209: remove right alongside the registry entry so a released session's handle is
// never leaked — but keep the local reference below, so the id can still be resolved for
// the ReleaseDetail this teardown notifies with.
PeerHandle removedHandle = handles.remove(paneId);
boolean preserveWorktree = cause == ReleaseCause.SHUTDOWN;
String snapshotRef = null;
if (removed != null) {
@@ -317,11 +303,8 @@ public final class SessionManager implements TurnListener {
// too, so a failed ticket's detail can point a lead at the same tree to re-dispatch.
// CB-584 (issue #65 criterion 5): carry agentSessionId alongside them, so a lead can
// also resume the member's conversation, not just re-dispatch onto its files.
// fleetd #209: a late-resolving adapter (opencode) may only now have an id — resolve
// one last time so a released member's detail carries the id it now has.
MemberSession resolved = resolveAgentSessionId(removed, removedHandle);
notifyReleased(new ReleaseDetail(resolved.terminalId(), resolved.worktree(),
resolved.branch(), snapshotRef, resolved.agentSessionId()));
notifyReleased(new ReleaseDetail(removed.terminalId(), removed.worktree(),
removed.branch(), snapshotRef, removed.agentSessionId()));
}
}
// CB-581: the pane must always stop, even if the dirty check above threw. A session removed
@@ -489,10 +472,6 @@ public final class SessionManager implements TurnListener {
try {
path = worktrees.add(repoRoot, branch, wt.baseRef());
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
// fleetd #185 stage 3: MUST run after overlayParity, not folded into add() — overlayParity
// copies more files into the worktree after add() returns, so sharing the group any earlier
// leaves those overlay files operator-owned and read-only for a different-uid member.
worktrees.shareWithGroup(repoRoot, path);
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, sessionName, resumeSessionId, memberRole));
} catch (RuntimeException e) {
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
@@ -525,7 +504,6 @@ public final class SessionManager implements TurnListener {
handle.charterReceipt(),
handle.agentSessionId());
registry.put(handle.id(), session);
handles.put(handle.id(), handle);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
@@ -557,73 +535,16 @@ public final class SessionManager implements TurnListener {
return launcher.defaultProfile();
}
/**
* The session for {@code paneId}, if it is still registered and not released. fleetd #209:
* resolves a still-unknown {@code agentSessionId} against the retained handle before returning,
* so {@code fleet_status} sees an id a lazy-resolving adapter has since written.
*/
/** The session for {@code paneId}, if it is still registered and not released. */
public Optional<MemberSession> get(String paneId) {
return Optional.ofNullable(registry.get(paneId)).map(this::resolveAgentSessionId);
return Optional.ofNullable(registry.get(paneId));
}
/**
* Fleet-owned roster: all registered sessions (acquired minus released). Deliberately does
* <strong>not</strong> resolve {@code agentSessionId} (fleetd #209 follow-up) — this is the
* roster supplier on the heartbeat and health-tick timers ({@code LeadHeartbeatLoop},
* {@code FleetHealthMonitor} in {@code Fleetd}), on the placement/exhaustion paths, and on the
* metrics scrape ({@code FleetMetrics}), all called far more often than any caller actually
* reads {@code agentSessionId}. Resolving here would mean every tick opens a lazy-resolving
* adapter's (opencode's) on-disk session store once per member whose id is still unknown — and
* for a member whose id never appears, that cost never stops, for the life of the process. Use
* {@link #rosterResolved()} instead wherever the id must be current.
*/
/** Fleet-owned roster: all registered sessions (acquired minus released). */
public List<MemberSession> roster() {
return List.copyOf(registry.values());
}
/**
* {@link #roster()}, with each session's still-unknown {@code agentSessionId} re-resolved
* against its retained handle (fleetd #209) — so a caller that actually reports the id (
* {@code fleet_list}, the REST roster) sees one a lazy-resolving adapter (opencode) has since
* written, rather than the null frozen in at spawn time. Reserved for caller-driven reads, not
* timers: see {@link #roster()}'s javadoc for why the plain roster must stay non-resolving.
*/
public List<MemberSession> rosterResolved() {
return registry.values().stream().map(this::resolveAgentSessionId).toList();
}
/**
* Resolve {@code session}'s {@code agentSessionId} if still unknown, re-polling the retained
* {@link PeerHandle} for this pane (fleetd #209). A no-op — returning {@code session} unchanged
* — once the id is already known, once no handle is retained for this pane (never spawned, or
* already released), or if the handle throws while answering. A resolved id is best-effort
* CAS-swapped into the registry via {@link #replace}; a lost race just means another caller
* already applied the same update, so the resolved value is returned either way.
*/
private MemberSession resolveAgentSessionId(MemberSession session) {
return resolveAgentSessionId(session, handles.get(session.paneId()));
}
private MemberSession resolveAgentSessionId(MemberSession session, PeerHandle handle) {
if (session.agentSessionId() != null || handle == null) {
return session;
}
String resolved;
try {
resolved = handle.agentSessionId();
} catch (RuntimeException e) {
log.debug("agentSessionId lookup failed for pane={} terminal={}: {}",
session.paneId(), session.terminalId(), e.toString());
return session;
}
if (resolved == null) {
return session;
}
MemberSession updated = session.withAgentSessionId(resolved);
replace(session, updated); // best-effort; a lost CAS just means the resolved value stands anyway
return updated;
}
/**
* CB-304 merged roster+live view. The registry is authoritative for worktree, branch,
* profile, owner, and state; the optional live agent supplies the herdr-reported status.
@@ -98,21 +98,4 @@ public interface Worktrees {
/** CB-586: the operator-visible census of {@code refs/wip/*} in one repository. */
record WipRefStats(int count, long costBytes) {
}
/**
* Make {@code repoRoot}'s git store and {@code worktreePath} writable by the configured group
* (fleetd #185 stage 3), so a member spawned as a different OS user (see
* {@code memberHerdrSocket}) can write its own worktree, its per-worktree git metadata, and
* its own commit objects. No-op when no group is configured.
*
* <p><strong>This isolates credentials, not the repository.</strong> A member in the group can
* still write the operator's git objects and refs in the shared repo — this only fixes file
* ownership/permissions so a different-uid member can work at all, it grants no narrower access
* than that.
*
* @param repoRoot the repository whose git store ({@code .git/objects}, {@code refs},
* {@code logs}, {@code worktrees}, {@code packed-refs}) needs sharing
* @param worktreePath the linked worktree's own directory
*/
void shareWithGroup(String repoRoot, String worktreePath);
}
@@ -1,31 +0,0 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-185: {@code ConnectionIdentity} must resolve a caller's pane on EITHER herdr daemon (a
* lead's MCP connection resolves against the lead daemon; a member's against the member daemon).
* Pinning {@code PaneLocator} to {@code memberHerdr} alone — the bug this guards against — leaves
* every lead's own connection unresolvable ({@code callerTerminal == null}) the moment
* {@code memberHerdrSocket} names a second daemon, which breaks {@code fleet_reply}/{@code
* fleet_ask} and {@code fleet_whoami} for a lead. A unit test on {@link
* dev.ltms.fleet.herdr.PaneLocator} alone (see {@code PaneLocatorTest}) proves the class CAN
* search two clients, but not that {@code Fleetd.main} actually wires it that way — hence this
* source-level assertion, the same technique {@code FleetdHerdrControlConstructionTest} uses.
*/
class FleetdConnectionIdentityConstructionTest {
@Test
void connectionIdentitySearchesBothDaemonsNotJustTheMemberOne() throws Exception {
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
assertFalse(source.contains("new PaneLocator(memberHerdr)"),
"PaneLocator must not be pinned to the member daemon alone — a lead's own "
+ "connection resolves against the LEAD daemon and would never be found");
assertTrue(source.contains("new PaneLocator(herdr, memberHerdr)"),
"PaneLocator must search the lead daemon first, then the member daemon");
}
}
@@ -1,29 +0,0 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-185: {@code FleetApp} must be constructed with BOTH herdr clients (the lead's and the
* member's), never the raw lead-only {@code herdr}. Passing only {@code herdr} — the bug this
* guards against — makes {@code GET /healthz} green while the member daemon is down (so every
* spawn fails invisibly) and silently drops every member workspace from {@code GET /sessions}.
* A behavioural test on {@code FleetApp} alone (see {@code FleetAppTwoDaemonTest}) proves the
* class merges/gates correctly when given two clients, but not that {@code Fleetd.main} actually
* passes it two — hence this source-level assertion, mirroring
* {@code FleetdHerdrControlConstructionTest}.
*/
class FleetdFleetAppConstructionTest {
@Test
void fleetAppIsConstructedWithBothHerdrDaemons() throws Exception {
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
assertFalse(source.contains("new FleetApp(herdr, workers,"),
"FleetApp must not be constructed with the lead-only herdr client");
assertTrue(source.contains("new FleetApp(herdr, memberHerdr, workers,"),
"FleetApp must be constructed with both the lead and the member herdr client");
}
}
@@ -1,17 +0,0 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
class FleetdHerdrControlConstructionTest {
@Test
void fleetdDelegatesStatefulControlsToTheRouter() throws Exception {
// AgentControl caches paneByTerminal, so the router must be its only production factory.
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
assertFalse(source.contains("new AgentControl("));
assertFalse(source.contains("new WorkspaceControl("));
}
}
@@ -1036,35 +1036,6 @@ class FleetConfigTest {
assertEquals(LeadMailbox.DEFAULT_PREFETCH, noEnv.prefetchOrDefault());
}
@Test
void absentWorktreeGroupLeavesItNull(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-worktree-group.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
FleetConfig cfg = FleetConfig.load(f);
assertNull(cfg.worktreeGroup(), "no worktreeGroup: key → null → GitWorktrees.shareWithGroup is a no-op");
}
@Test
void worktreeGroupKeyParses(@TempDir Path dir) throws Exception {
Path f = dir.resolve("worktree-group.yaml");
Files.writeString(f, "bind:\n port: 8080\nworktreeGroup: fleet-workers\n");
FleetConfig cfg = FleetConfig.load(f);
assertEquals("fleet-workers", cfg.worktreeGroup());
}
@Test
void absentWorktreeGroupSurvivesTheBackCompatConstructorChain() {
// fleetd #185 stage 3: withDefaults() (and every pre-existing call site) must not silently
// drop a live worktreeGroup by routing through a back-compat constructor that defaults it
// to null.
FleetConfig cfg = new FleetConfig(null, null, null, Map.of(), null, null, null, null, null,
null, null, null, null, null, null, null, null, null, null, null, "fleet-workers");
assertEquals("fleet-workers", cfg.withDefaults().worktreeGroup(),
"withDefaults() must carry a configured worktreeGroup through unchanged");
}
@Test
void absentPrimaryBlockLeavesPrimaryNull(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-primary.yaml");
@@ -31,8 +31,6 @@ public final class FakeHerdr implements HerdrClient {
*/
public final List<Call> calls = new CopyOnWriteArrayList<>();
private boolean healthy = true;
private String pingVersion = "0.8.0";
private int pingProtocol = 19;
private final List<String> extraWorkspaces = new ArrayList<>();
private final List<String> extraAgents = new ArrayList<>();
/** workspaceId → extra tabs that {@code tab.list} reports for it (CB-558 lead scans). */
@@ -42,7 +40,6 @@ public final class FakeHerdr implements HerdrClient {
private int workerTabPaneCount = 1;
private String paneCloseErrorCode = null;
private String agentSendErrorCode = null;
private boolean noPanes = false;
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String readText = "worker transcript tail"; // canned agent.read output
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
@@ -54,16 +51,6 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Make {@code ping} report this version/protocol instead of the default 0.8.0/19 — CB-185
* blocker 2's fixture for a lead and a member daemon running mismatched herdr versions.
*/
public FakeHerdr pingReports(String version, int protocol) {
this.pingVersion = version;
this.pingProtocol = protocol;
return this;
}
/** Reject the first {@code n} {@code agent.start} calls with {@code agent_name_taken}. */
public FakeHerdr agentNameTakenTimes(int n) {
this.agentNameTakenFor = n;
@@ -88,15 +75,6 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Make {@code pane.list} report no panes at all — models a second herdr daemon (CB-185) that
* simply does not host the pane a {@link PaneLocator} is searching for.
*/
public FakeHerdr withNoPanes() {
this.noPanes = true;
return this;
}
/** Set the {@code agent_status} that {@code agent.get} reports (drives the injector). */
public FakeHerdr agentStatus(String status) {
this.agentStatus = status;
@@ -178,8 +156,7 @@ public final class FakeHerdr implements HerdrClient {
try {
return switch (method) {
case "ping" -> mapper.readTree(
("{\"type\":\"pong\",\"version\":\"%s\",\"protocol\":%d}")
.formatted(pingVersion, pingProtocol));
"{\"type\":\"pong\",\"version\":\"0.8.0\",\"protocol\":19}");
case "workspace.list" -> mapper.readTree(("""
{"type":"workspace_list","workspaces":[
{"workspace_id":"w1","label":"dev-mgnl","focused":true,"pane_count":7,"agent_status":"unknown"},
@@ -291,9 +268,7 @@ public final class FakeHerdr implements HerdrClient {
case "pane.get" -> mapper.readTree("""
{"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9",
"tab_id":"w9:t2","agent_status":"idle"}}""");
case "pane.list" -> noPanes
? mapper.readTree("{\"type\":\"pane_list\",\"panes\":[]}")
: mapper.readTree("""
case "pane.list" -> mapper.readTree("""
{"type":"pane_list","panes":[
{"pane_id":"w2:p7","terminal_id":"term_a","workspace_id":"w2","tab_id":"w2:t7","agent":"claude"},
{"pane_id":"w2:p9","terminal_id":"term_shell","workspace_id":"w2","tab_id":"w2:t8"}]}""");
@@ -1,27 +0,0 @@
package dev.ltms.fleet.herdr;
import java.util.HashMap;
import java.util.Map;
import java.util.OptionalLong;
/**
* Fake {@link ParentResolver} backed by an explicit pid→parent map — lets {@link PaneLocatorTest}
* drive {@link PaneLocator}'s ancestry walk (grandchild pids, cycles) without spawning real OS
* processes.
*/
final class FakeParentResolver implements ParentResolver {
private final Map<Long, Long> parents = new HashMap<>();
/** {@code pid}'s parent is {@code parentPid}. A pid with no entry here has no known parent. */
FakeParentResolver parent(long pid, long parentPid) {
parents.put(pid, parentPid);
return this;
}
@Override
public OptionalLong parentOf(long pid) {
Long parent = parents.get(pid);
return parent == null ? OptionalLong.empty() : OptionalLong.of(parent);
}
}
@@ -1,31 +0,0 @@
package dev.ltms.fleet.herdr;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertNotSame;
class HerdrRouterTest {
@Test
void absentMemberClientSharesControlsForEveryTarget() {
FakeHerdr client = new FakeHerdr();
HerdrRouter router = new HerdrRouter(client, null, id -> id.equals("lead"));
assertSame(router.leadAgents(), router.memberAgents());
assertSame(router.leadSpaces(), router.memberSpaces());
assertSame(router.leadAgents(), router.agentsFor("lead"));
assertSame(router.leadAgents(), router.agentsFor("member"));
}
@Test
void separateClientsRouteLeadAndMemberTargets() {
FakeHerdr lead = new FakeHerdr();
FakeHerdr member = new FakeHerdr();
HerdrRouter router = new HerdrRouter(lead, member, id -> id.equals("lead"));
assertNotSame(router.leadAgents(), router.memberAgents());
assertNotSame(router.leadSpaces(), router.memberSpaces());
assertSame(router.leadAgents(), router.agentsFor("lead"));
assertSame(router.memberAgents(), router.agentsFor("member"));
}
}
@@ -1,11 +1,7 @@
package dev.ltms.fleet.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import org.junit.jupiter.api.Test;
import java.util.concurrent.atomic.AtomicInteger;
import static org.junit.jupiter.api.Assertions.*;
/** Unit tests for PID → pane resolution (the herdr half of connection-based MCP identity). */
@@ -28,162 +24,4 @@ class PaneLocatorTest {
assertNull(loc.terminalForPid(0));
assertNull(loc.terminalForPid(-1));
}
// --- two-daemon fallback (CB-185) -----------------------------------------
@Test
void fallsBackToTheMemberClientWhenTheLeadHasNoMatch() {
// The caller's pane lives on the member daemon only (e.g. the caller is a spawned
// member) — the lead client reports no panes at all, so the locator must fall back.
HerdrClient lead = new FakeHerdr().withNoPanes();
HerdrClient member = new FakeHerdr();
PaneLocator two = new PaneLocator(lead, member);
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID));
}
@Test
void searchesTheLeadClientBeforeTheMemberClient() {
// The caller's pane lives on the LEAD daemon (e.g. the caller is a peer lead) — with two
// daemons, resolving it must not depend on the member client having a matching pane.
HerdrClient lead = new FakeHerdr();
HerdrClient member = new FakeHerdr().withNoPanes();
PaneLocator two = new PaneLocator(lead, member);
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID));
}
@Test
void nullWhenNeitherClientHasTheMatch() {
PaneLocator two = new PaneLocator(new FakeHerdr().withNoPanes(), new FakeHerdr().withNoPanes());
assertNull(two.terminalForPid(FakeHerdr.WORKER_PID));
}
@Test
void collapsesToOneScanWhenLeadAndMemberAreTheSameClient() {
// The single-daemon deployment (no memberHerdrSocket configured): the two-arg constructor
// must behave exactly like the one-arg constructor, including making only one herdr call.
FakeHerdr shared = new FakeHerdr();
PaneLocator two = new PaneLocator(shared, shared);
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID));
long paneListCalls = shared.calls.stream().filter(c -> c.method().equals("pane.list")).count();
assertEquals(1, paneListCalls, "same-object lead/member must scan exactly once, not twice");
}
// --- ancestry walk (CB-161: grandchild pids matched no pane, resolving as primary) --------
@Test
void stillResolvesAPidThatIsExactlyThePaneShellPid() {
// Regression: a pid with no parent chain at all — no ancestry walk is needed to match it.
OnePaneHerdr pane = new OnePaneHerdr("term_x", "pX", 5000, 6000);
PaneLocator loc = new PaneLocator(pane, new FakeParentResolver());
assertEquals("term_x", loc.terminalForPid(5000));
}
@Test
void stillResolvesAPidThatIsExactlyAForegroundPid() {
// Regression: same as above, but matching via the foreground-processes list.
OnePaneHerdr pane = new OnePaneHerdr("term_x", "pX", 5000, 6000);
PaneLocator loc = new PaneLocator(pane, new FakeParentResolver());
assertEquals("term_x", loc.terminalForPid(6000));
}
@Test
void resolvesAGrandchildPidTwoLevelsBelowTheShellPid() {
// The bug: a helper process a worker spawns (python3, curl, ...) is a grandchild of the
// pane's shell — not the shell_pid and not a foreground pid directly. Before the fix,
// paneOwnsPid only checked direct pid equality, so this pid matched no pane and the
// caller fell through to loopback-trust as the primary.
OnePaneHerdr pane = new OnePaneHerdr("term_x", "pX", 5000, 6000);
FakeParentResolver parents = new FakeParentResolver()
.parent(7002, 7001) // grandchild -> child
.parent(7001, 5000); // child -> shell (the pane's shell_pid)
PaneLocator loc = new PaneLocator(pane, parents);
assertEquals("term_x", loc.terminalForPid(7002));
}
@Test
void nullForAPidWhoseAncestryMatchesNoPane() {
// Must not break the other direction: a pid that truly belongs to nothing here (e.g. the
// real primary) must still resolve to null. Resolving everything to a worker would demote
// the actual lead and refuse every orchestration call.
OnePaneHerdr pane = new OnePaneHerdr("term_x", "pX", 5000, 6000);
FakeParentResolver parents = new FakeParentResolver()
.parent(9002, 9001)
.parent(9001, 9000); // chain never reaches 5000 or 6000
PaneLocator loc = new PaneLocator(pane, parents);
assertNull(loc.terminalForPid(9002));
}
@Test
void ancestryWalkTerminatesOnACycleInsteadOfHanging() {
// A fake (or corrupted) parent map that cycles must not hang identity resolution, which
// runs on every MCP call. The walk must still terminate and correctly resolve to null.
OnePaneHerdr pane = new OnePaneHerdr("term_x", "pX", 5000, 6000);
FakeParentResolver parents = new FakeParentResolver()
.parent(100, 101)
.parent(101, 100); // cycle, never reaches the pane's pids
PaneLocator loc = new PaneLocator(pane, parents);
assertNull(loc.terminalForPid(100));
}
@Test
void ancestrySetIsComputedOnceAcrossBothClientsInTheTwoDaemonConstructor() {
// CB-185: the two-daemon constructor searches lead then member. The ancestor set is
// per-caller, not per-client — it must be walked once and reused, not recomputed for
// each client searched.
OnePaneHerdr pane = new OnePaneHerdr("term_x", "pX", 5000, 6000);
FakeParentResolver parents = new FakeParentResolver()
.parent(7002, 7001)
.parent(7001, 5000);
AtomicInteger calls = new AtomicInteger();
ParentResolver counting = pid -> {
calls.incrementAndGet();
return parents.parentOf(pid);
};
HerdrClient noPanes = new FakeHerdr().withNoPanes();
PaneLocator two = new PaneLocator(noPanes, pane, counting);
assertEquals("term_x", two.terminalForPid(7002));
assertEquals(3, calls.get(), "ancestry must be walked once (3 lookups: 7002, 7001, 5000), "
+ "not re-walked per herdr client");
}
/** Minimal single-pane {@link HerdrClient} fake, purpose-built for the ancestry tests above. */
private static final class OnePaneHerdr implements HerdrClient {
private final ObjectMapper mapper = new ObjectMapper();
private final String terminalId;
private final String paneId;
private final long shellPid;
private final long foregroundPid;
OnePaneHerdr(String terminalId, String paneId, long shellPid, long foregroundPid) {
this.terminalId = terminalId;
this.paneId = paneId;
this.shellPid = shellPid;
this.foregroundPid = foregroundPid;
}
@Override
public JsonNode call(String method, Object params) {
try {
return switch (method) {
case "pane.list" -> mapper.readTree(("""
{"type":"pane_list","panes":[
{"pane_id":"%s","terminal_id":"%s","workspace_id":"w1","tab_id":"w1:t1","agent":"claude"}]}""")
.formatted(paneId, terminalId));
case "pane.process_info" -> mapper.readTree(("""
{"type":"pane_process_info","process_info":{"pane_id":"%s","shell_pid":%d,
"foreground_processes":[{"pid":%d,"name":"node","argv0":"claude"}]}}""")
.formatted(paneId, shellPid, foregroundPid));
default -> throw new HerdrException("OnePaneHerdr has no canned response for " + method);
};
} catch (HerdrException e) {
throw e;
} catch (Exception e) {
throw new HerdrException("OnePaneHerdr decode failed for " + method, e);
}
}
@Override
public void close() {
}
}
}
@@ -197,15 +197,10 @@ class CompletionResolverTest {
void resolvesSynchronouslyBeforePostTurnContextClearing() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ previous answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
// fleetd#164: an injectable clock, so this real (non-crash) turn lands outside MIN_TURN_NANOS
// — captureBaseline and resolveBeforePostAction below run back-to-back with no real delay.
long[] clock = {1_000_000_000L};
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(), () -> clock[0]);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.captureBaseline("term_a", new TurnToken("term_a", waiter));
herdr.readText("⏺ answer that /clear would erase\n❯ ");
clock[0] += CompletionResolver.MIN_TURN_NANOS + 1; // this turn took longer than the floor
resolver.resolveBeforePostAction("term_a");
@@ -224,11 +219,7 @@ class CompletionResolverTest {
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
Rendezvous rendezvous = new Rendezvous();
// fleetd#164: an injectable clock so captureBaseline and resolve (back-to-back, no real
// delay) don't trip the too-fast-turn floor — this test is about the suppression guard, not timing.
long[] clock = {1_000_000_000L};
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(), () -> clock[0]);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
resolver.captureBaseline("term_a", new TurnToken("term_a", waiter)); // baseline is the clipped >cap block
@@ -236,7 +227,6 @@ class CompletionResolverTest {
assertEquals(CompletionResolver.MAX_SCRAPE_CHARS, turn.baseline().length(),
"the delivery baseline is clipped to the same cap resolve() applies to the tail");
clock[0] += CompletionResolver.MIN_TURN_NANOS + 1; // outside the floor
resolver.resolve("term_a", turn); // scrape unchanged → clipped tail == baseline → suppress
assertFalse(waiter.isDone(),
@@ -259,13 +249,12 @@ class CompletionResolverTest {
}
@Test
void aFailedScrapeResolvesAsAFailureEvenWithABaselinePresent() {
// fleetd#164: this test used to assert that a failed read resolved the send as a SUCCESS
// carrying an empty string ("an empty tail beats hanging until the caller's timeout") — that
// was the bug this ticket fixes: a lost turn and a genuine empty answer looked identical to
// every caller. This test encoded the bug and is changed here: a failed read must fail the
// send instead, naming the member, whether or not a baseline was captured (the baseline here
// is "", an empty pane at delivery — proof this isn't the CB-115 misattribution path either).
void resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent() {
// The most important branch of the CB-115 guard: a failed read means the resolver could not
// SEE the screen — "couldn't see", not "no change". It must still resolve the send (an empty
// tail beats hanging until the caller's timeout), even though a baseline was captured. The
// baseline here is "" (an empty pane at delivery), so without the !scrapeFailed clause the
// byte-identical guard would wrongly match the empty tail and suppress.
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.read throws HerdrException
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
@@ -276,82 +265,8 @@ class CompletionResolverTest {
assertTrue(waiter.isDone(),
"a failed scrape must still resolve the send, not hang until the caller's timeout");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
"a failed read is a lost turn, not a successful empty reply");
assertTrue(waiter.getNow(null).text().contains("term_a"),
"the failure names the member: " + waiter.getNow(null).text());
assertTrue(waiter.getNow(null).text().contains("could not be read"),
"the failure explains the scrape could not be read: " + waiter.getNow(null).text());
}
// --- fleetd#164: an empty (but readable) scrape must never resolve as a success --------------
@Test
void anEmptyScrapeResolvesAsAFailureNamingTheMember() {
// The core defect: a scrape that read CLEANLY but produced zero characters used to resolve
// the send as a SUCCESS carrying "" — indistinguishable, to every caller, from a worker that
// genuinely finished with nothing to say. A lost turn must never look like a real empty reply.
FakeHerdr herdr = new FakeHerdr().readText("");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "an empty scrape must still resolve the send, not hang");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
"an empty scrape is a lost turn, not a successful empty reply");
assertTrue(waiter.getNow(null).text().contains("term_a"),
"the failure names the member: " + waiter.getNow(null).text());
assertTrue(waiter.getNow(null).text().toLowerCase().contains("empty"),
"the failure says the scrape was empty: " + waiter.getNow(null).text());
}
// --- fleetd#164: a BUSY -> DONE transition inside the floor is a crash, not a fast answer ------
@Test
void aBusyToDoneTransitionInsideTheFloorResolvesAsAFailure() {
// The exact fleetd#164 scenario: the backend returned an HTTP 400 before the worker did
// anything, and the member went BUSY -> DONE in ~1s. That transition alone is indistinguishable
// from a genuine (if unusually fast) completion, so the resolver leans on the floor to catch it.
FakeHerdr herdr = new FakeHerdr().readText("⏺ HTTP 400: invalid request\n❯ ");
Rendezvous rendezvous = new Rendezvous();
long[] clock = {10_000_000_000L};
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(), () -> clock[0]);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null, clock[0]); // delivered "now"
clock[0] += CompletionResolver.MIN_TURN_NANOS - 1; // 1ns inside the floor
resolver.resolve("term_a", turn);
assertTrue(waiter.isDone(), "a suspiciously fast turn must still resolve (as a failure)");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
assertTrue(waiter.getNow(null).text().contains("term_a"),
"the failure names the member: " + waiter.getNow(null).text());
assertTrue(waiter.getNow(null).text().contains("HTTP 400"),
"the failure carries whatever was on screen: " + waiter.getNow(null).text());
}
@Test
void aBusyToDoneTransitionJustOutsideTheFloorResolvesNormally() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ a real, if quick, answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
long[] clock = {10_000_000_000L};
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(), () -> clock[0]);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null, clock[0]); // delivered "now"
clock[0] += CompletionResolver.MIN_TURN_NANOS + 1; // 1ns outside the floor
resolver.resolve("term_a", turn);
assertTrue(waiter.isDone());
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"a turn that took longer than the floor resolves normally");
assertEquals("a real, if quick, answer", waiter.getNow(null).text());
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
assertEquals("", waiter.getNow(null).text(), "the tail is empty because the screen was unreadable");
}
// --- CB-115/CB-116 fail guard: an already-done or absent waiter is left alone ---------
@@ -566,182 +481,6 @@ class CompletionResolverTest {
assertEquals("The usage limit has been reached.", waiter.getNow(null).text());
}
// --- fleetd#164 (part 2 addendum): narrow BACKEND_ERROR pattern classification ---------
@Test
void classifiesABackendErrorLineAsAFailureInsteadOfACompletedReply() {
String block = "⏺ API Error: 400 invalid request body\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "a backend-error scrape still resolves the blocked send");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
"a backend rejection is a failure, not a completed reply");
}
@Test
void theBackendErrorReasonNamesTheMemberAndCarriesTheMatchedLine() {
String block = "⏺ API Error: 400 invalid request body\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
String reason = waiter.getNow(null).text();
assertTrue(reason.contains("term_a"), "the failure names the member: " + reason);
assertTrue(reason.contains("API Error: 400 invalid request body"),
"the failure carries the matched backend-error line: " + reason);
}
@Test
void aCaseInsensitiveApiErrorLineIsStillClassifiedAsABackendError() {
String block = "⏺ api error: rate limited\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(), "the pattern is case-insensitive");
}
@Test
void aBackendErrorFailureStillCarriesTheRestOfTheScrape() {
// The pattern is a heuristic: a member that forgot fleet_reply while *reporting on* a backend
// error matches it too. Failing is still correct, but the report itself must survive — losing
// it would be the same information-destroying defect fleetd#164 exists to fix.
String block = "\u23fa I looked into the gateway problem.\n"
+ "The log line was: API Error: 400 invalid request body\n"
+ "The cause is a missing content-type header.\n\u276f ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
String reason = waiter.getNow(null).text();
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
assertTrue(reason.contains("The cause is a missing content-type header."),
"the failure carries the rest of the pane, not only the matched line: " + reason);
}
// --- fleetd#211: raw-scrape fallback classification when there is no usable assistant block ---
@Test
void anExhaustionLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrapeAndNotifiesTheSink() {
// No ⏺ anywhere, and the first visible line is TUI chrome (╭). lastAssistantBlock's boundary
// scan starts at the top of the raw screen and breaks immediately, so the trimmed block is "".
// The fix: fall back to matching the RAW scrape so this doesn't get lost as an empty scrape.
String block = """
╭──────────────────────────────────────╮
The usage limit has been reached. Try again later.
""";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "a raw-scrape match still resolves the blocked send");
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, waiter.getNow(null).kind(),
"classified from the raw scrape even though the trimmed block was empty");
assertEquals(1, notified.size(),
"the sink is the whole point of this ticket — it must be notified: " + notified);
assertTrue(notified.get(0).startsWith("term_a: "), "the sink is told which target exhausted");
assertTrue(notified.get(0).contains("The usage limit has been reached"),
"the sink is told the matched reason: " + notified.get(0));
}
@Test
void aBackendErrorLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrape() {
String block = """
╭──────────────────────────────────────╮
API Error: 400 invalid request body
""";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "a raw-scrape backend-error match still resolves the blocked send");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
"classified BACKEND_ERROR from the raw scrape even though the trimmed block was empty");
assertTrue(waiter.getNow(null).text().contains("API Error: 400 invalid request body"),
"the failure carries the matched line: " + waiter.getNow(null).text());
assertTrue(waiter.getNow(null).text().contains("--- pane tail ---"),
"fleetd#164: the failure must carry the pane, not only the matched line — with an "
+ "empty trimmed block the raw scrape is the only copy of what the member said: "
+ waiter.getNow(null).text());
assertTrue(waiter.getNow(null).text().contains("╭"),
"the carried pane is the raw scrape, chrome included: " + waiter.getNow(null).text());
}
@Test
void anOrdinaryPaneWithANormalAssistantBlockIsUnaffectedByTheRawScrapeFallback() {
// Pin: on a pane that already yields a usable block, the fallback branch is never reached —
// same outcome, same text, sink not called — even though the raw screen around the marker
// would itself match the configured exhausted pattern.
String block = "The usage limit has been reached, but this is a leading TUI line above the "
+ "marker.\n⏺ complete report\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone());
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"unchanged: a usable assistant block never reaches the raw-scrape fallback");
assertEquals("complete report", waiter.getNow(null).text(), "the reply text is unaffected");
assertTrue(notified.isEmpty(), "the fallback never runs, so the sink is never called");
}
@Test
void aGenuinelyEmptyScrapeStillFailsAsEmptyAndNeverNotifiesTheSink() {
// The false-positive pin: no exhaustion or backend-error text anywhere on the pane (just
// chrome, no marker) — the raw-scrape fallback must not manufacture a classification, and
// the sink must stay untouched.
String block = """
╭──────────────────────────────────────╮
│ > │
╰──────────────────────────────────────╯
""";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "an empty scrape must still resolve the send, not hang");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
"no exhaustion or backend-error text anywhere ⇒ this stays the ordinary empty-scrape failure");
assertTrue(waiter.getNow(null).text().toLowerCase().contains("empty"),
"the failure still says the scrape was empty: " + waiter.getNow(null).text());
assertTrue(notified.isEmpty(), "a genuinely empty pane must never quarantine a credential");
}
@Test
void coverageIsOffWhenNoProfileHasAPatternConfigured() {
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [terra])",
@@ -1,75 +0,0 @@
package dev.ltms.fleet.inject;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.msg.TestTurnTokens;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-185: with a router split across two herdr daemons, {@link StatusPoller} must refine a raw
* {@code UNKNOWN} status by reading the pane content from the SAME daemon the status was sampled
* from — the lead daemon for a lead target, the member daemon for a member target. Reading the
* wrong daemon never finds the pane, classification stays {@code UNKNOWN} forever, and the
* status-gated {@link Injector} wedges: a queued message is never delivered.
*
* <p>This exercises the real production classes ({@code StatusPoller(HerdrRouter, ...)},
* {@code Injector(HerdrRouter, ...)}) wired together, not a hand-built object graph — the earlier
* three CB-185 bugs all passed exactly that kind of test while the real wiring stayed broken.
*/
class StatusPollerRoutingTest {
private static final String LEAD_TARGET = "term_a";
@Test
void refinesALeadTargetFromTheLeadDaemonAndDelivers() throws Exception {
// The lead daemon's pane is at a settled idle prompt; the member daemon's pane content is
// unclassifiable garbage. A correct refiner reads the LEAD daemon and delivers.
FakeHerdr leadHerdr = new FakeHerdr().agentStatus("unknown").readText("⏺ answer\n❯ ");
FakeHerdr memberHerdr = new FakeHerdr().agentStatus("unknown")
.readText("garbled ansi noise with no prompt");
HerdrRouter router = new HerdrRouter(leadHerdr, memberHerdr, LEAD_TARGET::equals);
Injector injector = new Injector(router, TurnListener.NOOP, _ -> true, _ -> {
});
StatusPoller poller = new StatusPoller(router, injector, 10);
poller.start();
try {
CompletableFuture<Void> delivered =
injector.enqueue(LEAD_TARGET, "via-poller", TestTurnTokens.inert(LEAD_TARGET));
// Must resolve quickly: refining against the WRONG daemon (member) never classifies
// out of UNKNOWN, so this would time out under the bug.
delivered.get(2, TimeUnit.SECONDS);
} finally {
poller.stop();
}
assertTrue(leadHerdr.called("agent.read"), "refine must probe the LEAD daemon's pane content");
}
@Test
void aLeadTargetNeverDeliversWhenOnlyTheMemberDaemonIsClassifiable() throws Exception {
// Inverted control: the member daemon's content WOULD classify to idle, but this is a lead
// target — a correct implementation must not use it, so delivery must NOT happen.
FakeHerdr leadHerdr = new FakeHerdr().agentStatus("unknown")
.readText("garbled ansi noise with no prompt");
FakeHerdr memberHerdr = new FakeHerdr().agentStatus("unknown").readText("⏺ answer\n❯ ");
HerdrRouter router = new HerdrRouter(leadHerdr, memberHerdr, LEAD_TARGET::equals);
Injector injector = new Injector(router, TurnListener.NOOP, _ -> true, _ -> {
});
StatusPoller poller = new StatusPoller(router, injector, 10);
poller.start();
try {
CompletableFuture<Void> delivered =
injector.enqueue(LEAD_TARGET, "via-poller", TestTurnTokens.inert(LEAD_TARGET));
assertThrows(TimeoutException.class, () -> delivered.get(500, TimeUnit.MILLISECONDS),
"a lead target must never be refined from the member daemon's pane content");
} finally {
poller.stop();
}
}
}
@@ -85,34 +85,4 @@ class StatusRefinerTest {
assertEquals(AgentStatus.UNKNOWN, refiner.refine("term_a", AgentStatus.UNKNOWN));
}
// --- refine(target, raw, control) — CB-185 per-call routing ---------------
@Test
void threeArgRefineReadsThroughTheGivenControlNotTheConstructedOne() {
// The refiner is CONSTRUCTED with one control (standing in for "the member daemon"), but
// a call names a DIFFERENT control (standing in for "the lead daemon") — the read must go
// to the one passed to the call, since that is the daemon the raw status came from.
FakeHerdr constructedWith = new FakeHerdr().readText("nothing recognizable here");
FakeHerdr passedToCall = new FakeHerdr().readText("⏺ answer\n❯ ");
StatusRefiner refiner = new StatusRefiner(new AgentControl(constructedWith));
AgentStatus result = refiner.refine("term_a", AgentStatus.UNKNOWN, new AgentControl(passedToCall));
assertEquals(AgentStatus.IDLE, result, "must classify from the PASSED control's pane content");
assertTrue(passedToCall.called("agent.read"));
assertFalse(constructedWith.called("agent.read"),
"the control fixed at construction must not be read when a call-site control is given");
}
@Test
void twoArgRefineStillReadsTheConstructedControl() {
// The legacy 2-arg overload (single-daemon callers) must keep using the constructed
// control — this is refine(target, raw, control) called with the field as `control`.
FakeHerdr herdr = new FakeHerdr().readText("⏺ answer\n❯ ");
StatusRefiner refiner = new StatusRefiner(new AgentControl(herdr));
assertEquals(AgentStatus.IDLE, refiner.refine("term_a", AgentStatus.UNKNOWN));
assertTrue(herdr.called("agent.read"));
}
}
@@ -1,6 +1,5 @@
package dev.ltms.fleet.member;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
@@ -61,75 +60,8 @@ class ClaudeCodeLauncherTest {
assertTrue(args.stream().noneMatch(a -> a.contains("\"bridge\"")),
"the mount is named fleet since CB-632 — a member addresses its tools as "
+ "mcp__fleet__*, and CLAUDE.md's role-detection ladder names that prefix");
// fleetd #220: the charter travels as a FILE, never inline — charter prose on the command
// line overruns the byte cap of the pane line herdr types it into.
assertTrue(args.contains("--append-system-prompt-file"));
assertFalse(args.contains("--append-system-prompt"),
"the inline flag would put ~800 bytes of prose on the pane command line");
String charterFile = args.get(args.indexOf("--append-system-prompt-file") + 1);
assertTrue(readFile(charterFile).contains("fleet_reply"), "reply charter present in the file");
}
/** Read a charter file the launcher wrote, failing the test rather than the build on an IO error. */
private static String readFile(String path) {
try {
return java.nio.file.Files.readString(java.nio.file.Path.of(path));
} catch (java.io.IOException e) {
throw new AssertionError("charter file " + path + " is not readable", e);
}
}
/**
* fleetd #220 regression. herdr TYPES the launch command into the pane, and the pty line buffer
* holds 1024 bytes — past that the tail is dropped with no error anywhere, so the backend exits
* on a mangled argument and the spawn dies as an unexplained readiness timeout. That is what
* happened when #214 added --session-id to a command already 978 bytes long: --autocompact
* 250000 arrived as --autocompact 25. This asserts the whole assembled command still fits, with
* the flags a real spawn carries (MCP mount, charter, model, autocompact, session id).
*/
@Test
void theAssembledLaunchCommandFitsThePaneLineLimit() {
FakeHerdr herdr = new FakeHerdr();
// Shaped like the live sonnet profile, because the bug is a SUM: the charter alone fits,
// and so does every flag alone. Only model + autocompact + session id on top of the charter
// crossed the cap, which is why nothing caught it until a member failed to spawn.
FleetConfig.Profile cfg = new FleetConfig.Profile(
"sonnet", null, "claude-sonnet-5", null, null,
List.of("claude"), "tab", "fleetd-workers", "w #{n}", "http://127.0.0.1:8765/mcp",
null, null, null, null, null, null, null, null, true, null, null, null, null, null,
250000);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of()), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
.spawn();
List<String> args = spawnedArgs(herdr);
int bytes = "claude".length();
for (String arg : args) {
bytes += arg.getBytes(java.nio.charset.StandardCharsets.UTF_8).length + 3;
}
assertTrue(bytes <= HerdrPeerLauncher.PANE_COMMAND_BYTE_LIMIT,
"the launch command must fit the pane line: " + bytes + " bytes vs limit "
+ HerdrPeerLauncher.PANE_COMMAND_BYTE_LIMIT + " — args " + args);
}
/**
* fleetd #220: the guard refuses a command that cannot fit, instead of letting the pty drop the
* tail. The refusal must name the size and the argument to blame — a spawn that fails with
* "did not reach injectable state" tells the operator nothing, which is the whole reason this
* bug took a live pane scrape to find.
*/
@Test
void anOverlongLaunchCommandIsRefusedWithTheSizeAndTheCulprit() {
FakeHerdr herdr = new FakeHerdr();
String huge = "x".repeat(1500);
ClaudeCodeLauncher launcher = service(herdr, List.of("claude", huge), null);
PeerUnreachableException refused = assertThrows(PeerUnreachableException.class, launcher::spawn);
assertTrue(refused.getMessage().contains("1024"), "names the limit: " + refused.getMessage());
assertTrue(refused.getMessage().contains("1500"), "names the culprit's size: " + refused.getMessage());
assertFalse(herdr.called("agent.start"),
"nothing may be started — a truncated command is worse than no spawn");
assertTrue(args.contains("--append-system-prompt"));
assertTrue(args.stream().anyMatch(a -> a.contains("fleet_reply")), "reply charter present");
}
// CB-634: a profile with ideMcpUrl set mounts the IDE Index MCP as a second server and pins
@@ -329,16 +261,10 @@ class ClaudeCodeLauncherTest {
Map<?, ?> start = (Map<?, ?>) herdr.lastCall("agent.start").params();
assertEquals("claude", start.get("kind"), "herdr launches the canonical executable by kind");
// CB-533: the shared fixture pins model "coder", so the model flag is the whole args list.
// What this test guards is that argv[0] is NOT repeated — herdr supplies it from `kind`.
// fleetd #214: a plain spawn now also carries fleetd's own minted session id.
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("claude"), "the configured executable is not repeated in args");
int flag = args.indexOf("--session-id");
assertTrue(flag >= 0, "a plain spawn mints a session id: " + args);
assertDoesNotThrow(() -> UUID.fromString(args.get(flag + 1)), "the minted id is a valid UUID");
assertEquals(List.of("--model", "coder"),
List.of(args.get(args.size() - 2), args.get(args.size() - 1)),
"the CB-533 model flag still trails the launch flags: " + args);
assertEquals(List.of("--model", "coder"), start.get("args"),
"the configured executable is not repeated in args");
}
@Test
@@ -351,15 +277,7 @@ class ClaudeCodeLauncherTest {
assertFalse(args.contains("--append-system-prompt"), "no reply charter without mcpUrl");
// CB-533: the model flag is independent of the MCP mount — pinning the model is not part of
// "mount the bridge", so an unmounted worker still runs the model its profile names.
// fleetd #214: a plain spawn now carries fleetd's own minted session id; drop the two
// mint elements when checking the rest of the argv.
int flag = args.indexOf("--session-id");
assertTrue(flag >= 0, "a plain spawn mints a session id even without a bridge mount: " + args);
assertDoesNotThrow(() -> UUID.fromString(args.get(flag + 1)), "the minted id is a valid UUID");
List<String> rest = new java.util.ArrayList<>(args);
rest.remove(flag + 1);
rest.remove(flag);
assertEquals(List.of("--verbose", "--model", "coder"), rest,
assertEquals(List.of("--verbose", "--model", "coder"), args,
"the operator's own args are preserved, in order, ahead of the model flag");
}
@@ -730,7 +648,7 @@ class ClaudeCodeLauncherTest {
assertEquals("/work/proj", cwd, "effectiveCwd via SpawnRequest must match the three-arg resolution");
}
// --- CB-547a / fleetd #214: durable session identity (always mint / resume) -----------------
// --- CB-547a: durable session identity (mint / resume / no-identity legacy) -----------------
@Test
void freshSpawnMintsASessionIdAndPassesTheName() {
@@ -765,28 +683,18 @@ class ClaudeCodeLauncherTest {
}
@Test
void plainSpawnMintsASessionIdSoEveryMemberIsResumable() {
// fleetd #214: a plain spawn passes no sessionName and no resumeSessionId, yet the member
// must still be resumable — the id is minted unconditionally, and it is the ONLY resume
// handle a claude-code member has (unlike opencode, nothing resolves it after the launch).
// The binary requires a valid UUID (checked against claude 2.1.252: a non-UUID is refused
// at argument parsing with "Invalid session ID. Must be a valid UUID.").
void noIdentitySpawnKeepsTheLegacyArgvAndCarriesNoSessionHandle() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, null));
List<String> args = spawnedArgs(herdr);
int flag = args.indexOf("--session-id");
assertTrue(flag >= 0 && flag + 1 < args.size(),
"--session-id is minted even when no identity is requested: " + args);
String minted = args.get(flag + 1);
assertDoesNotThrow(() -> UUID.fromString(minted), "--session-id is a valid UUID: " + minted);
assertEquals(minted, handle.agentSessionId(),
"the resume handle is the minted id, so every member is resumable from fleet_list");
assertFalse(args.contains("-n"), "no sessionName was requested → no -n");
assertFalse(args.contains("-r"), "no resume was requested → no -r");
assertNull(handle.sessionName(), "no sessionName was requested → no logical name");
assertFalse(args.contains("--session-id"), "no identity → no --session-id");
assertFalse(args.contains("-n"), "no identity → no -n");
assertFalse(args.contains("-r"), "no identity → no -r");
assertNull(handle.agentSessionId(), "no identity → no resume handle");
assertNull(handle.sessionName(), "no identity → no logical name");
}
@Test
@@ -1117,7 +1025,7 @@ class ClaudeCodeLauncherTest {
* {@code System.getenv()} so the test is deterministic.
*/
@Test
void aCredentialShapedNameOnNeitherListIsLoggedAsAGap() {
void denyByDefaultKeepsTheExactCredentialGapWarn() {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
@@ -1137,79 +1045,39 @@ class ClaudeCodeLauncherTest {
logger.detachAppender(appender);
}
String expected = "memberCredentials gap: 1 credential-shaped env var name(s) are on neither "
+ "known: nor allow: — every member pane inherits them UNBLOCKED — "
+ "[A_BRAND_NEW_SECRET_TOKEN]. Add each to memberCredentials.known (blocked by default) "
+ "or .allow (if a member legitimately needs it).";
assertTrue(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN")),
"the gap must name the unrecognized credential-shaped var, never a value");
e.getLevel() == ch.qos.logback.classic.Level.WARN
&& e.getFormattedMessage().equals(expected)),
"deny-by-default must keep the exact existing WARN text");
assertNotNull(startEnv(herdr).get("GITEA_ACCESS_TOKEN"),
"deny-by-default must still overlay a known blocked credential");
assertFalse(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("PATH")),
"PATH/HOME are not credential-shaped and must not be reported as a gap");
assertFalse(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("AI_GATEWAY_TOKEN")),
"a name already on allow: is covered, not a gap");
}
/**
* CB-633 follow-up (#192): the deny-by-default WARN wording is a promise an operator relies on —
* pinned byte-for-byte so a future edit cannot drift it (e.g. while picking wording for the
* allow-list path) without a test noticing.
*/
@Test
void denyByDefaultKeepsTheExactCredentialGapWarn() {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("claude"), "tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null, 0, System::currentTimeMillis, () -> {}, null, () -> TEST_MEMBER_CREDENTIALS,
() -> Set.of("A_BRAND_NEW_SECRET_TOKEN"));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
svc.spawn();
} finally {
logger.detachAppender(appender);
}
assertTrue(appender.list.stream().anyMatch(e ->
("memberCredentials gap: 1 credential-shaped env var name(s) are on neither "
+ "known: nor allow: — every member pane inherits them UNBLOCKED — "
+ "[A_BRAND_NEW_SECRET_TOKEN]. Add each to memberCredentials.known "
+ "(blocked by default) or .allow (if a member legitimately needs it).")
.equals(e.getFormattedMessage())),
"the deny-by-default WARN text must not drift — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
}
/**
* CB-633 follow-up (#192), defect 1: under {@code policy: allow-list} on a zsh login shell the
* generated ZDOTDIR scrub genuinely blanks an unkept credential-shaped name, so the report must
* not say the pane inherits it UNBLOCKED — that claim is exactly what PR #174 got wrong. This
* goes through the real spawn path (not {@code HerdrPeerLauncherAllowListWiringTest}'s fixture,
* which overrides {@code buildLaunch} and bypasses none of the logic under test here — the
* shell-dependent branch lives in {@code applyEnvironmentAllowListPolicy}, which every spawn
* still passes through).
*/
@Test
void allowListPolicyOnZshReportsTheGapWithoutClaimingItIsUnblocked() {
void allowListReportsAnUnkeptCredentialAsBlankedWithoutSayingUnblocked() {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("claude"), "tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
FleetConfig.MemberCredentials creds = new FleetConfig.MemberCredentials(
FleetConfig.MemberCredentials.POLICY_ALLOW_LIST,
List.of("AI_GATEWAY_TOKEN"), List.of());
FleetConfig.MemberCredentials.POLICY_ALLOW_LIST, List.of(), List.of(), null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
name -> "SHELL".equals(name) ? "/bin/zsh" : null,
0, System::currentTimeMillis, () -> {}, null, () -> creds,
() -> Set.of("AI_GATEWAY_TOKEN", "A_BRAND_NEW_SECRET_TOKEN"));
() -> Set.of("PATH", "HOME", "A_BRAND_NEW_SECRET_TOKEN"));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.INFO);
ch.qos.logback.classic.Level original = logger.getLevel();
logger.setLevel(ch.qos.logback.classic.Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
@@ -1220,213 +1088,59 @@ class ClaudeCodeLauncherTest {
logger.setLevel(original);
}
assertTrue(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN")),
"the unkept name must still be reported — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
assertFalse(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().contains("UNBLOCKED")),
"on zsh the scrub genuinely blanks the name, so the report must not claim it is "
+ "inherited unblocked — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
ILoggingEvent gap = appender.list.stream()
.filter(e -> e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN"))
.findFirst().orElseThrow(() -> new AssertionError("allow-list must report the blanked name"));
assertEquals(ch.qos.logback.classic.Level.INFO, gap.getLevel(),
"a scrubbed credential name is the allow-list control working, not a WARN");
assertTrue(gap.getFormattedMessage().contains("will be blanked by the scrub"));
assertFalse(gap.getFormattedMessage().contains("UNBLOCKED"),
"the allow-list report must not claim that a blanked name is exposed");
}
/**
* CB-633 follow-up (#192): the mirror of the zsh test above. On a non-zsh login shell {@code
* ZDOTDIR} is ignored, so no scrub ever runs — the report must keep the WARN wording (a name here
* really is inherited unblocked) rather than claiming a scrub protects it. This is the trap PR
* #174 fell into the other direction: keying the wording on the shell, not on {@code
* creds.isAllowList()}, is what keeps this branch correct.
* Defect fix: {@code memberCredentials} is live-reloadable, so the policy really can change
* between spawns on the SAME launcher instance. An earlier INFO gap report (allow-list) must
* not permanently suppress a later, more serious WARN gap report (deny-by-default) — the two
* report kinds need separate guards, not one shared flag.
*/
@Test
void allowListPolicyOnNonZshKeepsTheWarnWording() {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("claude"), "tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
FleetConfig.MemberCredentials creds = new FleetConfig.MemberCredentials(
FleetConfig.MemberCredentials.POLICY_ALLOW_LIST,
List.of("AI_GATEWAY_TOKEN"), List.of());
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
name -> "SHELL".equals(name) ? "/bin/bash" : null,
0, System::currentTimeMillis, () -> {}, null, () -> creds,
() -> Set.of("AI_GATEWAY_TOKEN", "A_BRAND_NEW_SECRET_TOKEN"));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
svc.spawn();
} finally {
logger.detachAppender(appender);
}
assertTrue(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().contains("UNBLOCKED")
&& e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN")),
"no scrub runs on a non-zsh shell, so the WARN wording must be kept — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
assertFalse(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().toLowerCase(java.util.Locale.ROOT).contains("scrub")),
"nothing is scrubbed on this path, so the report must not claim otherwise — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
}
/**
* CB-633 follow-up (#192), defect 2: {@code memberCredentials} is a live, re-read-per-spawn
* supplier, so the policy can change between two spawns on the same launcher. Before this fix a
* single {@code AtomicBoolean} guarded both report kinds, so the harmless allow-list INFO on the
* first spawn would permanently suppress the real deny-by-default WARN a later spawn deserves.
* This goes through {@link ClaudeCodeLauncher#buildLaunch}'s real {@code baseEnv()} path — the
* WARN this test pins fires from {@code applyMemberCredentialPolicy}, which {@code
* HerdrPeerLauncherAllowListWiringTest}'s fixture never reaches at all (see its class javadoc).
*/
@Test
void secondSpawnStillWarnsAfterPolicyChangesFromAllowListToDenyByDefault() {
void anEarlierInfoGapDoesNotSuppressALaterWarnGapOnTheSameLauncher() {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("claude"), "tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
AtomicReference<FleetConfig.MemberCredentials> creds = new AtomicReference<>(
new FleetConfig.MemberCredentials(FleetConfig.MemberCredentials.POLICY_ALLOW_LIST,
List.of("AI_GATEWAY_TOKEN"), List.of()));
new FleetConfig.MemberCredentials(
FleetConfig.MemberCredentials.POLICY_ALLOW_LIST, List.of(), List.of(), null));
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
name -> "SHELL".equals(name) ? "/bin/zsh" : null,
0, System::currentTimeMillis, () -> {}, null, creds::get,
() -> Set.of("AI_GATEWAY_TOKEN", "A_BRAND_NEW_SECRET_TOKEN"));
() -> Set.of("PATH", "HOME", "A_BRAND_NEW_SECRET_TOKEN"));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.INFO);
ch.qos.logback.classic.Level original = logger.getLevel();
logger.setLevel(ch.qos.logback.classic.Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
logger.addAppender(appender);
svc.spawn(); // allow-list + zsh: harmless INFO, sets the allow-list guard only
creds.set(new FleetConfig.MemberCredentials(FleetConfig.MemberCredentials.POLICY_DENY_BY_DEFAULT,
List.of("AI_GATEWAY_TOKEN"), List.of()));
svc.spawn(); // policy reloaded to deny-by-default: this WARN must NOT be suppressed
svc.spawn(); // allow-list: logs the INFO gap report first
creds.set(new FleetConfig.MemberCredentials(
FleetConfig.MemberCredentials.POLICY_DENY_BY_DEFAULT, List.of(), List.of(), null));
svc.spawn(); // deny-by-default: must still log its own WARN gap report
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
assertTrue(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
e.getLevel() == ch.qos.logback.classic.Level.WARN
&& e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN")
&& e.getFormattedMessage().contains("UNBLOCKED")),
"the second spawn's deny-by-default WARN must still fire even though the first "
+ "spawn's allow-list INFO already logged the same underlying gap — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
}
/**
* CB-633 follow-up (#192), lead-review fix: {@code effectiveAllowed} is a SUPERSET of
* {@code known ∪ allow} — {@code MemberEnvAllowList.derive} also unions in every profile's
* {@code tokenEnv} (among other fields), so a credential-shaped name can be uncovered by
* {@code known:}/{@code allow:} and STILL survive the scrub because a profile's own
* {@code tokenEnv} names it. Here {@code tokenEnv} is deliberately set to a credential-shaped
* name the operator forgot to list — the misconfiguration this report exists to catch. The scrub
* genuinely keeps it, so the report must WARN, not claim (as the pre-lead-review cut of this fix
* did) that "no member pane keeps them".
*/
@Test
void allowListWarnsWhenTheDerivedAllowListKeepsAnUncoveredName() {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "SOME_LEAKY_TOKEN",
List.of("claude"), "tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
FleetConfig.MemberCredentials creds = new FleetConfig.MemberCredentials(
FleetConfig.MemberCredentials.POLICY_ALLOW_LIST, List.of(), List.of());
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
name -> "SHELL".equals(name) ? "/bin/zsh" : null,
0, System::currentTimeMillis, () -> {}, null, () -> creds,
() -> Set.of("SOME_LEAKY_TOKEN"));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
svc.spawn();
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
assertTrue(appender.list.stream().anyMatch(e ->
e.getLevel() == Level.WARN
&& e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().contains("SOME_LEAKY_TOKEN")
&& e.getFormattedMessage().contains("UNBLOCKED")),
"a name kept by the derived allow-list (via this profile's tokenEnv) must still WARN "
+ "— got: " + appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
assertFalse(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().contains("SOME_LEAKY_TOKEN")
&& e.getFormattedMessage().contains("blanks them anyway")),
"the scrub does NOT blank this name, so the INFO wording must not claim it does — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
}
/**
* CB-633 follow-up (#192), lead-review fix: a mixed gap — one name the derived allow-list keeps
* (this profile's {@code tokenEnv}), one it does not — must split cleanly: the WARN names only
* the kept one, the INFO names only the blanked one. Proves the split uses {@code
* MemberEnvAllowList.keeps} per-name rather than an all-or-nothing decision for the whole gap.
*/
@Test
void allowListSplitsAMixedGapBetweenTheWarnAndTheInfo() {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "SOME_LEAKY_TOKEN",
List.of("claude"), "tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
FleetConfig.MemberCredentials creds = new FleetConfig.MemberCredentials(
FleetConfig.MemberCredentials.POLICY_ALLOW_LIST, List.of(), List.of());
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
name -> "SHELL".equals(name) ? "/bin/zsh" : null,
0, System::currentTimeMillis, () -> {}, null, () -> creds,
() -> Set.of("SOME_LEAKY_TOKEN", "A_BRAND_NEW_SECRET_TOKEN"));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
svc.spawn();
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
boolean warnNamesOnlyKept = appender.list.stream().anyMatch(e ->
e.getLevel() == Level.WARN
&& e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().contains("SOME_LEAKY_TOKEN")
&& !e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN"));
boolean infoNamesOnlyBlanked = appender.list.stream().anyMatch(e ->
e.getLevel() == Level.INFO
&& e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN")
&& !e.getFormattedMessage().contains("SOME_LEAKY_TOKEN"));
assertTrue(warnNamesOnlyKept, "the WARN must name the derived-list-kept variable and only it "
+ "— got: " + appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
assertTrue(infoNamesOnlyBlanked, "the INFO must name the scrub-blanked variable and only it "
+ "— got: " + appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
"a prior INFO gap report must not suppress a later WARN gap report on the same "
+ "launcher instance");
}
/**
@@ -8,7 +8,6 @@ import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.CharterReceipt;
@@ -249,193 +248,6 @@ class CompositePeerLauncherTest {
"stop routes to the spawning adapter and closes exactly that worker's pane");
}
@Test
void stopKeepsSameHerdrPaneIdSeparateByOwningAdapter() {
// Separate herdr daemons can both issue w9:pRoot_1. The opaque handles identify their
// spawning adapters, so each stop reaches only its recorded owner.
FakeHerdr first = new FakeHerdr();
FakeHerdr second = new FakeHerdr();
PeerLauncher composite = new CompositePeerLauncher(
List.of(claudeAdapter(first), opencodeAdapter(second)), "claude");
PeerHandle claude = composite.spawn(new SpawnRequest("claude", null, null));
PeerHandle opencode = composite.spawn(new SpawnRequest("gemini", null, null));
assertNotEquals(claude.id(), opencode.id(), "each public paneId keeps its adapter owner");
composite.stop(claude.id());
assertTrue(first.calls.stream().anyMatch(c -> c.method().equals("pane.close")
&& "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id"))),
"the first daemon closes its own pane");
assertFalse(second.called("pane.close"), "the matching pane on the second daemon stays live");
composite.stop(opencode.id());
assertTrue(second.calls.stream().anyMatch(c -> c.method().equals("pane.close")
&& "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id"))),
"the second daemon then closes its own pane");
}
@Test
void stopAllowsALegacyBarePaneIdWithOneDaemon() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = new CompositePeerLauncher(List.of(claudeAdapter(herdr)), "claude");
composite.stop("w9:pRoot_1");
assertTrue(herdr.calls.stream().anyMatch(c -> c.method().equals("pane.close")
&& "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id"))),
"one daemon keeps the legacy bare-pane routing behaviour");
}
@Test
void stopAllowsAnUnownedPaneIdWithTwoAdaptersSharingOneDaemon() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
composite.stop("w9:pRoot_1");
assertTrue(herdr.calls.stream().anyMatch(c -> c.method().equals("pane.close")
&& "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id"))),
"two adapter kinds sharing one daemon keep the fallback route");
}
@Test
void listKeepsBothPanesWhenTwoDaemonsShareAPaneId() {
// herdr pane ids are per-daemon counters, so two daemons really can both hold w1:p1 on
// different panes. Deduplicating on the pane id alone dropped one of the two real agents.
FakeHerdr first = new FakeHerdr().withAgent("x", "term_x", "w1:p1", "w1:t1");
FakeHerdr second = new FakeHerdr().withAgent("y", "term_y", "w1:p1", "w1:t1");
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(claudeAdapter(first), opencodeAdapter(second)), "claude");
List<Agent> agents = composite.list();
assertEquals(2, agents.stream().filter(a -> "w1:p1".equals(a.paneId())).count(),
"one w1:p1 per daemon survives — the pane id alone is not a unique key");
assertTrue(agents.stream().anyMatch(a -> "term_x".equals(a.terminalId())));
assertTrue(agents.stream().anyMatch(a -> "term_y".equals(a.terminalId())));
}
@Test
void listStillDeduplicatesTwoAdaptersSharingOneDaemon() {
// Both adapters ask the SAME daemon, so both see the same agent set. Without the dedupe this
// would report every agent twice; the daemon key must not break that.
FakeHerdr herdr = new FakeHerdr().withAgent("x", "term_x", "w1:p1", "w1:t1");
CompositePeerLauncher composite = composite(herdr);
List<Agent> agents = composite.list();
assertEquals(1, agents.stream().filter(a -> "w1:p1".equals(a.paneId())).count(),
"one daemon still reports each of its agents once");
}
@Test
void stopKeepsTheOwnerRecordWhenTheDelegateRefusesTheStop() {
// Removing the record before the delegate accepted the stop lost the owner on failure: the
// pane was still alive, but the retry landed in the ambiguous branch and refused it for good.
FakeHerdr first = new FakeHerdr().paneCloseFailsWith("pane_busy");
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(claudeAdapter(first), opencodeAdapter(new FakeHerdr())), "claude");
PeerHandle claude = composite.spawn(new SpawnRequest("claude", null, null));
assertThrows(HerdrException.class, () -> composite.stop(claude.id()));
// The retry must still know its owner — a HerdrException, never "ambiguous paneId".
assertThrows(HerdrException.class, () -> composite.stop(claude.id()),
"the owner record survives a failed stop, so the retry is not ambiguous");
}
@Test
void stopRejectsAnUnownedPaneIdWhenMultipleDaemonsCouldOwnIt() {
// CB-185 blocker 1: genuine ambiguity — pane ids are per-daemon counters, so two daemons
// can each really hold an agent at "w1:p1". Neither claims ownership through spawnedBy
// (empty, as after a restart), so the probe must find BOTH and refuse rather than guess.
FakeHerdr first = new FakeHerdr().withAgent("x", "term_x", "w1:p1", "w1:t1");
FakeHerdr second = new FakeHerdr().withAgent("y", "term_y", "w1:p1", "w1:t1");
PeerLauncher composite = new CompositePeerLauncher(
List.of(claudeAdapter(first), opencodeAdapter(second)), "claude");
IllegalArgumentException error = assertThrows(IllegalArgumentException.class,
() -> composite.stop("w1:p1"));
assertEquals("ambiguous paneId 'w1:p1': 2 configured herdr daemons report this pane — "
+ "no way to tell which one the caller means", error.getMessage());
}
@Test
void stopOnAPaneNoConfiguredDaemonKnowsIsTreatedAsAlreadyStopped() {
// CB-185 blocker 1, the zero-owner branch: spawnedBy is empty (as after a restart) and
// neither daemon's agent.list mentions this pane at all — it is already gone. A retried
// stop() on an already-gone pane must succeed quietly, not refuse forever.
FakeHerdr first = new FakeHerdr();
FakeHerdr second = new FakeHerdr();
PeerLauncher composite = new CompositePeerLauncher(
List.of(claudeAdapter(first), opencodeAdapter(second)), "claude");
assertDoesNotThrow(() -> composite.stop("w1:p1"));
assertFalse(first.called("pane.close"), "no owner was found, so no delegate is told to close anything");
assertFalse(second.called("pane.close"), "no owner was found, so no delegate is told to close anything");
}
@Test
void stopWithEmptySpawnedByResolvesTheOwnerThroughAProbeAndSkipsTheOtherDaemon() {
// CB-185 blocker 1, the main fix: after a restart spawnedBy is empty for every surviving
// member. stop() must still find the one daemon that actually knows the pane and route
// only to it — never touching the daemon that never held it.
FakeHerdr first = new FakeHerdr().withAgent("x", "term_x", "w1:p1", "w1:t1");
FakeHerdr second = new FakeHerdr();
PeerLauncher composite = new CompositePeerLauncher(
List.of(claudeAdapter(first), opencodeAdapter(second)), "claude");
composite.stop("w1:p1");
assertTrue(first.calls.stream().anyMatch(c -> c.method().equals("pane.close")
&& "w1:p1".equals(((Map<?, ?>) c.params()).get("pane_id"))),
"the daemon that actually knows the pane closes it");
assertFalse(second.called("pane.close"), "the daemon that never held the pane is never touched");
}
@Test
void aProbeSurvivesOneUnreachableDaemonAndStillFindsTheOwnerOnTheOtherOne() {
// CB-185 blocker 1 (lead review): a daemon that is DOWN while we probe must not abort the
// whole probe — the pane the OPERATOR actually wants stopped can live on a different,
// healthy daemon, and that pane must not become un-stoppable because a third one is down.
FakeHerdr down = new FakeHerdr().healthy(false);
FakeHerdr owner = new FakeHerdr().withAgent("x", "term_x", "w1:p1", "w1:t1");
PeerLauncher composite = new CompositePeerLauncher(
List.of(claudeAdapter(down), opencodeAdapter(owner)), "claude");
assertDoesNotThrow(() -> composite.stop("w1:p1"),
"the unreachable daemon must be skipped, not fail the whole stop");
assertTrue(owner.calls.stream().anyMatch(c -> c.method().equals("pane.close")
&& "w1:p1".equals(((Map<?, ?>) c.params()).get("pane_id"))),
"the healthy daemon that actually owns the pane still closes it");
}
@Test
void aProbedOwnerIsCachedSoARetryAfterAFailedStopNeedsNoSecondProbe() {
// CB-185 blocker 1: the probe's whole point is to be cheap on repeat — a failed stop (e.g.
// "pane_busy") must not force another agent.list() round trip on every retry.
FakeHerdr first = new FakeHerdr().withAgent("x", "term_x", "w1:p1", "w1:t1")
.paneCloseFailsWith("pane_busy");
FakeHerdr second = new FakeHerdr();
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(claudeAdapter(first), opencodeAdapter(second)), "claude");
assertThrows(HerdrException.class, () -> composite.stop("w1:p1"));
long listCallsAfterFirst = first.calls.stream().filter(c -> c.method().equals("agent.list")).count()
+ second.calls.stream().filter(c -> c.method().equals("agent.list")).count();
assertTrue(listCallsAfterFirst > 0, "the first stop needed a probe");
assertThrows(HerdrException.class, () -> composite.stop("w1:p1"),
"still failing on the retry, but through the cached owner");
long listCallsAfterSecond = first.calls.stream().filter(c -> c.method().equals("agent.list")).count()
+ second.calls.stream().filter(c -> c.method().equals("agent.list")).count();
assertEquals(listCallsAfterFirst, listCallsAfterSecond,
"the retry is served from the cache — no additional agent.list probe");
}
@Test
void opencodeContextResetIsANoOpAndWarnsOnlyOnce() {
FakeHerdr herdr = new FakeHerdr();
@@ -1,6 +1,5 @@
package dev.ltms.fleet.member;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
@@ -15,16 +14,13 @@ import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.PosixFileAttributeView;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.function.Function;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
@@ -54,6 +50,8 @@ class HerdrPeerLauncherAllowListWiringTest {
/** A name the daemon itself injects — it must survive its own scrub, so it must be allowed. */
private static final String INJECTED = "ANTHROPIC_BASE_URL";
private static final String OPERATOR_KEEP = "CONTEXT7_TOKEN";
private static final String UNLISTED_CREDENTIAL = "A_BRAND_NEW_SECRET_TOKEN";
@Test
void spawningUnderAllowListPolicyGivesThePaneAGeneratedZdotdir() {
@@ -85,6 +83,43 @@ class HerdrPeerLauncherAllowListWiringTest {
+ "), or the daemon's own configuration is blanked by its own control");
}
/** Explicit operator keeps widen the derived base, but all other exported names stay blocked. */
@Test
void operatorKeepSurvivesWhileAnUnlistedCredentialIsBlanked(@TempDir Path home) throws Exception {
Path zsh = Path.of("/bin/zsh");
assumeTrue(Files.isExecutable(zsh), "/bin/zsh not present — nothing to prove here");
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr, allowList(List.of(OPERATOR_KEEP)));
var spawned = launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
try {
Path zdotdir = Path.of(launcher.env.get("ZDOTDIR"));
ProcessBuilder probe = new ProcessBuilder(zsh.toString(), "-i", "-c",
"[[ -n \"$" + OPERATOR_KEEP + "\" ]] && print -r -- kept || print -r -- missing; "
+ "[[ -z \"$" + UNLISTED_CREDENTIAL
+ "\" ]] && print -r -- blanked || print -r -- leaked");
probe.environment().clear();
probe.environment().putAll(Map.of(
"HOME", home.toString(),
"PATH", "/usr/bin:/bin",
"SHELL", zsh.toString(),
"ZDOTDIR", zdotdir.toString(),
OPERATOR_KEEP, "needed-by-member-binary",
UNLISTED_CREDENTIAL, "must-not-survive"));
probe.redirectError(ProcessBuilder.Redirect.DISCARD);
Process process = probe.start();
String output = new String(process.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
assertTrue(process.waitFor(60, java.util.concurrent.TimeUnit.SECONDS),
"the zsh scrub probe did not exit within 60 seconds");
assertEquals(0, process.exitValue(), "the zsh scrub probe must exit cleanly");
assertEquals("kept\nblanked\n", output,
"memberCredentials.allow must add a keep, while an unlisted credential stays blanked");
} finally {
launcher.stop(spawned.id());
}
}
/** The default policy must not generate anything — an upgrade changes nothing until asked. */
@Test
void spawningUnderTheDefaultPolicyGeneratesNoZdotdir() {
@@ -114,429 +149,118 @@ class HerdrPeerLauncherAllowListWiringTest {
"bash ignores ZDOTDIR; setting it would be protection theatre");
}
private static Supplier<FleetConfig.MemberCredentials> allowList() {
return () -> new FleetConfig.MemberCredentials(
FleetConfig.MemberCredentials.POLICY_ALLOW_LIST, List.of(), List.of(), null);
/**
* Defect fix: under {@code policy: allow-list} with a non-zsh login shell, nothing is
* scrubbed — the overlay-only fallback is the whole protection. The gap must still be
* reported, and with the deny-by-default WARN wording ("inherited UNBLOCKED"), never the
* allow-list INFO wording ("will be blanked by the scrub"), because nothing is scrubbed here.
*/
@Test
void nonZshAllowListFallbackStillReportsTheGapAsAWarn() {
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/bash",
() -> Set.of("PATH", "HOME", UNLISTED_CREDENTIAL));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
} finally {
logger.detachAppender(appender);
}
ILoggingEvent gap = appender.list.stream()
.filter(e -> e.getFormattedMessage().contains(UNLISTED_CREDENTIAL))
.findFirst().orElseThrow(() -> new AssertionError(
"the non-zsh allow-list fallback must still report the credential gap"));
assertEquals(ch.qos.logback.classic.Level.WARN, gap.getLevel(),
"nothing is scrubbed on this path, so the report must use the WARN wording, not "
+ "the allow-list INFO wording");
assertTrue(gap.getFormattedMessage().contains("UNBLOCKED"),
"the fallback path scrubs nothing, so the wording must say so");
assertFalse(gap.getFormattedMessage().contains("will be blanked by the scrub"),
"nothing is scrubbed on this path — that wording would be false here");
}
/** Same as {@link #allowList()} but with an operator-configured {@code allow:} list. */
private static Supplier<FleetConfig.MemberCredentials> allowListWithAllow(List<String> allow) {
/**
* The live operator config has {@code SSH_AUTH_SOCK} on BOTH {@code allow:} AND
* {@code sshAuthSock: block}. Defect fix: {@code sshAuthSock: block} must win — the allow:
* entry must not silently widen the derived allow-list to include it.
*/
@Test
void sshAuthSockBlockWinsOverAnAllowListEntry(@TempDir Path home) throws Exception {
Path zsh = Path.of("/bin/zsh");
assumeTrue(Files.isExecutable(zsh), "/bin/zsh not present");
FakeHerdr herdr = new FakeHerdr();
Supplier<FleetConfig.MemberCredentials> live = () -> new FleetConfig.MemberCredentials(
FleetConfig.MemberCredentials.POLICY_ALLOW_LIST,
List.of("AI_GATEWAY_TOKEN", "WORKER_GITEA_TOKEN", "CONTEXT7_TOKEN", "GITEA_HOST",
"OPENCODE_AUTOMODE_MODEL", "SSH_AUTH_SOCK", "CLAUDE_CODE_MESSAGING_TOKEN"),
List.of(), "block");
WiringLauncher launcher = new WiringLauncher(herdr, live);
var spawned = launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
try {
Path zdotdir = Path.of(launcher.env.get("ZDOTDIR"));
ProcessBuilder probe = new ProcessBuilder(zsh.toString(), "-i", "-c",
"[[ -n \"$SSH_AUTH_SOCK\" ]] && print -r -- LEAKED || print -r -- blanked");
probe.environment().clear();
probe.environment().putAll(Map.of(
"HOME", home.toString(), "PATH", "/usr/bin:/bin", "SHELL", zsh.toString(),
"ZDOTDIR", zdotdir.toString(), "SSH_AUTH_SOCK", "/tmp/agent.sock"));
probe.redirectError(ProcessBuilder.Redirect.DISCARD);
Process pr = probe.start();
String out = new String(pr.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
pr.waitFor(60, java.util.concurrent.TimeUnit.SECONDS);
assertEquals("blanked\n", out, "sshAuthSock: block must win over allow:");
} finally {
launcher.stop(spawned.id());
}
}
/**
* Mirror of {@link #sshAuthSockBlockWinsOverAnAllowListEntry}: with {@code sshAuthSock: allow}
* (no conflict with {@code allow:}), {@code SSH_AUTH_SOCK} IS kept. Without this case, a fix
* that simply deletes {@code SSH_AUTH_SOCK} everywhere would also pass the block test.
*/
@Test
void sshAuthSockAllowKeepsTheVariable(@TempDir Path home) throws Exception {
Path zsh = Path.of("/bin/zsh");
assumeTrue(Files.isExecutable(zsh), "/bin/zsh not present");
FakeHerdr herdr = new FakeHerdr();
Supplier<FleetConfig.MemberCredentials> creds = () -> new FleetConfig.MemberCredentials(
FleetConfig.MemberCredentials.POLICY_ALLOW_LIST,
List.of("SSH_AUTH_SOCK"), List.of(), "allow");
WiringLauncher launcher = new WiringLauncher(herdr, creds);
var spawned = launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
try {
Path zdotdir = Path.of(launcher.env.get("ZDOTDIR"));
ProcessBuilder probe = new ProcessBuilder(zsh.toString(), "-i", "-c",
"[[ -n \"$SSH_AUTH_SOCK\" ]] && print -r -- kept || print -r -- blanked");
probe.environment().clear();
probe.environment().putAll(Map.of(
"HOME", home.toString(), "PATH", "/usr/bin:/bin", "SHELL", zsh.toString(),
"ZDOTDIR", zdotdir.toString(), "SSH_AUTH_SOCK", "/tmp/agent.sock"));
probe.redirectError(ProcessBuilder.Redirect.DISCARD);
Process pr = probe.start();
String out = new String(pr.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
pr.waitFor(60, java.util.concurrent.TimeUnit.SECONDS);
assertEquals("kept\n", out, "sshAuthSock: allow must keep SSH_AUTH_SOCK");
} finally {
launcher.stop(spawned.id());
}
}
private static Supplier<FleetConfig.MemberCredentials> allowList() {
return allowList(List.of());
}
private static Supplier<FleetConfig.MemberCredentials> allowList(List<String> allow) {
return () -> new FleetConfig.MemberCredentials(
FleetConfig.MemberCredentials.POLICY_ALLOW_LIST, allow, List.of(), null);
}
/**
* Same as {@link #allowList()} but with an operator-configured {@code known:} list, so the
* fallback overlay (the non-zsh / no-memberLoginShell path) has something visible to shadow —
* {@link FleetConfig.MemberCredentials#blockedSet()} is {@code known - allow}, so an empty
* {@code known} (what {@link #allowList()} uses) blocks nothing and a fallback test would have
* no sentinel entry to assert on.
*/
private static Supplier<FleetConfig.MemberCredentials> allowListWithKnown(List<String> known) {
return () -> new FleetConfig.MemberCredentials(
FleetConfig.MemberCredentials.POLICY_ALLOW_LIST, List.of(), known, null);
}
/**
* CB-633 follow-up: a name that lives ONLY in {@code memberCredentials.allow:} — no profile
* mentions it — must survive the scrub the real spawn path generates. Calling {@code
* MemberEnvAllowList.derive} directly (as {@link MemberEnvAllowListTest} does) would pass even
* if {@code HerdrPeerLauncher} never threaded {@code allow:} into the derivation at all; this
* test goes through {@link HerdrPeerLauncher#spawn}, the method the daemon actually calls at
* spawn time, so it proves the union is wired in, not just correct in isolation.
*/
@Test
void spawningUnderAllowListPolicyIncludesAnOperatorConfiguredAllowName() {
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr,
allowListWithAllow(List.of("OPERATOR_ONLY_NAME")));
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
Path dir = Path.of(launcher.env.get("ZDOTDIR"));
assertTrue(readAll(dir.resolve(EnvAllowListScrub.SCRUB_FILE)).contains("OPERATOR_ONLY_NAME"),
"a name only in memberCredentials.allow: must reach the generated scrub through the "
+ "real launcher spawn path");
}
/**
* {@code SSH_AUTH_SOCK} is a live ssh-agent handle, not a value — it must stay blocked under
* {@code allow-list} even when the operator lists it under {@code allow:}, because {@code
* sshAuthSock} defaults to blocked. Governed ONLY by {@code memberCredentials.sshAuthSock}.
*/
@Test
void sshAuthSockStaysBlockedEvenWhenListedInMemberCredentialsAllow() {
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr,
allowListWithAllow(List.of("SSH_AUTH_SOCK")));
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
Path dir = Path.of(launcher.env.get("ZDOTDIR"));
String scrub = readAll(dir.resolve(EnvAllowListScrub.SCRUB_FILE));
assertFalse(scrub.contains("'SSH_AUTH_SOCK'"),
"SSH_AUTH_SOCK must not be on the derived allow-list just because the operator put "
+ "it under allow: — sshAuthSock is unset here, so it defaults to block");
}
@Test
void brokerUriEnvStaysBlockedWhenListedInMemberCredentialsAllow() {
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr,
allowListWithAllow(List.of("BROKER_CONNECTION_URI")), "/bin/zsh", null,
() -> config("BROKER_CONNECTION_URI"));
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
Path dir = Path.of(launcher.env.get("ZDOTDIR"));
assertFalse(readAll(dir.resolve(EnvAllowListScrub.SCRUB_FILE)).contains("'BROKER_CONNECTION_URI'"),
"broker.uriEnv must not reach a member even when listed in memberCredentials.allow:");
}
@Test
void brokerUriEnvIsDeniedUnderTheDenyListPolicyEvenWhenAllowed() {
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr,
() -> new FleetConfig.MemberCredentials(null, List.of("BROKER_CONNECTION_URI"), List.of(), null),
"/bin/bash", null,
() -> config("BROKER_CONNECTION_URI"));
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
assertEquals("blocked-by-fleetd-cb596-see-gitea-issue-82", launcher.env.get("BROKER_CONNECTION_URI"),
"the deny-list overlay must deny broker.uriEnv even when allow: names it");
}
/**
* CB-633 follow-up criterion 3: on every allow-list spawn the daemon logs one INFO line, shaped
* "member credentials: allowed N of M", with real counts — not constants. Real path: the count
* is asserted after a real {@link HerdrPeerLauncher#spawn} call, reading the log the production
* code actually emits.
*/
@Test
void logsAnAllowedCountLineAgainstTheHostEnvironmentOnEverySpawn() {
FakeHerdr herdr = new FakeHerdr();
// INJECTED is a key of this launch's own env map, so it always survives; the other two are
// neither derived from the profile nor configured anywhere, so they are blanked. Real
// N=1 (INJECTED), real M=3 (all three names) — neither number is hardcoded in the assertion
// by coincidence, they follow directly from this fixture.
Set<String> hostEnvNames = Set.of(INJECTED, "SOME_UNRELATED_NAME", "ANOTHER_UNRELATED_NAME");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames);
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
assertTrue(appender.list.stream()
.anyMatch(e -> "member credentials: allowed 1 of 3".equals(e.getFormattedMessage())),
"expected 'member credentials: allowed 1 of 3', got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
}
/**
* Lead-review fix: on a NON-zsh shell no scrub ever runs (bash ignores {@code ZDOTDIR}), so the
* "allowed N of M" line — which describes what the scrub does — must not be printed there either.
* Before this fix the line was logged BEFORE the zsh gate, so a non-zsh host printed e.g.
* "allowed 1 of 3" while blocking nothing at all, telling an operator a control ran when it did
* not. Real path: goes through {@link HerdrPeerLauncher#spawn}, same as the sibling test above,
* with the shell fixed to bash so the fallback branch is the one exercised.
*/
@Test
void noAllowedCountLineIsEmittedOnTheNonZshFallbackPath() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of(INJECTED, "SOME_UNRELATED_NAME", "ANOTHER_UNRELATED_NAME");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/bash", () -> hostEnvNames);
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
assertFalse(appender.list.stream()
.anyMatch(e -> e.getFormattedMessage().startsWith("member credentials: allowed ")),
"no scrub runs on a non-zsh shell, so no 'allowed N of M' count may be printed — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
}
/**
* fleetd #185 stage 2 pin: with {@code memberHerdrSocket:} absent (today's only mode, and the
* default — this host runs no other), the gap detector's WARN/INFO conclusions read exactly as
* they did before this fix. Real path: {@code FLEETD_WORKER_TOKEN} (the test profile's own
* {@code tokenEnv}) is a name the derived allow-list keeps, so it gets the "UNBLOCKED" WARN;
* {@code SOME_UNKNOWN_SECRET_TOKEN} is not derived from anywhere, so it gets the "scrub blanks
* them" INFO. This is the exact shape #185 stage 2 must not touch on this path.
*/
@Test
void gapConclusionsAreByteIdenticalWhenMemberHerdrSocketIsAbsent() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> config(null));
List<String> messages = spawnAndCaptureLogs(launcher);
assertTrue(messages.contains("memberCredentials gap: 1 credential-shaped env var name(s) are on "
+ "neither known: nor allow: — the derived allow-list keeps them anyway (a "
+ "profile's gitTokenEnv/gitHostEnv/tokenEnv/env: names one, or this spawn "
+ "injects it), so every member pane inherits them UNBLOCKED — [FLEETD_WORKER_TOKEN]. "
+ "Add each to memberCredentials.known (or .allow if a member legitimately needs "
+ "it), or remove it from whatever profile setting derives it in."),
"expected the pre-existing UNBLOCKED WARN unchanged, got: " + messages);
assertTrue(messages.contains("memberCredentials gap: 1 credential-shaped env var name(s) are on "
+ "neither known: nor allow: — [SOME_UNKNOWN_SECRET_TOKEN]. The allow-list scrub "
+ "blanks them anyway (they are not on the derived allow-list), so no member pane "
+ "keeps them; add each to memberCredentials.known or .allow to make that explicit."),
"expected the pre-existing 'scrub blanks them' INFO unchanged, got: " + messages);
}
/**
* fleetd #185 stage 2: with {@code memberHerdrSocket:} configured, member panes run under a
* different OS user — {@link HerdrPeerLauncher#hostEnvNames} describes fleetd's own process, not
* that user's. Neither "inherits them UNBLOCKED" nor "scrub blanks them" is evidence-backed
* there, so neither may print; the single unknown-environment WARN must, naming the config key.
*/
@Test
void gapDetectorReportsUnknownInsteadOfAConclusionWhenMemberHerdrSocketIsConfigured() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock"));
List<String> messages = spawnAndCaptureLogs(launcher);
assertTrue(messages.stream().anyMatch(m -> m.contains("memberHerdrSocket")
&& m.contains("UNKNOWN") && m.contains("cannot be verified")),
"expected the unknown-member-environment WARN naming memberHerdrSocket, got: " + messages);
assertFalse(messages.stream().anyMatch(m -> m.contains("UNBLOCKED")),
"the 'inherits them UNBLOCKED' conclusion must not print once the evidence is about "
+ "the wrong (daemon's own) environment — got: " + messages);
assertFalse(messages.stream().anyMatch(m -> m.contains("scrub blanks them")),
"the 'scrub blanks them' conclusion must not print once the evidence is about the "
+ "wrong (daemon's own) environment — got: " + messages);
}
/**
* fleetd #185 stage 2: the unknown-environment WARN is a standing fact about this launcher's
* configuration, not per-spawn news — it must fire once per launcher instance, the same shape as
* every other one-time WARN in this class (e.g. {@code warnNonZsh}).
*/
@Test
void theUnknownEnvironmentWarnFiresOnceNotOncePerSpawn() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock"));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.WARN);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
long count = appender.list.stream()
.filter(e -> e.getFormattedMessage().contains("cannot be verified from here"))
.count();
assertEquals(1, count, "the unknown-member-environment WARN must fire once per launcher "
+ "instance, not once per spawn — got " + count + " occurrence(s) among: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
}
/**
* Hard constraint: the gap detector must never log an env var VALUE, only its NAME. {@code
* SOME_UNKNOWN_SECRET_TOKEN} resolves to a distinctive canary value through the same {@code env}
* lookup the launcher uses elsewhere (SHELL, PATH, token resolution) — proving the value IS
* resolvable does not mean the detector reads it, since {@link HerdrPeerLauncher#hostEnvNames}
* (names only) is its data source, never {@code env.apply(name)} for those names.
*/
@Test
void theGapDetectorNeverLogsAnEnvVarValueOnlyItsName() {
FakeHerdr herdr = new FakeHerdr();
String canary = "sekrit-value-CANARY-9f3a1b7c";
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> config(null), Map.of("SOME_UNKNOWN_SECRET_TOKEN", canary));
List<String> messages = spawnAndCaptureLogs(launcher);
assertTrue(messages.stream().anyMatch(m -> m.contains("SOME_UNKNOWN_SECRET_TOKEN")),
"expected the credential-shaped NAME to appear in the log, got: " + messages);
assertFalse(messages.stream().anyMatch(m -> m.contains(canary)),
"the log must never contain an env var VALUE, only its NAME — got: " + messages);
}
/**
* fleetd #213 defect 1, acceptance criterion 1: {@code memberHerdrSocket} configured and {@code
* memberLoginShell} configured as non-zsh must fall back to the sentinel overlay exactly like a
* non-zsh {@code $SHELL} does today — and the "generated ZDOTDIR" INFO must not appear, since no
* scrub actually runs. The WiringLauncher's own {@code env("SHELL")} is deliberately set to
* {@code /bin/zsh} — the OPPOSITE of what {@code memberLoginShell} says — so a launcher that
* (incorrectly) fell back to fleetd's own {@code $SHELL} here would wrongly pass the gate and
* fail this test.
*/
@Test
void memberHerdrSocketWithNonZshMemberLoginShellFallsBackToTheOverlay() {
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr, allowListWithKnown(List.of("SOME_TOKEN")),
"/bin/zsh", null,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/bash"));
List<String> messages = spawnAndCaptureLogs(launcher);
assertEquals("blocked-by-fleetd-cb596-see-gitea-issue-82", launcher.env.get("SOME_TOKEN"),
"a non-zsh memberLoginShell must fall back to the CB-596 sentinel overlay, exactly "
+ "like a non-zsh $SHELL does when memberHerdrSocket is absent");
assertFalse(launcher.env.containsKey("ZDOTDIR"),
"no scrub directory may be generated when the configured member login shell is not zsh");
assertFalse(messages.stream().anyMatch(m -> m.contains("generated ZDOTDIR")),
"the 'generated ZDOTDIR' INFO must not appear when the scrub never runs — got: " + messages);
}
/**
* fleetd #213 defect 1, acceptance criterion 2: {@code memberHerdrSocket} configured and NO
* {@code memberLoginShell} configured must fall back exactly like criterion 1 above — AND
* fleetd's own {@code $SHELL} must never even be consulted (not merely "not decisive"). The
* fixture's {@code env} function reports {@code /bin/zsh} for {@code SHELL} — a value that would
* WRONGLY pass the zsh gate if the fix regressed to reading it — while flagging whether it was
* ever asked for at all, so this test fails loudly on either kind of regression.
*/
@Test
void memberHerdrSocketWithNoMemberLoginShellFallsBackAndNeverConsultsFleetdsOwnShell() {
FakeHerdr herdr = new FakeHerdr();
AtomicBoolean shellQueried = new AtomicBoolean(false);
Function<String, String> env = name -> {
if ("SHELL".equals(name)) {
shellQueried.set(true);
return "/bin/zsh"; // would wrongly pass the zsh gate if this ever leaked through
}
return null;
};
WiringLauncher launcher = new WiringLauncher(herdr, allowListWithKnown(List.of("SOME_TOKEN")), env,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", null));
List<String> messages = spawnAndCaptureLogs(launcher);
assertFalse(shellQueried.get(), "fleetd's own $SHELL must never be consulted once "
+ "memberHerdrSocket is configured — only memberLoginShell: may decide the gate");
assertEquals("blocked-by-fleetd-cb596-see-gitea-issue-82", launcher.env.get("SOME_TOKEN"),
"no memberLoginShell configured must fall back to the sentinel overlay, same as a "
+ "configured non-zsh shell");
assertFalse(messages.stream().anyMatch(m -> m.contains("generated ZDOTDIR")),
"no scrub may run without a configured memberLoginShell — got: " + messages);
}
/**
* fleetd #213 defect 2, acceptance criterion 3: with {@code memberHerdrSocket} configured, a
* zsh {@code memberLoginShell}, and {@code worktreeRoot}/{@code worktreeGroup} both configured,
* the generated scrub directory must live under {@code worktreeRoot} — NEVER under {@code
* java.io.tmpdir}, which is fleetd's own 0700 temp dir and unreadable by the member's different
* OS user. {@code worktreeGroup} is set to the CURRENT process's own primary group so {@code
* EnvAllowListScrub}'s group-sharing step resolves on whatever host runs this test, rather than
* hardcoding a group name that may not exist here.
*/
@Test
void memberHerdrSocketWithZshMemberLoginShellPutsTheScrubOutsideJavaIoTmpdir(@TempDir Path worktreeRoot)
throws IOException {
String group = currentUserGroup();
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/bash-should-be-ignored", null,
() -> configWithMemberHerdrSocketRootAndGroup("/tmp/other-user-herdr.sock", "/bin/zsh",
worktreeRoot.toString(), group));
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
String zdotdir = launcher.env.get("ZDOTDIR");
assertNotNull(zdotdir, "a zsh memberLoginShell with worktreeRoot+worktreeGroup configured "
+ "must still generate a ZDOTDIR");
Path dir = Path.of(zdotdir);
// The immediate PARENT is asserted (not merely startsWith(java.io.tmpdir)), because a
// JUnit @TempDir is itself carved out of the JVM's java.io.tmpdir — startsWith alone would
// pass by coincidence of the test fixture, not because the launcher used worktreeRoot.
assertEquals(worktreeRoot.toAbsolutePath().normalize(), dir.getParent(),
"the generated scrub directory's parent must be the configured worktreeRoot, not "
+ "System.getProperty(\"java.io.tmpdir\") — got parent " + dir.getParent());
}
/**
* fleetd #213, acceptance criterion 4: with {@code memberHerdrSocket} absent (today's only
* mode), behaviour must be byte-identical to before this fix — fleetd's own {@code $SHELL}
* still decides the gate (proven here, not merely assumed, by flagging the lookup), and the
* scrub still lands under {@code java.io.tmpdir}.
*/
@Test
void memberHerdrSocketAbsentStillConsultsFleetdsOwnShellAndBehavesAsBefore() {
FakeHerdr herdr = new FakeHerdr();
AtomicBoolean shellQueried = new AtomicBoolean(false);
Function<String, String> env = name -> {
if ("SHELL".equals(name)) {
shellQueried.set(true);
return "/bin/zsh";
}
return null;
};
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), env, null); // memberHerdrSocket absent
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
assertTrue(shellQueried.get(), "with memberHerdrSocket absent, fleetd's own $SHELL must "
+ "still decide the zsh gate, unchanged from before this fix");
String zdotdir = launcher.env.get("ZDOTDIR");
assertNotNull(zdotdir, "SHELL=/bin/zsh with memberHerdrSocket absent must still generate a "
+ "ZDOTDIR, as before this fix");
Path dir = Path.of(zdotdir);
assertTrue(dir.startsWith(Path.of(System.getProperty("java.io.tmpdir"))),
"with memberHerdrSocket absent the scrub directory must still be generated under "
+ "java.io.tmpdir, unchanged from before this fix: " + dir);
}
/** The current process's own primary group — resolvable on whatever host runs this test. */
private static String currentUserGroup() throws IOException {
PosixFileAttributeView view = Files.getFileAttributeView(Path.of("."), PosixFileAttributeView.class);
assumeTrue(view != null, "this host's filesystem does not support POSIX group ownership");
return view.readAttributes().group().getName();
}
/** Spawn once through the real launcher path, capturing every INFO+ line this class logs. */
private static List<String> spawnAndCaptureLogs(HerdrPeerLauncher launcher) {
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
return appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList();
}
private static String readAll(Path p) {
try {
return Files.readString(p);
@@ -566,54 +290,16 @@ class HerdrPeerLauncherAllowListWiringTest {
this(herdr, creds, shell, null);
}
/** Plus an injectable {@code hostEnvNames} source, for the "allowed N of M" log line test. */
WiringLauncher(FakeHerdr herdr, Supplier<FleetConfig.MemberCredentials> creds, String shell,
Supplier<Set<String>> hostEnvNames) {
this(herdr, creds, shell, hostEnvNames, null);
}
WiringLauncher(FakeHerdr herdr, Supplier<FleetConfig.MemberCredentials> creds, String shell,
Supplier<Set<String>> hostEnvNames, Supplier<FleetConfig> config) {
this(herdr, creds, shell, hostEnvNames, config, Map.of());
}
/**
* Plus a host-env value map (name → value), resolved through the same {@code env} lookup
* every adapter uses for {@code SHELL}/{@code PATH}/token resolution — fleetd #185 stage 2's
* "the gap detector never logs a value" tests use this to prove a value that IS resolvable
* for a credential-shaped name never reaches the log, since the detector only ever reads
* {@code hostEnvNames} (names), never {@code env.apply(name)} (values), for those names.
*/
WiringLauncher(FakeHerdr herdr, Supplier<FleetConfig.MemberCredentials> creds, String shell,
Supplier<Set<String>> hostEnvNames, Supplier<FleetConfig> config,
Map<String, String> extraEnvValues) {
super("test", new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("test", profile()), "test",
name -> "SHELL".equals(name) ? shell
: (extraEnvValues != null && extraEnvValues.containsKey(name))
? extraEnvValues.get(name) : null,
0, () -> 0L, () -> { }, null, creds, hostEnvNames, config);
}
/**
* Full control over the {@code env} lookup, bypassing the {@code shell}/{@code
* extraEnvValues} convenience above entirely — fleetd #213's "fleetd's own $SHELL must
* never be consulted" tests need to OBSERVE whether {@code SHELL} was ever looked up, which
* a plain value substitution cannot do.
*/
WiringLauncher(FakeHerdr herdr, Supplier<FleetConfig.MemberCredentials> creds,
Function<String, String> env, Supplier<FleetConfig> config) {
super("test", new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("test", profile()), "test",
env, 0, () -> 0L, () -> { }, null, creds, null, config);
name -> "SHELL".equals(name) ? shell : null,
0, () -> 0L, () -> { }, null, creds, hostEnvNames);
}
@Override
protected Launch buildLaunch(FleetConfig.Profile cfg, LaunchSpec spec) {
Map<String, String> launchEnv = baseEnv(cfg);
launchEnv.putAll(env);
env.clear();
env.putAll(launchEnv);
return new Launch(env, List.of("test"));
}
@@ -623,88 +309,6 @@ class HerdrPeerLauncherAllowListWiringTest {
}
}
private static FleetConfig config(String brokerUriEnv) {
return new FleetConfig(null, null, null, Map.of(), null, null, null, null, null,
new FleetConfig.Broker(null, brokerUriEnv, null), null, null, null, null, null,
null, null, null, null, null).withDefaults();
}
/**
* fleetd #185 stage 2: a config with {@code memberHerdrSocket:} set — member panes run on a
* second herdr owned by a different OS user, so {@link HerdrPeerLauncher#hostEnvNames} no
* longer describes what a member pane inherits.
*/
private static FleetConfig configWithMemberHerdrSocket(String memberHerdrSocket) {
return new FleetConfig(null, null, memberHerdrSocket, Map.of(), null, null, null, null, null,
null, null, null, null, null, null, null, null, null, null, null).withDefaults();
}
/**
* fleetd #213: as {@link #configWithMemberHerdrSocket(String)}, plus the {@code
* memberLoginShell:} the member's OS user actually runs — the config key {@link
* HerdrPeerLauncher#applyEnvironmentAllowListPolicy} must consult instead of fleetd's own
* {@code $SHELL} once {@code memberHerdrSocket} is configured.
*/
private static FleetConfig configWithMemberHerdrSocket(String memberHerdrSocket, String memberLoginShell) {
return new FleetConfig(
null, // bind
null, // herdrSocket
memberHerdrSocket, // memberHerdrSocket
Map.of(), // profiles
null, // guard
null, // worktreeRoot
null, // lifecycle
null, // spawnReadyTimeoutMs
null, // spawnReadyPollMs
null, // broker
null, // primary
null, // fleet
null, // leadHeartbeat
null, // health
null, // placement
null, // auth
null, // configReload
null, // quarantineCooldownSeconds
null, // memberCredentials
null, // coordinator
null, // worktreeGroup
memberLoginShell // memberLoginShell
).withDefaults();
}
/**
* fleetd #213: as above, plus {@code worktreeRoot:}/{@code worktreeGroup:} — both required for
* the ZDOTDIR scrub to run at all once {@code memberHerdrSocket} is configured; either missing
* falls back to the sentinel overlay, same as a non-zsh {@code memberLoginShell}.
*/
private static FleetConfig configWithMemberHerdrSocketRootAndGroup(String memberHerdrSocket,
String memberLoginShell, String worktreeRoot, String worktreeGroup) {
return new FleetConfig(
null, // bind
null, // herdrSocket
memberHerdrSocket, // memberHerdrSocket
Map.of(), // profiles
null, // guard
worktreeRoot, // worktreeRoot
null, // lifecycle
null, // spawnReadyTimeoutMs
null, // spawnReadyPollMs
null, // broker
null, // primary
null, // fleet
null, // leadHeartbeat
null, // health
null, // placement
null, // auth
null, // configReload
null, // quarantineCooldownSeconds
null, // memberCredentials
null, // coordinator
worktreeGroup, // worktreeGroup
memberLoginShell // memberLoginShell
).withDefaults();
}
/** The generated directory is a temp directory; make sure the test does not leave a pile. */
@Test
void theGeneratedDirectoryIsRemovedWhenThePaneIsStopped() {
@@ -88,67 +88,6 @@ class MemberEnvAllowListTest {
assertTrue(after.containsAll(Set.of("TOKEN_SECOND", "GIT_TOK", "SECOND_KEY")));
}
/**
* CB-633 follow-up: a name that appears ONLY in {@code memberCredentials.allow:} — no profile
* mentions it at all — must still survive the derivation. Before this fix {@code derive} never
* saw {@code allow:}, so setting {@code policy: allow-list} silently blanked exactly this name.
*/
@Test
void aNameOnlyInMemberCredentialsAllowSurvivesDerivation() {
FleetConfig.Profile p = profile("p", "TOKEN_A", null, null, Map.of("KEY_A", "v"));
Set<String> derived = MemberEnvAllowList.derive(List.of(p), Set.of("OPERATOR_ONLY_NAME"));
assertTrue(derived.contains("OPERATOR_ONLY_NAME"),
"memberCredentials.allow: must be unioned in, not ignored");
// and the profile-derived half must still be present — this is a union, not a replacement.
assertTrue(derived.contains("KEY_A"));
assertTrue(derived.contains("TOKEN_A"));
}
/**
* {@code SSH_AUTH_SOCK} is a live handle to the operator's ssh-agent, never a value — so it must
* stay excluded from the derived set even when the operator lists it under {@code allow:} for an
* unrelated reason. It is governed ONLY by {@code memberCredentials.sshAuthSock}, applied
* separately by the caller ({@code HerdrPeerLauncher}).
*/
@Test
void sshAuthSockInMemberCredentialsAllowIsStillExcluded() {
Set<String> derived = MemberEnvAllowList.derive(List.of(), Set.of("SSH_AUTH_SOCK", "OTHER_NAME"));
assertFalse(derived.contains("SSH_AUTH_SOCK"),
"SSH_AUTH_SOCK must never ride in on the generic allow: list");
assertTrue(derived.contains("OTHER_NAME"), "other allow: names are unaffected");
}
@Test
void configuredBrokerAndCoordinatorUriEnvNamesAreExcludedEvenWhenAllowed() {
FleetConfig config = config("BROKER_CONNECTION_URI", "COORDINATOR_CONNECTION_URI");
Set<String> excluded = MemberEnvAllowList.brokerUriEnvNames(config);
Set<String> derived = MemberEnvAllowList.derive(List.of(),
Set.of("BROKER_CONNECTION_URI", "COORDINATOR_CONNECTION_URI", "OTHER_NAME"), excluded);
assertFalse(derived.contains("BROKER_CONNECTION_URI"),
"broker.uriEnv is secret-bearing and must not ride in on allow:");
assertFalse(derived.contains("COORDINATOR_CONNECTION_URI"),
"coordinator.uriEnv has the same inline-password shape");
assertTrue(derived.contains("OTHER_NAME"), "unrelated allow: entries are unaffected");
}
@Test
void absentOrBlankBrokerUriEnvAddsNoExclusions() {
assertTrue(MemberEnvAllowList.brokerUriEnvNames(config(null, null)).isEmpty());
assertTrue(MemberEnvAllowList.brokerUriEnvNames(config(" ", "")).isEmpty());
}
private static FleetConfig config(String brokerUriEnv, String coordinatorUriEnv) {
return new FleetConfig(null, null, null, Map.of(), null, null, null, null, null,
new FleetConfig.Broker(null, brokerUriEnv, null), null, null, null, null, null,
null, null, null, null,
new FleetConfig.Coordinator(null, coordinatorUriEnv, null, null)).withDefaults();
}
/** {@code LC_*} categories are infrastructure by prefix; everything else needs an exact match. */
@Test
void keepsMatchesExactlyPlusTheLocalePrefixRule() {
@@ -1,9 +1,5 @@
package dev.ltms.fleet.member;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.config.FleetConfig;
@@ -18,13 +14,9 @@ import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.PosixFileAttributeView;
import java.nio.file.attribute.PosixFilePermissions;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ExecutorService;
@@ -33,7 +25,6 @@ import java.util.concurrent.Future;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* The opencode adapter's launch build: a file-based MCP mount + reply-charter instructions (no
@@ -343,7 +334,8 @@ class OpenCodeLauncherTest {
assertNull(handle.agentSessionId(), "no record yet → null, not a spawn-time block");
// Once the record appears (here: same cwd), lazy discovery resolves it — the handle's
// session id matches its own worktree, not another's.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_resolved", "/work/dir", 1000L);
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "p1", "ses_a.json",
"ses_resolved", "/work/dir", 1000L);
assertEquals("ses_resolved", handle.agentSessionId(),
"agentSessionId() re-scans and picks up a record that has since been written");
}
@@ -619,196 +611,4 @@ class OpenCodeLauncherTest {
assertTrue(json.path("mcp").path("intellij").isMissingNode(),
"no IDE server when ideMcpUrl is unset");
}
// --- fleetd #219: config root + discovery root under memberHerdrSocket ------------------------
/** A config with {@code memberHerdrSocket:} set, and optionally {@code worktreeRoot:}/{@code worktreeGroup:}. */
private static FleetConfig configWithMemberHerdrSocket(String worktreeRoot, String worktreeGroup) {
return new FleetConfig(
null, // bind
null, // herdrSocket
"/tmp/other-user.sock", // memberHerdrSocket
Map.of(), // profiles
null, // guard
worktreeRoot, // worktreeRoot
null, // lifecycle
null, // spawnReadyTimeoutMs
null, // spawnReadyPollMs
null, // broker
null, // primary
null, // fleet
null, // leadHeartbeat
null, // health
null, // placement
null, // auth
null, // configReload
null, // quarantineCooldownSeconds
null, // memberCredentials
null, // coordinator
worktreeGroup, // worktreeGroup
null // memberLoginShell
).withDefaults();
}
private static OpenCodeLauncher serviceWithConfig(FakeHerdr herdr, Path configRoot, Path discoveryRoot,
FleetConfig.Profile cfg, Supplier<FleetConfig> config) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
0, System::currentTimeMillis, () -> { }, configRoot, discoveryRoot, null, null, config);
}
/** The current process's own primary group — resolvable on whatever host runs this test. */
private static String currentUserGroup() throws IOException {
PosixFileAttributeView view = Files.getFileAttributeView(Path.of("."), PosixFileAttributeView.class);
assumeTrue(view != null, "this host's filesystem does not support POSIX group ownership");
return view.readAttributes().group().getName();
}
/**
* fleetd #219 site 1, acceptance criterion 1: with {@code memberHerdrSocket} configured and both
* {@code worktreeRoot}/{@code worktreeGroup} set, the generated {@code opencode.json} directory
* lives under {@code worktreeRoot} — NEVER under the injected {@code configRoot} (standing in for
* {@code java.io.tmpdir}, fleetd's own 0700 temp dir, unreadable by the member's different OS
* user) — and is shared read-only with the group via the SAME mechanism (fleetd #213's {@link
* EnvAllowListScrub#shareWithGroup}) the ZDOTDIR scrub uses.
*/
@Test
void memberHerdrSocketWithWorktreeRootAndGroupPutsConfigDirUnderWorktreeRootAndSharesIt(
@TempDir Path configRoot, @TempDir Path worktreeRoot) throws Exception {
String group = currentUserGroup();
FakeHerdr herdr = new FakeHerdr();
serviceWithConfig(herdr, configRoot, configRoot,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null),
() -> configWithMemberHerdrSocket(worktreeRoot.toString(), group)).spawn();
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(cfgPath, "the profile still needs a config file");
Path cfgFile = Path.of(cfgPath);
Path dir = cfgFile.getParent();
assertEquals(worktreeRoot.toAbsolutePath().normalize(), dir.getParent(),
"the generated directory's parent must be worktreeRoot, not the injected configRoot "
+ "standing in for java.io.tmpdir — got parent " + dir.getParent());
assertFalse(dir.startsWith(configRoot),
"the generated directory must NOT be created under configRoot when memberHerdrSocket "
+ "is configured: " + dir);
assertEquals("rwxr-x---", PosixFilePermissions.toString(Files.getPosixFilePermissions(dir)),
"the directory must be group-traversable+readable, owner-only writable");
assertEquals("rw-r-----", PosixFilePermissions.toString(Files.getPosixFilePermissions(cfgFile)),
"opencode.json must be group-readable, never group-writable");
}
/**
* fleetd #219 site 1, acceptance criterion 2: with {@code memberHerdrSocket} configured but
* NEITHER {@code worktreeRoot} nor {@code worktreeGroup} set, the launcher must refuse the spawn
* rather than write a config under {@code java.io.tmpdir} the member cannot read — that member
* would occupy a pane and never become deliverable, with nothing pointing at the real cause.
*/
@Test
void memberHerdrSocketWithoutWorktreeRootOrGroupRefusesTheSpawn(@TempDir Path configRoot) {
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = serviceWithConfig(herdr, configRoot, configRoot,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null),
() -> configWithMemberHerdrSocket(null, null));
IllegalStateException ex = assertThrows(IllegalStateException.class,
() -> launcher.spawn(new SpawnRequest(null, null, null)),
"a missing worktreeRoot/worktreeGroup must refuse the spawn, not write an unreadable config");
assertTrue(ex.getMessage().contains("worktreeRoot"),
"the refusal must name the missing config key — got: " + ex.getMessage());
assertFalse(herdr.called("tab.create"),
"the spawn must be refused BEFORE any pane is created — got calls: " + herdr.calls);
}
/**
* fleetd #219 site 1, acceptance criterion 2 (the other missing half): {@code worktreeRoot} set
* but {@code worktreeGroup} missing must ALSO refuse — either one alone is not enough to
* guarantee the member's OS user can read the directory.
*/
@Test
void memberHerdrSocketWithWorktreeRootButNoGroupRefusesTheSpawn(
@TempDir Path configRoot, @TempDir Path worktreeRoot) {
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = serviceWithConfig(herdr, configRoot, configRoot,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null),
() -> configWithMemberHerdrSocket(worktreeRoot.toString(), null));
IllegalStateException ex = assertThrows(IllegalStateException.class,
() -> launcher.spawn(new SpawnRequest(null, null, null)));
assertTrue(ex.getMessage().contains("worktreeGroup"),
"worktreeRoot alone is not enough — got: " + ex.getMessage());
}
/**
* fleetd #219 site 1, acceptance criterion 3: with {@code memberHerdrSocket} ABSENT — even when
* a live, non-null {@code config} supplier is threaded through (not merely {@code config == null},
* which every other test in this file already exercises) — the generated directory must still
* land directly under the injected {@code configRoot}, byte-identical to before this fix.
*/
@Test
void memberHerdrSocketAbsentStaysUnderConfigRootEvenWithALiveConfigSupplier(@TempDir Path configRoot)
throws Exception {
FakeHerdr herdr = new FakeHerdr();
FleetConfig config = new FleetConfig(null, null, null, Map.of(), null, null, null, null, null,
null, null, null, null, null, null, null, null, null, null, null, null, null).withDefaults();
serviceWithConfig(herdr, configRoot, configRoot,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null), () -> config)
.spawn();
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(cfgPath);
assertTrue(Path.of(cfgPath).startsWith(configRoot),
"with memberHerdrSocket absent, the config directory must still be created directly "
+ "under configRoot, unchanged from before this fix");
}
/**
* fleetd #219 site 2: under {@code memberHerdrSocket}, opencode session discovery must be
* declared unavailable rather than silently scanning fleetd's own {@code discoveryRoot} — which,
* under this config key, is NOT where the member's opencode actually writes its session
* database. This test proves the gate is real, not merely "no record yet": a matching record IS
* written to {@code discoveryRoot} (the exact fixture {@link
* #theHandleDiscoversTheSessionIdForTheWorkersCwdOnlyAfterItAppears} proves discovery would
* otherwise find), and {@code agentSessionId()} must still return {@code null} — proving the
* gate, not a coincidental absence of data, is what produced the null. One WARN is also logged,
* exactly once even across repeated calls.
*/
@Test
void discoveryIsUnavailableUnderMemberHerdrSocketEvenWhenARecordExists(
@TempDir Path configRoot, @TempDir Path worktreeRoot, @TempDir Path discRoot) throws Exception {
String group = currentUserGroup();
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_should_be_hidden", "/work/dir", 1000L);
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = new OpenCodeLauncher(new AgentControl(herdr),
new WorkspaceControl(herdr), Map.of("gemini", opencodeCfg(null, null, null)),
"gemini", _ -> null, 0, System::currentTimeMillis, () -> { }, configRoot, discRoot,
null, null, () -> configWithMemberHerdrSocket(worktreeRoot.toString(), group));
Logger logger = (Logger) LoggerFactory.getLogger(OpenCodeLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
PeerHandle handle;
try {
handle = launcher.spawn(new SpawnRequest(null, "/work/dir", null));
assertNull(handle.agentSessionId(),
"memberHerdrSocket configured: discovery must stay unavailable even though a "
+ "matching record exists in discoveryRoot");
assertNull(handle.agentSessionId(), "the gate must hold on a second call too");
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
List<String> warnings = appender.list.stream()
.filter(e -> e.getLevel() == Level.WARN)
.map(ILoggingEvent::getFormattedMessage)
.toList();
assertEquals(1, warnings.size(),
"exactly one WARN across two agentSessionId() calls — got: " + warnings);
assertTrue(warnings.get(0).contains("memberHerdrSocket"),
"the WARN must name memberHerdrSocket as the reason — got: " + warnings.get(0));
}
}
@@ -5,89 +5,68 @@ import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.sql.Connection;
import java.sql.DriverManager;
import java.sql.PreparedStatement;
import java.sql.SQLException;
import java.sql.Statement;
import java.nio.file.attribute.FileTime;
import static org.junit.jupiter.api.Assertions.*;
/**
* {@link OpenCodeSessionDiscovery} matches an opencode session row by the worker's cwd (its
* {@code directory}) against opencode's {@code opencode.db} SQLite database. These tests build a
* SYNTHETIC database themselves, in a JUnit temp directory — never the operator's real
* {@code ~/.local/share/opencode/opencode.db}, which a live opencode process may be writing.
* {@link OpenCodeSessionDiscovery} matches an opencode session record by the worker's cwd (its
* {@code directory}) against opencode's on-disk storage. These tests populate a TEMP storage root
* themselves — never the operator's real {@code ~/.local/share/opencode}.
*/
class OpenCodeSessionDiscoveryTest {
/**
* Create {@code <root>/opencode.db} with a minimal {@code session} table (just the columns
* {@link OpenCodeSessionDiscovery} reads: {@code id}, {@code directory}, {@code time_updated})
* and insert one row. Static so {@link OpenCodeLauncherTest} can reuse it.
* Write a session record {@code {"id":..., "directory":...}} under
* {@code <root>/session/<projectID>/<fileName>} and stamp it with a known last-modified time,
* so "most recently modified wins" is deterministic. Static so the launcher test can reuse it.
*/
static void writeRecord(Path root, String id, String directory, long timeUpdated) throws Exception {
Path db = root.resolve("opencode.db");
try (Connection connection = DriverManager.getConnection("jdbc:sqlite:" + db)) {
try (Statement statement = connection.createStatement()) {
statement.execute("CREATE TABLE IF NOT EXISTS session ("
+ "id TEXT PRIMARY KEY, directory TEXT, time_updated INTEGER)");
}
// Bound parameters, not string interpolation: the class under test uses a
// PreparedStatement, and a hand-escaped INSERT here is a pattern someone copies out.
try (PreparedStatement insert = connection.prepareStatement(
"INSERT INTO session (id, directory, time_updated) VALUES (?, ?, ?)")) {
insert.setString(1, id);
insert.setString(2, directory);
insert.setLong(3, timeUpdated);
insert.executeUpdate();
}
}
static void writeRecord(Path root, String projectId, String fileName, String id,
String directory, long lastModifiedEpochMillis) throws Exception {
Path dir = root.resolve("session").resolve(projectId);
Files.createDirectories(dir);
Path file = dir.resolve(fileName);
Files.writeString(file, "{\"id\":\"" + id + "\",\"directory\":\"" + directory
+ "\",\"projectID\":\"" + projectId + "\",\"version\":\"1.1.31\"}");
Files.setLastModifiedTime(file, FileTime.fromMillis(lastModifiedEpochMillis));
}
@Test
void findsTheRowWhoseDirectoryEqualsTheCwd(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L);
writeRecord(root, "ses_bbb", "/w/b", 2000L);
void findsTheRecordWhoseDirectoryEqualsTheCwd(@TempDir Path root) throws Exception {
writeRecord(root, "p1", "ses_a.json", "ses_aaa", "/w/a", 1000L);
writeRecord(root, "p2", "ses_b.json", "ses_bbb", "/w/b", 2000L);
assertEquals("ses_bbb", new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/b"),
"the row whose directory equals the cwd is the one found");
"the record whose directory equals the cwd is the one found");
assertEquals("ses_aaa", new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"));
}
@Test
void aNonMatchingDirectoryYieldsNullRatherThanAMismatch(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L);
writeRecord(root, "p1", "ses_a.json", "ses_aaa", "/w/a", 1000L);
assertNull(new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/other"),
"no row for this cwd yet → null, not a wrong session");
"no record for this cwd yet → null, not a wrong session");
}
@Test
void prefersTheMostRecentlyUpdatedRowWhenSeveralMatch(@TempDir Path root) throws Exception {
writeRecord(root, "ses_old", "/w/a", 1000L);
writeRecord(root, "ses_new", "/w/a", 5000L);
void prefersTheMostRecentlyModifiedRecordWhenSeveralMatch(@TempDir Path root) throws Exception {
writeRecord(root, "p1", "old.json", "ses_old", "/w/a", 1000L);
writeRecord(root, "p2", "new.json", "ses_new", "/w/a", 5000L);
assertEquals("ses_new", new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"),
"the row with the highest time_updated for the cwd wins");
"the freshest record for the cwd wins");
}
@Test
void aMissingDatabaseYieldsNullWithoutThrowing(@TempDir Path root) {
// No opencode.db at all under the root.
void aMissingOrEmptyStorageRootYieldsNullWithoutThrowing(@TempDir Path root) throws Exception {
// Missing: no session dir at all under the root.
assertNull(new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"));
}
@Test
void anEmptyDatabaseYieldsNullWithoutThrowing(@TempDir Path root) throws Exception {
Path db = root.resolve("opencode.db");
try (Connection connection = DriverManager.getConnection("jdbc:sqlite:" + db);
Statement statement = connection.createStatement()) {
statement.execute("CREATE TABLE session (id TEXT PRIMARY KEY, directory TEXT, "
+ "time_updated INTEGER)");
}
assertNull(new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"));
// Present but empty: a session dir with nothing in it produces no match, not a throw.
Path emptyRoot = root.resolve("empty");
Files.createDirectories(emptyRoot.resolve("session"));
assertNull(new OpenCodeSessionDiscovery(emptyRoot).sessionIdForDirectory("/w/a"));
}
@Test
@@ -97,41 +76,15 @@ class OpenCodeSessionDiscoveryTest {
assertNull(discovery.sessionIdForDirectory(" "));
}
/**
* The one line standing between fleetd and writing the operator's live {@code opencode.db} —
* 841MB, with a running opencode writing it — is {@code config.setReadOnly(true)} in
* {@link OpenCodeSessionDiscovery#openReadOnly()}. Delete it and every other test in this class
* still passes, so this is the test that guards it.
*
* <p>It asks the connection to write, and requires a refusal. The obvious alternative — make
* the database file unwritable and check the read still works — proves nothing: SQLite silently
* downgrades a read-write open of an unwritable file to read-only, so that test passes either
* way. It was tried and watched pass with the flag removed.
*/
@Test
void theDatabaseIsOpenedReadOnly(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L);
void aMalformedRecordIsSkippedRatherThanFatal(@TempDir Path root) throws Exception {
// A record that fails to parse must not abort the scan of its siblings.
Path dir = root.resolve("session").resolve("p1");
Files.createDirectories(dir);
Files.writeString(dir.resolve("broken.json"), "{not valid json");
writeRecord(root, "p1", "good.json", "ses_good", "/w/a", 1000L);
try (Connection connection = new OpenCodeSessionDiscovery(root).openReadOnly();
Statement statement = connection.createStatement()) {
SQLException refused = assertThrows(SQLException.class,
() -> statement.executeUpdate("INSERT INTO session (id, directory, time_updated) "
+ "VALUES ('ses_zzz', '/w/z', 1)"),
"the connection must REFUSE a write — opencode is writing this database live");
assertTrue(refused.getMessage().toLowerCase().contains("readonly")
|| refused.getMessage().toLowerCase().contains("read-only"),
"the refusal must be about read-only, not some other error: " + refused.getMessage());
}
}
@Test
void aCorruptDatabaseFileYieldsNullWithoutThrowing(@TempDir Path root) throws Exception {
// A file at opencode.db that is not a SQLite database at all — the open/query must fail
// safe, never fatal to a spawn.
Path db = root.resolve("opencode.db");
Files.writeString(db, "this is not a sqlite database");
assertNull(new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"),
"an unreadable database resolves to null, not an exception");
assertEquals("ses_good", new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"),
"an unreadable record is skipped; a later valid one still matches");
}
}
@@ -1,149 +0,0 @@
package dev.ltms.fleet.msg;
import com.rabbitmq.client.AMQP;
import com.rabbitmq.client.Channel;
import com.rabbitmq.client.Connection;
import com.rabbitmq.client.DeliverCallback;
import com.rabbitmq.client.Delivery;
import com.rabbitmq.client.Envelope;
import org.junit.jupiter.api.Test;
import java.io.IOException;
import java.lang.reflect.InvocationHandler;
import java.lang.reflect.Proxy;
import java.nio.charset.StandardCharsets;
import java.util.ArrayDeque;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.concurrent.atomic.AtomicInteger;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
/**
* Pins the manual-ack prefetch behaviour without a broker. The fake channel models a broker that
* sends no more than its QoS window of unacked deliveries. If {@link AmqpReplyInbox} starts acking
* messages while it adds them to {@code held}, this test drains the whole fake queue instead.
*/
class AmqpReplyInboxPrefetchTest {
@Test
void unackedDeliveriesKeepTheHeldBacklogAtThePrefetchWindow() {
int prefetch = 3;
int published = 8;
PrefetchBroker broker = new PrefetchBroker();
for (int i = 0; i < published; i++) {
broker.publish("m" + i, "payload " + i);
}
try (AmqpReplyInbox inbox = new AmqpReplyInbox(connectionFor(broker.channel()), prefetch)) {
inbox.own("worker");
assertEquals(prefetch, inbox.peek("worker").size(),
"held messages must stop at the unacked prefetch window");
assertEquals(published - prefetch, broker.queuedCount(),
"messages beyond the window must remain on the broker");
assertEquals(0, broker.ackCount(), "receipt must not ack a held message");
inbox.ack("worker", "m0");
assertEquals(prefetch, inbox.peek("worker").size(),
"one caller ack frees exactly one slot for the broker");
assertEquals(published - prefetch - 1, broker.queuedCount(),
"only one queued message may enter after one caller ack");
assertEquals(1, broker.ackCount(), "only the caller ack may reach the broker");
}
}
private static Connection connectionFor(Channel consumeChannel) {
Channel publishChannel = (Channel) Proxy.newProxyInstance(
AmqpReplyInboxPrefetchTest.class.getClassLoader(), new Class<?>[] {Channel.class},
(proxy, method, args) -> defaultValue(method.getReturnType()));
AtomicInteger channelCalls = new AtomicInteger();
InvocationHandler handler = (proxy, method, args) -> {
if (method.getName().equals("createChannel") && (args == null || args.length == 0)) {
return channelCalls.getAndIncrement() == 0 ? consumeChannel : publishChannel;
}
return defaultValue(method.getReturnType());
};
return (Connection) Proxy.newProxyInstance(AmqpReplyInboxPrefetchTest.class.getClassLoader(),
new Class<?>[] {Connection.class}, handler);
}
private static final class PrefetchBroker implements InvocationHandler {
private final ArrayDeque<Delivery> queued = new ArrayDeque<>();
private final Map<Long, Delivery> unacked = new LinkedHashMap<>();
private DeliverCallback consumer;
private int prefetch;
private int acks;
private long nextTag = 1;
Channel channel() {
return (Channel) Proxy.newProxyInstance(AmqpReplyInboxPrefetchTest.class.getClassLoader(),
new Class<?>[] {Channel.class}, this);
}
void publish(String msgId, String content) {
queued.add(new Delivery(new Envelope(nextTag++, false, "", ""),
new AMQP.BasicProperties.Builder().messageId(msgId).build(),
content.getBytes(StandardCharsets.UTF_8)));
}
int queuedCount() {
return queued.size();
}
int ackCount() {
return acks;
}
@Override
public Object invoke(Object proxy, java.lang.reflect.Method method, Object[] args) throws IOException {
switch (method.getName()) {
case "basicQos" -> {
prefetch = (int) args[0];
return null;
}
case "basicConsume" -> {
assertFalse((boolean) args[1], "the inbox consumer must use manual acknowledgements");
consumer = (DeliverCallback) args[2];
deliverAvailable();
return "consumer";
}
case "basicAck" -> {
unacked.remove((long) args[0]);
acks++;
deliverAvailable();
return null;
}
default -> {
return defaultValue(method.getReturnType());
}
}
}
private void deliverAvailable() throws IOException {
while (consumer != null && unacked.size() < prefetch && !queued.isEmpty()) {
Delivery delivery = queued.removeFirst();
unacked.put(delivery.getEnvelope().getDeliveryTag(), delivery);
consumer.handle("consumer", delivery);
}
}
}
private static Object defaultValue(Class<?> type) {
if (!type.isPrimitive() || type == void.class) {
return null;
}
if (type == boolean.class) {
return false;
}
if (type == long.class) {
return 0L;
}
if (type == int.class) {
return 0;
}
return 0;
}
}
@@ -33,17 +33,8 @@ class MessageServiceTest {
private final FakeHerdr herdr = new FakeHerdr().readText("BUILD GREEN: 391 files");
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
/**
* fleetd#164: this fixture drives a delivery and its completion back-to-back with no real time
* between them, so the real clock would trip {@link CompletionResolver#MIN_TURN_NANOS} on every
* completion-fallback test here. An ever-advancing fake clock stands in for the model-latency and
* herdr round-trips a real turn would spend, so each delivery-then-resolve pair still lands
* outside the floor.
*/
private final java.util.concurrent.atomic.AtomicLong resolverClock = new java.util.concurrent.atomic.AtomicLong();
private final CompletionResolver completion =
new CompletionResolver(agents, rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none(),
() -> resolverClock.addAndGet(CompletionResolver.MIN_TURN_NANOS + 1));
new CompletionResolver(agents, rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
private final Injector injector = new Injector(agents, completion);
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
private final MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
@@ -86,28 +77,6 @@ class MessageServiceTest {
assertTrue(reply.completed(), "a scraped completion still counts as completed");
}
@Test
void backendErrorScrapeThroughMessageServiceFailsInsteadOfBecomingReplyText() throws Exception {
// fleetd#164 (part 2 addendum): a scrape that reads cleanly but is only the backend's own
// rejection (e.g. an HTTP 400) must reach the caller as WORKER_FAILED, not as a completed
// reply whose text happens to be the error line.
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
herdr.readText("$ prompt");
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
herdr.readText("⏺ API Error: 400 invalid request body");
injector.onStatus(T, AgentStatus.IDLE);
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome(),
"a backend rejection must use the caller's failure outcome, not a completed reply");
assertFalse(reply.completed(), "plain backend errors are never fallback reply content");
assertTrue(reply.text().contains("API Error: 400 invalid request body"),
"the visible backend error is carried as the failure reason: " + reply.text());
}
@Test
void explicitFleetReplyResolvesAsReplied() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
@@ -688,32 +657,6 @@ class MessageServiceTest {
assertFailedTicket(third, "agent target term_a not found");
}
// --- #137 follow-up: abandon() must not guess when more than one task is open ---------------
//
// A test combining a genuine stranded reply (hasStrandedReply(T)==true) with two simultaneously
// open matching tasks was attempted here and removed after investigation showed the combination
// is not reachable through the public API today, not merely hard to time right:
//
// This class has exactly two call sites that ever hold a target's entry in the session-lock map
// (send() and answer()), and both open a Rendezvous waiter for that same target as the first thing
// they do after acquiring the lock, holding lock and waiter together for their whole critical
// section. So "the lock is held" and "a live waiter is open" are the same fact throughout this
// class. reply()'s fast path always resolves a currently-open waiter directly instead of
// stranding — so a strand can only be created while NO task is accepted (lock free), and the
// instant the lock is next taken (by any parked matching task's own send(), the moment it is
// scheduled), that acceptance clears strandedReplies again (see send()'s CB-640 comment) before
// abandon() can ever observe both facts together. Confirmed empirically too: an earlier version of
// this test stranded a reply, then created an "accepted" task (awaitWaiting()) followed by a
// "parked" one — and the accepted task's own acceptance silently cleared the strand it was
// supposed to be racing against, so the parked task came back WORKER_FAILED instead of DONE, not
// because the fix was missing but because the test's premise could not be constructed.
//
// The reachable half — matching.size() >= 2 alone, no strand — is exactly what
// abandonFailsEveryPendingAsyncTicketForTheReleasedTarget already covers (all fail, none guess).
// The oldest-wins code in abandon() stays as defence in depth (see its own javadoc) against a
// regression that would make the conjunction reachable, e.g. clearing strandedReplies on a
// narrower condition than "any acceptance" — not because this suite exercises it today.
@Test
void abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
@@ -736,68 +679,6 @@ class MessageServiceTest {
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
// --- #137: a fleet_ask round-trip must not orphan the ticket's own reply -------------------
//
// The primary's fleet_send{turnId} answer call is itself bounded (a real MCP call, capped well
// under a minute) — far shorter than a resumed turn can genuinely take to finish real work. These
// drive the exact real delegation path (async send -> worker asks -> primary answers -> primary's
// own wait gives up -> worker's real fleet_reply arrives afterwards) rather than calling a reply
// sink directly, since the bug is specifically about which sink the resumed turn's reply reaches.
@Test
void aReplyAfterAnswerTimesOutStillCompletesTheAsyncTicket() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
// The primary answers, but its own bounded wait for the worker's resumed turn is short and
// expires before the worker (still genuinely working) gets back to it.
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150);
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome(),
"the primary's own bounded wait gives up before the worker finishes resuming");
// The worker keeps working past that window and only now calls fleet_reply.
assertTrue(messages.reply(T, "PR opened: https://example/pulls/42"));
MessageService.TaskView done = awaitTicketPhase(ticket, MessageService.Phase.DONE);
assertEquals("PR opened: https://example/pulls/42", done.reply(),
"fleet_poll{ticket} must return the worker's real reply, not stay pending forever");
assertEquals("reply", done.replySource());
assertFalse(messages.hasStrandedReply(T),
"the reply completed its own ticket directly and never touched the inbox");
}
@Test
void fleetStopAfterAnOrphanedReplyDoesNotFailTheTicket() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150);
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome());
assertTrue(messages.reply(T, "PR opened: https://example/pulls/42"));
// fleet_stop tears the worker's session down right after the reply landed — this must never
// report the misleading "the worker session was released before it replied": a reply is
// exactly what happened.
assertFalse(messages.abandon(T, "the worker session was released before it replied"),
"a reply already arrived, so nothing here is a genuine failure");
MessageService.TaskView view = awaitTicketPhase(ticket, MessageService.Phase.DONE);
assertEquals("PR opened: https://example/pulls/42", view.reply());
}
@Test
void unansweredAsyncQuestionReturnsTheTicketToPendingAndReleasesItsTarget() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
@@ -1167,74 +1048,6 @@ class MessageServiceTest {
}
}
/**
* #197: the ticket TTL must run from COMPLETION, not from creation.
*
* <p>It used to compare the cutoff against {@code createdNanos}, so the real window to collect a
* reply was {@code TTL minus however long the task ran}. A delegation that ran longer than the
* TTL was already past the cutoff the moment it finished, so the very next prune destroyed its
* reply — and the reply lives only in the task's future, so nothing could get it back. That is
* the normal case for real work here, not an edge case: three workers in one session ran well
* past ten minutes and two of their complete reports were lost this way.
*
* <p>The task below runs for longer than the whole TTL before it replies, which is exactly the
* shape that used to lose everything. Remove the fix and this fails: {@code poll} returns
* {@code null} because the ticket was pruned on arrival.
*/
@Test
void aTaskRunningLongerThanTheTtlStillKeepsItsReport() throws Exception {
java.util.concurrent.atomic.AtomicLong clock = new java.util.concurrent.atomic.AtomicLong(1_000_000_000L);
try (var wiring = wireWithPushLoop(1, 50, clock::get)) {
String slow = wiring.service().sendAsync(T, "a task that takes longer than the TTL");
awaitWaiting();
injectDelivery();
// The worker is still working, and has been for longer than the entire TTL. Nothing may
// be pruned yet — the ticket has not finished, so there is no report to keep or lose.
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(30));
// Only now does it reply. Under the old clock this reply was born already expired.
assertTrue(rendezvous.resolve(T, "the long report"));
awaitTicketPhaseOn(wiring.service(), slow, MessageService.Phase.DONE);
// A second delegation runs pruneTerminalTickets before it returns.
wiring.service().sendAsync(T, "an unrelated second task");
MessageService.TaskView view = wiring.service().poll(slow);
assertNotNull(view, "a ticket that completed just now must survive the prune, however "
+ "long its task ran — the TTL is the window to COLLECT the report, not the "
+ "budget for producing it");
assertEquals(MessageService.Phase.DONE, view.phase());
assertEquals("the long report", view.reply(),
"the worker's actual report must still be there, not just the ticket");
}
}
/**
* The other half of #197: the TTL must still bound {@code tasks}. Measuring from completion
* would be a leak if a finished ticket were then kept forever, so this pins the eviction that
* still has to happen — the same ticket, left uncollected for longer than the TTL AFTER it
* finished, is gone.
*/
@Test
void aFinishedTicketIsStillPrunedOnceTheTtlPassesSinceItFinished() throws Exception {
java.util.concurrent.atomic.AtomicLong clock = new java.util.concurrent.atomic.AtomicLong(1_000_000_000L);
try (var wiring = wireWithPushLoop(1, 50, clock::get)) {
String done = wiring.service().sendAsync(T, "a quick task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "quick result"));
awaitTicketPhaseOn(wiring.service(), done, MessageService.Phase.DONE);
// Nobody collected it, and the TTL has now passed since it FINISHED.
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
wiring.service().sendAsync(T, "an unrelated second task");
assertNull(wiring.service().poll(done),
"the TTL must still evict an uncollected finished ticket, or tasks grows forever");
}
}
private MessageService.TaskView awaitTicketPhaseOn(MessageService svc, String ticket,
MessageService.Phase phase) throws Exception {
long deadline = System.currentTimeMillis() + 3000;
@@ -32,7 +32,6 @@ import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.function.Predicate;
import static org.junit.jupiter.api.Assertions.*;
@@ -65,11 +64,6 @@ class FleetAppTest {
}
private int start(FakeHerdr herdr, String workerBaseUrl, Set<String> allow, String placement, Worktrees worktrees) {
return start(herdr, workerBaseUrl, allow, placement, worktrees, ignored -> false);
}
private int start(FakeHerdr herdr, String workerBaseUrl, Set<String> allow, String placement,
Worktrees worktrees, Predicate<String> deliverable) {
FleetConfig.Profile wcfg = new FleetConfig.Profile(
"ltms-local", workerBaseUrl, "coder", null, "FLEETD_WORKER_TOKEN", null,
placement, "fleet", "worker: {profile} #{n}", null, null, null);
@@ -90,8 +84,7 @@ class FleetAppTest {
// it directly so the inbox contract holds for those endpoints.
inbox.own("term_a");
MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
app = new FleetApp(herdr, workers, sessions, messages, this.presence, null,
null, null, id -> this.presence.isPresent(id) || deliverable.test(id))
app = new FleetApp(herdr, workers, sessions, messages, this.presence, null)
.build().start("127.0.0.1", 0);
return app.port();
}
@@ -218,17 +211,9 @@ class FleetAppTest {
HttpResponse<String> res = req(port, "GET", "/members");
assertEquals(200, res.statusCode());
JsonNode body = mapper.readTree(res.body());
// fleetd #199: "members" is the canonical key. The endpoint is /members, so a caller that
// reads "members" must not see an empty fleet. "workers" is kept only as a deprecated alias
// and must carry the same rows — assert both, or the alias can silently drift.
JsonNode members = body.get("members");
assertNotNull(members, "GET /members must return its rows under \"members\"");
assertEquals(1, members.size());
JsonNode workers = body.get("workers");
assertNotNull(workers, "the deprecated \"workers\" alias is still emitted");
assertEquals(members, workers, "the alias must carry the same rows as \"members\"");
JsonNode w = members.get(0);
JsonNode workers = mapper.readTree(res.body()).get("workers");
assertEquals(1, workers.size());
JsonNode w = workers.get(0);
assertEquals(spawned.get("terminalId").asText(), w.get("sessionId").asText());
assertEquals(paneId, w.get("paneId").asText());
assertEquals("ltms-local", w.get("profile").asText());
@@ -499,16 +484,6 @@ class FleetAppTest {
assertTrue(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
}
@Test
void sessionStatusReportsRegisteredLeadAsReady() throws Exception {
Map<String, String> leads = Map.of("term_lead", "terra");
int port = start(new FakeHerdr(), "http://gx00.gw:8000", Set.of("gx00.gw"), "tab",
new GitWorktrees(), leads::containsKey);
JsonNode body = mapper.readTree(req(port, "GET", "/sessions/term_lead/status").body());
assertTrue(body.get("ready").asBoolean());
}
/**
* CB-582: a lead polling {@code GET /sessions/{id}/status} on its normal cadence — not the
* ticket-scoped {@code /tasks/{ticket}} — must also see a worker's open async {@code fleet_ask}
@@ -1,150 +0,0 @@
package dev.ltms.fleet.rest;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-185: with a router split across two herdr daemons (lead + {@code memberHerdrSocket}),
* {@link FleetApp#healthz} must require BOTH daemons to answer and {@link FleetApp#sessions}
* (which the {@code GET /sessions} route calls) must merge workspaces from both — the bug this
* guards against had {@code FleetApp} constructed with the raw lead-only client, so a down member
* daemon was invisible behind a green {@code /healthz} (every spawn then fails) and every member
* workspace was silently dropped from {@code GET /sessions}.
*
* <p>Builds the real {@link FleetApp} directly (not a hand-rolled stand-in) against only the two
* herdr clients — the other collaborators are unused by the two routes under test here.
*/
class FleetAppTwoDaemonTest {
private final HttpClient http = HttpClient.newHttpClient();
private Javalin app;
@AfterEach
void stop() {
if (app != null) app.stop();
}
private int start(HerdrClient lead, HerdrClient member) {
app = new FleetApp(lead, member, null, null, null, null, null, null, null, ignored -> false)
.build().start("127.0.0.1", 0);
return app.port();
}
private HttpResponse<String> get(int port, String path) throws Exception {
HttpRequest req = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path)).GET().build();
return http.send(req, HttpResponse.BodyHandlers.ofString());
}
@Test
void healthzIsGreenWhenBothDaemonsAnswer() throws Exception {
int port = start(new FakeHerdr(), new FakeHerdr());
assertEquals(200, get(port, "/healthz").statusCode());
}
@Test
void healthzIsDegradedWhenOnlyTheMemberDaemonIsDown() throws Exception {
int port = start(new FakeHerdr(), new FakeHerdr().healthy(false));
HttpResponse<String> res = get(port, "/healthz");
assertEquals(503, res.statusCode(),
"a down MEMBER daemon must not be masked by a healthy lead — every spawn goes "
+ "through the member daemon");
}
@Test
void healthzIsDegradedWhenOnlyTheLeadDaemonIsDown() throws Exception {
int port = start(new FakeHerdr().healthy(false), new FakeHerdr());
assertEquals(503, get(port, "/healthz").statusCode());
}
@Test
void healthzMakesExactlyOneCallWhenLeadAndMemberAreTheSameClient() throws Exception {
// Single-daemon deployment (no memberHerdrSocket) — must be byte-for-byte the old
// behaviour: one ping call, 200 on success.
FakeHerdr shared = new FakeHerdr();
int port = start(shared, shared);
assertEquals(200, get(port, "/healthz").statusCode());
long pings = shared.calls.stream().filter(c -> c.method().equals("ping")).count();
assertEquals(1, pings, "single-daemon deployment must make exactly one ping call");
}
@Test
void sessionsMergesWorkspacesFromBothDaemons() throws Exception {
FakeHerdr lead = new FakeHerdr();
FakeHerdr member = new FakeHerdr().withWorkspace("w9", "member-only-workspace");
int port = start(lead, member);
HttpResponse<String> res = get(port, "/sessions");
assertEquals(200, res.statusCode(), res.body());
assertTrue(res.body().contains("member-only-workspace"),
"GET /sessions must not silently drop the member daemon's workspaces");
}
@Test
void sessionsMakesExactlyOneWorkspaceListCallWhenLeadAndMemberAreTheSameClient() throws Exception {
FakeHerdr shared = new FakeHerdr();
int port = start(shared, shared);
assertEquals(200, get(port, "/sessions").statusCode());
long calls = shared.calls.stream().filter(c -> c.method().equals("workspace.list")).count();
assertEquals(1, calls, "single-daemon deployment must call workspace.list exactly once");
}
// ── CB-185 blocker 2: /healthz must report the MEMBER daemon's protocol too ────────────────
@Test
void healthzReportsBothDaemonsWhenTheirProtocolsDiffer() throws Exception {
FakeHerdr lead = new FakeHerdr().pingReports("0.8.0", 19);
FakeHerdr member = new FakeHerdr().pingReports("0.7.0", 18);
int port = start(lead, member);
HttpResponse<String> res = get(port, "/healthz");
assertEquals(200, res.statusCode(), res.body());
assertTrue(res.body().contains("\"protocol\":19"),
"the herdr key keeps reporting the LEAD's protocol, unchanged: " + res.body());
assertTrue(res.body().contains("\"member\""), "a separate member key is present: " + res.body());
assertTrue(res.body().contains("\"protocol\":18"),
"the member key reports the member daemon's own protocol: " + res.body());
assertTrue(res.body().contains("\"protocolMismatch\":true"),
"a differing protocol is called out explicitly, not left to be spotted by eye: " + res.body());
}
@Test
void healthzReportsBothDaemonsWithNoMismatchWhenProtocolsMatch() throws Exception {
int port = start(new FakeHerdr(), new FakeHerdr());
HttpResponse<String> res = get(port, "/healthz");
assertEquals(200, res.statusCode(), res.body());
assertTrue(res.body().contains("\"member\""), "the member key is present whenever a second daemon "
+ "is configured, even when the protocols happen to agree: " + res.body());
assertFalse(res.body().contains("protocolMismatch"),
"matching protocols must not raise a mismatch flag: " + res.body());
}
@Test
void healthzWithOneDaemonCarriesNoMemberOrMismatchKey() throws Exception {
// The single-daemon deployment (no memberHerdrSocket) must see no change at all beyond the
// historical body: no "member" key, no "protocolMismatch" key. (Map.of()'s own key order is
// JVM-salted regardless of this fix, so this checks content, not exact key order.)
FakeHerdr shared = new FakeHerdr();
int port = start(shared, shared);
HttpResponse<String> res = get(port, "/healthz");
assertEquals(200, res.statusCode());
assertTrue(res.body().contains("\"status\":\"ok\""), res.body());
assertTrue(res.body().contains("\"protocol\":19"), res.body());
assertTrue(res.body().contains("\"version\":\"0.8.0\""), res.body());
assertFalse(res.body().contains("\"member\""), "no second daemon configured, so no member key: " + res.body());
assertFalse(res.body().contains("protocolMismatch"), res.body());
}
}
@@ -30,19 +30,12 @@ public final class FakeWorktrees implements Worktrees {
public record PruneCall(String repoRoot, long minAgeMillis) {
}
public record ShareCall(String repoRoot, String worktreePath) {
}
private final List<AddCall> addCalls = new CopyOnWriteArrayList<>();
private final List<RemoveCall> removeCalls = new CopyOnWriteArrayList<>();
private final List<OverlayCall> overlayCalls = new CopyOnWriteArrayList<>();
private final List<RepoRootCall> repoRootCalls = new CopyOnWriteArrayList<>();
private final List<SnapshotCall> snapshotCalls = new CopyOnWriteArrayList<>();
private final List<PruneCall> pruneCalls = new CopyOnWriteArrayList<>();
private final List<ShareCall> shareCalls = new CopyOnWriteArrayList<>();
/** Tags every {@code overlayParity}/{@code shareWithGroup} call in call order, so a test can
* pin that sharing runs after the overlay copy (fleetd #185 stage 3). */
private final List<String> overlayShareOrder = new CopyOnWriteArrayList<>();
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
private final AtomicLong snapshotSeq = new AtomicLong();
@@ -143,13 +136,6 @@ public final class FakeWorktrees implements Worktrees {
}
overlayCalls.add(new OverlayCall(repoRoot, worktreePath, List.copyOf(overlay),
List.copyOf(copied), List.copyOf(skipped)));
overlayShareOrder.add("overlay:" + worktreePath);
}
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
shareCalls.add(new ShareCall(repoRoot, worktreePath));
overlayShareOrder.add("share:" + worktreePath);
}
@Override
@@ -217,17 +203,4 @@ public final class FakeWorktrees implements Worktrees {
public SnapshotCall lastSnapshot() {
return snapshotCalls.isEmpty() ? null : snapshotCalls.getLast();
}
public List<ShareCall> shareCalls() {
return List.copyOf(shareCalls);
}
public ShareCall lastShare() {
return shareCalls.isEmpty() ? null : shareCalls.getLast();
}
/** Call-order tags ({@code "overlay:<path>"}/{@code "share:<path>"}) — see field javadoc. */
public List<String> overlayShareOrder() {
return List.copyOf(overlayShareOrder);
}
}
@@ -1,23 +1,13 @@
package dev.ltms.fleet.session;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.IThrowableProxy;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.TimeUnit;
@@ -65,21 +55,12 @@ class GitWorktreesTest {
}
private static void git(Path cwd, String... args) throws Exception {
gitOutput(cwd, args);
}
private static String gitOutput(Path cwd, String... args) throws Exception {
List<String> cmd = new java.util.ArrayList<>(List.of("git"));
cmd.addAll(List.of(args));
ProcessBuilder pb = new ProcessBuilder(cmd).directory(cwd.toFile()).redirectErrorStream(true);
pb.environment().put("GIT_CONFIG_GLOBAL", "/dev/null");
pb.environment().put("GIT_CONFIG_SYSTEM", "/dev/null");
pb.environment().put("GIT_TERMINAL_PROMPT", "0");
Process p = pb.start();
Process p = new ProcessBuilder(cmd).directory(cwd.toFile()).redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git timed out: " + String.join(" ", cmd));
assertEquals(0, p.exitValue(), "git " + String.join(" ", args) + " failed:\n" + out);
return out;
}
/** Pending changes to {@code file} in {@code cwd}, empty when git considers it unmodified. */
@@ -224,289 +205,6 @@ class GitWorktreesTest {
"expected an explicitly empty server map, got:\n" + body);
}
/**
* The worktree command shares the primary checkout's config, so this checks the URL git actually
* reads after {@link GitWorktrees#add}, rather than checking only a URL formatting helper.
*/
@Test
void aProvisionedWorktreeUsesACleanHttpsOrigin(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
git(repo, "remote", "add", "origin", "https://synthetic-test-token@git.ltms.dev/akb/kb.git");
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "fleetd-157-safe-origin", "HEAD");
Path worktree = Path.of(wt);
String origin = gitOutput(worktree, "config", "--get", "remote.origin.url").trim();
assertEquals("https://git.ltms.dev/akb/kb.git", origin);
assertFalse(origin.contains("synthetic-test-token"), "provisioned worktree kept user info");
assertFalse(gitOutput(worktree, "remote", "-v").contains("synthetic-test-token"),
"git remote -v exposed user info");
assertFalse(gitOutput(worktree, "config", "--list").contains("synthetic-test-token"),
"git config --list exposed user info");
String helper = gitOutput(worktree, "config", "--worktree", "--get", "credential.helper");
assertTrue(helper.contains("WORKER_GITEA_TOKEN"), "credential helper does not read the member environment");
assertFalse(helper.contains("synthetic-test-token"), "credential helper stored user info");
}
@Test
void worktreeCredentialHelperCompletesWithoutUsingAnInheritedHelper(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
git(repo, "remote", "add", "origin", "https://git.ltms.dev/akb/kb.git");
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "fleetd-157-helper", "HEAD");
Path globalConfig = tmp.resolve("global.gitconfig");
Files.writeString(globalConfig, """
[credential]
helper = !f() { printf 'username=%s\\npassword=%s\\n\\n' operator operator-secret; }; f
""");
ProcessBuilder pb = new ProcessBuilder("git", "credential", "fill")
.directory(Path.of(wt).toFile()).redirectErrorStream(true);
pb.environment().put("GIT_CONFIG_GLOBAL", globalConfig.toString());
pb.environment().put("GIT_CONFIG_SYSTEM", "/dev/null");
pb.environment().put("GIT_TERMINAL_PROMPT", "0");
pb.environment().put("WORKER_GITEA_TOKEN", "synthetic-worker-value");
Process p = pb.start();
p.getOutputStream().write("protocol=https\nhost=git.ltms.dev\n\n".getBytes(StandardCharsets.UTF_8));
p.getOutputStream().close();
String credential = new String(p.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git credential fill timed out");
assertEquals(0, p.exitValue(), "git credential fill failed");
assertTrue(credential.contains("username=git"), "helper did not return its fixed username");
assertTrue(credential.contains("password=synthetic-worker-value"),
"helper did not return the worker token as the password");
assertFalse(credential.contains("operator-secret"), "Git used the inherited global helper");
}
/**
* fleetd #157 follow-up. An SSH origin never consults {@code credential.helper} — the environment
* credential helper set by {@link GitWorktrees#add} is therefore useless when a member's origin is
* SSH, which is exactly this repo's shape. The worktree must instead get a worktree-scoped
* {@code url.<https>.insteadOf <ssh>} rewrite so both fetch and push resolve to HTTPS, while the
* parent checkout — sharing the same repo-level origin config — must resolve the original SSH URL
* completely unchanged. The host/port here (a synthetic {@code forge.example.test:2222}, not
* {@code git.ltms.dev}) proves the rewrite is derived from the origin, not a hardcoded constant.
*/
@Test
void aProvisionedWorktreeRewritesAnSshOriginToHttpsWorktreeScoped(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
git(repo, "remote", "add", "origin", "ssh://git@forge.example.test:2222/acme/proj.git");
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "fleetd-157-ssh-rewrite", "HEAD");
Path worktree = Path.of(wt);
assertEquals("https://forge.example.test/acme/proj.git",
gitOutput(worktree, "remote", "get-url", "origin").trim(),
"worktree fetch URL was not rewritten to HTTPS");
assertEquals("https://forge.example.test/acme/proj.git",
gitOutput(worktree, "remote", "get-url", "--push", "origin").trim(),
"worktree push URL was not rewritten to HTTPS");
// The raw config value is unchanged — only the resolved URL is rewritten, via insteadOf.
assertEquals("ssh://git@forge.example.test:2222/acme/proj.git",
gitOutput(worktree, "config", "--get", "remote.origin.url").trim());
assertEquals("ssh://git@forge.example.test:2222/acme/proj.git",
gitOutput(repo, "remote", "get-url", "origin").trim(),
"the parent checkout's fetch URL must be untouched");
assertEquals("ssh://git@forge.example.test:2222/acme/proj.git",
gitOutput(repo, "remote", "get-url", "--push", "origin").trim(),
"the parent checkout's push URL must be untouched");
}
/** An origin already on HTTPS is left alone — the environment credential helper already covers it. */
@Test
void aProvisionedWorktreeLeavesAnHttpsOriginAlone(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
git(repo, "remote", "add", "origin", "https://git.ltms.dev/akb/kb.git");
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "fleetd-157-https-noop", "HEAD");
Path worktree = Path.of(wt);
assertEquals("https://git.ltms.dev/akb/kb.git",
gitOutput(worktree, "remote", "get-url", "origin").trim());
assertEquals(1, exitCode("git", "-C", wt, "config", "--worktree", "--get-regexp", "^url\\."),
"no url.*.insteadOf rewrite should be added for an already-HTTPS origin");
}
/** Test-local exit-code probe, mirroring {@link GitWorktrees#exitCode} for an assertion the
* production class does not expose. */
private static int exitCode(String... command) throws Exception {
ProcessBuilder pb = new ProcessBuilder(command).redirectErrorStream(true);
pb.environment().put("GIT_CONFIG_GLOBAL", "/dev/null");
pb.environment().put("GIT_CONFIG_SYSTEM", "/dev/null");
pb.environment().put("GIT_TERMINAL_PROMPT", "0");
Process p = pb.start();
p.getInputStream().readAllBytes();
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "command timed out: " + String.join(" ", command));
return p.exitValue();
}
@Test
void provisioningRefusesAWorktreeWhoseOriginStillHasHttpsUserInfo(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
git(repo, "remote", "add", "origin", "https://git.ltms.dev/akb/kb.git");
GitWorktrees worktrees = new GitWorktrees(tmp.resolve("wts").toString(), worktreePath -> {
try {
git(Path.of(worktreePath), "remote", "set-url", "origin",
"https://synthetic-test-token@git.ltms.dev/akb/kb.git");
} catch (Exception e) {
throw new RuntimeException(e);
}
});
WorktreeException error = assertThrows(WorktreeException.class,
() -> worktrees.add(repo.toString(), "fleetd-157-refuse-origin", "HEAD"));
assertEquals("worktree origin contains HTTPS user info; refusing provision", error.getMessage());
}
// ---- CB-189: broader remote-URL coverage — every remote, both fetch and push URLs, any
// non-SSH scheme. Reporting only, additive to the origin/https strip-and-refuse tests above. ----
private Logger reportingLogger;
private ListAppender<ILoggingEvent> reportingAppender;
/** {@link GitWorktrees}'s own logger, captured fresh for each test so assertions never see a
* message left over from a previous test. */
@BeforeEach
void attachReportingLogCapture() {
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
reportingLogger = ctx.getLogger(GitWorktrees.class);
reportingLogger.setLevel(Level.WARN);
reportingAppender = new ListAppender<>();
reportingAppender.setContext(ctx);
reportingAppender.start();
reportingLogger.addAppender(reportingAppender);
}
@AfterEach
void detachReportingLogCapture() {
reportingLogger.detachAppender(reportingAppender);
}
private List<String> capturedMessages() {
return reportingAppender.list.stream().map(ILoggingEvent::getFormattedMessage).toList();
}
/** Asserts {@code secret} appears in no captured message, and in no attached exception's
* message either — the constraint is that a credential must never reach a log, however it
* would have gotten there. */
private void assertNoLeak(String secret) {
for (ILoggingEvent event : reportingAppender.list) {
assertFalse(event.getFormattedMessage().contains(secret),
"log message leaked a credential (" + secret + "): " + event.getFormattedMessage());
IThrowableProxy thrown = event.getThrowableProxy();
if (thrown != null && thrown.getMessage() != null) {
assertFalse(thrown.getMessage().contains(secret),
"logged exception leaked a credential (" + secret + "): " + thrown.getMessage());
}
}
}
/** Gap 1: only {@code origin} was ever inspected. A credential on any other remote's fetch URL
* must now be reported. */
@Test
void aCredentialedUrlOnANonOriginRemoteIsReported(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
git(repo, "remote", "add", "origin", "https://git.ltms.dev/akb/kb.git");
git(repo, "remote", "add", "upstream", "https://leaky-upstream-token@git.ltms.dev/akb/kb.git");
new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-189-a", "HEAD");
List<String> messages = capturedMessages();
assertTrue(messages.stream().anyMatch(m -> m.contains("remote=upstream")),
"expected a report naming the leaking non-origin remote:\n" + messages);
assertNoLeak("leaky-upstream-token");
assertNoLeak("https://leaky-upstream-token@git.ltms.dev/akb/kb.git");
assertNoLeak("git.ltms.dev");
}
/** Gap 2: push URLs were never inspected. A credential visible only on {@code pushurl} — the
* fetch URL for the same remote stays clean — must now be reported. */
@Test
void aCredentialedPushUrlIsReported(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
git(repo, "remote", "add", "origin", "https://git.ltms.dev/akb/kb.git");
git(repo, "remote", "add", "mirror", "https://git.ltms.dev/akb/mirror.git");
git(repo, "remote", "set-url", "--push", "mirror",
"https://leaky-push-token@git.ltms.dev/akb/mirror.git");
new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-189-b", "HEAD");
List<String> messages = capturedMessages();
assertTrue(messages.stream().anyMatch(m -> m.contains("remote=mirror")),
"expected a report naming the remote with the leaking pushurl:\n" + messages);
assertNoLeak("leaky-push-token");
assertNoLeak("https://leaky-push-token@git.ltms.dev/akb/mirror.git");
assertNoLeak("git.ltms.dev");
}
/** Gap 3: only {@code https} was handled. A plain {@code http://user:pass@…} remote — worse
* than https, not better — must now be reported. */
@Test
void anHttpUrlWithCredentialsIsReported(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
git(repo, "remote", "add", "origin", "https://git.ltms.dev/akb/kb.git");
git(repo, "remote", "add", "insecure", "http://plainuser:plainpass@git.ltms.dev/akb/kb.git");
new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-189-c", "HEAD");
List<String> messages = capturedMessages();
assertTrue(messages.stream().anyMatch(m -> m.contains("remote=insecure")),
"expected a report for the credentialed plain-http remote:\n" + messages);
assertNoLeak("plainuser");
assertNoLeak("plainpass");
assertNoLeak("plainuser:plainpass");
assertNoLeak("git.ltms.dev");
}
/** A normal {@code ssh://} remote and a credential-free {@code https://} remote must produce no
* report at all — the check must not cry wolf on ordinary, safe configuration. */
@Test
void anSshRemoteAndACleanHttpsRemoteProduceNoReport(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
git(repo, "remote", "add", "origin", "ssh://git@git.ltms.dev:2224/akb/kb.git");
git(repo, "remote", "add", "clean", "https://git.ltms.dev/akb/kb.git");
new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-189-d", "HEAD");
assertTrue(reportingAppender.list.isEmpty(),
"expected no report for an ssh remote and a credential-free https remote, got:\n"
+ capturedMessages());
}
/**
* CB-189 review fix. Even a FAILING command that read a credential onto its stdout must never
* let that value reach the thrown {@link WorktreeException}'s message — this is gap 4 from the
* CB-189 issue, and the reason {@link GitWorktrees#execRedacted} exists at all. Drives the
* shared {@code exec}/{@code execRedacted} seam directly (it is package-private for exactly this,
* the same way the {@code afterWorktreeAdded} constructor parameter is a test seam) with a
* synthetic, non-git command whose stdout carries a marker — passed through the environment,
* never through argv, so the marker cannot leak via the command line that IS always printed
* unconditionally in the exception message — and which exits non-zero. This isolates the
* redaction guarantee itself rather than depending on a specific git failure mode that happens to
* echo a URL onto stdout before failing: none of the git subcommands this class actually runs was
* found to have one (a corrupted config makes {@code git config --get} fail before it ever reads
* the target key, so its output never carries the URL either). The marker is generated per-test
* run and injected only by the test, never a real-looking credential, so even a failing assertion
* could not itself print a secret.
*/
@Test
void execRedactedNeverCopiesFailingCommandOutputIntoTheExceptionMessage(@TempDir Path tmp) {
String marker = "cb189-marker-" + System.nanoTime();
GitWorktrees worktrees = new GitWorktrees(tmp.toString());
WorktreeException thrown = assertThrows(WorktreeException.class, () -> worktrees.exec(
Map.of("MARKER", marker), true, "sh", "-c", "echo \"$MARKER\"; exit 7"));
assertNotNull(thrown.getMessage());
assertFalse(thrown.getMessage().contains(marker),
"a failing command's captured stdout leaked into the exception message: "
+ thrown.getMessage());
}
/** Neutralizing must not look like work in progress, or a worker would commit it into its PR. */
@Test
void theNeutralizedConfigIsNotAPendingLocalModification(@TempDir Path tmp) throws Exception {
@@ -924,179 +622,4 @@ class GitWorktreesTest {
assertEquals(2, stats.count(), "two snapshot refs are reported");
assertTrue(stats.costBytes() > 0, "the cost of the snapshots is a positive byte count");
}
/**
* fleetd #185 stage 3: a recording {@link java.util.function.Function} test seam stands in for
* every process {@link GitWorktrees#shareWithGroup} would run — no real second OS user/group
* exists on this host, so these are unit tests against that seam, not a live-group integration
* test (out of scope per the ticket).
*/
private static List<String> joined(String[] command) {
return List.of(command);
}
/** Remove {@code path} and anything under it. Tolerates an already-absent path. */
private static void deleteRecursively(Path path) throws Exception {
if (!Files.exists(path)) {
return;
}
if (Files.isDirectory(path)) {
try (java.util.stream.Stream<Path> children = Files.list(path)) {
for (Path child : children.toList()) {
deleteRecursively(child);
}
}
}
Files.delete(path);
}
/** {@code worktreeGroup} absent ⇒ zero processes spawned and no git config written. */
@Test
void shareWithGroupIsNoopWhenNoGroupConfigured(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
return "";
};
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString(), null, _ -> {}, recordingRunner);
gitWorktrees.shareWithGroup(repo.toString(), repo.resolve("some-worktree").toString());
assertTrue(recorded.isEmpty(), "no group configured must spawn no process at all: " + recorded);
}
/** A configured group runs {@code git config core.sharedRepository group} first, then
* chgrp/chmod/setgid over every path {@link GitWorktrees#shareWithGroup} documents. */
@Test
void shareWithGroupRunsConfigThenChgrpChmodSetgidPerPath(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
String repoRoot = repo.toString();
Path worktree = Files.createDirectories(repo.resolve("some-worktree"));
String worktreePath = worktree.toString();
Files.createDirectories(repo.resolve(".git/worktrees"));
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
// What real git answers for an ordinary (non-linked) checkout: relative to repoRoot.
return List.of(cmd).contains("--git-common-dir") ? ".git\n" : "";
};
GitWorktrees gitWorktrees =
new GitWorktrees(tmp.resolve("wts").toString(), "devteam", _ -> {}, recordingRunner);
gitWorktrees.shareWithGroup(repoRoot, worktreePath);
assertEquals(List.of("git", "-C", repoRoot, "config", "core.sharedRepository", "group"), recorded.get(0),
"core.sharedRepository must be set first, so it keeps working after the one-time fix-up");
assertTrue(recorded.contains(List.of("git", "-C", repoRoot, "rev-parse", "--git-common-dir")),
"the git dir must be asked for, never hardcoded as <repoRoot>/.git — that is a FILE "
+ "when the checkout is itself a linked worktree: " + recorded);
for (String dir : List.of(worktreePath, repoRoot + "/.git/objects", repoRoot + "/.git/refs",
repoRoot + "/.git/logs", repoRoot + "/.git/worktrees")) {
assertTrue(recorded.contains(List.of("chgrp", "-R", "devteam", dir)), "missing chgrp -R for " + dir);
assertTrue(recorded.contains(List.of("chmod", "-R", "g+rwX", dir)), "missing chmod -R for " + dir);
assertTrue(recorded.contains(List.of("find", dir, "-type", "d", "-exec", "chmod", "g+s", "{}", "+")),
"missing setgid find pass for " + dir);
}
// packed-refs does not exist in a freshly-init'd repo (only git gc / pack-refs creates it) —
// tolerated absence, so it must not appear at all: no recursive/-R treatment for a plain file.
String packedRefs = repoRoot + "/.git/packed-refs";
assertTrue(recorded.stream().noneMatch(c -> c.contains(packedRefs)),
"packed-refs is absent here and must be skipped, not chgrp'd: " + recorded);
}
/**
* A path that does not exist is skipped, never handed to {@code chgrp}. {@code .git/logs} is
* absent whenever {@code core.logAllRefUpdates} is false or no ref has been updated yet, and
* {@code chgrp} on a missing path exits non-zero — which would fail EVERY provisioning spawn
* with a message blaming a group that is in fact fine.
*/
@Test
void shareWithGroupSkipsPathsThatDoNotExist(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
String repoRoot = repo.toString();
deleteRecursively(repo.resolve(".git/logs"));
assertFalse(Files.exists(repo.resolve(".git/logs")), "fixture: .git/logs must be gone");
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
return List.of(cmd).contains("--git-common-dir") ? ".git\n" : "";
};
GitWorktrees gitWorktrees =
new GitWorktrees(tmp.resolve("wts").toString(), "devteam", _ -> {}, recordingRunner);
gitWorktrees.shareWithGroup(repoRoot, repo.resolve("no-such-worktree").toString());
String logs = repoRoot + "/.git/logs";
assertTrue(recorded.stream().noneMatch(c -> c.contains(logs)),
"a missing .git/logs must be skipped, not chgrp'd: " + recorded);
assertTrue(recorded.stream().noneMatch(c -> c.contains(repo.resolve("no-such-worktree").toString())),
"a missing worktree path must be skipped too: " + recorded);
assertTrue(recorded.contains(List.of("chgrp", "-R", "devteam", repoRoot + "/.git/objects")),
"paths that DO exist are still shared: " + recorded);
}
/**
* The git store is located by {@code rev-parse --git-common-dir}, not by appending
* {@code /.git}. When git answers with an absolute path — what it does for a linked worktree,
* where {@code <repoRoot>/.git} is a file — every shared path must follow that answer.
*/
@Test
void shareWithGroupFollowsAnAbsoluteGitCommonDir(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
Path realGitDir = repo.resolve(".git");
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
return List.of(cmd).contains("--git-common-dir") ? realGitDir + "\n" : "";
};
GitWorktrees gitWorktrees =
new GitWorktrees(tmp.resolve("wts").toString(), "devteam", _ -> {}, recordingRunner);
gitWorktrees.shareWithGroup(tmp.resolve("some/linked/worktree").toString(),
repo.resolve("wt").toString());
assertTrue(recorded.contains(List.of("chgrp", "-R", "devteam", realGitDir + "/objects")),
"objects must be taken from the reported common dir, not <repoRoot>/.git: " + recorded);
}
/** {@code packed-refs}, when present, is chgrp/chmod'd but never setgid'd (it is a file, not a dir). */
@Test
void shareWithGroupIncludesPackedRefsWhenPresent(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
String repoRoot = repo.toString();
Path packedRefsPath = repo.resolve(".git/packed-refs");
Files.writeString(packedRefsPath, "");
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
return "";
};
GitWorktrees gitWorktrees =
new GitWorktrees(tmp.resolve("wts").toString(), "devteam", _ -> {}, recordingRunner);
gitWorktrees.shareWithGroup(repoRoot, repo.resolve("some-worktree").toString());
String packedRefs = packedRefsPath.toString();
assertTrue(recorded.contains(List.of("chgrp", "devteam", packedRefs)),
"packed-refs must be chgrp'd non-recursively when present: " + recorded);
assertTrue(recorded.contains(List.of("chmod", "g+rwX", packedRefs)),
"packed-refs must be chmod'd non-recursively when present: " + recorded);
assertTrue(recorded.stream().noneMatch(c -> c.contains("find") && c.contains(packedRefs)),
"packed-refs (a file) must never get the recursive setgid pass: " + recorded);
}
/** A group that does not exist (or that the operator is not a member of) fails loudly, naming it. */
@Test
void shareWithGroupThrowsNamingTheGroupWhenChgrpFails(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString(), "cb185-nonexistent-group-zz");
String wt = new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-185-share", "HEAD");
WorktreeException e = assertThrows(WorktreeException.class,
() -> gitWorktrees.shareWithGroup(repo.toString(), wt));
assertTrue(e.getMessage().contains("cb185-nonexistent-group-zz"),
"exception must name the missing/refused group: " + e.getMessage());
}
}
@@ -148,10 +148,6 @@ class SessionManagerTest {
return 0;
}
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
}
List<String> removeCalls() {
return List.copyOf(removeCalls);
}
@@ -887,18 +883,15 @@ class SessionManagerTest {
}
@Test
void acquireWithNeitherSessionFieldStillMintsAnAgentSessionId() {
// fleetd #214: the claude-code launcher mints a session id for EVERY spawn, so a member is
// resumable even when the spawn asked for no session identity.
void acquireWithNeitherSessionFieldLeavesAgentSessionIdNull() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession s = sessions.acquire("ltms-local", null, null, null);
assertNotNull(s.agentSessionId(),
"fleetd #214: a plain spawn mints a session id, so every member is resumable");
assertTrue(SessionManager.rosterView(s, null).containsKey("agentSessionId"),
"the minted id is in the roster, so fleet_list advertises every member's resume handle");
assertNull(s.agentSessionId(), "no identity requested — unchanged from before CB-584");
assertFalse(SessionManager.rosterView(s, null).containsKey("agentSessionId"),
"a null id is omitted from the roster, like charterSha256 for a receipt-less session");
}
@Test
@@ -934,235 +927,4 @@ class SessionManagerTest {
assertTrue(e.getMessage().contains("SESSION_RESUME"), e.getMessage());
assertTrue(e.getMessage().contains("stub-profile"), e.getMessage());
}
// ── fleetd #209: agentSessionId resolved lazily against the retained handle ─────────────────
//
// The opencode adapter cannot answer PeerHandle.agentSessionId() at spawn time — the on-disk
// session row is written only after the pane is live — so the id must be re-polled on a LATER
// call, against the SAME handle instance the launcher returned at spawn. SessionManager used
// to let that handle go out of scope at the end of the spawn method, so no caller ever re-asked
// it and fleet_list/fleet_status never saw the id. LazyIdHandle below reproduces exactly that
// shape: null on the first N calls (the spawn-time call included), a real id after.
/**
* A {@link PeerHandle} whose {@link #agentSessionId()} answers {@code null} for its first
* {@code nullCalls} invocations, then either a fixed id or a configured throw on every call
* after that — the shape of the opencode bug (fleetd #209): the session row is not written
* until after the pane is live, so early polls come back empty and a later one finds it.
*/
private static final class LazyIdHandle implements PeerHandle {
private final String id;
private final String terminalId;
private final int nullCalls;
private final String resolvedId;
private final java.util.concurrent.atomic.AtomicInteger calls =
new java.util.concurrent.atomic.AtomicInteger();
private volatile RuntimeException throwAfter;
LazyIdHandle(String id, String terminalId, int nullCalls, String resolvedId) {
this.id = id;
this.terminalId = terminalId;
this.nullCalls = nullCalls;
this.resolvedId = resolvedId;
}
/** After the null calls are exhausted, throw instead of answering the resolved id. */
LazyIdHandle throwing(RuntimeException e) {
this.throwAfter = e;
return this;
}
@Override
public String id() {
return id;
}
@Override
public String terminalId() {
return terminalId;
}
@Override
public String agentSessionId() {
int n = calls.incrementAndGet();
if (n <= nullCalls) {
return null;
}
if (throwAfter != null) {
throw throwAfter;
}
return resolvedId;
}
@Override
public CharterReceipt charterReceipt() {
return null;
}
int callCount() {
return calls.get();
}
}
/** A minimal {@link PeerLauncher} that hands out pre-built {@link LazyIdHandle}s, one per spawn. */
private static final class LazyIdLauncher implements PeerLauncher {
private final java.util.Deque<LazyIdHandle> queued = new java.util.ArrayDeque<>();
LazyIdLauncher queue(LazyIdHandle handle) {
queued.add(handle);
return this;
}
@Override
public Set<Capability> capabilities() {
return Set.of(Capability.WORKTREE, Capability.SESSION_RESUME);
}
@Override
public Set<Capability> capabilitiesFor(String profileName) {
return capabilities();
}
@Override
public PeerHandle spawn(SpawnRequest req) {
LazyIdHandle handle = queued.poll();
if (handle == null) {
throw new IllegalStateException("no queued LazyIdHandle for this spawn");
}
return handle;
}
@Override
public Set<String> profiles() {
return Set.of("lazy");
}
@Override
public String defaultProfile() {
return "lazy";
}
@Override
public String effectiveCwd(SpawnRequest req) {
return "/cwd";
}
@Override
public List<String> parityOverlay(String profileName) {
return List.of();
}
@Override
public List<?> list() {
return List.of();
}
@Override
public int reapOrphanWorkers() {
return 0;
}
@Override
public void stop(String id) {
}
@Override
public boolean clearContext(String id) {
return false;
}
}
@Test
void plainRosterDoesNotResolveAgentSessionId() {
// fleetd #209 follow-up: roster() sits on the heartbeat/health-tick timers (and the metrics
// scrape), so it must never trigger the resolve lookup — for opencode that lookup opens an
// on-disk session database, and a member whose id never appears would pay that cost forever.
// rosterResolved() is the one to use when a caller actually reports the id.
LazyIdHandle handle = new LazyIdHandle("p0", "t0", 1, "oc-session-0");
SessionManager sessions = new SessionManager(new LazyIdLauncher().queue(handle));
MemberSession acquired = sessions.acquire("lazy", "/cwd", "/caller", null);
assertEquals(1, handle.callCount(), "sanity: only the spawn-time call happened so far");
List<MemberSession> roster = sessions.roster();
assertEquals(1, roster.size());
assertEquals(acquired.paneId(), roster.getFirst().paneId());
assertNull(roster.getFirst().agentSessionId(), "the plain roster must not resolve the id");
assertEquals(1, handle.callCount(),
"roster() must never call agentSessionId() again — it sits on the heartbeat/health timers");
}
@Test
void rosterResolvedResolvesALateAgentSessionIdFromTheRetainedHandle() {
LazyIdHandle handle = new LazyIdHandle("p1", "t1", 1, "oc-session-1");
SessionManager sessions = new SessionManager(new LazyIdLauncher().queue(handle));
MemberSession acquired = sessions.acquire("lazy", "/cwd", "/caller", null);
assertNull(acquired.agentSessionId(),
"opencode has not written its session row yet at spawn time");
List<MemberSession> roster = sessions.rosterResolved();
assertEquals(1, roster.size());
assertEquals("oc-session-1", roster.getFirst().agentSessionId(),
"fleet_list must see the id once the adapter can answer it");
Map<String, Object> view = SessionManager.rosterView(roster.getFirst(), null);
assertEquals("oc-session-1", view.get("agentSessionId"),
"rosterView renders whatever rosterResolved() resolved");
}
@Test
void getResolvesALateAgentSessionIdFromTheRetainedHandle() {
LazyIdHandle handle = new LazyIdHandle("p2", "t2", 1, "oc-session-2");
SessionManager sessions = new SessionManager(new LazyIdLauncher().queue(handle));
MemberSession acquired = sessions.acquire("lazy", "/cwd", "/caller", null);
MemberSession resolved = sessions.get(acquired.paneId()).orElseThrow();
assertEquals("oc-session-2", resolved.agentSessionId(),
"fleet_status (single-session lookup) must also see the late-resolved id");
}
@Test
void releaseCarriesALateResolvedAgentSessionIdIntoTheReleaseDetail() {
LazyIdHandle handle = new LazyIdHandle("p3", "t3", 1, "oc-session-3");
SessionManager sessions = new SessionManager(new LazyIdLauncher().queue(handle));
MemberSession acquired = sessions.acquire("lazy", "/cwd", "/caller", null);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(detail -> released.add(detail.agentSessionId()));
sessions.release(acquired.paneId());
assertEquals(java.util.List.of("oc-session-3"), released,
"a released member's detail carries the id it has since resolved, not the null "
+ "frozen in at spawn time");
}
@Test
void aThrowingHandleDoesNotBreakRosterResolved() {
LazyIdHandle handle = new LazyIdHandle("p4", "t4", 1, "oc-session-4")
.throwing(new RuntimeException("sqlite locked"));
SessionManager sessions = new SessionManager(new LazyIdLauncher().queue(handle));
sessions.acquire("lazy", "/cwd", "/caller", null);
List<MemberSession> roster = assertDoesNotThrow(sessions::rosterResolved,
"a handle that throws resolving its id must not break the roster read");
assertEquals(1, roster.size());
assertNull(roster.getFirst().agentSessionId(), "the id stays unresolved when the lookup throws");
}
@Test
void aResolvedAgentSessionIdIsNotLookedUpAgain() {
LazyIdHandle handle = new LazyIdHandle("p5", "t5", 1, "oc-session-5");
SessionManager sessions = new SessionManager(new LazyIdLauncher().queue(handle));
sessions.acquire("lazy", "/cwd", "/caller", null);
assertEquals(1, handle.callCount(), "sanity: only the spawn-time call happened so far");
sessions.rosterResolved();
assertEquals(2, handle.callCount(), "the first rosterResolved() read resolves the id");
sessions.rosterResolved();
assertEquals(2, handle.callCount(), "once resolved, the id must not be looked up again");
}
}
@@ -163,37 +163,6 @@ class WorktreeSessionManagerTest {
"tracked copied paths are --skip-worktree'd");
}
/**
* fleetd #185 stage 3, THE TRAP: {@code overlayParity} copies more files into the worktree
* AFTER {@code add} returns, so {@code shareWithGroup} must run after it, not folded into
* {@code add()} — otherwise every overlay file lands operator-owned and unwritable for a
* different-uid member, with a green test suite hiding it.
*/
@Test
void shareWithGroupRunsAfterOverlayParityNotBeforeIt() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.track(".envrc");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-185", null));
assertEquals(1, worktrees.overlayCalls().size(), "overlayParity ran exactly once");
assertEquals(1, worktrees.shareCalls().size(), "shareWithGroup ran exactly once");
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
FakeWorktrees.ShareCall share = worktrees.lastShare();
assertEquals(s.worktree(), overlay.worktreePath());
assertEquals(s.worktree(), share.worktreePath());
List<String> order = worktrees.overlayShareOrder();
int overlayIndex = order.indexOf("overlay:" + s.worktree());
int shareIndex = order.indexOf("share:" + s.worktree());
assertTrue(overlayIndex >= 0 && shareIndex >= 0, "both calls must be recorded: " + order);
assertTrue(overlayIndex < shareIndex,
"shareWithGroup MUST run after overlayParity, not before/inside add(): " + order);
}
@Test
void releaseRemovesWorktreeButDoesNotDeleteBranch() {
FakeHerdr herdr = new FakeHerdr();