Compare commits

..

2 Commits

Author SHA1 Message Date
Dai Ha c50f5b2d61 fleetd #176 stage 2: make effectiveCredentialId() subscription-aware
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m44s
Stage 1's lead-seat matcher (leadSeatLookup) was correct but inert on
the live host: the lead runs on profile 'opus', members on 'sonnet',
both subscription:true with no explicit credentialId. Because
effectiveCredentialId() fell back to the profile's own name, opus and
sonnet never matched even though they share one Claude login, so the
matcher charged zero seats.

FleetConfig.Profile.effectiveCredentialId() now falls back to a shared
sentinel (SUBSCRIPTION_CREDENTIAL_ID = "<subscription>") instead of the
profile name when subscription:true and credentialId is unset. An
explicit credentialId still wins, so two separate Claude logins on one
host can still be kept apart.

This is also BackendQuarantine's and BackendOutagePolicy's grouping
key and CompositePeerLauncher's spawn-time enforcement key, so the fix
also links quarantine/cool-off across subscription profiles sharing an
account -- intentional: one usage limit really does take out every
profile on that login, mirroring credentialId: openai-shared already
doing this for off-subscription profiles. Every caller was reviewed;
none wants "this exact profile" over "this account".

Tests added:
- FleetdLeadSeatLookupTest: the live shape itself (lead on a
  DIFFERENT subscription profile than the target, same account,
  neither sets credentialId) -- the case stage 1's suite never covered
- FleetMcpTest: quarantining one subscription profile's shared
  account zeroes free on another sharing it, via the same
  effectiveCredentialId()-driven wiring Fleetd.main uses

Mutation-tested: reverting the subscription branch to the old
fall-back-to-profile-name behavior sends both new tests RED with 0
compile errors; reverting the mutation restores byte-identical
(diff -q) source and green tests.

fleetd.example.yaml's fleetd #176 notes are rewritten for the sentinel
semantics and when to override it with an explicit credentialId.
2026-09-03 13:10:10 +07:00
Dai Ha c796eac09c fleetd #176: subtract the lead's own subscription seat from free
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m16s
maxLoad counted panes, never subscription seats: a subscription:true
profile's lead is itself a live claude session on that same account,
so free overstated capacity by the lead's own seat (measured free:1
with a real ceiling of 0, and free:3 on an idle fleet with a real
ceiling of 2).

Add FleetMcp.LeadSeatSource (same shape as QuarantineSource/
OutageSource) and Fleetd.leadSeatLookup, which derives the seat count
from fleet.leaders.<name>.profile matched against the target profile
by effectiveCredentialId() - no hardcoded "-1", and no new config key:
profile: already exists for this exact "which account does this lead
share" question. maxLoad itself is left untouched; only free (and a
new, additive-only leadSeats field) changes.

Exhaustion quarantine (cause 2 in the ticket) already forced free to 0
via the same BackendQuarantine capacityView already reads - confirmed
by reading the exhaustionSink wiring, no code change needed there.
2026-09-03 12:55:08 +07:00
16 changed files with 1021 additions and 660 deletions
+5 -8
View File
@@ -215,14 +215,11 @@ must obey belongs in the charter, not here.
reads the list with `git config --worktree --get-all fleet.neutralizedConfig`, and the
consequence with `git config --worktree --get fleet.neutralizedConfigNote`. Never brief a worker
to edit one of these files: the edit cannot be committed, and it will not tell you so.
- **Flows and the error model** — rendezvous, `fleet_ask`, detached delivery, the turn-done
fallback and status gating — are diagrammed in `docs/MCP-Contract.md`. That page is now flows
only: its pre-build tool catalogue, parameter tables and REST paths were deleted rather than
corrected, because a hand-maintained second copy of the tool surface is what drifted for a month
while this line pointed every session at it (CB-609 / #114). **The live MCP schema is the tool
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
file because this file loads into every session's context.
- **Flows and the error model** — rendezvous, `fleet_ask`, detached delivery, turn-done fallback —
are diagrammed in `docs/MCP-Contract.md` **§6 only**. The rest of that page is a pre-build design
doc whose tool names, parameter names and REST paths never caught up with the code, so do not use
it as the tool reference (CB-609). Section 6 is kept out of this file because this file loads into
every session's context.
### Redeploying the daemon — the lead may do this (primary only)
+324 -123
View File
@@ -1,162 +1,311 @@
# MCP flows and error model — `fleetd`
# MCP Contract — `fleetd`'s unified gateway
> **What this page is.** The **flows**: how a delegation, a clarification, a detached task and a
> silent member each travel through `fleetd`. These shapes are what shipped, and they are hard to
> read off the code because they span the MCP face, the rendezvous registry, the `Injector` and
> herdr.
> **Status: 🔴 HISTORICAL DESIGN — do NOT use as the tool reference.** Written 2026-07-14, before
> any MCP code existed. The system shipped and this page never caught up, so **its tool names,
> parameter names and REST paths are wrong today**. Audited 2026-08-17; the specific drift:
>
> **What this page is NOT: a tool reference.** It deliberately holds no tool catalogue, no
> parameter tables and no REST paths. **The live MCP schema is the authority** — each tool's own
> description and parameters, as mounted — with the intent→tool table in `CLAUDE.md` as the short
> form.
> - **Tools it names that do not exist:** `fleet_read`, `fleet_cancel`.
> - **Shipped tools it omits:** `fleet_poll`, `fleet_ack`, `fleet_profiles`, `fleet_whoami`.
> - **Parameter names are wrong nearly everywhere** — it says `message`/`target`/`timeout_seconds`/
> `block` where the code takes `content`/`sessionId`/`timeoutMs`/`wait`; `text` where
> `fleet_reply` takes `content`; `target` where `fleet_stop` takes `paneId`.
> - **REST paths are wrong:** it says `POST /workers` and `DELETE /workers/{paneId}`; the daemon
> serves `POST /members` and `DELETE /members/{paneId}`.
>
> That absence is the fix for fleetd #114 (CB-609), and it is worth stating why. This page used to
> carry a full tool catalogue written in July 2026, before any MCP code existed. The code shipped;
> the page did not follow. By August it named two tools that do not exist, omitted five that do,
> had the wrong name for nearly every parameter, pointed at REST paths the daemon does not serve,
> and — worst — still described an identity model (*"any connection that does not map to a known
> worker is treated as a primary"*) that was a real privilege bug, fixed since by the ancestry
> walk in fleetd #161. Every one of those errors is the same error: **a second, hand-maintained
> copy of something the code already states**. So the second copy is gone rather than corrected.
> Only the flows remain, because a flow is a shape rather than a name, and shapes are what this
> page was ever good for.
> **The authoritative tool surface is the live MCP schema** (each tool's own description and
> parameters, as mounted), with the intent→tool table in `CLAUDE.md` as the short form. Both were
> checked against `mcp/FleetMcp.java` on 2026-08-17 and are accurate.
>
> The names that do appear below are checked by `McpContractDocTest`, which fails if this page
> names a `fleet_*` tool the server does not register. That test is the whole reason it is safe to
> write a tool name here at all.
> What is still worth reading here is **§6 — the flows and the error model** (rendezvous,
> `fleet_ask`, detached delivery, the turn-done fallback). The shapes it describes are the ones
> that shipped; only the names around them drifted. Rewriting this page is tracked as **CB-609**.
`fleetd` is the **sole communication gateway** for every Claude session in the bridge. Both
the **primary** (Opus, on subscription) and every **worker** (off-subscription Claude Code)
mount the *same* MCP server with a single `claude mcp add` line, and talk only through its
tools. No Claude session ever addresses a broker, a peer, or the network directly.
This document defines every MCP tool that face must expose, who may call it, its blocking
semantics, and how it maps onto the code already in the tree.
---
## 1. Rendezvous flows
## 1. Design constraints (non-negotiable)
### 1.1 Delegation — happy path
These come from the project's core invariants and bound every decision below.
One blocking call, zero polls. The lead's call is held open by `fleetd` until the member answers.
1. **One server, both roles.** The primary and all workers mount an identical server. The
catalog must serve both, and `fleetd` must decide *who is calling* from the connection —
never from a caller-supplied argument that could be spoofed.
2. **Subscription-safe by construction.** No MCP tool ever reads, sets, or forwards
`ANTHROPIC_BASE_URL`. Mounting the bridge cannot move a session off subscription.
Enforced today by [`SubscriptionGuard`](1-Architecture).
3. **Blocking rendezvous, no busy-poll.** The primary consumes a worker's reply through a
*single* MCP call that `fleetd` holds open — never a cross-turn poll loop that would burn
subscription quota.
4. **Status-gated delivery.** Anything that puts text into a worker flows through the existing
[`Injector`](1-Architecture): delivered only when the worker is `idle`/`blocked`, at most
one message per turn.
5. **`fleetd` owns policy; herdr owns PTYs.** MCP tools express *intent*; `fleetd`
translates it into guard checks, rendezvous bookkeeping, and herdr `agent.*` calls.
---
## 2. Topology
Both faces live in the one daemon. The **north face** is MCP (this document); the **south
face** is the herdr Unix socket. REST/SSE remains only for non-Claude clients and dashboards.
```mermaid
flowchart LR
OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
subgraph BD["fleetd — standalone daemon"]
MCP["MCP server (north face)<br/>fleet_send · fleet_reply<br/>fleet_ask · fleet_status · lifecycle"]
RDV["rendezvous registry<br/>(blocking-call waiters)"]
INJ["Injector + StatusPoller<br/>(status-gated writer)"]
SOCK["herdr socket client (south face)"]
MCP --> RDV
RDV --> INJ
INJ --> SOCK
MCP --> SOCK
end
HERDR["herdr<br/>panes · agent-status"]
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
OPUS -->|"fleet_send (blocks)"| MCP
W -.->|"fleet_reply / fleet_ask"| MCP
SOCK -->|"agent.start · agent.send<br/>agent.get · pane.close"| HERDR
HERDR -->|"drives PTY"| W
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
class OPUS,W ext
class MCP,RDV,INJ,SOCK core
```
---
## 3. Identity & addressing
Because the same server is mounted by everyone, `fleetd` resolves the caller's role on every
request — this is the linchpin of the whole contract and has no code yet.
- **Workers are known.** `fleetd` spawns every worker
([`WorkerService`](1-Architecture)) and records its herdr session UUID / `terminal_id` on
the returned [`Agent`]. When a call arrives on a connection that maps to a known worker,
the caller is *that* worker — so **workers never pass a target**; routing is implicit.
- **The primary is "not a worker".** Any connection that does not map to a known worker is
treated as a primary. It addresses workers **explicitly** by `target` — a session UUID,
a `terminal_id`, or a friendly `profile` name.
- **Turn correlation.** A blocking `fleet_send` registers a *waiter* keyed by worker
identity. A worker's later `fleet_reply` / `fleet_ask` on the same identity resolves that
waiter. A `turn_id` is minted per exchange so a clarification round-trip
(§6.2) rejoins the right turn.
---
## 4. Transport
`fleetd` is a long-lived daemon serving **multiple** concurrent clients (one primary + N
workers), so a per-client stdio child is the wrong shape. The recommended transport is
**streamable-HTTP / SSE** on the same bind as the REST face:
```bash
# identical on primary and every worker
claude mcp add --transport http fleetd http://127.0.0.1:8080/mcp
```
This adds an MCP-server dependency the pom does not yet carry. See [Open decisions](#10-open-decisions).
---
## 5. Tool catalog
| Tool | Caller | Blocks? | Backing (exists today?) |
|---|---|---|---|
| [`fleet_send`](#fleet_send) | primary | yes (default) | `Injector.enqueue` ✅ · rendezvous registry ❌ (CB-104) |
| [`fleet_reply`](#fleet_reply) | worker | no | rendezvous ❌ · pane injection via `Injector` ✅ |
| [`fleet_ask`](#fleet_ask) | worker | yes | reverse rendezvous ❌ |
| [`fleet_status`](#fleet_status) | either | no | `AgentControl.status` ✅ · `Injector.activeTargets` ✅ |
| [`fleet_spawn`](#lifecycle) | primary | no | `WorkerService.spawn` ✅ (`POST /workers`) |
| [`fleet_list`](#lifecycle) | either | no | `WorkerService.list` ✅ (`/agents`) |
| [`fleet_stop`](#lifecycle) | primary | no | `WorkerService.stop` ✅ (`DELETE /workers/{paneId}`) |
| [`fleet_read`](#fleet_read) | primary | no | `AgentControl.read` ✅ |
| [`fleet_cancel`](#fleet_cancel) | primary | no | — ❌ (future) |
### Core: delegation & rendezvous
#### `fleet_send`
*(primary → worker — the headline tool, CB-104)*
- **Params:** `message` (required); `target` (optional — defaults to the sole worker / default
profile); `timeout_seconds` (default 600); `block` (default `true`); `auto_spawn`
(default `true`); `turn_id` (optional — supplied when answering a worker's `fleet_ask`).
- **Blocking (`block:true`):** enqueue `message` via the `Injector`, then hold the call open
until exactly one of:
- worker calls `fleet_reply` → `{ outcome:"reply", text }`
- worker calls `fleet_ask` → `{ outcome:"question", text, turn_id }`
- worker's `agent_status` reaches done/idle with no reply → `{ outcome:"turn_done", text:<terminal tail> }`
- deadline elapses → `{ outcome:"timeout" }`
- worker gone → error `worker_gone`
- **Detached (`block:false`):** enqueue and return `{ outcome:"dispatched", dispatch_id }`
immediately. The eventual reply is injected into the primary's idle pane (§6.3), or drained
via `fleet_status` on a split-host primary.
#### `fleet_reply`
*(worker → primary)*
- **Params:** `text` (required); `final` (default `true`).
- **Behavior:** resolve the primary waiter registered against this worker with `text`. If no
waiter exists (detached delegation), `fleetd` **injects the primary's idle pane** instead.
Returns `{ delivered:true, mode:"resolved"|"injected" }`. No `target` — identity is implicit.
#### `fleet_ask`
*(worker → primary — the reverse rendezvous)*
- **Params:** `question` (required); `timeout_seconds`.
- **Behavior:** blocks the *worker's* call. Surfaces the question to the primary (resolving its
open `fleet_send` with `outcome:"question"`, or injecting its pane). When the primary
answers — a `fleet_send` carrying the matching `turn_id` — that unblocks this call and
returns `{ answer }` to the worker, which continues **in the same turn**.
### Worker lifecycle
<a id="lifecycle"></a>
Thin adapters over [`WorkerService`](1-Architecture) — parity with the existing REST routes.
- **`fleet_spawn`** — `{ profile? }` → worker view (`sessionId`, `terminalId`, `paneId`,
`status`). Guard-checked; a boundary breach returns error `subscription_boundary` (the
REST `403`).
- **`fleet_list`** — no params → all workers + `agent_status`. Read-only, either role.
- **`fleet_stop`** — `{ target }` → tears down the pane and its dedicated tab. Idempotent.
### Observability
#### `fleet_status`
*(either role — the README's 4th named tool)*
- **Params:** `target?`.
- **Behavior:** per-worker `agent_status`, queue depth (`Injector.activeTargets`), whether a
rendezvous is open, and ids. For the *calling* session it also reports/drains **pending
messages addressed to me** — the path a split-host primary's `Stop`-hook uses to wake and
collect replies without being injectable. Read-only, non-blocking.
#### `fleet_read`
*(primary)*
- **Params:** `target`; `source` ∈ `visible | recent | recent_unwrapped | detection`.
- **Behavior:** returns the worker's terminal text so the primary can peek at a *detached*
worker's progress. Adapter over `AgentControl.read`.
### Control (future)
#### `fleet_cancel`
*(primary)*
- **Params:** `target`. Interrupt the worker's current turn / abandon the rendezvous. No
backing code yet.
---
## 6. Rendezvous flows
### 6.1 Delegation — happy path
One blocking call, zero polls.
```mermaid
sequenceDiagram
participant P as "Lead (primary)"
participant B as "fleetd (MCP + Injector)"
participant P as Primary (Opus)
participant B as fleetd (MCP + Injector)
participant H as herdr
participant W as "Member"
participant W as Worker (Claude)
P->>B: "fleet_send{sessionId, content} — blocks"
B->>B: "register waiter(sessionId)"
B->>H: "agent.send — only in an injectable window"
H-->>W: "prompt injected"
W->>W: "works the turn"
W->>B: "fleet_reply{content}"
B->>B: "resolve waiter"
B-->>P: "{ outcome: reply }"
P->>B: fleet_send("do X", target=w) — blocks
B->>B: register waiter(w)
B->>H: agent.send(w, "do X") (idle window)
H-->>W: prompt injected
W->>W: works the turn
W->>B: fleet_reply("result")
B->>B: resolve waiter(w)
B-->>P: { outcome:"reply", text:"result" }
```
**The cap that matters:** a blocking `fleet_send` is bounded by the *caller's own* MCP client
timeout, about 60 seconds — not by the task. Anything slower than that must use the detached flow
in §1.3, or the lead's call returns while the member is still working.
### 6.2 Clarification — reverse rendezvous (`fleet_ask`)
### 1.2 Clarification — reverse rendezvous
The member pauses mid-turn to ask, the lead answers, and the member resumes **the same turn** with
its context intact.
The worker pauses mid-turn to ask; the primary answers; the worker resumes in the same turn.
```mermaid
sequenceDiagram
participant P as "Lead"
participant P as Primary
participant B as fleetd
participant W as "Member"
participant W as Worker
P->>B: "fleet_send{sessionId, content} — blocks"
B-->>W: "content injected"
W->>B: "fleet_ask{question} — member blocks"
B-->>P: "{ outcome: question, turnId }"
P->>B: "fleet_send{turnId, content} — answers THIS turn"
B-->>W: "fleet_ask returns the answer"
W->>W: "resumes the same turn"
W->>B: "fleet_reply{content}"
B-->>P: "{ outcome: reply }"
P->>B: fleet_send("do X", target=w) — blocks
B-->>W: "do X" (injected)
W->>B: fleet_ask("which config?") — worker blocks
B-->>P: { outcome:"question", text:"which config?", turn_id }
P->>B: fleet_send("config.yaml", target=w, turn_id) — blocks again
B-->>W: resolve fleet_ask → { answer:"config.yaml" }
W->>W: resumes same turn
W->>B: fleet_reply("done")
B-->>P: { outcome:"reply", text:"done" }
```
**Answer with `turnId`, never `sessionId`.** A `sessionId` send starts a new turn; it does not
resolve the waiting `fleet_ask`.
### 6.3 Detached delegation — pane injection
**The window is about 55 seconds and no nudge extends it.** So never brief a member to "ask me":
decide before delegating, or give the member an explicit default to fall back on.
### 1.3 Detached delegation — the lead does not block
The lead gets a ticket immediately and collects the answer later. This is the flow for any real
task, because of the ~60s cap in §1.1.
The primary does not block; the reply arrives later in its idle pane.
```mermaid
sequenceDiagram
participant P as "Lead"
participant P as Primary
participant B as fleetd
participant W as "Member"
participant W as Worker
P->>B: "fleet_send{sessionId, content, wait:false}"
B-->>P: "accepted — ticket"
P->>P: "continues its own work"
W->>B: "fleet_reply{content}"
Note over B: "no waiter is blocked — the reply is held"
B->>B: "nudge the lead's own pane (status-gated)"
P->>B: "fleet_poll{ticket}"
B-->>P: "the member's report"
P->>B: "fleet_ack{target, msgId}"
P->>B: fleet_send("do X", target=w, block=false)
B-->>P: { outcome:"dispatched", dispatch_id }
P->>P: continues its own work
W->>B: fleet_reply("result")
Note over B: no waiter → detached path
B->>B: Injector.enqueue(primary_pane, "result")
B-->>P: injected into idle pane (status-gated)
```
A terminal ticket nudges the lead's pane by itself, so a detached task does not need watching. The
nudge needs an injectable lead pane and is capped, so it is a convenience rather than a guarantee.
### 6.4 Uncooperative worker — turn-done fallback
### 1.4 The member never replies — turn-done fallback
A member that ends its turn without `fleet_reply` still produces something: `fleetd` reads its
pane tail. This is a **fallback, not a channel** — it is lossy in three separate ways, and every
one of them has produced a wrong answer in practice.
A worker that never calls `fleet_reply` still returns a result: `fleetd` reads its terminal
tail when the turn completes.
```mermaid
sequenceDiagram
participant P as "Lead"
participant P as Primary
participant B as fleetd
participant W as "Member"
participant W as Worker
P->>B: "fleet_send — blocks or detaches"
B-->>W: "content injected"
W->>W: "works, never calls fleet_reply"
B->>B: "StatusPoller sees the turn end"
B->>B: "read the pane tail"
B->>B: "classify: exhausted? echoed brief? real report?"
B-->>P: "{ outcome: turn_done } or a named failure"
P->>B: fleet_send("do X", target=w) — blocks
B-->>W: "do X" (injected)
W->>W: works, never calls fleet_reply
B->>B: StatusPoller sees agent_status → idle/done
B->>B: AgentControl.read(w, "recent")
B-->>P: { outcome:"turn_done", text:<terminal tail> }
```
The three ways it goes wrong, and what each looks like now:
| What happened | What the lead used to get | What it gets today |
|---|---|---|
| The report is longer than the scrape window | The **end** silently cut off | Still clipped, but marked partial |
| The member never started — spent credential | The lead's **own brief** echoed back as a report | A named failure: backend exhausted |
| The member is simply slow | A tail of work in progress | Unchanged — read it as a hint, not a result |
The echoed-brief case is the one to remember: it reads as a long, on-topic report with nothing in
it from the member. It is suppressed now, but the general rule stands — **check the member's
worktree with `git log` before believing a report you did not watch arrive.**
---
## 2. Status gating
## 7. Status gating
Delivery only happens in a safe window. `fleet_send` is a producer for the `Injector`, which
already enforces this through `AgentStatus.injectable()`.
Delivery only happens in a safe window. This is the state machine the `Injector` already
enforces via `AgentStatus.injectable()`; MCP `fleet_send` is simply its producer.
```mermaid
stateDiagram-v2
[*] --> IDLE
IDLE --> WORKING: "message delivered, picked up"
WORKING --> IDLE: "turn done"
WORKING --> BLOCKED: "awaits input"
BLOCKED --> WORKING: "input delivered"
IDLE --> UNKNOWN: "detection glitch"
BLOCKED --> UNKNOWN: "detection glitch"
UNKNOWN --> IDLE: "re-detected"
IDLE --> WORKING: message delivered / picks up
WORKING --> IDLE: turn done
WORKING --> BLOCKED: awaits input
BLOCKED --> WORKING: input delivered
IDLE --> UNKNOWN: detection glitch
BLOCKED --> UNKNOWN: detection glitch
UNKNOWN --> IDLE: re-detected
note right of IDLE
injectable — deliver head of FIFO
@@ -172,17 +321,69 @@ stateDiagram-v2
end note
```
**At most one message per turn.** After a send, the `Injector` waits for a `WORKING` pickup before
delivering the next, with a grace-poll fallback for turns that finish faster than the poll
interval.
At most one message is delivered per turn: after a send the `Injector` waits for a `WORKING`
pickup before delivering the next, with a `PICKUP_GRACE_POLLS` fallback for turns faster than
the poll interval. A herdr `events.subscribe` stream can later replace the sampling without
touching this state machine.
Two consequences a lead feels directly:
---
- **A second send to a busy member never lands.** It reports as queued and times out. The member
is fine; the message simply waits, and then restarts the member when it next goes idle.
- **A spawned member is not deliverable until it has mounted the MCP.** Until then a send waits on
that gate for about 60 seconds and then fails without ever reaching the pane.
## 8. Error model
`UNKNOWN` is deliberately neither injectable nor a pickup. A pane whose status cannot be read is
not a pane that is safe to write to — see fleetd #176 for what happens when a gate treats an
unreadable pane as a ready one.
| Condition | `fleet_send` result | Notes |
|---|---|---|
| Worker replies | `{ outcome:"reply" }` | normal |
| Worker asks | `{ outcome:"question", turn_id }` | answer with `fleet_send(turn_id)` |
| Turn ends, no reply | `{ outcome:"turn_done" }` | terminal tail as text |
| Deadline elapsed | `{ outcome:"timeout" }` | message may still be queued/delivered |
| Worker vanished | error `worker_gone` | `Injector.drop` fails the queued future |
| Guard breach on spawn | error `subscription_boundary` | REST `403` parity |
| Delivery failed at herdr | error, message dropped | poisoned message not left blocking the FIFO |
`fleet_reply` from a worker with no open waiter is **not** an error — it falls through to
detached pane injection (§6.3).
---
## 9. Mapping to existing code
The MCP face is a thin adapter layer; nearly every capability already exists behind the REST
seam. Only the **rendezvous registry** and the **caller-identity resolver** are new.
| MCP tool | Existing collaborator | New work |
|---|---|---|
| `fleet_send` | `Injector.enqueue`, `AgentControl.send` | waiter registry, timeout, outcome mux (CB-104) |
| `fleet_reply` / `fleet_ask` | `Injector` (pane injection) | reverse rendezvous, identity resolver |
| `fleet_status` | `AgentControl.status`, `Injector.activeTargets` | pending-drain projection |
| `fleet_spawn` / `list` / `stop` | `WorkerService.{spawn,list,stop}` | MCP adapter only |
| `fleet_read` | `AgentControl.read` | MCP adapter only |
Because the REST routes in `FleetApp` already exercise the collaborators, MCP tools are
validated by **parity** against those routes, not by re-testing behavior.
---
## 10. Open decisions
1. **`fleet_ask` direction.** This page defines it as *worker-asks-primary* (a genuine reverse
channel, matching the "inject the primary's pane" language). The alternative — a synonym for
a blocking primary→worker send — is weaker and produces different plumbing. **Recommend
worker-asks-primary.**
2. **Detached delivery shape.** A `block:false` param on `fleet_send` (keeps the catalog
small) vs. a separate `fleet_dispatch` tool. **Recommend the param.**
3. **Auto-spawn on send.** `fleet_send` provisions a worker per profile when none exists
(simplest primary UX) vs. requiring an explicit `fleet_spawn` first. **Recommend
auto-spawn, defaulting on.**
4. **Transport & SDK.** Streamable-HTTP/SSE co-located with the REST bind (recommended) vs.
stdio. Requires choosing a Java MCP server SDK and adding it to the pom.
---
## 11. Implementation staging
- **CB-104** — blocking `fleet_send` + rendezvous registry + caller-identity resolver
(the producer that finally drives the inert `StatusPoller`).
- **CB-1xx** — `fleet_reply` / `fleet_ask` reverse rendezvous + detached pane injection.
- **CB-1xx** — lifecycle + observability adapters (`fleet_spawn/list/stop/status/read`).
- **CB-1xx** — transport wiring + `claude mcp add` docs; parity tests vs. REST.
- **Later** — `fleet_cancel`; swap `StatusPoller` for herdr `events.subscribe`.
+45
View File
@@ -306,6 +306,35 @@ profiles:
# GOTCHA 2 — `maxLoad` is the ONLY throttle you have here. There is no metering, no budget
# and no refusal on cost; the cap on live members is the single thing standing between a
# fan-out and your monthly limit. Set it deliberately and keep it small.
#
# GOTCHA 3 (fleetd #176) — `maxLoad` counts members, never the lead itself. The lead is a live
# `claude` session on this SAME account (a lead is never moved off-subscription, whatever its
# own profile says), so it already holds one seat before any member spawns. If a lead's
# `fleet.leaders.<name>.profile` names THIS profile — or ANY OTHER `subscription: true`
# profile that shares this one's account (see THE SENTINEL, just below, next to
# `credentialId:`) — `fleet_list`'s `free` for this profile subtracts that lead's live
# seat(s) automatically; see `profile:` under THE FLEET below. If no lead entry names a
# profile sharing this account, fleetd has no way to know a lead holds a seat here, and `free`
# will overstate what a fresh `fleet_spawn` actually gets by exactly the seats the lead is
# quietly holding.
#
# THE SENTINEL (fleetd #176 stage 2, correcting an inert stage 1 fix): every `subscription:
# true` profile that leaves `credentialId` unset shares ONE implicit account-wide credential
# id with every other such profile on this host — because a subscription profile doesn't
# authenticate with a credential of its own, it authenticates as the operator's own Claude
# login, and there is exactly one of those. So on a typical host, `opus` (the lead's profile)
# and `sonnet` (the members' profile) are linked automatically, with NOTHING to set here — that
# is what makes GOTCHA 3 above work without also writing matching `credentialId:` values on
# both. This linkage is not just cosmetic: it is the same key `BackendQuarantine`/cool-off use,
# so a usage-limit hit on `opus` now quarantines `sonnet` too (and vice versa) — correct, since
# they are one Claude account, but worth knowing before you wonder why an unrelated-looking
# profile went quarantined.
#
# WHEN TO OVERRIDE — set explicit, DIFFERENT `credentialId:` values on two `subscription: true`
# profiles only when they are genuinely two separate Claude logins on the same host (a real,
# if unusual, setup). An explicit `credentialId` always wins over the sentinel, so this is the
# one way to keep two subscription profiles from being treated as one account for lead-seat
# counting AND for quarantine/cool-off grouping alike.
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
@@ -503,6 +532,22 @@ fleet:
# recognised: give it a `profile:` and the daemon launches the shortfall when fewer than
# `instances` are live. Omit `profile:` and it is recognise-only, as before.
#
# `profile:` has a SECOND job as of fleetd #176, even for a recognise-only lead you never want
# auto-launched: it is also how fleetd learns which account this lead's own session shares. A
# `subscription: true` profile bills the operator's Claude account, and the lead itself is always
# a live `claude` session on that same account — `maxLoad` never counted that seat. If a lead
# entry here names a profile that shares a worker profile's account, `fleet_list`'s `free` for
# that worker profile subtracts the lead's live seat(s) automatically. "Shares the account" is
# decided by matching `effectiveCredentialId()`, which (fleetd #176 stage 2 — see THE SENTINEL,
# next to `credentialId:`, in THE WORKERS above) means: an explicit, matching `credentialId:` on
# both, OR — the common case, needing NO extra config — both being `subscription: true` with
# `credentialId` left unset, since those all share one implicit account-wide id. A lead on `opus`
# and workers on `sonnet` link automatically this way; they do NOT need the same profile name.
# Setting `profile:` on an already-running, recognise-only lead is safe — the daemon only launches
# the SHORTFALL below `instances`, so naming a profile here does not, by itself, start anything.
# Omit it and fleetd has no way to derive the sharing — there is no other reliable signal on the
# daemon's side — so that lead's seat goes uncounted, exactly as before this ticket.
#
# `tab:` (CB-579) is REQUIRED and is the only field identity depends on — the exact label of the
# tab hosting the lead, matched case-insensitively. Label the tab yourself and put that same
# string here, and the pane is recognised on the next rescan. Reopen the tab later, or the session
@@ -640,7 +640,8 @@ public final class Fleetd {
new FleetMcp.OutageSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, outagePolicy));
}, outagePolicy),
new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), leaders, leads)));
// CB-637: the receive half. Only constructed when a lead mailbox actually opened — with no
// coordinator (or an unreachable one) there is nothing to deliver, so no scheduler is
@@ -765,6 +766,68 @@ public final class Fleetd {
.orElse(null);
}
/**
* fleetd #176: per-profile factory for {@link FleetMcp.LeadSeatSource} — how many seats a
* profile's own live LEAD session(s) hold on the same Claude subscription.
*
* <p>{@code maxLoad} counts only members; the lead itself is a live {@code claude} session that
* is never moved off-subscription ({@code LeadLauncher} strips {@code ANTHROPIC_BASE_URL}/
* {@code AUTH_TOKEN} from a lead's env whatever its profile says), so a {@code subscription:
* true} profile's real ceiling is lower than its configured {@code maxLoad} by exactly the
* number of lead seats sharing that same account.
*
* <p><b>The derivation, and why this route was chosen over a new config key.</b> The link is
* {@code fleet.leaders.<name>.profile} — the field the operator already sets to name which
* {@code profiles:} entry a lead runs on (CB-557; see {@code fleetd.example.yaml}) — matched
* against the profile passed in here via {@link FleetConfig.Profile#effectiveCredentialId()},
* the same grouping key {@link dev.ltms.fleet.placement.BackendQuarantine} already uses to say
* two profiles share one account. Nothing new is added to the config schema: this reuses a field
* that already exists and already means "the profile this lead's own session runs on". A lead
* entry that names no {@code profile:} (recognise-only, CB-558) says nothing about which account
* it shares, and there is no other reliable signal on the daemon's side to derive that from — so
* such a lead contributes no seats, exactly as before this ticket. Making that lead's seat count
* requires the operator to add one line (`profile: sonnet` under its {@code fleet.leaders} entry)
* — a config statement, not a code change, and the smallest one available given the field
* already exists for a closely related purpose.
*
* <p>Only counts leads {@code liveLeadTerminals} currently reports — CB-531's live tab scan (or
* the legacy {@code primary.terminal} pin) — never every configured lead: an entry whose
* {@code instances} nobody has actually started is not really competing for a seat, and must not
* shrink capacity for one that was never live.
*
* @param profiles the live profile map, normally {@code () -> config.get().profiles()}
* in {@code main} — hot, like every other {@code maxLoad}/
* {@code credentialId} read {@link FleetMcp.CapacitySource} already does
* @param leaders {@code fleet.leaders}, read once at startup like the rest of that
* block ({@code fleetd.example.yaml} notes it is not hot) — passed as a
* plain map, never re-read from {@code config.get()}
* @param liveLeadTerminals terminal_id → lead name for every CURRENTLY recognised lead, normally
* the same supplier {@link dev.ltms.fleet.auth.CallerResolver#leads()}
* and {@code LeadCoordLoop} already consult
*/
static Function<String, Integer> leadSeatLookup(Supplier<Map<String, FleetConfig.Profile>> profiles,
Map<String, FleetConfig.Leader> leaders, Supplier<Map<String, String>> liveLeadTerminals) {
return profileName -> {
FleetConfig.Profile target = profiles.get().get(profileName);
if (target == null || !target.isSubscription()) {
return 0;
}
String targetCredential = target.effectiveCredentialId();
int seats = 0;
for (String leadName : liveLeadTerminals.get().values()) {
FleetConfig.Leader lead = leaders.get(leadName);
if (lead == null || lead.profile() == null || lead.profile().isBlank()) {
continue;
}
FleetConfig.Profile leadProfile = profiles.get().get(lead.profile());
if (leadProfile != null && targetCredential.equals(leadProfile.effectiveCredentialId())) {
seats++;
}
}
return seats;
};
}
/**
* fleetd #248 / fleetd#201 Unit 5: package-private factory for the per-target backend-error
* pattern lookup {@link CompletionResolver} classifies a pane scrape against. Closes over the
@@ -92,15 +92,9 @@ import java.util.regex.PatternSyntaxException;
* fleetd's own process, so fleetd's own {@code $SHELL} says nothing about what
* that pane runs. There is no channel to ask herdr for another user's shell, so
* this must be told, never guessed. {@code null}/blank (or a value not ending
* in {@code zsh}) is treated the same as "not zsh". fleetd #155: what that means
* now depends on {@code memberCredentials.policy}. Under {@code deny-by-default}
* it stays a degraded control, never a refusal to spawn — the pane-creation
* overlay is unaffected by shell type, so the spawn proceeds with a WARN naming
* the shell. Under {@code allow-list} the spawn is REFUSED instead: that policy's
* whole point is a control a sourced file cannot undo, so silently falling back
* to the weaker overlay would be the same "control silently does nothing" defect
* #155 exists to remove — configure this field (or move the account to zsh, or
* switch policy back to {@code deny-by-default}) to unblock the spawn.
* in {@code zsh}) is treated the same as "not zsh": the {@code
* memberCredentials.policy: allow-list} ZDOTDIR scrub is skipped in favour of
* the CB-596 sentinel overlay — a degraded control, never a refusal to spawn.
* When {@code memberHerdrSocket} is NOT configured this field is never
* consulted at all; fleetd keeps reading its own {@code $SHELL}, exactly as
* before this field existed.
@@ -650,14 +644,46 @@ public record FleetConfig(
}
/**
* The credential group this profile quarantines with (CB-578 stage B): the configured
* {@link #credentialId} when set, else this profile's own name — so an unconfigured profile
* quarantines alone, exactly as it did before this field existed. Two profiles that set the
* same non-blank {@code credentialId} share one quarantine: a {@code BACKEND_EXHAUSTED}
* classification on either one quarantines both.
* The shared credential id every {@code subscription: true} profile falls back to when it
* sets no explicit {@link #credentialId} (fleetd #176 stage 2, correcting an inert first cut
* of that ticket). A subscription profile has no credential of its own to fall back to its
* name for: it authenticates as the operator's own Claude login, and a host has exactly one
* of those, whatever names the operator gives the profiles running on it. Falling back to the
* profile's own name (the way an ordinary off-subscription profile does) would keep two
* subscription profiles on one login apart from each other, which is the opposite of what
* "one account" means.
*
* <p>Measured live and what it broke: a lead on profile {@code opus}, members on profile
* {@code sonnet}, same Claude login, neither setting {@code credentialId}. Before this
* sentinel, {@code opus.effectiveCredentialId()} was {@code "opus"} and {@code sonnet
* .effectiveCredentialId()} was {@code "sonnet"} — so fleetd #176's lead-seat matcher (and,
* this sentinel now also fixes, {@code CompositePeerLauncher}'s quarantine/cool-off spawn
* refusal and {@code BackendOutagePolicy}'s incident grouping) silently never linked them: the
* fix shipped, and stayed inert on the one host it was written for.
*/
public static final String SUBSCRIPTION_CREDENTIAL_ID = "<subscription>";
/**
* The credential group this profile quarantines with (CB-578 stage B; extended fleetd #176
* stage 2 — see {@link #SUBSCRIPTION_CREDENTIAL_ID}): the configured {@link #credentialId}
* when set — that always wins, so an operator with two separate Claude logins on one host can
* still keep them apart. Otherwise, a {@code subscription: true} profile falls back to
* {@link #SUBSCRIPTION_CREDENTIAL_ID} rather than its own name; an ordinary off-subscription
* profile falls back to its own name, exactly as it did before this field existed, so an
* unconfigured off-subscription profile still quarantines alone.
*
* <p>A {@code BACKEND_EXHAUSTED} (or repeated backend-error) classification on any profile
* sharing the result quarantines/cools off every profile that shares it — including, now,
* every {@code subscription: true} profile with no explicit {@code credentialId}. That is
* intended, not incidental: one Claude subscription hitting a usage limit really does take out
* every profile running on it, the same way {@code credentialId: openai-shared} already lets
* two OpenAI-backed profiles share one quarantine.
*/
public String effectiveCredentialId() {
return (credentialId == null || credentialId.isBlank()) ? profile : credentialId;
if (credentialId != null && !credentialId.isBlank()) {
return credentialId;
}
return isSubscription() ? SUBSCRIPTION_CREDENTIAL_ID : profile;
}
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
@@ -952,7 +978,15 @@ public record FleetConfig(
* survives restarts of the agent inside it — so identity is now the tab label alone.
*
* @param profile the {@code profiles:} entry to launch this lead on when one must
* be created; {@code null} ⇒ recognise-only, never create
* be created; {@code null} ⇒ recognise-only, never create.
* <p>fleetd #176: also the field {@code Fleetd.leadSeatLookup} reads
* to learn which account this lead's own live session shares — set it
* (safely, even on an already-running recognise-only lead: naming a
* profile here never starts anything beyond {@code instances}) so a
* {@code subscription: true} worker profile sharing its
* {@code effectiveCredentialId()} has this lead's seat subtracted from
* {@code fleet_list}'s {@code free}. {@code null} here also means this
* lead's seat cannot be derived and is not counted.
* @param tab the exact tab label hosting this lead, matched case-insensitively;
* the only field identity depends on. Required — a lead with no
* {@code tab} can never be discovered, launched or not
@@ -96,6 +96,8 @@ public final class FleetMcp {
private final QuarantineSource quarantine;
/** fleetd #201 Unit 5: SEPARATE from {@link #quarantine} — see {@link OutageSource}'s doc. */
private final OutageSource outage;
/** fleetd #176: SEPARATE from both of the above — see {@link LeadSeatSource}'s doc. */
private final LeadSeatSource leadSeats;
/** CB-637: this daemon's lead-to-lead channel; {@code null} when no coordinator is configured. */
private final LeadChannel leadChannel;
@@ -135,6 +137,24 @@ public final class FleetMcp {
}
}
/**
* fleetd #176: the seats a profile's own live LEAD session(s) hold on the same Claude
* subscription — the third reason (alongside {@link QuarantineSource} and {@link OutageSource})
* {@code free} can overstate what a fresh {@code fleet_spawn} would actually get.
*
* <p>{@code maxLoad} counts only <em>members</em>, never the lead itself. But a
* {@code subscription: true} profile bills the operator's own Claude account, and the lead is
* always a live {@code claude} session on that same account (it is never moved off-subscription
* — see {@code LeadLauncher}). So a fan-out that fills every member slot still leaves the lead's
* own seat unaccounted for, and the daemon reports a slot that was never really free. See
* {@code Fleetd.leadSeatLookup} for how the count is derived — from {@code fleet.leaders.<name>
* .profile} and each profile's {@code effectiveCredentialId()}, never a hardcoded constant.
*/
public record LeadSeatSource(Function<String, Integer> seatsFor) {
/** Inert source — no profile is ever reported as sharing a seat with a lead. */
public static LeadSeatSource none() { return new LeadSeatSource(_ -> 0); }
}
/**
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
@@ -149,7 +169,7 @@ public final class FleetMcp {
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine) {
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
healthCoverage, quarantine, null, OutageSource.none());
healthCoverage, quarantine, null, OutageSource.none(), LeadSeatSource.none());
}
/**
@@ -163,12 +183,12 @@ public final class FleetMcp {
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, LeadChannel leadChannel) {
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
healthCoverage, quarantine, leadChannel, OutageSource.none());
healthCoverage, quarantine, leadChannel, OutageSource.none(), LeadSeatSource.none());
}
/**
* As above, with fleetd #201 Unit 5 cool-off facts for {@code fleet_list}/{@code fleet_profiles}
* (see {@link OutageSource}). This is what {@code Fleetd.main} actually wires up.
* (see {@link OutageSource}).
*
* @param outage required — pass {@link OutageSource#none()} for a caller that does not want the
* feature, never a defaulting overload (the same rule {@code quarantine} follows).
@@ -177,10 +197,28 @@ public final class FleetMcp {
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage) {
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
healthCoverage, quarantine, leadChannel, outage, LeadSeatSource.none());
}
/**
* As above, with fleetd #176 lead-seat facts (see {@link LeadSeatSource}). This is what
* {@code Fleetd.main} actually wires up.
*
* @param leadSeats required — pass {@link LeadSeatSource#none()} for a caller that does not want
* the feature, never a defaulting overload (the same rule {@code quarantine} and
* {@code outage} follow).
*/
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage,
LeadSeatSource leadSeats) {
this.leadChannel = leadChannel;
this.capacity = capacity;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
this.outage = Objects.requireNonNull(outage, "outage");
this.leadSeats = Objects.requireNonNull(leadSeats, "leadSeats");
this.healthCoverage = healthCoverage;
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
@@ -310,7 +348,7 @@ public final class FleetMcp {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
callers == null ? Map.of() : callers.leads(),
leadSeats, callers == null ? Map.of() : callers.leads(),
callerTerminal(exchange),
leadChannel == null ? null : leadChannel.selfCoordId());
};
@@ -994,7 +1032,8 @@ public final class FleetMcp {
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, OutageSource outage,
Map<String, String> leads, String selfTerm) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage, leads, selfTerm, null);
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
LeadSeatSource.none(), leads, selfTerm, null);
}
/**
@@ -1011,7 +1050,7 @@ public final class FleetMcp {
QuarantineSource quarantine, Map<String, String> leads, String selfTerm,
String selfCoordId) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, OutageSource.none(),
leads, selfTerm, selfCoordId);
LeadSeatSource.none(), leads, selfTerm, selfCoordId);
}
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
@@ -1019,6 +1058,16 @@ public final class FleetMcp {
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, OutageSource outage,
Map<String, String> leads, String selfTerm, String selfCoordId) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
LeadSeatSource.none(), leads, selfTerm, selfCoordId);
}
/** As above, plus fleetd #176 lead-seat facts (see {@link LeadSeatSource}). */
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, OutageSource outage,
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
String selfCoordId) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
@@ -1045,7 +1094,7 @@ public final class FleetMcp {
}
if (capacity.available()) result.put("capacity", profiles.stream()
.map(profile -> capacityView(profile, capacity.liveCount(), capacity.maxLoad(), roster, messages,
capacity.clock().getAsLong(), quarantine, outage)).toList());
capacity.clock().getAsLong(), quarantine, outage, leadSeats)).toList());
return text(json(result));
} catch (HerdrException e) {
return error("herdr error listing the fleet: " + e.getMessage());
@@ -1084,20 +1133,33 @@ public final class FleetMcp {
* {@code credentialId}/{@code coolingOffForSeconds}, but never {@code quarantinedForSeconds} —
* that key is added only when exhaustion quarantine is ALSO active for this profile, since the
* two checks are independent and either, both, or neither can be true.
*
* <p>fleetd #176: {@code maxLoad} counts panes, not subscription seats — it never counted the
* lead's own seat on a {@code subscription: true} profile's account. {@link LeadSeatSource}
* reports that count (0 for a non-subscription profile, or when no live lead shares its
* credential), and it is subtracted from {@code free} the same way {@code live} already is —
* {@code maxLoad} itself is left untouched, so the row still reports the configured cap. The
* {@code leadSeats} key is added only when the count is positive, for the same
* byte-identical-when-unused reason as the quarantine/cool-off keys above.
*/
private static Map<String, Object> capacityView(String profile, Function<String, Integer> liveCount,
Function<String, Integer> maxLoad, List<MemberSession> roster,
MessageService messages, long nowNanos, QuarantineSource quarantine,
OutageSource outage) {
OutageSource outage, LeadSeatSource leadSeats) {
Integer cap = maxLoad.apply(profile);
int live = liveCount.apply(profile);
int leadSeatCount = leadSeats.seatsFor().apply(profile);
int reclaimable = (int) roster.stream().filter(s -> profile.equals(s.profile()))
.filter(s -> (s.state() == MemberSession.State.READY || s.state() == MemberSession.State.DONE))
.filter(s -> messages == null || (!messages.hasAcceptedDelivery(s.terminalId()) && !messages.hasInboxMessage(s.terminalId())))
.count();
Map<String, Object> row = new LinkedHashMap<>();
row.put("profile", profile); row.put("maxLoad", cap); row.put("live", live);
row.put("free", cap == null ? null : Math.max(0, cap - live)); row.put("reclaimable", reclaimable);
row.put("free", cap == null ? null : Math.max(0, cap - live - leadSeatCount));
row.put("reclaimable", reclaimable);
if (leadSeatCount > 0) {
row.put("leadSeats", leadSeatCount);
}
String credentialId = quarantine.credentialIdFor().apply(profile);
if (credentialId != null) {
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
@@ -1253,11 +1315,7 @@ public final class FleetMcp {
+ "worktree:<ticket-slug> to provision an isolated git worktree. Pass resumeSessionId "
+ "to relaunch onto a prior conversation instead of starting cold — this requires an "
+ "explicit profile whose backend supports it (fleet_list shows agentSessionId for "
+ "resumable members; it is absent for a member fleetd cannot reliably re-identify, "
+ "e.g. an opencode member spawned without a worktree), and is refused otherwise "
+ "rather than silently starting fresh. For an opencode profile, resumeSessionId "
+ "itself also requires worktree:true/<slug> on THIS spawn — without one fleetd can "
+ "never re-verify which conversation it actually resumed (fleetd #249). "
+ "resumable members), and is refused otherwise rather than silently starting fresh. "
+ "sessionName gives the member a display name in its own UI when the backend supports "
+ "one. Returns the member's sessionId (use with fleet_send) and paneId (use with "
+ "fleet_stop).",
@@ -1268,7 +1326,7 @@ public final class FleetMcp {
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
"ticket", stringProp("Ticket slug when worktree:true"),
"sessionName", stringProp("Logical display name for the member's own session, when its backend supports one"),
"resumeSessionId", stringProp("A prior member's agentSessionId (from fleet_list) to resume — requires an explicit profile that supports it, and (for opencode) a worktree on this spawn too")),
"resumeSessionId", stringProp("A prior member's agentSessionId (from fleet_list) to resume — requires an explicit profile that supports it")),
List.of()));
}
@@ -1289,14 +1347,9 @@ public final class FleetMcp {
+ "discover a peer lead without being told its address. 'members' are the "
+ "sessions delegated to — each with sessionId, paneId, role (architect/dev/"
+ "reviewer), profile (the backend it runs on), state, optional "
+ "worktree/branch/owner/agentSessionId, and live herdr status. agentSessionId, "
+ "when present, is the id to pass as fleet_spawn's resumeSessionId to relaunch "
+ "onto that same conversation. It is ABSENT — not a guess — for a member fleetd "
+ "cannot reliably re-identify: some backends (e.g. opencode) resolve it from the "
+ "member's working directory, which only uniquely identifies a member when it "
+ "was spawned into its own fleetd-provisioned worktree (worktree:true/<slug>); a "
+ "member spawned without one shares its directory with others and never reports "
+ "an id, however long it runs (fleetd #249). An empty 'members' "
+ "worktree/branch/owner/agentSessionId (the id to pass as fleet_spawn's "
+ "resumeSessionId to relaunch onto that same conversation, when the backend "
+ "supports it), and live herdr status. An empty 'members' "
+ "means no members are spawned; it says nothing about peers. When capacity "
+ "facts are configured, a 'capacity' row per profile also reports free: 0 for "
+ "a quarantined profile's credential (see fleet_profiles), whatever its "
@@ -495,7 +495,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* <p><b>Additive, not a rewrite.</b> {@code .claude.json} is large (tens of KB, dozens of
* projects) and Claude Code itself rewrites it while running, so this reads the file as a JSON
* tree (missing or unreadable → treated as an empty object) and changes only
* {@code projects.<cwd>.hasTrustDialogAccepted} (that key alone — see fleetd #247) —
* {@code projects.<cwd>.hasTrustDialogAccepted} / {@code .hasCompletedProjectOnboarding} —
* every other top-level key and every other project entry is written back untouched. Only the
* one project entry for {@code cwd} is replaced/created; an existing entry for a DIFFERENT cwd
* (or the operator's own project history) is never touched.
@@ -505,7 +505,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* peer that starts without the seed still starts; it just may hit the dialog fleetd #149
* describes.
*
* <p><b>Gated to a provisioned worktree</b> ({@link HerdrPeerLauncher#isProvisionedWorktree}) — see that
* <p><b>Gated to a provisioned worktree</b> ({@link #isProvisionedWorktree}) — see that
* method's javadoc for the incident that made this gate mandatory, not optional: this must
* never run against a real checkout or an un-configured fallback cwd, only the exact
* always-fresh-directory population fleetd #149 describes.
@@ -516,7 +516,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* sibling-temp-file + {@code ATOMIC_MOVE}, never a truncate-in-place) so a crash mid-write or a
* concurrent reader never observes a half-written file, and through {@link #TRUST_JSON_LOCK} so
* two concurrent spawns' entries both survive instead of the second write silently discarding
* the first. Both exist because of a real incident: see {@link HerdrPeerLauncher#isProvisionedWorktree}'s javadoc
* the first. Both exist because of a real incident: see {@link #isProvisionedWorktree}'s javadoc
* and {@link #writeAtomically}'s javadoc.
*
* @param configDir the profile's {@code CLAUDE_CONFIG_DIR} ({@code cfg.configDir()}), or
@@ -558,18 +558,8 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
if (!(projectNode instanceof ObjectNode)) {
projects.set(cwd, project);
}
// fleetd #247: ONLY hasTrustDialogAccepted. We used to write
// hasCompletedProjectOnboarding beside it; do not put it back. Measured on
// 2026-09-03, minutes after a live spawn seeded this file: 28 of 28 project
// entries carried hasTrustDialogAccepted and 0 of 28 carried the onboarding key
// — including the 27 entries Claude Code wrote for itself. Claude Code
// normalises the whole file when it saves and drops that key every time, so
// writing it achieved nothing except making the next reader think it mattered.
// The member reached idle with the trust flag alone, which is the only outcome
// this seed exists for. If a future Claude Code needs the second flag the
// symptom returns as the trust dialog fleetd #149 describes — re-measure then,
// do not restore it on a guess.
project.put("hasTrustDialogAccepted", true);
project.put("hasCompletedProjectOnboarding", true);
writeAtomically(target, TRUST_JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root));
} catch (Exception e) {
log.debug("cannot seed workspace-trust entry for cwd '{}' into '{}'", cwd, target, e);
@@ -588,7 +578,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* <p><b>fleetd #149 incident.</b> The original implementation used
* {@code Files.writeString(target, content)} directly, which truncates {@code target} in place
* before writing the replacement bytes. Combined with an ungated {@code cwd} (see
* {@link HerdrPeerLauncher#isProvisionedWorktree}'s javadoc), a mutation-testing run hit that truncation window
* {@link #isProvisionedWorktree}'s javadoc), a mutation-testing run hit that truncation window
* against the operator's real {@code ~/.claude.json} and left it at 178 bytes. The gate closes
* <em>which file</em> this can ever target; this closes <em>how</em> the target is written, so
* that even a legitimate write against a real, live, concurrently-read {@code .claude.json}
@@ -638,6 +628,32 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
}
}
/**
* Whether {@code cwd} is a fleetd-provisioned git worktree — signalled the same way
* {@link #writeIdeOverlay} already gates on: a {@code .git} that is a <strong>regular
* file</strong> holding a {@code gitdir:} pointer, as opposed to a real checkout's {@code .git}
* <strong>directory</strong>. {@code null}/blank never qualifies.
*
* <p>Shared by every write that must land only in a worktree fleetd itself created for a
* member — never in a real checkout, an arbitrary configured directory, or (see the incident
* below) the daemon's own fallback cwd.
*
* <p><b>fleetd #149 incident.</b> {@link #seedTrustDialog} originally ran unconditionally on
* any non-blank {@code cwd}. Most of this launcher's OWN tests spawn a profile with no
* {@code cwd} configured, so the base class's {@code resolveCwd} falls through to the real
* {@code user.dir} — and with no {@code configDir} either (also the common case in this
* file's fixtures), the seed's target falls through the same way to the real
* {@code ~/.claude.json}. Running this repo's own test suite corrupted the operator's actual
* config file (it shrank from ~72 KB to a single seeded entry) the first time a mutation
* happened to make the write non-additive. Gating both cwd-targeted writes on "this is a
* worktree fleetd provisioned" — exactly the population fleetd #149 describes
* ({@code worktree: true} always lands in a brand-new directory) — makes that class of write
* impossible against a real checkout or an untouched fallback cwd, in production or in tests.
*/
private static boolean isProvisionedWorktree(String cwd) {
return cwd != null && !cwd.isBlank() && Files.isRegularFile(Path.of(cwd, ".git"));
}
/** {@code s}, or {@code null} when {@code s} is null/blank — the charter-presence test used above. */
private static String nonBlank(String s) {
return (s == null || s.isBlank()) ? null : s;
@@ -184,11 +184,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
*/
private final ConcurrentMap<String, Path> zdotdirByPane = new ConcurrentHashMap<>();
/**
* Guards {@link #warnNonZsh} to one WARN per launcher instance, not one per spawn. fleetd #155:
* only the {@code policy: deny-by-default} path still warns on a non-zsh shell — the {@code
* policy: allow-list} path refuses the spawn instead (see {@link #applyEnvironmentAllowListPolicy}).
*/
/** Guards {@link #warnNonZsh} to one WARN per launcher instance, not one per spawn. */
private final AtomicBoolean nonZshShellWarned = new AtomicBoolean();
/** Live config provides URI environment names that must never enter member panes. */
private final Supplier<FleetConfig> config;
@@ -362,42 +358,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
return Files.isRegularFile(candidate) ? candidate : null;
}
/**
* Whether {@code cwd} is a fleetd-provisioned git worktree — signalled the same way
* {@code ClaudeCodeLauncher#writeIdeOverlay} already gates on: a {@code .git} that is a
* <strong>regular file</strong> holding a {@code gitdir:} pointer, as opposed to a real
* checkout's {@code .git} <strong>directory</strong>. {@code null}/blank never qualifies.
*
* <p>Shared by every write (and, since fleetd #249, every identity read) that must land only
* in a worktree fleetd itself created for a member — never in a real checkout, an arbitrary
* configured directory, or (see the incident below) the daemon's own fallback cwd. Package-
* private (not {@code protected}) on purpose: {@link ClaudeCodeLauncher} and
* {@link OpenCodeLauncher} both call it, and same-package visibility is enough — no subclass
* outside this package needs it.
*
* <p><b>fleetd #149 incident.</b> {@code ClaudeCodeLauncher#seedTrustDialog} originally ran
* unconditionally on any non-blank {@code cwd}. Most of that launcher's OWN tests spawn a
* profile with no {@code cwd} configured, so the base class's {@code resolveCwd} falls
* through to the real {@code user.dir} — and with no {@code configDir} either (also the
* common case in that file's fixtures), the seed's target falls through the same way to the
* real {@code ~/.claude.json}. Running this repo's own test suite corrupted the operator's
* actual config file (it shrank from ~72 KB to a single seeded entry) the first time a
* mutation happened to make the write non-additive. Gating both cwd-targeted writes on "this
* is a worktree fleetd provisioned" — exactly the population fleetd #149 describes
* ({@code worktree: true} always lands in a brand-new directory) — makes that class of write
* impossible against a real checkout or an untouched fallback cwd, in production or in tests.
*
* <p><b>fleetd #249.</b> The same reasoning extends to a READ: {@code
* OpenCodeSessionDiscovery#sessionIdForDirectory} keys on {@code directory}, a heuristic that
* is only reliable when the directory is unique to this member — i.e., exactly the population
* this gate identifies. {@link OpenCodeLauncher} uses it to withhold {@code agentSessionId()}
* (report absence rather than a guess) and to refuse a {@code resumeSessionId} spawn that
* cannot be resolved reliably going forward.
*/
static boolean isProvisionedWorktree(String cwd) {
return cwd != null && !cwd.isBlank() && Files.isRegularFile(Path.of(cwd, ".git"));
}
// --- profile surface -----------------------------------------------------------------------
/** The configured peer profile names (what {@code spawn(profile)} accepts). */
@@ -1259,11 +1219,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
if (!creds.isAllowList()) {
overlayBlockedCredentials(workerEnv, creds);
logCredentialGap(creds, null);
// fleetd #155: this overlay itself does not depend on the shell (it lands in the
// pane-creation env map before any shell runs), so the spawn is never refused here —
// only policy=allow-list's ZDOTDIR scrub needs a login shell to run at all. Still worth
// telling the operator: the stronger post-shell control is unavailable on this shell.
warnNonZsh(memberLoginShell());
}
}
@@ -1316,16 +1271,10 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* dev.ltms.fleet.session.Worktrees#shareWithGroup} already uses, reused rather than
* inventing a second group key. Either one missing means the scrub cannot be guaranteed
* reachable by the member, which is the same "cannot guarantee the scrub runs" case as a
* non-zsh shell — that one still gets the fallback (see fleetd #155's javadoc note below
* on why a missing worktreeRoot/worktreeGroup is a different defect, out of that ticket's
* scope).</li>
* non-zsh shell, so it gets the identical fallback.</li>
* </ul>
* fleetd #155: the one exception to "never refuse to spawn" is a non-zsh login shell, right
* below. The operator asked for {@code policy: allow-list} specifically — its whole point is a
* control a sourced file cannot undo — so silently degrading to the weaker overlay is exactly
* the "silently does nothing" failure this ticket exists to remove. Every other branch in this
* method (worktreeRoot/worktreeGroup missing) keeps the old "degrade, never refuse" behaviour;
* that gap is real but is fleetd #213's scope, not this one.
* In every branch: never refuse to spawn. A degraded credential control must not become an
* outage for an opt-in feature.
*/
private Path applyEnvironmentAllowListPolicy(FleetConfig.Profile cfg, Launch launch) {
FleetConfig.MemberCredentials creds = memberCredentials == null ? null : memberCredentials.get();
@@ -1339,21 +1288,20 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// must not even be called on this path, only the explicit memberLoginShell: config can
// answer it. With memberHerdrSocket absent, nothing here changes: fleetd's own $SHELL is
// still the input, exactly as before this fix.
String loginShell = memberLoginShell();
String loginShell = memberHerdrSocket ? configuredMemberLoginShell() : resolveEnv("SHELL");
boolean zsh = isZshShell(loginShell);
if (!zsh) {
// fleetd #155: a non-zsh login shell ignores ZDOTDIR entirely — NO scrub would run. The
// operator explicitly asked for policy=allow-list's blocking control, so degrading to
// the weaker CB-596 overlay and spawning anyway would be the same "control silently does
// nothing" defect this ticket exists to close. Refuse instead — the caller (FleetMcp.spawn)
// catches IllegalArgumentException and surfaces the message to the operator.
throw new IllegalArgumentException(
"memberCredentials policy=allow-list requires the member's login shell to be "
+ "zsh, so the ZDOTDIR scrub can run after it — refusing to spawn under "
+ "login shell '" + (loginShell == null ? "<unset>" : loginShell)
+ "'. Configure memberLoginShell: as a zsh path (only read when "
+ "memberHerdrSocket is set), move the member's OS account onto zsh, or "
+ "set memberCredentials.policy: deny-by-default instead.");
// A non-zsh login shell ignores ZDOTDIR entirely: NO scrub would run, so pretending
// otherwise would be worse than saying so. Warn loudly and fall back to the CB-596
// sentinel overlay over the enumerated known: names — weaker (a sourced file can undo
// it), but strictly better than nothing. Deliberately no "allowed N of M" line here: the
// scrub this count describes does not run on this path, so printing it would tell an
// operator that a fraction of names were blocked when the real number blocked is zero.
// logCredentialGap's WARN (below) is the only signal for this path.
warnNonZsh(loginShell);
overlayBlockedCredentials(launch.env(), creds);
logCredentialGap(creds, null);
return null;
}
Path parentDir;
String group = null;
@@ -1408,20 +1356,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
return cfg == null ? null : cfg.memberLoginShell();
}
/**
* fleetd #155: the member pane's login shell, using the same {@code memberHerdrSocket} routing
* {@link #applyEnvironmentAllowListPolicy} already used — fleetd's own {@code $SHELL} decides it
* when {@code memberHerdrSocket} is absent (today's only mode); the configured {@code
* memberLoginShell:} decides it otherwise, since fleetd's own {@code $SHELL} names a different
* user's shell once member panes run under a different OS user. Shared by both {@link
* #applyEnvironmentAllowListPolicy} (allow-list: refuses on non-zsh) and {@link
* #applyMemberCredentialPolicy} (deny-by-default: warns on non-zsh) so the two policies agree on
* what "the member's shell" means.
*/
private String memberLoginShell() {
return memberHerdrSocketConfigured() ? configuredMemberLoginShell() : resolveEnv("SHELL");
}
/**
* fleetd #213: {@code worktreeRoot}, as the ZDOTDIR scrub's parent directory under {@code
* memberHerdrSocket}, or {@code null} when unconfigured — the same "cannot guarantee the scrub
@@ -1506,23 +1440,15 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
private static final String SSH_AUTH_SOCK = MemberEnvAllowList.SSH_AUTH_SOCK;
/**
* fleetd #155: called ONLY from the {@code policy: deny-by-default} branch of {@link
* #applyMemberCredentialPolicy} — under {@code policy: allow-list}, a non-zsh shell now REFUSES
* the spawn instead (see {@link #applyEnvironmentAllowListPolicy}), since that policy's whole
* point is a control a sourced file cannot undo. deny-by-default's own overlay does not depend
* on the shell (it lands in the pane-creation env map before any shell runs), so this is a
* heads-up, not a defect report: say once per launcher instance, naming the shell, that the
* stronger allow-list control is unavailable here — never silently.
* CB-633: a non-zsh login shell means the allow-list control CANNOT run — say so once per
* launcher instance, naming the shell, instead of failing silently.
*/
private void warnNonZsh(String shell) {
if (isZshShell(shell)) {
return;
}
if (nonZshShellWarned.compareAndSet(false, true)) {
log.warn("memberCredentials policy=deny-by-default: member login shell '{}' is NOT zsh — "
+ "the pane-creation credential overlay still applies here (it does not "
+ "depend on the shell), but policy=allow-list's stronger post-shell "
+ "ZDOTDIR scrub is unavailable on this shell.",
log.warn("memberCredentials policy=allow-list: member login shell '{}' is NOT zsh — "
+ "ZDOTDIR scrubbing cannot run, so members' inherited environment is "
+ "UNPROTECTED beyond the enumerated known: fallback. Move herdr onto a "
+ "zsh account or switch policy back to deny-by-default.",
shell == null ? "<unset>" : shell);
}
}
@@ -672,29 +672,13 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
/** Add lazy on-disk session discovery to the base handle. */
@Override
public PeerHandle spawn(SpawnRequest req) {
String cwd = effectiveCwd(req);
// fleetd #249: refuse rather than silently resume into unverifiable territory. opencode's
// `-s <id>` flag itself resumes precisely — the resolved id is what fails, not the resume —
// but resolvedSessionId() below can never confirm (or later re-report) this handle's own
// identity without a fleetd-provisioned worktree (isProvisionedWorktree(cwd)), because the
// directory is shared and sessionIdForDirectory's "most recently updated row" heuristic can
// pick a sibling's session. Refusing here, before anything spawns, beats letting the member
// start and only then discovering fleetd can never again verify who it actually is.
if (req.resumeSessionId() != null && !req.resumeSessionId().isBlank()
&& !isProvisionedWorktree(cwd)) {
throw new IllegalArgumentException("resumeSessionId requires a fleetd-provisioned "
+ "worktree for an opencode profile — without one, this member's cwd is shared "
+ "with other sessions, so fleetd can never reliably confirm (now or later) which "
+ "conversation it is actually running (fleetd #249). Pass fleet_spawn{worktree:"
+ "<ticket-slug>} to resume this member.");
}
PeerHandle inner = super.spawn(req);
// fleetd #175: the same profile config buildLaunch resolved for this spawn (requireProfile
// is deterministic on req.profileName(), so re-resolving here costs a map lookup, not a
// second decision) — SessionAwareHandle needs cfg.model() to know what THIS session should
// be running.
FleetConfig.Profile cfg = requireProfile(req.profileName());
return new SessionAwareHandle(inner, discovery, cwd, cfg,
return new SessionAwareHandle(inner, discovery, effectiveCwd(req), cfg,
this::memberHerdrSocketConfigured, discoveryUnavailableWarned, exhaustionSink);
}
@@ -734,17 +718,6 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* the directory right now."
*/
private final AtomicReference<String> resolvedSessionId = new AtomicReference<>();
/**
* fleetd #249: whether {@link #cwd} is a fleetd-provisioned git worktree
* ({@link HerdrPeerLauncher#isProvisionedWorktree}), computed once at spawn time since
* {@code cwd} never changes for this handle. When {@code false} the directory is shared
* with other sessions (the default no-worktree spawn inherits the lead's own cwd), so
* {@link OpenCodeSessionDiscovery#sessionIdForDirectory}'s "most recently updated row for
* this directory" heuristic can and does pick another session's row — see that class's
* javadoc. {@link #agentSessionId()} refuses to guess in that case: it reports absent
* rather than a possibly-foreign id.
*/
private final boolean worktreeProvisioned;
SessionAwareHandle(PeerHandle delegate, OpenCodeSessionDiscovery discovery, String cwd,
FleetConfig.Profile cfg,
@@ -758,7 +731,6 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
this.discoveryUnavailable = discoveryUnavailable;
this.discoveryUnavailableWarned = discoveryUnavailableWarned;
this.exhaustionSink = exhaustionSink;
this.worktreeProvisioned = isProvisionedWorktree(cwd);
}
@Override
@@ -788,9 +760,6 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
// built from (see OpenCodeLauncher#defaultDiscoveryRoot's javadoc for the full
// reasoning). Scanning fleetd's own $HOME under that config would only ever find "no
// row" and read as "resume unsupported" — declare it unavailable instead, once, loudly.
// Checked before the fleetd #249 worktree gate below: this OS-user mismatch makes
// discovery unusable regardless of whether cwd happens to be a provisioned worktree, so
// it earns the one-time WARN either way.
if (discoveryUnavailable.getAsBoolean()) {
if (discoveryUnavailableWarned.compareAndSet(false, true)) {
log.warn("opencode session discovery unavailable: memberHerdrSocket is "
@@ -802,16 +771,6 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
}
return null;
}
// fleetd #249: cwd is shared with other sessions unless fleetd itself provisioned this
// worktree, and sessionIdForDirectory's directory-keyed heuristic cannot tell this
// member's row apart from a sibling's in that case (measured: a three-day-old row from
// a different profile). Refuse to guess — absent is the honest answer, and it is what
// this codebase already returns elsewhere for absent evidence (fleetd #175's UNKNOWN).
// No WARN here: unlike discoveryUnavailable above, this is the ordinary, expected shape
// of the large majority of spawns (no worktree requested), not a configuration gap.
if (!worktreeProvisioned) {
return null;
}
// fleetd #234: once resolved, stay resolved. Re-deriving from `directory` on every call
// would let this handle's identity drift to a sibling session that later shares the
// same cwd and writes a newer row — see resolvedSessionId's javadoc.
@@ -0,0 +1,156 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #176: {@link Fleetd#leadSeatLookup} is the factory {@code Fleetd.main} wires into {@code
* FleetMcp.LeadSeatSource} so {@code fleet_list}'s {@code free} can subtract the seat(s) a
* {@code subscription: true} profile's own live LEAD session holds on that same account —
* {@code maxLoad} never counted the lead, only members. {@code FleetdLeadSeatWiringTest} proves
* {@code main} still passes this factory's result in; this class proves the factory's own matching
* logic: subscription-only, credential-matched, and counting only CURRENTLY LIVE leads.
*/
class FleetdLeadSeatLookupTest {
private static FleetConfig.Profile subscriptionProfile(String name, String credentialId) {
return new FleetConfig.Profile(name, null, "claude-sonnet-5", null, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, 3, true, null, credentialId, null);
}
private static FleetConfig.Profile offSubscriptionProfile(String name, String credentialId) {
return new FleetConfig.Profile(name, "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
null, "tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, 2, false, null, credentialId, null);
}
private static FleetConfig.Leader leadOnProfile(String profile) {
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
}
@Test
@DisplayName("a live lead sharing the target profile's credential counts as one seat")
void liveLeadSharingCredentialCountsAsOneSeat() {
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
() -> Map.of("term_primary", "primary"));
assertEquals(1, lookup.apply("sonnet"));
}
@Test
@DisplayName("no live lead names this profile ⇒ zero seats, exactly as before this ticket")
void noLiveLeadOnTheProfileCountsAsZero() {
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders, Map::of);
assertEquals(0, lookup.apply("sonnet"));
}
@Test
@DisplayName("a lead entry with no `profile:` (recognise-only) contributes no seats — cannot be derived")
void recogniseOnlyLeadWithNoProfileContributesNothing() {
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
FleetConfig.Leader recogniseOnly = new FleetConfig.Leader(null, "lead: primary", 1, "lead:", 10,
"claude", "claude-sonnet-5");
Map<String, FleetConfig.Leader> leaders = Map.of("primary", recogniseOnly);
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
() -> Map.of("term_primary", "primary"));
assertEquals(0, lookup.apply("sonnet"));
}
@Test
@DisplayName("a non-subscription profile never has a lead seat subtracted, whatever the credential match")
void nonSubscriptionProfileIsNeverAdjusted() {
Map<String, FleetConfig.Profile> profiles = Map.of(
"terra", offSubscriptionProfile("terra", "shared-openai"),
"sonnet", subscriptionProfile("sonnet", "shared-openai"));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
() -> Map.of("term_primary", "primary"));
assertEquals(0, lookup.apply("terra"), "terra is not subscription:true, so it must never be adjusted");
}
@Test
@DisplayName("explicit, different credentialIds still separate two subscription profiles (post fleetd #176 "
+ "stage 2 sentinel)")
void differentCredentialIsNotCounted() {
Map<String, FleetConfig.Profile> profiles = Map.of(
"sonnet", subscriptionProfile("sonnet", "claude-account-a"),
"opus", subscriptionProfile("opus", "claude-account-b"));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
() -> Map.of("term_primary", "primary"));
assertEquals(0, lookup.apply("sonnet"), "different accounts must never be conflated into one seat count "
+ "— an explicit credentialId on both sides must still win over the subscription sentinel, so an "
+ "operator with two separate Claude logins on one host can keep them apart");
}
/**
* fleetd #176 stage 2 — the exact live shape that shipped inert: a lead on subscription profile
* {@code opus}, members on a DIFFERENTLY NAMED subscription profile {@code sonnet}, same Claude
* login, and NEITHER profile sets {@code credentialId}. Every other test in this class puts the
* lead on the SAME profile name as the target, which happened to keep working even with the old
* fall-back-to-profile-name {@code effectiveCredentialId()} — this is the one that did not, and
* its absence is what let the stage-1 fix ship without ever catching the bug it was filed for.
*/
@Test
@DisplayName("[LIVE SHAPE] lead on a DIFFERENT subscription profile, same account, neither sets "
+ "credentialId ⇒ still counts as a seat")
void leadOnADifferentSubscriptionProfileSameAccountStillCountsAsASeat() {
Map<String, FleetConfig.Profile> profiles = Map.of(
"opus", subscriptionProfile("opus", null),
"sonnet", subscriptionProfile("sonnet", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
() -> Map.of("term_primary", "primary"));
assertEquals(1, lookup.apply("sonnet"), "opus and sonnet are both subscription:true with no explicit "
+ "credentialId, so they share one Claude login and the lead's live seat on opus must be charged "
+ "against sonnet too — this is the live host's actual shape (fleetd #176 stage 2)");
}
@Test
@DisplayName("two live instances of the same lead count as two seats")
void twoLiveInstancesOfTheSameLeadCountAsTwoSeats() {
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
() -> Map.of("term_a", "primary", "term_b", "primary"));
assertEquals(2, lookup.apply("sonnet"));
}
@Test
@DisplayName("an unconfigured target profile resolves to zero, not a thrown exception")
void unconfiguredTargetProfileIsZero() {
Function<String, Integer> lookup = Fleetd.leadSeatLookup(Map::of, Map.of(), Map::of);
assertEquals(0, lookup.apply("ghost"));
}
@Test
@DisplayName("live leads are read through the supplier on every call, not snapshotted")
void liveLeadsAreReadThroughOnEveryCall() {
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
java.util.Map<String, String> live = new java.util.HashMap<>();
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders, () -> live);
assertEquals(0, lookup.apply("sonnet"));
live.put("term_primary", "primary");
assertEquals(1, lookup.apply("sonnet"));
}
}
@@ -0,0 +1,43 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #176: {@code Fleetd.main} builds its {@code FleetMcp} from a 14-argument constructor whose
* last argument is a {@code FleetMcp.LeadSeatSource} wrapping {@link Fleetd#leadSeatLookup}. That
* argument is exactly the kind of wiring fleetd #248 warned about: dropping it (or swapping it for
* the inert {@code FleetMcp.LeadSeatSource.none()}) compiles with 0 errors and leaves every test
* that builds its own {@code FleetMcp}/{@code CapacitySource} directly — every test that predates
* this ticket — green, because none of them go through {@code main} at all.
*
* <p>{@link FleetdLeadSeatLookupTest} proves the factory's own matching logic; this class is the
* plain source-text assertion that proves {@code main} still passes its result in, mirroring
* {@code FleetdCompletionResolverWiringTest}'s approach for the same class of gap.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a
* {@code FleetMcp} and never runs {@code main}.
*/
class FleetdLeadSeatWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] FleetMcp's construction call still passes a LeadSeatSource built from leadSeatLookup(...)")
void fleetMcpConstructionStillWiresLeadSeatLookup() throws Exception {
String source = fleetdSource();
assertTrue(source.contains("new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), "
+ "leaders, leads))"),
"FleetMcp's construction call must still pass a LeadSeatSource built from "
+ "Fleetd.leadSeatLookup(...). Dropping it or swapping in "
+ "FleetMcp.LeadSeatSource.none() (fleetd #176's would-be silent regression, the same "
+ "shape as fleetd #248's measured mutations) compiles with 0 errors and leaves every "
+ "existing behavioural test green — this source check is what must go red instead.");
}
}
@@ -651,6 +651,48 @@ class FleetMcpTest {
assertEquals(2, out.split("\"free\":0", -1).length - 1, out);
}
/**
* fleetd #176 stage 2 (correcting the inert stage 1): two {@code subscription: true} profiles,
* {@code opus} and {@code sonnet}, neither setting an explicit {@code credentialId} — the exact
* shape measured on the live Mac fleet. This is INTENDED, not a regression: a real Claude usage
* limit on the one login behind both profiles really does take out every profile running on it,
* the same way {@code credentialId: openai-shared} already lets two OpenAI-backed profiles share
* one quarantine (see {@code everyProfileSharingTheQuarantinedCredentialReportsZeroFree} above).
* The {@code credentialIdFor} function here is built the same way {@code Fleetd.main} wires it —
* {@code profile -> profiles.get(profile).effectiveCredentialId()} — so this proves the actual
* config-driven behaviour, not just {@code capacityView}'s arithmetic with a hand-picked string.
*/
@Test
void quarantiningOneSubscriptionProfileZeroesFreeOnTheOtherSharingTheSameAccount() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
Map<String, FleetConfig.Profile> profiles = Map.of(
"opus", new FleetConfig.Profile("opus", null, "claude-opus-4", null, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, 3, true, null, null, null),
"sonnet", new FleetConfig.Profile("sonnet", null, "claude-sonnet-5", null, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, 3, true, null, null, null));
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine(FleetConfig.Profile.SUBSCRIPTION_CREDENTIAL_ID);
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(profile -> {
FleetConfig.Profile configured = profiles.get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, quarantine);
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 3,
() -> Set.of("opus", "sonnet"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
source, Map.of(), ""));
assertEquals(2, out.split("\"free\":0", -1).length - 1,
"opus and sonnet share one Claude login with neither setting credentialId, so quarantining "
+ "opus's account must also zero sonnet's free — this is intended, not a side effect: "
+ out);
assertEquals(2, out.split("\"credentialId\":\"" + FleetConfig.Profile.SUBSCRIPTION_CREDENTIAL_ID + "\"", -1)
.length - 1, out);
}
/** fleetd #201 Unit 5: cool-off forces {@code free:0} but never adds {@code quarantinedForSeconds}. */
@Test
void coolingOffProfileReportsZeroFreeButNeverQuarantinedForSeconds() {
@@ -767,6 +809,68 @@ class FleetMcpTest {
assertFalse(out.contains("quarantinedForSeconds"), out);
}
/**
* fleetd #176: this is the exact shape measured on the Mac fleet — {@code maxLoad:3, live:2},
* where one of the "free" three is really the lead's own seat on the same subscription. The old
* formula ({@code max(0, cap - live)}) reported {@code free:1}; the real ceiling is {@code 0}
* (two members plus the lead's own seat already fill all three).
*/
@Test
void leadSeatSubtractsFromFreeTheSameWayLiveDoes() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FleetMcp.LeadSeatSource leadSeats = new FleetMcp.LeadSeatSource(
profile -> "sonnet".equals(profile) ? 1 : 0);
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new FleetMcp.CapacitySource(profile -> 2, profile -> 3,
() -> Set.of("sonnet"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), leadSeats, Map.of(), "", null));
assertTrue(out.contains("\"maxLoad\":3"), "maxLoad itself must be left untouched: " + out);
assertTrue(out.contains("\"live\":2"), out);
assertTrue(out.contains("\"free\":0"), "2 live + 1 lead seat fills all 3: " + out);
assertTrue(out.contains("\"leadSeats\":1"), out);
}
/**
* fleetd #176: the OTHER measurement in the issue — a completely idle fleet still overstates
* {@code free} by the lead's own seat. {@code maxLoad:3, live:0} must report {@code free:2}, the
* real fan-out ceiling, not {@code 3}.
*/
@Test
void leadSeatLowersFreeOnAnOtherwiseIdleSubscriptionProfile() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FleetMcp.LeadSeatSource leadSeats = new FleetMcp.LeadSeatSource(
profile -> "sonnet".equals(profile) ? 1 : 0);
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 3,
() -> Set.of("sonnet"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), leadSeats, Map.of(), "", null));
assertTrue(out.contains("\"live\":0"), out);
assertTrue(out.contains("\"free\":2"), "an idle fleet's real ceiling is 3 minus the lead's own seat: " + out);
assertTrue(out.contains("\"leadSeats\":1"), out);
}
/** A profile with no lead seats reported must be byte-identical to before this ticket. */
@Test
void zeroLeadSeatsOmitsTheKeyAndLeavesFreeUnchanged() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 2,
() -> Set.of("terra"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", null));
assertTrue(out.contains("\"free\":2"), out);
assertFalse(out.contains("leadSeats"), "no lead shares this profile's credential: " + out);
}
@Test
void listReportsLeadsAndFlagsTheCallersOwnRow() {
FakeHerdr h = new FakeHerdr();
@@ -1,121 +0,0 @@
package dev.ltms.fleet.mcp;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashSet;
import java.util.Set;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #114 (CB-609): the guard that lets {@code docs/MCP-Contract.md} name a tool at all.
*
* <p>That page was written in July 2026, before any MCP code existed, and then did not follow the
* code. By August it named two tools that had never been built, omitted five that shipped, had the
* wrong name for nearly every parameter, and still described a caller-identity rule that was a
* privilege bug by then. Nothing failed, because nothing checked it — and {@code CLAUDE.md} sends
* every session in the fleet to that page.
*
* <p>The fix was to delete the tool catalogue rather than correct it: a hand-maintained second copy
* of the tool surface is the defect, not the particular errors it had accumulated. What survives is
* the flows, which are shapes rather than names. But the flows still have to say {@code fleet_send}
* somewhere to be readable, and that is exactly the sentence that rots. This test is what makes it
* safe to write.
*
* <p><b>It checks source text, not behaviour.</b> It reads the Markdown and reads {@link FleetMcp}'s
* source, and it only catches a name in the doc that the server does not register. It cannot catch a
* flow that describes the wrong order, or a parameter name in prose — those are not name-shaped. The
* doc's own header carries that caveat for its readers.
*/
class McpContractDocTest {
/** Tests run with the module directory as cwd, so the repo-root doc is one level up. */
private static final Path DOC = Path.of("../docs/MCP-Contract.md");
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
private static Set<String> matches(Path file, String regex) throws Exception {
Matcher m = Pattern.compile(regex).matcher(Files.readString(file));
Set<String> found = new LinkedHashSet<>();
while (m.find()) {
found.add(m.group(1));
}
return found;
}
/** Every {@code fleet_*} the doc mentions, in prose or in a diagram. */
private static Set<String> toolsNamedInTheDoc() throws Exception {
return matches(DOC, "(fleet_[a-z_]+)");
}
/** Every tool {@link FleetMcp} actually registers, read from its {@code tool("…")} calls. */
private static Set<String> toolsTheServerRegisters() throws Exception {
return matches(MCP_SOURCE, "tool\\(\"(fleet_[a-z_]+)\"");
}
@Test
@DisplayName("[SOURCE TEXT] every fleet_* tool named in MCP-Contract.md is one the server registers")
void theDocNamesNoToolThatDoesNotExist() throws Exception {
Set<String> registered = toolsTheServerRegisters();
Set<String> named = toolsNamedInTheDoc();
Set<String> unknown = new LinkedHashSet<>(named);
unknown.removeAll(registered);
assertTrue(unknown.isEmpty(),
"docs/MCP-Contract.md names " + unknown + ", which FleetMcp does not register. "
+ "Checked " + named.size() + " name(s) in the doc against " + registered.size()
+ " registered tool(s): " + registered + ". This is the fleetd #114 defect "
+ "recurring — the doc named fleet_read and fleet_cancel for weeks after the "
+ "code shipped without them. Either fix the name or drop it from the page; do "
+ "NOT weaken this test.");
}
/**
* The denominator guard. The check above passes trivially if the doc stops naming any tool at
* all — an empty set is a subset of everything. A checker that can silently check nothing is the
* fleetd #113 shape, so this pins that the doc really is still describing the flows, and that
* the registration scrape really did find the server's tools.
*/
@Test
@DisplayName("[SOURCE TEXT] the doc/server name check is not vacuous — both sides found names")
void theCheckActuallyHasSomethingToCheck() throws Exception {
Set<String> registered = toolsTheServerRegisters();
Set<String> named = toolsNamedInTheDoc();
assertTrue(registered.size() >= 10,
"scraped only " + registered.size() + " tool registrations from FleetMcp (" + registered
+ "); the server registers eleven, so the tool(\"…\") scrape has stopped matching "
+ "and the check above is now vacuous");
assertTrue(named.size() >= 4,
"docs/MCP-Contract.md names only " + named.size() + " fleet_* tool(s) (" + named + "). "
+ "The flows describe delegation, clarification, detached delivery and the "
+ "turn-done fallback, so it should name several. Too few means the page has been "
+ "gutted and this test is guarding nothing.");
}
/**
* fleetd #114's actual lesson. The catalogue was deleted on purpose; a well-meaning "let me just
* document the tools here" restores the exact second copy that drifted for a month.
*/
@Test
@DisplayName("[SOURCE TEXT] MCP-Contract.md still says it is not the tool reference")
void theDocStillDisclaimsBeingTheToolReference() throws Exception {
String doc = Files.readString(DOC);
assertTrue(doc.contains("**What this page is NOT: a tool reference.**"),
"docs/MCP-Contract.md must keep saying it is not the tool reference. That sentence is "
+ "the fix for fleetd #114: the page carried a hand-maintained tool catalogue that "
+ "drifted from the code for a month while CLAUDE.md pointed every session at it.");
assertEquals(0, countTables(doc.substring(0, doc.indexOf("## 1. Rendezvous flows"))),
"the header of docs/MCP-Contract.md must not grow a tool/parameter table — that is the "
+ "second copy fleetd #114 deleted");
}
private static int countTables(String markdown) {
return (int) markdown.lines().filter(l -> l.strip().startsWith("|")).count();
}
}
@@ -1409,16 +1409,14 @@ class ClaudeCodeLauncherTest {
}
/**
* fleetd #155 (was CB-633 follow-up (#192) "allowListPolicyOnNonZshKeepsTheWarnWording"): on a
* non-zsh login shell {@code ZDOTDIR} is ignored, so no scrub could ever run there. Before #155
* the launcher degraded to the weaker overlay and spawned anyway; that is exactly the "control
* silently does nothing" failure #155 exists to close, since {@code policy: allow-list} is the
* operator explicitly asking for a control a sourced file cannot undo. Now the spawn is REFUSED
* instead — this test asserts the refusal, naming the shell, on the real spawn path ({@link
* ClaudeCodeLauncher#spawn}), not merely on the launcher method that computes it.
* CB-633 follow-up (#192): the mirror of the zsh test above. On a non-zsh login shell {@code
* ZDOTDIR} is ignored, so no scrub ever runs — the report must keep the WARN wording (a name here
* really is inherited unblocked) rather than claiming a scrub protects it. This is the trap PR
* #174 fell into the other direction: keying the wording on the shell, not on {@code
* creds.isAllowList()}, is what keeps this branch correct.
*/
@Test
void allowListPolicyOnNonZshRefusesTheSpawn() {
void allowListPolicyOnNonZshKeepsTheWarnWording() {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
@@ -1432,12 +1430,27 @@ class ClaudeCodeLauncherTest {
0, System::currentTimeMillis, () -> {}, null, () -> creds,
() -> Set.of("AI_GATEWAY_TOKEN", "A_BRAND_NEW_SECRET_TOKEN"));
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, svc::spawn,
"policy=allow-list on a non-zsh shell must refuse the spawn, not silently degrade "
+ "to the weaker overlay");
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
svc.spawn();
} finally {
logger.detachAppender(appender);
}
assertTrue(e.getMessage().contains("allow-list") && e.getMessage().contains("/bin/bash"),
"the refusal must name the policy and the actual shell, got: " + e.getMessage());
assertTrue(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().contains("UNBLOCKED")
&& e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN")),
"no scrub runs on a non-zsh shell, so the WARN wording must be kept — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
assertFalse(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().toLowerCase(java.util.Locale.ROOT).contains("scrub")),
"nothing is scrubbed on this path, so the report must not claim otherwise — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
}
/**
@@ -2146,7 +2159,8 @@ class ClaudeCodeLauncherTest {
}
JsonNode root = new ObjectMapper().readTree(claudeJson.toFile());
JsonNode project = root.path("projects").path(worktree.toString());
seededBeforeStart.set(project.path("hasTrustDialogAccepted").asBoolean(false));
seededBeforeStart.set(project.path("hasTrustDialogAccepted").asBoolean(false)
&& project.path("hasCompletedProjectOnboarding").asBoolean(false));
} catch (IOException e) {
seededBeforeStart.set(false);
}
@@ -2164,7 +2178,7 @@ class ClaudeCodeLauncherTest {
}
@Test
void seedTrustDialogWritesOnlyTheTrustFlagAndNotTheOnboardingKey(
void seedTrustDialogWritesBothTrustFlagsForTheResolvedCwd(
@TempDir Path configDir, @TempDir Path worktree) throws Exception {
markAsProvisionedWorktree(worktree);
FakeHerdr herdr = new FakeHerdr();
@@ -2178,13 +2192,7 @@ class ClaudeCodeLauncherTest {
JsonNode project = new ObjectMapper().readTree(claudeJson.toFile())
.path("projects").path(worktree.toString());
assertTrue(project.path("hasTrustDialogAccepted").asBoolean(false));
// fleetd #247: the onboarding key must NOT be written. Claude Code strips it on every
// save (measured: 0 of 28 live entries had it, including its own), so writing it only
// adds a contested key to a file two processes share. This assertion is the guard that
// stops it coming back as a plausible-looking "completeness" fix.
assertFalse(project.has("hasCompletedProjectOnboarding"),
"hasCompletedProjectOnboarding must not be written — Claude Code drops it on "
+ "every save, and the member reaches idle on hasTrustDialogAccepted alone");
assertTrue(project.path("hasCompletedProjectOnboarding").asBoolean(false));
}
/**
@@ -2230,6 +2238,7 @@ class ClaudeCodeLauncherTest {
JsonNode mine = root.path("projects").path(worktree.toString());
assertTrue(mine.path("hasTrustDialogAccepted").asBoolean(false));
assertTrue(mine.path("hasCompletedProjectOnboarding").asBoolean(false));
}
/**
@@ -29,7 +29,6 @@ import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
@@ -100,26 +99,18 @@ class HerdrPeerLauncherAllowListWiringTest {
}
/**
* fleetd #155: a non-zsh shell cannot read {@code ZDOTDIR} at all, so under {@code
* policy: allow-list} — the policy the operator picked specifically for a control a sourced file
* cannot undo — the launcher must REFUSE the spawn rather than silently generate a directory
* nothing will ever read (protection theatre) or fall back to the weaker overlay (exactly the
* "control silently does nothing" defect this ticket exists to close). Real path: through {@link
* HerdrPeerLauncher#spawn}, the method the daemon actually calls.
* A non-zsh shell cannot read {@code ZDOTDIR} at all. The launcher must fall back rather than
* generate a directory nothing will ever read — a directory that would look like protection.
*/
@Test
void aNonZshShellUnderAllowListPolicyRefusesTheSpawn() {
void aNonZshShellGeneratesNothingAndFallsBack() {
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/bash");
IllegalArgumentException e = assertThrows(IllegalArgumentException.class,
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
"policy=allow-list on a non-zsh shell must refuse the spawn, not silently degrade");
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
assertTrue(e.getMessage().contains("allow-list") && e.getMessage().contains("/bin/bash"),
"the refusal must name the policy and the actual shell, got: " + e.getMessage());
assertFalse(launcher.env.containsKey("ZDOTDIR"),
"a refused spawn must not have generated (or wired in) a scrub directory: " + launcher.env);
"bash ignores ZDOTDIR; setting it would be protection theatre");
}
private static Supplier<FleetConfig.MemberCredentials> allowList() {
@@ -255,14 +246,11 @@ class HerdrPeerLauncherAllowListWiringTest {
* "allowed N of M" line — which describes what the scrub does — must not be printed there either.
* Before this fix the line was logged BEFORE the zsh gate, so a non-zsh host printed e.g.
* "allowed 1 of 3" while blocking nothing at all, telling an operator a control ran when it did
* not. fleetd #155: the spawn itself is now refused on this path (see {@code
* aNonZshShellUnderAllowListPolicyRefusesTheSpawn}) rather than falling back — this test's own
* concern still holds under the refusal: the "allowed N of M" line describes a scrub that never
* ran here, so it must still never appear. Real path: goes through {@link
* HerdrPeerLauncher#spawn}, same as the sibling test above, with the shell fixed to bash.
* not. Real path: goes through {@link HerdrPeerLauncher#spawn}, same as the sibling test above,
* with the shell fixed to bash so the fallback branch is the one exercised.
*/
@Test
void noAllowedCountLineIsEmittedOnTheNonZshRefusalPath() {
void noAllowedCountLineIsEmittedOnTheNonZshFallbackPath() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of(INJECTED, "SOME_UNRELATED_NAME", "ANOTHER_UNRELATED_NAME");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/bash", () -> hostEnvNames);
@@ -274,9 +262,7 @@ class HerdrPeerLauncherAllowListWiringTest {
appender.start();
logger.addAppender(appender);
try {
assertThrows(IllegalArgumentException.class,
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
"policy=allow-list on a non-zsh shell must refuse the spawn");
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
@@ -324,20 +310,13 @@ class HerdrPeerLauncherAllowListWiringTest {
* different OS user — {@link HerdrPeerLauncher#hostEnvNames} describes fleetd's own process, not
* that user's. Neither "inherits them UNBLOCKED" nor "scrub blanks them" is evidence-backed
* there, so neither may print; the single unknown-environment WARN must, naming the config key.
*
* <p>fleetd #155: {@code memberLoginShell} is now given explicitly as zsh, so the spawn reaches
* this WARN through the (still-degrading, not refusing) missing-{@code worktreeRoot}/{@code
* worktreeGroup} fallback rather than through the zsh gate, which now refuses instead of falling
* back — see {@code aNonZshShellUnderAllowListPolicyRefusesTheSpawn}. This test's own concern
* (the unknown-environment WARN) is orthogonal to which fallback reached {@code
* logCredentialGap}, so it still holds.
*/
@Test
void gapDetectorReportsUnknownInsteadOfAConclusionWhenMemberHerdrSocketIsConfigured() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/zsh"));
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock"));
List<String> messages = spawnAndCaptureLogs(launcher);
@@ -361,9 +340,8 @@ class HerdrPeerLauncherAllowListWiringTest {
void theUnknownEnvironmentWarnFiresOnceNotOncePerSpawn() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
// fleetd #155: memberLoginShell given explicitly as zsh — see the sibling test above for why.
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/zsh"));
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock"));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level original = logger.getLevel();
@@ -411,44 +389,42 @@ class HerdrPeerLauncherAllowListWiringTest {
}
/**
* fleetd #213 defect 1, acceptance criterion 1 — updated by fleetd #155: {@code
* memberHerdrSocket} configured and {@code memberLoginShell} configured as non-zsh must now
* REFUSE the spawn (not fall back to the sentinel overlay — see {@code
* aNonZshShellUnderAllowListPolicyRefusesTheSpawn}'s javadoc for why). The WiringLauncher's own
* {@code env("SHELL")} is deliberately set to {@code /bin/zsh} — the OPPOSITE of what {@code
* memberLoginShell} says — so a launcher that (incorrectly) fell back to fleetd's own {@code
* $SHELL} here would wrongly pass the gate and fail this test.
* fleetd #213 defect 1, acceptance criterion 1: {@code memberHerdrSocket} configured and {@code
* memberLoginShell} configured as non-zsh must fall back to the sentinel overlay exactly like a
* non-zsh {@code $SHELL} does today — and the "generated ZDOTDIR" INFO must not appear, since no
* scrub actually runs. The WiringLauncher's own {@code env("SHELL")} is deliberately set to
* {@code /bin/zsh} — the OPPOSITE of what {@code memberLoginShell} says — so a launcher that
* (incorrectly) fell back to fleetd's own {@code $SHELL} here would wrongly pass the gate and
* fail this test.
*/
@Test
void memberHerdrSocketWithNonZshMemberLoginShellRefusesTheSpawn() {
void memberHerdrSocketWithNonZshMemberLoginShellFallsBackToTheOverlay() {
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr, allowListWithKnown(List.of("SOME_TOKEN")),
"/bin/zsh", null,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/bash"));
IllegalArgumentException e = assertThrows(IllegalArgumentException.class,
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
"a non-zsh configured memberLoginShell under policy=allow-list must refuse the spawn");
List<String> messages = spawnAndCaptureLogs(launcher);
assertTrue(e.getMessage().contains("/bin/bash"),
"the refusal must name the configured memberLoginShell, got: " + e.getMessage());
assertFalse(launcher.env.containsKey("SOME_TOKEN"),
"a refused spawn must not have touched the pane env at all: " + launcher.env);
assertEquals("blocked-by-fleetd-cb596-see-gitea-issue-82", launcher.env.get("SOME_TOKEN"),
"a non-zsh memberLoginShell must fall back to the CB-596 sentinel overlay, exactly "
+ "like a non-zsh $SHELL does when memberHerdrSocket is absent");
assertFalse(launcher.env.containsKey("ZDOTDIR"),
"no scrub directory may be generated when the configured member login shell is not zsh");
assertFalse(messages.stream().anyMatch(m -> m.contains("generated ZDOTDIR")),
"the 'generated ZDOTDIR' INFO must not appear when the scrub never runs — got: " + messages);
}
/**
* fleetd #213 defect 1, acceptance criterion 2 — updated by fleetd #155: {@code
* memberHerdrSocket} configured and NO {@code memberLoginShell} configured must now REFUSE the
* spawn exactly like criterion 1 above — AND fleetd's own {@code $SHELL} must never even be
* consulted (not merely "not decisive"). The fixture's {@code env} function reports {@code
* /bin/zsh} for {@code SHELL} — a value that would WRONGLY pass the zsh gate if the fix
* regressed to reading it — while flagging whether it was ever asked for at all, so this test
* fails loudly on either kind of regression.
* fleetd #213 defect 1, acceptance criterion 2: {@code memberHerdrSocket} configured and NO
* {@code memberLoginShell} configured must fall back exactly like criterion 1 above — AND
* fleetd's own {@code $SHELL} must never even be consulted (not merely "not decisive"). The
* fixture's {@code env} function reports {@code /bin/zsh} for {@code SHELL} — a value that would
* WRONGLY pass the zsh gate if the fix regressed to reading it — while flagging whether it was
* ever asked for at all, so this test fails loudly on either kind of regression.
*/
@Test
void memberHerdrSocketWithNoMemberLoginShellRefusesAndNeverConsultsFleetdsOwnShell() {
void memberHerdrSocketWithNoMemberLoginShellFallsBackAndNeverConsultsFleetdsOwnShell() {
FakeHerdr herdr = new FakeHerdr();
AtomicBoolean shellQueried = new AtomicBoolean(false);
Function<String, String> env = name -> {
@@ -461,18 +437,15 @@ class HerdrPeerLauncherAllowListWiringTest {
WiringLauncher launcher = new WiringLauncher(herdr, allowListWithKnown(List.of("SOME_TOKEN")), env,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", null));
assertThrows(IllegalArgumentException.class,
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
"no configured memberLoginShell under policy=allow-list must refuse the spawn, same "
+ "as an explicit non-zsh one");
List<String> messages = spawnAndCaptureLogs(launcher);
assertFalse(shellQueried.get(), "fleetd's own $SHELL must never be consulted once "
+ "memberHerdrSocket is configured — only memberLoginShell: may decide the gate");
assertFalse(launcher.env.containsKey("SOME_TOKEN"),
"a refused spawn must not have touched the pane env at all, same as a "
assertEquals("blocked-by-fleetd-cb596-see-gitea-issue-82", launcher.env.get("SOME_TOKEN"),
"no memberLoginShell configured must fall back to the sentinel overlay, same as a "
+ "configured non-zsh shell");
assertFalse(launcher.env.containsKey("ZDOTDIR"),
"no scrub may run without a configured memberLoginShell");
assertFalse(messages.stream().anyMatch(m -> m.contains("generated ZDOTDIR")),
"no scrub may run without a configured memberLoginShell — got: " + messages);
}
/**
@@ -727,8 +700,7 @@ class HerdrPeerLauncherAllowListWiringTest {
/**
* fleetd #213: as above, plus {@code worktreeRoot:}/{@code worktreeGroup:} — both required for
* the ZDOTDIR scrub to run at all once {@code memberHerdrSocket} is configured; either missing
* falls back to the sentinel overlay (fleetd #155 left this branch alone — it is not a shell
* problem, so it is not this ticket's refusal).
* falls back to the sentinel overlay, same as a non-zsh {@code memberLoginShell}.
*/
private static FleetConfig configWithMemberHerdrSocketRootAndGroup(String memberHerdrSocket,
String memberLoginShell, String worktreeRoot, String worktreeGroup) {
@@ -323,39 +323,11 @@ class OpenCodeLauncherTest {
// --- CB-547: resume + post-hoc session discovery --------------------------------------------
/**
* Give {@code dir} the exact signature {@link HerdrPeerLauncher#isProvisionedWorktree} checks
* for: a {@code .git} REGULAR FILE, never a directory. Content is never parsed by that gate, so
* any {@code gitdir:} pointer is fine. Mirrors {@code ClaudeCodeLauncherTest}'s helper of the
* same shape (fleetd #249).
*/
private static void markAsProvisionedWorktree(Path dir) throws IOException {
Files.writeString(dir.resolve(".git"), "gitdir: /tmp/not-a-real-gitdir");
}
/**
* A fresh subdirectory of {@code configRoot}, marked as a provisioned worktree (fleetd #249),
* for tests that predate this gate and stood in a bare {@code "/work/dir"} string as their
* member's cwd — a directory that never existed on disk and, post-#249, would never pass
* {@link HerdrPeerLauncher#isProvisionedWorktree} either. Those tests are about the model
* mismatch / late-resolve machinery (fleetd #175/#234/#209), not about the worktree gate
* itself, so they need a cwd the gate accepts without changing what each test demonstrates.
*/
private static String provisionedWorkDir(Path configRoot) throws IOException {
Path dir = Files.createDirectories(configRoot.resolve("work-dir"));
markAsProvisionedWorktree(dir);
return dir.toString();
}
@Test
void aResumeSpawnIntoAProvisionedWorktreePassesTheSessionIdAsDashS(@TempDir Path root,
@TempDir Path worktree)
throws Exception {
markAsProvisionedWorktree(worktree);
void aResumeSpawnPassesTheSessionIdAsDashS(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null))
.spawn(new SpawnRequest(null, worktree.toString(), null, null,
"ses_41b79fc90ffeI9E8uZv6VprUn2"));
.spawn(new SpawnRequest(null, null, null, null, "ses_41b79fc90ffeI9E8uZv6VprUn2"));
List<String> args = startArgs(herdr);
int s = args.indexOf("-s");
@@ -364,27 +336,6 @@ class OpenCodeLauncherTest {
"the resume target id follows -s");
}
/**
* fleetd #249 acceptance criterion 3: without a fleetd-provisioned worktree, the member's cwd
* is shared with other sessions, so fleetd can never reliably confirm (now or later via {@link
* OpenCodeSessionDiscovery}) which conversation it is actually running. Refuse the spawn itself
* rather than silently launching opencode's {@code -s <id>} into unverifiable territory.
*/
@Test
void aResumeSpawnWithoutAProvisionedWorktreeIsRefused(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = service(herdr, root,
opencodeCfg("google/gemini-2.5-pro", null, null));
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, () ->
launcher.spawn(new SpawnRequest(null, null, null, null,
"ses_41b79fc90ffeI9E8uZv6VprUn2")));
assertTrue(e.getMessage().contains("worktree"), e.getMessage());
assertFalse(herdr.called("agent.start"),
"the refusal must happen before anything spawns — no pane, no process");
}
@Test
void aFreshSpawnCarriesNoSessionFlag(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
@@ -397,59 +348,25 @@ class OpenCodeLauncherTest {
@Test
void theHandleDiscoversTheSessionIdForTheWorkersCwdOnlyAfterItAppears(@TempDir Path root,
@TempDir Path discRoot,
@TempDir Path worktree)
@TempDir Path discRoot)
throws Exception {
markAsProvisionedWorktree(worktree);
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = new OpenCodeLauncher(new AgentControl(herdr),
new WorkspaceControl(herdr), Map.of("gemini", opencodeCfg(null, null, null)),
"gemini", _ -> null, 0, System::currentTimeMillis, () -> { }, root, discRoot);
PeerHandle handle = launcher.spawn(new SpawnRequest(null, worktree.toString(), null));
PeerHandle handle = launcher.spawn(new SpawnRequest(null, "/work/dir", null));
// opencode writes the record only when the session is first persisted — the instant the
// pane is ready it does not exist, so agentSessionId() is null (never a spawn failure).
assertNull(handle.agentSessionId(), "no record yet → null, not a spawn-time block");
// Once the record appears (here: same cwd), lazy discovery resolves it — the handle's
// session id matches its own worktree, not another's.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_resolved", worktree.toString(), 1000L);
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_resolved", "/work/dir", 1000L);
assertEquals("ses_resolved", handle.agentSessionId(),
"agentSessionId() re-scans and picks up a record that has since been written");
}
/**
* fleetd #249 acceptance criterion 1, exercised through the real caller path (the handle
* {@code fleet_list} actually reads), not {@link OpenCodeSessionDiscovery} directly. Without a
* fleetd-provisioned worktree the member's cwd is shared — the default no-worktree spawn
* inherits the lead's own long-lived cwd — so even once a matching row appears (here:
* simulating another profile's session that happens to share the directory) the handle must
* report absence rather than guess. Measured real-world case (2026-09-03): the row it would
* otherwise pick was three days old and belonged to a different profile.
*/
@Test
void theHandleNeverReportsAnIdForANonProvisionedCwdEvenAfterARowAppears(@TempDir Path root,
@TempDir Path discRoot)
throws Exception {
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = new OpenCodeLauncher(new AgentControl(herdr),
new WorkspaceControl(herdr), Map.of("gemini", opencodeCfg(null, null, null)),
"gemini", _ -> null, 0, System::currentTimeMillis, () -> { }, root, discRoot);
// No markAsProvisionedWorktree — this cwd has no .git file, the shared-cwd shape a
// no-worktree spawn (or a real checkout) actually has.
String sharedCwd = root.resolve("shared-cwd").toString();
PeerHandle handle = launcher.spawn(new SpawnRequest(null, sharedCwd, null));
assertNull(handle.agentSessionId(), "no record yet → null, same as the provisioned case");
// A row for this exact directory now appears — e.g. a sibling member, or a stale session
// from days earlier, sharing the same unprovisioned cwd.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_someone_elses", sharedCwd, 1000L);
assertNull(handle.agentSessionId(),
"a non-provisioned cwd must NEVER report an id, even once a row for it exists — "
+ "the row could belong to any other session sharing this directory");
}
@Test
void foreignWorkerMatchesOpencodePrefixButNotClaude() {
String nonce = "abc123";
@@ -1003,7 +920,6 @@ class OpenCodeLauncherTest {
@Test
void theRealSessionManagerLateResolvePathCatchesAModelMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
String workDir = provisionedWorkDir(configRoot);
FakeHerdr herdr = new FakeHerdr();
// xf's real shape (fleetd #175): weight:80, model "opencode/nemotron-3-ultra-free", no
// credentialId — the profile that actually escaped the fleet's accounting.
@@ -1013,7 +929,7 @@ class OpenCodeLauncherTest {
OpenCodeLauncher launcher = serviceWithSink(herdr, configRoot, discRoot, cfg, sink);
SessionManager sessions = new SessionManager(launcher);
MemberSession acquired = sessions.acquire(cfg.profile(), workDir, null, null);
MemberSession acquired = sessions.acquire(cfg.profile(), "/work/dir", null, null);
// Real late-resolve path, driven BEFORE opencode has written its session row — same shape
// as production the instant a pane goes ready.
@@ -1024,7 +940,7 @@ class OpenCodeLauncherTest {
// opencode writes its row late, running gpt-5.6-sol (a PAID credential) instead of the
// withdrawn free model the profile actually asked for — the exact fleetd #175 scenario.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
// Drive the SAME real late-resolve path again: sessions.get() -> resolveAgentSessionId ->
@@ -1043,13 +959,12 @@ class OpenCodeLauncherTest {
@Test
void aProviderPrefixedModelMatchingBothIdAndProviderIsNotAMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
String workDir = provisionedWorkDir(configRoot);
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, workDir, null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-terra\",\"providerID\":\"openai\"}");
assertEquals("ses_x", handle.agentSessionId());
@@ -1060,13 +975,12 @@ class OpenCodeLauncherTest {
@Test
void aGxProviderPrefixedModelMatchingBothIdAndProviderIsNotAMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
String workDir = provisionedWorkDir(configRoot);
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("gx/deepseek-v4-flash", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, workDir, null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"deepseek-v4-flash\",\"providerID\":\"gx\"}");
assertEquals("ses_x", handle.agentSessionId());
@@ -1083,13 +997,12 @@ class OpenCodeLauncherTest {
@Test
void aMissingProviderIdInTheEvidenceIsUnknownNotAMismatchWhenTheIdMatches(
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
String workDir = provisionedWorkDir(configRoot);
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, workDir, null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-terra\"}");
assertEquals("ses_x", handle.agentSessionId());
@@ -1105,13 +1018,12 @@ class OpenCodeLauncherTest {
@Test
void aMissingProviderIdInTheEvidenceStillCatchesARealIdMismatch(
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
String workDir = provisionedWorkDir(configRoot);
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, workDir, null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\"}");
assertEquals("ses_x", handle.agentSessionId());
@@ -1130,13 +1042,12 @@ class OpenCodeLauncherTest {
@Test
void aBareModelWithNoProviderPrefixMatchesOnIdAloneAndIsNotAMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
String workDir = provisionedWorkDir(configRoot);
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("deepseek-v4-flash", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, workDir, null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"deepseek-v4-flash\",\"providerID\":\"gx\"}");
assertEquals("ses_x", handle.agentSessionId());
@@ -1153,7 +1064,6 @@ class OpenCodeLauncherTest {
@Test
void aRealIdMismatchLogsAnErrorNamingBothModelsAndQuarantinesThroughTheSink(
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
String workDir = provisionedWorkDir(configRoot);
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(target + "|" + reason);
FleetConfig.Profile cfg = opencodeCfg("opencode/nemotron-3-ultra-free", null, null);
@@ -1165,8 +1075,8 @@ class OpenCodeLauncherTest {
PeerHandle handle;
try {
handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, workDir, null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
assertEquals("ses_x", handle.agentSessionId());
} finally {
@@ -1201,12 +1111,11 @@ class OpenCodeLauncherTest {
@Test
void unknownOrUnparseableModelEvidenceNeverQuarantines(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
String workDir = provisionedWorkDir(configRoot);
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, workDir, null));
.spawn(new SpawnRequest(null, "/work/dir", null));
// No row yet at all.
assertNull(handle.agentSessionId());
@@ -1228,13 +1137,12 @@ class OpenCodeLauncherTest {
@Test
void aProfileWithNoConfiguredModelIsNeverCheckedForAMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
String workDir = provisionedWorkDir(configRoot);
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg(null, null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, workDir, null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"anything-at-all\",\"providerID\":\"anyone\"}");
assertEquals("ses_x", handle.agentSessionId());
@@ -1258,22 +1166,21 @@ class OpenCodeLauncherTest {
@Test
void modelCheckReadsTheResolvedSessionsOwnRowNotWhateverIsNewestInTheSharedDirectory(
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
String workDir = provisionedWorkDir(configRoot);
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, workDir, null));
.spawn(new SpawnRequest(null, "/work/dir", null));
// Our own session's row, correctly matching the profile's requested model.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_ours", workDir, 1000L,
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_ours", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-terra\",\"providerID\":\"openai\"}");
assertEquals("ses_ours", handle.agentSessionId(), "resolves to our own session");
assertTrue(exhausted.isEmpty(), "matching model → no mismatch on first resolve: " + exhausted);
// A sibling member, spawned later into the SAME shared directory (no worktree, fleetd
// #234's default), writes a newer row running a totally different model.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_sibling", workDir, 9000L,
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_sibling", "/work/dir", 9000L,
"{\"id\":\"deepseek-v4-flash\",\"providerID\":\"gx\"}");
assertEquals("ses_ours", handle.agentSessionId(),
@@ -1305,7 +1212,6 @@ class OpenCodeLauncherTest {
@Test
void aSpawnTimeModelMismatchActuallyQuarantinesTheCredentialThroughTheRealAcquirePath(
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
String workDir = provisionedWorkDir(configRoot);
FleetConfig.Profile cfg = opencodeCfgWithCredential(
"terra", "opencode/nemotron-3-ultra-free", "openai-shared");
Map<String, FleetConfig.Profile> profiles = Map.of(cfg.profile(), cfg);
@@ -1326,14 +1232,14 @@ class OpenCodeLauncherTest {
// The mismatching row exists BEFORE the spawn — reproducing fleetd #234's exact timing:
// opencode's session table already carries evidence by the moment acquire() first asks.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
assertFalse(quarantine.isQuarantined("openai-shared"), "nothing quarantined before the spawn");
// The real production entrypoint: acquire() builds the MemberSession by calling
// handle.agentSessionId() BEFORE registry.put() runs.
MemberSession acquired = sessions.acquire(cfg.profile(), workDir, null, null);
MemberSession acquired = sessions.acquire(cfg.profile(), "/work/dir", null, null);
assertEquals("ses_x", acquired.agentSessionId(), "the id itself still resolves correctly");
assertTrue(quarantine.isQuarantined("openai-shared"),
@@ -1352,7 +1258,6 @@ class OpenCodeLauncherTest {
@Test
void aRosterOnlySinkSilentlyDropsTheSpawnTimeQuarantine(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
String workDir = provisionedWorkDir(configRoot);
FleetConfig.Profile cfg = opencodeCfgWithCredential(
"terra", "opencode/nemotron-3-ultra-free", "openai-shared");
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.SECONDS.toNanos(1800));
@@ -1365,10 +1270,10 @@ class OpenCodeLauncherTest {
OpenCodeLauncher launcher = serviceWithSink(herdr, configRoot, discRoot, cfg, rosterOnlySink);
SessionManager sessions = new SessionManager(launcher);
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
MemberSession acquired = sessions.acquire(cfg.profile(), workDir, null, null);
MemberSession acquired = sessions.acquire(cfg.profile(), "/work/dir", null, null);
assertEquals("ses_x", acquired.agentSessionId(), "the id itself still resolves correctly");
assertFalse(quarantine.isQuarantined("openai-shared"),
@@ -1396,7 +1301,6 @@ class OpenCodeLauncherTest {
@Test
void theSpawnTimeQuarantineSurvivesTheFleetdStyleForwardingHop(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
String workDir = provisionedWorkDir(configRoot);
FleetConfig.Profile cfg = opencodeCfgWithCredential(
"terra", "opencode/nemotron-3-ultra-free", "openai-shared");
Map<String, FleetConfig.Profile> profiles = Map.of(cfg.profile(), cfg);
@@ -1428,12 +1332,12 @@ class OpenCodeLauncherTest {
};
exhaustionSinkRef.set(realSink);
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
assertFalse(quarantine.isQuarantined("openai-shared"), "nothing quarantined before the spawn");
MemberSession acquired = sessions.acquire(cfg.profile(), workDir, null, null);
MemberSession acquired = sessions.acquire(cfg.profile(), "/work/dir", null, null);
assertEquals("ses_x", acquired.agentSessionId(), "the id itself still resolves correctly");
assertTrue(quarantine.isQuarantined("openai-shared"),