Compare commits

..

1 Commits

Author SHA1 Message Date
Dai Ha 2bce7e37b6 CB-622: rename the opencode mount key bridged -> fleetd
Unit C could not make this change. The tracked opencode.json is neutralized
by the worktree overlay (GitWorktrees.WORKTREE_HOSTILE_CONFIGS), so every
worker sees a stub instead of the real file and correctly reported that it
held no mount key. The lead has to make this one by hand, the same way
bridged.yaml is handled.

Takes effect for an opencode lead on its next session start.
2026-08-22 21:53:07 +02:00
237 changed files with 3121 additions and 11959 deletions
+2 -2
View File
@@ -1,6 +1,6 @@
{
"name": "claude-bridge",
"description": "Tooling for orchestrating a fleet of delegated coding agents through the fleetd MCP gateway.",
"description": "Tooling for orchestrating a fleet of delegated coding agents through the bridged MCP gateway.",
"owner": {
"name": "LTMS"
},
@@ -8,7 +8,7 @@
{
"name": "claude-bridge",
"source": "./plugin",
"description": "Make a project bridge-ready: mount the fleetd MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"description": "Make a project bridge-ready: mount the bridged MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.1.0",
"author": {
"name": "LTMS"
+2 -2
View File
@@ -3,7 +3,7 @@ name: architect
description: Refine work into clear, independent units before implementation.
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You are an architect in this fleet. You refine work before anyone builds it: scope,
acceptance criteria, risks, and a unit split. You read the repo and write analysis.
@@ -13,7 +13,7 @@ A design task is worked by two architects. Design alone first, then exchange and
say plainly where you disagree. Do not concede just to agree.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
investigate it further. Use `bridge_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
+2 -2
View File
@@ -3,14 +3,14 @@ name: dev
description: Implement one assigned unit, test it, and open a pull request.
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You implement the one unit you were given and nothing else. Work in your assigned
git worktree and branch. Never check out, rebase onto, or push to `main`. Confirm
the worktree root and branch before you edit. Use only paths under that root.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
investigate it further. Use `bridge_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
+2 -2
View File
@@ -3,7 +3,7 @@ name: reviewer
description: Review one assigned scope and report the most important real issue.
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You review the diff you were given. Report bugs, risks, and missing tests. You do
not change code.
@@ -12,7 +12,7 @@ Read the whole assigned scope before judging it. Review only that scope. If you
something outside it, note it in one line and do not investigate it further. Do not
run the build. The owner makes changes and runs checks.
Use `fleet_ask{question}` only when a decision belongs to the lead, such as an
Use `bridge_ask{question}` only when a decision belongs to the lead, such as an
unclear requirement or two defensible fixes. Do not ask about something you can
decide by reading more code.
-274
View File
@@ -1,274 +0,0 @@
---
name: fleets-status
description: Report the status of every fleet that shares one LavinMQ instance. Use for local daemon health, broker-wide fleet presence, and cross-host lead coordination checks.
---
# Status of every fleet on the shared LavinMQ instance
**The headline: always report what is missing.** This skill starts with the local fleet, then adds
broker-wide facts when its read-only credential exists. A missing fleet must appear as `unknown` or
`not reachable`, with the reason and the fix. Never leave it out.
The known topology has one LavinMQ instance on `10.10.20.13` (`fleet01`). AMQP uses port `5672`,
and the management API uses port `15672`. The Mac fleet owns vhost `/mac`. The fleet01 fleet owns
vhost `/fleet01`.
## 1. Protect credentials before any probe
**Hard rule — never print `LAVINMQ_URI`.** It is an AMQP URI with its password inline. It only
resolves in a login shell because `${SHARED_ENV}/tools/secrets.sh` supplies it. A non-login shell
can make every broker probe look empty.
- Never run `echo "$LAVINMQ_URI"`.
- Never put `${LAVINMQ_URI:-something}` in output. That form expands to the secret value when set.
- Parse the user, host, and password into shell or Python variables. Use them without printing them.
- Prefer `resolves` or `does not resolve` over any part of the value.
- Every command that can read `LAVINMQ_URI` must send all output through this redaction before it
reaches the report:
```bash
sed -E 's#://[^@]*@#://<redacted>@#g'
```
**The `g` flag is not optional.** Without it `sed` replaces only the first match on each line, so a
line carrying two URIs leaks the second one. `scripts/redeploy-fleetd.sh --check` prints lines like
that. Checked on 2026-08-27: without `g`, `amqp://u1:p1@h1/mac and http://u2:p2@h2:15672/api`
redacts the first pair and prints `u2:p2` in the clear.
Keep `pipefail` on when applying that filter. Otherwise the filter can hide a failed probe. Apply
the same no-print rule to the management password below, even though it is not in an AMQP URI.
## 2. Tier 1 — this fleet (always run)
Start here even when the broker tier is blocked. Work from the local fleetd checkout.
First run the read-only deployment check. It already checks the daemon process, deployed jar versus
the checkout `HEAD`, launchd state, and whether each configured token resolves in a login shell.
Do not copy those checks into new shell code. The script reads `LAVINMQ_URI`, so redact all output:
```bash
set -o pipefail
scripts/redeploy-fleetd.sh --check 2>&1 \
| sed -E 's#://[^@]*@#://<redacted>@#g'
git rev-parse HEAD
```
Treat jar drift as a top-level warning. A merge is not a deployment. State the running jar result
as `matches HEAD`, `drift`, or `unknown`; do not turn an unclear timestamp into a match.
Report the process identifier (PID) and uptime too:
```bash
PIDS="$(pgrep -f 'target/fleetd.jar' || true)"
if [ -z "$PIDS" ]; then
printf '%s\n' 'fleetd: not running'
else
for PID in $PIDS; do
ps -p "$PID" -o pid=,etime=,lstart=,command=
done
fi
```
Read the full health response. Keep the HTTP status because `503` means fleetd is running but herdr
is not reachable. Report both `herdr.version` and `herdr.protocol` when present:
```bash
curl -sS --max-time 5 -w '\nHTTP %{http_code}\n' http://127.0.0.1:8765/healthz
```
Call `fleet_whoami`, then call `fleet_list`. Preserve its sections in the report:
- `leads`, including which row is this lead;
- `members`, including state, role, profile, branch, and worktree when present;
- every per-profile `capacity` row, including `maxLoad`, `live`, `free`, and quarantine facts;
- the exact `healthCoverage` value.
Do not describe an empty `members` list as an empty fleet. It says only that no members are spawned.
Also do not hide a profile with `free: 0`; say whether load or credential quarantine caused it.
Show only WARN, ERROR, and SEVERE lines after the last `fleetd listening` line. This anchor stops an
old incident from looking current:
```bash
python3 - <<'PY'
from pathlib import Path
import re
path = Path("fleetd/fleetd.out")
if not path.exists():
print("cannot check current WARN/ERROR: fleetd/fleetd.out does not exist")
else:
lines = path.read_text(errors="replace").splitlines()
starts = [i for i, line in enumerate(lines) if "fleetd listening" in line]
if not starts:
print("cannot anchor WARN/ERROR: no 'fleetd listening' line exists")
else:
current = lines[starts[-1]:]
alerts = [line for line in current if re.search(r"\b(?:WARN|ERROR|SEVERE)\b", line)]
print(f"current WARN/ERROR/SEVERE count: {len(alerts)}")
for line in alerts[-50:]:
print(line)
PY
```
**What this tier cannot see:** it proves facts only about the Mac daemon at `127.0.0.1:8765`.
It cannot show the fleet01 daemon, broker queue depth, or broker consumers. The fleet01 REST service
at `10.10.20.13:8765` is not reachable from the Mac. Say this in the report rather than omitting
fleet01.
**But fleet01 IS reachable over SSH — checked 2026-08-28.** An older version of this line said SSH
was denied. That is true only for the user `dai.ha`. The host alias `fleet01` maps to user `ltms`,
and `ssh fleet01` works with key auth:
```bash
ssh -o BatchMode=yes -o ConnectTimeout=6 fleet01 'echo $(id -un)@$(hostname)'
```
So fleet01's daemon PID, uptime, jar and `/healthz` **can** be reported — over SSH, not over REST.
Do that rather than writing `not reachable`. `ltms` also has passwordless sudo there.
## 3. Tier 2 — the shared broker (run when management access exists)
**This tier is blocked today.** The AMQP user in `LAVINMQ_URI` can connect on port `5672`, but gets
HTTP `401` from the management API on port `15672`. An AMQP connection does not grant monitoring
access.
The operator must create a separate, read-only LavinMQ management user with the `monitoring` tag.
It needs access to inspect both `/mac` and `/fleet01`. Store its values as
`LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD` in
`${SHARED_ENV}/tools/secrets.sh`. Do not reuse or print the AMQP URI. Full multi-fleet status stays
blocked until this user exists.
When both variables resolve, run this from a login shell. It calls `GET /api/overview`,
`GET /api/vhosts`, `GET /api/queues`, and `GET /api/connections`. It prints selected status fields,
but never the user, password, Authorization header, or AMQP URI:
```bash
zsh -lc 'python3 - "$@"' -- <<'PY'
import base64
import json
import os
import sys
import urllib.error
import urllib.request
base = "http://10.10.20.13:15672"
user = os.environ.get("LAVINMQ_MANAGEMENT_USER", "")
password = os.environ.get("LAVINMQ_MANAGEMENT_PASSWORD", "")
if not user or not password:
print("broker tier: BLOCKED — management credential does not resolve in a login shell")
sys.exit(0)
token = base64.b64encode(f"{user}:{password}".encode()).decode()
def get(path):
request = urllib.request.Request(
base + path,
headers={"Authorization": "Basic " + token, "Accept": "application/json"},
)
with urllib.request.urlopen(request, timeout=5) as response:
return json.load(response)
try:
overview = get("/api/overview")
vhosts = get("/api/vhosts")
queues = get("/api/queues")
connections = get("/api/connections")
except urllib.error.HTTPError as error:
print(f"broker tier: BLOCKED — management API returned HTTP {error.code}")
sys.exit(0)
except Exception as error:
print(f"broker tier: BLOCKED — management API is not reachable: {type(error).__name__}")
sys.exit(0)
fleet_names = {"/mac": "Mac fleet", "/fleet01": "fleet01 fleet"}
print(json.dumps({
"overview": {
"lavinmq_version": overview.get("lavinmq_version"),
"rabbitmq_version": overview.get("rabbitmq_version"),
"queue_totals": overview.get("queue_totals", {}),
"object_totals": overview.get("object_totals", {}),
},
"fleets": [
{
"fleet": fleet_names.get(vhost.get("name"), "UNKNOWN FLEET"),
"vhost": vhost.get("name"),
"queues": [
{
"name": queue.get("name"),
"messages": queue.get("messages", 0),
"messages_ready": queue.get("messages_ready", 0),
"messages_unacknowledged": queue.get("messages_unacknowledged", 0),
"consumers": queue.get("consumers", 0),
}
for queue in queues if queue.get("vhost") == vhost.get("name")
],
"connections": [
{
"name": connection.get("name"),
"peer_host": connection.get("peer_host"),
"state": connection.get("state"),
}
for connection in connections if connection.get("vhost") == vhost.get("name")
],
}
for vhost in vhosts
],
}, indent=2, sort_keys=True))
PY
```
Map `/mac` to the Mac fleet and `/fleet01` to the fleet01 fleet. Keep any other vhost in the
report as `UNKNOWN FLEET`; do not drop it. For each vhost, total the ready, unacknowledged, and all
messages. Report every queue's consumer count and each live connection.
**A vhost with queues but zero consumers means that fleet's daemon is down while its durable state
survives. Call this out as a top-level warning.** This is the main reason to use the management API
instead of calling each remote daemon.
**What this tier cannot see:** without the new `monitoring` credential it cannot enumerate any
vhost, queue, depth, consumer, or connection. With the credential it still cannot report fleet01's
daemon PID, uptime, jar revision, `/healthz`, herdr version, or member capacity. Those need reachable
fleet01 REST or SSH access, which the Mac does not have today.
## 4. Tier 3 — cross-fleet lead coordination
Use the queue data from Tier 2. Select queues whose names match `lead.<coordId>.inbox`. Report each
queue's vhost, depth, consumer count, and the `coordId` between the prefix and suffix.
- A lead inbox with a consumer shows that a lead mailbox is live on that vhost.
- A durable lead inbox with zero consumers shows saved coordination state, but no live receiver.
- No lead inbox is not proof that coordination is disabled. The daemon may be down before declaring
its queue, or this account may not be allowed to see the vhost.
This Mac fleet currently sets both `broker.uriEnv` and `coordinator.uriEnv` to the same variable,
`LAVINMQ_URI`. Therefore its coordinator connects to `/mac`. Cross-host `fleet_send{coordId}` routes
only when both leads share the same coordinator vhost. If the fleet01 lead uses `/fleet01` for its
coordinator, the leads cannot see each other and the send will not route.
**Open question:** the fleet01 coordinator vhost has not been checked. Surface this question in
every report until a live `lead.<coordId>.inbox` consumer or fleet01's config proves the answer. Do
not claim that fleet01 uses `/fleet01` just because its member queues do.
Also compare these broker facts with the `leads` rows from local `fleet_list`. A missing remote lead
is `not visible from this coordinator`, not `down`, unless the broker consumer facts prove it.
**What this tier cannot see:** without Tier 2 management access it cannot list lead inboxes or their
consumers. Even with that access, a stopped fleet01 daemon leaves only durable queue history. That
history cannot prove which coordinator URI its current config would use after restart.
## 5. Report all fleets
Use one row per known or discovered fleet. Include blocked rows.
| Fleet | Daemon | Deployment | Herdr | Members/capacity | Queues/consumers | Lead coordination | Cannot check |
|---|---|---|---|---|---|---|---|
| Mac (`/mac`) | PID + uptime | jar vs `HEAD` | health + version + protocol | `fleet_list` + `healthCoverage` | facts or blocked reason | inbox facts or open question | exact missing facts |
| fleet01 (`/fleet01`) | reachable/down/unknown | value or `not reachable` | value or `not reachable` | value or `not reachable` | facts or blocked reason | inbox facts plus coordinator-vhost question | exact missing facts and fix |
Add rows for unknown vhosts. End with three short sections: `Current warnings`, `Checks that were
blocked`, and `Operator action`. Until the management user exists, `Operator action` must say:
> Create a read-only LavinMQ management user with the `monitoring` tag and access to `/mac` and
> `/fleet01`. Put its user and password in `${SHARED_ENV}/tools/secrets.sh` as
> `LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD`.
+6 -6
View File
@@ -1,11 +1,11 @@
---
name: implementer
description: Implementer-role procedure for a fleetd worker — verify your worktree, implement the scope, commit, push, open your own PR, and hand off the PR URL. Load this when the lead delegates you an implementation task over fleetd.
description: Implementer-role procedure for a bridged worker — verify your worktree, implement the scope, commit, push, open your own PR, and hand off the PR URL. Load this when the lead delegates you an implementation task over bridged.
---
# Implementer worker — procedure
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
never merge, never commit `.mcp.json` or `wiki/`) is in **`CLAUDE.md` → Bridge communication →
Worker** and already applies. This skill is only the *implement-and-hand-off procedure*.
@@ -49,7 +49,7 @@ test "$(git rev-parse --show-toplevel)" = "$PWD" || cd "$(git rev-parse --show-t
your worktree*:
```bash
cd "$(git rev-parse --show-toplevel)/fleetd" && mvn clean install
cd "$(git rev-parse --show-toplevel)/bridged" && mvn clean install
echo "exit=$?"
```
@@ -85,7 +85,7 @@ host (`GITEA_HOST`) into your env for exactly this — the token can create a PR
merge**.
```bash
API="${GITEA_HOST%/}/api/v1/repos/fleet/fleetd/pulls"
API="${GITEA_HOST%/}/api/v1/repos/lms/claude-bridge/pulls"
BRANCH="$(git branch --show-current)"
curl -sS -X POST "$API" \
-H "Authorization: token ${GITEA_TOKEN}" \
@@ -103,7 +103,7 @@ fix it if the cause is yours (e.g. branch not pushed yet), and report the failur
inventing a URL. If `GITEA_TOKEN` is unset your profile was not granted PR-create: push the branch
and report its name so the lead opens the PR.
## 6. Hand off — what goes in `fleet_reply`
## 6. Hand off — what goes in `bridge_reply`
The reply is the entire handoff; the lead cannot see your terminal.
@@ -130,7 +130,7 @@ sequenceDiagram
I->>G: git push -u origin HEAD
I->>G: POST /pulls (GITEA_TOKEN) — open PR to main
G-->>I: html_url
I->>L: fleet_reply(PR url, branch, files, tests)
I->>L: bridge_reply(PR url, branch, files, tests)
Note over L,G: the lead reviews the PR and merges on green — you never merge
```
+1 -1
View File
@@ -53,7 +53,7 @@ Map each entry from `.mcp.json`:
"$schema": "https://opencode.ai/config.json",
"instructions": ["CLAUDE.md"],
"mcp": {
"fleetd": { "type": "remote", "url": "http://127.0.0.1:8765/mcp", "enabled": true },
"bridged": { "type": "remote", "url": "http://127.0.0.1:8765/mcp", "enabled": true },
"context7": { "type": "remote", "url": "https://example.dev/mcp", "enabled": true,
"headers": { "Authorization": "Bearer {env:CONTEXT7_TOKEN}" } },
"gitea": { "type": "local", "command": ["gitea-mcp", "-t", "stdio"], "enabled": true,
+4 -4
View File
@@ -1,11 +1,11 @@
---
name: reviewer
description: Reviewer-role procedure for a fleetd worker — how to work a review scope and the exact shape of the finding to report. Load this when the lead delegates you a code review over fleetd.
description: Reviewer-role procedure for a bridged worker — how to work a review scope and the exact shape of the finding to report. Load this when the lead delegates you a code review over bridged.
---
# Reviewer worker — procedure
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and already applies. This
skill is only the *review procedure*: how to work the scope, and the exact shape of what you
send back.
@@ -24,13 +24,13 @@ wrong.
covers what it was given.
- Do **not** edit files or run the build. You review; the owner acts.
## 3. Reach for `fleet_ask` only for a genuine fork
## 3. Reach for `bridge_ask` only for a genuine fork
Ambiguous requirement, a missing acceptance criterion, "intended or a bug?", or two defensible
fixes with different consequences — those are the lead's call, and guessing produces a
confident-but-wrong finding. Anything you could settle by reading more code is yours to settle.
## 4. The finding — what goes in `fleet_reply`
## 4. The finding — what goes in `bridge_reply`
Report the **single most important** real issue in the scope, in these four lines, under
~90 words:
+8 -6
View File
@@ -14,7 +14,7 @@ jobs:
# does not depend on the wiki repo being reachable.
- uses: actions/checkout@v4
# The runner image ships an older default-jdk; fleetd sets maven.compiler.release=25, so
# The runner image ships an older default-jdk; bridged sets maven.compiler.release=25, so
# provision the JDK explicitly rather than apt-installing whatever "default" means today.
- name: Set up JDK 25
uses: actions/setup-java@v4
@@ -32,7 +32,7 @@ jobs:
mvn -version
- name: Build and test
working-directory: fleetd
working-directory: bridged
# This IS the mock-socket surface CB-503 asks for: the pom's `default-excludes` profile
# already sets excludedGroups=contract, so the @Tag("contract") tests — which need a live
# herdr socket and a RabbitMQ container — are excluded without any flag here. Everything
@@ -46,7 +46,7 @@ jobs:
# the log instead, where they are actually readable.
- name: Failing test output
if: failure()
working-directory: fleetd
working-directory: bridged
run: |
for f in target/surefire-reports/*.txt; do
[ -f "$f" ] || continue
@@ -59,7 +59,7 @@ jobs:
# AmqpReplyInboxContractTest reads AMQP_URI (set below to the service's network alias) and binds
# straight to it — no Docker, no skipped tests. This separation (build job hermetic and
# Docker-free; contract job broker-provided) is deliberate — see the default-excludes/contract
# profiles in fleetd/pom.xml. `setup-java` provides the JDK only; Maven is installed separately,
# profiles in bridged/pom.xml. `setup-java` provides the JDK only; Maven is installed separately,
# exactly as in the build job above.
contract:
runs-on: ubuntu-latest
@@ -69,6 +69,8 @@ jobs:
env:
RABBITMQ_DEFAULT_USER: guest
RABBITMQ_DEFAULT_PASS: guest
ports:
- 5672:5672
env:
# Service containers are reachable from the job by their network alias on their internal port.
AMQP_URI: amqp://guest:guest@rabbitmq:5672
@@ -91,12 +93,12 @@ jobs:
# runs against the RabbitMQ service container (AMQP_URI). Pinned to the one contract test to
# avoid re-running the unit suite already covered by the `build` job.
- name: Contract tests
working-directory: fleetd
working-directory: bridged
run: mvn -B -Pcontract test -Dtest=AmqpReplyInboxContractTest
- name: Failing test output
if: failure()
working-directory: fleetd
working-directory: bridged
run: |
for f in target/surefire-reports/*.txt; do
[ -f "$f" ] || continue
+4 -4
View File
@@ -14,8 +14,8 @@
.env
.envrc
# Daemon runtime artefacts. fleetd appends its log wherever it is launched from, so both the
# repo root and fleetd/ collect one; neither belongs in git.
fleetd.out
fleetd/fleetd.out
# Daemon runtime artefacts. bridged appends its log wherever it is launched from, so both the
# repo root and bridged/ collect one; neither belongs in git.
bridged.out
bridged/bridged.out
logs/
+1 -1
View File
@@ -1,3 +1,3 @@
[submodule "wiki"]
path = wiki
url = ssh://git@git.ltms.dev:2224/fleet/fleetd.wiki.git
url = ssh://git@git.ltms.dev:2224/lms/claude-bridge.wiki.git
+2 -2
View File
@@ -3,7 +3,7 @@ description: Refine work into clear, independent units before implementation.
mode: primary
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You are an architect in this fleet. You refine work before anyone builds it: scope,
acceptance criteria, risks, and a unit split. You read the repo and write analysis.
@@ -13,7 +13,7 @@ A design task is worked by two architects. Design alone first, then exchange and
say plainly where you disagree. Do not concede just to agree.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
investigate it further. Use `bridge_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
+2 -2
View File
@@ -3,14 +3,14 @@ description: Implement one assigned unit, test it, and open a pull request.
mode: primary
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You implement the one unit you were given and nothing else. Work in your assigned
git worktree and branch. Never check out, rebase onto, or push to `main`. Confirm
the worktree root and branch before you edit. Use only paths under that root.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
investigate it further. Use `bridge_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
+2 -2
View File
@@ -3,7 +3,7 @@ description: Review one assigned scope and report the most important real issue.
mode: primary
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You review the diff you were given. Report bugs, risks, and missing tests. You do
not change code.
@@ -12,7 +12,7 @@ Read the whole assigned scope before judging it. Review only that scope. If you
something outside it, note it in one line and do not investigate it further. Do not
run the build. The owner makes changes and runs checks.
Use `fleet_ask{question}` only when a decision belongs to the lead, such as an
Use `bridge_ask{question}` only when a decision belongs to the lead, such as an
unclear requirement or two defensible fixes. Do not ask about something you can
decide by reading more code.
+60 -62
View File
@@ -4,35 +4,35 @@
> **Canonical block.** Everything down to §Layering is the portable bridge charter, copied verbatim
> into every project that mounts the bridge MCP. Keep it byte-identical with the template in the
> wiki ([Use Cases](https://git.ltms.dev/fleet/fleetd/wiki/7-Use-Cases) → *The portable
> wiki ([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable
> CLAUDE.md block*); improvements go to the template first, then out to each project. Anything
> specific to *this* repo lives under §Project addendum below, never inline above it.
If no `fleet_*` MCP tools are mounted in this session, this section does not apply — skip it.
If no `bridge_*` MCP tools are mounted in this session, this section does not apply — skip it.
`fleetd` is the **sole communication gateway** between agents here. The orchestrating session (the
`bridged` is the **sole communication gateway** between agents here. The orchestrating session (the
**primary**) and every delegated peer (a **member**) mount the *same* MCP server and talk only
through its `fleet_*` tools. No session addresses a peer, a broker, or the network directly.
through its `bridge_*` tools. No session addresses a peer, a broker, or the network directly.
### Which role am I? — settle this before acting
**Every role reads this file.** A member runs in a git worktree of this same repo, so it inherits
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
**Call `fleet_whoami`.** It returns `primary`, `worker`, or `architect`, resolved by the daemon from
**Call `bridge_whoami`.** It returns `primary`, `worker`, or `architect`, resolved by the daemon from
your connection — unforgeable, and the same resolution its authorization gate uses. A worker also
carries its `sessionId`, `profile`, `worktree` and `branch`; an architect carries the slot name it
was bound to. Don't infer what you can ask.
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
fires: the reply charter in your system prompt (*"You are a spawned member in the
claude-bridge fleet"*) ⇒ **spawned member**; fleet tools prefixed `mcp__fleet__*` ⇒ **spawned
fires: the reply charter in your system prompt (*"You are an off-subscription worker in the
claude-bridge fleet"*) ⇒ **spawned member**; bridge tools prefixed `mcp__bridge__*` ⇒ **spawned
member** (the launcher fixes that mount name; a primary's mount is named by whoever wrote its
`.mcp.json`, so it varies — and a member spawned before CB-632 still says `mcp__bridge__*`); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
`.mcp.json`, so it varies); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect —
only `fleet_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
only `bridge_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate —
loud and self-correcting — while a member acting as the primary ends its turn with no `fleet_reply`,
loud and self-correcting — while a member acting as the primary ends its turn with no `bridge_reply`,
and the sender silently receives nothing. Fail toward the recoverable error.
### Invariants — both roles, no exceptions
@@ -41,7 +41,7 @@ and the sender silently receives nothing. Fail toward the recoverable error.
stays on subscription; only the bridge puts a member off it, at spawn. Mounting the bridge must
never move a session across that boundary.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `fleet_*` call is silently discarded.
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead or architect**;
reply/ask are only-as-itself — any peer may answer for its own pane, and for no other. A call
@@ -57,10 +57,10 @@ and the sender silently receives nothing. Fail toward the recoverable error.
**Delegate by default — that is the job.** With the bridge mounted you are an orchestrator on a
metered subscription, and workers are cheap, parallel, and disposable. The default answer to "who
does this?" is **a worker**, not you. Reach for `fleet_send` before you reach for `Edit`. The steps
does this?" is **a worker**, not you. Reach for `bridge_send` before you reach for `Edit`. The steps
below are the procedure — run them in order, every task, not only the big ones.
0. **Know your role** — `fleet_whoami`, once per session, before anything else.
0. **Know your role** — `bridge_whoami`, once per session, before anything else.
1. **Split.** Write the unit list. Every unit carries: scope · the files or PR in question ·
acceptance criteria · exactly what to report back. A unit with no acceptance criteria is not
ready to delegate — refine it or keep it.
@@ -70,17 +70,17 @@ below are the procedure — run them in order, every task, not only the big ones
delegate. The keep-list is closed: the conversation with the user, decomposition and planning,
the final judgment call, verification, merges, and anything that depends on context only you
hold. Nothing else is yours by default.
3. **Spawn every delegated unit first** — `fleet_spawn{profile, worktree:true, ticket}`, one per
3. **Spawn every delegated unit first** — `bridge_spawn{profile, worktree:true, ticket}`, one per
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model and cost, not in
tier, so the default is rarely what you want.
4. **Then send them all** — `fleet_send{sessionId, content, wait:false}`. Line 1 of every brief is
4. **Then send them all** — `bridge_send{sessionId, content, wait:false}`. Line 1 of every brief is
`Load the <name> skill.` naming the worker's playbook; those skills are opt-in and that line is
what makes them reliable. Where the project ships no such skill, spell the procedure out in the
brief instead. The brief is self-contained — the worker sees your message and the repo, nothing
of your context, your plan, or your screen.
5. **Collect** — `fleet_poll{ticket}` → `fleet_ack{target, msgId}`. Answer a worker's `fleet_ask`
with `fleet_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`fleet_status`, never by reading its terminal; it also reports an open question and the `turnId`
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{target, msgId}`. Answer a worker's `bridge_ask`
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`bridge_status`, never by reading its terminal; it also reports an open question and the `turnId`
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
@@ -93,11 +93,11 @@ below are the procedure — run them in order, every task, not only the big ones
read it yourself.
8. **Adjudicate, merge, tear down — yours alone.** Read the diff yourself: fully if it is small,
targeted at the reported findings and the risky paths if it is large. Reviewer findings direct
your attention; they never substitute for it. Then merge, then `fleet_stop{paneId}`.
your attention; they never substitute for it. Then merge, then `bridge_stop{paneId}`.
**Steps 3 and 4 are separate on purpose** — spawning and sending in one loop is how parallel work
silently becomes serial, and it is the most common way this layer is wasted. For the same reason,
prefer `wait:false` + `fleet_poll` for anything non-trivial: a blocking `fleet_send` is capped by
prefer `wait:false` + `bridge_poll` for anything non-trivial: a blocking `bridge_send` is capped by
*your own* MCP client call timeout (~60s), well below the task's real runtime.
**Delegating does not delegate responsibility.** Workers open PRs; you are the gate. Never delegate
@@ -105,22 +105,21 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
| Intent | Tool |
|---|---|
| Confirm your own role | `fleet_whoami` |
| See backends available | `fleet_profiles` |
| Start a member | `fleet_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` → `sessionId` + `paneId` |
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) · one peer's state: `fleet_status{sessionId}` |
| Delegate (blocking) | `fleet_send{sessionId, content}` |
| Delegate (long task) | `fleet_send{sessionId, content, wait:false}` → ticket → `fleet_poll{ticket}` |
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `fleet_reply{content}` — the one case a lead replies |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
| Tear down a member | `fleet_stop{paneId}` |
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) · one peer's state: `bridge_status{sessionId}` |
| Delegate (blocking) | `bridge_send{sessionId, content}` |
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** | `bridge_send{sessionId: <their terminal>, content}` — `bridge_list` → `leads` reports it. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `bridge_reply{content}` — the one case a lead replies |
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
| Tear down a member | `bridge_stop{paneId}` |
### Lead ↔ lead — coordinate, never delegate
`fleet_list` returns `leads` alongside `members`; your own row carries `self: true`. Every other row
`bridge_list` returns `leads` alongside `members`; your own row carries `self: true`. Every other row
is a peer — an orchestrator with its own context, its own members, and its own judgment. An empty
`members` array means no members are spawned; it says nothing about peers.
@@ -141,7 +140,7 @@ The traffic between leads is coordination and nothing else:
against the code, and re-run the build. A peer's correction gets the same treatment — right or
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
Being messaged by a peer does not make you its worker: answer with `fleet_reply`, and push back on
Being messaged by a peer does not make you its worker: answer with `bridge_reply`, and push back on
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
you.
@@ -149,11 +148,11 @@ you.
1. **Load the playbook skill the lead named** before doing anything else.
2. **Do the assigned scope only.** Note anything you spot outside it in one line; don't go hunt it.
3. **`fleet_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
3. **`bridge_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
defensible fixes, "bug or intended?"). It blocks and you resume the *same* turn with the answer.
Don't ask what you could decide yourself.
4. **End the turn with exactly one `fleet_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `fleet_reply` ⇒ the sender gets nothing and the exchange stalls.
4. **End the turn with exactly one `bridge_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `bridge_reply` ⇒ the sender gets nothing and the exchange stalls.
Do **not** lean on the completion fallback to carry your answer for you: when you end a turn
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
@@ -172,7 +171,7 @@ you.
| Layer | Scope | Reaches |
|---|---|---|
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `fleet_reply`* | every spawned member, at launch, every peer kind — never a lead |
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every spawned member, at launch, every peer kind — never a lead |
| **this section** | protocol + orchestration policy | primary **and** every member that reads the repo — tracked in git, so worktrees inherit it |
| role agent definition files | role contract and per-job procedure | a member whose launcher binds its role to the matching file in its worktree |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a member told to load one |
@@ -185,17 +184,16 @@ must obey belongs in the charter, not here.
## Project addendum — claude-bridge (not part of the canonical block)
- **This repo is the bridge.** The daemon is `fleetd`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/FleetMcp` (tools), `auth/Authz` (the role table),
- **This repo is the bridge.** The daemon is `bridged`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/BridgeMcp` (tools), `auth/Authz` (the role table),
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
`reviewer` (scoped review → one structured finding). Name one in every delegation.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace) and
`fleets-status` (report every fleet that shares one LavinMQ instance).
`port-to-opencode` (make an OpenCode session a participant in this workspace).
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **Flows and the error model** — rendezvous, `fleet_ask`, detached delivery, turn-done fallback —
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
are diagrammed in `docs/MCP-Contract.md` **§6 only**. The rest of that page is a pre-build design
doc whose tool names, parameter names and REST paths never caught up with the code, so do not use
it as the tool reference (CB-609). Section 6 is kept out of this file because this file loads into
@@ -203,7 +201,7 @@ must obey belongs in the charter, not here.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
**A merge is not a deployment.** The running `bridged` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
@@ -214,9 +212,9 @@ and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-fleetd.sh --check # report state, change nothing
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
scripts/redeploy-bridged.sh --check # report state, change nothing
scripts/redeploy-bridged.sh # build, confirm drain, restart, verify
scripts/redeploy-bridged.sh --yes # skip the drain prompt (fleet already checked)
```
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
@@ -233,17 +231,17 @@ if the script is unavailable or a step fails, this is what it was protecting you
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
2. **Drain live members first.** `bridge_list`, then `bridge_stop` each member, and collect anything
you still want with `bridge_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
4. **Re-check identity afterwards.** Call `bridge_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
5. **Prove the new jar is the one running.** Confirm a *fresh* `bridged listening` line at the end of
`bridged/bridged.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
@@ -252,7 +250,7 @@ it is one auditable command, so the operator allow-lists it once instead of appr
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
{ "permissions": { "allow": ["Bash(scripts/redeploy-bridged.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
@@ -270,15 +268,15 @@ Before you call any work done, check the row that matches what you touched:
| You changed… | Re-read and update… |
|---|---|
| a `fleet_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
| a `bridge_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
| `Authz` / the role table | invariant 3, and the primary-only vs worker-only claims |
| `ConnectionIdentity` / how a caller is resolved | the `fleet_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__fleet__*`), and the layering table's top row |
| `ConnectionIdentity` / how a caller is resolved | the `bridge_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__bridge__*`), and the layering table's top row |
| the injector / status gating | invariant 4 |
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
| **anything an operator can use, configure, or observe** — an MCP tool, a `fleetd.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
| **anything an operator can use, configure, or observe** — an MCP tool, a `bridged.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
That last row is not bookkeeping. Chapters 1–10 answer *how is this built* and *why this way*;
none of them has a home for *what can it do and how do I turn it on*, so for twenty tickets a
@@ -289,7 +287,7 @@ is a Roadmap line. A change that touches none of the three earns no entry, and t
outcome rather than an omission.
Then **propagate**: the block in this file and the template in the wiki
([Use Cases](https://git.ltms.dev/fleet/fleetd/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
block*) must stay byte-identical, and other projects carrying the block need the same edit. Verify
rather than trust:
@@ -312,10 +310,10 @@ PY
Two IDE MCP servers are connected: **intellij-index** (semantic code intelligence) and
**jetbrains** (file problems, reformat, debugger). IntelliJ has multiple projects open; our
module is **`fleetd`**. Always pass these to IDE MCP tools:
module is **`bridged`**. Always pass these to IDE MCP tools:
- `project_path` = `/Users/dai.ha/LTMS/claude-bridge/fleetd`
- IDE paths are relative to `fleetd/` (e.g. `src/main/java/dev/ltms/fleet/...`)
- `project_path` = `/Users/dai.ha/LTMS/claude-bridge/bridged`
- IDE paths are relative to `bridged/` (e.g. `src/main/java/dev/ltms/bridged/...`)
### After editing any file — mandatory
@@ -329,7 +327,7 @@ module is **`fleetd`**. Always pass these to IDE MCP tools:
whole-project gate before declaring work done or committing.
**Whenever dependencies change (or a `pom.xml` edit), validate CVEs with
`jetbrains get_file_problems{filePath: "fleetd/pom.xml"}`** — its Mend.io check reflects the
`jetbrains get_file_problems{filePath: "bridged/pom.xml"}`** — its Mend.io check reflects the
dependencies on disk. (Note: `ide_diagnostics` / intellij-index does NOT re-resolve dependencies
after a pom edit without a full Maven reimport, so it reports stale CVE results — don't trust it
for this.) Treat a CVE warning like any other: bump to a patched version and confirm
+27 -29
View File
@@ -9,23 +9,23 @@ Sibling of [`crush-bridge`](https://git.ltms.dev/systems/vms) (which drives a he
process*, so it inherits `CLAUDE.md`, hooks, skills, and MCP — just pointed at a
cheaper/local model.
## Leading approach — herdr-centric message server (`fleetd`)
## Leading approach — herdr-centric message server (`bridged`)
A small always-on message server, **`fleetd`**, controls
A small always-on message server, **`bridged`**, controls
[herdr](https://herdr.dev) (an agent multiplexer) over its Unix-socket API and exposes a
clean 2-way messaging API as an **MCP server that both the primary and the workers mount** —
one unified Claude setup and the **sole communication gateway** (REST/SSE stays for non-Claude
clients; any broker is `fleetd`-internal, below the gateway).
herdr owns the PTYs, multiplexing, persistence, and **agent-status events**; `fleetd` owns
clients; any broker is `bridged`-internal, below the gateway).
herdr owns the PTYs, multiplexing, persistence, and **agent-status events**; `bridged` owns
policy (subscription boundary, session lifecycle, status-gated delivery) and the client
contract. A Claude member launches with `ANTHROPIC_BASE_URL` pointed at the gateway,
`https://llm.ltms.dev/anthropic`, plus a bearer token; the lead stays env-clean and calls
`fleetd`'s MCP tools. See the wiki's **[13 User Guide](wiki/13-User-Guide.md)** to run it.
`bridged`'s MCP tools. See the wiki's **[13 User Guide](wiki/13-User-Guide.md)** to run it.
```mermaid
flowchart LR
OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
subgraph BD["fleetd — standalone daemon (not a claude process)"]
subgraph BD["bridged — standalone daemon (not a claude process)"]
SRV["SERVER face<br/>MCP · REST/SSE · policy"]
CLI["CLIENT face<br/>status-gated injector · herdr socket"]
SRV --> CLI
@@ -34,8 +34,8 @@ flowchart LR
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
M["llm.ltms.dev<br/>(the one gateway)"]
OPUS -->|"MCP fleet_send (blocks)"| SRV
W -.->|"MCP fleet_reply"| SRV
OPUS -->|"MCP bridge_send (blocks)"| SRV
W -.->|"MCP bridge_reply"| SRV
CLI -->|"Unix socket<br/>send_text · events.subscribe"| HERDR
HERDR -->|"drives PTY"| W
W -->|"inference"| M
@@ -47,26 +47,24 @@ flowchart LR
```
- **Subscription boundary:** the *primary* never sets `ANTHROPIC_BASE_URL` (stays on
Pro/Max). Only the *secondary* process is off-subscription — and `fleetd` itself is a
Pro/Max). Only the *secondary* process is off-subscription — and `bridged` itself is a
plain daemon (no Anthropic quota), so it may poll/subscribe freely.
- **One gateway (unified MCP setup):** `fleetd` is the **sole communication path** for every
- **One gateway (unified MCP setup):** `bridged` is the **sole communication path** for every
Claude session. Primary and workers each mount it as an MCP server (one `claude mcp add`
line, same on both) and talk over MCP tools — `fleet_send` / `fleet_reply` /
`fleet_status` (with `fleet_ask` planned for the blocked-worker path). **No Claude session
ever addresses a broker, a peer, or the network
directly**; any queue is `fleetd`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so
mounting the bridge is subscription-safe by construction.
**Tool naming:** the tools are `fleet_*` (renamed from `bridge_*` in CB-622). The old
`bridge_*` names were removed in CB-634 — only `fleet_*` answers now.
- **How the primary consumes a reply:** a single **blocking MCP call** (`fleet_send`);
`fleetd` holds it open until the worker calls `fleet_reply` or its turn hits
line, same on both) and talk over MCP tools — `bridge_send` / `bridge_reply` /
`bridge_status` (with `bridge_ask` planned for the blocked-worker path). **No Claude session
ever addresses a broker, a peer, or the network
directly**; any queue is `bridged`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so
mounting the bridge is subscription-safe by construction.
- **How the primary consumes a reply:** a single **blocking MCP call** (`bridge_send`);
`bridged` holds it open until the worker calls `bridge_reply` or its turn hits
`agent_status=done`, then returns the reply as the tool result. No cross-turn busy-poll, so
no quota burn. SSE is an optional side-channel for humans/dashboards watching status.
- **Worker → primary** rides `fleetd`'s **MCP rendezvous** — the reply resolves the primary's
blocking call (or, for detached work, `fleetd` **injects the primary's idle pane** when it's
- **Worker → primary** rides `bridged`'s **MCP rendezvous** — the reply resolves the primary's
blocking call (or, for detached work, `bridged` **injects the primary's idle pane** when it's
ready), so *no keystroke-into-primary and no broker are involved, even single-host*. The one
exception: a split-host primary that isn't a herdr pane wakes via its own `Stop`-hook, which
polls **`fleetd`** (never a broker). See the wiki for the two topologies.
polls **`bridged`** (never a broker). See the wiki for the two topologies.
- **Different model per process** sidesteps Claude Code's lack of per-subagent provider
routing — the worker isn't a subagent, it's its own configured process.
- **AgentAPI** ([`coder/agentapi`](https://github.com/coder/agentapi)) is retained only as a
@@ -75,11 +73,11 @@ flowchart LR
## Docs
Full design, setup, and operations live in the **[wiki](https://git.ltms.dev/fleet/fleetd/wiki)**,
Full design, setup, and operations live in the **[wiki](https://git.ltms.dev/lms/claude-bridge/wiki)**,
vendored here as a submodule under [`wiki/`](./wiki):
```bash
git clone --recurse-submodules ssh://git@git.ltms.dev:2224/fleet/fleetd.git
git clone --recurse-submodules ssh://git@git.ltms.dev:2224/lms/claude-bridge.git
# or, after a plain clone:
git submodule update --init
```
@@ -89,7 +87,7 @@ Gitea wiki.
## Status
🟢 **Implemented & dogfooded** — the herdr-centric **`fleetd`** message server is built and in
🟢 **Implemented & dogfooded** — the herdr-centric **`bridged`** message server is built and in
real use: an Opus primary delegates tasks to off-subscription workers that reply through the
bridge (code reviews delegated this way have produced committed bug fixes). Selected as the
primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a
@@ -100,9 +98,9 @@ tests run separately via `mvn test -Pcontract`):
- **Core gateway** — herdr socket client (contract-tested vs live 0.7.0); guard-checked worker
spawn with `ANTHROPIC_BASE_URL` injected only into the worker's env; status-gated injector;
blocking `fleet_send` with reply rendezvous; MCP server as a thin adapter over the REST core.
- **MCP tools** — `fleet_send` / `fleet_reply` / `fleet_status` (messaging) and `fleet_spawn`
/ `fleet_list` / `fleet_stop` / `fleet_profiles` / `fleet_poll` (fleet). Caller identity is
blocking `bridge_send` with reply rendezvous; MCP server as a thin adapter over the REST core.
- **MCP tools** — `bridge_send` / `bridge_reply` / `bridge_status` (messaging) and `bridge_spawn`
/ `bridge_list` / `bridge_stop` / `bridge_profiles` / `bridge_poll` (fleet). Caller identity is
connection-based (loopback peer PID → herdr pane), so the same mount serves primary and workers.
- **Delivery reliability** — completion fallback (a confirmed `working→idle` turn resolves a
send); async fire-and-poll (beats the caller's MCP call timeout for long tasks); and failure
@@ -110,7 +108,7 @@ tests run separately via `mvn test -Pcontract`):
- **Fleet** — multiple worker profiles, each with an independent base_url guard check; workers
inherit the primary's working directory (never `$HOME`); a readiness gate holds delivery until
a worker's Claude has connected the bridge MCP (no paste lost into its boot window).
- **Blocked-worker path** — `fleet_ask` reverse rendezvous: a worker pauses its delegated turn to
- **Blocked-worker path** — `bridge_ask` reverse rendezvous: a worker pauses its delegated turn to
ask the primary and resumes the *same* turn with the answer (CB-205).
- **Session lifecycle** — session manager with spawn/reuse/recycle, `idle_ttl` reaper, `context_cap`,
and graceful drain on shutdown (CB-301/CB-303); per-worker git worktrees on their own branch with
-166
View File
@@ -1,166 +0,0 @@
# CB-137 / fleetd issue #137 — report
## Real root cause (not the hypothesis in the ticket)
I read `MessageService.java` and `Rendezvous.java` before changing anything. The mechanism is real,
but the exact place it happens is `MessageService.answer()`, not "the reply goes to the inbox on
purpose" in general.
1. A lead delegates with `fleet_send{wait:false}` → `sendAsync()` creates a `Task` and runs `send()`
on a background virtual thread with a 30-minute internal budget (`ASYNC_TIMEOUT_MS`).
2. The worker calls `fleet_ask`. That resolves the open rendezvous waiter with `Kind.QUESTION`, so
`send()` returns immediately and the `Task` is left open (its `future` stays unresolved — see the
comment in `sendAsync`'s lambda: "Keep the accepted owner until answer() finishes it").
3. The lead answers with `fleet_send{turnId, content}`. This calls `FleetMcp.answer()` →
`MessageService.answer(turnId, content, timeout)`. The `timeout` here is **not** the generous
30-minute async budget — it is the MCP tool's own bounded wait: `DEFAULT_TIMEOUT_MS = 25_000`,
clamped to at most `MAX_TIMEOUT_MS = 120_000` (`FleetMcp.java:71-72,495,512`). This is the same
~60–120s window every blocking `fleet_send` call is capped at (documented elsewhere as "the
caller's own MCP client call timeout").
4. `answer()` opens a **fresh** rendezvous waiter for the worker session and blocks on it for at most
that window. If the worker's resumed turn takes longer than that to actually finish (very
plausible — the resumed turn can mean more edits, a build, a commit, a push, opening a PR), the
wait times out. On timeout, `answer()`'s `finally` block unconditionally calls
`rendezvous.close(workerSession, reply)`, **removing the waiter from the map**, and returns
`Outcome.TIMED_OUT_WORKING` to the lead.
5. The worker keeps working, unaware anything happened, and eventually calls `fleet_reply`. That
reaches `MessageService.reply(session, content)`, which tries `rendezvous.resolve(session,
content)` — but the waiter was already closed in step 4, so `resolve` returns `false`. `reply()`
then falls back to `inbox.publish(...)` and marks `strandedReplies.put(session, true)`
(CB-640 bookkeeping) — the reply is safely held, but **the async `Task`'s `future` is never
completed**.
6. `fleet_poll{ticket}` keeps returning `PENDING` forever (the `Task` never resolves) — until the
lead eventually calls `fleet_stop`. That fires `sessions.onRelease` → `messages.abandon(target,
reason)` (`Fleetd.java:481-497`), where `reason` is built with the exact text from the bug report
("the worker session was released before it replied; worktree=... branch=... snapshot=...",
`Fleetd.java:484-487`). `abandon()`'s loop finds the still-open `Task` (`question == null`, future
not done) and completes it as `WORKER_FAILED` with that misleading reason — even though the
worker's real reply is sitting, intact, in the inbox the whole time.
So: the reported behaviour is correct, and the specific trigger is `answer()`'s own bounded wait
being shorter than the worker's real resumed-turn time — not anything to do with the ~55s
`fleet_ask` window itself (that part, issue #61, is untouched).
## Fix
Two changes in `fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java`, both scoped to the
ticket/reply routing and the terminal-state text — `fleet_ask`'s own window and mechanics are
untouched.
**1. `reply()` — priority 1 (the ticket resolves with the real reply).**
Before falling back to the inbox, `reply()` now looks for an async `Task` that is specifically in the
"already answered but not yet resolved" state (`question == null`, `turnId != null` — set once
`answer()` has cleared the question but before anything completed the future, `!future.isDone()`).
If one exists for this `target`, the worker's reply completes that `Task`'s future directly as
`Outcome.REPLIED` with the real content, and the reply never touches the inbox at all. A task that
was never asked has `turnId == null` and can never match, so ordinary (no-`fleet_ask`) delegations
are unaffected — they already resolve through the pre-existing rendezvous fast path.
I chose this over leaving `answer()`'s own timeout behaviour untouched and instead keeping its
rendezvous waiter open in the background: that alternative works but reopens the "at most one
waiter per session" invariant (`Rendezvous.open` throws on a double-open) to a new class of races
with a fresh send arriving mid-window. The `send()` path already guards against sending into an
answered-but-still-resolving worker via `hasAsyncQuestion(target)` (checks `asyncTasksByTurn`,
which still holds the task until it resolves), so routing through `reply()` gets the same protection
without touching `answer()`'s waiter lifecycle at all — the smaller, safer diff.
**2. `abandon()` — priority 3 (required independently, "even if you fix (1)").**
Before marking any of a released target's still-open tasks `WORKER_FAILED`, `abandon()` now checks
`hasStrandedReply(target)` (the existing CB-640 fact — true whenever the *last* `reply()` for this
target fell through to the inbox). If true, it drains the inbox (`recoverStrandedReply`) and — if it
actually finds a message — completes the task as `REPLIED` with that real content instead of writing
the failure. This is deliberately a **separate** check from fix 1: fix 1 already prevents the
inbox-stranding from happening in the exact scenario this ticket describes, so by the time
`abandon()` runs the task is normally already resolved and `abandon()`'s `complete()` call is a
harmless no-op. This second check exists so that if some *other* future path ever strands a reply
in the inbox without resolving its ticket, `abandon()` still refuses to report a false failure —
"if a reply reached any sink for that turn, the terminal state is done," per the ticket. I verified
both are required by disabling each independently and confirming the two new tests fail (see below).
**Priority 4 (the snapshot/worktree hint).** Handled as a consequence of both fixes rather than a
separate branch: once a task resolves as `REPLIED` (via either fix), `abandon()` never calls
`new Reply(Outcome.WORKER_FAILED, reason)` for that task at all, so the "the worker session was
released before it replied; worktree=... branch=... snapshot=..." text is never constructed or
attached to that ticket's outcome. It still appears, correctly, for a task that never got a reply
(the existing `abandonFailsEveryPendingAsyncTicketForTheReleasedTarget` /
`anAbandonedAsyncTaskPollsAsFailedNotPending` tests still pass unchanged).
**Priority 2** was not needed — fix 1 makes `fleet_poll{ticket}` return the actual reply (the
higher-priority option), so I did not fall back to "the ticket merely resolves as done with no
content."
## Tests — driven through the real delegation path, not the reply sink directly
Both new tests in `fleetd/src/test/java/dev/ltms/fleet/msg/MessageServiceTest.java` go through
`sendAsync` → `injectDelivery` → `ask` → `answer` (with a short timeout, so it genuinely times out,
mirroring the ~25–120s real MCP-call bound vs. a longer resumed turn) → `reply` → `poll`/`abandon`.
No test constructs a `Reply` and hands it to a sink directly.
- `aReplyAfterAnswerTimesOutStillCompletesTheAsyncTicket` — asserts `fleet_poll{ticket}` (via
`messages.poll`) reaches `Phase.DONE` with the worker's actual reply text and
`replySource() == "reply"`, and that `hasStrandedReply(T)` stays `false` (proves the reply never
touched the inbox at all — fix 1 caught it).
- `fleetStopAfterAnOrphanedReplyDoesNotFailTheTicket` — same setup, then calls `abandon(T, "the
worker session was released before it replied")` (what `fleet_stop` triggers) and asserts it
returns `false` (no failure recorded) and the ticket still polls `DONE` with the real reply.
**Proof both fail without the change.** I temporarily short-circuited both new private methods
(`askAnsweredAsyncTask` → always `null`, `recoverStrandedReply` → always `null`) — i.e. disabled
both fixes — and ran just these two tests:
```
[ERROR] Tests run: 2, Failures: 2, Errors: 0, Skipped: 0
dev.ltms.fleet.msg.MessageServiceTest.aReplyAfterAnswerTimesOutStillCompletesTheAsyncTicket
org.opentest4j.AssertionFailedError: expected: <DONE> but was: <PENDING>
dev.ltms.fleet.msg.MessageServiceTest.fleetStopAfterAnOrphanedReplyDoesNotFailTheTicket
org.opentest4j.AssertionFailedError: a reply already arrived, so nothing here is a genuine failure
==> expected: <false> but was: <true>
```
This is the exact bug: the ticket stays `PENDING` forever, and `abandon()` reports `true` (a
failure) even though a reply had already arrived. I then restored both fixes (verified with
`grep -n "TEMP #137-proof"` finding nothing) and re-ran — both pass.
## Build
Ran from `fleetd/`, unpiped, full output read (not `| tail`):
```
mvn clean install
...
[INFO] Tests run: 1039, Failures: 0, Errors: 0, Skipped: 0
[INFO] BUILD SUCCESS
[INFO] Total time: 36.315 s
```
Main was at 1037 tests; this branch adds the 2 new tests above → 1039, all green, `exit=0`.
## What I could NOT check
- No IDE tooling is mounted for me (worker), so no `ide_diagnostics`/IntelliJ inspection pass — only
`mvn clean install` (compiler + full test suite), as the worker procedure allows.
- I cannot restart the daemon or dogfood this live — I have no forge/daemon control. This is
unverified against a real herdr pane, a real MCP client's ~60s call cap, or a real worker session;
everything above is verified only through the JUnit fixture's simulated timing
(`FakeHerdr`/`injector.onStatus`/direct `messages.answer(...,150)` calls), not a live fleet.
A primary should still consider a short live dogfood (an async delegation that asks, gets answered,
and takes longer than ~2 minutes to reply) before calling this closed.
- I did not touch, and did not re-verify, the `fleet_ask` ~55s window itself (issue #61) — out of
scope per the brief.
## Scope note (not investigated further)
`answer()`'s nested/double-`fleet_ask` case (the worker asks a second question before ever
replying to the first answer) has some pre-existing behaviour around which `turnId` a `QUESTION`
resolution gets attributed to that I did not fully untangle — it predates this change, my fix does
not touch it, and it is unrelated to the reported defect. Flagging only; not investigated further.
## Handoff
- Branch: `worker/cb-137-ask-ticket-e7760c-2`
- Worktree root: `/Users/dai.ha/LTMS/.bridged-worktrees/734324-2`
- Files changed:
- `fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java`
- `fleetd/src/test/java/dev/ltms/fleet/msg/MessageServiceTest.java`
- `REPORT-cb137.md` (this file)
- Build: `Tests run: 1039, Failures: 0, Errors: 0, Skipped: 0` / `BUILD SUCCESS` (verbatim above)
+14
View File
@@ -0,0 +1,14 @@
# Build output
target/
dependency-reduced-pom.xml
# Local runtime config (copy from bridged.example.yaml)
bridged.yaml
# CB-505 audit trail + daemon stdout/stderr — runtime records, never source
logs/
# Editor / OS
*.iml
.idea/
.DS_Store
@@ -1,10 +1,10 @@
# fleetd configuration (example). Copy to fleetd.yaml and adjust.
# bridged configuration (example). Copy to bridged.yaml and adjust.
#
# fleetd is the sole gateway between primary/worker Claude sessions and herdr.
# bridged is the sole gateway between primary/worker Claude sessions and herdr.
# It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL.
# REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token`
# below — fleetd REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
# below — bridged REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
bind:
host: 127.0.0.1
port: 8765
@@ -21,7 +21,7 @@ bind:
# bind — the daemon fails fast otherwise, because "unauthenticated ⇒
# primary" on a reachable port would hand spawn/stop/send to anyone.
# tokenEnv → host env var holding the token (never the literal value). Default
# FLEETD_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
# BRIDGED_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
#
# TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and
# let it own certificate lifecycle, e.g.
@@ -29,13 +29,13 @@ bind:
# The broker link gets TLS from its own URI (amqps://…) — see `broker` below.
# auth:
# mode: token
# tokenEnv: FLEETD_API_TOKEN
# tokenEnv: BRIDGED_API_TOKEN
# Optional pinned primary terminal (CB-307). Names the herdr pane the PRIMARY itself runs in:
# a caller whose connection maps to this pane resolves as the primary (no credential needed —
# the pane mapping is as unforgeable as a worker's), and reply nudges are pushed to it.
# REQUIRED when the primary runs inside a herdr pane — without it the pane match reads the
# primary as a worker and refuses spawn/send/stop. Get the id from fleet_whoami; re-pin if
# primary as a worker and refuses spawn/send/stop. Get the id from bridge_whoami; re-pin if
# the primary moves panes.
# primary:
# terminal: term_0123456789abcd
@@ -48,13 +48,13 @@ bind:
# every orchestration call. List each lead's pane here and all of them resolve as leads.
#
# tab → the ONLY field identity depends on (CB-579); the exact label of the tab hosting the lead.
# Label the tab yourself, or let fleetd label one it launches — see `fleet.leaders:` below.
# kind/model → descriptive; they document what runs in the pane and are echoed by fleet_whoami
# Label the tab yourself, or let bridged label one it launches — see `fleet.leaders:` below.
# kind/model → descriptive; they document what runs in the pane and are echoed by bridge_whoami
#
# A lead's tab must already carry its label (or be launched by fleetd, which labels it) — there is
# A lead's tab must already carry its label (or be launched by bridged, which labels it) — there is
# no terminal id to paste in and nothing to re-pin when the session restarts: the tab survives, so
# the same label resolves the same lead again on the next scan.
# `fleet_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
# `bridge_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
# IS a primary for authorization, so nothing that keys on the role breaks.
#
# KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single
@@ -63,13 +63,13 @@ bind:
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
#
# Two things stop the tab-name convention from becoming a way to claim leadership: the configured
# member spaces are excluded from the scan, so nothing fleetd places can land in a matching tab;
# member spaces are excluded from the scan, so nothing bridged places can land in a matching tab;
# and startup REFUSES a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel`
# override, also matches — so the two namespaces cannot overlap by accident. The label is a NAME,
# never a capability: what a pane may do is decided by the role the daemon resolves for it.
# CB-551: IDLE-LEAD HEARTBEAT — nudge the single lead back to work when it has been continuously
# idle (no open fleet_send driving it) past the quiet period. The fleet is one lead + architects +
# idle (no open bridge_send driving it) past the quiet period. The fleet is one lead + architects +
# workers, so a lead that stalls is a single point of failure; the ReplyPushLoop only nudges when a
# reply lands, and this timer catches the gap where nothing lands and the lead just sits idle.
#
@@ -93,20 +93,21 @@ bind:
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
# Math.max(15, intervalSeconds), so a lower value is silently raised, not
# rejected.
# workingSuspectAfterSeconds → age before a BUSY member is suspected of a stall (default 600).
# ENFORCED floor of 300: a lower value is silently raised.
# paneProbeIntervalSeconds → accepted and parsed, but NOT YET READ by anything. Setting it changes
# nothing right now. It exists so a later build can start honouring it without
# another config-shape change.
# notifications.mode → "webhook" flips what fleet_list REPORTS (healthCoverage: "full" instead
# of "detection-only") — it does NOT make fleetd send any webhook call; no
# workingSuspectAfterSeconds, paneProbeIntervalSeconds → accepted and parsed, but NOT YET READ by
# anything — the dormant monitor only consumes intervalSeconds today (CB-573
# shipped ahead of the evidence publishers these two knobs are for). Setting
# them changes nothing right now, and no minimum is enforced on either, because
# nothing reads them to enforce one. They exist so a later build can start
# honouring them without another config-shape change.
# notifications.mode → "webhook" flips what bridge_list REPORTS (healthCoverage: "full" instead
# of "detection-only") — it does NOT make bridged send any webhook call; no
# delivery mechanism is implemented yet. Any other value, or omitting the
# block, reports "detection-only".
# health:
# enabled: true
# intervalSeconds: 30 # floor 15
# workingSuspectAfterSeconds: 600 # floor 300 — how long BUSY with no activity means STALL_SUSPECTED
# paneProbeIntervalSeconds: 60 # parsed, but nothing reads it yet — changing it changes nothing
# intervalSeconds: 30
# workingSuspectAfterSeconds: 600
# paneProbeIntervalSeconds: 60
# notifications:
# mode: disabled
@@ -114,9 +115,6 @@ bind:
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
# Optional socket for member panes. Omit this to use herdrSocket for both leads and members.
# memberHerdrSocket: /Users/member/.config/herdr/herdr.sock
# How member sessions are spawned. Define one or more named profiles (backends) under
# `profiles`; each key is the profile name (also the ccs profile). A profile says only WHICH
# BACKEND — model, CLI adapter, credentials, cost. It says nothing about what a member spawned on
@@ -126,34 +124,8 @@ herdrSocket: ~/.config/herdr/herdr.sock
# Shared knobs (placement/workspace) can be repeated per profile; they usually match.
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
# Use `pane` for the legacy behaviour (split the focused tab).
# mcpUrl → fleetd mounts the bridge MCP (--mcp-config, inline) + reply charter
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
# (--append-system-prompt) as launch flags; nothing is written to the profile.
# ideMcpUrl → opt-in (CB-634), default off. When set, fleetd mounts the IDE Index MCP as a
# second inline server named `intellij`, and adds an IDE charter that pins every
# ide_* call to the member's own worktree. A URL, not a boolean — host and port
# are host-specific. Set it only on a host where the IDE actually runs.
# ideProjectDir → repo-relative module dir the IDE opens and the overlay pins (CB-634). Only read
# when ideMcpUrl is set. This repo's Maven pom lives in `fleetd/`, not at the
# worktree root, so opening the root imports no module and ide_* resolves nothing;
# set this to `fleetd`. Omit for a repo whose project is the worktree root.
# ideOpenCommand → host command that opens ideProjectDir in the IDE at spawn (CB-634 auto-open).
# Only read when ideMcpUrl is set. `{dir}` is replaced with the absolute module
# dir and the command runs through `/bin/sh -c`, so set env inline if needed —
# e.g. `env DISPLAY=:10.0 idea {dir}`. Best-effort: a failure is logged, never
# fails the spawn. Omit to open the member's module by hand. There is no close
# half yet — an opened module stays open until the operator closes it.
# autoCompactWindow → opt-in, default off. A bounded token window that forces a spawned member to
# compact its context instead of running on the backend's own default and dying
# mid-turn (losing its fleet_reply — the whole point of the turn — with it).
# Validated at config load to [100000, 1000000] — the band Claude Code's own
# --autocompact flag accepts.
# CROSS-BACKEND SEMANTICS DIFFER: on claude-code this is a launch-time
# `--autocompact <tokens>` flag — the member compacts AT this window. opencode
# has no equivalent flag (it only forces `compaction.auto: true`, unconditionally,
# already), so this is instead applied as the model's `limit.context` in the
# generated opencode.json — the member compacts WITHIN this window, not exactly
# at it — and only when this profile's `model:` is in `provider/model` form; if it
# isn't, fleetd logs a WARN naming the profile rather than silently doing nothing.
# tokenEnv → host env var holding the worker's auth token (value never stored in config);
# omit for a backend that needs no token (e.g. a local ollama).
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
@@ -163,8 +135,7 @@ herdrSocket: ~/.config/herdr/herdr.sock
# skills/MCP/hooks. Omit to leave the worker on the host default.
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
# worktree sees the same local config (CB-301-ext). Omit for the default set:
# [.env, .envrc]. (.claude/settings.local.json is NOT in the default — it
# pre-approves IDE/tool grants a member must not hold ambiently; CB-525/CB-634.)
# [.claude/settings.local.json, .env, .envrc].
#
# Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher
# mounts — the bridge, and nothing else. Replicating the primary's MCP config
@@ -172,7 +143,7 @@ herdrSocket: ~/.config/herdr/herdr.sock
# checkout, so its navigation returned paths OUTSIDE its own worktree: one
# worker made all 59 of its edits in the primary tree while compiling its
# worktree, and every build it ran was of code that did not contain them.
# fleetd neutralizes a provisioned worktree's .mcp.json for this reason;
# bridged neutralizes a provisioned worktree's .mcp.json for this reason;
# listing it here would copy the primary's back over that.
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
@@ -181,11 +152,11 @@ herdrSocket: ~/.config/herdr/herdr.sock
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
# exhaustedPattern → regex matched against a completion-fallback scrape (CB-578 stage A) to
# classify a turn that ended with no fleet_reply as the backend having
# classify a turn that ended with no bridge_reply as the backend having
# refused on a subscription usage limit, rather than a real answer. Opt-in —
# omit and this profile's completion fallback behaves exactly as before.
# Every backend words its refusal differently, so this is config, never a
# vendor string baked into fleetd itself.
# vendor string baked into bridged itself.
# DEFERRED: compiled once into a startup pattern map — editing it needs a
# daemon restart, same as this profile's model/baseUrl/argv.
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
@@ -200,15 +171,15 @@ herdrSocket: ~/.config/herdr/herdr.sock
# env → extra environment for this profile's workers, as a literal key/value map
# (CB-511). Use it to give workers a toolchain.
#
# A worker's environment does NOT come from your shell. fleetd hands herdr an
# A worker's environment does NOT come from your shell. bridged hands herdr an
# explicit env map and herdr merges it into ITS OWN process env — so before
# CB-511 a worker inherited whatever PATH the herdr server happened to be
# started with, which on a long-lived herdr can predate your toolchain entirely
# and leave workers unable to run `mvn` or `java` at all.
# fleetd now propagates ITS OWN PATH to every worker by default; set `env:`
# bridged now propagates ITS OWN PATH to every worker by default; set `env:`
# only to override that or add more (JAVA_HOME, …). Since the default is the
# daemon's PATH, make sure the daemon is started with a good one — see the PATH
# lines in deploy/dev.ltms.fleet.plist and deploy/fleetd.service.
# lines in deploy/dev.ltms.bridged.plist and deploy/bridged.service.
#
# Adapter-owned variables always win over `env:`: ANTHROPIC_BASE_URL and the
# rest of the ANTHROPIC_*/CLAUDE_* wiring are applied after it, so an `env:`
@@ -221,20 +192,20 @@ profiles:
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
model: coder
placement: tab
workspace: fleetd-workers
workspace: bridged-workers
# tabLabel: an optional per-profile override; the fleet template usually covers it
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: FLEETD_WORKER_TOKEN
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
# weight: relative selection weight for automatic placement (weighted, round-robin, and
# fixed's fallback walk). Absent defaults to 1.0. An explicit 0 or negative value means
# "never auto-select this profile" (CB-554) — it stays reachable via an explicit
# `fleet_spawn{profile:"gx10"}`, which bypasses placement entirely; only automatic
# `bridge_spawn{profile:"gx10"}`, which bypasses placement entirely; only automatic
# selection skips it.
weight: 0.5
# maxLoad: max live workers on this profile. Omit for unlimited. An explicit 0 (CB-585) caps
# the profile at zero live members — it is excluded from automatic placement and an explicit
# `fleet_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
# `bridge_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
# is named directly. Negative is refused at config load — there is no sane meaning for it.
maxLoad: 2
# subscription: true
@@ -253,7 +224,7 @@ profiles:
# subscription path no guard would vet the URL, so allowing both would be a way around the
# guard rather than a configuration.
#
# GOTCHA 1 — it is invisible to the startup secret check. `Fleetd.reportRequiredSecrets`
# GOTCHA 1 — it is invisible to the startup secret check. `Bridged.reportRequiredSecrets`
# skips subscription profiles on purpose (they need no token), so a boot log that reports
# every secret as fine says nothing about these profiles.
#
@@ -266,15 +237,11 @@ profiles:
# credentialId: shared-openai # opt-in: quarantine together with every other profile sharing this id (CB-578)
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".env", ".envrc"] # the default; never add .mcp.json or .claude/settings.local.json — see above
# ideMcpUrl: http://127.0.0.1:29170/index-mcp/streamable-http # opt-in (CB-634): IDE code intelligence, pinned to the worktree
# ideProjectDir: fleetd # CB-634: module dir the IDE opens + the overlay pins (this repo's pom is in fleetd/)
# ideOpenCommand: env DISPLAY=:10.0 idea {dir} # CB-634 auto-open: opens {dir} in the IDE at spawn; omit to open by hand
# autoCompactWindow: 250000 # opt-in: bound member context; claude-code compacts AT this, opencode within it (model limit.context)
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
gx11: # a second backend, so `placement: weighted` has a choice
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
placement: tab
workspace: fleetd-workers
workspace: bridged-workers
# tabLabel: an optional per-profile override; the fleet template usually covers it
mcpUrl: http://127.0.0.1:8765/mcp
argv: ["ccs", "gx11"]
@@ -292,15 +259,15 @@ profiles:
# The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a
# `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude.
#
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → fleet_send →
# structured fleet_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → bridge_send →
# structured bridge_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
# and need NO credentials — check `opencode models` for the current free list, since the names
# change. That also makes the worker off-subscription by construction.
# opencode-free:
# kind: opencode
# model: opencode/north-mini-code-free # `provider/model` selector, injected as `-m`
# placement: tab
# workspace: fleetd-workers
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
@@ -325,7 +292,7 @@ profiles:
# baseUrl: http://127.0.0.1:8000
# model: local-vllm/deepseek-v4-flash
# placement: tab
# workspace: fleetd-workers
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
@@ -363,8 +330,8 @@ placement: weighted
# quarantineCooldownSeconds: 1800
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
# upgraded fleetd keeps the old behaviour: the file is read once at boot and never again.
# enabled → turn the watch on. fleetd checks the file's modified time on a timer and
# upgraded bridged keeps the old behaviour: the file is read once at boot and never again.
# enabled → turn the watch on. bridged checks the file's modified time on a timer and
# reloads when it moves.
# intervalSeconds → how often to check (default 10). One `stat` per tick, so this is cheap.
#
@@ -375,7 +342,7 @@ placement: weighted
# / credentialId. Those are hot because the placement policy (and, for credentialId,
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
# not by itself enough to make a key hot.
# EXCEPT `fleet.leaders`: Fleetd.main reads it once at startup to build the lead tab
# EXCEPT `fleet.leaders`: Bridged.main reads it once at startup to build the lead tab
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
# with nothing in the deferred list — but has NO effect until you restart. Treat it
@@ -479,8 +446,8 @@ fleet:
# workspace: leads # where a launched lead's tab is created (default "leads").
# # MUST NOT be a member workspace — those are excluded from the
# # scan, so a lead placed in one is never found again.
# cwd: /path/to/repo # the launched lead's working directory (default: fleetd's own)
# kind: claude # descriptive; reported by fleet_whoami
# cwd: /path/to/repo # the launched lead's working directory (default: bridged's own)
# kind: claude # descriptive; reported by bridge_whoami
# gpt-sol-5.6:
# tab: "lead: gpt-sol-5.6"
# kind: opencode
@@ -508,7 +475,7 @@ guard:
# Member credential policy (CB-596, gitea issue #82). A herdr pane runs a LOGIN shell, and that
# shell re-sources the operator's own secret store — so a spawned member inherits every credential
# the operator's shell holds, not just the ones fleetd means to give it. Measured on this host:
# the operator's shell holds, not just the ones bridged means to give it. Measured on this host:
# 31 credential names, all set, with only ONE (GITEA_ACCESS_TOKEN) blocked before this — and that
# block was a single name hardcoded in HerdrPeerLauncher.java, not driven by this file. This block
# replaces that hardcoded shadow with a config-driven list of names.
@@ -538,51 +505,20 @@ guard:
# through either: the daemon logs a WARN naming any credential-shaped env var it finds on neither
# list (never its value), so a secret added to the store later does not go unnoticed forever.
#
# policy → "deny-by-default" (the default; also accepted spelled "deny-list") overlays each
# known-but-not-allowed name BEFORE the pane's login shell runs — real protection only
# where that shell does not re-export the name (see ROUND-2 CORRECTION above). An
# unrecognized value refuses to start, naming it.
# policy → "allow-list" (CB-633) moves the control to a per-spawn ZDOTDIR directory the daemon
# generates and passes through tab.create's env map. Each generated startup file sources
# its ~/ counterpart FIRST and then runs the scrub, so the scrub happens after the
# operator's whole chain and no sourced file can undo it.
# The scrub is sourced from BOTH the generated .zshrc and the generated .zlogin, because
# herdr does not open the same kind of shell everywhere: macOS panes run a LOGIN zsh (so
# .zlogin runs), Linux panes run a plain interactive zsh (so .zlogin never runs at all).
# A scrub in .zlogin alone would be a control that silently does nothing on Linux.
# The allow-list is DERIVED, never typed:
# every profile's tokenEnv/gitTokenEnv/gitHostEnv values and env-map keys, plus an
# infrastructure set (PATH HOME SHELL TERM LANG LC_* TMPDIR USER LOGNAME PWD SHLVL EDITOR
# PAGER JAVA_HOME XDG_* ZDOTDIR), plus whatever keys this spawn's own env overlay carries.
# Adding a profile can therefore only widen the list, never break another spawn's scrub.
# Under this policy `known`/`allow` below become REPORTING ONLY — they feed the gap WARN,
# they are no longer a control. If the member's login shell is NOT zsh, the daemon logs a
# loud WARN saying protection is off and falls back to deny-by-default's overlay.
# Each pane writes a scrub-report.txt naming how many variables it kept of how many it
# saw; the daemon logs that "allowed N of M" line when the pane stops. If the report is
# MISSING the daemon logs a WARN instead — the scrub cannot then be confirmed to have
# run, and a silently-dead control is exactly what this policy exists to prevent.
# allow → credential names a member legitimately needs. Under deny-by-default, left OUT of the
# pane's env overlay entirely, so the value the pane's own (login) shell exports passes
# through untouched. Under allow-list: reporting only.
# known → every credential name the operator's store is known to export. Under deny-by-default,
# every name here NOT also in `allow` is overlaid with a non-secret sentinel value before
# the pane's login shell runs — real protection only for names that shell does not itself
# re-export (see the ROUND-2 CORRECTION note above). Under allow-list: reporting only.
# sshAuthSock → whether SSH_AUTH_SOCK may pass through under allow-list ("allow") or must be
# blanked like any other non-derived name ("block", the default). This is a decision you
# have to make explicitly: SSH_AUTH_SOCK is a handle to YOUR ssh-agent, and a member
# holding it can sign with your keys — it sits in no secret file and looks like no
# credential, which is why it slipped past three earlier tickets (gitea #110). Blocking
# it breaks git over SSH inside members (push/fetch authenticate as you); use HTTPS
# remotes or scoped deploy keys instead of allowing it lightly.
# policy → only "deny-by-default" exists today (an operator-authored deny-list was deliberately
# rejected — see above). An unrecognized value refuses to start, naming it.
# allow → credential names a member legitimately needs. Left OUT of the pane's env overlay
# entirely, so the value the pane's own (login) shell exports passes through untouched.
# known → every credential name the operator's store is known to export. Every name here NOT
# also in `allow` is overlaid with a non-secret sentinel value before the pane's login
# shell runs — real protection only for names the login shell does not itself re-export
# (see the ROUND-2 CORRECTION note above for the ones it does).
#
# HOT-RELOADABLE the same way `fleet:` is (CB-559): read fresh on every spawn, so editing this list
# and reloading config (or restarting) changes what the NEXT spawn inherits; already-running members
# are unaffected either way.
# memberCredentials:
# policy: deny-by-default # or "deny-list", or "allow-list" (CB-633) — see above
# sshAuthSock: block # allow-list only; see the sshAuthSock note above
# policy: deny-by-default
# allow:
# - AI_GATEWAY_TOKEN # named in a profile's tokenEnv (local/gx) — a member reaching the
# # gateway is by design, not a leak
@@ -657,36 +593,14 @@ guard:
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
# uriEnv → CB-151: name of a host env var holding the AMQP URI, preferred over `uri` (wins
# whenever set). The URI carries `user:pass@` inline, so naming a variable keeps the
# password out of fleetd.yaml — same pattern as auth.tokenEnv/Profile.tokenEnv. A
# uriEnv that resolves to an unset or blank variable is treated as NOT configured and
# the daemon falls back to the in-memory inbox, warning loudly.
# prefetch → CB-527: consumer basicQos, capping how many unacked messages the inbox holds
# in-heap per owned target (the rest sits on the broker's durable queue instead of
# growing the JVM heap). Default 32 when omitted.
# broker:
# uriEnv: LAVINMQ_URI
# uri: amqp://guest:guest@127.0.0.1:5672
# prefetch: 32
# Shared cross-host LEADER coordination broker. OMIT this block to leave lead-to-lead messaging
# off entirely (config-only in this ticket — nothing here wires it into a live LeadMailbox yet).
# This is a SEPARATE AMQP vhost from `broker:` above: member/worker inboxes always stay on the
# per-fleet `broker:` vhost, and this vhost carries only leader-to-leader traffic, so two fleets
# whose members must never see each other can still share one coordination vhost for their leads.
# uriEnv → name of a host env var holding the coordination AMQP URI, same convention as
# broker.uriEnv (keeps the credential out of fleetd.yaml). Wins over `uri` when set.
# selfId → this daemon's own lead coord-id — the name its mailbox is owned under
# (lead.<selfId>.inbox), e.g. "mac-opus" or "fleet01-lead". Must be globally unique
# across every daemon sharing this vhost.
# prefetch → consumer basicQos, capping how many unacked messages the mailbox holds in-heap.
# Default 32 when omitted.
# coordinator:
# uriEnv: LEAD_COORD_URI
# selfId: mac-opus
# prefetch: 32
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open fleet_send,
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
# pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop
# stops as soon as the primary's inbox is empty.
@@ -697,11 +611,11 @@ guard:
#
# REQUIRED (CB-522) if the primary itself runs inside a herdr pane. Caller
# identity resolves a loopback PID to its herdr pane, and PaneLocator scans
# EVERY pane — not just fleetd-spawned ones — so such a primary is otherwise
# EVERY pane — not just bridged-spawned ones — so such a primary is otherwise
# classified as a WORKER and refused SPAWN/SEND/STOP. That failure is
# self-locking: the learned terminal is populated by the very orchestration
# calls being refused, so only this pinned value can break the cycle. Read the
# id off fleet_whoami (it reports the current terminal even while
# id off bridge_whoami (it reports the current terminal even while
# misclassified) and re-pin whenever the primary moves panes.
# pushReminders → max nudges before giving up (default 5)
# pushBackoffMs → delay between nudges in ms (default 15000)
@@ -5,10 +5,10 @@
## Why this exists
Stage 2 gave a worker→primary reply a **durable place to wait** when no `fleet_send` is
Stage 2 gave a worker→primary reply a **durable place to wait** when no `bridge_send` is
open: it lands in `agent.<target>.inbox` on the broker and survives a daemon bounce. But
delivery is still **pull** — the primary only sees the reply if it happens to call
`fleet_poll(target)` / `GET /sessions/{id}/replies`. A reply can sit indefinitely while
`bridge_poll(target)` / `GET /sessions/{id}/replies`. A reply can sit indefinitely while
the primary works on something else.
This layer makes delivery **active**: the bridge *pushes* a nudge to the primary the moment
@@ -26,11 +26,11 @@ pointed at the primary's pane instead.
```mermaid
flowchart LR
W["worker"] -->|"fleet_reply (no open send)"| MS["MessageService.reply"]
W["worker"] -->|"bridge_reply (no open send)"| MS["MessageService.reply"]
MS -->|"inbox.publish"| INBOX[("agent.&lt;target&gt;.inbox<br/>(durable, LavinMQ)")]
MS -->|"notify"| LOOP["ReplyPushLoop"]
LOOP -->|"status-gated inject"| PANE["primary's herdr pane"]
PANE -->|"primary drains"| DRAIN["fleet_poll(target)<br/>= peek + ack"]
PANE -->|"primary drains"| DRAIN["bridge_poll(target)<br/>= peek + ack"]
DRAIN -->|"inbox now empty"| LOOP
LOOP -.->|"still non-empty →<br/>re-inject on backoff"| PANE
classDef store fill:#2c5282,stroke:#1a365d,color:#ffffff;
@@ -51,13 +51,13 @@ the caller runs in a herdr pane on this host. Today it's discarded for the prima
(`presence.markPresent` is a no-op on it).
**Plan:** a single-slot `PrimaryRegistry` (thread-safe) holding the primary's `terminal_id`.
Populate it from the **orchestration-side** MCP tools — `fleet_send`, `fleet_spawn`,
`fleet_poll`, `fleet_list`, `fleet_status`, `fleet_profiles` — capturing
Populate it from the **orchestration-side** MCP tools — `bridge_send`, `bridge_spawn`,
`bridge_poll`, `bridge_list`, `bridge_status`, `bridge_profiles` — capturing
`callerTerminal(exchange)` when it is (a) non-null and (b) **not** a registered worker
session in `SessionManager`. That caller is, by construction, the primary. Worker-side tools
(`fleet_reply`, `fleet_ask`) never set it.
(`bridge_reply`, `bridge_ask`) never set it.
- **Config override / pin:** a `primary: { terminal: "<id>" }` block in `FleetConfig`
- **Config override / pin:** a `primary: { terminal: "<id>" }` block in `BridgedConfig`
(nested record, same shape as `Broker`). Lets an operator pin it, or supply it when
derivation can't (see degrade case).
- **Degrade:** if the primary is off-host or in a non-herdr terminal, `terminalForPid`
@@ -71,13 +71,13 @@ A `ReplyPushLoop` component, notified at the single no-waiter call site
(`MessageService.reply` → the `inbox.publish` branch, `MessageService.java:192`).
- **Inject a nudge, not the payload.** The injected turn tells the primary *to drain*
(e.g. "Worker `<target>` returned a reply — run `fleet_poll(target=<target>)` to collect
(e.g. "Worker `<target>` returned a reply — run `bridge_poll(target=<target>)` to collect
it"), it does **not** carry the reply text. Rationale: replies can be large/multiline and
terminal injection would mangle them; the drain response is the clean transport. Keeps the
push idempotent — re-nudging is harmless.
- **Ack = drain.** The primary draining (`drainReplies` = peek + ack) is the acknowledgement.
The loop's **stop condition is `inbox.peek(target).isEmpty()`** — the reply is gone from the
inbox because it was acked. No new `fleet_ack` tool needed for v1 (see Increment 3).
inbox because it was acked. No new `bridge_ack` tool needed for v1 (see Increment 3).
- **Status-gated injection (mechanism (b), chosen).** A dedicated lightweight scheduled loop,
**not** the worker `Injector`. It injects via `AgentControl.send(primaryTerminal, nudge)`
(the same herdr `agent.send` = `pane send-text` + submit that delivers to workers) only when
@@ -93,10 +93,10 @@ A `ReplyPushLoop` component, notified at the single no-waiter call site
the reply remains in the durable inbox and the next natural poll (or a later worker reply's
nudge) still surfaces it. Bounded so the bridge never spams the primary.
### Increment 3 — optional per-`msgId` `fleet_ack` tool (deferred)
### Increment 3 — optional per-`msgId` `bridge_ack` tool (deferred)
Drain-as-ack is coarse: it clears *all* pending replies for a target at once. If finer
control is ever needed (ack one reply, leave others held), add a `fleet_ack(msgId)` tool
control is ever needed (ack one reply, leave others held), add a `bridge_ack(msgId)` tool
mapping to `inbox.ack(target, msgId)` — the port already supports per-`msgId` ack. Not built
in v1; the stop-on-empty loop is sufficient.
@@ -106,7 +106,7 @@ in v1; the stop-on-empty loop is sufficient.
resolved terminal is non-null **and not a registered worker session**, seen on an
orchestration-side tool. This never mislabels a worker (workers are in `SessionManager`)
and needs no new env var or argument (identity stays connection-derived, per the existing
`FleetMcp` invariant).
`BridgeMcp` invariant).
2. **Readiness-gate mismatch → dedicated loop.** The existing `Injector` gates delivery on
`ready.test(target)` = `WorkerPresence` (the *worker's* MCP connected). The primary is not
@@ -132,7 +132,7 @@ boundary**. The bridge is signalling the primary that it has mail — not drivin
non-null terminal AND not a registered session" predicate; the loop's stop-on-empty and
bounded-reminder logic with an injected clock + a fake injector (no real herdr).
- **Live dogfood (primary-side):** with the daemon on the broker jar + a real worker,
delegate a task, let the worker reply after the `fleet_send` window closes, and observe the
delegate a task, let the worker reply after the `bridge_send` window closes, and observe the
bridge inject a drain nudge into *this* primary pane; confirm draining stops the reminders;
confirm an unreachable primary (registry empty) degrades to pull with no loss.
+7 -10
View File
@@ -5,17 +5,17 @@
<modelVersion>4.0.0</modelVersion>
<groupId>dev.ltms</groupId>
<artifactId>fleetd</artifactId>
<artifactId>bridged</artifactId>
<version>1.0.0</version>
<packaging>jar</packaging>
<name>fleetd</name>
<description>fleet message server: sole gateway between a lead session, its members, and herdr</description>
<name>bridged</name>
<description>claude-bridge message server: sole gateway between primary/worker Claude sessions and herdr</description>
<properties>
<maven.compiler.release>25</maven.compiler.release>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<mainClass>dev.ltms.fleet.Fleetd</mainClass>
<mainClass>dev.ltms.bridged.Bridged</mainClass>
<jackson.version>2.19.0</jackson.version>
<javalin.version>6.7.0</javalin.version>
@@ -106,7 +106,7 @@
</dependency>
<!-- MCP server: the SERVER face. Streamable-HTTP servlet mounted on Javalin's Jetty at
/mcp, exposing fleet_send/fleet_reply/fleet_status as thin adapters over REST. -->
/mcp, exposing bridge_send/bridge_reply/bridge_status as thin adapters over REST. -->
<dependency>
<groupId>io.modelcontextprotocol.sdk</groupId>
<artifactId>mcp</artifactId>
@@ -161,10 +161,7 @@
</dependencies>
<build>
<!-- CB-634: the cutover renamed the module dir (bridged/ -> fleetd/), the jar, and the
launchd plist together. The installed plist names fleetd/target/fleetd.jar and
KeepAlive is armed, so this name, the plist, and the wrapper must move as one. -->
<finalName>fleetd</finalName>
<finalName>bridged</finalName>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
@@ -206,7 +203,7 @@
</configuration>
</plugin>
<!-- Runnable fat jar: java -jar target/fleetd.jar -->
<!-- Runnable fat jar: java -jar target/bridged.jar -->
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-shade-plugin</artifactId>
@@ -1,117 +1,88 @@
package dev.ltms.fleet;
package dev.ltms.bridged;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.ConfigWatcher;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.herdr.LeadTabScanner;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.inject.StatusPoller;
import dev.ltms.fleet.inject.TurnListener;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.health.FleetHealthMonitor;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.mcp.LsofPeerPidLookup;
import dev.ltms.fleet.mcp.LsofProcessCwdLookup;
import dev.ltms.fleet.msg.AmqpReplyInbox;
import dev.ltms.fleet.msg.InMemoryReplyInbox;
import dev.ltms.fleet.msg.LeadChannel;
import dev.ltms.fleet.msg.LeadCoordLoop;
import dev.ltms.fleet.msg.LeadMailbox;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
import dev.ltms.fleet.msg.ReplyPushLoop;
import dev.ltms.fleet.rest.FleetApp;
import dev.ltms.fleet.session.GitWorktrees;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.session.SessionReaper;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.member.CompositePeerLauncher;
import dev.ltms.fleet.member.HerdrPeerLauncher;
import dev.ltms.fleet.member.OpenCodeLauncher;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.config.ConfigRef;
import dev.ltms.bridged.config.ConfigWatcher;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.LeadTabScanner;
import dev.ltms.bridged.lead.LeadLauncher;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.ExhaustedPatternLookup;
import dev.ltms.bridged.inject.ExhaustionSink;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.auth.MemberRegistry;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.health.FleetHealthMonitor;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
import dev.ltms.bridged.msg.AmqpReplyInbox;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.msg.LeadHeartbeatLoop;
import dev.ltms.bridged.msg.ReplyPushLoop;
import dev.ltms.bridged.rest.BridgedApp;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.session.SessionReaper;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.member.CompositePeerLauncher;
import dev.ltms.bridged.member.HerdrPeerLauncher;
import dev.ltms.bridged.member.OpenCodeLauncher;
import dev.ltms.bridged.placement.BackendQuarantine;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.regex.Pattern;
import java.util.stream.Collectors;
/**
* {@code fleetd} entry point. Wires the real herdr socket client to the REST app and
* {@code bridged} entry point. Wires the real herdr socket client to the REST app and
* starts listening. Before anything else it asserts its own environment is clean —
* {@code fleetd} is not a Claude process and must never carry a base_url.
* {@code bridged} is not a Claude process and must never carry a base_url.
*/
public final class Fleetd {
public final class Bridged {
private static final Logger log = LoggerFactory.getLogger(Fleetd.class);
private static final Logger log = LoggerFactory.getLogger(Bridged.class);
/** CB-504: how long to wait at startup for herdr's socket before serving degraded. */
private static final long HERDR_WAIT_SECONDS = 30;
/**
* CB-637: how often the lead coordination loop looks for peer messages. A few seconds — slow
* enough that an idle fleet is not polling a broker in a tight loop, fast enough that a peer
* lead's message is not left sitting once the local lead reaches a turn boundary. The mailbox
* pushes into the loop's held set on its own consumer thread, so this interval bounds only the
* pane delivery, never the receive.
*/
private static final long LEAD_COORD_INTERVAL_MS = 3_000L;
private static final long HERDR_WAIT_POLL_MILLIS = 500;
/**
* CB-632/CB-634: prefer {@code fleetd.yaml} in {@code dir}; fall back to the legacy
* {@code bridged.yaml} when the new name is not there. The product renamed to {@code fleetd},
* but a deployment whose local config file is still {@code bridged.yaml} keeps working until
* that file is renamed.
*/
static Path chooseDefaultConfigFile(Path dir) {
Path fleetd = dir.resolve("fleetd.yaml");
if (Files.exists(fleetd)) {
return fleetd;
}
return dir.resolve("bridged.yaml");
}
static void main(String[] args) {
Path configPath = args.length > 0 ? Path.of(args[0]) : chooseDefaultConfigFile(Path.of(""));
// The config file was renamed bridged.yaml -> fleetd.yaml. Name the file we actually
// loaded, whichever of the two names it carries.
log.info("Using configuration file {}", configPath);
FleetConfig cfg = FleetConfig.load(configPath);
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
BridgedConfig cfg = BridgedConfig.load(configPath);
// CB-594: report which secret env vars the config actually needs, by name, before anything
// else can fail on a silently-empty one. A daemon started without a login shell (launchd)
// boots fine either way — this is the only thing that says so out loud.
@@ -127,7 +98,7 @@ public final class Fleetd {
// contract; adding a reader here does not make a key reloadable by itself.
ConfigRef config = new ConfigRef(configPath, cfg);
// The primary/host env that launched fleetd must not be tainted.
// The primary/host env that launched bridged must not be tainted.
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
guard.assertPrimaryClean(System.getenv());
@@ -150,17 +121,14 @@ public final class Fleetd {
: UnixSocketHerdrClient.defaultSocketPath();
UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect(socket, new com.fasterxml.jackson.databind.ObjectMapper());
UnixSocketHerdrClient memberHerdr = cfg.memberHerdrSocket() != null && !cfg.memberHerdrSocket().isBlank()
? UnixSocketHerdrClient.connect(Path.of(cfg.memberHerdrSocket()), new com.fasterxml.jackson.databind.ObjectMapper())
: herdr;
AtomicReference<Supplier<Map<String, String>>> leadsRef = new AtomicReference<>(Map::of);
HerdrRouter router = new HerdrRouter(herdr, memberHerdr,
target -> leadsRef.get().get().containsKey(target));
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
Map<String, FleetConfig.Profile> claudeProfiles = new LinkedHashMap<>();
Map<String, FleetConfig.Profile> opencodeProfiles = new LinkedHashMap<>();
Map<String, BridgedConfig.Profile> claudeProfiles = new LinkedHashMap<>();
Map<String, BridgedConfig.Profile> opencodeProfiles = new LinkedHashMap<>();
cfg.profiles().forEach((name, w) -> {
if (w.isOpenCode()) {
opencodeProfiles.put(name, w);
@@ -173,14 +141,14 @@ public final class Fleetd {
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(new ClaudeCodeLauncher(router.memberAgents(), router.memberSpaces(), guard,
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(),
() -> config.get().memberCredentials()));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(router.memberAgents(), router.memberSpaces(),
adapters.add(new OpenCodeLauncher(agents, spaces,
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(),
@@ -189,7 +157,7 @@ public final class Fleetd {
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
// The cooldown is deferred (see FleetConfig#quarantineCooldownSeconds): it is read once
// The cooldown is deferred (see BridgedConfig#quarantineCooldownSeconds): it is read once
// here, at startup, and a config reload only changes it for a daemon restart.
BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
@@ -199,7 +167,7 @@ public final class Fleetd {
config,
profileName -> liveCountRef.get().apply(profileName),
quarantine);
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
@@ -258,39 +226,34 @@ public final class Fleetd {
final Supplier<Map<String, String>> leads;
var leaders = cfg.fleet().leaders();
if (!leaders.isEmpty()) {
Set<String> memberSpaces = cfg.profiles().values().stream()
.map(BridgedConfig.Profile::workspace)
.filter(Objects::nonNull)
.collect(Collectors.toSet());
Map<String, String> tabToName = new LinkedHashMap<>();
leaders.forEach((name, leader) -> {
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
tabToName.put(leader.tab(), name);
}
});
// The lead and the members now share ONE workspace (the operator asked for a single
// "session" with many tabs), so no workspace can be excluded — the lead lives in the
// members' space by design. A lead is told from a member by its exact tab label alone:
// a lead carries its configured `tab`, a member its `worker: {profile} #{n}` template,
// and the two never collide. (The scanner still supports an exclusion set for a split
// layout; the fleet's policy here is simply not to use one.)
//
// One shared rescan cadence: still taken from the first entry, as before — it is an
// operational cadence, not identity, so there is no correctness reason to give every
// lead its own scanner.
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
// This must use the lead daemon: scanning member tabs would demote the lead to a worker.
leads = new LeadTabScanner(herdr, tabToName, Set.of(),
leads = new LeadTabScanner(herdr, tabToName, memberSpaces,
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), System::nanoTime);
log.info("lead scan: tabs {} host a lead (rescan every {}s, shared fleet space)",
tabToName.keySet(), scanIntervalSeconds);
log.info("lead scan: tabs {} host a lead (rescan every {}s, member spaces {} excluded)",
tabToName.keySet(), scanIntervalSeconds, memberSpaces);
} else {
leads = () -> leadTerminals;
}
leadsRef.set(leads);
// CB-558: start any declared lead that is not already running. After the scanner is built,
// because both read the same tab labels and the ordering makes that dependency visible; and
// only when herdr answered, because the launcher's whole safety property is that it can
// count live leads first — it must never guess and risk a second orchestrator.
if (herdrUp && !leaders.isEmpty()) {
int launched = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg).ensureLeads();
int launched = new LeadLauncher(agents, spaces, cfg).ensureLeads();
if (launched > 0) {
log.info("lead auto-launch: {} lead(s) started", launched);
}
@@ -346,7 +309,6 @@ public final class Fleetd {
+ "BACKEND_EXHAUSTED): {}", credentialId,
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
});
AgentControl agents = router.memberAgents();
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns, exhaustionSink);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
@@ -370,7 +332,7 @@ public final class Fleetd {
}
@Override
public void onDelivered(String target, dev.ltms.fleet.msg.TurnToken token) {
public void onDelivered(String target, dev.ltms.bridged.msg.TurnToken token) {
completion.onDelivered(target, token);
sessions.onDelivered(target, token);
}
@@ -387,22 +349,23 @@ public final class Fleetd {
sessions.onTurnFailed(target);
}
};
Predicate<String> deliverable = deliverableTo(presence, leads);
Injector injector = new Injector(router, turnListener, deliverable,
Injector injector = new Injector(agents, turnListener, deliverableTo(presence, leads),
presence::forget);
StatusPoller poller = new StatusPoller(router, injector, Injector.POLL_INTERVAL_MILLIS);
StatusPoller poller = new StatusPoller(agents, injector, Injector.POLL_INTERVAL_MILLIS);
poller.start();
// CB-307: reply inbox. A broker: block selects the AMQP-backed durable adapter; absent (or
// unusable), fleetd stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
// CB-307: reply inbox. A broker: block (with a uri) selects the AMQP-backed durable adapter;
// absent, bridged stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
// connection, so keep the reference to close it in the ordered shutdown hook.
final ReplyInbox replyInbox = selectReplyInbox(cfg.broker(), System.getenv(), AmqpReplyInbox::open);
// CB-637: this daemon's lead-to-lead mailbox on the SHARED coordination vhost — a separate
// broker from the reply inbox by design (see FleetConfig.Coordinator). Absent a coordinator:
// block this is null and every lead path below is simply not wired, which is exactly the
// behaviour before this ticket. It owns a broker connection, so keep the reference for the
// ordered shutdown hook.
final LeadMailbox leadMailbox = openLeadMailbox(cfg.coordinator(), System.getenv(), LeadMailbox::open);
final ReplyInbox replyInbox;
if (cfg.broker() != null && cfg.broker().isConfigured()) {
replyInbox = AmqpReplyInbox.open(cfg.broker().uri(), cfg.broker().prefetchOrDefault());
log.info("reply inbox: AMQP broker (durable) at {} (prefetch={})",
cfg.broker().uri(), cfg.broker().prefetchOrDefault());
} else {
replyInbox = new InMemoryReplyInbox();
log.info("reply inbox: in-memory (soft-state)");
}
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
// The pin also feeds CallerResolver below: a primary running inside a herdr pane would
// otherwise resolve as a worker and be refused every orchestration tool.
@@ -418,7 +381,7 @@ public final class Fleetd {
+ "has lost the delegation map. Its pushReminders/pushBackoffMs stay valid.");
}
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
// open fleet_send. Uses its own lightweight scheduled executor, separate from the injector.
// open bridge_send. Uses its own lightweight scheduled executor, separate from the injector.
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
var pushScheduler = Executors.newSingleThreadScheduledExecutor(r ->
@@ -426,8 +389,8 @@ public final class Fleetd {
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
// CB-512: the push loop takes it too, so nudge outcomes (delivered|exhausted) are counted.
Metrics metrics = FleetMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, router.leadAgents(), replyInbox,
Metrics metrics = BridgedMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, agents, replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
@@ -437,7 +400,7 @@ public final class Fleetd {
Thread.ofVirtual().name("bridge-heartbeat-").unstarted(r));
if (cfg.leadHeartbeat() != null) {
var hb = cfg.leadHeartbeat();
heartbeat = new LeadHeartbeatLoop(primaryRegistry, router.leadAgents(), replyInbox, sessions::roster,
heartbeat = new LeadHeartbeatLoop(primaryRegistry, agents, replyInbox, sessions::roster,
pushLoop, heartbeatScheduler, System::nanoTime,
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
metrics);
@@ -446,7 +409,7 @@ public final class Fleetd {
heartbeat = null;
heartbeatScheduler.shutdownNow();
}
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
@@ -457,8 +420,7 @@ public final class Fleetd {
// CB-580: a member found GONE/NEVER_READY must fail whatever ticket is waiting on it,
// through the same idempotent target-wide operation CB-516 already uses on release.
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
System::nanoTime, cfg.health().intervalOrDefault(),
cfg.health().workingSuspectAfterOrDefault(), messages::abandon);
System::nanoTime, cfg.health().intervalOrDefault(), messages::abandon);
String coverage = FleetHealthMonitor.coverage(true,
cfg.health().notifications() != null && cfg.health().notifications().configured());
if ("detection-only".equals(coverage)) {
@@ -496,13 +458,10 @@ public final class Fleetd {
primaryRegistry.forgetDelegation(detail.terminalId()); // CB-532: don't leak the lead binding
});
// MCP server face (CB-105): fleet_send/fleet_reply/fleet_status, mounted at /mcp.
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
// CB-185: a caller's pane can live on either daemon (a lead's on the lead daemon, a
// member's on the member daemon) — search both, lead first. Collapses to one scan when
// memberHerdrSocket is unset (herdr == memberHerdr).
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr, memberHerdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
new PaneLocator(herdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
@@ -511,7 +470,7 @@ public final class Fleetd {
String token = System.getenv(cfg.auth().tokenEnv());
if (token == null || token.isBlank()) {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting fleetd");
+ " is unset or empty — export it before starting bridged");
}
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
@@ -521,41 +480,21 @@ public final class Fleetd {
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics, new FleetMcp.CapacitySource(profile -> liveCountRef.get().apply(profile),
BridgeMcp mcp = new BridgeMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics, new BridgeMcp.CapacitySource(profile -> liveCountRef.get().apply(profile),
profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.maxLoad();
}, () -> config.get().profiles().keySet(), System::nanoTime),
new FleetMcp.HealthCoverageSource(() -> {
new BridgeMcp.HealthCoverageSource(() -> {
var health = config.get().health();
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
health != null && health.notifications() != null && health.notifications().configured());
}),
new FleetMcp.QuarantineSource(profile -> {
new BridgeMcp.QuarantineSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, quarantine),
leadMailbox);
// CB-637: the receive half. Only constructed when a lead mailbox actually opened — with no
// coordinator (or an unreachable one) there is nothing to deliver, so no scheduler is
// created and no thread runs. It reads the SAME live lead supplier the injector's
// deliverability gate does, so a lead found by the tab scan after startup is reachable
// without a restart.
final LeadCoordLoop leadCoordLoop;
final ScheduledExecutorService leadCoordSchedulerRef;
if (leadMailbox != null) {
var leadCoordScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-leadcoord-").unstarted(r));
leadCoordLoop = new LeadCoordLoop(leadMailbox, router.leadAgents(), leads, leadCoordScheduler,
LEAD_COORD_INTERVAL_MS);
leadCoordLoop.start();
leadCoordSchedulerRef = leadCoordScheduler;
} else {
leadCoordLoop = null;
leadCoordSchedulerRef = null;
}
}, quarantine));
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed, so an
// upgraded daemon behaves exactly as before — the file is read once at boot and never again.
@@ -576,8 +515,6 @@ public final class Fleetd {
messages.close();
pushLoop.close();
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
if (leadCoordLoop != null) leadCoordLoop.close(); // CB-637: stop delivering peer-lead messages
if (leadCoordSchedulerRef != null) leadCoordSchedulerRef.shutdownNow();
if (healthMonitor != null) healthMonitor.stop();
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
mcp.close();
@@ -590,24 +527,13 @@ public final class Fleetd {
log.debug("reply inbox close: {}", e.toString());
}
}
// CB-637: the coordination connection goes with it — after the loop that reads it has
// stopped, so no tick can be mid-ack against a closed channel.
if (leadMailbox != null) {
try {
leadMailbox.close();
} catch (Exception e) {
log.debug("lead mailbox close: {}", e.toString());
}
}
router.close();
herdr.close();
}));
// CB-185: give FleetApp both daemons — /healthz must require both to answer and
// GET /sessions must merge across both, or a down/unpolled member daemon is invisible.
Javalin app = new FleetApp(herdr, memberHerdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics, deliverable).build();
Javalin app = new BridgedApp(herdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics).build();
app.start(cfg.bind().host(), cfg.bind().port());
log.info("fleetd listening on {}:{}, herdr socket {}",
log.info("bridged listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
}
@@ -620,7 +546,7 @@ public final class Fleetd {
* hazard is a property of spawning. A lead is never spawned: the operator started it and named it
* (or labelled its tab) only once it was up, so there is no boot window to guard.
*
* <p>A lead is also never enrolled in {@link MemberPresence} — {@code FleetMcp} marks presence
* <p>A lead is also never enrolled in {@link MemberPresence} — {@code BridgeMcp} marks presence
* for every spawned member (worker and architect), deliberately, since that map doubles as the
* member roster's availability signal and a lead counted there would show up as an available
* member. So without the second disjunct a lead is permanently un-deliverable: every
@@ -634,169 +560,22 @@ public final class Fleetd {
return target -> presence.isPresent(target) || leads.get().containsKey(target);
}
/** Injection seam for {@link #selectReplyInbox}: production binds {@link AmqpReplyInbox#open}. */
@FunctionalInterface
interface AmqpOpener {
ReplyInbox open(String uri, int prefetch);
}
/** Injection seam for {@link #openLeadMailbox}: production binds {@link LeadMailbox#open}. */
@FunctionalInterface
interface LeadMailboxOpener {
LeadMailbox open(String uri, String selfCoordId, int prefetch);
}
/**
* CB-637: open this daemon's lead-to-lead mailbox, or return {@code null} to leave the feature
* off. Package-private and env-injected for the same reason as {@link #selectReplyInbox}: the
* selection is then testable without a broker or a mutable process environment.
*
* <p>Every "off" path returns {@code null}, and each says why at the level it deserves:
*
* <ul>
* <li>no {@code coordinator:} block — silent. Lead coordination is opt-in; an operator who
* never configured it does not need to be told it is off on every boot.</li>
* <li>a block whose {@code uriEnv} does not resolve — INFO, the same "you moved to the secret
* store and the variable is not there" case {@code selectReplyInbox} warns about.</li>
* <li>a configured broker but no {@code selfId} — WARN. This one is a half-finished config: a
* mailbox is named after the coord-id that owns it, so with no id there is no queue to own
* and no {@code from} to send as. Loud, because the operator plainly intended the feature.</li>
* <li>the broker refuses at boot — WARN, and carry on. Mirrors {@code openAmqpOrFallback}: a
* coordination broker that is down must never take a whole fleet's daemon with it, and the
* fleet still works exactly as it did before this feature existed.</li>
* </ul>
*/
static LeadMailbox openLeadMailbox(FleetConfig.Coordinator coordinator, Map<String, String> env,
LeadMailboxOpener opener) {
if (coordinator == null) {
return null; // opt-in: nothing configured, nothing to say
}
String uri = coordinator.effectiveUri(env);
if (uri == null) {
log.info("lead coordination: OFF — coordinator{} has no usable broker uri",
coordinator.uriEnv() == null ? "" : ".uriEnv=" + coordinator.uriEnv());
return null;
}
if (coordinator.selfId() == null || coordinator.selfId().isBlank()) {
log.warn("coordinator.selfId is unset — lead coordination is OFF. A lead mailbox is the "
+ "queue named after the coord-id that owns it, so with no id there is nothing to "
+ "own and no sender identity to publish as. Set coordinator.selfId to a name that "
+ "is unique across every daemon sharing {} and restart fleetd.",
stripCredentials(uri));
return null;
}
try {
LeadMailbox mailbox = opener.open(uri, coordinator.selfId(), coordinator.prefetchOrDefault());
log.info("lead coordination: ON as coord-id {} (prefetch={})",
coordinator.selfId(), coordinator.prefetchOrDefault());
return mailbox;
} catch (IllegalStateException e) {
log.warn("cannot reach the AMQP coordination broker ({}) — lead-to-lead messaging is OFF "
+ "for this process lifetime. fleet_send{{coordId}} will report it as "
+ "not configured, and peer messages already queued stay on the broker until "
+ "a restart picks them up. Reason: {}",
stripCredentials(uri), reasonOf(e));
return null;
}
}
/**
* CB-151/152: pick the reply inbox. A usable broker — a literal {@code uri}, or a {@code
* uriEnv} whose variable resolves (both read from {@code env}) — selects the durable AMQP inbox.
* Everything else falls back to the in-memory inbox: no broker block, a blank {@code uri}, a
* {@code uriEnv} whose variable is unset or blank, or a broker unreachable at boot. The two
* lossy paths warn <em>loudly</em> — never silently — because what is lost is durable,
* cross-restart reply delivery. Package-private and env-injected so the selection is testable
* without a real broker or a mutable process environment.
*/
static ReplyInbox selectReplyInbox(FleetConfig.Broker broker, Map<String, String> env, AmqpOpener amqp) {
if (broker == null) {
log.info("reply inbox: in-memory (soft-state)");
return new InMemoryReplyInbox();
}
if (broker.hasUriEnv()) {
// uriEnv is authoritative whenever set (CB-151): the operator moved off clear text, so
// it must not quietly fall back onto a stale literal uri.
if (broker.uri() != null && !broker.uri().isBlank()) {
log.info("broker.uri is ignored because broker.uriEnv={} is set", broker.uriEnv());
}
String effectiveUri = broker.effectiveUri(env);
if (effectiveUri == null) {
log.warn("broker.uriEnv={} is unset or blank — durable AMQP reply inbox DISABLED. "
+ "Replies are soft-state and will not survive a restart. Set {} in the "
+ "daemon's environment (see scripts/redeploy-fleetd.sh) and restart to "
+ "use the durable broker inbox.",
broker.uriEnv(), broker.uriEnv());
log.info("reply inbox: in-memory (soft-state)");
return new InMemoryReplyInbox();
}
log.info("reply inbox: AMQP broker (durable) via env var {} (prefetch={})",
broker.uriEnv(), broker.prefetchOrDefault());
return openAmqpOrFallback(effectiveUri, broker.prefetchOrDefault(), "uriEnv " + broker.uriEnv(), amqp);
}
// No uriEnv: the literal uri path (existing behaviour).
if (broker.effectiveUri(env) == null) {
log.info("reply inbox: in-memory (soft-state)");
return new InMemoryReplyInbox();
}
log.info("reply inbox: AMQP broker (durable) (prefetch={})", broker.prefetchOrDefault());
return openAmqpOrFallback(broker.uri(), broker.prefetchOrDefault(), "uri", amqp);
}
/**
* Open the AMQP inbox, falling back to the in-memory inbox for this process lifetime if the
* broker cannot be reached at boot (CB-152). Not silent: the warning says durable delivery is
* off, replies are soft-state and will not survive a restart, plus the source that failed and
* the URI <em>with credentials stripped</em>. Never retries in the background — a broker that
* drops <em>after</em> startup already self-heals via the connection factory's automatic
* recovery; only the boot path is changed here.
*/
private static ReplyInbox openAmqpOrFallback(String effectiveUri, int prefetch, String source,
AmqpOpener amqp) {
try {
return amqp.open(effectiveUri, prefetch);
} catch (IllegalStateException e) {
log.warn("cannot reach AMQP broker ({}, {}) — falling back to the in-memory reply inbox "
+ "for this process lifetime. Durable, cross-restart reply delivery is OFF; "
+ "replies are soft-state and will not survive a restart. Reason: {}",
source, stripCredentials(effectiveUri), reasonOf(e));
log.info("reply inbox: in-memory (soft-state)");
return new InMemoryReplyInbox();
}
}
/** An AMQP URI carries {@code user:pass@} inline — show the host/port, never the credentials. */
static String stripCredentials(String uri) {
return uri == null ? null : uri.replaceAll("://[^@/]*@", "://");
}
/** The deepest cause's class and message — the outermost {@code IllegalStateException} echoes the URI (with password). */
private static String reasonOf(Throwable e) {
Throwable t = e;
while (t.getCause() != null && t.getCause() != t) {
t = t.getCause();
}
String msg = t.getMessage();
return t.getClass().getSimpleName() + (msg == null || msg.isBlank() ? "" : ": " + msg);
}
/**
* CB-594: which env vars the loaded config actually needs, and why — every non-{@code
* subscription} profile's {@code tokenEnv} (a subscription profile never reads one, see
* {@link FleetConfig.Profile#isSubscription()}), plus every profile's {@code gitTokenEnv}
* where set (opt-in), plus a configured {@code broker.uriEnv} (CB-151). Derived from the
* config, not hard-coded, so a new profile is covered for free. A var required by more than one
* profile is one entry naming every profile that needs it. Deliberately excludes {@code
* auth.tokenEnv}: that one is already enforced loudly, by a startup throw in {@code main()} —
* about 370 lines <em>below</em> this method's call site
* ({@link #reportRequiredSecrets(FleetConfig)}), not a few lines above it. That throw only
* {@link BridgedConfig.Profile#isSubscription()}), plus every profile's {@code gitTokenEnv}
* where set (opt-in). Derived from the config, not hard-coded, so a new profile is covered for
* free. A var required by more than one profile is one entry naming every profile that needs
* it. Deliberately excludes {@code auth.tokenEnv}: that one is already enforced loudly, by a
* startup throw in {@code main()} — about 370 lines <em>below</em> this method's call site
* ({@link #reportRequiredSecrets(BridgedConfig)}), not a few lines above it. That throw only
* fires when {@code auth.mode: token} is configured; under the default loopback-trust mode it
* never runs, and {@code auth.tokenEnv} is simply not required.
*
* <p>Package-private and pure (no I/O, no logging) so the derivation is unit-testable without
* capturing log output; {@link #reportRequiredSecrets(FleetConfig)} is the logging caller.
* capturing log output; {@link #reportRequiredSecrets(BridgedConfig)} is the logging caller.
*/
static Map<String, List<String>> requiredSecretEnvVars(FleetConfig cfg) {
static Map<String, List<String>> requiredSecretEnvVars(BridgedConfig cfg) {
Map<String, List<String>> requiredBy = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (!profile.isSubscription()) {
@@ -808,11 +587,6 @@ public final class Fleetd {
.add("profile '" + name + "' gitTokenEnv");
}
});
FleetConfig.Broker broker = cfg.broker();
if (broker != null && broker.hasUriEnv()) {
requiredBy.computeIfAbsent(broker.uriEnv(), _ -> new ArrayList<>())
.add("broker uriEnv");
}
return requiredBy;
}
@@ -825,7 +599,7 @@ public final class Fleetd {
* <p>A missing entry only warns — it must never refuse to start. A daemon that boots and says
* what is wrong is strictly more useful than one that will not boot at all.
*/
private static void reportRequiredSecrets(FleetConfig cfg) {
private static void reportRequiredSecrets(BridgedConfig cfg) {
Map<String, List<String>> requiredBy = requiredSecretEnvVars(cfg);
if (requiredBy.isEmpty()) {
log.info("startup secrets: no profile references a token env var — nothing to check");
@@ -839,8 +613,8 @@ public final class Fleetd {
} else {
log.warn("startup secret {}: MISSING ({}) — the daemon will start anyway, and this "
+ "failure stays invisible until a worker actually needs it. Fix "
+ "${SHARED_ENV}/tools/secrets.sh and restart fleetd from a LOGIN "
+ "shell (see scripts/redeploy-fleetd.sh).",
+ "${SHARED_ENV}/tools/secrets.sh and restart bridged from a LOGIN "
+ "shell (see scripts/redeploy-bridged.sh).",
varName, String.join(", ", sources));
}
});
@@ -848,7 +622,7 @@ public final class Fleetd {
/**
* CB-596: {@code known:} empty (block absent entirely, or present but empty) means {@link
* FleetConfig.MemberCredentials#blockedSet()} is empty too — every member pane inherits the
* BridgedConfig.MemberCredentials#blockedSet()} is empty too — every member pane inherits the
* operator's whole secret store, unblocked, exactly the defect this ticket fixes. Unlike a
* missing token ({@link #reportRequiredSecrets}), there is no name to point at: the point is
* that the block itself is missing. Warn once at startup and say what to add; never refuse to
@@ -858,21 +632,17 @@ public final class Fleetd {
* <p>Package-private so the test can capture the log directly, the same way {@link
* #requiredSecretEnvVars} is exposed for {@link #reportRequiredSecrets}'s own test.
*/
static void reportMemberCredentialsGap(FleetConfig cfg) {
FleetConfig.MemberCredentials creds = cfg.memberCredentials();
static void reportMemberCredentialsGap(BridgedConfig cfg) {
BridgedConfig.MemberCredentials creds = cfg.memberCredentials();
if (creds != null && !creds.known().isEmpty()) {
log.info("memberCredentials: policy={}, {} known name(s), {} allowed — blocking {} on "
+ "every spawn{}",
creds.policy(), creds.known().size(), creds.allow().size(), creds.blockedSet().size(),
creds.isAllowList()
? " (allow-list: known/allow are reporting only — the control is the derived ZDOTDIR scrub)"
: "");
log.info("memberCredentials: {} known name(s), {} allowed — blocking {} on every spawn",
creds.known().size(), creds.allow().size(), creds.blockedSet().size());
return;
}
log.warn("memberCredentials: absent or empty — the daemon will start anyway, and every "
+ "member pane inherits the operator's WHOLE secret store, unblocked (CB-592's "
+ "protection is lost). Add a memberCredentials: block (policy/allow/known) to "
+ "fleetd.yaml — see fleetd.example.yaml — and restart.");
+ "bridged.yaml — see bridged.example.yaml — and restart.");
}
/**
@@ -908,6 +678,6 @@ public final class Fleetd {
}
}
private Fleetd() {
private Bridged() {
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -1,9 +1,9 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
/**
* The authorization table (CB-505), stated once and enforced on both entry paths.
*
* <p>Most of these rules are already true de facto — {@code FleetMcp} derives a worker's identity
* <p>Most of these rules are already true de facto — {@code BridgeMcp} derives a worker's identity
* from the connection rather than reading it from an argument, so a worker has never been able to
* reply <em>as</em> another worker over MCP. What was missing is that the REST surface trusted the
* session id in the URL path, and neither surface checked role at all. This class makes the
@@ -1,7 +1,7 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.peer.MemberRole;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
@@ -12,7 +12,7 @@ import java.util.function.Supplier;
/**
* Resolves every caller to a {@link Principal}, for both entry paths into the core (CB-501).
*
* <p>There are two of them and they are not layered the way the docs suggest: {@code FleetMcp}
* <p>There are two of them and they are not layered the way the docs suggest: {@code BridgeMcp}
* calls the service layer directly and is mounted as a raw servlet (so it never passes through a
* Javalin filter), while the REST routes historically resolved no identity at all. Both now
* delegate here, so the authorization rules are stated once instead of drifting apart.
@@ -59,7 +59,7 @@ public final class CallerResolver {
* <p>Like {@link #leadTerminals}, a supplier rather than a fixed map, so a binding injected
* after startup — when the later spawn lifecycle establishes a live architect session, or an
* operator pins one — takes effect without a restart. Consulted per resolve; today's wiring
* in {@code Fleetd} reads a constant from config, which is the degenerate live case.
* in {@code Bridged} reads a constant from config, which is the degenerate live case.
*/
private final Supplier<Map<String, String>> architectTerminals;
private final Function<String, MemberRole> memberSlotRoles;
@@ -235,7 +235,7 @@ public final class CallerResolver {
// loopback-trust: same-host callers that are not workers are the primary. A non-loopback
// caller is anonymous even here — and startup refuses that combination anyway
// (FleetConfig.validateAuthExposure), so this is defence in depth, not the control.
// (BridgedConfig.validateAuthExposure), so this is defence in depth, not the control.
return isLoopback(remoteAddr) ? Principal.primary(c.pid()) : Principal.anonymous();
}
@@ -1,6 +1,6 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.bridged.peer.MemberRole;
/** Optional session lifecycle hook for live member-slot bindings. */
public interface MemberLifecycle {
@@ -1,7 +1,7 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.peer.MemberRole;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -59,7 +59,7 @@ public final class MemberRegistry implements MemberLifecycle {
private final Map<String, String> terminalToSlot = new HashMap<>();
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
public MemberRegistry(FleetConfig.Fleet fleet) {
public MemberRegistry(BridgedConfig.Fleet fleet) {
Map<String, Entry> flat = new LinkedHashMap<>();
if (fleet != null) {
for (MemberRole role : MemberRole.values()) {
@@ -94,7 +94,7 @@ public final class MemberRegistry implements MemberLifecycle {
* An immutable copy of the live {@code terminal_id → slot name} bindings.
*
* <p>Passed to {@link CallerResolver} as the source of architect identity, and what
* {@code fleet_whoami}/the roster will read to say which slot a pane hosts. Empty until the
* {@code bridge_whoami}/the roster will read to say which slot a pane hosts. Empty until the
* spawn lifecycle binds a slot.
*/
public Map<String, String> snapshot() {
@@ -1,4 +1,4 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
/**
* A resolved caller: its {@link Role}, and — for a worker — the herdr {@code terminal_id} that
@@ -39,14 +39,14 @@ public record Principal(Role role, String terminal, long pid, String name) {
*
* <p>Carries {@link Role#PRIMARY}: a lead <em>is</em> a primary as far as authorization goes,
* so every existing {@code isPrimary()} gate keeps working unchanged and the role table needed
* no new entry. The name is reporting only — it lets {@code fleet_whoami} say <em>which</em>
* no new entry. The name is reporting only — it lets {@code bridge_whoami} say <em>which</em>
* lead is asking once more than one is configured.
*
* <p><strong>CB-532: a lead now carries the terminal it was matched by.</strong> Under CB-530 it
* deliberately did not, because {@code terminal} meant "which worker pane" everywhere and a
* non-null one would have enrolled the lead in the worker presence map. That reading was what
* made a lead unaddressable: {@link #ownsSession} could never be true for it, so
* {@code fleet_reply} was refused and one lead could send to another but never be answered.
* {@code bridge_reply} was refused and one lead could send to another but never be answered.
* The terminal now means "which pane is this caller", the presence map keys on
* {@link #isSpawnedMember()} instead, and a lead is a peer that can both send and receive.
*/
@@ -63,7 +63,7 @@ public record Principal(Role role, String terminal, long pid, String name) {
* An architect (CB-548), identified by the slot it occupies and the pane bound to it.
*
* <p>Carries {@link Role#ARCHITECT}. {@code slotName} is reporting only — it lets
* {@code fleet_whoami} say <em>which</em> architect slot is asking, and it is the key the
* {@code bridge_whoami} say <em>which</em> architect slot is asking, and it is the key the
* (future) spawn lifecycle reads a profile back from. Identity is the {@code terminal}: like a
* worker's it comes from the connection and the live terminal→slot binding, so
* {@code ownsSession} works exactly as it does for a worker — an architect acts as its own
@@ -1,4 +1,4 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
/**
* What a caller is allowed to be on the bus (CB-501).
@@ -1,14 +1,13 @@
package dev.ltms.fleet.config;
package dev.ltms.bridged.config;
import com.fasterxml.jackson.annotation.JsonIgnoreProperties;
import com.fasterxml.jackson.core.JsonParser;
import com.fasterxml.jackson.core.JsonToken;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.dataformat.yaml.YAMLFactory;
import dev.ltms.fleet.msg.AmqpReplyInbox;
import dev.ltms.fleet.msg.LeadMailbox;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.placement.PlacementPolicies;
import dev.ltms.bridged.msg.AmqpReplyInbox;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.placement.PlacementPolicies;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -26,14 +25,13 @@ import java.util.Map;
import java.util.Set;
/**
* {@code fleetd} configuration, loaded from a YAML file (see
* {@code fleetd.example.yaml}). Unknown keys are ignored so config can grow ahead
* {@code bridged} configuration, loaded from a YAML file (see
* {@code bridged.example.yaml}). Unknown keys are ignored so config can grow ahead
* of the code — but an unknown <em>top-level</em> key is logged as a WARN at load (CB-530), because
* silently dropping a whole block is indistinguishable from honouring it.
*
* @param bind REST/MCP listen host:port
* @param herdrSocket path to the lead herdr Unix socket ({@code null} → client default)
* @param memberHerdrSocket optional member herdr Unix socket ({@code null}/blank → lead socket)
* @param herdrSocket path to herdr's Unix socket ({@code null} → client default)
* @param profiles named backend profiles, keyed by profile name (multi-backend fleet). A
* profile answers <em>which backend</em> — model, CLI adapter, credentials,
* cost. It says nothing about what the member spawned on it is for; that is
@@ -71,17 +69,11 @@ import java.util.Set;
* @param memberCredentials deny-by-default policy (CB-596) for which of the operator's own host
* credentials a spawned member's pane inherits. {@code null} (the block
* omitted) blocks nothing — see {@link MemberCredentials}.
* @param coordinator shared cross-host leader coordination broker: a SEPARATE AMQP vhost from
* {@link #broker} used only for lead-to-lead traffic (member/worker inboxes
* stay on {@code broker}'s vhost). {@code null} → no lead mailbox is opened.
* Config parsing + accessors only — nothing here wires it into a live
* {@code LeadMailbox}; that is a separate ticket. See {@link Coordinator}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record FleetConfig(
public record BridgedConfig(
Bind bind,
String herdrSocket,
String memberHerdrSocket,
Map<String, Profile> profiles,
Guard guard,
String worktreeRoot,
@@ -97,62 +89,49 @@ public record FleetConfig(
Auth auth,
ConfigReload configReload,
Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials,
Coordinator coordinator) {
/** Back-compat form before the {@code coordinator:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials) {
this(bind, herdrSocket, null, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, null);
}
MemberCredentials memberCredentials) {
/** Back-compat form before the CB-596 {@code memberCredentials:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds) {
this(bind, herdrSocket, null, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, null, null);
configReload, quarantineCooldownSeconds, null);
}
/** Default cooldown (CB-578 stage B) when {@code quarantineCooldownSeconds} is absent/non-positive. */
public static final int DEFAULT_QUARANTINE_COOLDOWN_SECONDS = 1800;
/** Back-compat 14-arg form — no {@code configReload:} block, so file watching stays off. */
public FleetConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, String placement, Auth auth) {
this(bind, herdrSocket, null, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, null, null, null, null);
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, null, null);
}
/** Back-compat form before the optional {@code health:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, String placement, Auth auth, ConfigReload configReload) {
this(bind, herdrSocket, null, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, configReload, null, null, null);
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, configReload, null);
}
/** Back-compat form before the CB-578 stage B {@code quarantineCooldownSeconds} field was added. */
public FleetConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload) {
this(bind, herdrSocket, null, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, null, null, null);
configReload, null);
}
/**
@@ -168,7 +147,7 @@ public record FleetConfig(
* {@code WeightedRoundRobinPolicy}), so losing it makes equal-weight placement
* non-reproducible across restarts. Unmodifiable-wrap instead of copy-and-scramble.
*/
public FleetConfig {
public BridgedConfig {
if (profiles != null && !profiles.isEmpty()) {
Map<String, Profile> normalized = new LinkedHashMap<>();
profiles.forEach((name, p) -> normalized.put(name,
@@ -198,14 +177,14 @@ public record FleetConfig(
* @param argv launch command; defaults to {@code ["claude"]}
* @param placement where a worker lands: {@code "tab"} (default — its own tab in the
* worker space) or {@code "pane"} (legacy — split the focused tab)
* @param workspace label of the shared fleet space; found-or-created on first spawn
* (default {@code "fleet"}). The lead and every member share this one
* space so the operator sees a single "session" with many tabs.
* @param workspace label of the dedicated worker space; found-or-created on first
* spawn (default {@code "bridged-workers"}). A future per-session
* layout is just a distinct label here — the shared space is default.
* @param tabLabel template for a worker tab's label; {@code {profile}}/{@code {model}}
* and {@code {n}} (per-worker number, to keep sibling tabs distinct)
* are substituted (default {@code "worker: {profile} #{n}"})
* @param mcpUrl bridge MCP URL to provision into the worker's {@code configDir} so it
* can call {@code fleet_reply} ({@code null}/blank → no provisioning; the
* can call {@code bridge_reply} ({@code null}/blank → no provisioning; the
* worker won't reply, only the fallback/timeout resolves the send)
* @param cwd fixed working directory for this profile's workers (CB-112 "told otherwise");
* {@code null}/blank → inherit the primary's cwd, else the daemon's
@@ -220,7 +199,7 @@ public record FleetConfig(
* worktrees then accumulate with no error anywhere.
* <p>Checked on 2026-08-15 (CB-581): inert as configured. Tracked overlay
* files carry {@code --skip-worktree} so {@code --porcelain} cannot see them,
* {@code fleetd.yaml} is gitignored, and the default pair {@code .env} /
* {@code bridged.yaml} is gitignored, and the default pair {@code .env} /
* {@code .envrc} does not exist in this repo. Note the default applies to
* <em>every</em> profile, so creating either file at the repo root is enough
* to make it live. Add a new overlay path to {@code .gitignore} in the same
@@ -245,7 +224,7 @@ public record FleetConfig(
* auto-select this profile": it is excluded from every automatic policy's
* pool the same way a quarantined candidate is (see
* {@code PlacementPolicyUtil}). This does not make the profile
* unreachable — an explicit {@code fleet_spawn{profile:"..."}} bypasses
* unreachable — an explicit {@code bridge_spawn{profile:"..."}} bypasses
* placement entirely and still resolves it. Weights among the remaining
* (non-excluded) candidates need not sum to 1.0; only their ratios matter.
* @param maxLoad max live workers allowed on this profile at one time. Absent
@@ -255,7 +234,7 @@ public record FleetConfig(
* the same way a {@code weight <= 0} profile is (see
* {@code PlacementPolicyUtil.available()}, which already treats "at cap"
* and "excluded" alike), and an explicit
* {@code fleet_spawn{profile:"..."}} against it is refused too (see
* {@code bridge_spawn{profile:"..."}} against it is refused too (see
* {@code CompositePeerLauncher.enforceMaxLoad}) — a cap is a capacity
* statement that does not stop being true just because the profile was
* named directly. A negative value has no sane meaning (there is no
@@ -263,7 +242,7 @@ public record FleetConfig(
* instead, naming the profile and the key. Live means any session the
* registry still owns (acquired and not yet released), in any state.
* @param kind which peer launcher spawns this profile: {@code "claude-code"} (default —
* the {@link dev.ltms.fleet.member.ClaudeCodeLauncher}) or {@code "opencode"}.
* the {@link dev.ltms.bridged.member.ClaudeCodeLauncher}) or {@code "opencode"}.
* The {@code CompositePeerLauncher} routes {@code spawn}/reap by this value, so
* each adapter drives only its own kind. Normalised to lower-case; blank ⇒ the
* default. It selects the adapter, not the transport — placement, tabs, cwd, and
@@ -280,7 +259,7 @@ public record FleetConfig(
* {@code ANTHROPIC_BASE_URL} or {@code ANTHROPIC_AUTH_TOKEN} is refused at
* config load (CB-542): on the subscription path no guard would vet it.
* @param exhaustedPattern regex matched against a completion-fallback scrape (CB-578 stage A) to
* classify a turn that ended with no {@code fleet_reply} as the backend
* classify a turn that ended with no {@code bridge_reply} as the backend
* having refused on a subscription usage limit, rather than a real answer.
* {@code null}/blank ⇒ the classification never fires for this profile and
* today's completion-fallback behaviour is unchanged. Every backend words
@@ -296,19 +275,6 @@ public record FleetConfig(
* profile that does not opt in. Read live off the current config, so it is
* HOT: a change takes effect on the next exhaustion classification / spawn,
* no restart needed.
* @param autoCompactWindow opt-in per-profile token window that forces a spawned member to
* auto-compact its context at (Claude Code) or within (opencode) a bound the
* operator chooses, instead of the backend's own default. {@code null} (the
* default) leaves today's behaviour exactly — opencode already forces
* {@code compaction.auto: true} unconditionally (CB-523) but has no absolute
* window, and Claude Code has neither. When set, validated at config load to
* {@code [100000, 1000000]} — the band Claude Code's own {@code --autocompact
* <tokens>} flag accepts. The two backends honour it differently: Claude Code
* compacts AT this window (a launch-time {@code --autocompact} flag);
* opencode has no such knob, so this is applied as the model's
* {@code limit.context} instead, which bounds the window opencode compacts
* <em>within</em>, and only when {@code model} resolves to a
* {@code provider/model} pair.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Profile(String profile, String baseUrl, String model,
@@ -323,13 +289,9 @@ public record FleetConfig(
Integer maxLoad,
Boolean subscription,
String exhaustedPattern,
String credentialId,
String ideMcpUrl,
String ideProjectDir,
String ideOpenCommand,
Integer autoCompactWindow) {
String credentialId) {
/** Peer kind spawned by {@link dev.ltms.fleet.member.ClaudeCodeLauncher} (the default). */
/** Peer kind spawned by {@link dev.ltms.bridged.member.ClaudeCodeLauncher} (the default). */
public static final String KIND_CLAUDE_CODE = "claude-code";
/** Peer kind spawned by the opencode adapter (CB-402). */
public static final String KIND_OPENCODE = "opencode";
@@ -343,9 +305,9 @@ public record FleetConfig(
? (KIND_CLAUDE_CODE.equals(k) ? List.of("claude") : List.of(k))
: List.copyOf(argv);
kind = k;
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "FLEETD_WORKER_TOKEN" : tokenEnv;
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_WORKER_TOKEN" : tokenEnv;
placement = (placement == null || placement.isBlank()) ? "tab" : placement.toLowerCase();
workspace = (workspace == null || workspace.isBlank()) ? "fleet" : workspace;
workspace = (workspace == null || workspace.isBlank()) ? "bridged-workers" : workspace;
// CB-557: no per-profile default any more. A label is generated from the member's ROLE
// ("dev: sonnet #2"), which a profile cannot know, so the template lives on `fleet:` and
// this field is only an override for a profile that wants its own. Blank ⇒ defer.
@@ -386,21 +348,6 @@ public record FleetConfig(
// "quarantines alone" fallback actually lives, so today's behaviour needs no defaulting
// here at all.
credentialId = (credentialId == null || credentialId.isBlank()) ? null : credentialId;
// CB-634: opt-in per profile, default off. A URL, not a boolean — host and port are
// host-specific, mirroring mcpUrl. When set, the member gets the IDE Index MCP mounted
// (pinned to its own worktree via the charter). Blank ⇒ off.
ideMcpUrl = (ideMcpUrl == null || ideMcpUrl.isBlank()) ? null : ideMcpUrl;
// CB-634 auto-open: both are only read when hasIdeMcp(). ideProjectDir is the repo-relative
// module dir IntelliJ must open (this repo's pom lives in `fleetd/`, not at the root), and
// it is also the project_path the overlay pins. Blank ⇒ the worktree root (unchanged before
// auto-open). ideOpenCommand is the host command that opens that dir in the IDE, with {dir}
// substituted; blank ⇒ no auto-open (the operator opens the module by hand).
ideProjectDir = (ideProjectDir == null || ideProjectDir.isBlank()) ? null : ideProjectDir;
ideOpenCommand = (ideOpenCommand == null || ideOpenCommand.isBlank()) ? null : ideOpenCommand;
// Opt-in per profile, default off (null). No clamping here — unlike ideMcpUrl/ideProjectDir
// there is no blank-string form to normalize (it's an Integer), and the [100000, 1000000]
// range is enforced eagerly at config load (rejectAutoCompactWindowOutOfRange), naming the
// profile, rather than silently clamped here. A profile that never sets it keeps null.
}
/**
@@ -445,47 +392,7 @@ public record FleetConfig(
public Profile withProfile(String p) {
return new Profile(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, subscription,
exhaustedPattern, credentialId, ideMcpUrl, ideProjectDir, ideOpenCommand, autoCompactWindow);
}
/**
* Backward-compatible constructor without the {@code autoCompactWindow} field — the profile
* leaves auto-compaction at the backend's own default (opencode's unconditional
* {@code compaction.auto: true}, or Claude Code's built-in threshold), exactly as before this
* key existed. This is the shape the canonical constructor had before the field was added —
* every pre-existing Java call site (and any YAML that omits the key) keeps compiling and
* behaving identically; Jackson binds the canonical (longest) constructor, so YAML omitting
* {@code autoCompactWindow:} still lands here as {@code null} via that path, not this one.
*/
public Profile(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind, Map<String, String> env, Float weight, Integer maxLoad,
Boolean subscription, String exhaustedPattern, String credentialId,
String ideMcpUrl, String ideProjectDir, String ideOpenCommand) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad,
subscription, exhaustedPattern, credentialId, ideMcpUrl, ideProjectDir, ideOpenCommand,
null);
}
/**
* Backward-compatible constructor without the CB-634 auto-open fields
* ({@code ideProjectDir}/{@code ideOpenCommand}) — a profile that opts into IDE MCP still
* pins the worktree root and does not auto-open. Keeps pre-auto-open call sites (and any YAML
* that omits the keys) compiling and behaving identically. This is the shape the canonical
* constructor had before the two fields were added.
*/
public Profile(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind, Map<String, String> env, Float weight, Integer maxLoad,
Boolean subscription, String exhaustedPattern, String credentialId, String ideMcpUrl) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad,
subscription, exhaustedPattern, credentialId, ideMcpUrl, null, null);
exhaustedPattern, credentialId);
}
/** True when this profile is served by the Claude Code adapter (the default kind). */
@@ -534,7 +441,7 @@ public record FleetConfig(
Boolean subscription, String exhaustedPattern) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad,
subscription, exhaustedPattern, null, null);
subscription, exhaustedPattern, null);
}
/** True when this profile's CB-578 stage A backend-exhausted classification is configured. */
@@ -565,7 +472,7 @@ public record FleetConfig(
* {@code ANTHROPIC_AUTH_TOKEN} is injected from {@code tokenEnv}. An {@code env:} entry for
* either is a bypass vector — it is what the subscription path would otherwise let survive
* unguarded — so it is rejected at config load (see
* {@link FleetConfig#validateSubscriptionProfiles()}).
* {@link BridgedConfig#validateSubscriptionProfiles()}).
*/
public boolean envCarriesAnthropicBinding() {
return env != null
@@ -582,19 +489,6 @@ public record FleetConfig(
return mcpUrl != null && !mcpUrl.isBlank();
}
/**
* True when the IDE Index MCP should be mounted into a spawned worker (CB-634), pinned to
* the worker's own worktree via the charter. Opt-in per profile, default off.
*/
public boolean hasIdeMcp() {
return ideMcpUrl != null && !ideMcpUrl.isBlank();
}
/** True when this profile mounts any MCP server into its member — the bridge, the IDE, or both. */
public boolean mountsAnyMcp() {
return hasMcp() || hasIdeMcp();
}
/**
* Render this member's tab label (CB-557): {@code {role}}, {@code {profile}},
* {@code {model}} and {@code {n}} are substituted.
@@ -644,9 +538,6 @@ public record FleetConfig(
Integer paneProbeIntervalSeconds, Notifications notifications) {
public boolean isEnabled() { return Boolean.TRUE.equals(enabled); }
public int intervalOrDefault() { return Math.max(15, intervalSeconds == null ? 30 : intervalSeconds); }
public int workingSuspectAfterOrDefault() {
return Math.max(300, workingSuspectAfterSeconds == null ? 600 : workingSuspectAfterSeconds);
}
public record Notifications(String mode) {
public boolean configured() { return "webhook".equalsIgnoreCase(mode); }
}
@@ -654,58 +545,22 @@ public record FleetConfig(
/**
* External AMQP broker for durable, cross-restart reply delivery (CB-307 Stage 2). Its mere
* presence swaps the in-memory {@code ReplyInbox} for the AMQP-backed adapter; absent, fleetd
* presence swaps the in-memory {@code ReplyInbox} for the AMQP-backed adapter; absent, bridged
* stays soft-state. Production default is LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1
* and is a URI-only swap.
*
* @param uri AMQP connection URI, e.g. {@code amqp://guest:guest@127.0.0.1:5672/}. Blank/
* {@code null} ⇒ the broker block is treated as absent (in-memory adapter).
* Ignored when {@code uriEnv} is set.
* @param uriEnv name of a host env var holding the AMQP URI (CB-151). The URI carries its
* credentials inline, so giving the variable <em>name</em> keeps the password
* out of the config file, same as {@code auth.tokenEnv}/{@code Profile.tokenEnv}.
* Wins over {@code uri} whenever set. Blank/{@code null} ⇒ ignored.
* @param prefetch CB-527: the consumer's {@code basicQos} prefetch count, bounding how many
* unacked messages the AMQP inbox holds in-heap per owned target. {@code null}/
* non-positive ⇒ {@link AmqpReplyInbox#DEFAULT_PREFETCH}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Broker(String uri, String uriEnv, Integer prefetch) {
public record Broker(String uri, Integer prefetch) {
/** True when a {@code uriEnv} is configured by name, whether or not its variable resolves. */
public boolean hasUriEnv() {
return uriEnv != null && !uriEnv.isBlank();
}
/**
* True when a usable broker URI is configured (an empty block does not enable AMQP).
* Honors {@code uriEnv} first: if it names a variable that is unset or blank, the broker is
* <em>not</em> configured (the daemon falls back to the in-memory inbox) — a bare {@code uri}
* is only consulted when no {@code uriEnv} is set. Env lookup makes this process-dependent;
* callers already reading {@link System#getenv} are the right ones to invoke it.
*/
/** True when a usable broker URI is configured (an empty block does not enable AMQP). */
public boolean isConfigured() {
return effectiveUri() != null;
}
/**
* The effective AMQP URI to connect with. {@code uriEnv} wins when set (both over {@code uri}
* and alone). When {@code uriEnv} names a variable that is unset or blank, returns {@code
* null} rather than falling back to {@code uri} — an operator who moved to the secret store
* must not silently drop back onto a stale clear-text URI. Returns the literal {@code uri}
* when no {@code uriEnv} is configured.
*/
public String effectiveUri() {
return effectiveUri(System.getenv());
}
/** As {@link #effectiveUri()}, reading the variable value from {@code env} (the injection seam). */
public String effectiveUri(Map<String, String> env) {
if (hasUriEnv()) {
String value = env.get(uriEnv);
return (value != null && !value.isBlank()) ? value : null;
}
return (uri != null && !uri.isBlank()) ? uri : null;
return uri != null && !uri.isBlank();
}
/** The prefetch to use, defaulting to {@link AmqpReplyInbox#DEFAULT_PREFETCH} when unset. */
@@ -714,82 +569,6 @@ public record FleetConfig(
}
}
/**
* Shared cross-host leader coordination broker: a {@link LeadMailbox} lets two leads on
* different daemons — possibly different hosts — exchange durable messages, which a herdr pane
* injection (how {@code fleet_send} reaches a lead today) cannot do at all. Its mere presence is
* config only in this ticket: nothing here opens a live {@code LeadMailbox} yet, that wiring is
* a separate ticket.
*
* <p>Deliberately a SEPARATE vhost from {@link Broker}, not a reuse of it. {@link Broker} is
* per-fleet — its queues are named by worker session id, and two fleets sharing one broker stay
* isolated by vhost (see {@code Two fleets share one LavinMQ}). Leader coordination is meant to
* cross exactly that boundary: two independently-owned fleets' leads talking to each other. Using
* the same vhost would either leak member traffic across the fleet boundary this is meant to
* cross, or force every fleet's members onto one shared vhost to get leader coordination — a
* second, dedicated vhost keeps "member inboxes stay fleet-local" true while still letting leads
* reach across fleets.
*
* @param uri AMQP connection URI for the coordination vhost, e.g.
* {@code amqp://guest:guest@127.0.0.1:5672/coord}. Blank/{@code null} ⇒ the
* coordinator block is treated as unconfigured. Ignored when {@code uriEnv} is set.
* @param uriEnv name of a host env var holding the AMQP URI, same convention as
* {@link Broker#uriEnv()} — keeps the credential out of the config file. Wins
* over {@code uri} whenever set. Blank/{@code null} ⇒ ignored.
* @param selfId this daemon's own lead coord-id — the name its {@code LeadMailbox} is owned
* under ({@code lead.<selfId>.inbox}), e.g. {@code "mac-opus"}. Blank/
* {@code null} ⇒ kept as {@code null} (no self id configured).
* @param prefetch the consumer's {@code basicQos} prefetch count. {@code null}/non-positive ⇒
* {@link LeadMailbox#DEFAULT_PREFETCH}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Coordinator(String uri, String uriEnv, String selfId, Integer prefetch) {
public Coordinator {
selfId = (selfId == null || selfId.isBlank()) ? null : selfId;
}
/** True when a {@code uriEnv} is configured by name, whether or not its variable resolves. */
public boolean hasUriEnv() {
return uriEnv != null && !uriEnv.isBlank();
}
/**
* True when a usable coordination broker URI is configured (an empty block does not enable
* it). Honors {@code uriEnv} first: if it names a variable that is unset or blank, the
* coordinator is <em>not</em> configured — a bare {@code uri} is only consulted when no
* {@code uriEnv} is set.
*/
public boolean isConfigured() {
return effectiveUri() != null;
}
/**
* The effective AMQP URI to connect with. {@code uriEnv} wins when set (both over
* {@code uri} and alone). When {@code uriEnv} names a variable that is unset or blank,
* returns {@code null} rather than falling back to {@code uri} — an operator who moved to
* the secret store must not silently drop back onto a stale clear-text URI. Returns the
* literal {@code uri} when no {@code uriEnv} is configured.
*/
public String effectiveUri() {
return effectiveUri(System.getenv());
}
/** As {@link #effectiveUri()}, reading the variable value from {@code env} (the injection seam). */
public String effectiveUri(Map<String, String> env) {
if (hasUriEnv()) {
String value = env.get(uriEnv);
return (value != null && !value.isBlank()) ? value : null;
}
return (uri != null && !uri.isBlank()) ? uri : null;
}
/** The prefetch to use, defaulting to {@link LeadMailbox#DEFAULT_PREFETCH} when unset. */
public int prefetchOrDefault() {
return (prefetch != null && prefetch > 0) ? prefetch : LeadMailbox.DEFAULT_PREFETCH;
}
}
/**
* Optional pinned primary terminal config (CB-307). When present with a non-blank
* {@code terminal}, the bridge uses this as the primary's herdr identity instead of
@@ -826,7 +605,7 @@ public record FleetConfig(
* work as peers — the second is silently demoted and refused every orchestration call.
*
* <p>{@code kind} and {@code model} are descriptive only: they document what runs in the pane
* and are reported back by {@code fleet_whoami}.
* and are reported back by {@code bridge_whoami}.
*
* <p><b>A lead is now also creatable (CB-557).</b> Before, nothing spawned one — a lead
* pre-existed, which is why it had to be recognised by configuration rather than created. With
@@ -862,12 +641,11 @@ public record FleetConfig(
String workspace, String cwd) {
/**
* Where an auto-launched lead's tab is created (CB-558). It defaults to the SAME shared
* {@code "fleet"} space the members use, so the operator sees one "session" with many tabs.
* The scanner no longer excludes member spaces — it tells a lead from a member by the exact
* tab label, so a lead sharing the members' space is still discovered (see LeadLauncher).
* Where an auto-launched lead's tab is created (CB-558). It must NOT be a member workspace:
* {@code LeadTabScanner} excludes those wholesale, so a lead placed in one would never be
* discovered and the daemon would relaunch it on every boot.
*/
public static final String DEFAULT_WORKSPACE = "fleet";
public static final String DEFAULT_WORKSPACE = "leads";
public Leader {
instances = (instances == null || instances < 0) ? 1 : instances;
@@ -1033,7 +811,7 @@ public record FleetConfig(
/**
* Opt-in idle-lead heartbeat (CB-551): when the single lead has been continuously idle past
* {@code idleAfterSeconds} with no open {@code fleet_send} driving it, nudge it back to work.
* {@code idleAfterSeconds} with no open {@code bridge_send} driving it, nudge it back to work.
*
* <p>Deliberately opt-in ({@code null} ⇒ off, exactly like {@code leadScan:}). The heartbeat
* spends the operator's model subscription on its own initiative — it prompts the lead to start
@@ -1066,7 +844,7 @@ public record FleetConfig(
}
/**
* Watch {@code fleetd.yaml} and re-read it when it changes (CB-559).
* Watch {@code bridged.yaml} and re-read it when it changes (CB-559).
*
* <p>Opt-in, like every other block that acts on its own initiative. A daemon that reloads
* whenever a file is saved would apply a half-finished edit the moment an editor writes it, and
@@ -1115,7 +893,7 @@ public record FleetConfig(
*
* <p>Member identity never depends on this block: a loopback peer PID that maps to a herdr
* pane is unforgeable and is always honoured (see
* {@link dev.ltms.fleet.mcp.ConnectionIdentity}). This only decides what happens for
* {@link dev.ltms.bridged.mcp.ConnectionIdentity}). This only decides what happens for
* <em>everyone else</em>.
*
* @param mode {@code "loopback-trust"} (default) — any loopback caller that is not a known
@@ -1124,7 +902,7 @@ public record FleetConfig(
* a caller must present {@code Authorization: Bearer <token>} or it is
* {@code ANONYMOUS} and authorized for nothing.
* @param tokenEnv name of the host env var holding the bearer token; the literal value is
* never stored in config. Defaults to {@code FLEETD_API_TOKEN}. Only read
* never stored in config. Defaults to {@code BRIDGED_API_TOKEN}. Only read
* when {@code mode} is {@code token}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
@@ -1137,7 +915,7 @@ public record FleetConfig(
public Auth {
mode = (mode == null || mode.isBlank()) ? MODE_LOOPBACK_TRUST : mode.toLowerCase();
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "FLEETD_API_TOKEN" : tokenEnv;
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_API_TOKEN" : tokenEnv;
}
/** True when a bearer token is required of every non-worker caller. */
@@ -1179,86 +957,27 @@ public record FleetConfig(
* UNLESS it is also in {@link #allow}. A name that shows up in neither list is not silently
* allowed — see {@code HerdrPeerLauncher}'s gap detector, which logs it.
*
* <p><b>CB-633: allow-list.</b> Deny-by-default's overlay is applied BEFORE the pane's login
* shell runs, so any file that chain sources can re-export over it — and did. The allow-list
* policy moves the control to a generated ZDOTDIR whose startup files run the scrub LAST, after
* the whole operator chain, and blank every exported variable not on an allow-list DERIVED from
* what the launcher itself injects (profiles' tokenEnv/gitTokenEnv/gitHostEnv/env keys plus an
* infrastructure set) — never hand-typed, so adding a profile cannot break a spawn. Under this
* policy {@link #allow} and {@link #known} stop being a control and become reporting only.
*
* @param policy how the block is computed. {@link #POLICY_DENY_BY_DEFAULT} (the default; also
* accepted as {@link #POLICY_DENY_LIST}) shadows each {@code known}-but-not-allowed
* name in the pane-creation env overlay — which a login shell that re-exports the
* name defeats (see CB-596's round-2 correction). {@link #POLICY_ALLOW_LIST}
* (CB-633) replaces the overlay with a per-spawn ZDOTDIR scrub that runs AFTER the
* pane's login shell has finished sourcing everything, blanking every variable not
* on the DERIVED allow-list (derived from what the launcher itself injects — never
* hand-typed). Under {@code allow-list}, {@link #allow} and {@link #known} are
* REPORTING ONLY: they feed the gap WARN, they are no longer a control.
* @param policy how the block is computed. Only {@link #POLICY_DENY_BY_DEFAULT} is understood
* today; {@code null}/blank defaults to it. An operator's own deny-list is
* deliberately not supported — see above.
* @param allow credential names a member legitimately needs (e.g. the gateway token it reaches
* the LLM through, the repo-scoped forge token it opens its own PR with). Under
* deny-list/deny-by-default every name here is left unmentioned in the pane's env
* overlay; under allow-list this list is reporting only — the control is derived,
* not configured.
* @param known every credential name the operator's store is known to export. Under
* deny-list/deny-by-default every name here that is NOT also in {@link #allow} is
* overlaid with a non-secret sentinel value. Under allow-list this list is
* reporting only.
* @param sshAuthSock whether the member may inherit {@code SSH_AUTH_SOCK} under the allow-list
* policy ({@code "allow"}) or must have it blanked ({@code "block"}, the default).
* This is a DECISION, never a default: {@code SSH_AUTH_SOCK} is a handle to the
* operator's ssh-agent, and a member holding it can sign with the operator's own
* keys — but it appears in no secret file and is credential-shaped like nothing on
* any list, which is why three earlier tickets missed it (gitea #110 / CB-607).
* Blocking it breaks git over SSH inside the member; allow it only when members do
* not need to authenticate as the operator over SSH. Ignored under deny-list /
* deny-by-default, which never touch the name.
* the LLM through, the repo-scoped forge token it opens its own PR with). Every
* name here is left unmentioned in the pane's env overlay, so the value the pane's
* own (login) shell exports passes through untouched.
* @param known every credential name the operator's store is known to export. Every name here
* that is NOT also in {@link #allow} is overlaid with a non-secret sentinel value,
* shadowing whatever the pane's login shell would otherwise export for it.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record MemberCredentials(String policy, List<String> allow, List<String> known,
String sshAuthSock) {
public record MemberCredentials(String policy, List<String> allow, List<String> known) {
/** Default policy: block every {@code known} name not in {@code allow}, via the env overlay. */
/** The only policy this build understands: block every {@code known} name not in {@code allow}. */
public static final String POLICY_DENY_BY_DEFAULT = "deny-by-default";
/** Alias of {@link #POLICY_DENY_BY_DEFAULT}, spelled the way CB-633 names the two policies. */
public static final String POLICY_DENY_LIST = "deny-list";
/**
* CB-633: derive the kept-name set from what the launcher itself injects, generate a
* per-spawn ZDOTDIR whose startup files blank every exported variable not on it AFTER the
* pane's shell has finished sourcing the operator's chain. The scrub is sourced from both
* the generated {@code .zshrc} and {@code .zlogin}, because herdr opens a LOGIN zsh on macOS
* and a plain interactive one on Linux — see {@code EnvAllowListScrub}.
*/
public static final String POLICY_ALLOW_LIST = "allow-list";
/** The pre-CB-633 three-field form — {@code sshAuthSock} defaults to blocked. */
public MemberCredentials(String policy, List<String> allow, List<String> known) {
this(policy, allow, known, null);
}
public MemberCredentials {
String normalizedPolicy = (policy == null || policy.isBlank())
? POLICY_DENY_BY_DEFAULT : policy.toLowerCase(java.util.Locale.ROOT);
// deny-list is an alias of deny-by-default, not a third behaviour — normalize to one
// spelling so every isDenyList()-style check has one value to compare against.
policy = POLICY_DENY_LIST.equals(normalizedPolicy) ? POLICY_DENY_BY_DEFAULT : normalizedPolicy;
policy = (policy == null || policy.isBlank()) ? POLICY_DENY_BY_DEFAULT : policy.toLowerCase();
allow = allow == null ? List.of() : List.copyOf(allow);
known = known == null ? List.of() : List.copyOf(known);
sshAuthSock = (sshAuthSock != null && "allow".equalsIgnoreCase(sshAuthSock.trim()))
? "allow" : "block";
}
/** True when this block selects the CB-633 derived-allow-list policy. */
public boolean isAllowList() {
return POLICY_ALLOW_LIST.equals(policy);
}
/** True when {@code SSH_AUTH_SOCK} may pass through under the allow-list policy. Default: no. */
public boolean sshAuthSockAllowed() {
return "allow".equals(sshAuthSock);
}
/** {@link #allow} as a set, for membership checks. */
@@ -1309,7 +1028,7 @@ public record FleetConfig(
return defaultProfileFor(MemberRole.DEV);
}
private static final Logger log = LoggerFactory.getLogger(FleetConfig.class);
private static final Logger log = LoggerFactory.getLogger(BridgedConfig.class);
private static final ObjectMapper YAML = new ObjectMapper(new YAMLFactory());
@@ -1318,17 +1037,17 @@ public record FleetConfig(
* {@link #warnUnknownTopLevelKeys}. Keep in step with the record components.
*
* <p>Package-private (not {@code private}) so a test can assert every key here is documented in
* {@code fleetd.example.yaml} — the only committed description of the config schema, since
* {@code fleetd.yaml} itself is gitignored.
* {@code bridged.example.yaml} — the only committed description of the config schema, since
* {@code bridged.yaml} itself is gitignored.
*/
static final Set<String> KNOWN_TOP_LEVEL_KEYS = Set.of(
"bind", "herdrSocket", "memberHerdrSocket", "profiles", "guard", "worktreeRoot",
"bind", "herdrSocket", "profiles", "guard", "worktreeRoot",
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
"memberCredentials", "coordinator");
"memberCredentials");
/** Load and validate config from {@code path}. */
public static FleetConfig load(Path path) {
public static BridgedConfig load(Path path) {
try {
String yaml = Files.readString(path);
rejectRenamedTopLevelKeys(yaml);
@@ -1336,19 +1055,18 @@ public record FleetConfig(
warnUnknownTopLevelKeys(yaml, path);
rejectDuplicateMemberSlots(yaml);
rejectNegativeMaxLoad(yaml);
rejectAutoCompactWindowOutOfRange(yaml);
rejectUnknownKind(yaml);
rejectUnknownAuthMode(yaml);
rejectUnknownPlacement(yaml);
rejectUnknownMemberCredentialsPolicy(yaml);
FleetConfig cfg = YAML.readValue(yaml, FleetConfig.class);
BridgedConfig cfg = YAML.readValue(yaml, BridgedConfig.class);
// CB-606: validated here, eagerly, using PlacementPolicies.fromName as the single source
// of truth — not lazily at first spawn (see CompositePeerLauncher's placementPolicy
// Supplier), where a bad name would still start a daemon that looks healthy.
rejectUnknownPlacementPolicy(cfg.placement());
return cfg.withDefaults();
} catch (IOException e) {
throw new UncheckedIOException("cannot read fleetd config at " + path, e);
throw new UncheckedIOException("cannot read bridged config at " + path, e);
}
}
@@ -1639,52 +1357,6 @@ public record FleetConfig(
}
}
/** Lowest {@code autoCompactWindow} Claude Code's {@code --autocompact <tokens>} flag accepts. */
static final int AUTO_COMPACT_WINDOW_MIN = 100_000;
/** Highest {@code autoCompactWindow} Claude Code's {@code --autocompact <tokens>} flag accepts. */
static final int AUTO_COMPACT_WINDOW_MAX = 1_000_000;
/**
* Reject a profile whose {@code autoCompactWindow:} is set but outside the token band Claude
* Code's own {@code --autocompact <tokens>} flag accepts (100k–1M), naming both the profile and
* the value.
*
* <p>Unset/{@code null} means "off" and passes silently — today's behaviour for every profile
* that does not opt in (see {@link Profile#autoCompactWindow()}). A profile that DOES set the key
* is validated eagerly, at config load, rather than failing later when Claude Code itself refuses
* the launch flag on spawn — the same "fail loud at load, not lazily at first spawn" reasoning as
* {@link #rejectNegativeMaxLoad} and {@link #rejectUnknownPlacementPolicy}.
*
* @param yaml the raw config text
* @throws IllegalStateException when any profile's {@code autoCompactWindow} is set and outside
* {@code [100000, 1000000]}
*/
static void rejectAutoCompactWindowOutOfRange(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return; // a malformed file is reported by the real parse, not here
}
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
return;
}
List<String> bad = profiles.entrySet().stream()
.filter(e -> e.getValue() instanceof Map<?, ?> p
&& p.get("autoCompactWindow") instanceof Number n
&& (n.doubleValue() < AUTO_COMPACT_WINDOW_MIN || n.doubleValue() > AUTO_COMPACT_WINDOW_MAX))
.map(e -> String.valueOf(e.getKey()))
.sorted()
.toList();
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: profile(s) [" + String.join(", ", bad)
+ "] set autoCompactWindow outside [" + AUTO_COMPACT_WINDOW_MIN + ", "
+ AUTO_COMPACT_WINDOW_MAX + "] — Claude Code's --autocompact flag accepts only "
+ "that band of tokens; omit the key to leave auto-compaction at each backend's "
+ "own default.");
}
}
/** The peer kinds this build has an adapter for — {@link Profile#kind()}'s only valid values. */
private static final Set<String> KNOWN_KINDS = Set.of(Profile.KIND_CLAUDE_CODE, Profile.KIND_OPENCODE);
@@ -1694,7 +1366,7 @@ public record FleetConfig(
*
* <p>{@link Profile}'s compact constructor only lower-cases {@code kind} and compares it against
* {@code KIND_OPENCODE} — anything else, including a typo like {@code opencod}, silently falls
* into the claude-code bucket ({@link dev.ltms.fleet.member.CompositePeerLauncher} routes by
* into the claude-code bucket ({@link dev.ltms.bridged.member.CompositePeerLauncher} routes by
* exact adapter claim, not by membership in a known set). With {@code argv:} also unset, the argv
* default special-cases only the exact string {@code "claude-code"}, so the launch command falls
* back to {@code List.of(kind)} — the daemon then tries to run a program literally named after the
@@ -1815,15 +1487,9 @@ public record FleetConfig(
}
}
/**
* The member-credential policies this build understands — {@link MemberCredentials#policy()}'s
* only valid values. Checked against the RAW yaml text (before {@link MemberCredentials}'s
* compact constructor normalizes {@code deny-list} onto {@code deny-by-default}), so the alias
* is listed explicitly.
*/
/** The member-credential policies this build understands — {@link MemberCredentials#policy()}'s only valid value. */
private static final Set<String> KNOWN_MEMBER_CREDENTIALS_POLICIES =
Set.of(MemberCredentials.POLICY_DENY_BY_DEFAULT, MemberCredentials.POLICY_DENY_LIST,
MemberCredentials.POLICY_ALLOW_LIST);
Set.of(MemberCredentials.POLICY_DENY_BY_DEFAULT);
/**
* Reject a {@code memberCredentials.policy} that is not {@link #KNOWN_MEMBER_CREDENTIALS_POLICIES}
@@ -1904,7 +1570,7 @@ public record FleetConfig(
}
/** Fill in nested defaults so callers never see nulls for structural fields. */
public FleetConfig withDefaults() {
public BridgedConfig withDefaults() {
Bind b = bind != null ? bind : new Bind(null, 0);
Guard g = guard != null ? guard : new Guard(List.of());
Lifecycle l = lifecycle != null ? lifecycle : new Lifecycle(null, null, null, false);
@@ -1935,17 +1601,12 @@ public record FleetConfig(
// guard/lifecycle, an absent block is not a safe "feature off" default here, it is a gap. It
// is deliberately not pre-populated with a Java-side name list (that would just reintroduce
// the hardcoded-list defect this record replaces); the block must be configured in
// fleetd.yaml to protect anything. See fleetd.example.yaml's memberCredentials: comment.
// CB-633: policy stays deny-by-default here — the allow-list scrub is opt-in, because it is
// stricter than today's behaviour (it blanks every non-derived name, not just known ones)
// and an upgrade must not change what a running deployment's members inherit.
// bridged.yaml to protect anything. See bridged.example.yaml's memberCredentials: comment.
MemberCredentials mc = memberCredentials != null ? memberCredentials
: new MemberCredentials(null, List.of(), List.of());
// coordinator is left as-is, like broker/primary above: null keeps no LeadMailbox opened,
// and this ticket's Coordinator is config-only anyway (nothing yet reads it at startup).
return new FleetConfig(b, herdrSocket, memberHerdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
return new BridgedConfig(b, herdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
quarantineCooldown, mc, coordinator);
quarantineCooldown, mc);
}
/**
@@ -1975,11 +1636,11 @@ public record FleetConfig(
/**
* Reject a lead-scan convention that a worker tab would also satisfy (CB-531).
*
* <p>The scan reads a tab label and concludes "a lead lives here". fleetd also <em>writes</em>
* <p>The scan reads a tab label and concludes "a lead lives here". bridged also <em>writes</em>
* tab labels — every member gets one rendered into its tab. Choose a lead {@code tabPrefix} that
* a member template matches and the daemon starts labelling its own members as leads, promoting
* the entire fleet to {@link dev.ltms.fleet.auth.Role#PRIMARY} with no message and no diff.
* The member-space exclusion in {@link dev.ltms.fleet.herdr.LeadTabScanner} already blocks the
* the entire fleet to {@link dev.ltms.bridged.auth.Role#PRIMARY} with no message and no diff.
* The member-space exclusion in {@link dev.ltms.bridged.herdr.LeadTabScanner} already blocks the
* realistic path, but defence that depends on one workspace label holding is not defence enough
* for a privilege boundary.
*
@@ -1,4 +1,4 @@
package dev.ltms.fleet.config;
package dev.ltms.bridged.config;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -16,7 +16,7 @@ import java.util.function.Supplier;
/**
* The daemon's live configuration, re-readable without a restart (CB-559).
*
* <p>Consumers hold this, not a {@link FleetConfig}, and read through {@link #get()} at the point
* <p>Consumers hold this, not a {@link BridgedConfig}, and read through {@link #get()} at the point
* of use. A component that captures {@code ref.get()} into a field at construction has opted out of
* reload — which is sometimes right (see <em>deferred</em> below), but it must then be a deliberate
* choice rather than an accident of where the field was initialised.
@@ -31,7 +31,7 @@ import java.util.function.Supplier;
* {@code placement:}, and an existing profile's {@code weight} / {@code maxLoad}. Those
* three are read through a supplier on {@code CompositePeerLauncher}, which is what makes
* them hot — not the fact that they are config. <strong>This does NOT include
* {@code fleet.leaders}</strong>: {@code Fleetd.main} reads {@code cfg.fleet().leaders()}
* {@code fleet.leaders}</strong>: {@code Bridged.main} reads {@code cfg.fleet().leaders()}
* once at startup to build the {@code LeadTabScanner} and the {@code LeadLauncher}, and
* neither is reconstructed on reload — so a lead added, removed, or re-{@code tab}'d under
* {@code fleet.leaders} needs a restart, the same as any deferred key below.</li>
@@ -42,7 +42,7 @@ import java.util.function.Supplier;
* {@code guard:}, {@code worktreeRoot:}, adding or removing a profile (a new backend needs its own launcher,
* which is constructed once), <em>and an existing profile's launch settings</em> —
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl},
* {@code exhaustedPattern} (CB-578 stage A — compiled once into {@code Fleetd.main}'s
* {@code exhaustedPattern} (CB-578 stage A — compiled once into {@code Bridged.main}'s
* pattern map at startup), and the rest. {@code credentialId} (CB-578 stage B) is NOT on
* this list — it is read live off the config supplier at every quarantine check and
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so it is hot instead.
@@ -65,30 +65,30 @@ import java.util.function.Supplier;
* running. A config file being edited is normally read once mid-save; degrading a working daemon
* because it caught a half-written file would be a bad trade.
*/
public final class ConfigRef implements Supplier<FleetConfig> {
public final class ConfigRef implements Supplier<BridgedConfig> {
private static final Logger log = LoggerFactory.getLogger(ConfigRef.class);
/** Keys that cannot change under a running daemon — see the class doc. */
private static final Set<String> COLD_KEYS =
Set.of("bind", "herdrSocket", "memberHerdrSocket", "broker", "auth");
Set.of("bind", "herdrSocket", "broker", "auth");
private final Path path;
private final AtomicReference<FleetConfig> current;
private final AtomicReference<BridgedConfig> current;
public ConfigRef(Path path, FleetConfig initial) {
public ConfigRef(Path path, BridgedConfig initial) {
this.path = path;
this.current = new AtomicReference<>(Objects.requireNonNull(initial, "initial config"));
}
/** A fixed reference that never reloads — for tests and for wiring built from a config in code. */
public static ConfigRef fixed(FleetConfig cfg) {
public static ConfigRef fixed(BridgedConfig cfg) {
return new ConfigRef(null, cfg);
}
/** The live configuration. Read this per use; do not cache it in a field. */
@Override
public FleetConfig get() {
public BridgedConfig get() {
return current.get();
}
@@ -128,7 +128,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
}
if (!applied) {
return "config reload refused — these keys cannot change under a running daemon: "
+ String.join(", ", coldKeys) + ". Restart fleetd to apply them.";
+ String.join(", ", coldKeys) + ". Restart bridged to apply them.";
}
if (!deferred.isEmpty()) {
return "config reloaded; these changes need a restart to take effect: "
@@ -149,10 +149,10 @@ public final class ConfigRef implements Supplier<FleetConfig> {
if (path == null) {
return Outcome.failed("this config was built in code and has no file to reload from");
}
FleetConfig old = current.get();
FleetConfig fresh;
BridgedConfig old = current.get();
BridgedConfig fresh;
try {
fresh = FleetConfig.load(path);
fresh = BridgedConfig.load(path);
// The same gate startup runs. A config that would have refused to boot must not be able
// to slip in through a reload — that is how a daemon ends up in a state it could never
// have started in, which is the hardest kind to debug.
@@ -182,7 +182,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
}
/** Cold keys whose value differs between the running config and the candidate. */
private static List<String> changedColdKeys(FleetConfig old, FleetConfig fresh) {
private static List<String> changedColdKeys(BridgedConfig old, BridgedConfig fresh) {
List<String> changed = new ArrayList<>();
if (!Objects.equals(old.bind(), fresh.bind())) {
changed.add("bind");
@@ -190,9 +190,6 @@ public final class ConfigRef implements Supplier<FleetConfig> {
if (!Objects.equals(old.herdrSocket(), fresh.herdrSocket())) {
changed.add("herdrSocket");
}
if (!Objects.equals(old.memberHerdrSocket(), fresh.memberHerdrSocket())) {
changed.add("memberHerdrSocket");
}
if (!Objects.equals(old.broker(), fresh.broker())) {
changed.add("broker");
}
@@ -205,7 +202,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
}
/** Changed keys that were accepted but whose effect waits for a restart. */
private static List<String> changedDeferredKeys(FleetConfig old, FleetConfig fresh) {
private static List<String> changedDeferredKeys(BridgedConfig old, BridgedConfig fresh) {
List<String> changed = new ArrayList<>();
if (!Objects.equals(old.lifecycle(), fresh.lifecycle())) {
changed.add("lifecycle");
@@ -229,9 +226,9 @@ public final class ConfigRef implements Supplier<FleetConfig> {
if (!Objects.equals(old.quarantineCooldownSeconds(), fresh.quarantineCooldownSeconds())) {
changed.add("quarantineCooldownSeconds");
}
Map<String, FleetConfig.Profile> before =
Map<String, BridgedConfig.Profile> before =
old.profiles() == null ? Map.of() : old.profiles();
Map<String, FleetConfig.Profile> after =
Map<String, BridgedConfig.Profile> after =
fresh.profiles() == null ? Map.of() : fresh.profiles();
// Adding or removing a profile is deferred: a new backend needs its own launcher, and
// launchers are built once at startup.
@@ -250,7 +247,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
// a reload can produce, because the operator has no reason to doubt it.
List<String> relaunch = new ArrayList<>();
before.forEach((name, was) -> {
FleetConfig.Profile now = after.get(name);
BridgedConfig.Profile now = after.get(name);
if (now != null && !sameLaunchSettings(was, now)) {
relaunch.add(name);
}
@@ -269,7 +266,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
* {@code CompositePeerLauncher}/the CB-578 stage B exhaustion sink) and really do take effect on
* the next spawn.
*/
private static boolean sameLaunchSettings(FleetConfig.Profile a, FleetConfig.Profile b) {
private static boolean sameLaunchSettings(BridgedConfig.Profile a, BridgedConfig.Profile b) {
return Objects.equals(a.baseUrl(), b.baseUrl())
&& Objects.equals(a.model(), b.model())
&& Objects.equals(a.configDir(), b.configDir())
@@ -279,9 +276,6 @@ public final class ConfigRef implements Supplier<FleetConfig> {
&& Objects.equals(a.workspace(), b.workspace())
&& Objects.equals(a.tabLabel(), b.tabLabel())
&& Objects.equals(a.mcpUrl(), b.mcpUrl())
// CB-634: the IDE MCP mount is a launch flag, fixed at spawn like mcpUrl — a
// reload changes it only for members spawned after, so a changed value is deferred.
&& Objects.equals(a.ideMcpUrl(), b.ideMcpUrl())
&& Objects.equals(a.cwd(), b.cwd())
&& Objects.equals(a.parityOverlay(), b.parityOverlay())
&& Objects.equals(a.gitTokenEnv(), b.gitTokenEnv())
@@ -289,7 +283,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
&& Objects.equals(a.kind(), b.kind())
&& Objects.equals(a.env(), b.env())
&& Objects.equals(a.subscription(), b.subscription())
// CB-578 stage B: exhaustedPattern is compiled once into Fleetd.main's pattern map
// CB-578 stage B: exhaustedPattern is compiled once into Bridged.main's pattern map
// at startup (see ExhaustedPatternLookup wiring) — a reload never re-reads it, so a
// changed pattern must be reported as deferred, exactly like model/baseUrl/argv.
&& Objects.equals(a.exhaustedPattern(), b.exhaustedPattern());
@@ -1,4 +1,4 @@
package dev.ltms.fleet.config;
package dev.ltms.bridged.config;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -11,7 +11,7 @@ import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
/**
* Polls {@code fleetd.yaml}'s modified time and asks {@link ConfigRef} to reload when it moves
* Polls {@code bridged.yaml}'s modified time and asks {@link ConfigRef} to reload when it moves
* (CB-559). Opt-in through {@code configReload.enabled}.
*
* <p><strong>Why polling and not a filesystem watch.</strong> {@code WatchService} on macOS has no
@@ -1,4 +1,4 @@
package dev.ltms.fleet.guard;
package dev.ltms.bridged.guard;
/** Thrown when the subscription boundary would be violated. Never swallow this. */
public class GuardException extends RuntimeException {
@@ -1,4 +1,4 @@
package dev.ltms.fleet.guard;
package dev.ltms.bridged.guard;
import java.net.URI;
import java.net.URISyntaxException;
@@ -16,7 +16,7 @@ import java.util.Set;
* means its traffic would leave the subscription. That is a hard stop.</li>
* </ul>
*
* Both checks throw {@link GuardException} on violation. {@code fleetd} calls
* Both checks throw {@link GuardException} on violation. {@code bridged} calls
* {@link #assertWorker} before spawning a worker and {@link #assertPrimaryClean}
* against its own environment at startup.
*/
@@ -1,7 +1,7 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.session.MemberSession;
/**
* Pure classifier. Collection and repair are deliberately outside this package.
@@ -1,11 +1,10 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.session.MemberSession;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -26,8 +25,6 @@ public final class FleetHealthMonitor {
/** Bounded attempts to run {@link #failTarget} for one transition. Never retried tick-to-tick (CB-580). */
static final int MAX_FAIL_TARGET_ATTEMPTS = 3;
// CB-641: Match the injector's 60s readiness gate so health allows a full first boot.
static final long READINESS_GRACE_NANOS = TimeUnit.SECONDS.toNanos(60);
private final AgentControl agents;
private final Supplier<List<MemberSession>> roster;
@@ -35,48 +32,29 @@ public final class FleetHealthMonitor {
private final ScheduledExecutorService scheduler;
private final LongSupplier clock;
private final long intervalSeconds;
private final long workingSuspectAfterNanos;
private final BiConsumer<String, String> failTarget;
private final Map<String, HealthPrior> priors = new HashMap<>();
private final Map<String, HealthState> states = new HashMap<>();
/**
* CB-643: consecutive ticks on which a target looked like an orphaned delegation. The fact
* {@link MessageService#hasOrphanedDelegation} reports is a true snapshot, but it can read true
* for one tick during an ordinary race — an async ticket exists before its virtual thread has
* reached {@code rendezvous.open()}, so for that instant nothing is accepted or queued behind
* it. {@code decide} maps the field straight to {@code DELEGATION_ORPHANED} with no cross-tick
* smoothing of its own, so a single racy read would log a fault that clears on the next tick.
* Requiring two consecutive observations costs one interval of latency on a real orphan and
* removes that false positive entirely.
*/
private final Map<String, Integer> orphanStreaks = new HashMap<>();
/** How many consecutive ticks a target must look orphaned before health reports it (CB-643). */
static final int ORPHAN_CONFIRM_TICKS = 2;
// CB-643: every HealthSnapshot field now carries real evidence. The NOT_YET_OBSERVED placeholder
// that stood in for 7 of the 12 is gone, and with it the reason 8 of the 9 fault states were
// unreachable — GONE and NEVER_READY included, which is what kept CB-580's failTarget from ever
// firing. Do not reintroduce a constant here: a field with no publisher is a dead state, and the
// tests pass either way, so nothing else will tell you.
// These facts need the evidence publishers introduced by later M4 units. They are not negatives.
private static final boolean NOT_YET_OBSERVED = false;
/**
* @param failTarget CB-568's idempotent target-wide failure operation (e.g. {@code messages::abandon}),
* invoked once when a member transitions into a terminal health state. Required —
* there is deliberately no defaulting overload; a caller that does not want the
* fail-tickets-on-terminal-health behavior must pass an explicit inert value (see
* {@code TestTurnTokens.inert} / {@code FleetMcp.CapacitySource.none()} for the pattern).
* {@code TestTurnTokens.inert} / {@code BridgeMcp.CapacitySource.none()} for the pattern).
*/
public FleetHealthMonitor(AgentControl agents, Supplier<List<MemberSession>> roster, MessageService messages,
ScheduledExecutorService scheduler, LongSupplier clock, long intervalSeconds,
long workingSuspectAfterSeconds, BiConsumer<String, String> failTarget) {
BiConsumer<String, String> failTarget) {
this.agents = agents;
this.roster = roster;
this.messages = messages;
this.scheduler = scheduler;
this.clock = clock;
this.intervalSeconds = intervalSeconds;
this.workingSuspectAfterNanos = TimeUnit.SECONDS.toNanos(workingSuspectAfterSeconds);
this.failTarget = Objects.requireNonNull(failTarget, "failTarget");
}
@@ -91,50 +69,28 @@ public final class FleetHealthMonitor {
// Package-private so tests can run one tick without waiting.
void tick() {
try {
List<Agent> agentsNow = agents.list(); // Exactly one list call for this complete observation.
List<MemberSession> rosterNow = roster.get(); // One in-memory roster snapshot for this tick.
List<Agent> agentsNow;
boolean controlLinkDown = false;
try {
agentsNow = agents.list(); // Exactly one list call for this complete observation.
} catch (HerdrException error) {
agentsNow = List.of();
controlLinkDown = true;
log.warn("fleet health control link unavailable; classifying roster", error);
}
Map<String, Agent> live = new HashMap<>();
for (Agent agent : agentsNow) live.put(agent.terminalId(), agent);
HashSet<String> current = new HashSet<>();
long nowNanos = clock.getAsLong();
for (MemberSession session : rosterNow) {
current.add(session.terminalId());
Agent agent = live.get(session.terminalId());
AgentStatus status = agent == null ? AgentStatus.UNKNOWN : agent.status();
boolean accepted = messages.hasAcceptedDelivery(session.terminalId());
boolean present = agent != null;
boolean targetNotFound = !controlLinkDown && !present
&& session.state() != MemberSession.State.SPAWNING;
boolean readinessGraceElapsed = nowNanos - session.spawnedAtNanos() >= READINESS_GRACE_NANOS;
boolean stalled = session.state() == MemberSession.State.BUSY
&& nowNanos - session.lastActivityAtNanos() >= workingSuspectAfterNanos;
// CB-643: the three message-layer facts CB-640 published. Read them here rather than
// leaving them false — that constant is what made 8 of the 9 fault states dead.
boolean queuedDelivery = messages.hasQueuedDelivery(session.terminalId());
boolean replyStranded = messages.hasStrandedReply(session.terminalId());
boolean orphanedDelegation = confirmOrphan(session.terminalId(),
messages.hasOrphanedDelegation(session.terminalId()));
HealthSnapshot snapshot = new HealthSnapshot(session.state(), status, accepted, queuedDelivery,
messages.hasInboxMessage(session.terminalId()), present, targetNotFound, controlLinkDown,
readinessGraceElapsed, orphanedDelegation, replyStranded, stalled);
HealthSnapshot snapshot = new HealthSnapshot(session.state(), status, accepted, NOT_YET_OBSERVED,
messages.hasInboxMessage(session.terminalId()), agent != null, NOT_YET_OBSERVED,
NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED);
HealthDecision decision = decide(snapshot, priors.getOrDefault(session.terminalId(), HealthPrior.NONE),
nowNanos);
clock.getAsLong());
priors.put(session.terminalId(), decision.prior());
reportTransition(session.terminalId(), decision.state());
}
priors.keySet().retainAll(current);
states.keySet().retainAll(current);
orphanStreaks.keySet().retainAll(current);
} catch (Throwable error) {
// Any unclassified collection failure must never kill the monitor's only scheduler task.
// A list failure is health evidence, and must never kill the monitor's only scheduler task.
log.warn("fleet health collection failed; will retry next tick", error);
} finally {
if (!scheduler.isShutdown()) {
@@ -143,20 +99,6 @@ public final class FleetHealthMonitor {
}
}
/**
* Debounce {@link MessageService#hasOrphanedDelegation} across ticks (CB-643). Returns true only
* once {@code observed} has held for {@link #ORPHAN_CONFIRM_TICKS} consecutive ticks; a single
* false reading resets the streak, so a transient race never reaches the classifier.
*/
private boolean confirmOrphan(String target, boolean observed) {
if (!observed) {
orphanStreaks.remove(target);
return false;
}
int streak = orphanStreaks.merge(target, 1, Integer::sum);
return streak >= ORPHAN_CONFIRM_TICKS;
}
void reportTransition(String target, HealthState next) {
HealthState previous = states.put(target, next);
if (previous == next) return;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
/** Classification plus the private fact that the next pure decision needs. */
public record HealthDecision(HealthState state, HealthPrior prior) { }
@@ -1,4 +1,4 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
/** Private cross-tick observation. It is deliberately not a reported health value. */
public record HealthPrior(boolean busyButDone) {
@@ -1,7 +1,7 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.session.MemberSession;
/** Read-only facts from one fleet collection tick. */
public record HealthSnapshot(MemberSession.State sessionState, AgentStatus liveStatus,
@@ -1,4 +1,4 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
/** Health classifications reported for a member. */
public enum HealthState {
@@ -1,12 +1,12 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.bridged.msg.Rendezvous;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
/**
* Counts turns that ended via the completion fallback instead of {@code fleet_reply}.
* Counts turns that ended via the completion fallback instead of {@code bridge_reply}.
* MUTE is an observation by target and profile, not a classifier state and never suppresses faults.
*/
public final class MuteCounter {
@@ -1,4 +1,4 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
import java.util.ArrayList;
import java.util.HashMap;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -38,11 +38,6 @@ public final class AgentControl {
this.herdr = herdr;
}
/** The herdr daemon this control object sends its agent calls to. */
public HerdrClient herdr() {
return herdr;
}
/** One agent-targeted call, translating a terminal id to its pane id (retrying once fresh). */
private JsonNode agentCall(String method, String target, Map<String, Object> extra) {
String resolved = resolveTarget(target);
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
/**
* A herdr agent's lifecycle state, as reported by {@code agent_status}. Drives the
@@ -32,7 +32,7 @@ public enum AgentStatus {
};
}
/** Whether {@code fleetd} may inject a message now without stepping on a live turn. */
/** Whether {@code bridged} may inject a message now without stepping on a live turn. */
public boolean injectable() {
return this == IDLE || this == BLOCKED || this == DONE;
}
@@ -1,17 +1,17 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
/**
* Client face onto the herdr daemon (protocol 14, herdr 0.7.0).
*
* <p>This is the ONLY thing in {@code fleetd} that speaks to herdr. Every method
* <p>This is the ONLY thing in {@code bridged} that speaks to herdr. Every method
* maps to a herdr JSON-RPC call over its Unix domain socket. Requests are
* newline-delimited JSON with a <em>string</em> id; responses carry either a
* {@code result} object (whose {@code type} field discriminates the payload) or an
* {@code error} object.
*
* <p>Higher layers ({@code fleetd}'s policy brain, REST endpoints, MCP adapters)
* <p>Higher layers ({@code bridged}'s policy brain, REST endpoints, MCP adapters)
* depend on this interface, not on the socket. Tests substitute a fake; the
* {@code contract}-tagged suite exercises the real implementation against a running
* herdr to catch protocol drift.
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.JsonNode;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
/**
* Raised when a herdr call fails: transport error, or an {@code error} envelope
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.slf4j.Logger;
@@ -14,7 +14,7 @@ import java.util.function.Supplier;
/**
* Discovers which panes host a lead by scanning herdr for tabs the operator labelled by convention
* (CB-531), and hands {@link dev.ltms.fleet.auth.CallerResolver} the resulting
* (CB-531), and hands {@link dev.ltms.bridged.auth.CallerResolver} the resulting
* {@code terminal_id → lead name} map.
*
* <p><strong>Why scan at all.</strong> A lead is never spawned — a human opens a tab and starts an
@@ -39,18 +39,18 @@ import java.util.function.Supplier;
* <li>Worker spaces are excluded wholesale ({@code excludedWorkspaceLabels}), so a worker cannot
* become a lead by being placed — as a split, say — inside a matching tab.</li>
* <li>A worker cannot rename a tab: {@code tab.rename} is reachable only through
* {@link WorkspaceControl}, which no {@code fleet_*} tool exposes. The label is writable by
* {@link WorkspaceControl}, which no {@code bridge_*} tool exposes. The label is writable by
* the human at the terminal and by nobody the bridge is defending against.</li>
* <li>The label is a <em>name</em>, not a capability. What a pane may do is decided by
* {@code Authz} against the role {@code CallerResolver} returns; a tab that calls itself a
* lead still cannot act as one unless the daemon's own registry agrees.</li>
* </ol>
*
* <p><strong>CB-558 — fleetd now writes lead labels too.</strong> This class used to be able to say
* that fleetd never renames a lead tab, so the label was always the human's own writing and there
* <p><strong>CB-558 — bridged now writes lead labels too.</strong> This class used to be able to say
* that bridged never renames a lead tab, so the label was always the human's own writing and there
* was no round-trip from the daemon's rename back into its next decision.
* {@code dev.ltms.fleet.lead.LeadLauncher} ends that: an auto-launched lead is labelled by the
* daemon and found again by this scan. The trust direction above is unaffected — fleetd writing a
* {@code dev.ltms.bridged.lead.LeadLauncher} ends that: an auto-launched lead is labelled by the
* daemon and found again by this scan. The trust direction above is unaffected — bridged writing a
* name for a lead it just started is not a pane promoting itself — but <em>staleness</em> becomes
* real: a label left behind by a session that has since died would read as a live lead forever.
* This scanner does not solve that (its job is naming, and a stale name costs nothing here); the
@@ -58,7 +58,7 @@ import java.util.function.Supplier;
* ever make a decision that <em>removes</em> something based on this map, add the same check.
* The remaining hazard is an <em>operator</em> one — a worker {@code tabLabel} template that
* happens to start with the same prefix would promote the whole fleet — and that is refused at
* startup by {@code FleetConfig.validateLeadTabPrefixes} rather than documented here.
* startup by {@code BridgedConfig.validateLeadTabPrefixes} rather than documented here.
*
* <p><strong>Caching.</strong> {@link #get()} is on the request path (every resolve), so the scan
* is TTL-cached and a stale-but-valid map is preferred to a herdr round-trip. A failed scan keeps
@@ -0,0 +1,59 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import java.util.Map;
/**
* Resolves which herdr pane a process belongs to — the herdr half of connection-based MCP
* identity (CB-105). Given the PID that opened an MCP connection, {@link #terminalForPid} finds
* the agent pane whose process tree contains it, so {@code bridged} can tell <em>which worker</em>
* is calling without the worker sending anything spoofable.
*
* <p>herdr owns the PID→pane truth: {@code pane.process_info} reports each pane's {@code shell_pid}
* and foreground process PIDs. This scans agent panes; a spawn-time {@code pid→terminal} cache is
* the obvious optimization once wired into {@code ClaudeCodeLauncher}.
*/
public final class PaneLocator {
private final HerdrClient herdr;
public PaneLocator(HerdrClient herdr) {
this.herdr = herdr;
}
/**
* The {@code terminal_id} of the agent pane whose process tree contains {@code pid}, or
* {@code null} if no agent pane owns it (e.g. the caller is the primary, or off-host).
*/
public String terminalForPid(long pid) {
if (pid <= 0) {
return null;
}
for (JsonNode pane : herdr.call("pane.list", Map.of()).path("panes")) {
String paneId = pane.path("pane_id").asText(null);
if (paneId != null && paneOwnsPid(paneId, pid)) {
return pane.path("terminal_id").asText(null);
}
}
return null;
}
private boolean paneOwnsPid(String paneId, long pid) {
JsonNode info;
try {
info = herdr.call("pane.process_info", Map.of("pane_id", paneId)).path("process_info");
} catch (HerdrException e) {
return false; // pane vanished mid-scan — just skip it
}
if (info.path("shell_pid").asLong(-1) == pid) {
return true;
}
for (JsonNode p : info.path("foreground_processes")) {
if (p.path("pid").asLong(-1) == pid) {
return true;
}
}
return false;
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -7,7 +7,7 @@ import com.fasterxml.jackson.databind.JsonNode;
* dedicated worker space so they never split or clutter the user's real work spaces.
*
* @param workspaceId herdr's stable id (e.g. {@code "w4"})
* @param label display label shown in herdr's UI (e.g. {@code "fleetd-workers"})
* @param label display label shown in herdr's UI (e.g. {@code "bridged-workers"})
* @param activeTabId the workspace's currently-focused tab, or {@code null}
*/
public record Workspace(String workspaceId, String label, String activeTabId) {
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.slf4j.Logger;
@@ -1,31 +1,29 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.TurnToken;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.TurnToken;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.time.Duration;
import java.util.List;
import java.util.Objects;
import java.util.Set;
import java.util.TreeSet;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.LongSupplier;
import java.util.regex.Pattern;
/**
* The CB-106 completion fallback: bridges the {@link Injector}'s turn-completion signal to the
* {@link Rendezvous} so a blocking {@code fleet_send} resolves even when the worker finishes its
* task without ever calling {@code fleet_reply} — the common case for a real delegated coding task.
* {@link Rendezvous} so a blocking {@code bridge_send} resolves even when the worker finishes its
* task without ever calling {@code bridge_reply} — the common case for a real delegated coding task.
*
* <p>On a confirmed {@code working → idle} boundary it scrapes the worker's recent transcript and
* resolves the awaiting send with that tail (a {@link Rendezvous.Kind#COMPLETION} resolution, so the
* caller can tell a scrape from a structured reply). It scrapes only when a send is actually waiting
* — a fleet worker's own turns, or a send that already timed out, cost no herdr traffic. An explicit
* {@code fleet_reply} that raced in first wins; {@link Rendezvous#resolveCompletion} is then a no-op.
* {@code bridge_reply} that raced in first wins; {@link Rendezvous#resolveCompletion} is then a no-op.
*
* <p>It also handles the CB-109 stall signal ({@link #onTurnFailed}): a worker that ran a turn then
* wedged in an {@code unknown} state resolves the send as a failure (with the error screen as
@@ -40,7 +38,7 @@ import java.util.regex.Pattern;
* <p><strong>Waiter-specific resolution (CB-116).</strong> On delivery we also capture the exact
* {@link Rendezvous} waiter this turn belongs to, and the completion/failure fallbacks resolve
* <em>that</em> waiter — never "whatever send is waiting now". A completion fallback runs on a virtual
* thread and can land after the worker's {@code fleet_reply} already resolved the turn and the
* thread and can land after the worker's {@code bridge_reply} already resolved the turn and the
* <em>next</em> send opened its own waiter on the same session; resolving the current waiter would
* then deliver turn N's stale scrape as turn N+1's answer. Targeting the captured waiter makes a late
* completion a harmless no-op (its waiter is already done) instead of a cross-turn stale reply.
@@ -63,33 +61,13 @@ public final class CompletionResolver implements TurnListener {
/** Cap the scraped tail so a long transcript can't return an unbounded blob. */
static final int MAX_SCRAPE_CHARS = 4000;
/**
* fleetd#164: the floor below which a {@code BUSY -> DONE} transition cannot be real work. A
* backend that rejects a turn outright (e.g. an HTTP 400 from the model, before the worker read
* a single file or produced a token) drives the exact same confirmed {@code working -> idle}
* transition a genuine completion does — just in about a second instead of the many seconds a
* real turn costs. {@link #onTurnComplete} cannot tell those two cases apart from the transition
* alone, so a turn that settles inside this floor is treated as a crash signature and resolved
* as a failure, never as a (possibly empty) success.
*/
public static final long MIN_TURN_NANOS = Duration.ofSeconds(2).toNanos();
/**
* fleetd#164 (part 2): one stable, explicit backend-failure marker seen on a Claude Code pane
* when the backend itself rejected the turn (e.g. {@code "API Error: 400 invalid request body"}).
* Kept deliberately narrow — a growing list of ad-hoc error strings rots as backends change their
* wording; broader backend-error surfacing is out of scope here (fleetd#164 point 3).
*/
private static final Pattern BACKEND_ERROR = Pattern.compile("(?i)\\bAPI Error\\s*:");
private static final String CLIPPED_PANE_TAIL_MARKER =
"[Pane tail clipped: member did not call fleet_reply.]";
"[Pane tail clipped: member did not call bridge_reply.]";
private final AgentControl agents;
private final Rendezvous rendezvous;
private final ExhaustedPatternLookup exhaustedPatterns;
private final ExhaustionSink exhaustionSink;
private final LongSupplier nowNanos;
/**
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
@@ -101,21 +79,8 @@ public final class CompletionResolver implements TurnListener {
* reference: a completion scrape equal to it means the worker produced no new output (the previous
* turn's wind-down sampled as this boundary), so it is suppressed. Overwritten on each delivery;
* cleared when the turn resolves. Package-private so tests can capture and replay a specific turn.
*
* <p>{@code deliveredAtNanos} (fleetd#164) is the {@link #nowNanos} reading taken at delivery —
* the other half of the {@link #MIN_TURN_NANOS} floor check, compared against a fresh reading at
* resolution time.
*/
record InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline, long deliveredAtNanos) {
/**
* Convenience for tests exercising scrape/suppression logic that don't care about turn
* timing: back-dates the delivery far enough that {@link #MIN_TURN_NANOS} can never fire.
* Not used by production code — {@link #captureBaseline} always records a real reading.
*/
InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline) {
this(waiter, baseline, Long.MIN_VALUE / 2);
}
record InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline) {
}
private final ConcurrentHashMap<String, InFlight> inFlight = new ConcurrentHashMap<>();
@@ -132,25 +97,10 @@ public final class CompletionResolver implements TurnListener {
*/
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
ExhaustionSink exhaustionSink) {
this(agents, rendezvous, exhaustedPatterns, exhaustionSink, System::nanoTime);
}
/**
* Test constructor with an injectable clock (fleetd#164), matching the {@code LongSupplier}
* pattern {@link dev.ltms.fleet.session.SessionManager} and {@link dev.ltms.fleet.msg.MessageService}
* already use: lets a test place a turn's delivery and its resolution at an exact, controllable
* distance apart around the {@link #MIN_TURN_NANOS} floor, without a real sleep. Public (rather
* than package-private like those two) because callers that wire a full {@code MessageService}
* fixture — e.g. {@code MessageServiceTest} — construct this resolver directly from another
* package.
*/
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
ExhaustionSink exhaustionSink, LongSupplier nowNanos) {
this.agents = agents;
this.rendezvous = rendezvous;
this.exhaustedPatterns = Objects.requireNonNull(exhaustedPatterns, "exhaustedPatterns");
this.exhaustionSink = Objects.requireNonNull(exhaustionSink, "exhaustionSink");
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
}
@Override
@@ -180,7 +130,7 @@ public final class CompletionResolver implements TurnListener {
baseline = null; // fail open: no baseline ⇒ no suppression
log.debug("delivery baseline for {} failed: {}", target, e.getMessage());
}
inFlight.put(target, new InFlight(waiter, baseline, nowNanos.getAsLong()));
inFlight.put(target, new InFlight(waiter, baseline));
}
/** The turn currently baselined for {@code target}, or {@code null} — a test hook for the captureBaseline path. */
@@ -221,20 +171,11 @@ public final class CompletionResolver implements TurnListener {
void resolve(String target, InFlight turn) {
CompletableFuture<Rendezvous.Resolution> waiter = turn == null ? null : turn.waiter();
if (waiter == null || waiter.isDone()) {
// Nobody is blocked on THIS turn (it had no send, or its fleet_reply already won). Skip
// Nobody is blocked on THIS turn (it had no send, or its bridge_reply already won). Skip
// the scrape; resolving the current waiter here would be the CB-116 cross-turn stale reply.
inFlight.remove(target, turn);
return;
}
// fleetd#164: a BUSY -> DONE transition inside the floor cannot be real work — it's a crash
// signature (e.g. a backend HTTP 400 before the worker did anything), not a fast answer. Fail
// it before spending a scrape on the ordinary path; the reason still carries whatever is on
// screen, since that is usually the backend's own error.
long elapsedNanos = nowNanos.getAsLong() - turn.deliveredAtNanos();
if (elapsedNanos < MIN_TURN_NANOS) {
fail(target, turn, tooFastReason(target, elapsedNanos));
return;
}
String tail;
String assistantBlock = null;
int originalLength = 0;
@@ -246,66 +187,49 @@ public final class CompletionResolver implements TurnListener {
clipped = originalLength > MAX_SCRAPE_CHARS;
tail = clip(assistantBlock);
} catch (RuntimeException e) {
log.warn("completion scrape for {} failed: {}", target, e.getMessage());
// The worker finished but we couldn't read its screen — still resolve the send so the
// caller unblocks; an empty tail beats hanging until the caller's timeout.
log.warn("completion scrape for {} failed; resolving with an empty tail: {}",
target, e.getMessage());
tail = "";
scrapeFailed = true;
}
// fleetd#164: a scrape nobody could read, and a scrape that read cleanly but produced nothing,
// both used to resolve the send as a SUCCESS carrying "" — indistinguishable from a worker that
// genuinely finished with nothing to say. That is the defect: fail loudly instead, naming the
// member, so a caller (including a lead deciding whether to delegate again) can tell a lost
// turn from a real empty answer.
if (scrapeFailed || tail.isEmpty()) {
fail(target, turn, emptyScrapeReason(target, scrapeFailed));
return;
}
// Misattribution guard (CB-115): if the scrape is byte-identical to the pane content at
// delivery, this turn produced no new output — the boundary belongs to the previous turn's
// wind-down (common on rapid back-to-back sends). Suppress rather than resolve the send with
// a stale answer; the real fleet_reply (or a later genuine completion) resolves it instead.
// a stale answer; the real bridge_reply (or a later genuine completion) resolves it instead.
// A scrape that failed to read is exempt — an empty tail there is "couldn't see", not "no change".
String baseline = turn.baseline();
if (baseline != null && baseline.equals(tail)) {
if (!scrapeFailed && baseline != null && baseline.equals(tail)) {
log.debug("suppressing misattributed completion for {} (no output change since delivery)",
target);
return; // keep the in-flight record: a later genuine completion still needs it
}
// CB-578 stage A: a turn that ended with no fleet_reply AND whose scrape matches the
// CB-578 stage A: a turn that ended with no bridge_reply AND whose scrape matches the
// backend's configured usage-limit pattern is a refusal, not an answer. Classify it as
// BACKEND_EXHAUSTED rather than handing the caller a scrape that reads like a real reply.
Pattern exhausted = exhaustedPatterns.patternFor(target);
String matchedLine = exhausted == null ? null : firstMatchingLine(assistantBlock, exhausted);
if (matchedLine != null) {
String reason = "backend exhausted (usage limit): " + matchedLine;
if (rendezvous.resolveExhausted(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("completion for {} classified BACKEND_EXHAUSTED (no fleet_reply; scrape "
+ "matched the profile's exhausted pattern): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
if (!scrapeFailed) {
Pattern exhausted = exhaustedPatterns.patternFor(target);
String matchedLine = exhausted == null ? null : firstMatchingLine(assistantBlock, exhausted);
if (matchedLine != null) {
String reason = "backend exhausted (usage limit): " + matchedLine;
if (rendezvous.resolveExhausted(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("completion for {} classified BACKEND_EXHAUSTED (no bridge_reply; scrape "
+ "matched the profile's exhausted pattern): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
}
return;
}
return;
}
// fleetd#164 (part 2): a scrape that read cleanly and produced content still isn't a real
// reply when that content is the backend's own rejection (e.g. an HTTP 400 before the worker
// did any work). Classify it as a failure naming the member, rather than handing the caller a
// scrape that reads like a completed answer.
String backendError = firstMatchingLine(assistantBlock, BACKEND_ERROR);
if (backendError != null) {
// Carry the whole scrape, not just the matched line. The pattern is a heuristic: a member
// that forgot fleet_reply while reporting *about* a backend error matches it too. Failing
// is still right — the caller must not read a scrape as an answer — but dropping the rest
// of the pane would destroy the report, which is the same defect fleetd#164 is about.
fail(target, turn, "member " + target + " ended on a backend error: " + backendError
+ "\n--- pane tail ---\n" + tail);
return;
}
String completion = clipped ? tail + "\n" + CLIPPED_PANE_TAIL_MARKER : tail;
if (rendezvous.resolveCompletion(waiter, completion)) {
inFlight.remove(target, turn);
if (clipped) {
log.warn("completion scrape for {} clipped from {} chars to the {} char cap; "
+ "member did not call fleet_reply, so the pane tail is partial",
+ "member did not call bridge_reply, so the pane tail is partial",
target, originalLength, MAX_SCRAPE_CHARS);
}
log.debug("resolved send to {} via turn-completion fallback ({} chars scraped)",
@@ -348,36 +272,6 @@ public final class CompletionResolver implements TurnListener {
}
}
/**
* fleetd#164: the failure reason for a turn that settled inside {@link #MIN_TURN_NANOS} — names
* the member and both timings, and appends whatever the pane shows (usually the backend's own
* error) so the caller sees the cause, not just "it failed".
*/
private String tooFastReason(String target, long elapsedNanos) {
String scrape;
try {
scrape = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
scrape = "";
}
String reason = String.format(
"member %s went BUSY -> DONE in %dms (floor %dms) — too fast to be real work, most "
+ "likely a backend error before any work started",
target, elapsedNanos / 1_000_000, MIN_TURN_NANOS / 1_000_000);
return scrape.isBlank() ? reason : reason + ": " + scrape;
}
/**
* fleetd#164: the failure reason for a scrape that produced zero characters — names the member
* and says plainly that the turn produced nothing, so a caller (a lead deciding whether to
* delegate again included) never mistakes a lost turn for a genuinely empty reply.
*/
private static String emptyScrapeReason(String target, boolean scrapeFailed) {
return "member " + target + " turn completed with an empty scrape (0 chars) — "
+ (scrapeFailed ? "its pane could not be read; " : "")
+ "treating as a lost turn, not a real answer";
}
/**
* The first line of {@code text} matching {@code pattern}, stripped — the CB-578 stage A
* evidence carried in a {@code BACKEND_EXHAUSTED} reason so the operator sees the real refusal
@@ -395,7 +289,7 @@ public final class CompletionResolver implements TurnListener {
/**
* Coverage summary for the CB-578 stage A exhausted-pattern classification, logged at startup
* the way {@link dev.ltms.fleet.health.FleetHealthMonitor#coverage} is — so an operator can
* the way {@link dev.ltms.bridged.health.FleetHealthMonitor#coverage} is — so an operator can
* see whether the classification is on, and for which profiles, without reading every
* profile's config by hand.
*
@@ -1,4 +1,4 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import java.util.regex.Pattern;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
/**
* Notified when {@link CompletionResolver} actually delivers a {@code BACKEND_EXHAUSTED}
@@ -8,7 +8,7 @@ package dev.ltms.fleet.inject;
*
* <p>{@link CompletionResolver} knows only {@code target} (a herdr terminal id); it has no notion of
* profiles or credentials, so mapping {@code target} to whatever should be quarantined is entirely
* the sink's job — see {@code Fleetd.main}'s wiring, which resolves target → session → profile →
* the sink's job — see {@code Bridged.main}'s wiring, which resolves target → session → profile →
* {@code effectiveCredentialId()} and calls {@code BackendQuarantine.quarantine} on it.
*/
@FunctionalInterface
@@ -1,9 +1,8 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.msg.TurnToken;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.msg.TurnToken;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -82,7 +81,7 @@ public final class Injector {
/**
* The single source for the injector poll cadence — how often the {@link StatusPoller} drives
* {@link #onStatus} at. {@code Fleetd} passes this to every {@link StatusPoller} it constructs,
* {@link #onStatus} at. {@code Bridged} passes this to every {@link StatusPoller} it constructs,
* and this class reads it to state the readiness grace in seconds on the CB-562 expiry log
* instead of hardcoding "60s". One constant, so a cadence change cannot silently desync a log
* that claims a grace duration.
@@ -90,7 +89,6 @@ public final class Injector {
public static final long POLL_INTERVAL_MILLIS = 250;
private final AgentControl agents;
private final HerdrRouter router;
private final TurnListener turnListener;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
private final Consumer<String> forget; // CB-114: clear a gone worker's readiness/presence
@@ -126,25 +124,11 @@ public final class Injector {
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
this.agents = agents;
this.router = null;
this.turnListener = turnListener;
this.ready = ready;
this.forget = forget;
}
public Injector(HerdrRouter router, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
this.agents = null;
this.router = router;
this.turnListener = turnListener;
this.ready = ready;
this.forget = forget;
}
private AgentControl agentsFor(String target) {
return router != null ? router.agentsFor(target) : agents;
}
/** A pending message and the future that completes when it has been delivered. */
private record Pending(String text, TurnToken token, CompletableFuture<Void> delivered) {
}
@@ -269,7 +253,7 @@ public final class Injector {
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
try {
agentsFor(target).send(target, p.text());
agents.send(target, p.text());
t.queue.poll();
t.awaitingPickup = true;
t.awaitingCompletion = true;
@@ -329,7 +313,7 @@ public final class Injector {
// thread while it holds the target lock.
if (resubmit) {
try {
agentsFor(target).submit(target); // nudge a raced Enter so the pending paste submits
agents.submit(target); // nudge a raced Enter so the pending paste submits
} catch (RuntimeException e) {
log.debug("resubmit to {} failed (will retry next poll): {}", target, e.getMessage());
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import java.util.concurrent.ConcurrentHashMap;
import java.util.Set;
@@ -1,9 +1,8 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.HerdrException;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -23,7 +22,6 @@ public final class StatusPoller {
private static final Logger log = LoggerFactory.getLogger(StatusPoller.class);
private final AgentControl agents;
private final HerdrRouter router;
private final Injector injector;
private final StatusRefiner refiner;
private final long intervalMillis;
@@ -37,24 +35,11 @@ public final class StatusPoller {
public StatusPoller(AgentControl agents, Injector injector, StatusRefiner refiner,
long intervalMillis) {
this.agents = agents;
this.router = null;
this.injector = injector;
this.refiner = refiner;
this.intervalMillis = intervalMillis;
}
public StatusPoller(HerdrRouter router, Injector injector, long intervalMillis) {
this.agents = null;
this.router = router;
this.injector = injector;
// CB-185: this refiner's own AgentControl (member) is only a default for the legacy 2-arg
// refine() overload — the loop below always calls the 3-arg refine(target, raw, control)
// with the per-target control from router.agentsFor(target), so a lead target is refined
// against the LEAD daemon even though this field points at the member one.
this.refiner = new StatusRefiner(router.memberAgents());
this.intervalMillis = intervalMillis;
}
/** Start the polling loop on a virtual thread. Idempotent. */
public synchronized void start() {
if (running) return;
@@ -71,11 +56,7 @@ public final class StatusPoller {
try {
// herdr's agent_status can misreport a settled worker as `unknown`; refine it
// against the pane content before it drives delivery/completion (CB-115).
// CB-185: refine THROUGH the same control the raw status came from — a router
// splits lead/member targets across two herdr daemons, and reading a lead's pane
// through the (fixed) member refiner never finds it, wedging that lead at UNKNOWN.
AgentControl control = router != null ? router.agentsFor(target) : agents;
AgentStatus status = refiner.refine(target, control.status(target), control);
AgentStatus status = refiner.refine(target, agents.status(target));
injector.onStatus(target, status);
} catch (HerdrException e) {
// The worker's agent is gone — stop trying and unblock its waiters.
@@ -1,7 +1,7 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -41,32 +41,15 @@ public final class StatusRefiner {
}
/**
* Return a trustworthy status for {@code target}, reading its pane through this refiner's own
* {@link AgentControl}. Equivalent to {@link #refine(String, AgentStatus, AgentControl)} with
* that control — kept for callers that only ever talk to one herdr daemon.
* Return a trustworthy status for {@code target}. Any non-{@code UNKNOWN} {@code raw} is returned
* unchanged; an {@code UNKNOWN} triggers a pane read and content classification. A read failure
* leaves it {@code UNKNOWN} (the safe default: no delivery, and the stall path still applies).
*/
public AgentStatus refine(String target, AgentStatus raw) {
return refine(target, raw, agents);
}
/**
* Return a trustworthy status for {@code target}. Any non-{@code UNKNOWN} {@code raw} is returned
* unchanged; an {@code UNKNOWN} triggers a pane read (through {@code control}) and content
* classification. A read failure leaves it {@code UNKNOWN} (the safe default: no delivery, and
* the stall path still applies).
*
* <p>CB-185: {@code control} must be the {@link AgentControl} for the <em>same</em> daemon the
* raw status was sampled from — a router splits lead and member targets across two herdr
* daemons, and reading a lead's pane through the member client (or vice versa) fails to find
* the pane and leaves the target wedged at {@code UNKNOWN} forever. Callers that route per
* target (e.g. {@code StatusPoller}) must pass that target's control explicitly rather than
* relying on the control fixed at construction.
*/
public AgentStatus refine(String target, AgentStatus raw, AgentControl control) {
if (raw != AgentStatus.UNKNOWN) return raw;
String pane;
try {
pane = control.read(target, PROBE_SOURCE);
pane = agents.read(target, PROBE_SOURCE);
} catch (RuntimeException e) {
log.debug("status refine read for {} failed; leaving UNKNOWN: {}", target, e.getMessage());
return AgentStatus.UNKNOWN;
@@ -1,12 +1,12 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import dev.ltms.fleet.msg.TurnToken;
import dev.ltms.bridged.msg.TurnToken;
/**
* Notified when a worker's delegated turn is observed to complete — a confirmed
* {@code WORKING → IDLE} transition after a delivery. This is the CB-106 completion signal the
* {@code CompletionResolver} uses to resolve a blocked send whose worker never called
* {@code fleet_reply}. Kept as a seam so the {@link Injector} needs no dependency on the message
* {@code bridge_reply}. Kept as a seam so the {@link Injector} needs no dependency on the message
* layer and stays unit-testable with a capturing fake.
*/
@FunctionalInterface
@@ -1,12 +1,12 @@
package dev.ltms.fleet.lead;
package dev.ltms.bridged.lead;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -15,7 +15,8 @@ import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import dev.ltms.fleet.peer.PeerLauncher;
import java.util.Set;
import java.util.stream.Collectors;
/**
* Starts the leads {@code fleet.leaders:} declares, when none is already running (CB-558).
@@ -25,7 +26,7 @@ import dev.ltms.fleet.peer.PeerLauncher;
* must never receive:
* <ol>
* <li>it appends the <em>reply charter</em> — "you are an off-subscription worker … end every turn
* with {@code fleet_reply}". A lead is the orchestrator; telling it that it is a worker is
* with {@code bridge_reply}". A lead is the orchestrator; telling it that it is a worker is
* exactly backwards.</li>
* <li>it registers the session with {@code SessionManager}, which subjects it to the idle reaper,
* the context cap and the shutdown drain. An idle lead is the normal state of a lead, so the
@@ -39,7 +40,7 @@ import dev.ltms.fleet.peer.PeerLauncher;
* broken later, quietly.
*
* <p><strong>Liveness, and the label round-trip.</strong> {@code LeadTabScanner} used to be able to
* promise that fleetd never writes a lead label. That is no longer true — an auto-launched lead is
* promise that bridged never writes a lead label. That is no longer true — an auto-launched lead is
* labelled by this class, and the scanner reads that label back. The risk this opens is not
* privilege escalation (the tab label never granted anything a pane could take for itself; see that
* class's javadoc), but <em>staleness</em>: a label left behind by a crashed session would otherwise
@@ -53,14 +54,14 @@ public final class LeadLauncher {
private final AgentControl agents;
private final WorkspaceControl spaces;
private final FleetConfig cfg;
private final BridgedConfig cfg;
/**
* @param agents herdr agent control (start, list)
* @param spaces workspace / tab control (ensure, create, label, list)
* @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and each lead's tab
*/
public LeadLauncher(AgentControl agents, WorkspaceControl spaces, FleetConfig cfg) {
public LeadLauncher(AgentControl agents, WorkspaceControl spaces, BridgedConfig cfg) {
this.agents = agents;
this.spaces = spaces;
this.cfg = cfg;
@@ -74,7 +75,7 @@ public final class LeadLauncher {
* the remaining leads are still attempted.
*/
public int ensureLeads() {
Map<String, FleetConfig.Leader> leaders = cfg.fleet().leaders();
Map<String, BridgedConfig.Leader> leaders = cfg.fleet().leaders();
if (leaders.isEmpty()) {
return 0;
}
@@ -90,9 +91,9 @@ public final class LeadLauncher {
}
int started = 0;
for (Map.Entry<String, FleetConfig.Leader> e : leaders.entrySet()) {
for (Map.Entry<String, BridgedConfig.Leader> e : leaders.entrySet()) {
String name = e.getKey();
FleetConfig.Leader lead = e.getValue();
BridgedConfig.Leader lead = e.getValue();
int running = live.getOrDefault(name, 0);
int wanted = lead.instances();
@@ -104,12 +105,12 @@ public final class LeadLauncher {
// A lead with a `tab:` but no `profile:` is recognise-only by design: the operator
// opens it by hand. Say so once rather than looking like a silent failure.
log.info("lead '{}' is not live, and names no profile — it can be recognised but not "
+ "launched. Add `profile:` under fleet.leaders.{} to have fleetd start it.",
+ "launched. Add `profile:` under fleet.leaders.{} to have bridged start it.",
name, name);
continue;
}
FleetConfig.Profile profile = cfg.profiles().get(lead.profile());
BridgedConfig.Profile profile = cfg.profiles().get(lead.profile());
if (profile == null) {
log.warn("lead '{}' names profile '{}', which is not configured — not launching",
name, lead.profile());
@@ -137,16 +138,16 @@ public final class LeadLauncher {
* already carries {@link Agent#tabId()} directly, so a hand-opened lead is found the same way an
* auto-launched one is — by labelling its tab to match.
*/
private Map<String, Integer> liveLeads(Map<String, FleetConfig.Leader> leaders) {
// A lead and the members share ONE workspace now (the operator asked for a single "session"
// with many tabs), so a workspace can no longer be excluded wholesale — the lead lives in the
// member workspace by design. The sole discriminator is the exact tab label: a lead carries
// its configured `fleet.leaders.<name>.tab` ("lead: opus"), while a member carries its
// profile's `worker: {profile} #{n}` template. These never collide, so an exact-label match
// separates them without needing to know which workspace anyone is in.
private Map<String, Integer> liveLeads(Map<String, BridgedConfig.Leader> leaders) {
Set<String> memberSpaces = cfg.profiles().values().stream()
.map(BridgedConfig.Profile::workspace)
.filter(w -> w != null && !w.isBlank())
.collect(Collectors.toSet());
// tabId → the lead name its label declares.
Map<String, String> nameByTab = new LinkedHashMap<>();
for (Workspace ws : spaces.listWorkspaces()) {
if (ws.workspaceId() == null) {
if (ws.workspaceId() == null || memberSpaces.contains(ws.label())) {
continue;
}
for (Tab tab : spaces.listTabs(ws.workspaceId())) {
@@ -173,12 +174,12 @@ public final class LeadLauncher {
* <p>Matched exactly (case-insensitively) against each lead's configured {@code tab}, so an
* operator's {@code "lead: something-else"} tab is not mistaken for a configured lead.
*/
private String leadNameOf(String label, Map<String, FleetConfig.Leader> leaders) {
private String leadNameOf(String label, Map<String, BridgedConfig.Leader> leaders) {
if (label == null) {
return null;
}
String l = label.strip();
for (Map.Entry<String, FleetConfig.Leader> e : leaders.entrySet()) {
for (Map.Entry<String, BridgedConfig.Leader> e : leaders.entrySet()) {
String tab = e.getValue().tabLabel();
if (tab != null && l.equalsIgnoreCase(tab.strip())) {
return e.getKey();
@@ -188,7 +189,7 @@ public final class LeadLauncher {
}
/** Start one lead. Returns false (having logged) rather than throwing on any failure. */
private boolean launch(String name, FleetConfig.Leader lead, FleetConfig.Profile profile) {
private boolean launch(String name, BridgedConfig.Leader lead, BridgedConfig.Profile profile) {
String label = lead.tabLabel();
String cwd = (lead.cwd() == null || lead.cwd().isBlank())
? System.getProperty("user.dir") : lead.cwd();
@@ -235,7 +236,7 @@ public final class LeadLauncher {
* The herdr agent kind for this profile — the same value the matching member adapter passes, so
* herdr resolves the same executable for a lead as it does for a member on that backend.
*/
private static String herdrKind(FleetConfig.Profile profile) {
private static String herdrKind(BridgedConfig.Profile profile) {
return profile.isOpenCode() ? "opencode" : "claude";
}
@@ -247,12 +248,11 @@ public final class LeadLauncher {
* any other primary. This is the single most important difference from the member launchers;
* do not "unify" it back.
*/
private List<String> leadArgv(FleetConfig.Profile profile) {
private List<String> leadArgv(BridgedConfig.Profile profile) {
List<String> argv = new ArrayList<>(profile.argv());
if (profile.hasMcp()) {
argv.add("--mcp-config");
argv.add("{\"mcpServers\":{\"" + PeerLauncher.MCP_MOUNT_NAME
+ "\":{\"type\":\"http\",\"url\":\""
argv.add("{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ profile.mcpUrl() + "\"}}}");
}
// Appended last, for the same reason the member launcher does it (CB-533): the argv is
@@ -274,7 +274,7 @@ public final class LeadLauncher {
* definition, so there is no configuration under which pointing it elsewhere is correct, and a
* profile that carries them (a member profile reused as a lead's backend) must not leak them in.
*/
private Map<String, String> leadEnv(FleetConfig.Profile profile) {
private Map<String, String> leadEnv(BridgedConfig.Profile profile) {
Map<String, String> out = new LinkedHashMap<>();
if (profile.env() != null) {
out.putAll(profile.env());
@@ -1,4 +1,4 @@
package dev.ltms.fleet.logging;
package dev.ltms.bridged.logging;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
@@ -1,28 +1,26 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
import dev.ltms.fleet.auth.AuditLog;
import dev.ltms.fleet.auth.Authz;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.auth.Role;
import dev.ltms.fleet.guard.GuardException;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.msg.LeadChannel;
import dev.ltms.fleet.msg.LeadMessage;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementException;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.WorktreeRequest;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.bridged.auth.AuditLog;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.auth.Role;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.placement.BackendQuarantine;
import dev.ltms.bridged.placement.PlacementException;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerLauncher;
import io.modelcontextprotocol.common.McpTransportContext;
import io.modelcontextprotocol.json.McpJsonMapper;
import io.modelcontextprotocol.json.jackson3.JacksonMcpJsonMapperSupplier;
@@ -39,38 +37,30 @@ import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.UUID;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.BiFunction;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
import java.util.stream.Collectors;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
/**
* The MCP SERVER face (CB-105): a Streamable-HTTP MCP server whose tools are <em>thin adapters</em>
* over the same {@link MessageService}/{@link Rendezvous} the REST routes use — so the two are
* validated by parity, not by re-implementing behaviour. The primary Opus calls {@code fleet_send}
* / {@code fleet_status}; the worker calls {@code fleet_reply}.
* validated by parity, not by re-implementing behaviour. The primary Opus calls {@code bridge_send}
* / {@code bridge_status}; the worker calls {@code bridge_reply}.
*
* <p>Beyond delegation the primary also manages the fleet here (CB-108): {@code fleet_spawn} /
* {@code fleet_list} / {@code fleet_stop} drive the {@link PeerLauncher} SPI so a worker's whole
* <p>Beyond delegation the primary also manages the fleet here (CB-108): {@code bridge_spawn} /
* {@code bridge_list} / {@code bridge_stop} drive the {@link PeerLauncher} SPI so a worker's whole
* lifecycle is managed through MCP, with each adapter's subscription boundary enforced inside it.
*
* <p>The tool <em>logic</em> lives in package-private static methods returning a
* {@link McpSchema.CallToolResult}, so it is unit-testable without standing up the HTTP transport;
* the SDK owns the wire protocol. Mount {@link #servlet()} at {@code /mcp} on the daemon's Jetty.
*/
public final class FleetMcp {
private static final Logger log = LoggerFactory.getLogger(FleetMcp.class);
public final class BridgeMcp {
private static final long DEFAULT_TIMEOUT_MS = 25_000;
private static final long MAX_TIMEOUT_MS = 120_000;
// fleet_ask blocks the WORKER's own MCP call, which its client caps near 60s — default under
// bridge_ask blocks the WORKER's own MCP call, which its client caps near 60s — default under
// that so the bridge returns a clean timeout before the client severs the call (CB-205).
private static final long ASK_DEFAULT_TIMEOUT_MS = 55_000;
private static final long ASK_MAX_TIMEOUT_MS = 115_000;
@@ -92,10 +82,8 @@ public final class FleetMcp {
private final CapacitySource capacity;
private final HealthCoverageSource healthCoverage;
private final QuarantineSource quarantine;
/** CB-637: this daemon's lead-to-lead channel; {@code null} when no coordinator is configured. */
private final LeadChannel leadChannel;
/** Capacity facts used by {@code fleet_list}; production must supply the placement live count. */
/** Capacity facts used by {@code bridge_list}; production must supply the placement live count. */
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
Supplier<Set<String>> configuredProfiles, LongSupplier clock) {
/** Inert test-only source. It omits capacity rather than inventing zero live counts. */
@@ -107,7 +95,7 @@ public final class FleetMcp {
public record HealthCoverageSource(Supplier<String> value) { }
/**
* CB-578 stage B quarantine facts used by {@code fleet_profiles}: a profile → credential id
* CB-578 stage B quarantine facts used by {@code bridge_profiles}: a profile → credential id
* lookup, plus the shared {@link BackendQuarantine} to read remaining cooldowns off.
*/
public record QuarantineSource(Function<String, String> credentialIdFor, BackendQuarantine quarantine) {
@@ -121,28 +109,13 @@ public final class FleetMcp {
* Jetty's context handler and never passes through Javalin's {@code before}
* filter, so the REST guard does not cover it.
* @param metrics registry for auth-failure counting; may be {@code null}
* @param quarantine CB-578 stage B facts for {@code fleet_profiles}; required — pass
* @param quarantine CB-578 stage B facts for {@code bridge_profiles}; required — pass
* {@link QuarantineSource#none()} for a caller that does not want the feature
*/
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
public BridgeMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine) {
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
healthCoverage, quarantine, null);
}
/**
* As above, with this daemon's lead-to-lead channel (CB-637). {@code leadChannel} is
* {@code null} whenever no {@code coordinator:} block is configured or its broker could not be
* reached at boot — cross-daemon lead messaging is simply off, and {@code fleet_send{coordId}}
* says so rather than failing obscurely.
*/
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, LeadChannel leadChannel) {
this.leadChannel = leadChannel;
this.capacity = capacity;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
this.healthCoverage = healthCoverage;
@@ -151,7 +124,7 @@ public final class FleetMcp {
.jsonMapper(json)
.mcpEndpoint("/mcp")
// Resolve the caller from the connection (peer PID → herdr pane) in one lookup: the
// worker terminal for fleet_reply (no spoofable arg), and the PID so fleet_spawn can
// worker terminal for bridge_reply (no spoofable arg), and the PID so bridge_spawn can
// inherit the primary's cwd (CB-112). Any contact from a worker marks it available
// (CB-113) — its MCP initialize is the reliable "the agent is up" signal.
.contextExtractor(req -> {
@@ -172,9 +145,10 @@ public final class FleetMcp {
CALLER_NAME, orEmpty(p.name())));
})
.build();
// Each handler is built once and wired to its fleet_* tool below.
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> sendHandler =
(exchange, req) -> {
this.server = McpServer.sync(transport)
.serverInfo("bridge", "0.1.0")
.capabilities(McpSchema.ServerCapabilities.builder().tools(true).build())
.toolCall(sendTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SEND,
str(req.arguments(), "sessionId"));
if (denied != null) return denied;
@@ -188,15 +162,8 @@ public final class FleetMcp {
String target = str(a, "sessionId");
String content = str(a, "content");
String turnId = str(a, "turnId");
String coordId = str(a, "coordId");
if (coordId != null && !coordId.isBlank()) {
// CB-637: a peer LEAD on another daemon, addressed by coord-id over the shared
// coordination broker. Checked before the turnId branch so a call that sets both
// is rejected as the conflict it is, rather than silently taking one route.
return sendToLead(leadChannel, coordId, content, target, turnId);
}
if (turnId != null && !turnId.isBlank()) {
// Answering a worker's fleet_ask (CB-205): resolve its blocked question and
// Answering a worker's bridge_ask (CB-205): resolve its blocked question and
// block for the worker's reply as it resumes the same turn. This is the same
// delegation, so ownership is left untouched (CB-548) — never re-recorded.
return answer(messages, turnId, content, timeoutMs(a));
@@ -211,49 +178,43 @@ public final class FleetMcp {
return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, target, content, onAccepted, workers.profiles())
: send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles());
};
// fleet_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> replyHandler =
(exchange, req) -> {
})
// bridge_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
.toolCall(replyTool(), (exchange, req) -> {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.REPLY, self);
if (denied != null) return denied;
return reply(messages, self, str(req.arguments(), "content"));
};
// fleet_ask (CB-205): a worker's mid-turn question — identity from the CONNECTION.
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> askHandler =
(exchange, req) -> {
})
// bridge_ask (CB-205): a worker's mid-turn question — identity from the CONNECTION.
.toolCall(askTool(), (exchange, req) -> {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.ASK, self);
if (denied != null) return denied;
return ask(messages, self, str(req.arguments(), "question"), timeoutMs(req.arguments()));
};
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> statusHandler =
(exchange, req) -> {
})
.toolCall(statusTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return status(messages, str(req.arguments(), "sessionId"));
};
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> pollHandler =
(exchange, req) -> {
})
.toolCall(pollTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
Map<String, Object> a = req.arguments();
return poll(messages, str(a, "ticket"), str(a, "target"));
};
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
// Acking removes a reply from the inbox, so it is a drain, not a read.
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> ackHandler =
(exchange, req) -> {
})
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
// Acking removes a reply from the inbox, so it is a drain, not a read.
.toolCall(ackTool(), (exchange, req) -> {
Map<String, Object> a = req.arguments();
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.DRAIN, str(a, "target"));
if (denied != null) return denied;
return ack(messages, str(a, "target"), str(a, "msgId"));
};
// Fleet management (CB-108): spawn/list/stop over ClaudeCodeLauncher.
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> spawnHandler =
(exchange, req) -> {
})
// Fleet management (CB-108): spawn/list/stop over ClaudeCodeLauncher.
.toolCall(spawnTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SPAWN, null);
if (denied != null) return denied;
String caller = callerTerminal(exchange);
@@ -268,62 +229,30 @@ public final class FleetMcp {
return spawn(sessions, str(a, "profile"), str(a, "role"), str(a, "cwd"), callerCwd,
callerTerminal(exchange), worktreeRequest(a),
str(a, "sessionName"), str(a, "resumeSessionId"));
};
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> listHandler =
(exchange, _) -> {
})
.toolCall(listTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine,
callers == null ? Map.of() : callers.leads(),
callerTerminal(exchange),
leadChannel == null ? null : leadChannel.selfCoordId());
};
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> stopHandler =
(exchange, req) -> {
callerTerminal(exchange));
})
.toolCall(stopTool(), (exchange, req) -> {
String paneId = str(req.arguments(), "paneId");
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.STOP, paneId);
if (denied != null) return denied;
return stop(sessions, paneId);
};
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> profilesHandler =
(exchange, _) -> {
})
.toolCall(profilesTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return profiles(workers, quarantine);
};
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> whoamiHandler =
(exchange, _) -> {
})
.toolCall(whoamiTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return whoami(principal(exchange), sessions);
};
McpSchema.Tool fleetSend = sendTool();
McpSchema.Tool fleetReply = replyTool();
McpSchema.Tool fleetAsk = askTool();
McpSchema.Tool fleetStatus = statusTool();
McpSchema.Tool fleetPoll = pollTool();
McpSchema.Tool fleetAck = ackTool();
McpSchema.Tool fleetSpawn = spawnTool();
McpSchema.Tool fleetList = listTool();
McpSchema.Tool fleetStop = stopTool();
McpSchema.Tool fleetProfiles = profilesTool();
McpSchema.Tool fleetWhoami = whoamiTool();
this.server = McpServer.sync(transport)
.serverInfo("fleet", "0.1.0")
.capabilities(McpSchema.ServerCapabilities.builder().tools(true).build())
.toolCall(fleetSend, sendHandler)
.toolCall(fleetReply, replyHandler)
.toolCall(fleetAsk, askHandler)
.toolCall(fleetStatus, statusHandler)
.toolCall(fleetPoll, pollHandler)
.toolCall(fleetAck, ackHandler)
.toolCall(fleetSpawn, spawnHandler)
.toolCall(fleetList, listHandler)
.toolCall(fleetStop, stopHandler)
.toolCall(fleetProfiles, profilesHandler)
.toolCall(fleetWhoami, whoamiHandler)
})
.build();
this.authz = callers;
this.metrics = metrics;
@@ -385,7 +314,7 @@ public final class FleetMcp {
* without fabricating an SDK {@code McpSyncServerExchange}.
*
* <p>This surface exists because the enforcement was previously unreachable from a test: no
* test constructs a {@code FleetMcp}, so the whole MCP-side gate ran zero times in the suite
* test constructs a {@code BridgeMcp}, so the whole MCP-side gate ran zero times in the suite
* while the REST-side equivalent had ten tests. A security control nothing exercises is a
* claim, not a control.
*
@@ -406,7 +335,7 @@ public final class FleetMcp {
String reason = Authz.isUnauthenticated(caller) ? "unauthenticated" : "forbidden";
AuditLog.denied(caller, action, target, reason);
if (metrics != null) {
metrics.inc(FleetMetrics.AUTH_FAILURES, "reason", reason);
metrics.inc(BridgedMetrics.AUTH_FAILURES, "reason", reason);
}
return error(reason + ": " + caller.describe() + " may not " + action);
}
@@ -477,7 +406,7 @@ public final class FleetMcp {
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
/**
* {@code fleet_send}: delegate {@code content} to a worker session and block for its reply.
* {@code bridge_send}: delegate {@code content} to a worker session and block for its reply.
* The configured profiles are required so a profile name can never bypass target validation.
*
* (CB-548): {@code onAccepted} records delegator ownership the instant the send is accepted, so
@@ -501,8 +430,8 @@ public final class FleetMcp {
}
/**
* {@code fleet_send} carrying a {@code turnId}: the primary's answer to a worker's
* {@code fleet_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as
* {@code bridge_send} carrying a {@code turnId}: the primary's answer to a worker's
* {@code bridge_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as
* it resumes the same turn — surfaced to the primary identically to a normal send.
*/
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs) {
@@ -514,13 +443,13 @@ public final class FleetMcp {
}
/**
* {@code fleet_ask} (CB-205): a worker pauses its delegated turn to ask the primary, blocking
* {@code bridge_ask} (CB-205): a worker pauses its delegated turn to ask the primary, blocking
* until the primary answers. The worker is identified by its connection ({@code callerTerminal}),
* never an argument — a {@code null} means the caller is not a known worker.
*/
static McpSchema.CallToolResult ask(MessageService messages, String callerTerminal, String question, Long timeoutMs) {
if (callerTerminal == null) {
return error("fleet_ask is for workers only — could not identify the calling worker "
return error("bridge_ask is for workers only — could not identify the calling worker "
+ "from the connection");
}
if (isBlank(question)) {
@@ -530,10 +459,10 @@ public final class FleetMcp {
MessageService.AskResult r = messages.ask(callerTerminal, question, timeout);
return switch (r.outcome()) {
case ANSWERED -> text(r.answer());
case NO_WAITER -> error("no primary is awaiting this turn — fleet_ask only works while a "
+ "fleet_send delegation is open to answer it");
case NO_WAITER -> error("no primary is awaiting this turn — bridge_ask only works while a "
+ "bridge_send delegation is open to answer it");
case TIMED_OUT -> text("[no answer within " + timeout + "ms — the primary did not respond; "
+ "proceed on your best judgement, then call fleet_reply to end the turn]");
+ "proceed on your best judgement, then call bridge_reply to end the turn]");
};
}
@@ -541,10 +470,10 @@ public final class FleetMcp {
private static McpSchema.CallToolResult formatReply(MessageService.Reply r, long timeout) {
return switch (r.outcome()) {
case REPLIED -> text(r.text());
// The worker's turn finished but it never called fleet_reply — hand back the scraped
// The worker's turn finished but it never called bridge_reply — hand back the scraped
// transcript tail, flagged so the primary knows it isn't a structured reply.
case COMPLETED_UNREPLIED -> text(
"[worker finished without a structured fleet_reply — transcript tail follows]\n" + r.text());
"[worker finished without a structured bridge_reply — transcript tail follows]\n" + r.text());
// The worker ran the turn then wedged (CB-109) — surface the error context.
case WORKER_FAILED -> text("[worker failed — turn ended in an unrecoverable state]\n" + r.text());
// The backend refused on a subscription usage limit (CB-578 stage A) — the worker's
@@ -554,7 +483,7 @@ public final class FleetMcp {
+ "usage limit]\n" + r.text());
// The worker paused mid-turn to ask (CB-205) — tell the primary how to answer in-turn.
case QUESTION -> text("[question] the worker paused to ask before it can finish:\n" + r.text()
+ "\n\nAnswer it by calling fleet_send again with turnId=\"" + r.turnId()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + r.turnId()
+ "\" and content set to your answer; the worker resumes the same turn.");
case STALE_TURN -> error("that question is no longer open — it timed out or was already "
+ "answered (turnId stale)");
@@ -564,7 +493,7 @@ public final class FleetMcp {
}
/**
* {@code fleet_send} with {@code wait:false}: delegate {@code content} and return a ticket
* {@code bridge_send} with {@code wait:false}: delegate {@code content} and return a ticket
* immediately (fire-and-poll), so a long task isn't cut off by the caller's MCP call timeout.
* The configured profiles are required so a profile name can never bypass target validation.
*
@@ -581,67 +510,19 @@ public final class FleetMcp {
return targetError;
}
String ticket = messages.sendAsync(sessionId, content, onAccepted);
return text("accepted — task delegated. Poll fleet_poll with ticket=" + ticket);
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
}
/** A configured profile is never a send target; other unknown values may be herdr-owned panes. */
private static McpSchema.CallToolResult profileTargetError(String sessionId, Set<String> profiles) {
if (profiles.contains(sessionId)) {
return error("unknown send target \"" + sessionId + "\": it is a configured profile name, not a "
+ "session id. Call fleet_list to find a member or lead sessionId.");
+ "session id. Call bridge_list to find a member or lead sessionId.");
}
return null;
}
/**
* {@code fleet_send} carrying a {@code coordId} (CB-637): a message to a PEER LEAD, published to
* that lead's durable mailbox on the shared coordination broker. This is the only lead→lead path
* that crosses hosts — the existing pane-injection route can only reach a lead whose herdr socket
* this daemon shares.
*
* <p>{@code coordId} is mutually exclusive with {@code sessionId} and {@code turnId}: those two
* address a worker session owned by <em>this</em> daemon, a coord-id addresses a lead owned by
* another one, and there is no sensible reading of a call that sets both. Rejected by name rather
* than resolved by precedence, so a caller that meant the other route learns it instead of having
* its message quietly go somewhere else.
*
* <p>The publish is synchronous and confirmed by the broker, so the result is a real delivery
* receipt rather than a hopeful one. Its failure — nobody owns {@code coordId}'s mailbox, the
* broker nacked, or the confirm timed out — arrives as {@link IllegalStateException} and is
* turned into a tool error naming the coord-id. It is never allowed to escape as a crash: an
* unreachable peer is an ordinary outcome of addressing a fleet you do not control.
*
* @param leadChannel this daemon's channel, or {@code null} when no coordinator is configured
*/
static McpSchema.CallToolResult sendToLead(LeadChannel leadChannel, String coordId, String content,
String sessionId, String turnId) {
if (!isBlank(sessionId) || !isBlank(turnId)) {
String conflict = !isBlank(sessionId) ? "sessionId" : "turnId";
return error("coordId and " + conflict + " are mutually exclusive: coordId addresses a peer "
+ "LEAD on another daemon over the coordination broker, while " + conflict
+ " addresses a worker session on this one. Pass exactly one.");
}
if (isBlank(content)) {
return error("content is required");
}
if (leadChannel == null) {
return error("lead coordination is not configured (no coordinator: block) — cannot send to "
+ "peer lead \"" + coordId + "\". Add a coordinator: block with a shared broker uri "
+ "and this daemon's selfId, then restart fleetd.");
}
LeadMessage msg = new LeadMessage(UUID.randomUUID().toString(), leadChannel.selfCoordId(),
coordId, content);
try {
leadChannel.publish(coordId, msg);
} catch (IllegalStateException e) {
return error("cannot deliver to peer lead \"" + coordId + "\": " + e.getMessage()
+ ". Check that a daemon is running with coordinator.selfId=\"" + coordId
+ "\" and is connected to the same coordination broker.");
}
return text("delivered to peer lead " + coordId + " (msgId " + msg.msgId() + ")");
}
/** {@code fleet_poll}: check an async delegation by ticket, or drain a worker's inbox by target. */
/** {@code bridge_poll}: check an async delegation by ticket, or drain a worker's inbox by target. */
static McpSchema.CallToolResult poll(MessageService messages, String ticket, String target) {
if (!isBlank(target)) {
var replies = messages.drainReplies(target);
@@ -659,25 +540,25 @@ public final class FleetMcp {
}
return switch (v.phase()) {
case DONE -> text(v.replySource() != null && v.replySource().equals("transcript")
? "[done — worker finished without a structured fleet_reply; transcript tail follows]\n" + v.reply()
? "[done — worker finished without a structured bridge_reply; transcript tail follows]\n" + v.reply()
: v.reply());
case PENDING -> text("[pending — " + v.detail() + "]");
case ASKING -> text("[question — worker is waiting for your answer]\n" + v.reply()
+ "\n\nAnswer it by calling fleet_send again with turnId=\"" + v.turnId()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + v.turnId()
+ "\" and content set to your answer; the worker resumes the same turn.");
case FAILED -> text("[failed — " + v.detail() + "]");
};
}
/**
* {@code fleet_reply}: the worker returns its structured answer, resolving the awaiting send
* {@code bridge_reply}: the worker returns its structured answer, resolving the awaiting send
* or — when no send is open — queueing the reply in the inbox for later drain (CB-307).
* {@code callerTerminal} is resolved from the connection (never an argument); a {@code null}
* means the caller is not a known worker (e.g. the primary called it by mistake).
*/
static McpSchema.CallToolResult reply(MessageService messages, String callerTerminal, String content) {
if (callerTerminal == null) {
return error("fleet_reply is for workers only — could not identify the calling worker "
return error("bridge_reply is for workers only — could not identify the calling worker "
+ "from the connection");
}
if (content == null) {
@@ -687,7 +568,7 @@ public final class FleetMcp {
return text("delivered");
}
/** {@code fleet_ack}: acknowledge (remove) a specific reply from the inbox. */
/** {@code bridge_ack}: acknowledge (remove) a specific reply from the inbox. */
static McpSchema.CallToolResult ack(MessageService messages, String target, String msgId) {
if (isBlank(target) || isBlank(msgId)) {
return error("target and msgId are required");
@@ -697,8 +578,8 @@ public final class FleetMcp {
}
/**
* {@code fleet_status}: the live lifecycle status of a worker session, plus — when the worker
* is paused mid-turn in an async {@code fleet_ask} (CB-582) — the open question and how to
* {@code bridge_status}: the live lifecycle status of a worker session, plus — when the worker
* is paused mid-turn in an async {@code bridge_ask} (CB-582) — the open question and how to
* answer it, so a lead on its normal poll cadence does not need the ticket to notice.
*/
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
@@ -712,7 +593,7 @@ public final class FleetMcp {
return text(base);
}
return text(base + "\n\n[question — worker is waiting for your answer]\n" + ask.question()
+ "\n\nAnswer it by calling fleet_send again with turnId=\"" + ask.turnId()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + ask.turnId()
+ "\" and content set to your answer; the worker resumes the same turn."
+ " (ticket " + ask.ticket() + ")");
} catch (HerdrException e) {
@@ -721,7 +602,7 @@ public final class FleetMcp {
}
/**
* {@code fleet_whoami}: the caller's own identity, as the daemon already resolved it.
* {@code bridge_whoami}: the caller's own identity, as the daemon already resolved it.
*
* <p>Every other tool <em>consumes</em> this identity — the authorization gate, the reply
* rendezvous, the cwd inherit — but none reported it, so an agent had to infer its own role
@@ -729,7 +610,7 @@ public final class FleetMcp {
* name its MCP mount happens to carry, or {@code ANTHROPIC_BASE_URL} (which Claude-model
* workers do not set). The failure mode of guessing is asymmetric and silent: a primary that
* mistakes itself for a worker is refused by {@link Authz} and learns immediately, while a
* worker that mistakes itself for the primary ends its turn without {@code fleet_reply} and
* worker that mistakes itself for the primary ends its turn without {@code bridge_reply} and
* the sender simply receives nothing. This tool removes the guess.
*
* <p>For a worker the session registry adds what it knows about that session. A worker the
@@ -787,13 +668,13 @@ public final class FleetMcp {
// --- fleet management logic (CB-108 / CB-301) --------------------------------------------
/** {@code fleet_spawn} without cwd/caller context (default resolution). */
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
return spawn(sessions, profile, null, null, null, null, null, null, null);
}
/**
* {@code fleet_spawn}: launch a guard-checked member for {@code profile} (blank → the default
* {@code bridge_spawn}: launch a guard-checked member for {@code profile} (blank → the default
* profile) under {@code role} (blank → {@code dev}), and return its session id + pane id. The
* member's cwd is {@code requestedCwd} if given, else the profile's config, else
* {@code callerCwd} (the primary's directory), else the daemon's.
@@ -801,7 +682,7 @@ public final class FleetMcp {
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
* CB-584: {@code sessionName}/{@code resumeSessionId} request agent session identity — a resumed
* conversation requires an explicit {@code profile} whose adapter declares
* {@link dev.ltms.fleet.peer.Capability#SESSION_RESUME}, or the spawn is refused rather than
* {@link dev.ltms.bridged.peer.Capability#SESSION_RESUME}, or the spawn is refused rather than
* silently starting a cold session.
*
* <p>{@code role} and {@code profile} are independent: the role picks the contract, the profile
@@ -837,7 +718,7 @@ public final class FleetMcp {
}
}
/** Build a {@link WorktreeRequest} from {@code fleet_spawn}'s optional {@code worktree}/{@code ticket} args. */
/** Build a {@link WorktreeRequest} from {@code bridge_spawn}'s optional {@code worktree}/{@code ticket} args. */
private static WorktreeRequest worktreeRequest(Map<String, Object> a) {
Object w = a.get("worktree");
if (w == null || Boolean.FALSE.equals(w)) {
@@ -866,7 +747,7 @@ public final class FleetMcp {
}
/**
* {@code fleet_profiles}: the configured worker profiles, the default, and — CB-578 stage B —
* {@code bridge_profiles}: the configured worker profiles, the default, and — CB-578 stage B —
* which of them are currently quarantined (backend exhausted) and for how much longer. The
* {@code quarantined} key is present only when at least one profile is, so a fleet where nothing
* has ever been quarantined gets exactly the pre-stage-B shape.
@@ -895,7 +776,7 @@ public final class FleetMcp {
}
/**
* {@code fleet_list}: the whole fleet — {@code leads} and {@code workers} — each merged with
* {@code bridge_list}: the whole fleet — {@code leads} and {@code workers} — each merged with
* live herdr status. CB-519 decoupled the registry key (a host-unique id) from the herdr pane
* coordinate, so the join is on the terminal id, which both the session and the live agent carry.
*
@@ -908,10 +789,10 @@ public final class FleetMcp {
* <p>Leads are drawn from the resolver rather than from a second registry, so an address listed
* here is one that would actually resolve as a lead — see {@link CallerResolver#leads()}. The
* caller's own row is flagged {@code "self": true}: a peer needs to tell its own pane apart from
* a peer's, and the alternative is every lead calling {@code fleet_whoami} to subtract itself.
* a peer's, and the alternative is every lead calling {@code bridge_whoami} to subtract itself.
*
* <p>CB-583: the {@code capacity} rows reuse {@code quarantine} (the same {@link QuarantineSource}
* {@code fleet_profiles} reads) so the two surfaces cannot disagree about which profile is
* {@code bridge_profiles} reads) so the two surfaces cannot disagree about which profile is
* quarantined — see {@link #capacityView}.
*
* @param leads terminal_id → lead name, live from the resolver
@@ -926,22 +807,6 @@ public final class FleetMcp {
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, Map<String, String> leads, String selfTerm) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, leads, selfTerm, null);
}
/**
* As above, additionally reporting this daemon's own lead coordination id (CB-637) when one is
* configured and its channel opened. There is no peer-discovery surface yet — a lead addresses a
* peer by a coord-id it was told — so this row exists to answer the one question the operator
* cannot answer any other way: what is MY coord-id, the one a peer must use to reach me. It is
* omitted entirely when no coordinator is configured, so an ordinary fleet's output is unchanged.
*
* @param selfCoordId this daemon's coord-id, or {@code null} when lead coordination is off
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, Map<String, String> leads, String selfTerm,
String selfCoordId) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
@@ -960,9 +825,6 @@ public final class FleetMcp {
Map<String, Object> result = new LinkedHashMap<>();
result.put("leads", leadRows); result.put("members", out);
result.put("healthCoverage", healthCoverage.value().get());
if (selfCoordId != null && !selfCoordId.isBlank()) {
result.put("coordinator", Map.of("selfId", selfCoordId, "configured", true));
}
if (capacity.available()) result.put("capacity", profiles.stream()
.map(profile -> capacityView(profile, capacity.liveCount(), capacity.maxLoad(), roster, messages,
capacity.clock().getAsLong(), quarantine)).toList());
@@ -973,7 +835,7 @@ public final class FleetMcp {
}
/**
* Capacity is advisory only. {@code reclaimable} says there is no bridge work, not that fleetd
* Capacity is advisory only. {@code reclaimable} says there is no bridge work, not that bridged
* may stop the member: the bridge has capacity facts but no work list, and choosing work needs
* authority it does not have. {@code idleForSeconds} is derived from monotonic nanoTime and has
* no wall-clock meaning across a daemon restart.
@@ -994,7 +856,7 @@ public final class FleetMcp {
* CB-583: {@code free} alone cannot tell a lead "busy, will free up" from "refusing, and
* nothing changes for N seconds" — those need different decisions. So a quarantined profile
* forces {@code free} to 0, whatever its {@code maxLoad}/{@code live} say, and the row carries
* the same {@code credentialId}/{@code quarantinedForSeconds} facts {@code fleet_profiles}
* the same {@code credentialId}/{@code quarantinedForSeconds} facts {@code bridge_profiles}
* reports, reusing {@link QuarantineSource} rather than a second lookup. Both new keys are
* added only when the profile is actually quarantined, so an ordinary fleet's rows are
* byte-identical to before this change.
@@ -1027,7 +889,7 @@ public final class FleetMcp {
*
* <p>{@code status} is herdr's live view, and {@code unknown} when herdr is not tracking that
* pane as an agent — the honest answer, and the one that matters: a lead whose pane herdr cannot
* see is a lead a {@code fleet_send} cannot be typed into. It is reported rather than hidden,
* see is a lead a {@code bridge_send} cannot be typed into. It is reported rather than hidden,
* because a peer that has gone unreachable is exactly what the sender needs to know.
*/
private static Map<String, Object> leadView(String terminal, String name, Agent live,
@@ -1043,7 +905,7 @@ public final class FleetMcp {
return m;
}
/** {@code fleet_stop}: tear a worker down by its pane id. */
/** {@code bridge_stop}: tear a worker down by its pane id. */
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
if (isBlank(paneId)) {
return error("paneId is required");
@@ -1086,31 +948,25 @@ public final class FleetMcp {
// --- tool schemas --------------------------------------------------------------------------
private static McpSchema.Tool sendTool() {
return tool("fleet_send",
return tool("bridge_send",
"Delegate a task to a worker session. By default blocks until the worker replies and "
+ "returns its reply (or a 'still working / queued' note on timeout). Pass wait:false "
+ "for a long task to return a ticket immediately, then poll it with fleet_poll. To "
+ "answer a worker's fleet_ask, pass its turnId (with content) instead of sessionId. "
+ "To message a PEER LEAD on another daemon — possibly another host — pass its "
+ "coordId instead; that is coordination, never a task.",
+ "for a long task to return a ticket immediately, then poll it with bridge_poll. To "
+ "answer a worker's bridge_ask, pass its turnId (with content) instead of sessionId.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id (herdr terminal_id) to delegate to"),
"content", stringProp("The task/message to send to the worker (or your answer, with turnId)"),
"timeoutMs", Map.of("type", "integer", "description", "Max ms to wait for a reply (blocking mode)"),
"wait", Map.of("type", "boolean",
"description", "Block for the reply (default true); false returns a ticket to poll"),
"turnId", stringProp("When answering a worker's fleet_ask, its question turnId — "
+ "routes your answer back into the same turn (omit for a normal delegation)"),
"coordId", stringProp("A peer LEAD's coordination id — delivers content to that "
+ "lead's durable mailbox on the shared coordination broker, which works "
+ "across hosts. Mutually exclusive with sessionId and turnId. Your own "
+ "coordId is reported by fleet_list.")),
"turnId", stringProp("When answering a worker's bridge_ask, its question turnId — "
+ "routes your answer back into the same turn (omit for a normal delegation)")),
List.of("content")));
}
private static McpSchema.Tool askTool() {
// No target/session arg — the worker's identity is resolved from the connection.
return tool("fleet_ask",
return tool("bridge_ask",
"Pause your current delegated turn to ask the primary a question, blocking until it "
+ "answers — then resume the same turn with the answer. Use this when only the "
+ "primary has a decision or detail you need to continue. You do not address the "
@@ -1123,19 +979,19 @@ public final class FleetMcp {
}
private static McpSchema.Tool pollTool() {
return tool("fleet_poll",
"Check an async delegation (a fleet_send with wait:false) by its ticket: "
return tool("bridge_poll",
"Check an async delegation (a bridge_send with wait:false) by its ticket: "
+ "pending, done (with the worker's reply), or failed. When target (a worker "
+ "session id) is present instead of ticket, drain that worker's inbox of "
+ "replies delivered when no send was open.",
objectSchema(Map.of(
"ticket", stringProp("The ticket returned by fleet_send wait:false"),
"ticket", stringProp("The ticket returned by bridge_send wait:false"),
"target", stringProp("Worker session id to drain pending replies from (optional)")),
List.of()));
}
private static McpSchema.Tool ackTool() {
return tool("fleet_ack",
return tool("bridge_ack",
"Acknowledge (remove) a specific reply from a worker's inbox. Use when the primary "
+ "has processed a reply and wants to confirm it, leaving other pending replies "
+ "in the inbox for later drain.",
@@ -1146,22 +1002,22 @@ public final class FleetMcp {
}
private static McpSchema.Tool spawnTool() {
return tool("fleet_spawn",
return tool("bridge_spawn",
"Spawn a new off-subscription member session. A member has two independent attributes: "
+ "role (what it is for) and profile (which backend it runs on). Pass role to pick "
+ "the contract — 'dev' implements a unit and opens its own PR, 'reviewer' reviews a "
+ "diff it did not write, 'architect' refines a ticket before anyone builds it; omit "
+ "it for 'dev'. Pass profile (from fleet_profiles) to pick the backend, or omit it "
+ "it for 'dev'. Pass profile (from bridge_profiles) to pick the backend, or omit it "
+ "for the default. The two are independent: a reviewer may run on the same profile "
+ "as the dev it reviews. The member opens your current directory by default; pass "
+ "cwd to pin a different one. Pass worktree:true (with ticket) or "
+ "worktree:<ticket-slug> to provision an isolated git worktree. Pass resumeSessionId "
+ "to relaunch onto a prior conversation instead of starting cold — this requires an "
+ "explicit profile whose backend supports it (fleet_list shows agentSessionId for "
+ "explicit profile whose backend supports it (bridge_list shows agentSessionId for "
+ "resumable members), and is refused otherwise rather than silently starting fresh. "
+ "sessionName gives the member a display name in its own UI when the backend supports "
+ "one. Returns the member's sessionId (use with fleet_send) and paneId (use with "
+ "fleet_stop).",
+ "one. Returns the member's sessionId (use with bridge_send) and paneId (use with "
+ "bridge_stop).",
objectSchema(Map.of(
"role", stringProp("What the member is for: architect, dev or reviewer (default dev)"),
"profile", stringProp("Which backend to run it on (omit for the default profile)"),
@@ -1169,33 +1025,33 @@ public final class FleetMcp {
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
"ticket", stringProp("Ticket slug when worktree:true"),
"sessionName", stringProp("Logical display name for the member's own session, when its backend supports one"),
"resumeSessionId", stringProp("A prior member's agentSessionId (from fleet_list) to resume — requires an explicit profile that supports it")),
"resumeSessionId", stringProp("A prior member's agentSessionId (from bridge_list) to resume — requires an explicit profile that supports it")),
List.of()));
}
private static McpSchema.Tool profilesTool() {
return tool("fleet_profiles",
"List the configured worker profiles (backends) and which one fleet_spawn uses by "
return tool("bridge_profiles",
"List the configured worker profiles (backends) and which one bridge_spawn uses by "
+ "default. A 'quarantined' map is present when a backend-exhausted refusal put "
+ "a profile's credential on cooldown — fleet_spawn onto it is refused until "
+ "a profile's credential on cooldown — bridge_spawn onto it is refused until "
+ "quarantinedForSeconds elapses; a profile sharing that credential is listed too.",
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool listTool() {
return tool("fleet_list",
return tool("bridge_list",
"List the whole fleet the bridge tracks, in two parts. 'leads' are your PEERS — other "
+ "orchestrators, each with its sessionId (the address to fleet_send to), "
+ "orchestrators, each with its sessionId (the address to bridge_send to), "
+ "name, live status, and 'self': true on your own row; this is how you "
+ "discover a peer lead without being told its address. 'members' are the "
+ "sessions delegated to — each with sessionId, paneId, role (architect/dev/"
+ "reviewer), profile (the backend it runs on), state, optional "
+ "worktree/branch/owner/agentSessionId (the id to pass as fleet_spawn's "
+ "worktree/branch/owner/agentSessionId (the id to pass as bridge_spawn's "
+ "resumeSessionId to relaunch onto that same conversation, when the backend "
+ "supports it), and live herdr status. An empty 'members' "
+ "means no members are spawned; it says nothing about peers. When capacity "
+ "facts are configured, a 'capacity' row per profile also reports free: 0 for "
+ "a quarantined profile's credential (see fleet_profiles), whatever its "
+ "a quarantined profile's credential (see bridge_profiles), whatever its "
+ "maxLoad/live — with credentialId and quarantinedForSeconds naming the "
+ "quarantine, so 'free: 0, busy' can be told apart from 'free: 0, refusing "
+ "for N seconds'.",
@@ -1203,8 +1059,8 @@ public final class FleetMcp {
}
private static McpSchema.Tool stopTool() {
return tool("fleet_stop",
"Tear down a worker session by its paneId (from fleet_spawn or fleet_list).",
return tool("bridge_stop",
"Tear down a worker session by its paneId (from bridge_spawn or bridge_list).",
objectSchema(Map.of(
"paneId", stringProp("The worker's paneId to stop")),
List.of("paneId")));
@@ -1212,9 +1068,9 @@ public final class FleetMcp {
private static McpSchema.Tool replyTool() {
// No session/target arg — the caller's identity is resolved from the connection.
return tool("fleet_reply",
return tool("bridge_reply",
"Return your structured answer for a message you were sent, resolving the sender's "
+ "blocked fleet_send. A worker MUST end every delegated turn with exactly "
+ "blocked bridge_send. A worker MUST end every delegated turn with exactly "
+ "one of these. A lead uses it only to answer another lead that messaged "
+ "it — never to answer a worker, whose turn it is not.",
objectSchema(Map.of(
@@ -1223,7 +1079,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool statusTool() {
return tool("fleet_status",
return tool("bridge_status",
"Get the live lifecycle status (idle/working/blocked/unknown) of a worker session.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id to query")),
@@ -1231,14 +1087,14 @@ public final class FleetMcp {
}
private static McpSchema.Tool whoamiTool() {
return tool("fleet_whoami",
return tool("bridge_whoami",
"Report who YOU are on the bridge — your role is resolved from your connection "
+ "(unforgeable), never from anything you claim. Returns role 'primary' (you "
+ "orchestrate: spawn/send/stop; reply ONLY to answer a peer lead that "
+ "messaged you, never to answer a worker), 'architect' (you delegate turns "
+ "and reply/ask as your own pane, but cannot spawn/stop/drain), or 'worker' "
+ "(you were delegated to: you must end every turn with exactly one "
+ "fleet_reply, and cannot spawn or send), plus 'leader'/'architect' naming "
+ "bridge_reply, and cannot spawn or send), plus 'leader'/'architect' naming "
+ "which one you are, your own sessionId, and profile/worktree/branch when "
+ "you are a worker. Call this first when following role-conditional "
+ "instructions rather than guessing.",
@@ -1,6 +1,6 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.bridged.herdr.PaneLocator;
/**
* Resolves <em>who is calling</em> an MCP tool from the connection alone — the anti-spoofing
@@ -10,7 +10,7 @@ import dev.ltms.fleet.herdr.PaneLocator;
*
* <p>Both sources are authoritative and unforgeable: the OS reports the real connecting PID, and
* herdr owns the PID→pane mapping. A worker cannot claim to be another worker, nor the primary.
* Single-host only (the herd shares the {@code fleetd} host); the token path is the split-host
* Single-host only (the herd shares the {@code bridged} host); the token path is the split-host
* fallback.
*/
public final class ConnectionIdentity {
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
/**
* Resolves the OS PID that owns a loopback TCP source port — the OS half of connection-based MCP
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -11,7 +11,7 @@ import java.util.concurrent.atomic.AtomicReference;
* Single-slot, thread-safe registry for the primary's herdr {@code terminal_id}.
*
* <p>Populated from the caller terminal of orchestration-side MCP tools
* ({@code fleet_send}, {@code fleet_spawn}) — tools that only the primary calls.
* ({@code bridge_send}, {@code bridge_spawn}) — tools that only the primary calls.
* A pinned terminal (from config) seeds the registry at construction and makes
* subsequent {@link #record(String)} calls no-ops.
*
@@ -29,7 +29,7 @@ public final class PrimaryRegistry {
/**
* CB-532: worker terminal → the lead that delegated to it. The single slot above answers "who is
* THE primary", a question with no correct answer once two leads orchestrate the same fleet:
* whichever called {@code fleet_send} first captured every nudge, including nudges for the
* whichever called {@code bridge_send} first captured every nudge, including nudges for the
* other lead's delegations. This map answers the question that actually matters — "who is
* waiting on THIS worker" — and is what lets {@code primary.terminal} be retired.
*/
@@ -71,7 +71,7 @@ public final class PrimaryRegistry {
*
* <p>Called from the {@code MessageService} accepted-delivery hook — only after a send has won
* the session's send lock and queued delivery — where both halves are known (CB-548). It is
* deliberately <em>not</em> called at {@code fleet_send} request time: a concurrent sender that
* deliberately <em>not</em> called at {@code bridge_send} request time: a concurrent sender that
* times out {@code BUSY} must not steal a live delegation's reply routing without ever owning
* the turn. Last writer wins — if a second lead's later send is accepted, replies follow the
* lead that most recently delegated to it, which is the one waiting.
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
/**
* Resolves a process's current working directory from its PID — the OS half of CB-112's
@@ -1,12 +1,11 @@
package dev.ltms.fleet.member;
package dev.ltms.bridged.member;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -36,7 +35,7 @@ import java.util.function.Supplier;
* <li>{@link #buildLaunch}, which encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em> touching
* herdr, and mount the bridge MCP + reply charter as inline launch flags. A worker's base_url
* lives in the env map handed to herdr and nowhere else; {@code fleetd}'s own environment is
* lives in the env map handed to herdr and nowhere else; {@code bridged}'s own environment is
* never mutated, and nothing is written to the worker's profile.</li>
* </ul>
*/
@@ -54,7 +53,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* existing deployments and tests keep the legacy non-blocking spawn semantics.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, guard, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300));
@@ -65,7 +64,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* until the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, guard, profiles, defaultProfile, env,
@@ -78,10 +77,10 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* <em>role</em>; a profile may still override it with its own {@code tabLabel}.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<FleetConfig.Fleet> fleet) {
Supplier<BridgedConfig.Fleet> fleet) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
@@ -92,11 +91,11 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* Production constructor, plus the CB-596 {@code memberCredentials} policy supplier.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials) {
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
@@ -121,7 +120,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* (poll ms) is accepted for API symmetry but otherwise unused here
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
@@ -135,11 +134,11 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* @param fleet live fleet config, read once for each spawn
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<FleetConfig.Fleet> fleet) {
Supplier<BridgedConfig.Fleet> fleet) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet);
this.guard = guard;
@@ -149,12 +148,12 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* Full testability constructor, plus the CB-596 {@code memberCredentials} policy supplier.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials) {
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials);
this.guard = guard;
@@ -166,12 +165,12 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* default (the real {@code System.getenv()} key set) via the constructor above.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials,
Supplier<Set<String>> hostEnvNames) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials, hostEnvNames);
@@ -188,7 +187,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* applied here — see {@link #applySessionIdentity}.
*/
@Override
protected Launch buildLaunch(FleetConfig.Profile cfg, LaunchSpec spec) {
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
// CB-539: a profile may deliberately opt into the subscription (subscription: true) when no
// off-subscription endpoint exists for it — e.g. `sonnet` on `ccs`. That profile gets no
// ANTHROPIC_BASE_URL/AUTH_TOKEN (there is nothing to point them at) and the guard's base_url
@@ -212,17 +211,6 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
guard.assertWorker(baseUrl); // hard stop before we spawn anything
}
// CB-634: the IDE guidance is delivered as an on-disk CLAUDE.local.md overlay, NOT through
// the charter — the charter returns to role -> reply only. Best-effort: a failed overlay
// must never fail the spawn, and `writeIdeOverlay` no-ops unless the cwd is a provisioned
// worktree (see its .git-file safety gate). The overlay pins, and the auto-open opens, the
// module dir (this repo's pom is in `fleetd/`, not at the worktree root) — see ideProjectPath.
if (cfg.hasIdeMcp()) {
String projectPath = PeerLauncher.ideProjectPath(spec.cwd(), cfg.ideProjectDir());
writeIdeOverlay(spec.cwd(), projectPath);
PeerLauncher.openInIde(projectPath, cfg.ideOpenCommand(), log);
}
Map<String, String> workerEnv = baseEnv(cfg);
if (onSubscription) {
// CB-542 belt-and-braces: on the subscription path no guard vets these two keys, and the
@@ -239,16 +227,16 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
applyGitToken(workerEnv, cfg);
// CB-547a: Claude Code can MINT its own session id, so fleetd chooses it — a fresh spawn
// CB-547a: Claude Code can MINT its own session id, so bridged chooses it — a fresh spawn
// gets a UUID we pass as --session-id and return from agentSessionId(), so the resume
// handle is known BEFORE the agent has written anything; a resume spawn adopts its prior
// id via -r and passes no --session-id (the two conflict). Both are injected before the
// model flag so --model keeps outranking the operator's own argv.
// mutableArgv: argvWithFleet may hand back the profile's own (immutable) List.of when it
// mutableArgv: argvWithBridge may hand back the profile's own (immutable) List.of when it
// has neither MCP nor a charter — session flags must be added into a list we own.
List<String> argv = mutableArgv(argvWithFleet(cfg, spec));
List<String> argv = mutableArgv(argvWithBridge(cfg, spec));
String agentSessionId = applySessionIdentity(argv, spec.sessionName(), spec.resumeSessionId());
return new Launch(workerEnv, argvWithAutoCompact(argvWithModel(argv, cfg), cfg), agentSessionId);
return new Launch(workerEnv, argvWithModel(argv, cfg), agentSessionId);
}
/**
@@ -301,33 +289,27 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* is present it keeps its proven inline {@code --append-system-prompt} delivery, which is also
* the only form that reaches a member with no repo checkout.
*/
private List<String> argvWithFleet(FleetConfig.Profile cfg, LaunchSpec spec) {
private List<String> argvWithBridge(BridgedConfig.Profile cfg, LaunchSpec spec) {
String roleCharter = nonBlank(spec.roleCharter());
// CB-634: the IDE guidance is delivered as an on-disk overlay (writeIdeOverlay), not through
// the charter. The charter file is role -> reply only.
String replyCharter = nonBlank(spec.replyCharter());
Path agentFile = agentDefinitionFile(spec.cwd(), spec.role(), ".claude", "agents");
if (!cfg.mountsAnyMcp() && roleCharter == null
&& replyCharter == null && agentFile == null) {
if (!cfg.hasMcp() && roleCharter == null && replyCharter == null && agentFile == null) {
return cfg.argv();
}
List<String> argv = mutableArgv(cfg.argv());
if (cfg.mountsAnyMcp()) {
if (cfg.hasMcp()) {
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
argv.add("--mcp-config");
argv.add(mcpConfigJson(cfg));
argv.add(mcpJson);
}
// Combine the charters in order role -> reply, dropping any that are absent. When two
// or more survive they must ride one --append-system-prompt-file (CB-618 forbids the inline
// flag and the file flag together). A lone reply charter keeps its proven inline delivery.
List<String> charters = new java.util.ArrayList<>(2);
if (roleCharter != null) charters.add(roleCharter);
if (replyCharter != null) charters.add(replyCharter);
if (charters.size() == 1 && replyCharter != null && roleCharter == null) {
if (roleCharter != null) {
String combined = replyCharter == null ? roleCharter : roleCharter + "\n\n" + replyCharter;
argv.add("--append-system-prompt-file");
argv.add(writeCharterFile(combined).toString());
} else if (replyCharter != null) {
argv.add("--append-system-prompt");
argv.add(replyCharter);
} else if (!charters.isEmpty()) {
argv.add("--append-system-prompt-file");
argv.add(writeCharterFile(String.join("\n\n", charters)).toString());
}
if (agentFile != null) {
argv.add("--agent");
@@ -336,83 +318,6 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
return argv;
}
/**
* The {@code --mcp-config} JSON for this member: always the bridge mount when {@link
* FleetConfig.Profile#hasMcp()}, plus the IDE Index MCP as a second server named {@code
* intellij} when {@link FleetConfig.Profile#hasIdeMcp()} (CB-634). At least one is present —
* the caller only reaches here when {@link FleetConfig.Profile#mountsAnyMcp()} is true.
*/
private static String mcpConfigJson(FleetConfig.Profile cfg) {
StringBuilder servers = new StringBuilder();
if (cfg.hasMcp()) {
servers.append('"').append(PeerLauncher.MCP_MOUNT_NAME)
.append("\":{\"type\":\"http\",\"url\":\"").append(cfg.mcpUrl()).append("\"}");
}
if (cfg.hasIdeMcp()) {
if (servers.length() > 0) servers.append(',');
servers.append("\"intellij\":{\"type\":\"http\",\"url\":\"")
.append(cfg.ideMcpUrl()).append("\"}");
}
return "{\"mcpServers\":{" + servers + "}}";
}
/**
* Deliver the shared IDE guidance ({@link PeerLauncher#ideOverlayText}) as an on-disk
* {@code CLAUDE.local.md} overlay beside the project's own {@code CLAUDE.md} (CB-634), and
* register the overlay in the repository's common {@code info/exclude} so it never shows as
* untracked (git reads a worktree's excludes from the common dir, not the per-worktree gitdir).
*
* <p><strong>Safety gate:</strong> the overlay is written ONLY when {@code cwd/.git} is a
* <em>regular file</em> — a provisioned worktree keeps a {@code .git} FILE holding a
* {@code gitdir: <path>} pointer, while the primary's real checkout has a {@code .git}
* DIRECTORY. Returning without writing when {@code .git} is a directory is the whole safety of
* the feature: it must never write into a non-worktree cwd, i.e. never clobber a project that
* does not want the overlay.
*
* <p>Best-effort: a failure is logged at debug and swallowed — a failed overlay must never fail
* the spawn.
*
* @param cwd the member's worktree root, where the {@code CLAUDE.local.md} file is written
* @param projectPath the module dir the overlay pins {@code project_path} to (see
* {@link PeerLauncher#ideProjectPath}); equals {@code cwd} when no module subdir
*/
private static void writeIdeOverlay(String cwd, String projectPath) {
try {
Path dotGit = Path.of(cwd, ".git");
if (!Files.isRegularFile(dotGit)) {
// Not a provisioned worktree (primary's real checkout has a .git directory, or the
// cwd is not a repo at all). Never write into it.
return;
}
// The overlay FILE lives at the worktree root (claude-code's cwd), but its CONTENT pins
// project_path to the module dir the IDE opened (projectPath), not the worktree root.
Files.writeString(Path.of(cwd, "CLAUDE.local.md"), PeerLauncher.ideOverlayText(projectPath));
String gitdirLine = Files.readString(dotGit).trim();
Path gitDir = Path.of(gitdirLine.replaceFirst("^gitdir:\\s*", ""));
if (!gitDir.isAbsolute()) {
gitDir = Path.of(cwd).resolve(gitDir).normalize();
}
// git reads info/exclude from the COMMON dir, never the per-worktree gitdir (only
// info/sparse-checkout is per-worktree). A provisioned worktree's gitdir is
// <common>/worktrees/<name>, so the common dir is two levels up; writing the entry into
// the per-worktree gitdir leaves it un-honoured and the overlay shows as untracked.
Path commonDir = gitDir;
if (gitDir.getParent() != null && gitDir.getParent().getFileName() != null
&& "worktrees".equals(gitDir.getParent().getFileName().toString())) {
commonDir = gitDir.getParent().getParent();
}
Path exclude = commonDir.resolve("info").resolve("exclude");
Files.createDirectories(exclude.getParent());
String overlayLine = "CLAUDE.local.md";
if (!Files.exists(exclude) || Files.readAllLines(exclude).stream().noneMatch(overlayLine::equals)) {
Files.writeString(exclude, (Files.exists(exclude) ? System.lineSeparator() : "")
+ overlayLine + System.lineSeparator());
}
} catch (Exception e) {
log.debug("cannot write IDE overlay into worktree '{}'", cwd, e);
}
}
/** {@code s}, or {@code null} when {@code s} is null/blank — the charter-presence test used above. */
private static String nonBlank(String s) {
return (s == null || s.isBlank()) ? null : s;
@@ -428,7 +333,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
*/
private static Path writeCharterFile(String charterText) {
try {
Path file = Files.createTempFile("fleetd-role-charter-", ".md");
Path file = Files.createTempFile("bridged-role-charter-", ".md");
Files.writeString(file, charterText);
file.toFile().deleteOnExit();
return file;
@@ -453,7 +358,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* {@code gx10} does) are untouched — this adds nothing when there is nothing to add. This is
* the {@code kind: claude} counterpart of the opencode adapter's {@code -m provider/model}.
*/
private static List<String> argvWithModel(List<String> argv, FleetConfig.Profile cfg) {
private static List<String> argvWithModel(List<String> argv, BridgedConfig.Profile cfg) {
if (cfg.model() == null || cfg.model().isBlank()) {
return argv;
}
@@ -463,30 +368,6 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
return withModel;
}
/**
* Pin a bounded auto-compaction window on the command line via {@code --autocompact <tokens>},
* opt-in per profile (CB-634's sibling ticket: a member that runs out of context dies mid-turn
* and its {@code fleet_reply} — the whole point of the turn — is lost with it; opencode already
* forces {@code compaction.auto: true} unconditionally, CB-523, but Claude Code has no equivalent
* and runs at the backend's own default window).
*
* <p>Mirrors {@link #argvWithModel}: appended after it, so it survives the {@code ccs <profile>}
* wrapper the same way {@code --model} does, and outranks env/settings and the operator's own
* {@code argv}. Verified: {@code claude 2.1.241 --help} lists {@code --autocompact <auto|tokens>}
* (either the literal {@code auto}, or an integer 100k–1M) — {@link FleetConfig#load} rejects a
* configured value outside that band before this ever runs, so the flag Claude Code receives here
* is always in range.
*/
private static List<String> argvWithAutoCompact(List<String> argv, FleetConfig.Profile cfg) {
if (cfg.autoCompactWindow() == null) {
return argv;
}
List<String> withAutoCompact = mutableArgv(argv);
withAutoCompact.add("--autocompact");
withAutoCompact.add(String.valueOf(cfg.autoCompactWindow()));
return withAutoCompact;
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
@@ -530,7 +411,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(FleetConfig.Profile::hasGitToken);
return profileConfigs().stream().anyMatch(BridgedConfig.Profile::hasGitToken);
}
// --- CB-117 reap predicate (Claude prefix), kept for direct unit testing -------------------
@@ -1,21 +1,19 @@
package dev.ltms.fleet.member;
package dev.ltms.bridged.member;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementCandidate;
import dev.ltms.fleet.placement.PlacementContext;
import dev.ltms.fleet.placement.PlacementException;
import dev.ltms.fleet.placement.PlacementPolicies;
import dev.ltms.fleet.placement.PlacementPolicy;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.BackendQuarantine;
import dev.ltms.bridged.placement.PlacementCandidate;
import dev.ltms.bridged.placement.PlacementContext;
import dev.ltms.bridged.placement.PlacementException;
import dev.ltms.bridged.placement.PlacementPolicies;
import dev.ltms.bridged.placement.PlacementPolicy;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -23,7 +21,6 @@ import java.util.ArrayList;
import java.util.Collections;
import java.util.EnumSet;
import java.util.HashSet;
import java.util.IdentityHashMap;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
@@ -47,17 +44,12 @@ import java.util.stream.Collectors;
* the single adapter that declares it. Profiles partition cleanly across adapters: the
* constructor rejects a name claimed by two.</li>
* <li><strong>By pane id</strong> — {@link #stop} routes to the adapter that spawned that pane
* (recorded at spawn time). A pane the composite never spawned, or one whose record was lost
* to a daemon restart (CB-185 blocker 1 — {@link #spawnedBy} is in-memory only), can use the
* fallback route in a one-daemon fleet. With more than one herdr daemon, {@link #probeOwner}
* asks each configured daemon which one actually knows the pane: exactly one match routes
* (and caches); no match is treated as already-gone; more than one match is a genuine
* ambiguity (pane ids are per-daemon counters, so two daemons really can both hold, say,
* {@code w1:p1}) and stop refuses rather than closing a pane on an arbitrary herdr daemon.</li>
* (recorded at spawn time). A pane the composite never spawned (only real for a caller that
* hand-rolls an id) falls back to the first delegate; teardown is pane-id addressed and
* tab cleanup is single-occupant guarded, so it is safe either way.</li>
* <li><strong>Fleet-wide</strong> — {@link #reapOrphanWorkers} and {@link #capabilities} fan out
* and combine. {@link #list} is deduplicated by (owning daemon, pane id): delegates that share
* one herdr connection report the same global agent set, but two daemons can each hold a pane
* called {@code w1:p1}, so the daemon has to be part of the key.</li>
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
* shares one herdr connection and so reports the same global agent set.</li>
* </ul>
*
* <p>CB-518: an unqualified spawn is routed through a {@link PlacementPolicy}. The default
@@ -86,7 +78,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
* this way is the set of adapters ({@link #byProfile}), because a new backend needs a launcher
* and launchers are built once; {@code ConfigRef} classifies that as deferred and says so.
*/
private final Supplier<Map<String, FleetConfig.Profile>> profileConfigs;
private final Supplier<Map<String, BridgedConfig.Profile>> profileConfigs;
private final Supplier<PlacementPolicy> placementPolicy;
/** CB-578 stage B: credential cooldown, checked before an explicit spawn and filtered into placement. */
@@ -97,7 +89,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
* {@code null}, and an empty pool for a role, both fall back to every configured profile — the
* pre-CB-557 behaviour.
*/
private final Supplier<FleetConfig.Fleet> fleet;
private final Supplier<BridgedConfig.Fleet> fleet;
/**
* Backward-compatible constructor: fixed placement, no live-counting. Use this for tests and
@@ -127,7 +119,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, FleetConfig.Profile> profileConfigs,
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount) {
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, null, BackendQuarantine.none());
@@ -144,27 +136,27 @@ public final class CompositePeerLauncher implements PeerLauncher {
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, FleetConfig.Profile> profileConfigs,
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
FleetConfig.Fleet fleet) {
BridgedConfig.Fleet fleet) {
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, BackendQuarantine.none());
}
/**
* Production constructor with role pools and quarantine (CB-578 stage B). The full-featured
* non-reloading form; {@link #CompositePeerLauncher(List, String, Supplier, Function, BackendQuarantine)}
* is what {@code Fleetd.main} actually wires up.
* is what {@code Bridged.main} actually wires up.
*
* @param quarantine required — pass {@link BackendQuarantine#none()} for a caller that does not
* want the feature, never a defaulting overload (CB-578 stage B's own rule).
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, FleetConfig.Profile> profileConfigs,
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
FleetConfig.Fleet fleet,
BridgedConfig.Fleet fleet,
BackendQuarantine quarantine) {
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
@@ -183,7 +175,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Supplier<FleetConfig> config,
Supplier<BridgedConfig> config,
Function<String, Integer> liveCount,
BackendQuarantine quarantine) {
this(delegates, defaultProfile,
@@ -197,10 +189,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
/** The all-suppliers form every other constructor funnels into. */
private CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Supplier<Map<String, FleetConfig.Profile>> profileConfigs,
Supplier<Map<String, BridgedConfig.Profile>> profileConfigs,
Supplier<PlacementPolicy> placementPolicy,
Function<String, Integer> liveCount,
Supplier<FleetConfig.Fleet> fleet,
Supplier<BridgedConfig.Fleet> fleet,
BackendQuarantine quarantine) {
this.fleet = fleet;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
@@ -222,7 +214,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
}
}
}
// Order-preserving for the same reason, and because profiles() is user-visible (fleet_profiles).
// Order-preserving for the same reason, and because profiles() is user-visible (bridge_profiles).
this.byProfile = Collections.unmodifiableMap(index);
}
@@ -237,8 +229,8 @@ public final class CompositePeerLauncher implements PeerLauncher {
* <p>Read fresh on every call so a reload is visible; a caller that needs two consistent reads
* takes one local, as {@link #poolFor} does.
*/
private Map<String, FleetConfig.Profile> profiles0() {
Map<String, FleetConfig.Profile> m = profileConfigs.get();
private Map<String, BridgedConfig.Profile> profiles0() {
Map<String, BridgedConfig.Profile> m = profileConfigs.get();
return m == null ? Map.of() : m;
}
@@ -262,7 +254,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
String requestedProfile = req.profileName();
if (requestedProfile != null && !requestedProfile.isBlank()) {
// An explicit profile bypasses the placement policy, but not the capacity cap: maxLoad
// is documented as an unconditional limit on this profile (FleetConfig.Profile), and
// is documented as an unconditional limit on this profile (BridgedConfig.Profile), and
// the charter makes explicit-profile spawns the normal path — so skipping the check
// here would leave the cap dead config in real operation.
HerdrPeerLauncher d = route(requestedProfile);
@@ -275,7 +267,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
// CB-557: an unqualified spawn is placed inside the pool of the role it asked for, not across
// the whole profile list. An EXPLICIT profile (above) is left alone on purpose — it is the
// operator overriding, and refusing it would break `fleet_spawn{profile:"opus"}`, which
// operator overriding, and refusing it would break `bridge_spawn{profile:"opus"}`, which
// carries no role and so would be judged against the dev pool it was never meant for.
List<PlacementCandidate> candidates = candidates(req.role());
String roleDefault = defaultProfileFor(req.role());
@@ -325,7 +317,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
/**
* Refuse an explicit-profile spawn when the profile is at its {@code maxLoad} cap.
*
* <p>maxLoad is a documented, unconditional capacity limit (see {@code FleetConfig.Profile#maxLoad}),
* <p>maxLoad is a documented, unconditional capacity limit (see {@code BridgedConfig.Profile#maxLoad}),
* and the charter makes explicit-profile spawns the normal path — so enforcing it only in placement
* ({@code PlacementPolicyUtil}, package-private, hence not linked) would leave the cap dead config
* on every call that names a profile. Same rule as placement: {@code live >= cap} is at capacity.
@@ -367,7 +359,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
/** {@code profile}'s credential group (CB-578 stage B), or the profile's own name if unconfigured. */
private String credentialIdFor(String profile) {
FleetConfig.Profile cfg = profiles0().get(profile);
BridgedConfig.Profile cfg = profiles0().get(profile);
return cfg == null ? profile : cfg.effectiveCredentialId();
}
@@ -385,7 +377,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
// until CB-585: an explicit `maxLoad: 0` now survives as 0 and is a real cap of zero, so the
// check below refuses every spawn on that profile, and a negative value is refused at config
// load rather than normalized away.
FleetConfig.Profile cfg = profiles0().get(profile);
BridgedConfig.Profile cfg = profiles0().get(profile);
Integer cap = (cfg == null) ? null : cfg.maxLoad();
if (cap == null) {
return;
@@ -406,8 +398,8 @@ public final class CompositePeerLauncher implements PeerLauncher {
* a pool entry with no profile, so a survivor is a profile this particular composite does not own.
*/
private List<String> poolFor(MemberRole role) {
Map<String, FleetConfig.Profile> configured = profiles0();
FleetConfig.Fleet f = fleet.get();
Map<String, BridgedConfig.Profile> configured = profiles0();
BridgedConfig.Fleet f = fleet.get();
List<String> pool = (f == null) ? List.of() : f.profilesFor(role);
List<String> known = pool.stream().filter(configured::containsKey).toList();
return known.isEmpty() ? List.copyOf(configured.keySet()) : known;
@@ -423,7 +415,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
private List<PlacementCandidate> candidates(MemberRole role) {
List<PlacementCandidate> out = new ArrayList<>();
for (String name : poolFor(role)) {
FleetConfig.Profile w = profiles0().get(name);
BridgedConfig.Profile w = profiles0().get(name);
if (w != null) {
out.add(new PlacementCandidate(name, null, w.weight(), w.maxLoad()));
}
@@ -443,98 +435,12 @@ public final class CompositePeerLauncher implements PeerLauncher {
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.get(id);
HerdrPeerLauncher d = spawnedBy.remove(id);
if (d == null) {
if (herdrDaemonCount() == 1) {
log.debug("stop({}) — no recorded owner in a single-daemon fleet", id);
d = delegates.getFirst();
} else {
d = probeOwner(id);
if (d == null) {
// No configured herdr daemon has ever heard of this pane. CB-185 blocker 1: this
// is the normal case right after a daemon restart empties spawnedBy for a member
// that has ALREADY been torn down since — the caller retried a stop that already
// succeeded. Nothing to close and no owner to cache; matching the tolerance
// HerdrPeerLauncher#stop already gives an already-gone pane (agent.close swallows
// that as success), stop() here is a no-op rather than a refusal.
log.debug("stop({}) — no configured herdr daemon knows this pane; "
+ "treating as already stopped", id);
return;
}
}
log.debug("stop({}) — no recorded owner, routing to the first adapter (pane-addressed)", id);
d = delegates.getFirst();
}
// Drop the owner record only after the delegate accepted the stop. Removing it first meant a
// delegate that threw left the pane alive with its owner forgotten, so the retry fell into
// the ambiguous branch above and refused the id for good.
d.stop(id);
spawnedBy.remove(id);
}
/**
* CB-185 blocker 1: recover a spawnedBy cache miss by asking every distinct herdr daemon which
* one actually knows {@code id} — the fix for "after a restart, every surviving member becomes
* un-stoppable" (spawnedBy is in-memory only, so a restart empties it, and members intentionally
* outlive the daemon).
*
* <p>Grouped by daemon identity, not by delegate, for the same reason {@link #list()} groups
* that way: two adapters (claude-code, opencode) sharing one herdr connection would otherwise be
* probed twice, and a pane on their shared daemon would look owned by two adapters instead of
* one daemon.
*
* <p>A daemon that fails to answer {@code list()} (e.g. it is down) is treated as "does not know
* this pane" rather than aborting the whole probe — one unreachable daemon must never make a
* pane that a <em>different</em>, healthy daemon actually owns un-stoppable too, which would
* resurrect the exact bug this method exists to fix.
*
* @return the owning delegate — cached into {@link #spawnedBy} so the next call is free — or
* {@code null} when no daemon knows the pane
* @throws IllegalArgumentException when more than one daemon claims the pane: pane ids are
* per-daemon counters, so two daemons really can both hold, say, {@code w1:p1}, and there
* is no way to tell which one the caller means
*/
private HerdrPeerLauncher probeOwner(String id) {
Map<HerdrClient, HerdrPeerLauncher> byDaemon = new IdentityHashMap<>();
for (HerdrPeerLauncher delegate : delegates) {
byDaemon.putIfAbsent(delegate.herdr(), delegate);
}
List<HerdrPeerLauncher> owners = new ArrayList<>();
for (HerdrPeerLauncher representative : byDaemon.values()) {
List<Agent> agents;
try {
agents = representative.list();
} catch (HerdrException e) {
log.warn("stop({}) probe: a configured herdr daemon was unreachable ({}); "
+ "treating it as not knowing this pane", id, e.getClass().getSimpleName());
continue;
}
boolean knows = agents.stream().anyMatch(a -> id.equals(a.paneId()));
if (knows) {
owners.add(representative);
}
}
if (owners.size() > 1) {
throw new IllegalArgumentException("ambiguous paneId '" + id + "': "
+ owners.size() + " configured herdr daemons report this pane — "
+ "no way to tell which one the caller means");
}
if (owners.isEmpty()) {
return null;
}
HerdrPeerLauncher owner = owners.get(0);
spawnedBy.put(id, owner);
return owner;
}
/**
* Count actual herdr daemons, not peer adapter kinds. Identity is intentional: separate client
* objects may represent different daemons even if a client later implements value equality.
*/
private int herdrDaemonCount() {
Set<HerdrClient> daemons = Collections.newSetFromMap(new IdentityHashMap<>());
for (HerdrPeerLauncher delegate : delegates) {
daemons.add(delegate.herdr());
}
return daemons.size();
}
@Override
@@ -569,24 +475,14 @@ public final class CompositePeerLauncher implements PeerLauncher {
return route(profileName).capabilities();
}
/**
* Every herdr agent, deduplicated by (owning daemon, pane id).
*
* <p>Delegates that share one {@link HerdrClient} see the same global agent set, so listing them
* both would report every agent twice — that is what the dedupe is for. But pane ids are
* per-daemon counters, so two daemons really can both hold {@code w1:p1} on different panes.
* Keying on the pane id alone would silently drop one of them from {@code fleet_list} and from
* every status view built on it. The daemon is part of the key for exactly that reason.
*/
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
Map<HerdrClient, Integer> daemonIndex = new IdentityHashMap<>();
Map<String, Agent> byPane = new LinkedHashMap<>();
for (HerdrPeerLauncher d : delegates) {
int daemon = daemonIndex.computeIfAbsent(d.herdr(), _ -> daemonIndex.size());
for (Agent a : d.list()) {
if (a.paneId() != null) {
byPane.putIfAbsent(daemon + "\u0000" + a.paneId(), a);
byPane.putIfAbsent(a.paneId(), a);
}
}
}
@@ -1,20 +1,19 @@
package dev.ltms.fleet.member;
package dev.ltms.bridged.member;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.CharterReceipt;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -50,7 +49,7 @@ import java.util.regex.Pattern;
* <li>{@code namePrefix} (constructor arg) — the label prefix ({@code claude}, {@code opencode})
* that drives both unique naming and the orphan-reap pattern, so each adapter reaps only its
* own kind of pane and never another's.</li>
* <li>{@link #buildLaunch(FleetConfig.Profile, LaunchSpec)} — the peer-specific env map + argv, including any
* <li>{@link #buildLaunch(BridgedConfig.Profile, LaunchSpec)} — the peer-specific env map + argv, including any
* subscription/guard check, MCP mount, and instruction injection. The base never sees how the
* peer is configured; it only places and starts the returned {@link Launch}.</li>
* </ul>
@@ -77,7 +76,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
private final String namePrefix; // label prefix: naming + reap scheme
private final AgentControl agents;
private final WorkspaceControl spaces;
private final Map<String, FleetConfig.Profile> profiles; // profile name → spawn settings
private final Map<String, BridgedConfig.Profile> profiles; // profile name → spawn settings
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
/** Host env lookup (injectable for tests); adapters read it in {@link #buildLaunch}. */
@@ -92,14 +91,14 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* <p>CB-559: a supplier rather than a snapshot, so a config reload affects the next launch
* without a restart. Existing tabs keep the label they were given.
*/
private final Supplier<FleetConfig.Fleet> fleet;
private final Supplier<BridgedConfig.Fleet> fleet;
/**
* CB-596: the live {@code memberCredentials:} policy, read once per spawn (same hot-reload shape
* as {@link #fleet}). {@code null} — either the supplier itself, or what it returns — means no
* policy is configured and {@link #applyMemberCredentialPolicy} shadows nothing.
*/
private final Supplier<FleetConfig.MemberCredentials> memberCredentials;
private final Supplier<BridgedConfig.MemberCredentials> memberCredentials;
/**
* Enumerates the daemon's own process environment variable NAMES ONLY, never values — the CB-596
@@ -118,14 +117,14 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
/** The final instruction always requires a bridge reply when the bridge MCP is mounted. */
protected static final String REPLY_CHARTER =
"You are a spawned member in the claude-bridge fleet. Every message you receive arrives "
+ "through the bridge, and the ONLY channel back to the sender is the fleet_reply MCP tool. "
+ "through the bridge, and the ONLY channel back to the sender is the bridge_reply MCP tool. "
+ "Text you write in your terminal is NOT sent anywhere — the sender cannot see your screen, "
+ "so an in-terminal answer is silently discarded. Therefore you MUST end EVERY turn by calling "
+ "fleet_reply with `content` set to your complete response. This holds for every message without "
+ "bridge_reply with `content` set to your complete response. This holds for every message without "
+ "exception — tasks, questions, clarifications, acknowledgements, and ordinary back-and-forth "
+ "conversation. Call fleet_reply exactly once, as the final action of your turn, with your full "
+ "conversation. Call bridge_reply exactly once, as the final action of your turn, with your full "
+ "answer in `content`; never wait for confirmation first. If you end a turn without calling "
+ "fleet_reply, the sender receives nothing and the exchange stalls.";
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/**
* Tab numbers, counted per {@code role/profile} pair (CB-557).
@@ -157,20 +156,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
private final ConcurrentMap<String, String> paneByAgentId = new ConcurrentHashMap<>();
private final AtomicBoolean resetUnsupportedLogged = new AtomicBoolean();
/**
* CB-633: the per-spawn ZDOTDIR directory generated for a pane under
* {@code memberCredentials.policy: allow-list}, keyed by herdr pane id so every teardown exit
* ({@link #stop} is reached from explicit DELETE, orphan reap, and the spawn-readiness gate
* timeout alike) can read the scrub's own report and then remove the directory. A pane that
* never reaches {@code stop} (spawn failure) leaks its directory only until JVM exit, where the
* generator's {@code deleteOnExit} hooks are the backstop — the same cleanup shape the existing
* charter/config temp files use.
*/
private final ConcurrentMap<String, Path> zdotdirByPane = new ConcurrentHashMap<>();
/** Guards {@link #warnNonZsh} to one WARN per launcher instance, not one per spawn. */
private final AtomicBoolean nonZshShellWarned = new AtomicBoolean();
/**
* @param namePrefix label prefix for this peer kind (drives naming and reap)
* @param agents herdr agent control (start, status, close)
@@ -184,7 +169,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* interval is baked into this hook, so the base needs no poll field
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
@@ -200,11 +185,11 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* above, so every existing call site keeps the default without an edit.
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<FleetConfig.Fleet> fleet) {
Supplier<BridgedConfig.Fleet> fleet) {
this(namePrefix, agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
nowMillis, sleeper, fleet, null);
}
@@ -218,12 +203,12 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* existing call site keeps the pre-CB-596 default without an edit.
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials) {
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
this(namePrefix, agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
nowMillis, sleeper, fleet, memberCredentials, null);
}
@@ -233,12 +218,12 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* seam only — every production call site leaves this {@code null} and gets the real host env.
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials,
Supplier<Set<String>> hostEnvNames) {
this.fleet = fleet;
this.namePrefix = namePrefix;
@@ -261,7 +246,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* Any subscription/guard check, MCP mount, and instruction injection happen here. The env map
* and argv are adapter-private; the base only places and starts what is returned.
*/
protected abstract Launch buildLaunch(FleetConfig.Profile cfg, LaunchSpec spec);
protected abstract Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec);
/** Direct transport access for peer-specific, non-turn control operations. */
protected final AgentControl agents() {
@@ -342,7 +327,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
if (name == null || name.isBlank()) {
return List.of();
}
FleetConfig.Profile cfg = profiles.get(name);
BridgedConfig.Profile cfg = profiles.get(name);
return cfg == null ? List.of() : cfg.parityOverlay();
}
@@ -366,18 +351,18 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
protected Collection<FleetConfig.Profile> profileConfigs() {
protected Collection<BridgedConfig.Profile> profileConfigs() {
return profiles.values();
}
/** Resolve {@code profileName} (null/blank → default) to its config, or throw with the options. */
protected FleetConfig.Profile requireProfile(String profileName) {
protected BridgedConfig.Profile requireProfile(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
throw new IllegalArgumentException("no default worker profile is configured — "
+ "pass a profile; configured: " + profiles.keySet());
}
FleetConfig.Profile cfg = profiles.get(name);
BridgedConfig.Profile cfg = profiles.get(name);
if (cfg == null) {
throw new IllegalArgumentException("unknown worker profile '" + name
+ "' — configured: " + profiles.keySet());
@@ -414,48 +399,30 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
/**
* Spawn a peer with session identity (CB-547a). The session values, role, and charter are
* threaded from the {@link SpawnRequest} into {@link #buildLaunch(FleetConfig.Profile,
* threaded from the {@link SpawnRequest} into {@link #buildLaunch(BridgedConfig.Profile,
* LaunchSpec)}, and the launch's resolved agent-session id is returned alongside the agent so
* the caller can put it on the {@link PeerHandle}.
*/
protected Spawned spawnInternal(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId, MemberRole role) {
FleetConfig.Profile cfg = requireProfile(profileName);
FleetConfig.Fleet liveFleet = fleet == null ? null : fleet.get();
BridgedConfig.Profile cfg = requireProfile(profileName);
BridgedConfig.Fleet liveFleet = fleet == null ? null : fleet.get();
String roleCharter = liveFleet == null ? null : liveFleet.charterFor(role);
String replyCharter = cfg.hasMcp() ? REPLY_CHARTER : null;
String charter = roleCharter == null ? replyCharter
: replyCharter == null ? roleCharter : roleCharter + "\n\n" + replyCharter;
// CB-571: fingerprint the exact composed charter bytes once, here in the base, before the
// string leaves for an adapter — so Claude and OpenCode derive the same digest. A failed
// start has no fleet_spawn result and no roster row, so the failure log below is the only
// start has no bridge_spawn result and no roster row, so the failure log below is the only
// surface the byte count can appear on. The charter text itself is never logged.
CharterReceipt receipt = CharterReceipt.compose(role, cfg.profile(), roleCharter, charter);
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
try {
Launch launch = buildLaunch(cfg, new LaunchSpec(sessionName, resumeSessionId, role, charter,
roleCharter, replyCharter, cwd));
// CB-633: applied AFTER buildLaunch so the generated scrub's allow-list can also cover
// the exact env-map keys this launch injects (ANTHROPIC_*, OPENCODE_CONFIG, GITEA_TOKEN, …)
// — anything the daemon deliberately sets must survive its own control.
Path zdotdir = applyEnvironmentAllowListPolicy(cfg, launch);
Agent agent;
try {
agent = cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd, role, liveFleet)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd, charter);
} catch (RuntimeException e) {
// A failed spawn has no pane id, so nothing would ever key this directory for
// teardown and it would sit in the temp dir until the JVM exits cleanly — which,
// for a daemon, may be never. Remove it on the way out.
if (zdotdir != null) {
EnvAllowListScrub.deleteRecursively(zdotdir);
}
throw e;
}
if (zdotdir != null) {
zdotdirByPane.put(agent.paneId(), zdotdir);
}
Agent agent = cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd, role, liveFleet)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd, charter);
logCharterReceipt(receipt, true);
return new Spawned(agent, launch.agentSessionId(), receipt);
} catch (RuntimeException e) {
@@ -521,11 +488,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
req.sessionName(), spawned.agentSessionId(), spawned.receipt());
}
/** The herdr daemon that owns this launcher's pane coordinates. */
public HerdrClient herdr() {
return agents.herdr();
}
@Override
public String effectiveCwd(SpawnRequest req) {
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
@@ -545,7 +507,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
*/
private static String resolveCwd(String requestedCwd, FleetConfig.Profile cfg, String callerCwd) {
private static String resolveCwd(String requestedCwd, BridgedConfig.Profile cfg, String callerCwd) {
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
}
@@ -557,8 +519,8 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
/** Dedicated worker space → own tab (carrying cwd+env) → start the peer into the seed pane. */
private Agent spawnInTab(FleetConfig.Profile cfg, Map<String, String> workerEnv,
List<String> argv, String cwd, MemberRole role, FleetConfig.Fleet liveFleet) {
private Agent spawnInTab(BridgedConfig.Profile cfg, Map<String, String> workerEnv,
List<String> argv, String cwd, MemberRole role, BridgedConfig.Fleet liveFleet) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId(), cwd, workerEnv);
log.info("spawning {} profile={} space={} tab={} cwd={}",
@@ -616,7 +578,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* composed charter, if any; its argv element is replaced by its digest so the log still shows
* which args were passed without exposing the charter prose.
*/
private Agent spawnAsPane(FleetConfig.Profile cfg, Map<String, String> workerEnv,
private Agent spawnAsPane(BridgedConfig.Profile cfg, Map<String, String> workerEnv,
List<String> argv, String cwd, String charter) {
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
namePrefix, cfg.profile(), cwd, redactCharter(argv, charter));
@@ -658,7 +620,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* backstop for the astronomically unlikely nonce+seq clash; the name is a label only — herdr
* detects kind and status from terminal output, not from it.
*/
private Started startUniquelyNamed(FleetConfig.Profile cfg, List<String> argv, String paneId) {
private Started startUniquelyNamed(BridgedConfig.Profile cfg, List<String> argv, String paneId) {
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
// argv[0] — the configured executable — is dropped and only the extra args are passed.
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
@@ -812,19 +774,11 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
log.debug("not closing tab {} — it holds {} panes (not a dedicated peer tab)",
loc.tabId(), loc.tabPaneCount());
}
// CB-633: log this pane's allowed-N-of-M scrub report, then remove the generated ZDOTDIR.
// Last in, best-effort — a failure here must not mask a real teardown failure above.
try {
releaseZdotdir(paneId);
} catch (RuntimeException e) {
log.warn("memberCredentials allow-list: releasing ZDOTDIR for pane {} failed: {}",
paneId, e.getMessage());
}
}
/** Whether any configured profile places peers in their own tab (so tabs may need cleanup). */
private boolean usesTabPlacement() {
return profiles.values().stream().anyMatch(FleetConfig.Profile::tabPlacement);
return profiles.values().stream().anyMatch(BridgedConfig.Profile::tabPlacement);
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
@@ -890,7 +844,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* {@code GITEA_HOST}. Push over SSH is unaffected; the only incremental grant is PR-create.
* Peer-neutral, so every herdr adapter reuses it unchanged.
*/
protected void applyGitToken(Map<String, String> workerEnv, FleetConfig.Profile cfg) {
protected void applyGitToken(Map<String, String> workerEnv, BridgedConfig.Profile cfg) {
if (!cfg.hasGitToken()) {
return;
}
@@ -906,7 +860,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* herdr's own login-shell process environment (gitea issue #82, superseding CB-592's single
* hardcoded {@code GITEA_ACCESS_TOKEN} name — see {@link #applyMemberCredentialPolicy}). herdr
* spawns a pane from its <em>own</em> process environment and layers our map on top —
* {@link dev.ltms.fleet.herdr.WorkspaceControl#createTab} and {@code #splitPane} send only
* {@link dev.ltms.bridged.herdr.WorkspaceControl#createTab} and {@code #splitPane} send only
* the keys we put in that map, so any key we never mention passes straight through from
* herdr's own shell, admin credentials included.
*
@@ -932,10 +886,10 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* start a login shell, and it makes the intent explicit at the one place every adapter passes.
*/
private static final String BLOCKED_CREDENTIAL_SENTINEL =
"blocked-by-fleetd-cb596-see-gitea-issue-82";
"blocked-by-bridged-cb596-see-gitea-issue-82";
/**
* CB-592: marks a pane as a fleetd member so a shell startup file can decline to export
* CB-592: marks a pane as a bridged member so a shell startup file can decline to export
* operator-only credentials into it (gitea issue #77).
*
* <p>This name is deliberately one that {@code secrets.sh} never exports, which is exactly why
@@ -958,7 +912,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
* {@code env:} entries, then the CB-592 admin-token shadow.
*
* <p>Why this exists: fleetd passes herdr an explicit env map, and herdr merges it into
* <p>Why this exists: bridged passes herdr an explicit env map, and herdr merges it into
* <em>its own</em> process environment. So before this, a worker inherited whatever PATH the
* herdr server happened to be started with — on this host, one from weeks earlier with no JDK
* and no Maven, which left workers unable to run the build they were being asked to run. The
@@ -968,7 +922,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* <p>Adapter-specific variables are layered on top of this by {@code buildLaunch} and therefore
* win. That ordering is deliberate and load-bearing: it stops a profile's {@code env:} from
* overriding {@code ANTHROPIC_BASE_URL} and slipping past {@link
* dev.ltms.fleet.guard.SubscriptionGuard}, which is checked against the profile's
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
* {@code baseUrl} and nothing else.
*
* <p>The CB-596 credential shadow and the CB-592 marker are put in <em>last</em>, after the
@@ -977,7 +931,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* place both are applied: every {@code buildLaunch} in every adapter calls this first, so a new
* profile, and a peer kind not yet written, gets them for free.
*/
protected Map<String, String> baseEnv(FleetConfig.Profile cfg) {
protected Map<String, String> baseEnv(BridgedConfig.Profile cfg) {
Map<String, String> workerEnv = new LinkedHashMap<>();
String path = env.apply("PATH");
if (path != null && !path.isBlank()) {
@@ -1001,243 +955,36 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
*
* <p>No {@code memberCredentials} configured — the supplier is {@code null}, or it resolves to
* one whose {@code known} list is empty — shadows nothing. This is a real, config-driven gap
* (see {@link FleetConfig.MemberCredentials}'s javadoc), not a safe default: deny-by-default
* (see {@link BridgedConfig.MemberCredentials}'s javadoc), not a safe default: deny-by-default
* only defends names the operator has actually enumerated in {@code known}.
*
* <p>CB-633: under {@code policy: allow-list} this overlay is NOT the control anymore — it is
* applied before the login shell runs and a sourced file can (and did) undo it. The control is
* the ZDOTDIR scrub ({@link #applyEnvironmentAllowListPolicy}); {@code known}/{@code allow}
* remain as reporting only via {@link #logCredentialGap}.
*
* <p>CB-633 follow-up (#192): under {@code allow-list} this method does NOT call {@link
* #logCredentialGap} itself — at this point (called from {@link #baseEnv}, before {@link
* #applyEnvironmentAllowListPolicy} runs) we do not yet know whether the pane's shell is zsh, so
* we cannot yet pick correct wording. That decision, and the call, are deferred entirely to
* {@link #applyEnvironmentAllowListPolicy}, which knows by then whether the scrub will actually
* run.
*/
private void applyMemberCredentialPolicy(Map<String, String> workerEnv) {
FleetConfig.MemberCredentials creds = memberCredentials == null ? null : memberCredentials.get();
BridgedConfig.MemberCredentials creds = memberCredentials == null ? null : memberCredentials.get();
if (creds == null) {
return;
}
if (!creds.isAllowList()) {
overlayBlockedCredentials(workerEnv, creds);
logCredentialGap(creds, null);
}
}
/** Put {@link #BLOCKED_CREDENTIAL_SENTINEL} over every blocked name in the pane-creation env map. */
private static void overlayBlockedCredentials(Map<String, String> workerEnv,
FleetConfig.MemberCredentials creds) {
for (String name : creds.blockedSet()) {
workerEnv.put(name, BLOCKED_CREDENTIAL_SENTINEL);
}
}
/**
* CB-633: under {@code memberCredentials.policy: allow-list}, generate the per-spawn ZDOTDIR
* directory whose startup files blank every exported variable not on the DERIVED allow-list —
* running AFTER the pane's shell has finished sourcing the operator's chain, which is what no
* pre-shell env overlay can achieve. {@code EnvAllowListScrub} sources the scrub from both the
* generated {@code .zshrc} and {@code .zlogin}, since a herdr pane is a login shell on macOS
* and a plain interactive one on Linux. Mutates {@code launch.env()} to carry
* {@code ZDOTDIR=<dir>}, so both placement paths ({@link #spawnInTab}, {@link #spawnAsPane})
* pass it through {@code tab.create}/{@code pane.split}. Returns the directory for teardown
* registration, or {@code null} when the policy does not apply.
*
* <p>The allow-list handed to the generator is the derived profile set UNIONed with the
* operator's own {@code memberCredentials.allow:} names ({@link MemberEnvAllowList#derive(
* Collection, Set)} — CB-633 follow-up) and with the exact keys of THIS launch's env map —
* names the daemon itself injects must survive its own control. {@code SSH_AUTH_SOCK} is added
* ONLY when the config explicitly allows it, EVEN IF the operator also listed it under
* {@code allow:}; by default it is absent, so the scrub blanks it like any other non-derived
* name. It stays a one-off decision because it is a live handle to the operator's ssh-agent, not
* a value — a member holding it can sign with every key the agent holds, so letting it ride in
* on the generic {@code allow:} list would hand that out for an unrelated reason.
*/
private Path applyEnvironmentAllowListPolicy(FleetConfig.Profile cfg, Launch launch) {
FleetConfig.MemberCredentials creds = memberCredentials == null ? null : memberCredentials.get();
if (creds == null || !creds.isAllowList()) {
return null;
}
Set<String> allowed = derivedAllowedNames(creds, launch);
String loginShell = resolveEnv("SHELL");
boolean zsh = loginShell != null && (loginShell.endsWith("/zsh") || loginShell.equals("zsh"));
if (!zsh) {
// A non-zsh login shell ignores ZDOTDIR entirely: NO scrub would run, so pretending
// otherwise would be worse than saying so. Warn loudly and fall back to the CB-596
// sentinel overlay over the enumerated known: names — weaker (a sourced file can undo
// it), but strictly better than nothing. Deliberately no "allowed N of M" line here: the
// scrub this count describes does not run on this path, so printing it would tell an
// operator that a fraction of names were blocked when the real number blocked is zero.
// logCredentialGap's WARN (below) is the only signal for this path.
warnNonZsh(loginShell);
overlayBlockedCredentials(launch.env(), creds);
logCredentialGap(creds, null);
return null;
}
// Only reached when the scrub is actually about to run — the count below describes that
// scrub, so it must not be logged before this gate (see the non-zsh branch above). Same
// reasoning gates logCredentialGap's wording: passing the derived `allowed` set (non-null)
// here, and ONLY here, is what tells it the scrub will really blank an unkept name — #192.
logAllowListCoverage(allowed);
logCredentialGap(creds, allowed);
Path dir = EnvAllowListScrub.generate(Path.of(System.getProperty("java.io.tmpdir")), allowed);
launch.env().put("ZDOTDIR", dir.toAbsolutePath().toString());
log.info("memberCredentials policy=allow-list: profile={} generated ZDOTDIR {} — derived "
+ "allow-list holds {} name(s); the pane reports allowed N of M at release",
cfg.profile(), dir.getFileName(), allowed.size());
return dir;
}
/**
* The full kept-name set for this spawn: the profile-derived names, unioned with {@code
* memberCredentials.allow:} (CB-633 follow-up — previously ignored by this whole policy), the
* ssh-agent handle when explicitly allowed, and the exact keys of THIS launch's own env map.
*/
private Set<String> derivedAllowedNames(FleetConfig.MemberCredentials creds, Launch launch) {
Set<String> allowed = new java.util.TreeSet<>(
MemberEnvAllowList.derive(profiles.values(), creds.allowSet()));
if (creds.sshAuthSockAllowed()) {
allowed.add(SSH_AUTH_SOCK);
} // blocked by default: absent from the set ⇒ blanked by the scrub like any other name
allowed.addAll(launch.env().keySet());
return allowed;
}
/**
* CB-633 follow-up: one INFO line per allow-list spawn WHOSE SCRUB ACTUALLY RUNS, so an operator
* can read a single log line and know the scrub ran and how much of the visible environment it
* will keep. Callable ONLY from the zsh branch of {@link #applyEnvironmentAllowListPolicy}, after
* the shell gate — logging it before that gate (or on the non-zsh fallback, where nothing is
* scrubbed) would tell an operator a fraction of names were blocked when the real number blocked
* is zero, which is worse than not logging at all. {@code M} is {@link #hostEnvNames}' size (the
* daemon's own environment — see that field's javadoc for why it stands in for the pane's, which
* the daemon has no channel to inspect at spawn time) and {@code N} is how many of those names
* survive {@code allowed} (including the {@code LC_*} prefix rule). Neither number is a constant:
* both come from the actual derived set and the actual environment this spawn sees. Never logs a
* variable NAME or VALUE — only the counts.
*/
private void logAllowListCoverage(Set<String> allowed) {
Set<String> hostNames = hostEnvNames.get();
long kept = hostNames.stream().filter(name -> MemberEnvAllowList.keeps(allowed, name)).count();
log.info("member credentials: allowed {} of {}", kept, hostNames.size());
}
/** The operator ssh-agent handle — kept ONLY by explicit config decision, never by default. */
private static final String SSH_AUTH_SOCK = MemberEnvAllowList.SSH_AUTH_SOCK;
/**
* CB-633: a non-zsh login shell means the allow-list control CANNOT run — say so once per
* launcher instance, naming the shell, instead of failing silently.
*/
private void warnNonZsh(String shell) {
if (nonZshShellWarned.compareAndSet(false, true)) {
log.warn("memberCredentials policy=allow-list: member login shell '{}' is NOT zsh — "
+ "ZDOTDIR scrubbing cannot run, so members' inherited environment is "
+ "UNPROTECTED beyond the enumerated known: fallback. Move herdr onto a "
+ "zsh account or switch policy back to deny-by-default.",
shell == null ? "<unset>" : shell);
}
}
/**
* CB-633 teardown half: read the pane's scrub report (the denominator report the generated
* scrub wrote) and delete the directory. Called from {@link #stop}, which is the one funnel
* every teardown exit already goes through.
*
* <p><b>A missing report is a WARN, not a debug line.</b> The report is the only evidence that
* the scrub ran at all in that pane. Its absence has an innocent reading — the pane died before
* its shell finished starting — and a serious one: the shell was not zsh, or it read its
* startup files from somewhere other than the directory we generated, in which case the member
* ran for its whole life with the operator's full secret store in its environment and nothing
* said so. We cannot tell those two apart from here, so the line says what is and is not known
* rather than picking one. Logging this at debug is how a control that silently stopped working
* stays unnoticed — the failure mode this whole class exists to remove.
*/
private void releaseZdotdir(String paneId) {
Path dir = zdotdirByPane.remove(paneId);
if (dir == null) {
return;
}
EnvAllowListScrub.ScrubReport report = EnvAllowListScrub.readReport(dir);
if (report == null) {
log.warn("memberCredentials allow-list: pane {} left no scrub report in {} — the "
+ "environment scrub cannot be confirmed to have run. Either the pane ended "
+ "before its shell finished starting, or its shell never read our generated "
+ "startup files, in which case that member saw the full host environment.",
paneId, dir);
} else {
log.info("memberCredentials allow-list: pane {} allowed {} of {} environment variables",
paneId, report.allowed(), report.total());
List<String> shaped = report.blanked().stream()
.filter(name -> CREDENTIAL_SHAPED_NAME.matcher(name).matches())
.toList();
if (!shaped.isEmpty()) {
log.warn("memberCredentials allow-list: pane {} blanked credential-shaped variable(s) "
+ "{} — confirm none of them was something a member legitimately needed",
paneId, shaped);
}
}
EnvAllowListScrub.deleteRecursively(dir);
logCredentialGap(creds);
}
/** Credential-shaped env var name heuristic for {@link #logCredentialGap} — case-insensitive. */
private static final Pattern CREDENTIAL_SHAPED_NAME =
Pattern.compile("(?i).*(TOKEN|SECRET|_KEY|APIKEY|PASSWORD|CREDENTIAL|AUTH).*");
/**
* Guards the {@code effectiveAllowed == null} branch of {@link #logCredentialGap} — the
* genuinely-unprotected report (deny-by-default, and the allow-list non-zsh fallback) — to one
* WARN per launcher instance, not one per spawn.
*
* <p>CB-633 follow-up (#192): kept SEPARATE from {@link #allowListGapLogged} on purpose.
* {@code memberCredentials} is a live, re-read-per-spawn supplier, so the policy can change
* between two spawns on the same launcher. A single shared flag would let a harmless allow-list
* INFO on spawn 1 permanently suppress the real deny-by-default WARN a later spawn deserves —
* the report that matters most getting hidden by the report that doesn't. Two flags mean each
* report kind fires exactly once, independent of what the other kind already logged.
*/
private final AtomicBoolean unprotectedGapLogged = new AtomicBoolean();
/**
* Guards the {@code effectiveAllowed != null} branch of {@link #logCredentialGap} — the
* allow-list-scrub-covered report — to one INFO per launcher instance. See {@link
* #unprotectedGapLogged}'s javadoc for why this is a separate flag rather than a shared one.
*/
private final AtomicBoolean allowListGapLogged = new AtomicBoolean();
/** Guards {@link #logCredentialGap} to one WARN per launcher instance, not one per spawn. */
private final AtomicBoolean credentialGapLogged = new AtomicBoolean();
/**
* CB-596 criterion 4: a credential-shaped host env var name on neither {@code known} nor
* {@code allow} is not silently allowed — it is reported. {@link #hostEnvNames} enumerates the
* daemon's own environment (see that field's javadoc for why the daemon's env is read rather
* than the spawned pane's, which the daemon has no channel to inspect at spawn time); this logs
* every such NAME — never a value, a prefix of a value, or a hash of a value, so the log itself
* cannot leak anything.
*
* <p>CB-633 follow-up (#192): {@code effectiveAllowed} picks the wording, and it must NOT be
* picked from {@code creds.isAllowList()} — see {@link #applyEnvironmentAllowListPolicy}'s
* javadoc for the reasoning this mirrors. {@code null} means no scrub-derived allow-list was
* computed for this call — true on the deny-by-default path AND on the allow-list non-zsh
* fallback, where nothing is ever scrubbed — so the whole gap is real and gets the WARN,
* unchanged from before this fix. Non-null means this call came from the zsh branch of {@link
* #applyEnvironmentAllowListPolicy}, reachable ONLY after that method's own zsh gate — but
* {@code effectiveAllowed} is a SUPERSET of {@code known ∪ allow}: {@link MemberEnvAllowList#derive}
* also unions in every profile's {@code gitTokenEnv}/{@code gitHostEnv}/{@code tokenEnv}/
* {@code env:} keys, and {@link #derivedAllowedNames} further unions in this very spawn's own
* env keys — so a name can be in the gap (uncovered by {@code known}/{@code allow}) AND still be
* kept by the derived allow-list, in which case the scrub does NOT blank it and the member DOES
* inherit it. Lead review on #192 caught this: the first cut of this fix reported the WHOLE gap
* as scrub-blanked without checking that, which reported a real leak as safe — the exact
* inversion #192 exists to remove. So on this path the gap is split with {@link
* MemberEnvAllowList#keeps}, the SAME predicate the generated scrub itself evaluates, so this
* split cannot drift from what the scrub actually does: the names it says are kept get the WARN
* (same severity, and same guard, as the deny-by-default case — a name genuinely reaching a
* member unprotected is equally serious whichever path put it there), and the names it says are
* blanked keep the INFO.
* every such NAME, at WARN, at most once per launcher instance — never a value, a prefix of a
* value, or a hash of a value, so the log itself cannot leak anything.
*/
private void logCredentialGap(FleetConfig.MemberCredentials creds, Set<String> effectiveAllowed) {
private void logCredentialGap(BridgedConfig.MemberCredentials creds) {
Set<String> covered = new HashSet<>(creds.known());
covered.addAll(creds.allow());
List<String> gap = hostEnvNames.get().stream()
@@ -1248,37 +995,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
if (gap.isEmpty()) {
return;
}
if (effectiveAllowed == null) {
warnGapUnprotected(gap);
return;
}
List<String> keptByDerivedList = gap.stream()
.filter(name -> MemberEnvAllowList.keeps(effectiveAllowed, name))
.toList();
List<String> blankedByScrub = gap.stream()
.filter(name -> !MemberEnvAllowList.keeps(effectiveAllowed, name))
.toList();
if (!keptByDerivedList.isEmpty() && unprotectedGapLogged.compareAndSet(false, true)) {
log.warn("memberCredentials gap: {} credential-shaped env var name(s) are on neither "
+ "known: nor allow: — the derived allow-list keeps them anyway (a profile's "
+ "gitTokenEnv/gitHostEnv/tokenEnv/env: names one, or this spawn injects it), "
+ "so every member pane inherits them UNBLOCKED — {}. Add each to "
+ "memberCredentials.known (or .allow if a member legitimately needs it), or "
+ "remove it from whatever profile setting derives it in.",
keptByDerivedList.size(), keptByDerivedList);
}
if (!blankedByScrub.isEmpty() && allowListGapLogged.compareAndSet(false, true)) {
log.info("memberCredentials gap: {} credential-shaped env var name(s) are on neither "
+ "known: nor allow: — {}. The allow-list scrub blanks them anyway (they "
+ "are not on the derived allow-list), so no member pane keeps them; add "
+ "each to memberCredentials.known or .allow to make that explicit.",
blankedByScrub.size(), blankedByScrub);
}
}
/** The deny-by-default (and allow-list non-zsh fallback) WARN — unchanged byte-for-byte by #192. */
private void warnGapUnprotected(List<String> gap) {
if (unprotectedGapLogged.compareAndSet(false, true)) {
if (credentialGapLogged.compareAndSet(false, true)) {
log.warn("memberCredentials gap: {} credential-shaped env var name(s) are on neither "
+ "known: nor allow: — every member pane inherits them UNBLOCKED — {}. "
+ "Add each to memberCredentials.known (blocked by default) or .allow "
@@ -1,16 +1,15 @@
package dev.ltms.fleet.member;
package dev.ltms.bridged.member;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.node.ObjectNode;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.CharterReceipt;
import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.SpawnRequest;
import java.io.IOException;
import java.io.UncheckedIOException;
@@ -24,9 +23,6 @@ import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>opencode</strong> — an open-source,
* provider-agnostic terminal coding agent. Its whole reason for existing is to prove the
@@ -37,7 +33,7 @@ import org.slf4j.LoggerFactory;
* <p>The divergences from {@link ClaudeCodeLauncher}, all confined to {@link #buildLaunch}:
* <ul>
* <li><strong>No subscription boundary.</strong> opencode carries no {@code ANTHROPIC_BASE_URL}
* and there is no {@link dev.ltms.fleet.guard.SubscriptionGuard} — the guard is a
* and there is no {@link dev.ltms.bridged.guard.SubscriptionGuard} — the guard is a
* Claude-private concern, not part of the SPI. opencode reads the operator's own provider
* credentials from its global {@code auth.json}; the bridge injects none.</li>
* <li><strong>File-based MCP mount + instructions.</strong> opencode has no inline
@@ -53,8 +49,6 @@ import org.slf4j.LoggerFactory;
*/
public final class OpenCodeLauncher extends HerdrPeerLauncher {
private static final Logger log = LoggerFactory.getLogger(OpenCodeLauncher.class);
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "opencode";
@@ -76,7 +70,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300),
@@ -88,7 +82,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, profiles, defaultProfile, env,
@@ -101,10 +95,10 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* <em>role</em>; a profile may still override it with its own {@code tabLabel}.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<FleetConfig.Fleet> fleet) {
Supplier<BridgedConfig.Fleet> fleet) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot(), defaultDiscoveryRoot(), fleet);
@@ -114,11 +108,11 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* Production constructor, plus the CB-596 {@code memberCredentials} policy supplier.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials) {
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot(), defaultDiscoveryRoot(), fleet, memberCredentials);
@@ -144,7 +138,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* {@link OpenCodeSessionDiscovery})
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
@@ -159,12 +153,12 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* @param fleet live fleet config, read once for each spawn
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Path configRoot, Path discoveryRoot,
Supplier<FleetConfig.Fleet> fleet) {
Supplier<BridgedConfig.Fleet> fleet) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet);
this.configRoot = configRoot;
@@ -175,13 +169,13 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* Full testability constructor, plus the CB-596 {@code memberCredentials} policy supplier.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Path configRoot, Path discoveryRoot,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials) {
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials);
this.configRoot = configRoot;
@@ -207,22 +201,11 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* grant; and select the model with {@code -m}.
*/
@Override
protected Launch buildLaunch(FleetConfig.Profile cfg, LaunchSpec spec) {
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
Map<String, String> workerEnv = baseEnv(cfg);
// autoCompactWindow's opencode lever (limit.context) only targets a specific provider/model
// entry, so it needs model: in "provider/model" form. A profile that opts in without that
// shape gets no silent no-op — log it, once, here, whether or not writeConfig ends up running.
boolean wantsContextLimit = cfg.autoCompactWindow() != null && splitProviderModel(cfg.model()) != null;
if (cfg.autoCompactWindow() != null && !wantsContextLimit) {
log.warn("profile '{}' sets autoCompactWindow but model '{}' is not \"<provider>/<model>\" "
+ "form — opencode's per-model context limit could not be applied for this profile",
cfg.profile(), cfg.model());
}
// A config file is needed for the bridge MCP mount, a member charter, the IDE MCP (+ its
// guidance overlay, CB-634), a pinned endpoint (CB-508), or a resolvable autoCompactWindow.
if (cfg.hasMcp() || cfg.hasIdeMcp() || spec.charter() != null || hasCustomProvider(cfg)
|| wantsContextLimit) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg, spec.charter(), spec.cwd()).toString());
// A config file is needed for the bridge MCP mount, a member charter, or a pinned endpoint (CB-508).
if (cfg.hasMcp() || spec.charter() != null || hasCustomProvider(cfg)) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg, spec.charter()).toString());
}
applyGitToken(workerEnv, cfg);
List<String> argv = argvWithResume(argvWithModel(argvWithAuto(cfg), cfg), spec.resumeSessionId());
@@ -255,7 +238,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* primary's Anthropic subscription, and an opencode process has no Anthropic credential path
* at all. Pointing it at a local vLLM cannot leak the subscription.
*/
private static boolean hasCustomProvider(FleetConfig.Profile cfg) {
private static boolean hasCustomProvider(BridgedConfig.Profile cfg) {
return cfg.baseUrl() != null && !cfg.baseUrl().isBlank();
}
@@ -264,12 +247,12 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* permissions opencode does not explicitly deny. It is unconditional, not a preference: a
* spawned peer has no human at its pane — the bridge spawned it — so one that stops at an
* approval prompt is a wedged agent, indistinguishable from a legitimate mid-turn wait and
* unable to end its turn with {@code fleet_reply}. opencode's own help calls this
* unable to end its turn with {@code bridge_reply}. opencode's own help calls this
* "dangerous!", but the blast radius here is already bounded by design: a worker runs in its
* own git worktree on its own branch, is off-subscription, and cannot merge — the lead is the
* gate.
*/
private List<String> argvWithAuto(FleetConfig.Profile cfg) {
private List<String> argvWithAuto(BridgedConfig.Profile cfg) {
List<String> argv = mutableArgv(cfg.argv());
argv.add("--auto");
return argv;
@@ -292,7 +275,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
}
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
private List<String> argvWithModel(List<String> argv, FleetConfig.Profile cfg) {
private List<String> argvWithModel(List<String> argv, BridgedConfig.Profile cfg) {
if (cfg.model() != null && !cfg.model().isBlank()) {
argv.add("-m");
argv.add(cfg.model());
@@ -306,16 +289,16 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
*/
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd) {
private Path writeConfig(BridgedConfig.Profile cfg, String charterText) {
try {
Path dir = Files.createTempDirectory(configRoot, "fleetd-opencode-");
Path dir = Files.createTempDirectory(configRoot, "bridged-opencode-");
dir.toFile().deleteOnExit();
ObjectNode root = JSON.createObjectNode();
root.put("$schema", "https://opencode.ai/config.json");
// CB-523, opencode side: a worker that runs out of context dies mid-turn, and its reply
// — the entire point of the turn — is lost with it. Auto-compaction is therefore not an
// operator preference for a fleetd worker, it is a condition of the turn contract.
// operator preference for a bridged worker, it is a condition of the turn contract.
//
// Stated deliberately even though it is redundant today: OPENCODE_CONFIG is MERGED over
// ~/.config/opencode/config.json rather than replacing it, so a worker already inherits
@@ -324,7 +307,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
//
// Know the cost before removing it: this key WINS over the home config (verified — an
// OPENCODE_CONFIG value overrides the home value, it does not defer to it), so an
// operator who sets `compaction.auto: false` at home cannot turn it off for fleetd
// operator who sets `compaction.auto: false` at home cannot turn it off for bridged
// workers. That is the intended trade for peers we spawn and whose turns we must land;
// if per-profile control is ever wanted, add a profile knob rather than dropping this.
root.putObject("compaction").put("auto", true);
@@ -337,42 +320,15 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
root.putArray("instructions").add(charter.toAbsolutePath().toString());
}
if (cfg.hasMcp() || cfg.hasIdeMcp()) {
// One shared mcp node for both servers — putObject would replace the node (and thus
// the other server) on the second call, so build into a single get-or-create node.
ObjectNode mcp = root.withObject("mcp");
if (cfg.hasMcp()) {
ObjectNode mount = mcp.putObject(PeerLauncher.MCP_MOUNT_NAME);
mount.put("type", "remote");
mount.put("url", cfg.mcpUrl());
mount.put("enabled", true);
}
// CB-634: mount the IDE Index MCP in the same shape as the bridge remote server, and
// deliver its guidance via the instructions array (opencode does not read
// CLAUDE.local.md) rather than any system-prompt string.
if (cfg.hasIdeMcp()) {
ObjectNode ide = mcp.putObject("intellij");
ide.put("type", "remote");
ide.put("url", cfg.ideMcpUrl());
ide.put("enabled", true);
// CB-634: pin the overlay and open the IDE at the module dir (this repo's pom is
// in `fleetd/`, not at the worktree root) — see PeerLauncher.ideProjectPath.
String projectPath = PeerLauncher.ideProjectPath(cwd, cfg.ideProjectDir());
Path rules = dir.resolve("ide-rules.md");
Files.writeString(rules, PeerLauncher.ideOverlayText(projectPath));
rules.toFile().deleteOnExit();
// The array may already hold the member-charter path; withArray gets-or-creates.
root.withArray("instructions").add(rules.toAbsolutePath().toString());
PeerLauncher.openInIde(projectPath, cfg.ideOpenCommand(), log);
}
if (cfg.hasMcp()) {
ObjectNode bridge = root.putObject("mcp").putObject("bridge");
bridge.put("type", "remote");
bridge.put("url", cfg.mcpUrl());
bridge.put("enabled", true);
}
if (hasCustomProvider(cfg)) {
addCustomProvider(root, cfg);
}
if (cfg.autoCompactWindow() != null) {
applyContextLimit(root, cfg);
}
Path cfgFile = dir.resolve("opencode.json");
// Built with Jackson rather than string concatenation: the provider block is nested and
@@ -393,20 +349,20 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* <p>The provider id comes from the {@code provider/model} selector in {@code model:}, so one
* field drives both the declaration and the {@code -m} flag and they cannot drift apart.
*/
private void addCustomProvider(ObjectNode root, FleetConfig.Profile cfg) {
private void addCustomProvider(ObjectNode root, BridgedConfig.Profile cfg) {
String[] parts = splitModelSelector(cfg);
String providerId = parts[0];
String modelId = parts[1];
ObjectNode provider = root.putObject("provider").putObject(providerId);
provider.put("npm", "@ai-sdk/openai-compatible");
provider.put("name", providerId + " (fleetd)");
provider.put("name", providerId + " (bridged)");
ObjectNode options = provider.putObject("options");
options.put("baseURL", openAiBaseUrl(cfg.baseUrl()));
// vLLM and friends usually ignore the key, but the AI SDK still requires a non-empty one.
String token = resolveEnv(cfg.tokenEnv());
options.put("apiKey", (token == null || token.isBlank()) ? "fleetd-local-noauth" : token);
options.put("apiKey", (token == null || token.isBlank()) ? "bridged-local-noauth" : token);
provider.putObject("models").putObject(modelId).put("name", modelId);
}
@@ -416,69 +372,19 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* needs both, so a bare model name is rejected loudly rather than silently falling back to the
* default gateway — a worker quietly talking to the wrong endpoint is the failure this avoids.
*/
private static String[] splitModelSelector(FleetConfig.Profile cfg) {
String[] parts = splitProviderModel(cfg.model());
if (parts == null) {
private static String[] splitModelSelector(BridgedConfig.Profile cfg) {
String model = cfg.model();
int slash = model == null ? -1 : model.indexOf('/');
if (model == null || model.isBlank() || slash <= 0 || slash == model.length() - 1) {
throw new IllegalArgumentException(
"profile " + cfg.profile() + " sets baseUrl (a pinned opencode endpoint) so"
+ " model: must be \"<provider>/<model>\", e.g."
+ " \"local-vllm/deepseek-v4-flash\"; got "
+ (cfg.model() == null ? "null" : '"' + cfg.model() + '"'));
}
return parts;
}
/**
* Split {@code model} into its {@code provider} and {@code model} halves, or {@code null} when
* it is not in that shape (unset/blank, or no non-trailing {@code /}). Unlike
* {@link #splitModelSelector}, non-throwing — callers that only *optionally* need the split
* (autoCompactWindow's context-limit application) use this to fall back to a WARN rather than an
* exception, since a profile without {@code baseUrl} is not required to name a provider/model.
*/
private static String[] splitProviderModel(String model) {
int slash = model == null ? -1 : model.indexOf('/');
if (model == null || model.isBlank() || slash <= 0 || slash == model.length() - 1) {
return null;
+ (model == null ? "null" : '"' + model + '"'));
}
return new String[]{model.substring(0, slash), model.substring(slash + 1)};
}
/**
* Apply the per-profile {@code autoCompactWindow} as opencode's per-model context limit.
*
* <p>opencode has no absolute "compact at N tokens" knob — its {@code compaction} block only
* exposes {@code auto}/{@code prune}/{@code reserved}/{@code tail_turns}/
* {@code preserve_recent_tokens} — so the real lever is the model's own
* {@code provider.<p>.models.<m>.limit.context}, which bounds the window opencode compacts
* <em>within</em> rather than compacting exactly AT it the way Claude Code's {@code --autocompact}
* does.
*
* <p>Uses get-or-create nodes ({@code withObject}) at every level so this MERGES with any provider
* block {@link #addCustomProvider} already wrote for a custom-provider (pinned-endpoint) profile —
* it must never overwrite that block's {@code npm}/{@code name}/{@code options}. For a gateway
* profile (no {@code baseUrl}, so no prior provider block) this writes a partial
* {@code provider.<p>.models.<m>.limit} override, which opencode merges over its own built-in
* provider definition.
*
* <p>opencode's {@code limit} schema requires both {@code context} and {@code output}; there is no
* independent signal for the latter here, so 16384 is written as a safe default (documented in
* {@code fleetd.example.yaml}).
*
* <p>Silently does nothing when {@code model:} is not in {@code provider/model} form — a warning
* for that case is already logged once in {@code buildLaunch}, so this stays quiet rather than
* duplicating it.
*/
private static void applyContextLimit(ObjectNode root, FleetConfig.Profile cfg) {
String[] parts = splitProviderModel(cfg.model());
if (parts == null) {
return;
}
ObjectNode limit = root.withObject("provider").withObject(parts[0])
.withObject("models").withObject(parts[1]).withObject("limit");
limit.put("context", cfg.autoCompactWindow());
limit.put("output", 16384);
}
/**
* The OpenAI-compatible base URL for {@code baseUrl}. A bare {@code host:port} gets {@code /v1}
* appended (where these servers put the API); a URL that already carries a path is taken as-is,
@@ -590,7 +496,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(FleetConfig.Profile::hasGitToken);
return profileConfigs().stream().anyMatch(BridgedConfig.Profile::hasGitToken);
}
// --- CB-117 reap predicate (opencode prefix), kept for direct unit testing -----------------
@@ -1,4 +1,4 @@
package dev.ltms.fleet.member;
package dev.ltms.bridged.member;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
@@ -9,7 +9,7 @@ import java.nio.file.Path;
import java.util.stream.Stream;
/**
* Resolves the opencode session id for a fleetd worker from opencode's on-disk storage — the
* Resolves the opencode session id for a bridged worker from opencode's on-disk storage — the
* only place this adapter touches opencode's private layout, and deliberately the <em>only</em>
* class that does.
*
@@ -22,7 +22,7 @@ import java.util.stream.Stream;
* its shape, naming, and field names — lives here, so a layout change, or a switch to the HTTP
* server, changes exactly one class and nothing in {@link OpenCodeLauncher}.
*
* <p>The determinism that makes this useful is structural, not a guess: every fleetd worker runs
* <p>The determinism that makes this useful is structural, not a guess: every bridged worker runs
* in its own unique git worktree, so the record's {@code directory} (its project root) equals the
* worker's cwd identifies <em>its</em> session unambiguously. We match on {@code directory} rather
* than diffing {@code opencode session list} before/after — that races under concurrent spawns, and
@@ -50,7 +50,7 @@ final class OpenCodeSessionDiscovery {
*
* <p>Never throws: a missing {@code storageRoot}, an unreadable/malformed record, or a
* directory that has not been persisted yet all resolve to {@code null} rather than failing a
* spawn. A fleetd worker's session record is written lazily (when the session is first
* spawn. A bridged worker's session record is written lazily (when the session is first
* persisted), so {@code null} here is the normal answer right after the pane is ready, and the
* caller retries later.
*
@@ -1,8 +1,8 @@
package dev.ltms.fleet.metrics;
package dev.ltms.bridged.metrics;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.MemberSession;
import java.util.LinkedHashMap;
import java.util.Map;
@@ -13,33 +13,33 @@ import java.util.Map;
*
* <p>The set is deliberately small: each series maps to a failure mode this project has actually
* hit, not to whatever was easy to count. The two worth watching in practice are
* {@code fleet_sends_total{outcome="completion_fallback"}} — a rising share means turn detection
* {@code bridged_sends_total{outcome="completion_fallback"}} — a rising share means turn detection
* is degrading, the CB-115/116/118 failure family — and
* {@code fleet_push_nudges_total{outcome="exhausted"}}, which means the primary stopped draining
* {@code bridged_push_nudges_total{outcome="exhausted"}}, which means the primary stopped draining
* its inbox and CB-307's active push gave up.
*/
public final class FleetMetrics {
public final class BridgedMetrics {
/** Counter: delegated sends by terminal outcome. */
public static final String SENDS = "fleet_sends_total";
public static final String SENDS = "bridged_sends_total";
/** Counter: worker replies by the path that carried them (rendezvous vs stranded-to-inbox). */
public static final String REPLIES = "fleet_replies_total";
public static final String REPLIES = "bridged_replies_total";
/** Counter: push-loop nudges to the primary, by outcome. */
public static final String PUSH_NUDGES = "fleet_push_nudges_total";
public static final String PUSH_NUDGES = "bridged_push_nudges_total";
/** Counter: idle-lead heartbeat nudges to the lead, by outcome (CB-551). */
public static final String HEARTBEAT_NUDGES = "fleet_lead_heartbeat_nudges_total";
public static final String HEARTBEAT_NUDGES = "bridged_lead_heartbeat_nudges_total";
/** Counter: spawn attempts by peer kind and outcome. */
public static final String SPAWNS = "fleet_spawns_total";
public static final String SPAWNS = "bridged_spawns_total";
/** Counter: herdr socket calls by method and outcome. */
public static final String HERDR_CALLS = "fleet_herdr_calls_total";
public static final String HERDR_CALLS = "bridged_herdr_calls_total";
/** Counter: rejected requests by reason (CB-501). */
public static final String AUTH_FAILURES = "fleet_auth_failures_total";
public static final String AUTH_FAILURES = "bridged_auth_failures_total";
/** Gauge: session census by lifecycle state. */
public static final String SESSIONS = "fleet_sessions";
public static final String SESSIONS = "bridged_sessions";
/** Gauge: undrained replies held per target. */
public static final String INBOX_DEPTH = "fleet_inbox_depth";
public static final String INBOX_DEPTH = "bridged_inbox_depth";
private FleetMetrics() {
private BridgedMetrics() {
}
/**
@@ -1,4 +1,4 @@
package dev.ltms.fleet.metrics;
package dev.ltms.bridged.metrics;
import java.util.Map;
import java.util.NavigableMap;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import com.rabbitmq.client.AMQP;
import com.rabbitmq.client.Channel;
@@ -29,9 +29,8 @@ import java.util.concurrent.TimeoutException;
*
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target has a durable
* queue {@code agent.<target>.inbox}. The gateway that owns the target starts a manual-ack consumer
* ({@link #own}) that pulls persistent messages, up to its prefetch window, off that queue into an
* in-memory <em>held</em> map (keyed by {@code msgId}) but does <em>not</em> ack them.
* {@link #peek} returns that snapshot;
* ({@link #own}) that pulls persistent messages off that queue into an in-memory <em>held</em> map
* (keyed by {@code msgId}) but does <em>not</em> ack them. {@link #peek} returns that snapshot;
* {@link #ack} acks the broker delivery-tag and drops the entry. Because messages stay unacked until
* the owning gateway actually drains them, a crash (or a {@code java -jar} bounce) before caller-ack
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
@@ -142,7 +141,7 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
return new AmqpReplyInbox(factory.newConnection("fleetd-reply-inbox"), prefetch);
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"), prefetch);
} catch (Exception e) {
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import java.util.LinkedHashMap;
import java.util.List;
@@ -16,7 +16,7 @@ import java.util.concurrent.ConcurrentHashMap;
* non-broker path stays interchangeable.
*
* <p><strong>This is soft-state, NOT persistence.</strong> Lost on a {@code java -jar} bounce — that
* is correct and consistent with "fleetd stays soft-state." The Stage-2 AMQP adapter replaces this.
* is correct and consistent with "bridged stays soft-state." The Stage-2 AMQP adapter replaces this.
*/
public final class InMemoryReplyInbox implements ReplyInbox {
@@ -1,11 +1,11 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.session.MemberSession;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -18,7 +18,7 @@ import java.util.function.Supplier;
/**
* CB-551: an opt-in heartbeat that nudges the single idle lead back to work once it has been
* continuously idle past a quiet period with no open {@code fleet_send} driving it.
* continuously idle past a quiet period with no open {@code bridge_send} driving it.
*
* <p>Why this exists: the fleet is ONE lead + architects + workers, so an idle, stalled lead is a
* single point of failure for the fleet's progress. {@link ReplyPushLoop} nudges the lead only when
@@ -99,7 +99,7 @@ public final class LeadHeartbeatLoop {
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(FleetMetrics.HEARTBEAT_NUDGES, "outcome", outcome);
metrics.inc(BridgedMetrics.HEARTBEAT_NUDGES, "outcome", outcome);
}
}
@@ -290,7 +290,7 @@ public final class LeadHeartbeatLoop {
/** The nudge body, phrased for the two cases the heartbeat distinguishes. */
String nudgeText() {
StringBuilder sb = new StringBuilder(
"Heartbeat: you are idle and no fleet_send is waiting on you.");
"Heartbeat: you are idle and no bridge_send is waiting on you.");
if (hasPending()) {
sb.append(" The fleet has state to collect: ").append(pendingDetail());
} else {
@@ -309,10 +309,10 @@ public final class LeadHeartbeatLoop {
.append(" pending collection");
if (!replyTargets.isEmpty()) {
// Render each as the exact command so the lead can act without parsing: the nearest
// analogue to ReplyPushLoop's fleet_poll(target=...) nudge.
// analogue to ReplyPushLoop's bridge_poll(target=...) nudge.
sb.append(" (")
.append(String.join(", ",
replyTargets.stream().map(t -> "fleet_poll(target=" + t + ")").toList()))
replyTargets.stream().map(t -> "bridge_poll(target=" + t + ")").toList()))
.append(")");
}
sb.append(", ").append(doneSessions).append(" DONE session").append(doneSessions == 1 ? "" : "s")
@@ -1,11 +1,10 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -25,7 +24,7 @@ import java.util.function.LongSupplier;
/**
* The blocking delegation feature (CB-104): deliver {@code content} into a worker and block until
* the worker returns a <em>structured reply</em> via {@code fleet_reply} (the {@link Rendezvous}),
* the worker returns a <em>structured reply</em> via {@code bridge_reply} (the {@link Rendezvous}),
* then hand that reply back. Delivery is the {@link Injector}'s job (the background poller sends it
* when the worker is injectable); this service never drives the injector or scrapes the terminal —
* completion is the worker's explicit reply, not a guess about {@code agent_status}.
@@ -63,10 +62,10 @@ public final class MessageService {
/** Outcome of a blocking send. */
public enum Outcome {
/** The worker called {@code fleet_reply}; {@code text} holds the structured answer. */
/** The worker called {@code bridge_reply}; {@code text} holds the structured answer. */
REPLIED,
/**
* The worker's delegated turn finished without a {@code fleet_reply} (CB-106 fallback);
* The worker's delegated turn finished without a {@code bridge_reply} (CB-106 fallback);
* {@code text} is the scraped transcript tail rather than a structured answer.
*/
COMPLETED_UNREPLIED,
@@ -76,7 +75,7 @@ public final class MessageService {
*/
WORKER_FAILED,
/**
* The turn finished without a {@code fleet_reply} and the scrape matched the backend's
* The turn finished without a {@code bridge_reply} and the scrape matched the backend's
* configured usage-limit refusal pattern (CB-578 stage A); {@code text} is the reason,
* carrying the matched line. The worker's pane is healthy — only its account is refusing —
* so this is never reported as a completed reply, and is kept distinct from
@@ -97,14 +96,14 @@ public final class MessageService {
BUSY,
/**
* An answer ({@link #answer(String, String, long)}) referenced a {@code turnId} that is no
* longer open — the worker's {@code fleet_ask} already timed out or was answered.
* longer open — the worker's {@code bridge_ask} already timed out or was answered.
*/
STALE_TURN
}
/**
* @param outcome how the send ended (or paused)
* @param text the worker's answer when {@link #completed()} (a structured {@code fleet_reply}
* @param text the worker's answer when {@link #completed()} (a structured {@code bridge_reply}
* for {@link Outcome#REPLIED}, a scraped transcript tail for
* {@link Outcome#COMPLETED_UNREPLIED}), or the question for {@link Outcome#QUESTION},
* else {@code null}
@@ -123,7 +122,7 @@ public final class MessageService {
}
}
/** How a worker's {@code fleet_ask} (CB-205) resolved. */
/** How a worker's {@code bridge_ask} (CB-205) resolved. */
public enum AskOutcome {
/** The primary answered; {@link AskResult#answer} carries it. */
ANSWERED,
@@ -133,7 +132,7 @@ public final class MessageService {
TIMED_OUT
}
/** The outcome of a worker's {@code fleet_ask}: how it resolved and (if answered) the answer. */
/** The outcome of a worker's {@code bridge_ask}: how it resolved and (if answered) the answer. */
public record AskResult(AskOutcome outcome, String answer) {
}
@@ -141,7 +140,7 @@ public final class MessageService {
public enum Phase {
/** Delegated and in flight — queued for the worker or being worked. */
PENDING,
/** The worker is paused in {@code fleet_ask}; {@link TaskView#reply} and {@link TaskView#turnId} identify it. */
/** The worker is paused in {@code bridge_ask}; {@link TaskView#reply} and {@link TaskView#turnId} identify it. */
ASKING,
/** The worker's turn finished; {@link TaskView#reply} holds the answer. */
DONE,
@@ -154,7 +153,7 @@ public final class MessageService {
*
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, or the question when
* {@link #phase} is {@link Phase#ASKING}; otherwise {@code null}
* @param replySource {@code "reply"} (structured {@code fleet_reply}) or {@code "transcript"}
* @param replySource {@code "reply"} (structured {@code bridge_reply}) or {@code "transcript"}
* (completion scrape) when {@link Phase#DONE}, else {@code null}
* @param detail a human note (live worker status while pending, ask state, or failure reason)
* @param turnId correlation id for an {@link Phase#ASKING} ticket, else {@code null}
@@ -168,38 +167,28 @@ public final class MessageService {
private final String ticket;
private final String target;
private final CompletableFuture<Reply> future = new CompletableFuture<>();
/**
* When {@link #future} resolved, or {@code null} while it is still pending — the clock
* {@link #pruneTerminalTickets} measures the TTL from (#197). Deliberately a boxed
* {@code Long} rather than a {@code long} with a sentinel: {@link System#nanoTime} may
* legitimately return any value, zero and negatives included, so no numeric sentinel can mean
* "not stamped yet". Stamped by a {@code whenComplete} hook registered in the constructor, so
* every completion path stamps it — a reply, the completion fallback, a timeout, a failure,
* or an abandon on teardown — without each of those having to remember to.
*/
private volatile Long completedNanos;
private final long createdNanos;
private volatile Reply question;
private volatile String turnId;
private Task(String ticket, String target, LongSupplier nowNanos) {
private Task(String ticket, String target, long createdNanos) {
this.ticket = ticket;
this.target = target;
future.whenComplete((reply, ex) -> completedNanos = nowNanos.getAsLong());
this.createdNanos = createdNanos;
}
}
/**
* A worker session's currently-open {@code fleet_ask} question, surfaced so {@code fleet_status}
* A worker session's currently-open {@code bridge_ask} question, surfaced so {@code bridge_status}
* can show it without the caller needing the ticket first (CB-582). Only covers async
* (fire-and-poll) delegations, which track the question on their {@link Task}; a blocking
* ({@code wait:true}) send already hands the question straight back to its own caller, so there is
* nothing hidden left for {@code fleet_status} to surface in that case.
* nothing hidden left for {@code bridge_status} to surface in that case.
*/
public record PendingAsk(String ticket, String question, String turnId) {
}
private final AgentControl agents;
private final HerdrRouter router;
private final Injector injector;
private final Rendezvous rendezvous;
private final ReplyInbox inbox;
@@ -213,27 +202,8 @@ public final class MessageService {
/** Async task that owns each exact forward rendezvous waiter. */
private final ConcurrentHashMap<CompletableFuture<Rendezvous.Resolution>, Task> asyncTasksByWaiter =
new ConcurrentHashMap<>();
/** Async tickets paused on a specific {@code fleet_ask} turn. */
/** Async tickets paused on a specific {@code bridge_ask} turn. */
private final ConcurrentHashMap<String, Task> asyncTasksByTurn = new ConcurrentHashMap<>();
/**
* Targets whose most recent {@code fleet_reply} arrived with no send awaiting it (CB-640) —
* {@link Rendezvous#resolve} returned {@code false} and the reply was queued into the inbox
* instead (see {@link #reply}). The reply itself is not lost (it sits in the inbox for a
* later drain), but the stranding is a fact the health layer needs to see. Bounded by the
* target's own lifecycle rather than a TTL: an entry is cleared the next time this target's
* delivery is accepted ({@link #send}) or the target is torn down ({@link #abandon}), so the
* map holds at most one entry per session with an unresolved stranding right now.
*/
private final ConcurrentHashMap<String, Boolean> strandedReplies = new ConcurrentHashMap<>();
/**
* Targets whose last send timed out with {@link Outcome#TIMED_OUT_QUEUED} (CB-640) — the
* message never reached the {@link Injector} delivery window before the caller's deadline, so
* it is still sitting in the injector's own per-target queue. Set where {@link #send} already
* computes {@code wasDelivered} for that outcome; no new queue is kept here, only the fact.
* Cleared the same way as {@link #strandedReplies}: the next accepted delivery for the target
* ({@link #send} opening a fresh waiter) or a teardown ({@link #abandon}).
*/
private final ConcurrentHashMap<String, Boolean> queuedDeliveries = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
Thread.ofVirtual().name("bridge-async-", 0).factory());
@@ -246,7 +216,7 @@ public final class MessageService {
* (CB-588) whenever an async ticket started by {@link #sendAsync} reaches a
* terminal phase, whenever {@link #poll} hands a terminal ticket to its caller,
* and (CB-582) whenever an async ticket's worker pauses mid-turn in
* {@code fleet_ask} or that pause ends (answered or lapsed)
* {@code bridge_ask} or that pause ends (answered or lapsed)
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop) {
@@ -268,7 +238,6 @@ public final class MessageService {
MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox,
ReplyPushLoop pushLoop, Metrics metrics, LongSupplier nowNanos) {
this.agents = agents;
this.router = null;
this.injector = injector;
this.rendezvous = rendezvous;
this.inbox = inbox;
@@ -277,20 +246,6 @@ public final class MessageService {
this.nowNanos = nowNanos;
}
public MessageService(HerdrRouter router, Injector injector, Rendezvous rendezvous, ReplyInbox inbox,
ReplyPushLoop pushLoop, Metrics metrics) {
this.agents = null;
this.router = router;
this.injector = injector;
this.rendezvous = rendezvous;
this.inbox = inbox;
this.pushLoop = pushLoop;
this.metrics = metrics;
this.nowNanos = System::nanoTime;
}
private AgentControl agentsFor(String target) { return router != null ? router.agentsFor(target) : agents; }
/** Create with an explicit {@link ReplyInbox} and no push loop. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox) {
this(agents, injector, rendezvous, inbox, null);
@@ -303,7 +258,7 @@ public final class MessageService {
/** Current lifecycle status of a worker (the {@code GET /sessions/{id}/status} surface). */
public AgentStatus status(String target) {
return agentsFor(target).status(target);
return agents.status(target);
}
/** Read-only delegation fact for fleet views. */
@@ -317,120 +272,31 @@ public final class MessageService {
}
/**
* Read-only delegation fact for fleet views (CB-640): {@code target}'s last send timed out
* before the {@link Injector} ever delivered it — the caller saw
* {@link Outcome#TIMED_OUT_QUEUED} (see the {@code TimeoutException} branch of {@link #send}),
* and the message is still sitting in the injector's per-target queue waiting for the worker
* to go idle. Distinct from {@link Outcome#TIMED_OUT_WORKING}, where delivery already happened
* and only the reply is outstanding. Cleared the next time this target's delivery is accepted
* or the target is abandoned — see {@link #queuedDeliveries}.
*/
public boolean hasQueuedDelivery(String target) {
return target != null && queuedDeliveries.containsKey(target);
}
/**
* Read-only delegation fact for fleet views (CB-640): {@code target}'s last {@code fleet_reply}
* arrived while no send was waiting for it, so {@link Rendezvous#resolve} returned
* {@code false} and {@link #reply} fell back to queueing it in the inbox (see the CB-307
* javadoc there and {@code docs/CB-307-Reliable-Delivery.md} §1). Cleared the next time this
* target's delivery is accepted or the target is abandoned — see {@link #strandedReplies}.
*/
public boolean hasStrandedReply(String target) {
return target != null && strandedReplies.containsKey(target);
}
/**
* Read-only delegation fact for fleet views (CB-640): an async ticket is still
* {@link Phase#PENDING} against {@code target}, yet nothing is actually in flight for it — no
* open rendezvous waiter ({@link #hasAcceptedDelivery}) and no message still sitting in the
* injector's queue ({@link #hasQueuedDelivery}). A healthy PENDING ticket can briefly look this
* way while its virtual thread has not yet been scheduled or is blocked on the session lock
* behind another send to the same target, so this is a snapshot fact for the health classifier
* to weigh across ticks, not proof on its own that the ticket is stuck. It also genuinely
* persists — not just as a passing race — once an async {@code fleet_ask} lapses unanswered:
* {@link #ask} clears the ticket's question and returns it to {@code PENDING}, but {@link #send}
* already closed the forward waiter the instant the question surfaced, so the target has
* neither an accepted nor a queued delivery left to show for it.
*/
public boolean hasOrphanedDelegation(String target) {
if (target == null || hasAcceptedDelivery(target) || hasQueuedDelivery(target)) {
return false;
}
for (Task task : tasks.values()) {
if (target.equals(task.target) && task.question == null && !task.future.isDone()) {
return true;
}
}
return false;
}
/**
* Route a worker's explicit {@code fleet_reply}: resolve an open send, complete an async ticket
* still parked waiting on this exact turn's answer, or — only once neither applies — queue it in
* the inbox. Unlike the bare {@link Rendezvous#resolve}, a no-waiter result is <em>not</em> a
* failure — the reply is held for later drain.
* Route a worker's explicit {@code bridge_reply}: resolve an open send, or queue it in the
* inbox if no send is currently open. Unlike the bare {@link Rendezvous#resolve}, a no-waiter
* result is <em>not</em> a failure — the reply is held for later drain.
*
* <p><strong>Do NOT use this for mid-turn questions.</strong> {@code fleet_ask} /
* <p><strong>Do NOT use this for mid-turn questions.</strong> {@code bridge_ask} /
* {@link Rendezvous#resolveQuestion} must keep today's {@code NO_WAITER} behaviour — questions
* are interactive and must never be queued.
*
* @return always {@code true} — the reply resolved a live send, completed a parked ticket, or
* was queued
* @return always {@code true} — the reply either resolved a live send or was queued
*/
public boolean reply(String session, String content) {
if (rendezvous.resolve(session, content)) {
count(FleetMetrics.REPLIES, "path", "rendezvous");
count(BridgedMetrics.REPLIES, "path", "rendezvous");
return true; // a live send took it — unchanged fast path
}
// #137: no live rendezvous waiter, but this may be the worker's real fleet_reply resuming a
// turn that {@link #answer} already gave up waiting on. answer()'s own bounded wait (the
// primary's fleet_send{turnId} call, capped well under a minute) can time out and close its
// waiter long before the worker — now actually resuming real work — finishes and replies. That
// reply used to have nowhere to land but the session inbox, leaving the async ticket's future
// unresolved forever: fleet_poll{ticket} stayed PENDING until fleet_stop's abandon() forced it
// FAILED with a misleading "session released before it replied" reason, even though the reply
// had, in fact, arrived. Completing the matching ticket directly here means fleet_poll{ticket}
// sees the real reply instead.
Task orphan = askAnsweredAsyncTask(session);
if (orphan != null && orphan.future.complete(new Reply(Outcome.REPLIED, content))) {
if (orphan.turnId != null) {
asyncTasksByTurn.remove(orphan.turnId, orphan);
}
count(FleetMetrics.REPLIES, "path", "async-recovered");
return true; // the ticket itself took it — no inbox stranding at all
}
inbox.publish(session, UUID.randomUUID().toString(), content);
// CB-640: record the stranding itself (not just the reply text) so fleet health can see a
// worker whose replies keep missing their waiter, not only the queue depth this leaves behind.
strandedReplies.put(session, Boolean.TRUE);
// A rising inbox share is the signal CB-307 exists to make visible: the worker finished but
// nobody was waiting, so delivery now depends on the push loop and a drain.
count(FleetMetrics.REPLIES, "path", "inbox");
count(BridgedMetrics.REPLIES, "path", "inbox");
if (pushLoop != null) {
pushLoop.onReplyQueued(session);
}
return true; // held, not lost
}
/**
* The still-open async task on {@code target} whose {@code fleet_ask} was already answered — its
* {@link Task#turnId} is stamped but its {@link Task#question} was cleared by {@link #answer} —
* yet whose future is not resolved yet (#137). {@code null} if no such task exists, including the
* common case where {@code target}'s worker never used {@code fleet_ask} at all (a task that was
* never asked has {@code turnId == null}, so it can never match here and only ever completes
* through the ordinary rendezvous fast path in {@link #reply}).
*/
private Task askAnsweredAsyncTask(String target) {
for (Task task : tasks.values()) {
if (target.equals(task.target) && task.question == null && task.turnId != null
&& !task.future.isDone()) {
return task;
}
}
return null;
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
if (metrics != null) {
@@ -442,7 +308,7 @@ public final class MessageService {
private Reply recorded(Reply r) {
String label = sendOutcomeLabel(r.outcome());
if (label != null) {
count(FleetMetrics.SENDS, "outcome", label);
count(BridgedMetrics.SENDS, "outcome", label);
}
return r;
}
@@ -463,7 +329,7 @@ public final class MessageService {
* Abandon any send still waiting on {@code target} because its session has gone away (CB-516).
*
* <p>Without this, tearing a worker down left its rendezvous waiter open: a blocking
* {@code fleet_send} kept blocking, and an async one kept reporting {@code PENDING} until
* {@code bridge_send} kept blocking, and an async one kept reporting {@code PENDING} until
* {@link #ASYNC_TIMEOUT_MS} — thirty minutes — even though the worker provably no longer
* existed and the delegation could never complete. Worse, {@code poll} already had the evidence
* (it calls {@code liveStatus} to build its detail string and gets back {@code "unknown"}) and
@@ -472,43 +338,16 @@ public final class MessageService {
* <p>Resolving the waiter as a failure — rather than letting it time out — also means the
* outcome is counted, so a torn-down delegation stops being invisible to {@code /metrics}.
*
* <p><strong>#137 defence in depth.</strong> {@link #reply} already hands a worker's real
* {@code fleet_reply} straight to the async ticket it belongs to whenever one is still parked
* waiting for it (see {@link #askAnsweredAsyncTask}), so by the time a session is released its
* tasks are normally already resolved — this loop's {@code complete} calls are then harmless
* no-ops (a {@link CompletableFuture} can only resolve once). But should some other path someday
* strand a reply in the inbox without completing its ticket, checking
* {@link #hasStrandedReply(String)} here — before ever writing a failure — means a torn-down
* session whose worker in fact replied is still reported {@code REPLIED} with that reply's own
* text, never the misleading "the worker session was released before it replied" (which also
* means the snapshot/worktree recovery hint that follows it never prints once a reply exists).
*
* @return true if a live waiter or an async task was failed (never true for one recovered as a
* reply — see the note above)
* @return true if a live waiter was failed
*/
public boolean abandon(String target, String reason) {
boolean hadStrandedReply = hasStrandedReply(target);
// CB-640: the session is gone — nothing will ever accept or deliver into it now.
strandedReplies.remove(target);
queuedDeliveries.remove(target);
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
boolean failed = waiter != null && !waiter.isDone() && rendezvous.resolveFailure(waiter, reason);
boolean asyncFailed = false;
Reply recovered = null; // lazily drained at most once, only if a task actually needs it
for (Task task : tasks.values()) {
if (!target.equals(task.target) || task.question != null || task.future.isDone()) {
continue;
}
if (hadStrandedReply && recovered == null) {
recovered = recoverStrandedReply(target);
}
Reply outcome = recovered != null ? recovered : new Reply(Outcome.WORKER_FAILED, reason);
if (task.future.complete(outcome)) {
if (outcome.outcome() == Outcome.WORKER_FAILED) {
asyncFailed = true;
} else if (task.turnId != null) {
asyncTasksByTurn.remove(task.turnId, task);
}
if (target.equals(task.target) && task.question == null
&& task.future.complete(new Reply(Outcome.WORKER_FAILED, reason))) {
asyncFailed = true;
}
}
if (failed) {
@@ -517,21 +356,6 @@ public final class MessageService {
return failed || asyncFailed;
}
/**
* Drain {@code target}'s inbox and hand its content back as a {@link Outcome#REPLIED} result
* (#137 defence in depth for {@link #abandon}) — {@code null} if it turned out empty (the
* stranding fact raced away, e.g. a lead's own {@code fleet_poll} on the raw session already
* drained it first). When more than one message is queued, only the newest is the worker's actual
* final answer ({@link #drainReplies} returns them oldest-first).
*/
private Reply recoverStrandedReply(String target) {
var messages = drainReplies(target);
if (messages.isEmpty()) {
return null;
}
return new Reply(Outcome.REPLIED, messages.get(messages.size() - 1).content());
}
/**
* Acknowledge a specific reply by {@code msgId} for {@code target}. Removes it from the inbox
* so that a subsequent drain or peek no longer returns it.
@@ -622,11 +446,6 @@ public final class MessageService {
// enqueue failure is safely closed by the finally below: nothing is left queued, and the
// failed send leaves no stale waiter behind.
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
// CB-640: this send now owns target's delivery, so any earlier stranded-reply or
// still-queued fact no longer describes the live state — clear both rather than let
// them outlive the send that supersedes them.
strandedReplies.remove(target);
queuedDeliveries.remove(target);
try {
if (task != null) {
asyncTasksByWaiter.put(reply, task);
@@ -645,11 +464,6 @@ public final class MessageService {
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
if (!wasDelivered) {
// CB-640: still sitting in the injector's queue, waiting for the member to
// go idle — record the fact for fleet health (see queuedDeliveries).
queuedDeliveries.put(target, Boolean.TRUE);
}
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
} catch (ExecutionException e) {
@@ -670,7 +484,7 @@ public final class MessageService {
/**
* A worker's mid-turn question (CB-205 reverse rendezvous): surface {@code question} to the
* primary by resolving its open blocking {@code fleet_send}, then block this (worker) call until
* primary by resolving its open blocking {@code bridge_send}, then block this (worker) call until
* the primary answers via {@link #answer} or {@code timeoutMillis} elapses. Identity is the
* worker's own session — it does not address the primary.
*
@@ -695,10 +509,10 @@ public final class MessageService {
rendezvous.closeAsk(ticket.turnId());
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
}
// CB-582: the question just became visible via fleet_poll (Phase.ASKING) for an async
// CB-582: the question just became visible via bridge_poll (Phase.ASKING) for an async
// (wait:false) delegation — nudge the lead's own pane the same way a terminal ticket does
// (CB-588), since the lead's normal poll cadence is minutes away and the reverse-rendezvous
// window (~55s, see FleetMcp/FleetApp) is far shorter. A blocking (wait:true) send has
// window (~55s, see BridgeMcp/BridgedApp) is far shorter. A blocking (wait:true) send has
// no Task and gets the question directly in its own reply, so task == null there — nothing
// to nudge.
if (task != null && pushLoop != null) {
@@ -709,7 +523,7 @@ public final class MessageService {
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
return new AskResult(AskOutcome.ANSWERED, answer);
} catch (TimeoutException e) {
log.debug("fleet_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
log.debug("bridge_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
clearAsyncQuestion(ticket.turnId(), true);
return new AskResult(AskOutcome.TIMED_OUT, null);
} catch (ExecutionException e) {
@@ -735,13 +549,13 @@ public final class MessageService {
}
/**
* The primary's answer to a worker's {@code fleet_ask} (CB-205): resolve the worker's blocked
* The primary's answer to a worker's {@code bridge_ask} (CB-205): resolve the worker's blocked
* question identified by {@code turnId}, then — like a fresh {@link #send} — block for the worker's
* eventual {@code fleet_reply} as it finishes the resumed turn. The worker session is derived from
* eventual {@code bridge_reply} as it finishes the resumed turn. The worker session is derived from
* {@code turnId}, never a caller argument.
*
* <p>Unlike {@link #send} this does not re-inject through the {@link Injector}: the worker is
* mid-turn (already picked up), so the answer flows back through its own open {@code fleet_ask}
* mid-turn (already picked up), so the answer flows back through its own open {@code bridge_ask}
* call, not a new status-gated delivery. The forward waiter is opened <em>before</em> the worker
* is unblocked so a reply that lands the instant it resumes is not lost.
*/
@@ -806,13 +620,13 @@ public final class MessageService {
*/
public String sendAsync(String target, String content, Runnable onAccepted) {
String ticket = "task-" + ticketSeq.incrementAndGet();
Task task = new Task(ticket, target, nowNanos);
Task task = new Task(ticket, target, nowNanos.getAsLong());
tasks.put(ticket, task);
if (pushLoop != null) {
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
// worker paused in fleet_ask leaves it running, per finishAsyncTask's own contract — so
// worker paused in bridge_ask leaves it running, per finishAsyncTask's own contract — so
// this fires exactly once, from whichever path completes it: finishAsyncTask(task, result)
// below on any non-QUESTION outcome of send() — a worker's fleet_reply, the CB-106
// below on any non-QUESTION outcome of send() — a worker's bridge_reply, the CB-106
// completion fallback, a CB-109 wedge, TIMED_OUT, BUSY, or BACKEND_EXHAUSTED — the same
// finishAsyncTask reached via answer()'s finishAsyncTask(turnId, result) once a QUESTION
// is resolved, completeExceptionally(t) just below when send() itself throws, or a CB-516
@@ -890,7 +704,7 @@ public final class MessageService {
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
private String liveStatus(String target) {
try {
return agentsFor(target).status(target).name().toLowerCase();
return agents.status(target).name().toLowerCase();
} catch (RuntimeException e) {
return "unknown";
}
@@ -905,26 +719,13 @@ public final class MessageService {
* too, its own {@code pendingTickets} entry would outlive the ticket it names: an unpolled ticket
* (or one the reminder cap already gave up on) is pruned here but never collected there, so it
* lingers in {@code pendingTickets} forever and rides along on every later nudge to the same lead
* — naming a ticket {@code fleet_poll} can no longer find (CB-588 follow-up).
*
* <p>The TTL runs from **completion**, not from creation (#197). It used to compare against
* {@code createdNanos}, which made the real collection window {@code TTL minus however long the
* task ran}: a delegation that took longer than the TTL had its reply destroyed on the first
* sweep after it landed, every time. That is the normal case here — real work runs well past ten
* minutes — and the reply lives only in {@code future}, so pruning it discards the worker's whole
* report with nothing to fall back on. Measuring from completion gives every ticket the same full
* window whatever its runtime, and still bounds {@code tasks}.
*
* <p>A task whose future is done but whose {@code completedNanos} is not stamped yet is left
* alone. That window is the few instructions between {@code complete()} and the constructor's
* {@code whenComplete} hook running; the next sweep collects it.
* — naming a ticket {@code bridge_poll} can no longer find (CB-588 follow-up).
*/
private void pruneTerminalTickets() {
long cutoff = nowNanos.getAsLong() - TICKET_TTL_NANOS;
tasks.entrySet().removeIf(e -> {
Task t = e.getValue();
Long completed = t.completedNanos;
boolean expired = t.future.isDone() && completed != null && completed < cutoff;
boolean expired = t.future.isDone() && t.createdNanos < cutoff;
if (expired && pushLoop != null) {
pushLoop.ticketCollected(e.getKey());
}
@@ -983,10 +784,10 @@ public final class MessageService {
}
/**
* The question {@code workerSession} is currently paused on via {@code fleet_ask}, if any
* (CB-582) — {@code fleet_status} uses this to show a pending question without the caller
* The question {@code workerSession} is currently paused on via {@code bridge_ask}, if any
* (CB-582) — {@code bridge_status} uses this to show a pending question without the caller
* needing the ticket. {@code null} when the session has no open async question (including a
* session mid a <em>blocking</em> {@code fleet_ask}, which has no {@link Task} to look up — see
* session mid a <em>blocking</em> {@code bridge_ask}, which has no {@link Task} to look up — see
* {@link PendingAsk}).
*/
public PendingAsk pendingAsk(String workerSession) {
@@ -1,13 +1,13 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicLong;
/**
* The reply rendezvous: where a blocking {@code fleet_send} awaits how the worker's delegated turn
* The reply rendezvous: where a blocking {@code bridge_send} awaits how the worker's delegated turn
* ends. The sending (primary) request thread {@link #open}s a waiter; it is resolved either by the
* worker's explicit {@code fleet_reply} ({@link #resolve}, arriving on a different thread via
* worker's explicit {@code bridge_reply} ({@link #resolve}, arriving on a different thread via
* {@code POST /sessions/{id}/reply}) or — the CB-106 fallback — by the injector observing the
* worker's delegated turn return to idle without a reply ({@link #resolveCompletion}).
*
@@ -27,14 +27,14 @@ public final class Rendezvous {
/** How a delegated turn ended (or paused). */
public enum Kind {
/** The worker called {@code fleet_reply} with a structured answer. */
/** The worker called {@code bridge_reply} with a structured answer. */
REPLY,
/** The worker's turn finished without a {@code fleet_reply}; {@code text} is a scrape. */
/** The worker's turn finished without a {@code bridge_reply}; {@code text} is a scrape. */
COMPLETION,
/** The worker ran the turn then wedged (CB-109); {@code text} is the failure context. */
FAILED,
/**
* The turn finished without a {@code fleet_reply}, and the scrape matched the backend's
* The turn finished without a {@code bridge_reply}, and the scrape matched the backend's
* configured usage-limit refusal pattern (CB-578 stage A); {@code text} is the reason,
* carrying the matched line. The pane is healthy — only the account is refusing — so this
* is kept separate from a session simply going {@code GONE}.
@@ -43,7 +43,7 @@ public final class Rendezvous {
/**
* The worker paused mid-turn to ask the primary a question (CB-205 reverse rendezvous);
* {@code text} is the question and {@code turnId} correlates the primary's answer back to
* the worker's blocked {@code fleet_ask}. Not terminal — the turn resumes after the answer.
* the worker's blocked {@code bridge_ask}. Not terminal — the turn resumes after the answer.
*/
QUESTION
}
@@ -75,7 +75,7 @@ public final class Rendezvous {
/** Reverse rendezvous (CB-205): worker questions awaiting the primary's answer, keyed by {@code turnId}. */
private final ConcurrentHashMap<String, AskWaiter> asks = new ConcurrentHashMap<>();
private final AtomicLong askSeq = new AtomicLong();
/** Per-session index of the currently-open ask, so duplicate fleet_ask calls coalesce onto one turn. */
/** Per-session index of the currently-open ask, so duplicate bridge_ask calls coalesce onto one turn. */
private final ConcurrentHashMap<String, String> openAsksBySession = new ConcurrentHashMap<>();
/**
@@ -132,7 +132,7 @@ public final class Rendezvous {
return complete(session, new Resolution(Kind.REPLY, content));
}
// --- reverse rendezvous (CB-205 fleet_ask) ------------------------------------------------
// --- reverse rendezvous (CB-205 bridge_ask) ------------------------------------------------
/**
* Open a reverse-rendezvous waiter for a worker's mid-turn question. If {@code session} already has
@@ -166,7 +166,7 @@ public final class Rendezvous {
}
/**
* Surface a worker's mid-turn {@code question} by resolving the primary's open {@code fleet_send}
* Surface a worker's mid-turn {@code question} by resolving the primary's open {@code bridge_send}
* with a {@link Kind#QUESTION} carrying {@code turnId}. Same session-keyed semantics as
* {@link #resolve}: the one outstanding send for {@code session} unblocks with the question.
*
@@ -184,7 +184,7 @@ public final class Rendezvous {
}
/**
* Resolve a worker's blocked {@code fleet_ask} with the primary's {@code answer}, unblocking it
* Resolve a worker's blocked {@code bridge_ask} with the primary's {@code answer}, unblocking it
* to resume its turn.
*
* @return {@code true} if the ask was still open and got the answer; {@code false} if the
@@ -195,7 +195,7 @@ public final class Rendezvous {
return w != null && w.answer().complete(answer);
}
/** Drop a reverse-rendezvous turn once its {@code fleet_ask} has resolved (answered or lapsed). */
/** Drop a reverse-rendezvous turn once its {@code bridge_ask} has resolved (answered or lapsed). */
public void closeAsk(String turnId) {
AskWaiter w = asks.get(turnId);
if (w == null) {
@@ -209,9 +209,9 @@ public final class Rendezvous {
/**
* Resolve a specific captured {@code waiter} as a completion (the delegated turn finished with no
* {@code fleet_reply}); {@code text} is the scraped transcript tail. The waiter is the one
* {@code bridge_reply}); {@code text} is the scraped transcript tail. The waiter is the one
* captured when this turn was delivered, so a late completion for turn N cannot land on turn N+1's
* send (CB-116). A no-op if that waiter was already resolved — a raced {@code fleet_reply} wins.
* send (CB-116). A no-op if that waiter was already resolved — a raced {@code bridge_reply} wins.
*
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
*/
@@ -233,7 +233,7 @@ public final class Rendezvous {
/**
* Resolve a specific captured {@code waiter} as {@link Kind#BACKEND_EXHAUSTED} (CB-578 stage A):
* the turn finished with no {@code fleet_reply} and the scrape matched the backend's configured
* the turn finished with no {@code bridge_reply} and the scrape matched the backend's configured
* usage-limit pattern; {@code reason} carries the matched line. Like
* {@link #resolveCompletion(CompletableFuture, String)} it targets the exact captured send
* (CB-116). A no-op if that waiter was already resolved — first resolution wins.
@@ -1,4 +1,4 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import java.util.List;
@@ -1,10 +1,10 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -19,9 +19,9 @@ import java.util.stream.Collectors;
/**
* A status-gated push loop that nudges a lead's own herdr pane when it has uncollected work
* waiting: a worker reply queued with no live {@code fleet_send} to resolve it (CB-307), an
* async delegation ticket ({@code fleet_send(wait:false)}) that reached a terminal phase
* (CB-588), or an async ticket's worker pausing mid-turn in {@code fleet_ask} to await an answer
* waiting: a worker reply queued with no live {@code bridge_send} to resolve it (CB-307), an
* async delegation ticket ({@code bridge_send(wait:false)}) that reached a terminal phase
* (CB-588), or an async ticket's worker pausing mid-turn in {@code bridge_ask} to await an answer
* (CB-582).
*
* <p><strong>CB-590: one schedule per lead.</strong> All three kinds of work are triggered
@@ -50,24 +50,24 @@ import java.util.stream.Collectors;
public final class ReplyPushLoop {
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
static final String NUDGE_FORMAT = "Worker %s returned a reply — run fleet_poll(target=%s) to collect it";
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
/** Coalesced form, several uncollected replies for the same lead. */
static final String REPLIES_NUDGE_FORMAT =
"%d workers returned replies — run fleet_poll(target=...) for each to collect them: %s";
"%d workers returned replies — run bridge_poll(target=...) for each to collect them: %s";
/** Singular form, one uncollected ticket. */
static final String TICKET_NUDGE_FORMAT =
"Ticket %s finished%s — run fleet_poll(ticket=%s) to collect it";
"Ticket %s finished%s — run bridge_poll(ticket=%s) to collect it";
/** Coalesced form, several uncollected tickets for the same lead. */
static final String TICKETS_NUDGE_FORMAT =
"%d tickets finished%s — run fleet_poll(ticket=...) for each to collect them: %s";
/** Singular form, one worker paused mid-turn in fleet_ask (CB-582) — names the answer call directly. */
"%d tickets finished%s — run bridge_poll(ticket=...) for each to collect them: %s";
/** Singular form, one worker paused mid-turn in bridge_ask (CB-582) — names the answer call directly. */
static final String QUESTION_NUDGE_FORMAT =
"Worker %s asked a question (ticket %s) — answer it with fleet_send(turnId=\"%s\", "
"Worker %s asked a question (ticket %s) — answer it with bridge_send(turnId=\"%s\", "
+ "content=...) to resume its turn:\n%s";
/** Coalesced form, several open questions for the same lead. */
static final String QUESTIONS_NUDGE_FORMAT =
"%d workers are paused on a question — run fleet_poll(ticket=...) for each, then answer "
+ "with fleet_send(turnId=..., content=...): %s";
"%d workers are paused on a question — run bridge_poll(ticket=...) for each, then answer "
+ "with bridge_send(turnId=..., content=...): %s";
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
@@ -87,7 +87,7 @@ public final class ReplyPushLoop {
/** Tickets that have gone terminal but not yet been polled, keyed by ticket. */
private final ConcurrentHashMap<String, PendingTicket> pendingTickets = new ConcurrentHashMap<>();
/**
* Open {@code fleet_ask} questions not yet answered or lapsed, keyed by {@code turnId}
* Open {@code bridge_ask} questions not yet answered or lapsed, keyed by {@code turnId}
* (CB-582). A question's own nudge count is tracked the same per-item way as
* {@link #pendingTickets} (CB-598): a fresh question keeps its source eligible regardless of
* how depleted an older, still-open question's count is.
@@ -118,7 +118,7 @@ public final class ReplyPushLoop {
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(FleetMetrics.PUSH_NUDGES, "outcome", outcome);
metrics.inc(BridgedMetrics.PUSH_NUDGES, "outcome", outcome);
}
}
@@ -130,7 +130,7 @@ public final class ReplyPushLoop {
/**
* Reply targets still pending for {@code lead} — registered via {@link #onReplyQueued} and
* whose inbox still holds an unacked message. A target whose inbox has since drained (acked,
* or collected via a live {@code fleet_send} rendezvous instead) is dropped from
* or collected via a live {@code bridge_send} rendezvous instead) is dropped from
* {@link #pendingReplies} here rather than lingering forever; there is no explicit "reply
* collected" callback the way {@link #ticketCollected} exists for tickets, so the inbox itself
* is the only signal.
@@ -323,7 +323,7 @@ public final class ReplyPushLoop {
}
/**
* Called when an async delegation ticket ({@code fleet_send(wait:false)}, CB-107) reaches a
* Called when an async delegation ticket ({@code bridge_send(wait:false)}, CB-107) reaches a
* terminal phase — DONE or a failure. Unlike {@link #onReplyQueued}, which nudges about the
* durable-inbox no-waiter path, this covers the path {@code MessageService.reply} takes when a
* fire-and-poll send's own rendezvous waiter resolves the reply directly: that path returns
@@ -352,7 +352,7 @@ public final class ReplyPushLoop {
}
/**
* Called when a ticket's terminal state has been collected via {@code fleet_poll}. Removes it
* Called when a ticket's terminal state has been collected via {@code bridge_poll}. Removes it
* from the pending set so a scheduled tick — and any nudge it sends — never names a ticket the
* lead already has. A ticket that was never pending (unknown ticket, or one nudged with no push
* loop configured) is a no-op.
@@ -362,16 +362,16 @@ public final class ReplyPushLoop {
}
/**
* Called when an async ticket's worker pauses mid-turn in {@code fleet_ask} (CB-582): the
* question is now visible via {@code fleet_poll} (Phase.ASKING), but the reverse-rendezvous
* window it opened with (~55s default, see {@code FleetMcp}/{@code FleetApp}) is far shorter
* Called when an async ticket's worker pauses mid-turn in {@code bridge_ask} (CB-582): the
* question is now visible via {@code bridge_poll} (Phase.ASKING), but the reverse-rendezvous
* window it opened with (~55s default, see {@code BridgeMcp}/{@code BridgedApp}) is far shorter
* than a lead's normal minutes-long poll cadence — exactly the gap this closes. Resolves the
* delegating lead the same way {@link #onTicketTerminal} does and coalesces onto the same
* per-lead schedule (CB-590).
*
* @param ticket the async ticket the question belongs to (for {@code fleet_poll})
* @param ticket the async ticket the question belongs to (for {@code bridge_poll})
* @param target the worker session that asked
* @param turnId correlation id the lead answers with ({@code fleet_send turnId=...})
* @param turnId correlation id the lead answers with ({@code bridge_send turnId=...})
* @param question the question text
*/
public void onQuestionOpened(String ticket, String target, String turnId, String question) {
@@ -386,7 +386,7 @@ public final class ReplyPushLoop {
}
/**
* Called when a worker's {@code fleet_ask} resolves — answered or lapsed unanswered — so a
* Called when a worker's {@code bridge_ask} resolves — answered or lapsed unanswered — so a
* scheduled tick never nudges about a question the lead already handled. A {@code turnId} that
* was never pending (never nudged, or already closed) is a no-op.
*/
@@ -1,4 +1,4 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import java.util.concurrent.CompletableFuture;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.peer;
package dev.ltms.bridged.peer;
/**
* Declared capabilities of a {@link PeerLauncher}. The protocol is the union across all
@@ -9,7 +9,7 @@ package dev.ltms.fleet.peer;
public enum Capability {
/**
* The peer supports {@code fleet_ask} rendezvous — pausing its delegated turn to ask
* The peer supports {@code bridge_ask} rendezvous — pausing its delegated turn to ask
* the primary a question, then resuming once answered. All Claude Code peers support this.
*/
MID_TURN_ASK,
@@ -1,4 +1,4 @@
package dev.ltms.fleet.peer;
package dev.ltms.bridged.peer;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.peer;
package dev.ltms.bridged.peer;
import java.util.Locale;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.peer;
package dev.ltms.bridged.peer;
/**
* An opaque handle returned by {@link PeerLauncher#spawn(SpawnRequest)}. The core routes on
@@ -0,0 +1,107 @@
package dev.ltms.bridged.peer;
import java.util.List;
import java.util.Set;
/**
* SPI for materializing a connected peer — the only way the bridge core creates or tears down
* a peer process. Every launcher is a first-party, in-tree adapter selected by (future) profile
* config; today's single adapter is the {@code ClaudeCodeLauncher} / Claude Code over herdr.
*
* <p>The core delegates spawn and teardown to this interface without knowing how the peer is set
* up. Environment variables, CLI flags, subscription guards, transport (herdr tab/pane) layout,
* and naming conventions are all adapter-private — the core sees only the returned
* {@link PeerHandle} whose {@code id()} is the registry/routing key.
*
* <p>The interface is a superset of what {@code SessionManager} and {@code Bridged.main} call
* on the concrete launcher today.
*/
public interface PeerLauncher {
/**
* The set of {@link Capability capabilities} this launcher declares. A peer whose profile
* opts into a git-forge token should include {@link Capability#SELF_PR}; the base set for
* the Claude Code herdr adapter is always {@code MID_TURN_ASK, WORKTREE, ORPHAN_REAP}.
*/
Set<Capability> capabilities();
/**
* The capabilities of the adapter that {@code profileName} resolves to (null/blank → the
* default profile, the same resolution {@link #spawn} uses). Distinct from {@link
* #capabilities()}, which unions every configured adapter: a caller that must know whether
* <em>this</em> profile's backend supports a capability — e.g. {@link Capability#SESSION_RESUME}
* before honoring {@link SpawnRequest#resumeSessionId()} — needs the per-profile answer, not
* the fleet-wide union, or a mixed fleet could OK a resume that lands on a non-supporting
* adapter (CB-584).
*
* @throws IllegalArgumentException if the profile is unknown and no default is configured
*/
Set<Capability> capabilitiesFor(String profileName);
/**
* {@code profileName}/requestedCwd null/blank → default resolution. Returns after the peer
* process is live (env + argv + placement complete). Never returns {@code null}.
*
* @param req the spawn parameters (profile, requested cwd, caller cwd)
* @return a handle whose {@link PeerHandle#id()} is the registry/routing key
* @throws IllegalArgumentException if the profile is unknown and no default is configured
*/
PeerHandle spawn(SpawnRequest req);
/**
* The configured worker profile names — the set of names {@code spawn(profileName)} accepts.
*/
Set<String> profiles();
/**
* The profile a no-argument {@link #spawn(SpawnRequest)} uses, or {@code null} if none is configured.
*/
String defaultProfile();
/**
* Resolve the effective working directory for a spawn {@code req} without actually spawning.
* Resolution order: requestedCwd → profile cwd → callerCwd → daemon cwd.
*
* @return the resolved absolute path, never null/blank
*/
String effectiveCwd(SpawnRequest req);
/**
* The parity-overlay file list for {@code profileName} (default list when unset). Used by
* worktree provisioning to copy config files into the isolated checkout before spawning.
*/
List<String> parityOverlay(String profileName);
/**
* The set of all agents this launcher currently tracks, transport-specific. Each element
* exposes at minimum a pane-like {@code id()} matching this launcher's {@link PeerHandle}
* scheme, plus transport-level status. Callers merge this set with the session registry to
* build a live roster view.
*/
List<?> list();
/**
* Reap orphaned peers left behind by a prior daemon process. Only peers whose naming scheme
* matches this launcher's and whose nonce differs from the current process are eligible.
* Best-effort: a failure to list or to stop any one peer is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
int reapOrphanWorkers();
/**
* Tear a peer down by its registry/routing key ({@link PeerHandle#id()}). Tolerates an
* already-gone peer. Also cleans up launcher-private resources (e.g. empty dedicated tabs)
* when safe to do so.
*/
void stop(String id);
/**
* Discard the context of the peer identified by {@code id}. Implementations must bypass normal
* bridge delivery/turn accounting. Unsupported peer kinds return {@code false} without sending
* a guessed command.
*
* @return {@code true} when a reset was sent and its status transition must settle before reuse
*/
boolean clearContext(String id);
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.peer;
package dev.ltms.bridged.peer;
/**
* Thrown when a {@link PeerLauncher} starts a peer process but the peer
@@ -1,4 +1,4 @@
package dev.ltms.fleet.peer;
package dev.ltms.bridged.peer;
/**
* What a caller asks for when spawning a peer.
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
import java.util.LinkedHashMap;
import java.util.Map;
@@ -8,7 +8,7 @@ import java.util.concurrent.ConcurrentHashMap;
import java.util.function.LongSupplier;
/**
* Where a credential (not a profile — see {@code FleetConfig.Profile#effectiveCredentialId()})
* Where a credential (not a profile — see {@code BridgedConfig.Profile#effectiveCredentialId()})
* sits out a cooldown after a {@code BACKEND_EXHAUSTED} classification (CB-578 stage B), so a fresh
* spawn does not walk straight back onto the account that just refused on a usage limit.
*
@@ -16,7 +16,7 @@ import java.util.function.LongSupplier;
* models on the same OpenAI account) share one quarantine — {@link #quarantine} one credential id
* and every profile whose {@code effectiveCredentialId()} equals it is quarantined too, without this
* class knowing anything about profiles at all. That mapping is the caller's job (see
* {@code CompositePeerLauncher} and {@code dev.ltms.fleet.inject.ExhaustionSink}).
* {@code CompositePeerLauncher} and {@code dev.ltms.bridged.inject.ExhaustionSink}).
*
* <p>The clock is injected ({@link LongSupplier}, conventionally {@code System::nanoTime} like
* {@code FleetHealthMonitor}), never read inline, so a quarantine's expiry is testable without a
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
/**
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
@@ -11,7 +11,7 @@ package dev.ltms.fleet.placement;
* a usage limit, not a transient capacity or reachability concern.
* <li>Weight 0 (CB-554): {@code fixed} is still automatic selection, so a profile the operator
* marked "never auto-select me" ({@code weight <= 0}) must be skipped here exactly as
* {@code weighted}/{@code round-robin} skip it — an explicit {@code fleet_spawn} naming
* {@code weighted}/{@code round-robin} skip it — an explicit {@code bridge_spawn} naming
* the profile is unaffected, only this automatic fallback walk.
* </ul>
* A fleet where nothing is ever quarantined or weight-0 never exercises either path, so today's
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
/**
* A profile (and, in CB-308, a host) that can be chosen by a {@link PlacementPolicy}.
@@ -21,12 +21,12 @@ public record PlacementCandidate(String profile, String host, float weight, Inte
/**
* True when this candidate carries an explicit {@code weight <= 0} (CB-554) and must be
* skipped by every automatic policy — the same way a quarantined or unreachable candidate is
* skipped. {@code FleetConfig.Profile}'s compact constructor already normalises "absent" to
* skipped. {@code BridgedConfig.Profile}'s compact constructor already normalises "absent" to
* {@code 1.0} and "negative" to {@code 0.0}, so this is a plain threshold check here; it does
* not need to distinguish "explicit 0" from "absent" itself.
*
* <p>Exclusion is about <em>automatic</em> selection only — an explicit
* {@code fleet_spawn{profile:"..."}} bypasses placement entirely and is unaffected.
* {@code bridge_spawn{profile:"..."}} bypasses placement entirely and is unaffected.
*/
public boolean excluded() {
return weight <= 0.0f;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.Set;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
/**
* Thrown when a {@link PlacementPolicy} has no candidate available. Kept as a distinct type so
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
/**
* Factory for the built-in placement policies.
@@ -1,7 +1,7 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
/**
* How {@code fleetd} chooses a worker profile when a spawn names none. Implementations are
* How {@code bridged} chooses a worker profile when a spawn names none. Implementations are
* deterministic and unit-testable; the caller (the composite launcher) handles failover retries.
*/
public interface PlacementPolicy {

Some files were not shown because too many files have changed in this diff Show More