Compare commits

..

2 Commits

Author SHA1 Message Date
Dai Ha 6eaa7aced3 CB-528a: spawn an OpenAI Codex peer via CodexLauncher 2026-08-10 21:09:18 +02:00
Dai Ha ef49835c4f CB-528: the CodexHome seam, ahead of the adapter that uses it
Codex reads everything from CODEX_HOME — config, credentials, sessions, skills,
plugins, state. Pointing a peer at the operator's own ~/.codex would hand it the
operator's tool surface and let it write into the operator's session history:
the same failure CB-525 exists to prevent on the Claude side, in a runtime where
there is no --mcp-config to neutralize.

Extracted as an interface rather than a launcher method because provisioning is
filesystem work with its own failure modes. The common one is a missing
credential, which Codex surfaces as an opaque 401 mid-turn instead of a spawn
error — so the contract says provision() must fail loudly there. Splitting it
also lets the launcher be tested without touching a real home directory.

Lands before the adapter so the launcher and the provisioner can be built
against a fixed seam instead of against each other.
2026-08-10 20:47:26 +02:00
252 changed files with 10822 additions and 35153 deletions
-24
View File
@@ -1,24 +0,0 @@
---
name: architect
description: Refine work into clear, independent units before implementation.
---
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You are an architect in this fleet. You refine work before anyone builds it: scope,
acceptance criteria, risks, and a unit split. You read the repo and write analysis.
You never commit production code and never open a pull request.
A design task is worked by two architects. Design alone first, then exchange and
say plainly where you disagree. Do not concede just to agree.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
Report only work you actually did and the real output of checks you ran. Do not
claim a result from a tool you could not use. The primary's IDE tools are not yours.
A mounted forge tool may use a blocked credential and fail by design.
The launcher provides the required bridge reply instructions for every member.
-29
View File
@@ -1,29 +0,0 @@
---
name: dev
description: Implement one assigned unit, test it, and open a pull request.
---
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You implement the one unit you were given and nothing else. Work in your assigned
git worktree and branch. Never check out, rebase onto, or push to `main`. Confirm
the worktree root and branch before you edit. Use only paths under that root.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
Implement the change and run the full required build in your worktree. Read the
complete output and report its real result. Do not hide failures with a pipe. State
only checks you actually ran. The primary's IDE tools are not yours. A mounted forge
tool may use a blocked credential and fail by design.
Stage only files you changed. Never use `git add -A` or `git add .`. Never commit
`.mcp.json` or `wiki/`. Commit with a clear message, push your branch, and open your
own pull request against `main`. Never merge.
Your handoff must name the pull request or why it was not created, the branch, the
files changed, the build result, and any caveat for review.
The launcher provides the required bridge reply instructions for every member.
-34
View File
@@ -1,34 +0,0 @@
---
name: reviewer
description: Review one assigned scope and report the most important real issue.
---
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You review the diff you were given. Report bugs, risks, and missing tests. You do
not change code.
Read the whole assigned scope before judging it. Review only that scope. If you see
something outside it, note it in one line and do not investigate it further. Do not
run the build. The owner makes changes and runs checks.
Use `fleet_ask{question}` only when a decision belongs to the lead, such as an
unclear requirement or two defensible fixes. Do not ask about something you can
decide by reading more code.
Report the single most important real issue in this form:
```
1. <path>:<line>
2. issue: <one sentence: what is wrong and why it matters>
3. fix: <one line: the concrete change>
4. severity: high | medium | low
```
If there is no real issue, report `NO ISSUE` and one line saying why. A clean review
is valid. Do not invent an issue. Use high for a wrong result, data loss, security,
or a hang or crash on a real path. Use medium for an edge-path bug or a correctness
risk under load or concurrency. Use low for clarity, a latent foot-gun, or a smell
with no current failure.
The launcher provides the required bridge reply instructions for every member.
+4 -4
View File
@@ -5,7 +5,7 @@ description: Implementer-role procedure for a bridged worker — verify your wor
# Implementer worker — procedure
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
never merge, never commit `.mcp.json` or `wiki/`) is in **`CLAUDE.md` → Bridge communication →
Worker** and already applies. This skill is only the *implement-and-hand-off procedure*.
@@ -85,7 +85,7 @@ host (`GITEA_HOST`) into your env for exactly this — the token can create a PR
merge**.
```bash
API="${GITEA_HOST%/}/api/v1/repos/fleet/fleetd/pulls"
API="${GITEA_HOST%/}/api/v1/repos/lms/claude-bridge/pulls"
BRANCH="$(git branch --show-current)"
curl -sS -X POST "$API" \
-H "Authorization: token ${GITEA_TOKEN}" \
@@ -103,7 +103,7 @@ fix it if the cause is yours (e.g. branch not pushed yet), and report the failur
inventing a URL. If `GITEA_TOKEN` is unset your profile was not granted PR-create: push the branch
and report its name so the lead opens the PR.
## 6. Hand off — what goes in `fleet_reply`
## 6. Hand off — what goes in `bridge_reply`
The reply is the entire handoff; the lead cannot see your terminal.
@@ -130,7 +130,7 @@ sequenceDiagram
I->>G: git push -u origin HEAD
I->>G: POST /pulls (GITEA_TOKEN) — open PR to main
G-->>I: html_url
I->>L: fleet_reply(PR url, branch, files, tests)
I->>L: bridge_reply(PR url, branch, files, tests)
Note over L,G: the lead reviews the PR and merges on green — you never merge
```
-180
View File
@@ -1,180 +0,0 @@
---
name: port-to-opencode
description: Make an OpenCode session a first-class participant in a Claude Code workspace — instructions, MCP servers, and secrets — without duplicating config. Use when onboarding opencode to a project that already has CLAUDE.md and .mcp.json, or when an opencode peer needs the same tools and rules as the Claude session.
---
# Porting a Claude Code workspace to OpenCode
**The headline: there is almost nothing to port.** OpenCode reads `CLAUDE.md` natively. The only
artifact you create is one `opencode.json` mapping MCP servers. Do not translate instructions, do
not generate a second rules file, and do not install a sync tool — every one of those makes the
workspace worse.
Everything below was verified against `opencode 1.18.16` and the OpenCode docs.
## 1. Know what you get for free
OpenCode's instruction search order:
```
1. walking up from cwd: AGENTS.md , then CLAUDE.md
2. global: ~/.config/opencode/AGENTS.md
3. Claude Code global: ~/.claude/CLAUDE.md (unless disabled)
```
*"The first matching file wins in each category."*
Consequences that decide the whole procedure:
- **A project `CLAUDE.md` is already read.** No port needed.
- **Your user-level `~/.claude/CLAUDE.md` is already read too.** Global preferences carry over.
- **An `AGENTS.md` in the repo SHADOWS `CLAUDE.md`.** If one exists from a previous Codex port,
**delete it** — otherwise opencode reads the stale translated copy instead of the real rules.
This is the single most likely way to get this wrong.
## 2. Create `opencode.json` for MCP servers only
Project config lives at `opencode.json` in the repo root; the global one is
`~/.config/opencode/opencode.json`. **Configs merge, they do not replace** — so machine-local
servers belong in the global file and shared ones in the project file.
Map each entry from `.mcp.json`:
| `.mcp.json` | `opencode.json` |
|---|---|
| `"type": "http"` / `"sse"` | `"type": "remote"`, `"url"` |
| `"type": "stdio"` | `"type": "local"`, `"command": ["bin", "arg"]` |
| `"command"` + `"args"` | single `"command"` array |
| `"env"` | `"environment"` |
| `"headers"` | `"headers"` |
```json
{
"$schema": "https://opencode.ai/config.json",
"instructions": ["CLAUDE.md"],
"mcp": {
"fleetd": { "type": "remote", "url": "http://127.0.0.1:8765/mcp", "enabled": true },
"context7": { "type": "remote", "url": "https://example.dev/mcp", "enabled": true,
"headers": { "Authorization": "Bearer {env:CONTEXT7_TOKEN}" } },
"gitea": { "type": "local", "command": ["gitea-mcp", "-t", "stdio"], "enabled": true,
"environment": { "GITEA_ACCESS_TOKEN": "{env:GITEA_ACCESS_TOKEN}" } }
}
}
```
Set `instructions` explicitly even though `CLAUDE.md` is found anyway — the fallback only applies
while no `AGENTS.md` exists, and being explicit survives someone adding one later.
## 3. Reference secrets, never embed them
OpenCode substitutes at load time, in both `headers` and `environment`:
```
{env:VARIABLE_NAME} value from the environment
{file:~/.secrets/token} value read from a file
```
**No credential ever belongs in `opencode.json`.** With `{env:…}` there is no reason to write one,
which is what makes this file safe to commit — and it must be committable, because a peer running
in a git worktree receives tracked files only.
**But `{env:…}` reads OpenCode's *process* environment — and OpenCode has no env store of its
own.** A Claude Code `env` block in `~/.claude/settings.json` or `.claude/settings.local.json` does
**not** reach it: those files are Claude Code's, and the variables exist only inside processes
Claude Code spawned. Verify with a clean login shell, not the shell your agent hands you:
```bash
env -u MY_TOKEN zsh -lc 'echo "${MY_TOKEN:-NOT IN PROFILE}"'
```
So `{env:…}` only works if something puts the variable in the environment first. Pick the supply
route by who launches opencode:
| Launcher | Route |
|---|---|
| a human, from a terminal | one central store, sourced by the login shell |
| a spawner (bridge, CI, IDE) | `{env:…}`, with the spawner injecting the variable |
**For the human case, keep every credential in one file the login shell sources.** Here that file
is `${SHARED_ENV}/tools/secrets.sh`, sourced from `${SHARED_ENV}/.ltms`, kept at mode 600 and never
committed. `opencode.json` then names variables and holds no values:
```json
"headers": { "Authorization": "Bearer {env:CONTEXT7_TOKEN}" }
```
Both routes end at the same syntax, and that is the point. The file does not change when a human
launches opencode instead of the bridge.
**`{file:…}` also works, and this project moved away from it.** Opencode resolves a relative
`{file:}` path against the project root, so a gitignored `.secrets/` beside `opencode.json` needs no
shell setup at all. It did not fail; the problem is that it makes a second copy of the token. The
same secret then lives in two places, and the copy you forget is the one that leaks or goes stale.
One store with many references is easier to rotate and to audit.
**One catch survives either choice, so state it out loud:** what a human's shell exports does not
reach a spawned peer, and neither does a gitignored `.secrets/` — a git worktree receives tracked
files only. Spawned peers must be fed through `{env:…}` by whatever launches them. Check the
variable *names* match: a spawner often injects under a different name than your shell uses, and
the config has no fallback. In this repo the bridge goes further: it neutralizes a worktree's
`opencode.json`, so a member cannot inherit the primary's credentials by accident.
## 4. Do not port machine-local MCP servers
IDE indexes, language servers, editor bridges — anything bound to *your* checkout — stay out of the
project file. Put them in `~/.config/opencode/opencode.json` if you want them personally.
A committed project config reaches every worktree. A worker that mounts servers whose paths point
into the primary's checkout will edit the primary's files while building in its own — every build
passes, every change lands in the wrong tree.
Port the servers the work needs. Leave the rest.
## 5. Verify against the running agent, not the file
A file on disk proves nothing about what the agent loaded.
```bash
opencode run "In one line: state a rule from this project's instructions."
```
The answer must reflect the actual `CLAUDE.md`. If it answers generically, the instructions did not
reach the model and everything after this is built on sand.
Then confirm the tools are mounted:
```bash
opencode mcp list
```
**"connected" does not mean "working".** A stdio server with a missing credential still completes
the MCP handshake and reports green; only a real tool call reveals it. Verified: `gitea` showed
`✓ connected` with no token, then failed the first call with `token is required`. A remote server
is more honest (`⚠ needs authentication`), but do not rely on that difference — **exercise one
authenticated tool per server**:
```bash
opencode run "Call <server>'s <tool>. Report the result or the exact error. One line."
```
Check too that no server you deliberately withheld is present.
## 6. Report
State what you changed, which servers crossed and which you withheld and why, and quote the
verification answer verbatim. If any server failed to connect, say so plainly — a partially mounted
peer is worse than a missing one, because it looks configured.
## Gotchas
- **A leftover `AGENTS.md` silently wins over `CLAUDE.md`.** Check for one before anything else.
- **`opencode.json` is merged, not overridden** — a global entry and a project entry with the same
server name both matter; keep names distinct unless you intend to layer them.
- **`enabled: false`** turns a server off without deleting its config — prefer it over removal when
you may want the server back.
- **OAuth-based servers** store tokens in `~/.local/share/opencode/mcp-auth.json` after
`opencode mcp auth <server>`; that is machine state, never config to commit.
- **Skills and subagents do not port.** OpenCode uses its own agent markdown under
`.opencode/agents/`; `.claude/skills/**` is not read. If a delegation brief tells a peer to load a
skill by name, that instruction has no effect on an opencode peer — spell the procedure out in the
brief, or author the equivalent agent file.
+3 -3
View File
@@ -5,7 +5,7 @@ description: Reviewer-role procedure for a bridged worker — how to work a revi
# Reviewer worker — procedure
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and already applies. This
skill is only the *review procedure*: how to work the scope, and the exact shape of what you
send back.
@@ -24,13 +24,13 @@ wrong.
covers what it was given.
- Do **not** edit files or run the build. You review; the owner acts.
## 3. Reach for `fleet_ask` only for a genuine fork
## 3. Reach for `bridge_ask` only for a genuine fork
Ambiguous requirement, a missing acceptance criterion, "intended or a bug?", or two defensible
fixes with different consequences — those are the lead's call, and guessing produces a
confident-but-wrong finding. Anything you could settle by reading more code is yours to settle.
## 4. The finding — what goes in `fleet_reply`
## 4. The finding — what goes in `bridge_reply`
Report the **single most important** real issue in the scope, in these four lines, under
~90 words:
+2
View File
@@ -69,6 +69,8 @@ jobs:
env:
RABBITMQ_DEFAULT_USER: guest
RABBITMQ_DEFAULT_PASS: guest
ports:
- 5672:5672
env:
# Service containers are reachable from the job by their network alias on their internal port.
AMQP_URI: amqp://guest:guest@rabbitmq:5672
-21
View File
@@ -1,21 +0,0 @@
# No secret belongs in this repo any more: every credential lives in one shell-level store
# (${SHARED_ENV}/tools/secrets.sh), and opencode.json reads it as {env:...}. This line stays as a
# backstop, so a workspace-scoped copy that someone re-creates by habit still cannot be committed.
.secrets/
# Settings backups inherit the env block — and secrets with it.
.claude/settings.local.json.bak*
# The default profile parityOverlay copies these primary→worktree, so they appear in EVERY worker
# worktree. Two reasons they must be ignored. They hold environment values, which is reason enough.
# And since CB-576 a release preserves any worktree that `git status --porcelain` calls dirty —
# untracked files included, deliberately. An untracked overlay file would therefore make every
# COMPLETED release preserve its worktree, and worktrees would pile up with no error to notice.
.env
.envrc
# Daemon runtime artefacts. bridged appends its log wherever it is launched from, so both the
# repo root and bridged/ collect one; neither belongs in git.
bridged.out
bridged/bridged.out
logs/
+1 -1
View File
@@ -1,3 +1,3 @@
[submodule "wiki"]
path = wiki
url = ssh://git@git.ltms.dev:2224/fleet/fleetd.wiki.git
url = ssh://git@git.ltms.dev:2224/lms/claude-bridge.wiki.git
-24
View File
@@ -1,24 +0,0 @@
---
description: Refine work into clear, independent units before implementation.
mode: primary
---
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You are an architect in this fleet. You refine work before anyone builds it: scope,
acceptance criteria, risks, and a unit split. You read the repo and write analysis.
You never commit production code and never open a pull request.
A design task is worked by two architects. Design alone first, then exchange and
say plainly where you disagree. Do not concede just to agree.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
Report only work you actually did and the real output of checks you ran. Do not
claim a result from a tool you could not use. The primary's IDE tools are not yours.
A mounted forge tool may use a blocked credential and fail by design.
The launcher provides the required bridge reply instructions for every member.
-29
View File
@@ -1,29 +0,0 @@
---
description: Implement one assigned unit, test it, and open a pull request.
mode: primary
---
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You implement the one unit you were given and nothing else. Work in your assigned
git worktree and branch. Never check out, rebase onto, or push to `main`. Confirm
the worktree root and branch before you edit. Use only paths under that root.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
Implement the change and run the full required build in your worktree. Read the
complete output and report its real result. Do not hide failures with a pipe. State
only checks you actually ran. The primary's IDE tools are not yours. A mounted forge
tool may use a blocked credential and fail by design.
Stage only files you changed. Never use `git add -A` or `git add .`. Never commit
`.mcp.json` or `wiki/`. Commit with a clear message, push your branch, and open your
own pull request against `main`. Never merge.
Your handoff must name the pull request or why it was not created, the branch, the
files changed, the build result, and any caveat for review.
The launcher provides the required bridge reply instructions for every member.
-34
View File
@@ -1,34 +0,0 @@
---
description: Review one assigned scope and report the most important real issue.
mode: primary
---
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You review the diff you were given. Report bugs, risks, and missing tests. You do
not change code.
Read the whole assigned scope before judging it. Review only that scope. If you see
something outside it, note it in one line and do not investigate it further. Do not
run the build. The owner makes changes and runs checks.
Use `fleet_ask{question}` only when a decision belongs to the lead, such as an
unclear requirement or two defensible fixes. Do not ask about something you can
decide by reading more code.
Report the single most important real issue in this form:
```
1. <path>:<line>
2. issue: <one sentence: what is wrong and why it matters>
3. fix: <one line: the concrete change>
4. severity: high | medium | low
```
If there is no real issue, report `NO ISSUE` and one line saying why. A clean review
is valid. Do not invent an issue. Use high for a wrong result, data loss, security,
or a hang or crash on a real path. Use medium for an edge-path bug or a correctness
risk under load or concurrency. Use low for clarity, a latent foot-gun, or a smell
with no current failure.
The launcher provides the required bridge reply instructions for every member.
+63 -171
View File
@@ -4,52 +4,47 @@
> **Canonical block.** Everything down to §Layering is the portable bridge charter, copied verbatim
> into every project that mounts the bridge MCP. Keep it byte-identical with the template in the
> wiki ([Use Cases](https://git.ltms.dev/fleet/fleetd/wiki/7-Use-Cases) → *The portable
> wiki ([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable
> CLAUDE.md block*); improvements go to the template first, then out to each project. Anything
> specific to *this* repo lives under §Project addendum below, never inline above it.
If no `fleet_*` MCP tools are mounted in this session, this section does not apply — skip it.
If no `bridge_*` MCP tools are mounted in this session, this section does not apply — skip it.
`bridged` is the **sole communication gateway** between agents here. The orchestrating session (the
**primary**) and every delegated peer (a **member**) mount the *same* MCP server and talk only
through its `fleet_*` tools. No session addresses a peer, a broker, or the network directly.
**primary**) and every delegated peer (a **worker**) mount the *same* MCP server and talk only
through its `bridge_*` tools. No session addresses a peer, a broker, or the network directly.
### Which role am I? — settle this before acting
**Every role reads this file.** A member runs in a git worktree of this same repo, so it inherits
**Both roles read this file.** A worker runs in a git worktree of this same repo, so it inherits
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
**Call `fleet_whoami`.** It returns `primary`, `worker`, or `architect`, resolved by the daemon from
your connection — unforgeable, and the same resolution its authorization gate uses. A worker also
carries its `sessionId`, `profile`, `worktree` and `branch`; an architect carries the slot name it
was bound to. Don't infer what you can ask.
**Call `bridge_whoami`.** It returns `{"role":"primary"}` or `{"role":"worker","sessionId":…,
"profile":…,"worktree":…,"branch":…}`, resolved by the daemon from your connection — unforgeable,
and the same resolution its authorization gate uses. Don't infer what you can ask.
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
fires: the reply charter in your system prompt (*"You are a spawned member in the
claude-bridge fleet"*) ⇒ **spawned member**; fleet tools prefixed `mcp__fleet__*` ⇒ **spawned
member** (the launcher fixes that mount name; a primary's mount is named by whoever wrote its
`.mcp.json`, so it varies — and a member spawned before CB-632 still says `mcp__bridge__*`); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect —
only `fleet_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate —
loud and self-correcting — while a member acting as the primary ends its turn with no `fleet_reply`,
and the sender silently receives nothing. Fail toward the recoverable error.
fires: the reply charter in your system prompt (*"You are an off-subscription worker in the
claude-bridge fleet"*) ⇒ **worker**; bridge tools prefixed `mcp__bridge__*` ⇒ **worker** (the
launcher fixes that mount name; a primary's mount is named by whoever wrote its `.mcp.json`, so it
varies); `ANTHROPIC_BASE_URL` set ⇒ **worker** (Claude-model workers run on a clean env, so its
*absence* proves nothing). **Still unsure ⇒ act as a worker.** The two mistakes are not symmetric: a
primary acting as a worker is refused by the authorization gate — loud and self-correcting — while a
worker acting as the primary ends its turn with no `bridge_reply`, and the sender silently receives
nothing. Fail toward the recoverable error.
### Invariants — both roles, no exceptions
1. **Never set, export, or forward `ANTHROPIC_BASE_URL`** (or `ANTHROPIC_AUTH_TOKEN`). The primary
stays on subscription; only the bridge puts a member off it, at spawn. Mounting the bridge must
stays on subscription; only the bridge puts a worker off it, at spawn. Mounting the bridge must
never move a session across that boundary.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `fleet_*` call is silently discarded.
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead or architect**;
reply/ask are only-as-itself — any peer may answer for its own pane, and for no other. A call
outside your role is refused, not queued.
cannot act as another session. Spawn/stop/send/drain are primary-only; reply/ask are
worker-only-and-only-as-itself. A call outside your role is refused, not queued.
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
deliverable, and a send waits on that gate for ~60s and then fails without ever reaching its pane.
re-send because a call looks slow — the bridge delivers when the peer is `idle`/`blocked`.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
@@ -57,10 +52,10 @@ and the sender silently receives nothing. Fail toward the recoverable error.
**Delegate by default — that is the job.** With the bridge mounted you are an orchestrator on a
metered subscription, and workers are cheap, parallel, and disposable. The default answer to "who
does this?" is **a worker**, not you. Reach for `fleet_send` before you reach for `Edit`. The steps
does this?" is **a worker**, not you. Reach for `bridge_send` before you reach for `Edit`. The steps
below are the procedure — run them in order, every task, not only the big ones.
0. **Know your role** — `fleet_whoami`, once per session, before anything else.
0. **Know your role** — `bridge_whoami`, once per session, before anything else.
1. **Split.** Write the unit list. Every unit carries: scope · the files or PR in question ·
acceptance criteria · exactly what to report back. A unit with no acceptance criteria is not
ready to delegate — refine it or keep it.
@@ -70,22 +65,20 @@ below are the procedure — run them in order, every task, not only the big ones
delegate. The keep-list is closed: the conversation with the user, decomposition and planning,
the final judgment call, verification, merges, and anything that depends on context only you
hold. Nothing else is yours by default.
3. **Spawn every delegated unit first** — `fleet_spawn{profile, worktree:true, ticket}`, one per
3. **Spawn every delegated unit first** — `bridge_spawn{profile, worktree:true, ticket}`, one per
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model and cost, not in
tier, so the default is rarely what you want.
4. **Then send them all** — `fleet_send{sessionId, content, wait:false}`. Line 1 of every brief is
4. **Then send them all** — `bridge_send{sessionId, content, wait:false}`. Line 1 of every brief is
`Load the <name> skill.` naming the worker's playbook; those skills are opt-in and that line is
what makes them reliable. Where the project ships no such skill, spell the procedure out in the
brief instead. The brief is self-contained — the worker sees your message and the repo, nothing
of your context, your plan, or your screen.
5. **Collect** — `fleet_poll{ticket}` → `fleet_ack{target, msgId}`. Answer a worker's `fleet_ask`
with `fleet_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`fleet_status`, never by reading its terminal; it also reports an open question and the `turnId`
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
forge tools it appears to have hold a blocked credential and fail, and a piped command
(`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a fact.
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`bridge_status`, never by reading its terminal.
6. **Verify yourself.** Re-run the build and the checks. A worker mounts only the bridge MCP and
cannot run your other tooling, and a piped command (`… | tail`) hides failures behind a zero
exit — never promote a worker's "clean" to a fact.
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
@@ -93,11 +86,11 @@ below are the procedure — run them in order, every task, not only the big ones
read it yourself.
8. **Adjudicate, merge, tear down — yours alone.** Read the diff yourself: fully if it is small,
targeted at the reported findings and the risky paths if it is large. Reviewer findings direct
your attention; they never substitute for it. Then merge, then `fleet_stop{paneId}`.
your attention; they never substitute for it. Then merge, then `bridge_stop{paneId}`.
**Steps 3 and 4 are separate on purpose** — spawning and sending in one loop is how parallel work
silently becomes serial, and it is the most common way this layer is wasted. For the same reason,
prefer `wait:false` + `fleet_poll` for anything non-trivial: a blocking `fleet_send` is capped by
prefer `wait:false` + `bridge_poll` for anything non-trivial: a blocking `bridge_send` is capped by
*your own* MCP client call timeout (~60s), well below the task's real runtime.
**Delegating does not delegate responsibility.** Workers open PRs; you are the gate. Never delegate
@@ -105,65 +98,28 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
| Intent | Tool |
|---|---|
| Confirm your own role | `fleet_whoami` |
| See backends available | `fleet_profiles` |
| Start a member | `fleet_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` → `sessionId` + `paneId` |
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) · one peer's state: `fleet_status{sessionId}` |
| Delegate (blocking) | `fleet_send{sessionId, content}` |
| Delegate (long task) | `fleet_send{sessionId, content, wait:false}` → ticket → `fleet_poll{ticket}` |
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `fleet_reply{content}` — the one case a lead replies |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
| Tear down a member | `fleet_stop{paneId}` |
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a worker | `bridge_spawn{profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` · one worker's state: `bridge_status{sessionId}` |
| Delegate (blocking) | `bridge_send{sessionId, content}` |
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
| Answer a worker's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
| Tear down | `bridge_stop{paneId}` |
### Lead ↔ lead — coordinate, never delegate
`fleet_list` returns `leads` alongside `members`; your own row carries `self: true`. Every other row
is a peer — an orchestrator with its own context, its own members, and its own judgment. An empty
`members` array means no members are spawned; it says nothing about peers.
**A lead never assigns a task to another lead.** Work goes to members — only ever downward, never
sideways. Sending a peer a brief with acceptance criteria is a category error: a brief is a member's
artefact, and a peer is not yours to task. If a unit needs doing and it falls in your area, spawn a
member and delegate it yourself; if it falls in the peer's area, say so and let the peer assign it.
The traffic between leads is coordination and nothing else:
1. **Divide the map, not the work.** Agree who owns which area, then each of you assigns inside your
own. Split by **context ownership** — whoever already holds the context owns that area — and say
who takes what, in one message, before either of you starts. Two leads silently working the same
unit is the failure mode here, and neither notices until the merge.
2. **Share findings, hazards, and corrections.** What you have already discovered, what broke, what
the next person will trip on. This is the traffic that actually pays for the channel: it costs one
message and saves a peer a rediscovery.
3. **Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
against the code, and re-run the build. A peer's correction gets the same treatment — right or
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
Being messaged by a peer does not make you its worker: answer with `fleet_reply`, and push back on
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
you.
### Member (worker or architect) — the turn contract
### Worker — the turn contract
1. **Load the playbook skill the lead named** before doing anything else.
2. **Do the assigned scope only.** Note anything you spot outside it in one line; don't go hunt it.
3. **`fleet_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
3. **`bridge_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
defensible fixes, "bug or intended?"). It blocks and you resume the *same* turn with the answer.
Don't ask what you could decide yourself.
4. **End the turn with exactly one `fleet_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `fleet_reply` ⇒ the sender gets nothing and the exchange stalls.
Do **not** lean on the completion fallback to carry your answer for you: when you end a turn
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
lead with its end cut off.
5. **Report honestly.** State only what you actually ran and its real output, including failures,
and never claim the result of a check you had no way to run. **Measure your own tools; do not
assume them.** What you mount depends on your backend: an opencode member gets the bridge and
nothing else, while a Claude Code member also inherits the operator's user-scope MCP servers,
which the bridge never chose for you. Two rules follow. The primary's IDE tooling is still not
yours, whatever you see. And **a mounted tool is not a working tool** — the forge server you may
find there holds a deliberately blocked credential and fails every call, by design.
4. **End the turn with exactly one `bridge_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `bridge_reply` ⇒ the sender gets nothing and the exchange stalls.
5. **Report honestly.** State only what you actually ran and its real output, including failures.
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
so never claim the result of a check you had no way to run.
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
@@ -171,91 +127,27 @@ you.
| Layer | Scope | Reaches |
|---|---|---|
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `fleet_reply`* | every spawned member, at launch, every peer kind — never a lead |
| **this section** | protocol + orchestration policy | primary **and** every member that reads the repo — tracked in git, so worktrees inherit it |
| role agent definition files | role contract and per-job procedure | a member whose launcher binds its role to the matching file in its worktree |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a member told to load one |
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every worker, at launch, every peer kind |
| **this section** | protocol + orchestration policy | primary **and** every Claude worker — tracked in git, so worktrees inherit it |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a worker told to load one |
| the bridge's own docs | design detail, flows, error model | on demand |
A rule belongs in **exactly one** layer — the outermost one that must obey it. A member without a
repo checkout still gets the launcher's reply charter, which is why that one rule stays there.
Peers that don't read `CLAUDE.md` (non-Claude adapters) get the charter only, so any rule *they*
must obey belongs in the charter, not here.
A rule belongs in **exactly one** layer — the outermost one that must obey it. Peers that don't read
`CLAUDE.md` (non-Claude adapters) get the charter only, so any rule *they* must obey belongs in the
charter, not here.
## Project addendum — claude-bridge (not part of the canonical block)
- **This repo is the bridge.** The daemon is `bridged`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/FleetMcp` (tools), `auth/Authz` (the role table),
and the code behind the rules above is `mcp/BridgeMcp` (tools), `auth/Authz` (the role table),
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
`reviewer` (scoped review → one structured finding). Name one in every delegation.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace).
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **Flows and the error model** — rendezvous, `fleet_ask`, detached delivery, turn-done fallback —
are diagrammed in `docs/MCP-Contract.md` **§6 only**. The rest of that page is a pre-build design
doc whose tool names, parameter names and REST paths never caught up with the code, so do not use
it as the tool reference (CB-609). Section 6 is kept out of this file because this file loads into
every session's context.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `bridged` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-bridged.sh --check # report state, change nothing
scripts/redeploy-bridged.sh # build, confirm drain, restart, verify
scripts/redeploy-bridged.sh --yes # skip the drain prompt (fleet already checked)
```
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `bridged listening` line at the end of
`bridged/bridged.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-bridged.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
session's context.
### The prompt is part of the product — update it with the code (mandatory)
@@ -268,10 +160,10 @@ Before you call any work done, check the row that matches what you touched:
| You changed… | Re-read and update… |
|---|---|
| a `fleet_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
| a `bridge_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
| `Authz` / the role table | invariant 3, and the primary-only vs worker-only claims |
| `ConnectionIdentity` / how a caller is resolved | the `fleet_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__fleet__*`), and the layering table's top row |
| `ConnectionIdentity` / how a caller is resolved | the `bridge_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__bridge__*`), and the layering table's top row |
| the injector / status gating | invariant 4 |
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
@@ -287,7 +179,7 @@ is a Roadmap line. A change that touches none of the three earns no entry, and t
outcome rather than an omission.
Then **propagate**: the block in this file and the template in the wiki
([Use Cases](https://git.ltms.dev/fleet/fleetd/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
block*) must stay byte-identical, and other projects carrying the block need the same edit. Verify
rather than trust:
+29 -33
View File
@@ -9,33 +9,32 @@ Sibling of [`crush-bridge`](https://git.ltms.dev/systems/vms) (which drives a he
process*, so it inherits `CLAUDE.md`, hooks, skills, and MCP — just pointed at a
cheaper/local model.
## Leading approach — herdr-centric message server (`fleetd`)
## Leading approach — herdr-centric message server (`bridged`)
A small always-on message server, **`fleetd`**, controls
A small always-on message server, **`bridged`**, controls
[herdr](https://herdr.dev) (an agent multiplexer) over its Unix-socket API and exposes a
clean 2-way messaging API as an **MCP server that both the primary and the workers mount** —
one unified Claude setup and the **sole communication gateway** (REST/SSE stays for non-Claude
clients; any broker is `fleetd`-internal, below the gateway).
herdr owns the PTYs, multiplexing, persistence, and **agent-status events**; `fleetd` owns
clients; any broker is `bridged`-internal, below the gateway).
herdr owns the PTYs, multiplexing, persistence, and **agent-status events**; `bridged` owns
policy (subscription boundary, session lifecycle, status-gated delivery) and the client
contract. A Claude member launches with `ANTHROPIC_BASE_URL` pointed at the gateway,
`https://llm.ltms.dev/anthropic`, plus a bearer token; the lead stays env-clean and calls
`fleetd`'s MCP tools. See the wiki's **[13 User Guide](wiki/13-User-Guide.md)** to run it.
contract. The worker `claude` launches with `ANTHROPIC_BASE_URL=https://ollama.ltms.dev` + a
bearer token; the primary Opus stays env-clean and calls `bridged`'s MCP tools.
```mermaid
flowchart LR
OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
subgraph BD["fleetd — standalone daemon (not a claude process)"]
subgraph BD["bridged — standalone daemon (not a claude process)"]
SRV["SERVER face<br/>MCP · REST/SSE · policy"]
CLI["CLIENT face<br/>status-gated injector · herdr socket"]
SRV --> CLI
end
HERDR["herdr<br/>panes · agent-status"]
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
M["llm.ltms.dev<br/>(the one gateway)"]
M["ollama.ltms.dev<br/>(worker model)"]
OPUS -->|"MCP fleet_send (blocks)"| SRV
W -.->|"MCP fleet_reply"| SRV
OPUS -->|"MCP bridge_send (blocks)"| SRV
W -.->|"MCP bridge_reply"| SRV
CLI -->|"Unix socket<br/>send_text · events.subscribe"| HERDR
HERDR -->|"drives PTY"| W
W -->|"inference"| M
@@ -47,27 +46,24 @@ flowchart LR
```
- **Subscription boundary:** the *primary* never sets `ANTHROPIC_BASE_URL` (stays on
Pro/Max). Only the *secondary* process is off-subscription — and `fleetd` itself is a
Pro/Max). Only the *secondary* process is off-subscription — and `bridged` itself is a
plain daemon (no Anthropic quota), so it may poll/subscribe freely.
- **One gateway (unified MCP setup):** `fleetd` is the **sole communication path** for every
- **One gateway (unified MCP setup):** `bridged` is the **sole communication path** for every
Claude session. Primary and workers each mount it as an MCP server (one `claude mcp add`
line, same on both) and talk over MCP tools — `fleet_send` / `fleet_reply` /
`fleet_status` (with `fleet_ask` planned for the blocked-worker path). **No Claude session
ever addresses a broker, a peer, or the network
directly**; any queue is `fleetd`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so
mounting the bridge is subscription-safe by construction.
**Tool naming:** the tools were renamed from `bridge_*` to `fleet_*` (CB-622). The daemon
still answers the old `bridge_*` names for one release, but they are deprecated — use the
`fleet_*` names.
- **How the primary consumes a reply:** a single **blocking MCP call** (`fleet_send`);
`fleetd` holds it open until the worker calls `fleet_reply` or its turn hits
line, same on both) and talk over MCP tools — `bridge_send` / `bridge_reply` /
`bridge_status` (with `bridge_ask` planned for the blocked-worker path). **No Claude session
ever addresses a broker, a peer, or the network
directly**; any queue is `bridged`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so
mounting the bridge is subscription-safe by construction.
- **How the primary consumes a reply:** a single **blocking MCP call** (`bridge_send`);
`bridged` holds it open until the worker calls `bridge_reply` or its turn hits
`agent_status=done`, then returns the reply as the tool result. No cross-turn busy-poll, so
no quota burn. SSE is an optional side-channel for humans/dashboards watching status.
- **Worker → primary** rides `fleetd`'s **MCP rendezvous** — the reply resolves the primary's
blocking call (or, for detached work, `fleetd` **injects the primary's idle pane** when it's
- **Worker → primary** rides `bridged`'s **MCP rendezvous** — the reply resolves the primary's
blocking call (or, for detached work, `bridged` **injects the primary's idle pane** when it's
ready), so *no keystroke-into-primary and no broker are involved, even single-host*. The one
exception: a split-host primary that isn't a herdr pane wakes via its own `Stop`-hook, which
polls **`fleetd`** (never a broker). See the wiki for the two topologies.
polls **`bridged`** (never a broker). See the wiki for the two topologies.
- **Different model per process** sidesteps Claude Code's lack of per-subagent provider
routing — the worker isn't a subagent, it's its own configured process.
- **AgentAPI** ([`coder/agentapi`](https://github.com/coder/agentapi)) is retained only as a
@@ -76,11 +72,11 @@ flowchart LR
## Docs
Full design, setup, and operations live in the **[wiki](https://git.ltms.dev/fleet/fleetd/wiki)**,
Full design, setup, and operations live in the **[wiki](https://git.ltms.dev/lms/claude-bridge/wiki)**,
vendored here as a submodule under [`wiki/`](./wiki):
```bash
git clone --recurse-submodules ssh://git@git.ltms.dev:2224/fleet/fleetd.git
git clone --recurse-submodules ssh://git@git.ltms.dev:2224/lms/claude-bridge.git
# or, after a plain clone:
git submodule update --init
```
@@ -90,7 +86,7 @@ Gitea wiki.
## Status
🟢 **Implemented & dogfooded** — the herdr-centric **`fleetd`** message server is built and in
🟢 **Implemented & dogfooded** — the herdr-centric **`bridged`** message server is built and in
real use: an Opus primary delegates tasks to off-subscription workers that reply through the
bridge (code reviews delegated this way have produced committed bug fixes). Selected as the
primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a
@@ -101,9 +97,9 @@ tests run separately via `mvn test -Pcontract`):
- **Core gateway** — herdr socket client (contract-tested vs live 0.7.0); guard-checked worker
spawn with `ANTHROPIC_BASE_URL` injected only into the worker's env; status-gated injector;
blocking `fleet_send` with reply rendezvous; MCP server as a thin adapter over the REST core.
- **MCP tools** — `fleet_send` / `fleet_reply` / `fleet_status` (messaging) and `fleet_spawn`
/ `fleet_list` / `fleet_stop` / `fleet_profiles` / `fleet_poll` (fleet). Caller identity is
blocking `bridge_send` with reply rendezvous; MCP server as a thin adapter over the REST core.
- **MCP tools** — `bridge_send` / `bridge_reply` / `bridge_status` (messaging) and `bridge_spawn`
/ `bridge_list` / `bridge_stop` / `bridge_profiles` / `bridge_poll` (fleet). Caller identity is
connection-based (loopback peer PID → herdr pane), so the same mount serves primary and workers.
- **Delivery reliability** — completion fallback (a confirmed `working→idle` turn resolves a
send); async fire-and-poll (beats the caller's MCP call timeout for long tasks); and failure
@@ -111,7 +107,7 @@ tests run separately via `mvn test -Pcontract`):
- **Fleet** — multiple worker profiles, each with an independent base_url guard check; workers
inherit the primary's working directory (never `$HOME`); a readiness gate holds delivery until
a worker's Claude has connected the bridge MCP (no paste lost into its boot window).
- **Blocked-worker path** — `fleet_ask` reverse rendezvous: a worker pauses its delegated turn to
- **Blocked-worker path** — `bridge_ask` reverse rendezvous: a worker pauses its delegated turn to
ask the primary and resumes the *same* turn with the answer (CB-205).
- **Session lifecycle** — session manager with spawn/reuse/recycle, `idle_ttl` reaper, `context_cap`,
and graceful drain on shutdown (CB-301/CB-303); per-worker git worktrees on their own branch with
+1 -3
View File
@@ -2,9 +2,7 @@
target/
dependency-reduced-pom.xml
# Local runtime config (copy from fleetd.example.yaml). Both names: the live file is still
# bridged.yaml until the cutover, and Fleetd reads either one.
fleetd.yaml
# Local runtime config (copy from bridged.example.yaml)
bridged.yaml
# CB-505 audit trail + daemon stdout/stderr — runtime records, never source
+244
View File
@@ -0,0 +1,244 @@
# bridged configuration (example). Copy to bridged.yaml and adjust.
#
# bridged is the sole gateway between primary/worker Claude sessions and herdr.
# It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL.
# REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token`
# below — bridged REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
bind:
host: 127.0.0.1
port: 8765
# API authentication (CB-501). Governs how a caller that is NOT an on-host worker pane proves it
# is the primary. Worker identity never depends on this: a loopback peer PID that maps to a herdr
# pane is unforgeable and is always honoured, so turning auth on cannot lock the fleet out.
#
# mode: loopback-trust → DEFAULT, and the historical behaviour: any loopback caller that is not
# a worker is the primary, no credential needed. Sound ONLY because the
# OS refuses remote connections to a loopback socket.
# mode: token → such a caller must send `Authorization: Bearer <token>`; without it it
# is anonymous and authorized for nothing. REQUIRED for a non-loopback
# bind — the daemon fails fast otherwise, because "unauthenticated ⇒
# primary" on a reachable port would hand spawn/stop/send to anyone.
# tokenEnv → host env var holding the token (never the literal value). Default
# BRIDGED_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
#
# TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and
# let it own certificate lifecycle, e.g.
# location / { proxy_pass http://127.0.0.1:8765; proxy_set_header Authorization $http_authorization; }
# The broker link gets TLS from its own URI (amqps://…) — see `broker` below.
# auth:
# mode: token
# tokenEnv: BRIDGED_API_TOKEN
# Optional pinned primary terminal (CB-307). Names the herdr pane the PRIMARY itself runs in:
# a caller whose connection maps to this pane resolves as the primary (no credential needed —
# the pane mapping is as unforgeable as a worker's), and reply nudges are pushed to it.
# REQUIRED when the primary runs inside a herdr pane — without it the pane match reads the
# primary as a worker and refuses spawn/send/stop. Get the id from bridge_whoami; re-pin if
# the primary moves panes.
# primary:
# terminal: term_0123456789abcd
# pushReminders: 5 # max nudges before giving up (default 5)
# pushBackoffMs: 15000 # delay between nudges (default 15000)
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
# How worker sessions are spawned. Define one or more named profiles (backends) under
# `workers`; each key is the profile name (also the ccs profile). `defaultWorker` picks
# which one a no-argument spawn uses (bridge_spawn with no profile / POST /workers).
#
# Shared knobs (placement/workspace/tabLabel) can be repeated per profile; they usually match.
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
# Use `pane` for the legacy behaviour (split the focused tab).
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
# (--append-system-prompt) as launch flags; nothing is written to the profile.
# tokenEnv → host env var holding the worker's auth token (value never stored in config);
# omit for a backend that needs no token (e.g. a local ollama).
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
# cwd on an MCP spawn, else the daemon's cwd — never $HOME. See
# docs/Worker-Startup-and-Trust.md.
# configDir → CLAUDE_CONFIG_DIR for the worker, so it inherits that profile's
# skills/MCP/hooks. Omit to leave the worker on the host default.
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
# worktree sees the same local config (CB-301-ext). Omit for the default set:
# [.claude/settings.local.json, .env, .envrc].
#
# Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher
# mounts — the bridge, and nothing else. Replicating the primary's MCP config
# handed a worker the primary's IDE servers, which are bound to the primary's
# checkout, so its navigation returned paths OUTSIDE its own worktree: one
# worker made all 59 of its edits in the primary tree while compiling its
# worktree, and every build it ran was of code that did not contain them.
# bridged neutralizes a provisioned worktree's .mcp.json for this reason;
# listing it here would copy the primary's back over that.
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
# Opt-in by design — omit and the worker gets no PR-create grant (push over
# SSH is unaffected). The token value itself is never stored in this file.
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
# env → extra environment for this profile's workers, as a literal key/value map
# (CB-511). Use it to give workers a toolchain.
#
# A worker's environment does NOT come from your shell. bridged hands herdr an
# explicit env map and herdr merges it into ITS OWN process env — so before
# CB-511 a worker inherited whatever PATH the herdr server happened to be
# started with, which on a long-lived herdr can predate your toolchain entirely
# and leave workers unable to run `mvn` or `java` at all.
# bridged now propagates ITS OWN PATH to every worker by default; set `env:`
# only to override that or add more (JAVA_HOME, …). Since the default is the
# daemon's PATH, make sure the daemon is started with a good one — see the PATH
# lines in deploy/dev.ltms.bridged.plist and deploy/bridged.service.
#
# Adapter-owned variables always win over `env:`: ANTHROPIC_BASE_URL and the
# rest of the ANTHROPIC_*/CLAUDE_* wiring are applied after it, so an `env:`
# entry cannot repoint a worker past the SubscriptionGuard — which is checked
# against `baseUrl` alone.
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
workers:
gx10: # ccs profile name (NOT a hostname)
kind: claude-code # which adapter spawns this profile (default; may omit)
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
model: coder
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
weight: 0.5 # relative selection weight for placement: weighted
maxLoad: 2 # max live workers on this profile (omit for unlimited)
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
gx11: # a second backend, so `placement: weighted` has a choice
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}"
mcpUrl: http://127.0.0.1:8765/mcp
argv: ["ccs", "gx11"]
weight: 0.5
maxLoad: 2
# Pin an auto-compact window BELOW the served model's context ceiling. The global
# ~/.claude/settings.json value is shared by every ccs instance and the primary, so the
# per-profile override belongs here. Equal to the ceiling means auto-compact never fires
# before the server rejects the prompt, which kills a worker mid-turn (CB-523).
env:
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "280000"
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
# The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a
# `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude.
#
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → bridge_send →
# structured bridge_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
# and need NO credentials — check `opencode models` for the current free list, since the names
# change. That also makes the worker off-subscription by construction.
# opencode-free:
# kind: opencode
# model: opencode/north-mini-code-free # `provider/model` selector, injected as `-m`
# placement: tab
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
#
# CB-508: point an opencode profile at your OWN OpenAI-compatible endpoint (local vLLM, llama.cpp,
# LM Studio, TGI…) instead of opencode's gateway. Setting `baseUrl` on a `kind: opencode` profile
# makes the bridge emit a custom `provider` block into the generated opencode.json — opencode has
# no ANTHROPIC_BASE_URL seam, so this is how the endpoint is pinned.
# baseUrl → a bare host:port gets `/v1` appended (where these servers mount the API); a URL that
# already has a path is used verbatim, so a custom mount point still works.
# model → MUST be "<provider>/<model>". The provider half names the generated block; the model
# half must match an id the server reports at /v1/models. One field drives both the
# declaration and the `-m` flag, so they cannot drift apart. A bare model name with a
# baseUrl set is rejected at spawn rather than silently using the default gateway.
# tokenEnv → optional; its value becomes the provider apiKey. Most local servers ignore the key,
# so a placeholder is used when unset (the AI SDK still requires a non-empty one).
# NOTE: no `guard` entry is needed even with a baseUrl set. The SubscriptionGuard exists to stop a
# worker borrowing the primary's Anthropic subscription, and an opencode process has no Anthropic
# credential path at all.
# opencode-local:
# kind: opencode
# baseUrl: http://127.0.0.1:8000
# model: local-vllm/deepseek-v4-flash
# placement: tab
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
placement: weighted
defaultWorker: gx10
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
# must carry none. Every profile above must have its host listed here.
guard:
offSubscriptionHosts:
- gx00.gw
- gx01.gw
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
# unknown keys are ignored, so a snake_case key would be silently dropped (default kept).
# spawnReadyTimeoutMs: 20000
# spawnReadyPollMs: 300
# Worktree provisioning root (CB-301-ext). Where per-worker git worktrees are checked out so
# each worker owns an isolated branch instead of sharing the primary's tree. Omit to default
# to a sibling directory of the repo root.
# worktreeRoot: /Users/me/src/.bridged-worktrees
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
# contextCap → force-release a session after this many delegated turns
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
# lifecycle:
# idleTtlSeconds: 300
# contextCap: 10
# drainTimeoutSeconds: 5
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
# Set a broker uri to swap in the AMQP-backed inbox: worker replies with no open send are held
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
# broker:
# uri: amqp://guest:guest@127.0.0.1:5672
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
# pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop
# stops as soon as the primary's inbox is empty.
# terminal → pin the primary's herdr terminal id. Omit to learn it from the connection on
# the first orchestration-side MCP call (the normal case). An off-host or
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
# degrades to pull; the reply is still never lost.
#
# REQUIRED (CB-522) if the primary itself runs inside a herdr pane. Caller
# identity resolves a loopback PID to its herdr pane, and PaneLocator scans
# EVERY pane — not just bridged-spawned ones — so such a primary is otherwise
# classified as a WORKER and refused SPAWN/SEND/STOP. That failure is
# self-locking: the learned terminal is populated by the very orchestration
# calls being refused, so only this pinned value can break the cycle. Read the
# id off bridge_whoami (it reports the current terminal even while
# misclassified) and re-pin whenever the primary moves panes.
# pushReminders → max nudges before giving up (default 5)
# pushBackoffMs → delay between nudges in ms (default 15000)
# primary:
# terminal: term_65619bd6174568
# pushReminders: 5
# pushBackoffMs: 15000
+14 -14
View File
@@ -5,10 +5,10 @@
## Why this exists
Stage 2 gave a worker→primary reply a **durable place to wait** when no `fleet_send` is
Stage 2 gave a worker→primary reply a **durable place to wait** when no `bridge_send` is
open: it lands in `agent.<target>.inbox` on the broker and survives a daemon bounce. But
delivery is still **pull** — the primary only sees the reply if it happens to call
`fleet_poll(target)` / `GET /sessions/{id}/replies`. A reply can sit indefinitely while
`bridge_poll(target)` / `GET /sessions/{id}/replies`. A reply can sit indefinitely while
the primary works on something else.
This layer makes delivery **active**: the bridge *pushes* a nudge to the primary the moment
@@ -26,11 +26,11 @@ pointed at the primary's pane instead.
```mermaid
flowchart LR
W["worker"] -->|"fleet_reply (no open send)"| MS["MessageService.reply"]
W["worker"] -->|"bridge_reply (no open send)"| MS["MessageService.reply"]
MS -->|"inbox.publish"| INBOX[("agent.&lt;target&gt;.inbox<br/>(durable, LavinMQ)")]
MS -->|"notify"| LOOP["ReplyPushLoop"]
LOOP -->|"status-gated inject"| PANE["primary's herdr pane"]
PANE -->|"primary drains"| DRAIN["fleet_poll(target)<br/>= peek + ack"]
PANE -->|"primary drains"| DRAIN["bridge_poll(target)<br/>= peek + ack"]
DRAIN -->|"inbox now empty"| LOOP
LOOP -.->|"still non-empty →<br/>re-inject on backoff"| PANE
classDef store fill:#2c5282,stroke:#1a365d,color:#ffffff;
@@ -51,13 +51,13 @@ the caller runs in a herdr pane on this host. Today it's discarded for the prima
(`presence.markPresent` is a no-op on it).
**Plan:** a single-slot `PrimaryRegistry` (thread-safe) holding the primary's `terminal_id`.
Populate it from the **orchestration-side** MCP tools — `fleet_send`, `fleet_spawn`,
`fleet_poll`, `fleet_list`, `fleet_status`, `fleet_profiles` — capturing
Populate it from the **orchestration-side** MCP tools — `bridge_send`, `bridge_spawn`,
`bridge_poll`, `bridge_list`, `bridge_status`, `bridge_profiles` — capturing
`callerTerminal(exchange)` when it is (a) non-null and (b) **not** a registered worker
session in `SessionManager`. That caller is, by construction, the primary. Worker-side tools
(`fleet_reply`, `fleet_ask`) never set it.
(`bridge_reply`, `bridge_ask`) never set it.
- **Config override / pin:** a `primary: { terminal: "<id>" }` block in `FleetConfig`
- **Config override / pin:** a `primary: { terminal: "<id>" }` block in `BridgedConfig`
(nested record, same shape as `Broker`). Lets an operator pin it, or supply it when
derivation can't (see degrade case).
- **Degrade:** if the primary is off-host or in a non-herdr terminal, `terminalForPid`
@@ -71,13 +71,13 @@ A `ReplyPushLoop` component, notified at the single no-waiter call site
(`MessageService.reply` → the `inbox.publish` branch, `MessageService.java:192`).
- **Inject a nudge, not the payload.** The injected turn tells the primary *to drain*
(e.g. "Worker `<target>` returned a reply — run `fleet_poll(target=<target>)` to collect
(e.g. "Worker `<target>` returned a reply — run `bridge_poll(target=<target>)` to collect
it"), it does **not** carry the reply text. Rationale: replies can be large/multiline and
terminal injection would mangle them; the drain response is the clean transport. Keeps the
push idempotent — re-nudging is harmless.
- **Ack = drain.** The primary draining (`drainReplies` = peek + ack) is the acknowledgement.
The loop's **stop condition is `inbox.peek(target).isEmpty()`** — the reply is gone from the
inbox because it was acked. No new `fleet_ack` tool needed for v1 (see Increment 3).
inbox because it was acked. No new `bridge_ack` tool needed for v1 (see Increment 3).
- **Status-gated injection (mechanism (b), chosen).** A dedicated lightweight scheduled loop,
**not** the worker `Injector`. It injects via `AgentControl.send(primaryTerminal, nudge)`
(the same herdr `agent.send` = `pane send-text` + submit that delivers to workers) only when
@@ -93,10 +93,10 @@ A `ReplyPushLoop` component, notified at the single no-waiter call site
the reply remains in the durable inbox and the next natural poll (or a later worker reply's
nudge) still surfaces it. Bounded so the bridge never spams the primary.
### Increment 3 — optional per-`msgId` `fleet_ack` tool (deferred)
### Increment 3 — optional per-`msgId` `bridge_ack` tool (deferred)
Drain-as-ack is coarse: it clears *all* pending replies for a target at once. If finer
control is ever needed (ack one reply, leave others held), add a `fleet_ack(msgId)` tool
control is ever needed (ack one reply, leave others held), add a `bridge_ack(msgId)` tool
mapping to `inbox.ack(target, msgId)` — the port already supports per-`msgId` ack. Not built
in v1; the stop-on-empty loop is sufficient.
@@ -106,7 +106,7 @@ in v1; the stop-on-empty loop is sufficient.
resolved terminal is non-null **and not a registered worker session**, seen on an
orchestration-side tool. This never mislabels a worker (workers are in `SessionManager`)
and needs no new env var or argument (identity stays connection-derived, per the existing
`FleetMcp` invariant).
`BridgeMcp` invariant).
2. **Readiness-gate mismatch → dedicated loop.** The existing `Injector` gates delivery on
`ready.test(target)` = `WorkerPresence` (the *worker's* MCP connected). The primary is not
@@ -132,7 +132,7 @@ boundary**. The bridge is signalling the primary that it has mail — not drivin
non-null terminal AND not a registered session" predicate; the loop's stop-on-empty and
bounded-reminder logic with an injected clock + a fake injector (no real herdr).
- **Live dogfood (primary-side):** with the daemon on the broker jar + a real worker,
delegate a task, let the worker reply after the `fleet_send` window closes, and observe the
delegate a task, let the worker reply after the `bridge_send` window closes, and observe the
bridge inject a drain nudge into *this* primary pane; confirm draining stops the reminders;
confirm an unreachable primary (registry empty) degrades to pull with no loss.
-696
View File
@@ -1,696 +0,0 @@
# bridged configuration (example). Copy to bridged.yaml and adjust.
#
# bridged is the sole gateway between primary/worker Claude sessions and herdr.
# It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL.
# REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token`
# below — bridged REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
bind:
host: 127.0.0.1
port: 8765
# API authentication (CB-501). Governs how a caller that is NOT an on-host worker pane proves it
# is the primary. Worker identity never depends on this: a loopback peer PID that maps to a herdr
# pane is unforgeable and is always honoured, so turning auth on cannot lock the fleet out.
#
# mode: loopback-trust → DEFAULT, and the historical behaviour: any loopback caller that is not
# a worker is the primary, no credential needed. Sound ONLY because the
# OS refuses remote connections to a loopback socket.
# mode: token → such a caller must send `Authorization: Bearer <token>`; without it it
# is anonymous and authorized for nothing. REQUIRED for a non-loopback
# bind — the daemon fails fast otherwise, because "unauthenticated ⇒
# primary" on a reachable port would hand spawn/stop/send to anyone.
# tokenEnv → host env var holding the token (never the literal value). Default
# BRIDGED_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
#
# TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and
# let it own certificate lifecycle, e.g.
# location / { proxy_pass http://127.0.0.1:8765; proxy_set_header Authorization $http_authorization; }
# The broker link gets TLS from its own URI (amqps://…) — see `broker` below.
# auth:
# mode: token
# tokenEnv: BRIDGED_API_TOKEN
# Optional pinned primary terminal (CB-307). Names the herdr pane the PRIMARY itself runs in:
# a caller whose connection maps to this pane resolves as the primary (no credential needed —
# the pane mapping is as unforgeable as a worker's), and reply nudges are pushed to it.
# REQUIRED when the primary runs inside a herdr pane — without it the pane match reads the
# primary as a worker and refuses spawn/send/stop. Get the id from fleet_whoami; re-pin if
# the primary moves panes.
# primary:
# terminal: term_0123456789abcd
# pushReminders: 5 # max nudges before giving up (default 5)
# pushBackoffMs: 15000 # delay between nudges (default 15000)
# CB-530: MORE THAN ONE LEAD. `primary:` above is singular by construction — every other pane
# resolves as a worker — which is right for one lead driving a fleet and wrong the moment two leads
# (say a Claude lead and an opencode lead) work as peers: the second is silently demoted and refused
# every orchestration call. List each lead's pane here and all of them resolve as leads.
#
# tab → the ONLY field identity depends on (CB-579); the exact label of the tab hosting the lead.
# Label the tab yourself, or let bridged label one it launches — see `fleet.leaders:` below.
# kind/model → descriptive; they document what runs in the pane and are echoed by fleet_whoami
#
# A lead's tab must already carry its label (or be launched by bridged, which labels it) — there is
# no terminal id to paste in and nothing to re-pin when the session restarts: the tab survives, so
# the same label resolves the same lead again on the next scan.
# `fleet_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
# IS a primary for authorization, so nothing that keys on the role breaks.
#
# KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single
# destination for its nudges, and is a separate mechanism from lead identity — see `fleet.leaders:`.
#
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
#
# Two things stop the tab-name convention from becoming a way to claim leadership: the configured
# member spaces are excluded from the scan, so nothing bridged places can land in a matching tab;
# and startup REFUSES a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel`
# override, also matches — so the two namespaces cannot overlap by accident. The label is a NAME,
# never a capability: what a pane may do is decided by the role the daemon resolves for it.
# CB-551: IDLE-LEAD HEARTBEAT — nudge the single lead back to work when it has been continuously
# idle (no open fleet_send driving it) past the quiet period. The fleet is one lead + architects +
# workers, so a lead that stalls is a single point of failure; the ReplyPushLoop only nudges when a
# reply lands, and this timer catches the gap where nothing lands and the lead just sits idle.
#
# Opt-in on purpose — it SPENDS the operator's subscription on its own initiative (each nudge starts
# a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent
# block = feature off, exactly as before.
#
# Three knobs, each with a default that errs on the side of not burning context:
# idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 —
# # absorbs normal post-turn pauses; re-prompting every pause burns context)
# backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000)
# quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops
# # until real state appears (default 3 — never nag an empty fleet forever)
# leadHeartbeat:
# idleAfterSeconds: 300
# backoffMs: 60000
# quietNudgeCap: 3
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
# per tick.
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
# Math.max(15, intervalSeconds), so a lower value is silently raised, not
# rejected.
# workingSuspectAfterSeconds, paneProbeIntervalSeconds → accepted and parsed, but NOT YET READ by
# anything — the dormant monitor only consumes intervalSeconds today (CB-573
# shipped ahead of the evidence publishers these two knobs are for). Setting
# them changes nothing right now, and no minimum is enforced on either, because
# nothing reads them to enforce one. They exist so a later build can start
# honouring them without another config-shape change.
# notifications.mode → "webhook" flips what fleet_list REPORTS (healthCoverage: "full" instead
# of "detection-only") — it does NOT make bridged send any webhook call; no
# delivery mechanism is implemented yet. Any other value, or omitting the
# block, reports "detection-only".
# health:
# enabled: true
# intervalSeconds: 30
# workingSuspectAfterSeconds: 600
# paneProbeIntervalSeconds: 60
# notifications:
# mode: disabled
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
# How member sessions are spawned. Define one or more named profiles (backends) under
# `profiles`; each key is the profile name (also the ccs profile). A profile says only WHICH
# BACKEND — model, CLI adapter, credentials, cost. It says nothing about what a member spawned on
# it is for; that is the member's role, and roles live under `fleet:` below. Which profile an
# unqualified spawn lands on comes from that role's pool, not from a global default.
#
# Shared knobs (placement/workspace) can be repeated per profile; they usually match.
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
# Use `pane` for the legacy behaviour (split the focused tab).
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
# (--append-system-prompt) as launch flags; nothing is written to the profile.
# ideMcpUrl → opt-in (CB-634), default off. When set, bridged mounts the IDE Index MCP as a
# second inline server named `intellij`, and adds an IDE charter that pins every
# ide_* call to the member's own worktree. A URL, not a boolean — host and port
# are host-specific. Set it only on a host where the IDE actually runs.
# ideProjectDir → repo-relative module dir the IDE opens and the overlay pins (CB-634). Only read
# when ideMcpUrl is set. This repo's Maven pom lives in `bridged/`, not at the
# worktree root, so opening the root imports no module and ide_* resolves nothing;
# set this to `bridged`. Omit for a repo whose project is the worktree root.
# ideOpenCommand → host command that opens ideProjectDir in the IDE at spawn (CB-634 auto-open).
# Only read when ideMcpUrl is set. `{dir}` is replaced with the absolute module
# dir and the command runs through `/bin/sh -c`, so set env inline if needed —
# e.g. `env DISPLAY=:10.0 idea {dir}`. Best-effort: a failure is logged, never
# fails the spawn. Omit to open the member's module by hand. There is no close
# half yet — an opened module stays open until the operator closes it.
# tokenEnv → host env var holding the worker's auth token (value never stored in config);
# omit for a backend that needs no token (e.g. a local ollama).
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
# cwd on an MCP spawn, else the daemon's cwd — never $HOME. See
# docs/Worker-Startup-and-Trust.md.
# configDir → CLAUDE_CONFIG_DIR for the worker, so it inherits that profile's
# skills/MCP/hooks. Omit to leave the worker on the host default.
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
# worktree sees the same local config (CB-301-ext). Omit for the default set:
# [.env, .envrc]. (.claude/settings.local.json is NOT in the default — it
# pre-approves IDE/tool grants a member must not hold ambiently; CB-525/CB-634.)
#
# Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher
# mounts — the bridge, and nothing else. Replicating the primary's MCP config
# handed a worker the primary's IDE servers, which are bound to the primary's
# checkout, so its navigation returned paths OUTSIDE its own worktree: one
# worker made all 59 of its edits in the primary tree while compiling its
# worktree, and every build it ran was of code that did not contain them.
# bridged neutralizes a provisioned worktree's .mcp.json for this reason;
# listing it here would copy the primary's back over that.
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
# Opt-in by design — omit and the worker gets no PR-create grant (push over
# SSH is unaffected). The token value itself is never stored in this file.
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
# exhaustedPattern → regex matched against a completion-fallback scrape (CB-578 stage A) to
# classify a turn that ended with no fleet_reply as the backend having
# refused on a subscription usage limit, rather than a real answer. Opt-in —
# omit and this profile's completion fallback behaves exactly as before.
# Every backend words its refusal differently, so this is config, never a
# vendor string baked into bridged itself.
# DEFERRED: compiled once into a startup pattern map — editing it needs a
# daemon restart, same as this profile's model/baseUrl/argv.
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
# BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME
# credentialId share one quarantine — the case this exists for is two models
# on one account (e.g. sol and terra both billing one OpenAI credential): an
# exhaustion on either one must lock out both, or the fleet just walks onto
# the same dead account under the sibling's name. Opt-in — omit and this
# profile quarantines alone, under its own name, exactly as if the field did
# not exist. Cooldown length is the top-level quarantineCooldownSeconds below.
# HOT: read live at every spawn/exhaustion check — no restart needed.
# env → extra environment for this profile's workers, as a literal key/value map
# (CB-511). Use it to give workers a toolchain.
#
# A worker's environment does NOT come from your shell. bridged hands herdr an
# explicit env map and herdr merges it into ITS OWN process env — so before
# CB-511 a worker inherited whatever PATH the herdr server happened to be
# started with, which on a long-lived herdr can predate your toolchain entirely
# and leave workers unable to run `mvn` or `java` at all.
# bridged now propagates ITS OWN PATH to every worker by default; set `env:`
# only to override that or add more (JAVA_HOME, …). Since the default is the
# daemon's PATH, make sure the daemon is started with a good one — see the PATH
# lines in deploy/dev.ltms.fleet.plist and deploy/bridged.service.
#
# Adapter-owned variables always win over `env:`: ANTHROPIC_BASE_URL and the
# rest of the ANTHROPIC_*/CLAUDE_* wiring are applied after it, so an `env:`
# entry cannot repoint a worker past the SubscriptionGuard — which is checked
# against `baseUrl` alone.
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
profiles:
gx10: # ccs profile name (NOT a hostname)
kind: claude-code # which adapter spawns this profile (default; may omit)
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
model: coder
placement: tab
workspace: bridged-workers
# tabLabel: an optional per-profile override; the fleet template usually covers it
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
# weight: relative selection weight for automatic placement (weighted, round-robin, and
# fixed's fallback walk). Absent defaults to 1.0. An explicit 0 or negative value means
# "never auto-select this profile" (CB-554) — it stays reachable via an explicit
# `fleet_spawn{profile:"gx10"}`, which bypasses placement entirely; only automatic
# selection skips it.
weight: 0.5
# maxLoad: max live workers on this profile. Omit for unlimited. An explicit 0 (CB-585) caps
# the profile at zero live members — it is excluded from automatic placement and an explicit
# `fleet_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
# is named directly. Negative is refused at config load — there is no sane meaning for it.
maxLoad: 2
# subscription: true
# THE KNOB THAT DECIDES WHO PAYS (CB-539). Default false. When true, this profile's members
# run on the OPERATOR'S OWN Claude subscription instead of a metered endpoint — every spawn
# bills your plan and eats your usage limit. Off-subscription is the whole point of this
# daemon, so treat `true` as a deliberate exception, not a convenience.
#
# What changes when it is set (ClaudeCodeLauncher):
# - no ANTHROPIC_BASE_URL and no ANTHROPIC_AUTH_TOKEN are injected — the member inherits
# the operator's own Claude Code auth, which is exactly why it bills the plan;
# - SubscriptionGuard never vets it, because there is no baseUrl to vet;
# - no token is required, so `tokenEnv` is irrelevant here.
#
# MUTUALLY EXCLUSIVE with `baseUrl` — setting both is refused at config load (CB-542). On the
# subscription path no guard would vet the URL, so allowing both would be a way around the
# guard rather than a configuration.
#
# GOTCHA 1 — it is invisible to the startup secret check. `Fleetd.reportRequiredSecrets`
# skips subscription profiles on purpose (they need no token), so a boot log that reports
# every secret as fine says nothing about these profiles.
#
# GOTCHA 2 — `maxLoad` is the ONLY throttle you have here. There is no metering, no budget
# and no refusal on cost; the cap on live members is the single thing standing between a
# fan-out and your monthly limit. Set it deliberately and keep it small.
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
# credentialId: shared-openai # opt-in: quarantine together with every other profile sharing this id (CB-578)
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".env", ".envrc"] # the default; never add .mcp.json or .claude/settings.local.json — see above
# ideMcpUrl: http://127.0.0.1:29170/index-mcp/streamable-http # opt-in (CB-634): IDE code intelligence, pinned to the worktree
# ideProjectDir: bridged # CB-634: module dir the IDE opens + the overlay pins (this repo's pom is in bridged/)
# ideOpenCommand: env DISPLAY=:10.0 idea {dir} # CB-634 auto-open: opens {dir} in the IDE at spawn; omit to open by hand
gx11: # a second backend, so `placement: weighted` has a choice
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
placement: tab
workspace: bridged-workers
# tabLabel: an optional per-profile override; the fleet template usually covers it
mcpUrl: http://127.0.0.1:8765/mcp
argv: ["ccs", "gx11"]
weight: 0.5
maxLoad: 2
# Pin an auto-compact window BELOW the served model's context ceiling. The global
# ~/.claude/settings.json value is shared by every ccs instance and the primary, so the
# per-profile override belongs here. Equal to the ceiling means auto-compact never fires
# before the server rejects the prompt, which kills a worker mid-turn (CB-523).
env:
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "280000"
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
# The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a
# `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude.
#
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → fleet_send →
# structured fleet_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
# and need NO credentials — check `opencode models` for the current free list, since the names
# change. That also makes the worker off-subscription by construction.
# opencode-free:
# kind: opencode
# model: opencode/north-mini-code-free # `provider/model` selector, injected as `-m`
# placement: tab
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
#
# CB-508: point an opencode profile at your OWN OpenAI-compatible endpoint (local vLLM, llama.cpp,
# LM Studio, TGI…) instead of opencode's gateway. Setting `baseUrl` on a `kind: opencode` profile
# makes the bridge emit a custom `provider` block into the generated opencode.json — opencode has
# no ANTHROPIC_BASE_URL seam, so this is how the endpoint is pinned.
# baseUrl → a bare host:port gets `/v1` appended (where these servers mount the API); a URL that
# already has a path is used verbatim, so a custom mount point still works.
# model → MUST be "<provider>/<model>". The provider half names the generated block; the model
# half must match an id the server reports at /v1/models. One field drives both the
# declaration and the `-m` flag, so they cannot drift apart. A bare model name with a
# baseUrl set is rejected at spawn rather than silently using the default gateway.
# tokenEnv → optional; its value becomes the provider apiKey. Most local servers ignore the key,
# so a placeholder is used when unset (the AI SDK still requires a non-empty one).
# NOTE: no `guard` entry is needed even with a baseUrl set. The SubscriptionGuard exists to stop a
# worker borrowing the primary's Anthropic subscription, and an opencode process has no Anthropic
# credential path at all.
# opencode-local:
# kind: opencode
# baseUrl: http://127.0.0.1:8000
# model: local-vllm/deepseek-v4-flash
# placement: tab
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
#
# `weighted` IS NOT "cheapest first" — read this before you set weights (CB-589).
# It is smooth weighted round-robin: it spreads spawns across EVERY profile that has a free slot,
# in weight ratio. It has no idea which profile costs money. So with local:10 / paid:2 you do not
# get "use local, overflow to paid" — you get roughly one spawn in six going to the paid profile
# while the local box still has a free slot.
#
# There is a sharper second effect. The policy's running score map lives for the daemon's whole
# life. While a profile is at maxLoad it is filtered out and its score FREEZES, so the paid
# profiles keep accumulating against it. When the local slot frees up it returns with a stale
# score and can LOSE the next pick — a paid spawn while the free box sits idle.
#
# Until a real cost-first policy exists, the workaround is to make the ratio decisive rather than
# proportional: give the free profile a weight so large that it wins every pick it is eligible
# for, and paid profiles only ever take genuine overflow. On this host that is local weight 100
# against paid weights of ~1.
#
# The gotcha with that workaround: it expresses a PREFERENCE ORDER through a RATIO knob. Add a
# future profile at weight 150 and it silently outranks the free box, with nothing to warn you.
# Re-check the weights whenever you add a profile.
placement: weighted
# How long a credential sits out after a BACKEND_EXHAUSTED classification (CB-578 stage B), in
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
# 1800 (30 minutes) when omitted or non-positive.
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
# restart. Editing this needs a daemon restart to take effect.
# quarantineCooldownSeconds: 1800
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
# upgraded bridged keeps the old behaviour: the file is read once at boot and never again.
# enabled → turn the watch on. bridged checks the file's modified time on a timer and
# reloads when it moves.
# intervalSeconds → how often to check (default 10). One `stat` per tick, so this is cheap.
#
# Not every key can move under a running daemon, and the difference is about what already exists
# when the reload happens — not about how important the key is:
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad
# / credentialId. Those are hot because the placement policy (and, for credentialId,
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
# not by itself enough to make a key hot.
# EXCEPT `fleet.leaders`: Fleetd.main reads it once at startup to build the lead tab
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
# with nothing in the deferred list — but has NO effect until you restart. Treat it
# as deferred in practice, even though today's reload output does not say so.
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
# until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`,
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, `quarantineCooldownSeconds` (CB-578
# stage B — baked once into the quarantine tracker built at startup), ADDING or
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
# configDir, mcpUrl, tabLabel, exhaustedPattern. The launcher takes a copy of
# `profiles:` at startup and resolves every spawn out of that copy, so those never
# reach a launch until you restart. The reload logs them by name rather than
# pretending they applied.
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
# bound, the broker connection is open, and the auth mode decides who may reach the
# port that is already listening.
#
# A changed COLD key refuses the WHOLE reload — not the hot half applied and the cold half warned
# about. A half-applied reload would leave the daemon matching no file on disk, which is the worst
# thing a reload can do to an operator debugging one. A file that fails to parse or fails a startup
# validator is refused the same way, and the running config stays live.
# configReload:
# enabled: true
# intervalSeconds: 10
# THE FLEET (CB-557) — who the daemon may run, and under which role. This one block replaced four
# older keys: `leaders:`, `members:`, `leadScan:` and `defaultProfile:`.
#
# A member is anything a lead spawns, and every member has two INDEPENDENT attributes:
# role — which contract: architect, dev or reviewer. It picks the launch charter, the role
# file, the playbook skill and the authz row.
# profile — which backend: one of the `profiles:` keys above (model, CLI adapter, cost).
# They vary on their own. A reviewer may run on the same profile as the dev whose diff it reads,
# which is why the two cannot be one field.
#
# The ROLE IS THE CONTAINING KEY, not a `role:` field. That is not only tidier: a misspelled role
# used to parse into a member with no contract at all, while a misspelled pool name here simply
# declares nothing.
#
# Each pool lists the profiles that role MAY run on — these are pools, not identities. That is also
# what replaced `defaultProfile:`: an unqualified spawn names a role, and that role's pool supplies
# the candidates, in definition order. A dev and a reviewer staying anonymous is exactly compatible
# with being listed here; the entry key just names the entry.
fleet:
# Optional launch-charter text, keyed only by the singular role wire names: architect, dev,
# reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put secrets
# here: a later launch step writes this text to a world-readable temp file, and ${ENV} interpolation
# is deliberately not supported.
charters:
architect: |-
You are an architect in this fleet. You refine work before anyone builds it:
scope, acceptance criteria, risks, and a unit split. You read the repo and
write analysis. You never commit production code and never open a PR.
A design task is worked by two architects. Design alone first, then exchange
and say plainly where you disagree. Do not concede just to agree.
dev: |-
You implement the one unit you were given, and nothing else. You test it,
commit it, and open your own pull request. You never merge.
reviewer: |-
You review the diff you were given. You report bugs, risks and missing tests.
You do not change code.
# Optional. Template for a member tab's label; {role}, {profile}, {model} and {n} are substituted.
# {n} counts per role+profile, so `dev: sonnet #2` really is the second sonnet dev. Because {role}
# comes from a closed enum, a generated label can never begin with a lead's tabPrefix.
# tabLabel: "{role}: {profile} #{n}"
# Panes that orchestrate rather than are orchestrated. A lead may now be CREATED as well as
# recognised: give it a `profile:` and the daemon launches the shortfall when fewer than
# `instances` are live. Omit `profile:` and it is recognise-only, as before.
#
# `tab:` (CB-579) is REQUIRED and is the only field identity depends on — the exact label of the
# tab hosting the lead, matched case-insensitively. Label the tab yourself and put that same
# string here, and the pane is recognised on the next rescan. Reopen the tab later, or the session
# inside it restarts — the terminal id changes; the tab, and its label, do not, so no config edit
# follows a restart.
#
# A lead the daemon launches is labelled BY the daemon with this same `tab:` value, so it is found
# by the same scan. A lead counts as live only when herdr also reports a running agent in that
# tab — a label left behind by a session that died does not block the relaunch, and a tab that is
# gone entirely drops out of the next scan rather than being remembered forever.
#
# An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with
# the session lifecycle (the idle reaper would kill your orchestrator), and stays on the
# subscription — ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.
#
# GET THE `tab:` VALUE RIGHT. A pane that does not match any configured `tab:` (a typo, a renamed
# tab, a pane no entry names at all) is not recognised as a lead — it resolves as an ordinary
# WORKER instead, silently, and every orchestration call it makes (spawn/stop/send/drain) is
# refused. There is no error at startup for this: an unmatched pane is simply not a lead. If your
# primary suddenly can't spawn or send, check this section first.
# leaders:
# opus-5.0:
# profile: opus # omit to never create this lead, only recognise it
# instances: 1 # desired live count; only the shortfall is launched. 0 = off
# tab: "lead: opus-5.0" # REQUIRED — the exact tab label this lead lives in
# tabPrefix: "lead:" # only used to guard against a worker tabLabel colliding with
# # this convention at startup; plays no part in matching a lead
# scanIntervalSeconds: 10 # rescan cadence, and the worst case before a new tab is seen
# workspace: leads # where a launched lead's tab is created (default "leads").
# # MUST NOT be a member workspace — those are excluded from the
# # scan, so a lead placed in one is never found again.
# cwd: /path/to/repo # the launched lead's working directory (default: bridged's own)
# kind: claude # descriptive; reported by fleet_whoami
# gpt-sol-5.6:
# tab: "lead: gpt-sol-5.6"
# kind: opencode
# model: openai/gpt-5.6-terra
# architects:
# architect-1:
# profile: opus # a strong model, on the operator's subscription
# architect-2:
# profile: sol # a different vendor on purpose — two architects that share a
# # model share its blind spots
developers:
gx10:
profile: gx10
# reviewers:
# gx10:
# profile: gx10 # the same backend may serve two roles; that is the point
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
# must carry none. Every profile above must have its host listed here.
guard:
offSubscriptionHosts:
- gx00.gw
- gx01.gw
# Member credential policy (CB-596, gitea issue #82). A herdr pane runs a LOGIN shell, and that
# shell re-sources the operator's own secret store — so a spawned member inherits every credential
# the operator's shell holds, not just the ones bridged means to give it. Measured on this host:
# 31 credential names, all set, with only ONE (GITEA_ACCESS_TOKEN) blocked before this — and that
# block was a single name hardcoded in HerdrPeerLauncher.java, not driven by this file. This block
# replaces that hardcoded shadow with a config-driven list of names.
#
# ROUND-2 CORRECTION, measured live: the pane-creation env overlay below (applied at tab.create /
# pane.split, BEFORE the pane's login shell runs) does NOT survive that login shell for any name
# secrets.sh actually exports — the shell re-exports it afterwards and overwrites the sentinel.
# Proof: GITEA_ACCESS_TOKEN comes back blocked only because secrets.sh itself carries a guarded
# export (`[ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...`) — that guard, not this
# file, is what wins. No other name in `known` below has a matching guard in secrets.sh yet (1
# guard measured against 33 export lines there). So today this block's overlay is REAL protection
# only for a name secrets.sh does not export, or a peer kind whose pane never runs a login shell —
# for everything secrets.sh exports and guards, the guard in secrets.sh (out of scope for this
# ticket) is what actually blocks it, not this list. An exec-time fix (winning after the login
# shell finishes, before the agent process starts) was attempted and found to have no seam in the
# current herdr protocol — AgentControl.start takes a fixed `kind` (herdr resolves the executable)
# plus trailing CLI args for that binary, not an arbitrary argv or an env map; only tab.create /
# pane.split accept `env`, and that is this same pane-creation overlay. See gitea #82 for the open
# design question this leaves.
#
# DENY-BY-DEFAULT, NOT A DENY-LIST. A deny-list (name the bad ones, let everything else through) is
# silently wrong the moment the operator's store gains a new secret — nothing would ever report it.
# Deny-by-default inverts that: `known` bounds the blast radius to names actually enumerated below,
# and EVERY one of them is blocked UNLESS it is also in `allow`. Omitting this block entirely (the
# shipped default) blocks NOTHING — unlike most optional blocks in this file, absence here is a real
# gap, not a safe "feature off". A name that is neither `known` nor `allow`-ed is not silently let
# through either: the daemon logs a WARN naming any credential-shaped env var it finds on neither
# list (never its value), so a secret added to the store later does not go unnoticed forever.
#
# policy → "deny-by-default" (the default; also accepted spelled "deny-list") overlays each
# known-but-not-allowed name BEFORE the pane's login shell runs — real protection only
# where that shell does not re-export the name (see ROUND-2 CORRECTION above). An
# unrecognized value refuses to start, naming it.
# policy → "allow-list" (CB-633) moves the control to a per-spawn ZDOTDIR directory the daemon
# generates and passes through tab.create's env map. Each generated startup file sources
# its ~/ counterpart FIRST and then runs the scrub, so the scrub happens after the
# operator's whole chain and no sourced file can undo it.
# The scrub is sourced from BOTH the generated .zshrc and the generated .zlogin, because
# herdr does not open the same kind of shell everywhere: macOS panes run a LOGIN zsh (so
# .zlogin runs), Linux panes run a plain interactive zsh (so .zlogin never runs at all).
# A scrub in .zlogin alone would be a control that silently does nothing on Linux.
# The allow-list is DERIVED, never typed:
# every profile's tokenEnv/gitTokenEnv/gitHostEnv values and env-map keys, plus an
# infrastructure set (PATH HOME SHELL TERM LANG LC_* TMPDIR USER LOGNAME PWD SHLVL EDITOR
# PAGER JAVA_HOME XDG_* ZDOTDIR), plus whatever keys this spawn's own env overlay carries.
# Adding a profile can therefore only widen the list, never break another spawn's scrub.
# Under this policy `known`/`allow` below become REPORTING ONLY — they feed the gap WARN,
# they are no longer a control. If the member's login shell is NOT zsh, the daemon logs a
# loud WARN saying protection is off and falls back to deny-by-default's overlay.
# Each pane writes a scrub-report.txt naming how many variables it kept of how many it
# saw; the daemon logs that "allowed N of M" line when the pane stops. If the report is
# MISSING the daemon logs a WARN instead — the scrub cannot then be confirmed to have
# run, and a silently-dead control is exactly what this policy exists to prevent.
# allow → credential names a member legitimately needs. Under deny-by-default, left OUT of the
# pane's env overlay entirely, so the value the pane's own (login) shell exports passes
# through untouched. Under allow-list: reporting only.
# known → every credential name the operator's store is known to export. Under deny-by-default,
# every name here NOT also in `allow` is overlaid with a non-secret sentinel value before
# the pane's login shell runs — real protection only for names that shell does not itself
# re-export (see the ROUND-2 CORRECTION note above). Under allow-list: reporting only.
# sshAuthSock → whether SSH_AUTH_SOCK may pass through under allow-list ("allow") or must be
# blanked like any other non-derived name ("block", the default). This is a decision you
# have to make explicitly: SSH_AUTH_SOCK is a handle to YOUR ssh-agent, and a member
# holding it can sign with your keys — it sits in no secret file and looks like no
# credential, which is why it slipped past three earlier tickets (gitea #110). Blocking
# it breaks git over SSH inside members (push/fetch authenticate as you); use HTTPS
# remotes or scoped deploy keys instead of allowing it lightly.
#
# HOT-RELOADABLE the same way `fleet:` is (CB-559): read fresh on every spawn, so editing this list
# and reloading config (or restarting) changes what the NEXT spawn inherits; already-running members
# are unaffected either way.
# memberCredentials:
# policy: deny-by-default # or "deny-list", or "allow-list" (CB-633) — see above
# sshAuthSock: block # allow-list only; see the sshAuthSock note above
# allow:
# - AI_GATEWAY_TOKEN # named in a profile's tokenEnv (local/gx) — a member reaching the
# # gateway is by design, not a leak
# - WORKER_GITEA_TOKEN # the repo-scoped forge token a member needs to open its own PR (CB-302)
# - CONTEXT7_TOKEN # already decided as allowed by CB-593
# - GITEA_HOST # not a credential — a hostname, paired with the forge token above
# known:
# - AI_GATEWAY_TOKEN
# - BESZEL_ADMIN_EMAIL
# - BESZEL_ADMIN_PASSWORD
# - BESZEL_HUB_URL
# - BESZEL_KEY
# - BESZEL_UNIVERSAL_TOKEN
# - BRAIN_MCP_TOKEN
# - CF_ACCOUNT_ID
# - CF_API_TOKEN
# - CF_USER_TOKEN
# - CONFLUENCE_API_TOKEN
# - CONFLUENCE_USERNAME
# - CONTEXT7_TOKEN
# - GITEA_ACCESS_TOKEN
# - GITEA_HOST
# - GITLAB_OAUTH_CLIENT_SECRET
# - GITLAB_PERSONAL_ACCESS_TOKEN
# - GRAFANA_ADMIN_PASSWORD
# - GRAFANA_ADMIN_USER
# - HASS_TOKEN
# - HW_PASSWORD
# - HW_USER
# - LTMS_API_KEY
# - MEMORY_MCP_TOKEN
# - METRICS_PUSH_TOKEN
# - OPENCODE_AUTOMODE_MODEL
# - TELEGRAM_BOT_TOKEN
# - TELEGRAM_CHAT_ID
# - TS_API_KEY
# - TS_AUTHKEY
# - WORKER_GITEA_TOKEN
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
# unknown keys are ignored, so a snake_case key would be silently dropped (default kept).
# spawnReadyTimeoutMs: 20000
# spawnReadyPollMs: 300
# Worktree provisioning root (CB-301-ext). Where per-worker git worktrees are checked out so
# each worker owns an isolated branch instead of sharing the primary's tree. Omit to default
# to a sibling directory of the repo root.
# worktreeRoot: /Users/me/src/.bridged-worktrees
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
# contextCap → force-release a session after this many delegated turns
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
# clearAfterTurn → whether a reusable worker discards its conversation context after every
# completed delegated turn (default false). Works for claude-code workers
# only — any other peer kind (e.g. opencode) logs "context reset is
# unsupported for peer kind …" once and the reset is a no-op.
# lifecycle:
# idleTtlSeconds: 300
# contextCap: 10
# drainTimeoutSeconds: 5
# clearAfterTurn: false
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
# Set a broker uri to swap in the AMQP-backed inbox: worker replies with no open send are held
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
# uriEnv → CB-151: name of a host env var holding the AMQP URI, preferred over `uri` (wins
# whenever set). The URI carries `user:pass@` inline, so naming a variable keeps the
# password out of fleetd.yaml — same pattern as auth.tokenEnv/Profile.tokenEnv. A
# uriEnv that resolves to an unset or blank variable is treated as NOT configured and
# the daemon falls back to the in-memory inbox, warning loudly.
# prefetch → CB-527: consumer basicQos, capping how many unacked messages the inbox holds
# in-heap per owned target (the rest sits on the broker's durable queue instead of
# growing the JVM heap). Default 32 when omitted.
# broker:
# uriEnv: LAVINMQ_URI
# prefetch: 32
# Shared cross-host LEADER coordination broker. OMIT this block to leave lead-to-lead messaging
# off entirely (config-only in this ticket — nothing here wires it into a live LeadMailbox yet).
# This is a SEPARATE AMQP vhost from `broker:` above: member/worker inboxes always stay on the
# per-fleet `broker:` vhost, and this vhost carries only leader-to-leader traffic, so two fleets
# whose members must never see each other can still share one coordination vhost for their leads.
# uriEnv → name of a host env var holding the coordination AMQP URI, same convention as
# broker.uriEnv (keeps the credential out of fleetd.yaml). Wins over `uri` when set.
# selfId → this daemon's own lead coord-id — the name its mailbox is owned under
# (lead.<selfId>.inbox), e.g. "mac-opus" or "fleet01-lead". Must be globally unique
# across every daemon sharing this vhost.
# prefetch → consumer basicQos, capping how many unacked messages the mailbox holds in-heap.
# Default 32 when omitted.
# coordinator:
# uriEnv: LEAD_COORD_URI
# selfId: mac-opus
# prefetch: 32
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open fleet_send,
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
# pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop
# stops as soon as the primary's inbox is empty.
# terminal → pin the primary's herdr terminal id. Omit to learn it from the connection on
# the first orchestration-side MCP call (the normal case). An off-host or
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
# degrades to pull; the reply is still never lost.
#
# REQUIRED (CB-522) if the primary itself runs inside a herdr pane. Caller
# identity resolves a loopback PID to its herdr pane, and PaneLocator scans
# EVERY pane — not just bridged-spawned ones — so such a primary is otherwise
# classified as a WORKER and refused SPAWN/SEND/STOP. That failure is
# self-locking: the learned terminal is populated by the very orchestration
# calls being refused, so only this pinned value can break the cycle. Read the
# id off fleet_whoami (it reports the current terminal even while
# misclassified) and re-pin whenever the primary moves panes.
# pushReminders → max nudges before giving up (default 5)
# pushBackoffMs → delay between nudges in ms (default 15000)
# primary:
# terminal: term_65619bd6174568
# pushReminders: 5
# pushBackoffMs: 15000
+4 -7
View File
@@ -5,17 +5,17 @@
<modelVersion>4.0.0</modelVersion>
<groupId>dev.ltms</groupId>
<artifactId>fleetd</artifactId>
<artifactId>bridged</artifactId>
<version>1.0.0</version>
<packaging>jar</packaging>
<name>fleetd</name>
<description>fleet message server: sole gateway between a lead session, its members, and herdr</description>
<name>bridged</name>
<description>claude-bridge message server: sole gateway between primary/worker Claude sessions and herdr</description>
<properties>
<maven.compiler.release>25</maven.compiler.release>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<mainClass>dev.ltms.fleet.Fleetd</mainClass>
<mainClass>dev.ltms.bridged.Bridged</mainClass>
<jackson.version>2.19.0</jackson.version>
<javalin.version>6.7.0</javalin.version>
@@ -161,9 +161,6 @@
</dependencies>
<build>
<!-- CB-632: stays 'bridged' until the cutover renames the module dir and the launchd
plist together. The installed plist names bridged/target/bridged.jar and
KeepAlive is armed, so renaming the jar alone strands a restart. -->
<finalName>bridged</finalName>
<plugins>
<plugin>
@@ -0,0 +1,325 @@
package dev.ltms.bridged;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
import dev.ltms.bridged.msg.AmqpReplyInbox;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.msg.ReplyPushLoop;
import dev.ltms.bridged.rest.BridgedApp;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.session.SessionReaper;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.placement.PlacementPolicies;
import dev.ltms.bridged.worker.CompositePeerLauncher;
import dev.ltms.bridged.worker.HerdrPeerLauncher;
import dev.ltms.bridged.worker.OpenCodeLauncher;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
/**
* {@code bridged} entry point. Wires the real herdr socket client to the REST app and
* starts listening. Before anything else it asserts its own environment is clean —
* {@code bridged} is not a Claude process and must never carry a base_url.
*/
public final class Bridged {
private static final Logger log = LoggerFactory.getLogger(Bridged.class);
/** How often the injector samples a busy worker's status while it has queued work. */
private static final long INJECT_POLL_MILLIS = 250;
/** CB-504: how long to wait at startup for herdr's socket before serving degraded. */
private static final long HERDR_WAIT_SECONDS = 30;
private static final long HERDR_WAIT_POLL_MILLIS = 500;
static void main(String[] args) {
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
BridgedConfig cfg = BridgedConfig.load(configPath);
// The primary/host env that launched bridged must not be tainted.
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
guard.assertPrimaryClean(System.getenv());
// CB-501: refuse to start if the bind is wider than the auth mode can defend. Under
// loopback-trust, "not a known worker" means "the primary" — sound only because the OS
// refuses remote connections to a loopback socket. This throws rather than warns so the
// dangerous configuration cannot be reached by ignoring a log line.
cfg.validateAuthExposure();
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
: UnixSocketHerdrClient.defaultSocketPath();
UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect(socket, new com.fasterxml.jackson.databind.ObjectMapper());
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
Map<String, BridgedConfig.Worker> claudeProfiles = new LinkedHashMap<>();
Map<String, BridgedConfig.Worker> opencodeProfiles = new LinkedHashMap<>();
cfg.workerProfiles().forEach((name, w) -> {
if (w.isOpenCode()) {
opencodeProfiles.put(name, w);
} else {
claudeProfiles.put(name, w);
}
});
List<HerdrPeerLauncher> adapters = new ArrayList<>();
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
claudeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(agents, spaces,
opencodeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(name -> 0);
PeerLauncher workers = new CompositePeerLauncher(
adapters,
cfg.defaultProfile(),
cfg.workerProfiles(),
PlacementPolicies.fromName(cfg.placement()),
profileName -> liveCountRef.get().apply(profileName));
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
// with /healthz reporting "degraded" is strictly more useful than exiting.
if (awaitHerdr(herdr)) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
// with the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
} else {
log.warn("herdr did not answer within {}s — starting anyway; /healthz will report "
+ "degraded until it comes up. Orphaned worker panes (if any) were NOT reaped.",
HERDR_WAIT_SECONDS);
}
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
int contextCap = 0;
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
&& cfg.lifecycle().contextCap() > 0) {
contextCap = cfg.lifecycle().contextCap();
}
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()), contextCap);
liveCountRef.set(profileName -> (int) sessions.roster().stream()
.filter(s -> profileName.equals(s.profile()))
.count());
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
final SessionReaper reaper;
if (cfg.lifecycle() != null
&& cfg.lifecycle().idleTtlSeconds() != null
&& cfg.lifecycle().idleTtlSeconds() > 0) {
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
reaper.start();
} else {
reaper = null;
}
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
Rendezvous rendezvous = new Rendezvous();
CompletionResolver completion = new CompletionResolver(agents, rendezvous);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
WorkerPresence presence = sessions.asPresence();
TurnListener turnListener = new TurnListener() {
@Override
public void onTurnComplete(String target) {
completion.onTurnComplete(target);
sessions.onTurnComplete(target);
}
@Override
public void onDelivered(String target) {
completion.onDelivered(target);
sessions.onDelivered(target);
}
@Override
public void onTurnFailed(String target) {
completion.onTurnFailed(target);
sessions.onTurnFailed(target);
}
};
Injector injector = new Injector(agents, turnListener, presence::isPresent, presence::forget);
StatusPoller poller = new StatusPoller(agents, injector, INJECT_POLL_MILLIS);
poller.start();
// CB-307: reply inbox. A broker: block (with a uri) selects the AMQP-backed durable adapter;
// absent, bridged stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
// connection, so keep the reference to close it in the ordered shutdown hook.
final ReplyInbox replyInbox;
if (cfg.broker() != null && cfg.broker().isConfigured()) {
replyInbox = AmqpReplyInbox.open(cfg.broker().uri());
log.info("reply inbox: AMQP broker (durable) at {}", cfg.broker().uri());
} else {
replyInbox = new InMemoryReplyInbox();
log.info("reply inbox: in-memory (soft-state)");
}
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
// The pin also feeds CallerResolver below: a primary running inside a herdr pane would
// otherwise resolve as a worker and be refused every orchestration tool.
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
// open bridge_send. Uses its own lightweight scheduled executor, separate from the injector.
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
var pushScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-push-").unstarted(r));
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
// CB-512: the push loop takes it too, so nudge outcomes (delivered|exhausted) are counted.
Metrics metrics = BridgedMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, agents, replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
sessions.onAcquire(replyInbox::own);
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
sessions.onRelease(terminal -> {
messages.abandon(terminal, "the worker session was released before it replied");
replyInbox.release(terminal);
});
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
final CallerResolver callers;
if (cfg.auth().tokenMode()) {
String token = System.getenv(cfg.auth().tokenEnv());
if (token == null || token.isBlank()) {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting bridged");
}
callers = new CallerResolver(identity, true, token, pinnedPrimaryTerminal);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = new CallerResolver(identity, false, null, pinnedPrimaryTerminal);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
BridgeMcp mcp = new BridgeMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics);
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
// last. This replaces the earlier independent hooks that could race and close herdr early.
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
poller.stop();
messages.close();
pushLoop.close();
mcp.close();
if (reaper != null) reaper.stop();
// Release the broker connection last among message resources (no-op for the in-memory inbox).
if (replyInbox instanceof AutoCloseable closeable) {
try {
closeable.close();
} catch (Exception e) {
log.debug("reply inbox close: {}", e.toString());
}
}
herdr.close();
}));
Javalin app = new BridgedApp(herdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics).build();
app.start(cfg.bind().host(), cfg.bind().port());
log.info("bridged listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
*
* @return true if herdr answered, false if it never did
*/
private static boolean awaitHerdr(HerdrClient herdr) {
long deadline = System.nanoTime() + HERDR_WAIT_SECONDS * 1_000_000_000L;
boolean waited = false;
while (true) {
try {
herdr.call("ping");
if (waited) {
log.info("herdr is up");
}
return true;
} catch (HerdrException e) {
if (System.nanoTime() >= deadline) {
return false;
}
if (!waited) {
log.info("waiting up to {}s for the herdr socket…", HERDR_WAIT_SECONDS);
waited = true;
}
try {
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
return false;
}
}
}
}
private Bridged() {
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -1,9 +1,9 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
/**
* The authorization table (CB-505), stated once and enforced on both entry paths.
*
* <p>Most of these rules are already true de facto — {@code FleetMcp} derives a worker's identity
* <p>Most of these rules are already true de facto — {@code BridgeMcp} derives a worker's identity
* from the connection rather than reading it from an argument, so a worker has never been able to
* reply <em>as</em> another worker over MCP. What was missing is that the REST surface trusted the
* session id in the URL path, and neither surface checked role at all. This class makes the
@@ -46,28 +46,18 @@ public final class Authz {
return false; // authenticated as nothing ⇒ authorized for nothing
}
return switch (action) {
// Fleet lifecycle is the primary's alone — spawn, stop, drain. An architect
// deliberately does NOT get these (CB-548), so it cannot tear down or stand up workers
// even though it coordinates them; and a worker driving any of these would be a worker
// escalating into the orchestrator role.
case SPAWN, STOP, DRAIN -> caller.isPrimary();
// Orchestration is the primary's alone. A worker driving spawn/stop/send would be a
// worker escalating into the orchestrator role.
case SPAWN, STOP, SEND, DRAIN -> caller.isPrimary();
// Delivering a turn is open to the primary and the architect: an architect delegates
// to workers (that is the role's point) but still has no lifecycle rights. A worker is
// excluded — sending would be it escalating.
case SEND -> caller.isPrimary() || caller.isArchitect();
// The load-bearing rule: a caller acts only as the pane it occupies. CB-532 widened who
// that can be — a lead answering another lead is replying for its OWN terminal, which
// this already permits — while the rule itself is unchanged, and is what stops anyone
// forging a reply for a rendezvous someone else is waiting on. An architect's own pane
// passes through the same check, so it can answer a funnel that delegated to it. An
// unnamed primary (token/loopback, no pane) owns nothing and is still excluded.
// The load-bearing rule: a worker acts only as itself. The primary is deliberately
// excluded — a reply/ask is a worker's own turn output, and letting the primary forge
// one would corrupt the rendezvous correlation it is itself waiting on.
case REPLY, ASK -> caller.ownsSession(targetSession);
// Observation is open to every authenticated role: a worker legitimately polls its own
// Observation is open to both authenticated roles: a worker legitimately polls its own
// status, and the roster carries no secrets.
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
case READ, METRICS -> caller.isPrimary() || caller.isWorker();
};
}
@@ -0,0 +1,142 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
/**
* Resolves every caller to a {@link Principal}, for both entry paths into the core (CB-501).
*
* <p>There are two of them and they are not layered the way the docs suggest: {@code BridgeMcp}
* calls the service layer directly and is mounted as a raw servlet (so it never passes through a
* Javalin filter), while the REST routes historically resolved no identity at all. Both now
* delegate here, so the authorization rules are stated once instead of drifting apart.
*
* <p><strong>Resolution order</strong> — connection identity first, token second, nothing third:
* <ol>
* <li>A loopback peer PID that maps to the pinned {@code primary.terminal} pane (CB-307) ⇒
* {@link Role#PRIMARY}. The pane mapping is as unforgeable as a worker's, and the config
* explicitly names that pane as the primary's own — without this rule a primary running
* <em>inside</em> a herdr pane is misread as a worker and locked out of orchestration.</li>
* <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#WORKER}. This is
* unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured
* regardless of auth mode, so enabling auth never breaks the fleet.</li>
* <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li>
* <li>Otherwise, under {@code loopback-trust}, a loopback caller ⇒ {@link Role#PRIMARY}
* (the historical behaviour, now an explicit configured choice).</li>
* <li>Otherwise {@link Role#ANONYMOUS}.</li>
* </ol>
*/
public final class CallerResolver {
private final ConnectionIdentity identity;
private final boolean tokenMode;
private final byte[] expectedToken; // null unless tokenMode
private final String pinnedPrimaryTerminal; // null unless primary.terminal is configured
/** Loopback-trust resolver: no token required, historical behaviour. */
public CallerResolver(ConnectionIdentity identity) {
this(identity, false, null, null);
}
/** As {@link #CallerResolver(ConnectionIdentity, boolean, String, String)} with no pin. */
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
this(identity, tokenMode, token, null);
}
/**
* @param identity connection-based worker identification
* @param tokenMode when true, a non-worker caller must present a valid bearer token
* @param token the expected bearer token; required (non-blank) when
* {@code tokenMode}
* @param pinnedPrimaryTerminal the primary's own herdr {@code terminal_id} from
* {@code primary.terminal} ({@code null}/blank = unpinned); a
* caller resolving to this pane is the primary, not a worker
*/
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
String pinnedPrimaryTerminal) {
if (tokenMode && (token == null || token.isBlank())) {
throw new IllegalArgumentException(
"auth.mode=token requires a non-empty token; check that the env var named by "
+ "auth.tokenEnv is exported to the daemon's environment");
}
this.identity = identity;
this.tokenMode = tokenMode;
this.expectedToken = tokenMode ? token.getBytes(StandardCharsets.UTF_8) : null;
this.pinnedPrimaryTerminal =
pinnedPrimaryTerminal == null || pinnedPrimaryTerminal.isBlank()
? null : pinnedPrimaryTerminal;
}
/**
* Resolve the caller of a request.
*
* @param remoteAddr the connection's remote address
* @param remotePort the connection's remote port (used for the peer-PID lookup)
* @param authorizationHeader the raw {@code Authorization} header, or {@code null}
*/
public Principal resolve(String remoteAddr, int remotePort, String authorizationHeader) {
ConnectionIdentity.Caller c = identity.resolve(remoteAddr, remotePort);
if (c.terminal() != null) {
if (c.terminal().equals(pinnedPrimaryTerminal)) {
// The config names this pane as the primary's own. The pane mapping is exactly as
// unforgeable as a worker's, so it outranks the token path — no credential needed.
return Principal.primary(c.pid());
}
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
}
if (tokenMode) {
return presentedTokenMatches(authorizationHeader)
? Principal.primary(c.pid())
: Principal.anonymous();
}
// loopback-trust: same-host callers that are not workers are the primary. A non-loopback
// caller is anonymous even here — and startup refuses that combination anyway
// (BridgedConfig.validateAuthExposure), so this is defence in depth, not the control.
return isLoopback(remoteAddr) ? Principal.primary(c.pid()) : Principal.anonymous();
}
/** The working directory of the calling process (CB-112 spawn cwd inheritance), or {@code null}. */
public String cwdForPid(long pid) {
return identity.cwdForPid(pid);
}
/** True when auth requires a bearer token of non-worker callers. */
public boolean tokenMode() {
return tokenMode;
}
private boolean presentedTokenMatches(String authorizationHeader) {
String presented = bearerValue(authorizationHeader);
if (presented == null) {
return false;
}
// Constant-time: MessageDigest.isEqual does not short-circuit on the first differing byte,
// so a token cannot be recovered a byte at a time by timing the response.
return MessageDigest.isEqual(presented.getBytes(StandardCharsets.UTF_8), expectedToken);
}
/** Extract the credential from {@code Authorization: Bearer <token>}, or {@code null}. */
private static String bearerValue(String header) {
if (header == null) {
return null;
}
String h = header.trim();
if (h.length() < 7 || !h.regionMatches(true, 0, "Bearer ", 0, 7)) {
return null;
}
String token = h.substring(7).trim();
return token.isEmpty() ? null : token;
}
private static boolean isLoopback(String remoteAddr) {
if (remoteAddr == null) {
return false;
}
return remoteAddr.equals("127.0.0.1") || remoteAddr.equals("::1")
|| remoteAddr.equals("0:0:0:0:0:0:0:1") || remoteAddr.startsWith("127.");
}
}
@@ -0,0 +1,57 @@
package dev.ltms.bridged.auth;
/**
* A resolved caller: its {@link Role}, and — for a worker — the herdr {@code terminal_id} that
* identifies which worker it is (CB-501).
*
* @param role what this caller is authorized to act as
* @param terminal the worker's herdr terminal id; {@code null} for {@code PRIMARY}/{@code ANONYMOUS}
* @param pid the connecting process id, or {@code -1} when not resolvable (audit context)
*/
public record Principal(Role role, String terminal, long pid) {
/** A caller authenticated as nothing — the default when no check establishes anything else. */
public static Principal anonymous() {
return new Principal(Role.ANONYMOUS, null, -1);
}
/** The orchestrating session. */
public static Principal primary(long pid) {
return new Principal(Role.PRIMARY, null, pid);
}
/** A worker peer, identified by its herdr pane. */
public static Principal worker(String terminal, long pid) {
return new Principal(Role.WORKER, terminal, pid);
}
public boolean isPrimary() {
return role == Role.PRIMARY;
}
public boolean isWorker() {
return role == Role.WORKER;
}
public boolean isAnonymous() {
return role == Role.ANONYMOUS;
}
/**
* Whether this caller may act <em>as</em> {@code sessionId} — the "own session only" rule that
* keeps one worker from replying or asking on another's behalf. Only a worker can own a
* session, and only its own.
*/
public boolean ownsSession(String sessionId) {
return isWorker() && terminal != null && terminal.equals(sessionId);
}
/** Short, non-sensitive description for audit lines and error details. */
public String describe() {
return switch (role) {
case WORKER -> "worker:" + terminal;
case PRIMARY -> "primary";
case ANONYMOUS -> "anonymous";
};
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
/**
* What a caller is allowed to be on the bus (CB-501).
@@ -24,16 +24,6 @@ public enum Role {
*/
WORKER,
/**
* A config-declared architect slot (CB-548): a gateway-local named session on a strong-model
* profile that coordinates and delegates turns but does not own the fleet. Unforgeable like a
* worker's — derived from the connection's pane and the live terminal→slot binding, never from
* a request argument. May {@code SEND} a turn, {@code REPLY}/{@code ASK} only as its own pane,
* and {@code READ}/{@code METRICS}; may <em>not</em> {@code SPAWN}/{@code STOP}/{@code DRAIN}
* (those stay the primary's, to keep lifecycle in one pair of hands).
*/
ARCHITECT,
/** Authenticated as nothing. Authorized for nothing but {@code /healthz}. */
ANONYMOUS
}
@@ -0,0 +1,455 @@
package dev.ltms.bridged.config;
import com.fasterxml.jackson.annotation.JsonIgnoreProperties;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.dataformat.yaml.YAMLFactory;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Collections;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
/**
* {@code bridged} configuration, loaded from a YAML file (see
* {@code bridged.example.yaml}). Unknown keys are ignored so config can grow ahead
* of the code.
*
* @param bind REST/MCP listen host:port
* @param herdrSocket path to herdr's Unix socket ({@code null} → client default)
* @param worker single worker profile (legacy; superseded by {@code workers})
* @param workers named worker profiles, keyed by profile name (multi-backend fleet)
* @param defaultWorker which {@code workers} key a no-argument spawn uses ({@code null} → the
* single {@code worker}, or the sole/first profile)
* @param guard subscription-boundary allowlist
* @param worktreeRoot nullable root directory for provisioned worktrees; defaults to a sibling
* of the repo root
* @param lifecycle session lifecycle limits ({@code null} = all disabled)
* @param spawnReadyTimeoutMs max ms to wait for a spawned worker to reach an injectable state
* ({@code null} / 0 disables the poll gate — legacy non-blocking behaviour)
* @param spawnReadyPollMs poll interval while waiting for the worker to become injectable
* @param broker external AMQP broker for durable reply delivery ({@code null} → in-memory,
* soft-state {@code ReplyInbox}; present → the AMQP-backed adapter, CB-307 Stage 2)
* @param primary optional pinned primary terminal config ({@code null} → derived from connection);
* a non-blank {@code terminal} seeds {@code PrimaryRegistry} and prevents
* connection-derived overrides, CB-307
* @param placement how to choose a worker profile for an unqualified spawn:
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
* historical behaviour), CB-501
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record BridgedConfig(
Bind bind,
String herdrSocket,
Worker worker,
Map<String, Worker> workers,
String defaultWorker,
Guard guard,
String worktreeRoot,
Lifecycle lifecycle,
Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs,
Broker broker,
Primary primary,
String placement,
Auth auth) {
@JsonIgnoreProperties(ignoreUnknown = true)
public record Bind(String host, int port) {
public Bind {
if (host == null || host.isBlank()) host = "127.0.0.1";
if (port <= 0) port = 8765;
}
}
/**
* @param profile ccs profile a worker is spawned under (Stage-1: {@code ltms-local})
* @param baseUrl the off-subscription endpoint injected as {@code ANTHROPIC_BASE_URL}
* @param model model alias, injected as {@code ANTHROPIC_MODEL} (may be {@code null})
* @param configDir {@code CLAUDE_CONFIG_DIR} so the worker inherits the profile's
* skills/MCP/hooks (may be {@code null})
* @param tokenEnv name of the host env var holding the worker's auth token; its value
* is injected as {@code ANTHROPIC_AUTH_TOKEN} (never stored in config)
* @param argv launch command; defaults to {@code ["claude"]}
* @param placement where a worker lands: {@code "tab"} (default — its own tab in the
* worker space) or {@code "pane"} (legacy — split the focused tab)
* @param workspace label of the dedicated worker space; found-or-created on first
* spawn (default {@code "bridged-workers"}). A future per-session
* layout is just a distinct label here — the shared space is default.
* @param tabLabel template for a worker tab's label; {@code {profile}}/{@code {model}}
* and {@code {n}} (per-worker number, to keep sibling tabs distinct)
* are substituted (default {@code "worker: {profile} #{n}"})
* @param mcpUrl bridge MCP URL to provision into the worker's {@code configDir} so it
* can call {@code bridge_reply} ({@code null}/blank → no provisioning; the
* worker won't reply, only the fallback/timeout resolves the send)
* @param cwd fixed working directory for this profile's workers (CB-112 "told otherwise");
* {@code null}/blank → inherit the primary's cwd, else the daemon's
* @param parityOverlay repo-relative paths copied primary→worktree for config parity; null/empty
* defaults to a sensible set of local config files
* @param gitTokenEnv name of the host env var holding the git-forge API token; when set, its
* value is injected as {@code GITEA_TOKEN} so the worker can open its own PR
* at checkpoint (CB-302). {@code null}/blank ⇒ no token is injected
* (minimal-grant default — push over SSH stays free, PR-create is opt-in)
* @param gitHostEnv name of the host env var holding the forge host (default {@code GITEA_HOST});
* injected as {@code GITEA_HOST} <em>only</em> when {@code gitTokenEnv} is set
* @param weight relative selection weight for {@code placement: weighted}. Absent or
* non-positive ⇒ 1.0. Weights are normalised by the policy, so they need
* not sum to 1.0.
* @param maxLoad max live workers allowed on this profile at one time; absent or
* non-positive ⇒ unlimited. Live means any session the registry still owns
* (acquired and not yet released), in any state.
* @param kind which peer launcher spawns this profile: {@code "claude-code"} (default —
* the {@link dev.ltms.bridged.worker.ClaudeCodeLauncher}) or {@code "opencode"}.
* The {@code CompositePeerLauncher} routes {@code spawn}/reap by this value, so
* each adapter drives only its own kind. Normalised to lower-case; blank ⇒ the
* default. It selects the adapter, not the transport — placement, tabs, cwd, and
* the readiness gate are kind-independent and stay in the shared base.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd,
List<String> parityOverlay,
String gitTokenEnv, String gitHostEnv,
String kind,
Map<String, String> env,
Float weight,
Integer maxLoad) {
/** Peer kind spawned by {@link dev.ltms.bridged.worker.ClaudeCodeLauncher} (the default). */
public static final String KIND_CLAUDE_CODE = "claude-code";
/** Peer kind spawned by the opencode adapter (CB-402). */
public static final String KIND_OPENCODE = "opencode";
/**
* Peer kind spawned by the Codex adapter (CB-528). Like {@link #KIND_OPENCODE} it carries
* its own argv and never inherits the Claude binary, and it sits outside the
* {@code ANTHROPIC_BASE_URL} subscription guard because Codex has no such seam.
*/
public static final String KIND_CODEX = "codex";
public Worker {
// A claude-code worker defaults its launch command to `claude`; other kinds carry their own
// argv (e.g. `opencode`) and must not inherit the Claude binary — so only default when unset
// AND this is the claude-code kind.
String k = (kind == null || kind.isBlank()) ? KIND_CLAUDE_CODE : kind.toLowerCase();
argv = (argv == null || argv.isEmpty())
? (KIND_CLAUDE_CODE.equals(k) ? List.of("claude") : List.of(k))
: List.copyOf(argv);
kind = k;
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_WORKER_TOKEN" : tokenEnv;
placement = (placement == null || placement.isBlank()) ? "tab" : placement.toLowerCase();
workspace = (workspace == null || workspace.isBlank()) ? "bridged-workers" : workspace;
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? "worker: {profile} #{n}" : tabLabel;
// CB-525: .mcp.json is deliberately NOT here. Replicating the primary's MCP config gave a
// worker the primary's IDE servers, which are bound to the primary's checkout — so its
// navigation returned paths outside its own worktree. GitWorktrees now neutralizes that
// file instead; a worker's tools are whatever its launcher mounts.
parityOverlay = (parityOverlay == null || parityOverlay.isEmpty())
? List.of(".claude/settings.local.json", ".env", ".envrc")
: List.copyOf(parityOverlay);
// gitTokenEnv stays null when unset (opt-in). gitHostEnv defaults so operators enabling
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
gitHostEnv = (gitHostEnv == null || gitHostEnv.isBlank()) ? "GITEA_HOST" : gitHostEnv;
env = (env == null) ? Map.of() : Map.copyOf(env);
weight = (weight == null || weight <= 0.0f) ? 1.0f : weight;
maxLoad = (maxLoad == null || maxLoad <= 0) ? null : maxLoad;
}
/**
* Backward-compatible constructor without the CB-302 git-forge fields — the worker is
* granted no PR-create token (push over SSH is unaffected). Keeps pre-CB-302 call sites
* (and any {@code workers:} YAML that omits the git keys) working unchanged.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, null, null, null, null, null, null);
}
/**
* Backward-compatible constructor with the CB-302 git-forge fields but no explicit peer
* {@code kind} — defaults to {@link #KIND_CLAUDE_CODE}. Keeps pre-CB-402 call sites working.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, null, null, null, null);
}
/**
* Backward-compatible constructor without the CB-511 {@code env:} passthrough — the worker
* gets the daemon's PATH and nothing else. Keeps pre-CB-511 call sites working.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, null, null, null);
}
/** A copy with {@code profile} set — used to default a profile to its {@code workers} key. */
public Worker withProfile(String p) {
return new Worker(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad);
}
/** True when this profile is served by the Claude Code adapter (the default kind). */
public boolean isClaudeCode() {
return KIND_CLAUDE_CODE.equals(kind);
}
/** True when this profile is served by the opencode adapter (CB-402). */
public boolean isOpenCode() {
return KIND_OPENCODE.equals(kind);
}
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
public boolean hasGitToken() {
return gitTokenEnv != null && !gitTokenEnv.isBlank();
}
/** True when workers should land in their own tab in the worker space. */
public boolean tabPlacement() {
return "tab".equals(placement);
}
/** True when the bridge MCP should be mounted into a spawned worker (via launch flags). */
public boolean hasMcp() {
return mcpUrl != null && !mcpUrl.isBlank();
}
/**
* Render {@link #tabLabel} for the {@code n}-th worker (substitutes
* {@code {profile}}/{@code {model}}/{@code {n}}), so sibling worker tabs are distinct.
*/
public String renderTabLabel(long n) {
return tabLabel
.replace("{profile}", profile == null ? "" : profile)
.replace("{model}", model == null ? "" : model)
.replace("{n}", Long.toString(n));
}
}
/**
* Session lifecycle limits. All knobs are opt-in: {@code null} or {@code 0} disables the
* feature so existing configs keep the previous behaviour.
*
* @param idleTtlSeconds max seconds a {@code READY}/{@code DONE} session may sit idle
* before it is reaped ({@code null} → disabled)
* @param contextCap max delegated turns a session serves before force-release
* ({@code null} → disabled)
* @param drainTimeoutSeconds seconds to wait for {@code BUSY} sessions to finish before
* forced teardown on shutdown (default 5 when unset)
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Lifecycle(Integer idleTtlSeconds, Integer contextCap, Integer drainTimeoutSeconds) {
}
/**
* External AMQP broker for durable, cross-restart reply delivery (CB-307 Stage 2). Its mere
* presence swaps the in-memory {@code ReplyInbox} for the AMQP-backed adapter; absent, bridged
* stays soft-state. Production default is LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1
* and is a URI-only swap.
*
* @param uri AMQP connection URI, e.g. {@code amqp://guest:guest@127.0.0.1:5672/}. Blank/{@code null}
* ⇒ the broker block is treated as absent (in-memory adapter).
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Broker(String uri) {
/** True when a usable broker URI is configured (an empty block does not enable AMQP). */
public boolean isConfigured() {
return uri != null && !uri.isBlank();
}
}
/**
* Optional pinned primary terminal config (CB-307). When present with a non-blank
* {@code terminal}, the bridge uses this as the primary's herdr identity instead of
* deriving it from the MCP connection. It feeds two consumers: the push loop (where to nudge
* when replies land), and caller resolution — a caller whose connection maps to this pane is
* the primary, where the pane match would otherwise classify it as a worker. Pin it when the
* primary runs <em>inside</em> a herdr pane; it also helps off-host or non-herdr primaries,
* where connection-derived identity is unavailable and only the nudge target matters.
*
* @param terminal the primary's herdr {@code terminal_id} ({@code null}/blank → derive)
* @param pushReminders max reminder nudges before giving up (default 5)
* @param pushBackoffMs delay between reminders in ms (default 15000)
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Primary(String terminal, Integer pushReminders, Integer pushBackoffMs) {
/** @return configured reminder cap, or 5 */
public int remindersOrDefault() {
return pushReminders != null ? pushReminders : 5;
}
/** @return configured backoff in ms, or 15000 */
public long backoffMsOrDefault() {
return pushBackoffMs != null ? pushBackoffMs.longValue() : 15_000L;
}
}
/**
* API authentication (CB-501). Governs how a caller that is <em>not</em> an on-host worker
* pane proves it is the primary.
*
* <p>Worker identity never depends on this block: a loopback peer PID that maps to a herdr
* pane is unforgeable and is always honoured (see
* {@link dev.ltms.bridged.mcp.ConnectionIdentity}). This only decides what happens for
* <em>everyone else</em>.
*
* @param mode {@code "loopback-trust"} (default) — any loopback caller that is not a known
* worker is the primary, no credential needed; this is the historical
* behaviour, now chosen explicitly rather than implied. {@code "token"} — such
* a caller must present {@code Authorization: Bearer <token>} or it is
* {@code ANONYMOUS} and authorized for nothing.
* @param tokenEnv name of the host env var holding the bearer token; the literal value is
* never stored in config. Defaults to {@code BRIDGED_API_TOKEN}. Only read
* when {@code mode} is {@code token}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Auth(String mode, String tokenEnv) {
/** Historical behaviour: loopback non-worker ⇒ primary, no credential. */
public static final String MODE_LOOPBACK_TRUST = "loopback-trust";
/** A non-worker caller must present a valid bearer token to be the primary. */
public static final String MODE_TOKEN = "token";
public Auth {
mode = (mode == null || mode.isBlank()) ? MODE_LOOPBACK_TRUST : mode.toLowerCase();
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_API_TOKEN" : tokenEnv;
}
/** True when a bearer token is required of every non-worker caller. */
public boolean tokenMode() {
return MODE_TOKEN.equals(mode);
}
}
/**
* Subscription boundary. Only these hosts may back a worker's
* {@code ANTHROPIC_BASE_URL}; the primary must carry none.
*
* @param offSubscriptionHosts hostnames allowed for worker base_urls
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Guard(List<String> offSubscriptionHosts) {
public Guard {
offSubscriptionHosts = offSubscriptionHosts == null ? List.of() : List.copyOf(offSubscriptionHosts);
}
public Set<String> hostSet() {
return Set.copyOf(offSubscriptionHosts);
}
}
/**
* The effective worker profiles, keyed by profile name. Prefers the {@code workers} map (each
* value's {@code profile} defaulted to its key); falls back to the legacy singular {@code worker}
* (keyed by its own profile). Empty if neither is configured.
*/
public Map<String, Worker> workerProfiles() {
if (workers != null && !workers.isEmpty()) {
Map<String, Worker> out = new LinkedHashMap<>();
workers.forEach((name, w) -> out.put(name,
(w.profile() == null || w.profile().isBlank()) ? w.withProfile(name) : w));
// Deliberately NOT Map.copyOf: its iteration order is salted per JVM run, which would
// discard the YAML definition order built above. Placement tie-breaks on candidate
// order (see WeightedRoundRobinPolicy), so losing it makes equal-weight placement
// non-reproducible across restarts. Unmodifiable-wrap instead of copy-and-scramble.
return Collections.unmodifiableMap(out);
}
if (worker != null) {
String name = (worker.profile() == null || worker.profile().isBlank()) ? "default" : worker.profile();
return Map.of(name, worker);
}
return Map.of();
}
/**
* The profile a no-argument spawn uses: {@code defaultWorker} if set, else the legacy single
* {@code worker}'s profile, else the sole/first configured profile, else {@code null}.
*/
public String defaultProfile() {
if (defaultWorker != null && !defaultWorker.isBlank()) {
return defaultWorker;
}
if (worker != null && worker.profile() != null && !worker.profile().isBlank()) {
return worker.profile();
}
Map<String, Worker> p = workerProfiles();
return p.isEmpty() ? null : p.keySet().iterator().next();
}
private static final ObjectMapper YAML = new ObjectMapper(new YAMLFactory());
/** Load and validate config from {@code path}. */
public static BridgedConfig load(Path path) {
try {
BridgedConfig cfg = YAML.readValue(Files.readString(path), BridgedConfig.class);
return cfg.withDefaults();
} catch (IOException e) {
throw new UncheckedIOException("cannot read bridged config at " + path, e);
}
}
/** Fill in nested defaults so callers never see nulls for structural fields. */
public BridgedConfig withDefaults() {
Bind b = bind != null ? bind : new Bind(null, 0);
Guard g = guard != null ? guard : new Guard(List.of());
Lifecycle l = lifecycle != null ? lifecycle : new Lifecycle(null, null, null);
Integer timeout = (spawnReadyTimeoutMs != null) ? spawnReadyTimeoutMs : 20000;
Integer pollMs = (spawnReadyPollMs != null) ? spawnReadyPollMs : 300;
Auth a = auth != null ? auth : new Auth(null, null);
String placementOrDefault = (placement != null && !placement.isBlank()) ? placement : "fixed";
// broker is left as-is: null (or an empty/blank uri) keeps the in-memory soft-state inbox.
// primary is left as-is: null defaults to connection-derived identity.
return new BridgedConfig(b, herdrSocket, worker, workers, defaultWorker, g, worktreeRoot, l, timeout, pollMs, broker, primary, placementOrDefault, a);
}
/**
* Reject a configuration whose network exposure outruns its authentication (CB-501).
*
* <p>{@code loopback-trust} means "any caller that is not a known worker pane is the primary" —
* safe only because the OS refuses non-local connections to a loopback bind. Widen
* {@code bind.host} without switching to {@code token} mode and that sentence becomes "any
* client that can reach this port is the primary", which is the most privileged role on the
* bus. Rather than document the hazard, make it unrepresentable: fail fast at startup.
*
* @throws IllegalStateException when a non-loopback bind is paired with {@code loopback-trust}
*/
public void validateAuthExposure() {
String host = bind().host();
if (isLoopbackBind(host) || auth().tokenMode()) {
return;
}
throw new IllegalStateException(
"refusing to start: bind.host=" + host + " is not loopback, but auth.mode="
+ auth().mode() + ". A non-loopback bind treats every unauthenticated "
+ "caller as the primary (spawn/stop/send/drain on any session). Set "
+ "auth.mode: token (with auth.tokenEnv) before exposing this port, or "
+ "bind to 127.0.0.1 and put a reverse proxy in front.");
}
/** True for the loopback addresses and the unspecified-but-local forms we treat as same-host. */
private static boolean isLoopbackBind(String host) {
if (host == null || host.isBlank()) {
return true; // Bind's own default is 127.0.0.1
}
String h = host.trim().toLowerCase();
return h.equals("127.0.0.1") || h.equals("::1") || h.equals("localhost")
|| h.startsWith("127.");
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.guard;
package dev.ltms.bridged.guard;
/** Thrown when the subscription boundary would be violated. Never swallow this. */
public class GuardException extends RuntimeException {
@@ -1,4 +1,4 @@
package dev.ltms.fleet.guard;
package dev.ltms.bridged.guard;
import java.net.URI;
import java.net.URISyntaxException;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
/**
* A herdr agent's lifecycle state, as reported by {@code agent_status}. Drives the
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.JsonNode;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
/**
* Raised when a herdr call fails: transport error, or an {@code error} envelope
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.slf4j.Logger;
@@ -41,16 +41,6 @@ public final class WorkspaceControl {
return out;
}
/** Every tab in {@code workspaceId}, in herdr's order. */
public List<Tab> listTabs(String workspaceId) {
JsonNode result = herdr.call("tab.list", Map.of("workspace_id", workspaceId));
List<Tab> out = new ArrayList<>();
for (JsonNode t : result.path("tabs")) {
out.add(Tab.from(t));
}
return out;
}
/** The first workspace with this exact label, if any. */
public Optional<Workspace> findByLabel(String label) {
return listWorkspaces().stream()
@@ -1,29 +1,23 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.TurnToken;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.msg.Rendezvous;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Objects;
import java.util.Set;
import java.util.TreeSet;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.regex.Pattern;
/**
* The CB-106 completion fallback: bridges the {@link Injector}'s turn-completion signal to the
* {@link Rendezvous} so a blocking {@code fleet_send} resolves even when the worker finishes its
* task without ever calling {@code fleet_reply} — the common case for a real delegated coding task.
* {@link Rendezvous} so a blocking {@code bridge_send} resolves even when the worker finishes its
* task without ever calling {@code bridge_reply} — the common case for a real delegated coding task.
*
* <p>On a confirmed {@code working → idle} boundary it scrapes the worker's recent transcript and
* resolves the awaiting send with that tail (a {@link Rendezvous.Kind#COMPLETION} resolution, so the
* caller can tell a scrape from a structured reply). It scrapes only when a send is actually waiting
* — a fleet worker's own turns, or a send that already timed out, cost no herdr traffic. An explicit
* {@code fleet_reply} that raced in first wins; {@link Rendezvous#resolveCompletion} is then a no-op.
* {@code bridge_reply} that raced in first wins; {@link Rendezvous#resolveCompletion} is then a no-op.
*
* <p>It also handles the CB-109 stall signal ({@link #onTurnFailed}): a worker that ran a turn then
* wedged in an {@code unknown} state resolves the send as a failure (with the error screen as
@@ -38,7 +32,7 @@ import java.util.regex.Pattern;
* <p><strong>Waiter-specific resolution (CB-116).</strong> On delivery we also capture the exact
* {@link Rendezvous} waiter this turn belongs to, and the completion/failure fallbacks resolve
* <em>that</em> waiter — never "whatever send is waiting now". A completion fallback runs on a virtual
* thread and can land after the worker's {@code fleet_reply} already resolved the turn and the
* thread and can land after the worker's {@code bridge_reply} already resolved the turn and the
* <em>next</em> send opened its own waiter on the same session; resolving the current waiter would
* then deliver turn N's stale scrape as turn N+1's answer. Targeting the captured waiter makes a late
* completion a harmless no-op (its waiter is already done) instead of a cross-turn stale reply.
@@ -61,13 +55,8 @@ public final class CompletionResolver implements TurnListener {
/** Cap the scraped tail so a long transcript can't return an unbounded blob. */
static final int MAX_SCRAPE_CHARS = 4000;
private static final String CLIPPED_PANE_TAIL_MARKER =
"[Pane tail clipped: member did not call fleet_reply.]";
private final AgentControl agents;
private final Rendezvous rendezvous;
private final ExhaustedPatternLookup exhaustedPatterns;
private final ExhaustionSink exhaustionSink;
/**
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
@@ -85,36 +74,23 @@ public final class CompletionResolver implements TurnListener {
private final ConcurrentHashMap<String, InFlight> inFlight = new ConcurrentHashMap<>();
/**
* @param exhaustedPatterns CB-578 stage A: per-target lookup for a profile's configured
* usage-limit refusal pattern. Required — there is deliberately no
* defaulting overload; a caller that does not want the classification
* must pass an explicit inert value ({@link ExhaustedPatternLookup#none()}).
* @param exhaustionSink CB-578 stage B: notified when a {@code BACKEND_EXHAUSTED}
* classification actually resolves a waiter. Required for the same
* reason as {@code exhaustedPatterns} — pass {@link ExhaustionSink#none()}
* to opt out.
*/
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
ExhaustionSink exhaustionSink) {
public CompletionResolver(AgentControl agents, Rendezvous rendezvous) {
this.agents = agents;
this.rendezvous = rendezvous;
this.exhaustedPatterns = Objects.requireNonNull(exhaustedPatterns, "exhaustedPatterns");
this.exhaustionSink = Objects.requireNonNull(exhaustionSink, "exhaustionSink");
}
@Override
public void onDelivered(String target, TurnToken token) {
public void onDelivered(String target) {
// Capture the exact waiter this turn belongs to (CB-116) and snapshot the pane's pre-turn
// content — what it shows *before* the just-delivered turn produces output — as the staleness
// reference (CB-115). Done synchronously (like the delivering send itself) so both are in
// place before this turn's completion can fire.
captureBaseline(target, token);
captureBaseline(target);
}
/** Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link #onDelivered}). */
void captureBaseline(String target, TurnToken token) {
CompletableFuture<Rendezvous.Resolution> waiter = token.waiter();
void captureBaseline(String target) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null) {
inFlight.remove(target); // no send is waiting on this delivery — nothing to resolve later
return;
@@ -146,46 +122,25 @@ public final class CompletionResolver implements TurnListener {
Thread.ofVirtual().name("completion-" + target).start(() -> resolve(target, turn));
}
/**
* Resolve the completed turn before adapter housekeeping can erase its rendered output. This is
* intentionally synchronous and used only when a post-turn context reset is enabled; the normal
* path remains off-loaded so polling is not blocked by a scrape.
*/
public void resolveBeforePostAction(String target) {
resolve(target, inFlight.get(target));
}
@Override
public void onTurnFailed(String target) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, null));
}
@Override
public void onTurnFailed(String target, String reason) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, reason));
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn));
}
/** Synchronous resolve (the unit-testable core of {@link #onTurnComplete}). */
void resolve(String target, InFlight turn) {
CompletableFuture<Rendezvous.Resolution> waiter = turn == null ? null : turn.waiter();
if (waiter == null || waiter.isDone()) {
// Nobody is blocked on THIS turn (it had no send, or its fleet_reply already won). Skip
// Nobody is blocked on THIS turn (it had no send, or its bridge_reply already won). Skip
// the scrape; resolving the current waiter here would be the CB-116 cross-turn stale reply.
inFlight.remove(target, turn);
return;
}
String tail;
String assistantBlock = null;
int originalLength = 0;
boolean clipped = false;
boolean scrapeFailed = false;
try {
assistantBlock = lastAssistantBlock(agents.read(target, SCRAPE_SOURCE));
originalLength = assistantBlock.strip().length();
clipped = originalLength > MAX_SCRAPE_CHARS;
tail = clip(assistantBlock);
tail = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
} catch (RuntimeException e) {
// The worker finished but we couldn't read its screen — still resolve the send so the
// caller unblocks; an empty tail beats hanging until the caller's timeout.
@@ -197,7 +152,7 @@ public final class CompletionResolver implements TurnListener {
// Misattribution guard (CB-115): if the scrape is byte-identical to the pane content at
// delivery, this turn produced no new output — the boundary belongs to the previous turn's
// wind-down (common on rapid back-to-back sends). Suppress rather than resolve the send with
// a stale answer; the real fleet_reply (or a later genuine completion) resolves it instead.
// a stale answer; the real bridge_reply (or a later genuine completion) resolves it instead.
// A scrape that failed to read is exempt — an empty tail there is "couldn't see", not "no change".
String baseline = turn.baseline();
if (!scrapeFailed && baseline != null && baseline.equals(tail)) {
@@ -205,33 +160,8 @@ public final class CompletionResolver implements TurnListener {
target);
return; // keep the in-flight record: a later genuine completion still needs it
}
// CB-578 stage A: a turn that ended with no fleet_reply AND whose scrape matches the
// backend's configured usage-limit pattern is a refusal, not an answer. Classify it as
// BACKEND_EXHAUSTED rather than handing the caller a scrape that reads like a real reply.
if (!scrapeFailed) {
Pattern exhausted = exhaustedPatterns.patternFor(target);
String matchedLine = exhausted == null ? null : firstMatchingLine(assistantBlock, exhausted);
if (matchedLine != null) {
String reason = "backend exhausted (usage limit): " + matchedLine;
if (rendezvous.resolveExhausted(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("completion for {} classified BACKEND_EXHAUSTED (no fleet_reply; scrape "
+ "matched the profile's exhausted pattern): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
}
return;
}
}
String completion = clipped ? tail + "\n" + CLIPPED_PANE_TAIL_MARKER : tail;
if (rendezvous.resolveCompletion(waiter, completion)) {
if (rendezvous.resolveCompletion(waiter, tail)) {
inFlight.remove(target, turn);
if (clipped) {
log.warn("completion scrape for {} clipped from {} chars to the {} char cap; "
+ "member did not call fleet_reply, so the pane tail is partial",
target, originalLength, MAX_SCRAPE_CHARS);
}
log.debug("resolved send to {} via turn-completion fallback ({} chars scraped)",
target, tail.length());
}
@@ -239,11 +169,6 @@ public final class CompletionResolver implements TurnListener {
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
void fail(String target, InFlight turn) {
fail(target, turn, null);
}
/** Synchronous fail with an optional reason supplied by a dropped worker queue. */
void fail(String target, InFlight turn, String explicitReason) {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fall back to the currently-registered waiter (unambiguous — that send never completed, so
// no next turn exists to confuse it with).
@@ -253,64 +178,23 @@ public final class CompletionResolver implements TurnListener {
inFlight.remove(target, turn); // nobody blocked on this worker — nothing to fail
return;
}
String reason = explicitReason;
if (reason == null || reason.isBlank()) {
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
}
String reason;
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
}
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("failing send to {} via turn-stall fallback: {}", target, reason);
log.debug("failed send to {} via turn-stall fallback", target);
}
}
/**
* The first line of {@code text} matching {@code pattern}, stripped — the CB-578 stage A
* evidence carried in a {@code BACKEND_EXHAUSTED} reason so the operator sees the real refusal
* text, never a generic label. {@code null} if no line matches.
*/
static String firstMatchingLine(String text, Pattern pattern) {
if (text == null || text.isEmpty()) return null;
for (String line : text.split("\n", -1)) {
if (pattern.matcher(line).find()) {
return line.strip();
}
}
return null;
}
/**
* Coverage summary for the CB-578 stage A exhausted-pattern classification, logged at startup
* the way {@link dev.ltms.fleet.health.FleetHealthMonitor#coverage} is — so an operator can
* see whether the classification is on, and for which profiles, without reading every
* profile's config by hand.
*
* @param allProfiles every configured profile name
* @param configuredProfiles the subset of {@code allProfiles} that carry an exhausted pattern
*/
public static String coverage(Set<String> allProfiles, Set<String> configuredProfiles) {
if (configuredProfiles.isEmpty()) {
return "off (no profile has an exhaustedPattern configured; profiles: " + sorted(allProfiles) + ")";
}
Set<String> unconfigured = new TreeSet<>(allProfiles);
unconfigured.removeAll(configuredProfiles);
return unconfigured.isEmpty()
? "full (all profiles configured: " + sorted(allProfiles) + ")"
: "partial (configured: " + sorted(configuredProfiles) + "; not configured: " + sorted(unconfigured) + ")";
}
private static List<String> sorted(Set<String> names) {
return names.stream().sorted().toList();
}
private static String clip(String s) {
if (s == null) return "";
String trimmed = s.strip();
@@ -1,8 +1,7 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.msg.TurnToken;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -79,15 +78,6 @@ public final class Injector {
*/
private static final int READINESS_GRACE_POLLS = 240;
/**
* The single source for the injector poll cadence — how often the {@link StatusPoller} drives
* {@link #onStatus} at. {@code Fleetd} passes this to every {@link StatusPoller} it constructs,
* and this class reads it to state the readiness grace in seconds on the CB-562 expiry log
* instead of hardcoding "60s". One constant, so a cadence change cannot silently desync a log
* that claims a grace duration.
*/
public static final long POLL_INTERVAL_MILLIS = 250;
private final AgentControl agents;
private final TurnListener turnListener;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
@@ -130,7 +120,7 @@ public final class Injector {
}
/** A pending message and the future that completes when it has been delivered. */
private record Pending(String text, TurnToken token, CompletableFuture<Void> delivered) {
private record Pending(String text, CompletableFuture<Void> delivered) {
}
/** Per-worker delivery state, guarded by its own monitor (single writer per worker). */
@@ -142,10 +132,6 @@ public final class Injector {
boolean turnObserved; // saw a real `working` sample since that delivery (turn ran)
int unknownSinceTurn; // consecutive `unknown` samples while a delegation is outstanding (CB-109)
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
boolean postTurnPending; // completion observed; adapter housekeeping has not started yet
boolean awaitingPostTurnPickup;
boolean postTurnObserved;
int injectableSincePostTurnPickup;
synchronized void add(Pending p) {
queue.add(p);
@@ -160,9 +146,9 @@ public final class Injector {
* <p>Uses an atomic map update so a concurrent {@link #drop} cannot slip between "find the
* target" and "queue the message" and orphan it in a target it just removed.
*/
public CompletableFuture<Void> enqueue(String target, String text, TurnToken token) {
public CompletableFuture<Void> enqueue(String target, String text) {
CompletableFuture<Void> delivered = new CompletableFuture<>();
Pending p = new Pending(text, token, delivered);
Pending p = new Pending(text, delivered);
targets.compute(target, (_, existing) -> {
Target t = (existing != null) ? existing : new Target();
t.add(p); // synchronized on the Target monitor — atomic with a concurrent drop
@@ -187,15 +173,9 @@ public final class Injector {
boolean turnCompleted = false;
boolean turnFailed = false;
boolean resubmit = false;
boolean startPostTurn = false;
List<Pending> notReady = null; // queued messages failed because the worker never became ready
synchronized (t) {
if (status == AgentStatus.WORKING) {
if (t.awaitingPostTurnPickup) {
t.awaitingPostTurnPickup = false;
t.injectableSincePostTurnPickup = 0;
t.postTurnObserved = true;
}
// Definitive pickup: the worker is busy on our last message, and (if a delivery is
// outstanding) a real turn is now confirmed to be running.
t.awaitingPickup = false;
@@ -205,16 +185,6 @@ public final class Injector {
if (t.awaitingCompletion) t.turnObserved = true;
} else if (status.injectable()) { // IDLE or BLOCKED
t.unknownSinceTurn = 0;
if (t.awaitingPostTurnPickup) {
if (++t.injectableSincePostTurnPickup >= PICKUP_GRACE_POLLS) {
t.awaitingPostTurnPickup = false;
t.injectableSincePostTurnPickup = 0;
} else {
resubmit = true;
}
} else if (t.postTurnObserved) {
t.postTurnObserved = false;
}
if (t.awaitingPickup) {
if (++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
// Pickup edge was never sampled (turn faster than the poll, or status lag).
@@ -239,16 +209,11 @@ public final class Injector {
t.awaitingCompletion = false;
t.turnObserved = false;
turnCompleted = true;
if (turnListener.hasPostTurnAction(target)) {
t.postTurnPending = true;
startPostTurn = true;
}
}
// Deliver the next queued message only once the prior turn is fully settled, so a
// completion is never confused with the pickup of the following message — and only
// once the worker is available (CB-113), so we never paste into its boot window.
if (!t.awaitingCompletion && !t.postTurnPending
&& !t.awaitingPostTurnPickup && !t.postTurnObserved) {
if (!t.awaitingCompletion) {
Pending p = t.queue.peek();
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
@@ -274,11 +239,6 @@ public final class Injector {
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
log.warn("readiness grace for {} expired after {} polls ({}s): target never "
+ "became deliverable, so failing {} queued message(s) that never "
+ "reached its pane",
target, READINESS_GRACE_POLLS,
READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000, notReady.size());
t.queue.clear();
t.notReadySincePoll = 0;
}
@@ -303,8 +263,7 @@ public final class Injector {
// Reclaim the entry once the worker is fully quiescent (nothing queued, no pickup or
// completion awaited), so the map cannot grow without bound across short-lived workers.
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion
&& !t.postTurnPending && !t.awaitingPostTurnPickup && !t.postTurnObserved) {
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion) {
targets.remove(target, t);
}
}
@@ -331,21 +290,7 @@ public final class Injector {
turnListener.onTurnFailed(target);
}
if (turnCompleted) {
if (startPostTurn) {
boolean started = turnListener.onTurnCompleteWithPostAction(target);
synchronized (t) {
t.postTurnPending = false;
if (started) {
t.awaitingPostTurnPickup = true;
t.injectableSincePostTurnPickup = 0;
}
if (t.queue.isEmpty() && !t.awaitingPostTurnPickup) {
targets.remove(target, t);
}
}
} else {
turnListener.onTurnComplete(target);
}
turnListener.onTurnComplete(target);
}
if (turnFailed) {
turnListener.onTurnFailed(target);
@@ -357,7 +302,7 @@ public final class Injector {
} else {
// Baseline the pane's pre-turn content so a misattributed completion (no new output)
// can't resolve this send with the previous turn's stale answer (CB-115).
turnListener.onDelivered(target, sent.token());
turnListener.onDelivered(target);
sent.delivered().complete(null);
}
}
@@ -372,8 +317,7 @@ public final class Injector {
.filter(e -> {
synchronized (e.getValue()) {
Target t = e.getValue();
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion
|| t.postTurnPending || t.awaitingPostTurnPickup || t.postTurnObserved;
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion;
}
})
.map(java.util.Map.Entry::getKey)
@@ -399,20 +343,13 @@ public final class Injector {
hadDeliveredTurn = t.awaitingCompletion;
t.awaitingCompletion = false;
t.awaitingPickup = false;
t.postTurnPending = false;
t.awaitingPostTurnPickup = false;
t.postTurnObserved = false;
}
log.warn("{} is gone, dropping its queue: {} message(s) failed{}; cause: {}", target,
pending.size(),
hadDeliveredTurn ? " (including one turn already in flight whose completion was never confirmed)" : "",
cause.getMessage());
forget.accept(target); // the worker is gone — clear its readiness/presence too (CB-114)
for (Pending p : pending) {
p.delivered().completeExceptionally(cause);
}
// A queued send has no in-flight record, while a delivered turn does. CompletionResolver
// handles both forms and resolves its waiter at most once.
turnListener.onTurnFailed(target, cause.getMessage());
if (hadDeliveredTurn) {
turnListener.onTurnFailed(target);
}
}
}
@@ -1,8 +1,8 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.HerdrException;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -1,7 +1,7 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -1,12 +1,10 @@
package dev.ltms.fleet.inject;
import dev.ltms.fleet.msg.TurnToken;
package dev.ltms.bridged.inject;
/**
* Notified when a worker's delegated turn is observed to complete — a confirmed
* {@code WORKING → IDLE} transition after a delivery. This is the CB-106 completion signal the
* {@code CompletionResolver} uses to resolve a blocked send whose worker never called
* {@code fleet_reply}. Kept as a seam so the {@link Injector} needs no dependency on the message
* {@code bridge_reply}. Kept as a seam so the {@link Injector} needs no dependency on the message
* layer and stays unit-testable with a capturing fake.
*/
@FunctionalInterface
@@ -15,25 +13,6 @@ public interface TurnListener {
/** A worker's delegated turn finished (worker returned to idle after visibly working). */
void onTurnComplete(String target);
/**
* Whether completion must pause ordinary delivery while adapter-specific housekeeping starts.
* This is queried before the injector considers the next queued message, closing the same-tick
* ordering gap. Implementations must not mutate state here.
*/
default boolean hasPostTurnAction(String target) {
return false;
}
/**
* Complete the delegated turn and start its non-turn housekeeping operation.
*
* @return true when the injector must observe that operation settle before the next delivery
*/
default boolean onTurnCompleteWithPostAction(String target) {
onTurnComplete(target);
return false;
}
/**
* A worker that visibly ran a delegated turn then wedged in a non-idle, non-working state
* (CB-109) — e.g. an error screen herdr classifies as {@code unknown} — so no
@@ -43,14 +22,6 @@ public interface TurnListener {
default void onTurnFailed(String target) {
}
/**
* As {@link #onTurnFailed(String)}, carrying the reason a worker became unreachable. The default
* keeps existing listeners working while allowing the completion resolver to report a useful cause.
*/
default void onTurnFailed(String target, String reason) {
onTurnFailed(target);
}
/**
* A message was just delivered into {@code target}'s pane (CB-115). Fired so the completion
* resolver can snapshot the pane's pre-turn content: a later {@link #onTurnComplete} whose
@@ -59,7 +30,7 @@ public interface TurnListener {
* resolve the send with the previous turn's stale answer. A default no-op keeps the interface
* functional for callers that don't scrape.
*/
default void onDelivered(String target, TurnToken token) {
default void onDelivered(String target) {
}
/** No-op default for callers that only need delivery, not completion signalling. */
@@ -1,4 +1,4 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import java.util.concurrent.ConcurrentHashMap;
import java.util.Set;
@@ -15,7 +15,7 @@ import java.util.Set;
* mounts the bridge MCP is never marked present — its sends stay queued until they time out, which is
* correct (it could not have replied anyway).
*/
public class MemberPresence {
public class WorkerPresence {
private final Set<String> present = ConcurrentHashMap.newKeySet();
@@ -0,0 +1,797 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.AuditLog;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.auth.Role;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.peer.PeerLauncher;
import io.modelcontextprotocol.common.McpTransportContext;
import io.modelcontextprotocol.json.McpJsonMapper;
import io.modelcontextprotocol.json.jackson3.JacksonMcpJsonMapperSupplier;
import io.modelcontextprotocol.server.McpServer;
import io.modelcontextprotocol.server.McpSyncServer;
import io.modelcontextprotocol.server.McpSyncServerExchange;
import io.modelcontextprotocol.server.transport.HttpServletStreamableServerTransportProvider;
import io.modelcontextprotocol.spec.McpSchema;
import com.fasterxml.jackson.databind.ObjectMapper;
import jakarta.servlet.http.HttpServlet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.function.Function;
import java.util.stream.Collectors;
/**
* The MCP SERVER face (CB-105): a Streamable-HTTP MCP server whose tools are <em>thin adapters</em>
* over the same {@link MessageService}/{@link Rendezvous} the REST routes use — so the two are
* validated by parity, not by re-implementing behaviour. The primary Opus calls {@code bridge_send}
* / {@code bridge_status}; the worker calls {@code bridge_reply}.
*
* <p>Beyond delegation the primary also manages the fleet here (CB-108): {@code bridge_spawn} /
* {@code bridge_list} / {@code bridge_stop} drive the {@link PeerLauncher} SPI so a worker's whole
* lifecycle is managed through MCP, with each adapter's subscription boundary enforced inside it.
*
* <p>The tool <em>logic</em> lives in package-private static methods returning a
* {@link McpSchema.CallToolResult}, so it is unit-testable without standing up the HTTP transport;
* the SDK owns the wire protocol. Mount {@link #servlet()} at {@code /mcp} on the daemon's Jetty.
*/
public final class BridgeMcp {
private static final long DEFAULT_TIMEOUT_MS = 25_000;
private static final long MAX_TIMEOUT_MS = 120_000;
// bridge_ask blocks the WORKER's own MCP call, which its client caps near 60s — default under
// that so the bridge returns a clean timeout before the client severs the call (CB-205).
private static final long ASK_DEFAULT_TIMEOUT_MS = 55_000;
private static final long ASK_MAX_TIMEOUT_MS = 115_000;
private static final ObjectMapper MAPPER = new ObjectMapper(); // worker-view JSON projections
/** Transport-context key under which the extractor stashes the resolved caller identity. */
static final String CALLER_TERMINAL = "callerTerminal";
/** Transport-context key under which the extractor stashes the caller's PID (for cwd inherit). */
static final String CALLER_PID = "callerPid";
/** Transport-context key under which the extractor stashes the resolved {@link Role} (CB-501). */
static final String CALLER_ROLE = "callerRole";
private final HttpServletStreamableServerTransportProvider transport;
private final McpSyncServer server;
private final CallerResolver authz; // CB-501: null → authorization not enforced (legacy)
private final Metrics metrics; // CB-502: null → auth failures not counted
/**
* Legacy constructor — no authorization. Retained so existing tests exercise tool behaviour
* without an auth fixture.
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
PrimaryRegistry primaryRegistry) {
this(messages, workers, sessions, identity, presence, primaryRegistry, null, null);
}
/**
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
* Jetty's context handler and never passes through Javalin's {@code before}
* filter, so the REST guard does not cover it.
* @param metrics registry for auth-failure counting; may be {@code null}
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
PrimaryRegistry primaryRegistry, CallerResolver callers, Metrics metrics) {
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
.jsonMapper(json)
.mcpEndpoint("/mcp")
// Resolve the caller from the connection (peer PID → herdr pane) in one lookup: the
// worker terminal for bridge_reply (no spoofable arg), and the PID so bridge_spawn can
// inherit the primary's cwd (CB-112). Any contact from a worker marks it available
// (CB-113) — its MCP initialize is the reliable "the agent is up" signal.
.contextExtractor(req -> {
// One resolution per call, shared with the REST surface via CallerResolver so
// the two paths cannot drift on who a caller is.
Principal p = callers != null
? callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization"))
: legacyPrincipal(identity, req.getRemoteAddr(), req.getRemotePort());
presence.markPresent(p.terminal()); // no-op for the primary (null terminal)
return McpTransportContext.create(Map.of(
CALLER_TERMINAL, orEmpty(p.terminal()),
CALLER_PID, Long.toString(p.pid()),
CALLER_ROLE, p.role().name()));
})
.build();
this.server = McpServer.sync(transport)
.serverInfo("bridge", "0.1.0")
.capabilities(McpSchema.ServerCapabilities.builder().tools(true).build())
.toolCall(sendTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SEND,
str(req.arguments(), "sessionId"));
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
Map<String, Object> a = req.arguments();
String turnId = str(a, "turnId");
if (turnId != null && !turnId.isBlank()) {
// Answering a worker's bridge_ask (CB-205): resolve its blocked question and
// block for the worker's reply as it resumes the same turn.
return answer(messages, turnId, str(a, "content"), timeoutMs(a));
}
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, str(a, "sessionId"), str(a, "content"))
: send(messages, str(a, "sessionId"), str(a, "content"), timeoutMs(a));
})
// bridge_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
.toolCall(replyTool(), (exchange, req) -> {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.REPLY, self);
if (denied != null) return denied;
return reply(messages, self, str(req.arguments(), "content"));
})
// bridge_ask (CB-205): a worker's mid-turn question — identity from the CONNECTION.
.toolCall(askTool(), (exchange, req) -> {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.ASK, self);
if (denied != null) return denied;
return ask(messages, self, str(req.arguments(), "question"), timeoutMs(req.arguments()));
})
.toolCall(statusTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return status(messages, str(req.arguments(), "sessionId"));
})
.toolCall(pollTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
Map<String, Object> a = req.arguments();
return poll(messages, str(a, "ticket"), str(a, "target"));
})
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
// Acking removes a reply from the inbox, so it is a drain, not a read.
.toolCall(ackTool(), (exchange, req) -> {
Map<String, Object> a = req.arguments();
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.DRAIN, str(a, "target"));
if (denied != null) return denied;
return ack(messages, str(a, "target"), str(a, "msgId"));
})
// Fleet management (CB-108): spawn/list/stop over ClaudeCodeLauncher.
.toolCall(spawnTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SPAWN, null);
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
Map<String, Object> a = req.arguments();
// CB-112: worker inherits the primary's cwd unless the call pins one.
// CB-301: carry the caller's identity as the session owner (null for the primary).
// CB-301-ext: optional isolated worktree for parallel implementers.
String callerCwd = identity.cwdForPid(callerPid(exchange));
return spawn(sessions, str(a, "profile"), str(a, "cwd"), callerCwd,
callerTerminal(exchange), worktreeRequest(a));
})
.toolCall(listTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listWorkers(workers, sessions);
})
.toolCall(stopTool(), (exchange, req) -> {
String paneId = str(req.arguments(), "paneId");
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.STOP, paneId);
if (denied != null) return denied;
return stop(sessions, paneId);
})
.toolCall(profilesTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return profiles(workers);
})
.toolCall(whoamiTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return whoami(principal(exchange), sessions);
})
.build();
this.authz = callers;
this.metrics = metrics;
}
/**
* Pre-CB-501 identity: worker if the connection maps to a pane, otherwise the primary. Used
* only by the legacy constructor, where authorization is not enforced anyway.
*/
private static Principal legacyPrincipal(ConnectionIdentity identity, String addr, int port) {
ConnectionIdentity.Caller c = identity.resolve(addr, port);
return c.terminal() != null
? Principal.worker(c.terminal(), c.pid())
: Principal.primary(c.pid());
}
/** The caller reconstructed from the transport context. */
private static Principal principal(McpSyncServerExchange exchange) {
return principalFrom(exchange.transportContext().get(CALLER_ROLE),
callerTerminal(exchange), callerPid(exchange));
}
/**
* Rebuild a {@link Principal} from the three values the context extractor stashed.
*
* <p>Split out from {@link #principal(McpSyncServerExchange)} so the identity rules are
* reachable without an {@code McpSyncServerExchange} — that is an SDK type this project has no
* mocking library to fabricate, which is why this logic had no test at all until CB-513.
*
* @param role the stashed {@link Role} name, or {@code null} on the legacy path
* @param terminal the worker terminal, or {@code null} for a non-worker
* @param pid the calling pid, or {@code -1}
*/
static Principal principalFrom(Object role, String terminal, long pid) {
if (role == null) {
// No role stashed (legacy path): fall back to the historical interpretation.
return terminal != null ? Principal.worker(terminal, pid) : Principal.primary(pid);
}
return new Principal(Role.valueOf(role.toString()), terminal, pid);
}
/**
* Gate a tool call on the CB-505 table. Returns {@code null} when the call may proceed, or the
* error result to return when it may not.
*/
private McpSchema.CallToolResult deny(McpSyncServerExchange exchange, Authz.Action action,
String target) {
return denyFor(principal(exchange), action, target);
}
/**
* The policy half of {@link #deny}: everything except pulling the caller out of the MCP
* exchange. Kept separate so the authorization decision — the actual control — is unit-testable
* without fabricating an SDK {@code McpSyncServerExchange}.
*
* <p>This surface exists because the enforcement was previously unreachable from a test: no
* test constructs a {@code BridgeMcp}, so the whole MCP-side gate ran zero times in the suite
* while the REST-side equivalent had ten tests. A security control nothing exercises is a
* claim, not a control.
*
* @return {@code null} when the call may proceed, or the error result to return when it may not
*/
McpSchema.CallToolResult denyFor(Principal caller, Authz.Action action, String target) {
// The enforcement switch lives HERE rather than in the exchange-facing wrapper: any future
// tool that calls this directly must not be able to skip the gate by accident.
if (authz == null) {
return null; // legacy constructor: authorization not enforced
}
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
return null;
}
String reason = Authz.isUnauthenticated(caller) ? "unauthenticated" : "forbidden";
AuditLog.denied(caller, action, target, reason);
if (metrics != null) {
metrics.inc(BridgedMetrics.AUTH_FAILURES, "reason", reason);
}
return error(reason + ": " + caller.describe() + " may not " + action);
}
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
private static String callerTerminal(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_TERMINAL);
String s = v == null ? null : v.toString();
return (s == null || s.isBlank()) ? null : s;
}
/** The caller's PID resolved from this call's connection, or {@code -1} if unknown. */
private static long callerPid(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_PID);
try {
return v == null ? -1 : Long.parseLong(v.toString());
} catch (NumberFormatException e) {
return -1;
}
}
private static String orEmpty(String s) {
return s == null ? "" : s;
}
/** The Streamable-HTTP servlet to mount at {@code /mcp} on the daemon's Jetty. */
public HttpServlet servlet() {
return transport;
}
/** Graceful shutdown of the MCP server. */
public void close() {
server.closeGracefully();
}
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
/** {@code bridge_send}: delegate {@code content} to a worker session and block for its reply. */
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content, Long timeoutMs) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
try {
return formatReply(messages.send(sessionId, content, timeout), timeout);
} catch (HerdrException e) {
return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
}
}
/**
* {@code bridge_send} carrying a {@code turnId}: the primary's answer to a worker's
* {@code bridge_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as
* it resumes the same turn — surfaced to the primary identically to a normal send.
*/
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs) {
if (isBlank(turnId) || isBlank(content)) {
return error("turnId and content are required to answer a worker's question");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
return formatReply(messages.answer(turnId, content, timeout), timeout);
}
/**
* {@code bridge_ask} (CB-205): a worker pauses its delegated turn to ask the primary, blocking
* until the primary answers. The worker is identified by its connection ({@code callerTerminal}),
* never an argument — a {@code null} means the caller is not a known worker.
*/
static McpSchema.CallToolResult ask(MessageService messages, String callerTerminal, String question, Long timeoutMs) {
if (callerTerminal == null) {
return error("bridge_ask is for workers only — could not identify the calling worker "
+ "from the connection");
}
if (isBlank(question)) {
return error("question is required");
}
long timeout = Math.clamp(timeoutMs == null ? ASK_DEFAULT_TIMEOUT_MS : timeoutMs, 1, ASK_MAX_TIMEOUT_MS);
MessageService.AskResult r = messages.ask(callerTerminal, question, timeout);
return switch (r.outcome()) {
case ANSWERED -> text(r.answer());
case NO_WAITER -> error("no primary is awaiting this turn — bridge_ask only works while a "
+ "bridge_send delegation is open to answer it");
case TIMED_OUT -> text("[no answer within " + timeout + "ms — the primary did not respond; "
+ "proceed on your best judgement, then call bridge_reply to end the turn]");
};
}
/** Render a {@link MessageService.Reply} as a tool result — shared by {@link #send} and {@link #answer}. */
private static McpSchema.CallToolResult formatReply(MessageService.Reply r, long timeout) {
return switch (r.outcome()) {
case REPLIED -> text(r.text());
// The worker's turn finished but it never called bridge_reply — hand back the scraped
// transcript tail, flagged so the primary knows it isn't a structured reply.
case COMPLETED_UNREPLIED -> text(
"[worker finished without a structured bridge_reply — transcript tail follows]\n" + r.text());
// The worker ran the turn then wedged (CB-109) — surface the error context.
case WORKER_FAILED -> text("[worker failed — turn ended in an unrecoverable state]\n" + r.text());
// The worker paused mid-turn to ask (CB-205) — tell the primary how to answer in-turn.
case QUESTION -> text("[question] the worker paused to ask before it can finish:\n" + r.text()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + r.turnId()
+ "\" and content set to your answer; the worker resumes the same turn.");
case STALE_TURN -> error("that question is no longer open — it timed out or was already "
+ "answered (turnId stale)");
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> text("[no reply within " + timeout + "ms — worker "
+ r.outcome().name().toLowerCase().replace("timed_out_", "") + "; retry or poll status]");
};
}
/**
* {@code bridge_send} with {@code wait:false}: delegate {@code content} and return a ticket
* immediately (fire-and-poll), so a long task isn't cut off by the caller's MCP call timeout.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
String ticket = messages.sendAsync(sessionId, content);
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
}
/** {@code bridge_poll}: check an async delegation by ticket, or drain a worker's inbox by target. */
static McpSchema.CallToolResult poll(MessageService messages, String ticket, String target) {
if (!isBlank(target)) {
var replies = messages.drainReplies(target);
if (replies.isEmpty()) {
return text("[]");
}
return text(json(replies));
}
if (isBlank(ticket)) {
return error("ticket (or target) is required");
}
MessageService.TaskView v = messages.poll(ticket);
if (v == null) {
return error("unknown ticket: " + ticket + " (never issued, or expired)");
}
return switch (v.phase()) {
case DONE -> text(v.replySource() != null && v.replySource().equals("transcript")
? "[done — worker finished without a structured bridge_reply; transcript tail follows]\n" + v.reply()
: v.reply());
case PENDING -> text("[pending — " + v.detail() + "]");
case FAILED -> text("[failed — " + v.detail() + "]");
};
}
/**
* {@code bridge_reply}: the worker returns its structured answer, resolving the awaiting send
* or — when no send is open — queueing the reply in the inbox for later drain (CB-307).
* {@code callerTerminal} is resolved from the connection (never an argument); a {@code null}
* means the caller is not a known worker (e.g. the primary called it by mistake).
*/
static McpSchema.CallToolResult reply(MessageService messages, String callerTerminal, String content) {
if (callerTerminal == null) {
return error("bridge_reply is for workers only — could not identify the calling worker "
+ "from the connection");
}
if (content == null) {
return error("content is required");
}
messages.reply(callerTerminal, content);
return text("delivered");
}
/** {@code bridge_ack}: acknowledge (remove) a specific reply from the inbox. */
static McpSchema.CallToolResult ack(MessageService messages, String target, String msgId) {
if (isBlank(target) || isBlank(msgId)) {
return error("target and msgId are required");
}
messages.ackReply(target, msgId);
return text("acknowledged " + msgId);
}
/** {@code bridge_status}: the live lifecycle status of a worker session. */
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
if (isBlank(sessionId)) {
return error("sessionId is required");
}
try {
return text(messages.status(sessionId).name().toLowerCase());
} catch (HerdrException e) {
return error("herdr error for session " + sessionId + ": " + e.getMessage());
}
}
/**
* {@code bridge_whoami}: the caller's own identity, as the daemon already resolved it.
*
* <p>Every other tool <em>consumes</em> this identity — the authorization gate, the reply
* rendezvous, the cwd inherit — but none reported it, so an agent had to infer its own role
* from side channels the daemon does not control: a charter string in its system prompt, the
* name its MCP mount happens to carry, or {@code ANTHROPIC_BASE_URL} (which Claude-model
* workers do not set). The failure mode of guessing is asymmetric and silent: a primary that
* mistakes itself for a worker is refused by {@link Authz} and learns immediately, while a
* worker that mistakes itself for the primary ends its turn without {@code bridge_reply} and
* the sender simply receives nothing. This tool removes the guess.
*
* <p>For a worker the session registry adds what it knows about that session. A worker the
* registry has no record of — one that outlived a daemon restart — still gets its role and
* {@code sessionId}, which is the load-bearing part.
*/
static McpSchema.CallToolResult whoami(Principal caller, SessionManager sessions) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("role", caller.role().name().toLowerCase());
if (!caller.isWorker()) {
return text(json(m));
}
m.put("sessionId", caller.terminal());
sessions.roster().stream()
.filter(s -> caller.terminal().equals(s.terminalId()))
.findFirst()
.ifPresent(s -> {
m.put("paneId", s.paneId());
m.put("profile", s.profile());
m.put("state", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
if (s.ownerTerminal() != null) {
m.put("owner", s.ownerTerminal());
}
});
return text(json(m));
}
// --- fleet management logic (CB-108 / CB-301) --------------------------------------------
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
return spawn(sessions, profile, null, null, null, null);
}
/**
* {@code bridge_spawn}: launch a guard-checked worker for {@code profile} (blank → the default
* profile) and return its session id + pane id. The worker's cwd is {@code requestedCwd} if given,
* else the profile's config, else {@code callerCwd} (the primary's directory), else the daemon's.
* CB-301: the session is registered with {@code ownerTerminal} as its owner.
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
*/
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile,
String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest worktreeRequest) {
try {
WorkerSession worker = sessions.acquire(isBlank(profile) ? null : profile,
requestedCwd, callerCwd, ownerTerminal, worktreeRequest);
return text(json(workerView(worker)));
} catch (GuardException e) {
return error("subscription boundary: " + e.getMessage());
} catch (IllegalArgumentException e) {
return error(e.getMessage()); // unknown / no-default profile
} catch (PeerUnreachableException e) {
return error("spawn timed out — worker pane never reached injectable state: " + e.getMessage());
} catch (HerdrException e) {
return error("herdr error spawning worker: " + e.getMessage());
}
}
/** Build a {@link WorktreeRequest} from {@code bridge_spawn}'s optional {@code worktree}/{@code ticket} args. */
private static WorktreeRequest worktreeRequest(Map<String, Object> a) {
Object w = a.get("worktree");
if (w == null || Boolean.FALSE.equals(w)) {
return null;
}
String ticket = str(a, "ticket");
if (w instanceof String s) {
if (s.isBlank() || "false".equalsIgnoreCase(s)) {
return null;
}
if ("true".equalsIgnoreCase(s)) {
if (isBlank(ticket)) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return new WorktreeRequest(s, null);
}
if (w instanceof Boolean b && b) {
if (isBlank(ticket)) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return null;
}
/** {@code bridge_profiles}: the configured worker profiles and the default. */
static McpSchema.CallToolResult profiles(PeerLauncher workers) {
return text(json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
}
/**
* {@code bridge_list}: bridge-owned roster merged with live herdr status. CB-519 decoupled the
* registry key (a host-unique id) from the herdr pane coordinate, so the join is on the
* terminal id, which both the session and the live agent carry.
*/
static McpSchema.CallToolResult listWorkers(PeerLauncher workers, SessionManager sessions) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
return text(json(Map.of("workers", out)));
} catch (HerdrException e) {
return error("herdr error listing workers: " + e.getMessage());
}
}
/** {@code bridge_stop}: tear a worker down by its pane id. */
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
if (isBlank(paneId)) {
return error("paneId is required");
}
try {
sessions.release(paneId);
return text("stopped " + paneId);
} catch (HerdrException e) {
return error("herdr error stopping " + paneId + ": " + e.getMessage());
}
}
/** CB-301 projection from the authoritative session registry. */
private static Map<String, Object> workerView(WorkerSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", s.terminalId());
m.put("paneId", s.paneId());
m.put("status", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
return m;
}
private static String json(Object o) {
try {
return MAPPER.writeValueAsString(o);
} catch (Exception e) {
return String.valueOf(o);
}
}
// --- tool schemas --------------------------------------------------------------------------
private static McpSchema.Tool sendTool() {
return tool("bridge_send",
"Delegate a task to a worker session. By default blocks until the worker replies and "
+ "returns its reply (or a 'still working / queued' note on timeout). Pass wait:false "
+ "for a long task to return a ticket immediately, then poll it with bridge_poll. To "
+ "answer a worker's bridge_ask, pass its turnId (with content) instead of sessionId.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id (herdr terminal_id) to delegate to"),
"content", stringProp("The task/message to send to the worker (or your answer, with turnId)"),
"timeoutMs", Map.of("type", "integer", "description", "Max ms to wait for a reply (blocking mode)"),
"wait", Map.of("type", "boolean",
"description", "Block for the reply (default true); false returns a ticket to poll"),
"turnId", stringProp("When answering a worker's bridge_ask, its question turnId — "
+ "routes your answer back into the same turn (omit for a normal delegation)")),
List.of("content")));
}
private static McpSchema.Tool askTool() {
// No target/session arg — the worker's identity is resolved from the connection.
return tool("bridge_ask",
"Pause your current delegated turn to ask the primary a question, blocking until it "
+ "answers — then resume the same turn with the answer. Use this when only the "
+ "primary has a decision or detail you need to continue. You do not address the "
+ "primary; identity is resolved from your connection.",
objectSchema(Map.of(
"question", stringProp("The question to put to the primary"),
"timeoutMs", Map.of("type", "integer",
"description", "Max ms to wait for the primary's answer")),
List.of("question")));
}
private static McpSchema.Tool pollTool() {
return tool("bridge_poll",
"Check an async delegation (a bridge_send with wait:false) by its ticket: "
+ "pending, done (with the worker's reply), or failed. When target (a worker "
+ "session id) is present instead of ticket, drain that worker's inbox of "
+ "replies delivered when no send was open.",
objectSchema(Map.of(
"ticket", stringProp("The ticket returned by bridge_send wait:false"),
"target", stringProp("Worker session id to drain pending replies from (optional)")),
List.of()));
}
private static McpSchema.Tool ackTool() {
return tool("bridge_ack",
"Acknowledge (remove) a specific reply from a worker's inbox. Use when the primary "
+ "has processed a reply and wants to confirm it, leaving other pending replies "
+ "in the inbox for later drain.",
objectSchema(Map.of(
"target", stringProp("Worker session id whose inbox to ack from"),
"msgId", stringProp("The message id to acknowledge")),
List.of("target", "msgId")));
}
private static McpSchema.Tool spawnTool() {
return tool("bridge_spawn",
"Spawn a new off-subscription worker session. Pass a profile (from bridge_profiles) to "
+ "pick the backend, or omit it for the default. The worker opens your current "
+ "directory by default; pass cwd to pin a different one. Pass worktree:true (with "
+ "ticket) or worktree:<ticket-slug> to provision an isolated git worktree. "
+ "Returns the worker's sessionId (use with bridge_send) and paneId (use with bridge_stop).",
objectSchema(Map.of(
"profile", stringProp("Worker profile to spawn (omit for the default profile)"),
"cwd", stringProp("Working directory for the worker (omit to inherit yours)"),
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
"ticket", stringProp("Ticket slug when worktree:true")),
List.of()));
}
private static McpSchema.Tool profilesTool() {
return tool("bridge_profiles",
"List the configured worker profiles (backends) and which one bridge_spawn uses by default.",
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool listTool() {
return tool("bridge_list",
"List the worker sessions the bridge tracks — each with its sessionId, paneId, profile, "
+ "state, optional worktree/branch/owner, and live herdr status.",
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool stopTool() {
return tool("bridge_stop",
"Tear down a worker session by its paneId (from bridge_spawn or bridge_list).",
objectSchema(Map.of(
"paneId", stringProp("The worker's paneId to stop")),
List.of("paneId")));
}
private static McpSchema.Tool replyTool() {
// No session/target arg — the worker's identity is resolved from the connection.
return tool("bridge_reply",
"Return your structured answer for the task you were delegated, "
+ "resolving the caller's blocked bridge_send.",
objectSchema(Map.of(
"content", stringProp("Your reply/answer")),
List.of("content")));
}
private static McpSchema.Tool statusTool() {
return tool("bridge_status",
"Get the live lifecycle status (idle/working/blocked/unknown) of a worker session.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id to query")),
List.of("sessionId")));
}
private static McpSchema.Tool whoamiTool() {
return tool("bridge_whoami",
"Report who YOU are on the bridge — your role is resolved from your connection "
+ "(unforgeable), never from anything you claim. Returns role 'primary' (you "
+ "orchestrate: spawn/send/stop, and you must never call bridge_reply) or "
+ "'worker' (you were delegated to: you must end every turn with exactly one "
+ "bridge_reply, and cannot spawn or send), plus your own sessionId, profile, "
+ "worktree and branch when you are a worker. Call this first when following "
+ "role-conditional instructions rather than guessing your role.",
objectSchema(Map.of(), List.of()));
}
// --- small helpers -------------------------------------------------------------------------
// The SDK 2.0.0 deprecates its own Tool builders without a stable replacement — isolate it here.
@SuppressWarnings("deprecation")
private static McpSchema.Tool tool(String name, String description, Map<String, Object> inputSchema) {
return McpSchema.Tool.builder(name).description(description).inputSchema(inputSchema).build();
}
private static Map<String, Object> objectSchema(Map<String, Object> properties, List<String> required) {
return Map.of("type", "object", "properties", properties, "required", required);
}
private static Map<String, Object> stringProp(String description) {
return Map.of("type", "string", "description", description);
}
private static McpSchema.CallToolResult text(String s) {
return McpSchema.CallToolResult.builder().addTextContent(s == null ? "" : s).build();
}
private static McpSchema.CallToolResult error(String s) {
return McpSchema.CallToolResult.builder().addTextContent(s).isError(true).build();
}
private static String str(Map<String, Object> args, String key) {
Object v = args.get(key);
return v == null ? null : v.toString();
}
private static Long timeoutMs(Map<String, Object> args) {
Object v = args.get("timeoutMs");
return v instanceof Number n ? n.longValue() : null;
}
private static long clamp(long ms) {
return Math.clamp(ms, 1, MAX_TIMEOUT_MS);
}
private static boolean isBlank(String s) {
return s == null || s.isBlank();
}
}
@@ -1,6 +1,6 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.bridged.herdr.PaneLocator;
/**
* Resolves <em>who is calling</em> an MCP tool from the connection alone — the anti-spoofing
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
/**
* Resolves the OS PID that owns a loopback TCP source port — the OS half of connection-based MCP
@@ -0,0 +1,68 @@
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Optional;
import java.util.concurrent.atomic.AtomicReference;
/**
* Single-slot, thread-safe registry for the primary's herdr {@code terminal_id}.
*
* <p>Populated from the caller terminal of orchestration-side MCP tools
* ({@code bridge_send}, {@code bridge_spawn}) — tools that only the primary calls.
* A pinned terminal (from config) seeds the registry at construction and makes
* subsequent {@link #record(String)} calls no-ops.
*
* <p>The push loop ({@code ReplyPushLoop}) uses {@link #isKnown()} to decide
* whether active nudging is possible; an empty registry means the primary is
* off-host or non-herdr and delivery falls back to pull.
*/
public final class PrimaryRegistry {
private static final Logger log = LoggerFactory.getLogger(PrimaryRegistry.class);
private final AtomicReference<String> terminal = new AtomicReference<>();
private final boolean pinned;
/**
* @param pinnedTerminal an optional pinned terminal from config ({@code null}/blank = unpinned)
*/
public PrimaryRegistry(String pinnedTerminal) {
if (pinnedTerminal != null && !pinnedTerminal.isBlank()) {
this.terminal.set(pinnedTerminal);
this.pinned = true;
log.info("primary terminal pinned: {}", pinnedTerminal);
} else {
this.pinned = false;
}
}
/**
* Record a terminal_id. No-op when:
* <ul>
* <li>the registry is pinned (config override),
* <li>{@code terminalId} is {@code null} or blank (non-herdr caller).
* </ul>
*/
public void record(String terminalId) {
if (pinned) return;
if (terminalId == null || terminalId.isBlank()) return;
String prev = terminal.getAndSet(terminalId);
if (prev == null) {
log.debug("primary terminal learned: {}", terminalId);
} else if (!prev.equals(terminalId)) {
log.debug("primary terminal changed: {} -> {}", prev, terminalId);
}
}
/** The known primary terminal, or empty if not yet learned (and not pinned). */
public Optional<String> primaryTerminal() {
return Optional.ofNullable(terminal.get());
}
/** {@code true} once a terminal has been recorded (or was pinned at construction). */
public boolean isKnown() {
return terminal.get() != null;
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
/**
* Resolves a process's current working directory from its PID — the OS half of CB-112's
@@ -1,8 +1,8 @@
package dev.ltms.fleet.metrics;
package dev.ltms.bridged.metrics;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import java.util.LinkedHashMap;
import java.util.Map;
@@ -13,33 +13,31 @@ import java.util.Map;
*
* <p>The set is deliberately small: each series maps to a failure mode this project has actually
* hit, not to whatever was easy to count. The two worth watching in practice are
* {@code fleet_sends_total{outcome="completion_fallback"}} — a rising share means turn detection
* {@code bridged_sends_total{outcome="completion_fallback"}} — a rising share means turn detection
* is degrading, the CB-115/116/118 failure family — and
* {@code fleet_push_nudges_total{outcome="exhausted"}}, which means the primary stopped draining
* {@code bridged_push_nudges_total{outcome="exhausted"}}, which means the primary stopped draining
* its inbox and CB-307's active push gave up.
*/
public final class FleetMetrics {
public final class BridgedMetrics {
/** Counter: delegated sends by terminal outcome. */
public static final String SENDS = "fleet_sends_total";
public static final String SENDS = "bridged_sends_total";
/** Counter: worker replies by the path that carried them (rendezvous vs stranded-to-inbox). */
public static final String REPLIES = "fleet_replies_total";
public static final String REPLIES = "bridged_replies_total";
/** Counter: push-loop nudges to the primary, by outcome. */
public static final String PUSH_NUDGES = "fleet_push_nudges_total";
/** Counter: idle-lead heartbeat nudges to the lead, by outcome (CB-551). */
public static final String HEARTBEAT_NUDGES = "fleet_lead_heartbeat_nudges_total";
public static final String PUSH_NUDGES = "bridged_push_nudges_total";
/** Counter: spawn attempts by peer kind and outcome. */
public static final String SPAWNS = "fleet_spawns_total";
public static final String SPAWNS = "bridged_spawns_total";
/** Counter: herdr socket calls by method and outcome. */
public static final String HERDR_CALLS = "fleet_herdr_calls_total";
public static final String HERDR_CALLS = "bridged_herdr_calls_total";
/** Counter: rejected requests by reason (CB-501). */
public static final String AUTH_FAILURES = "fleet_auth_failures_total";
public static final String AUTH_FAILURES = "bridged_auth_failures_total";
/** Gauge: session census by lifecycle state. */
public static final String SESSIONS = "fleet_sessions";
public static final String SESSIONS = "bridged_sessions";
/** Gauge: undrained replies held per target. */
public static final String INBOX_DEPTH = "fleet_inbox_depth";
public static final String INBOX_DEPTH = "bridged_inbox_depth";
private FleetMetrics() {
private BridgedMetrics() {
}
/**
@@ -57,9 +55,6 @@ public final class FleetMetrics {
"Worker replies by delivery path (rendezvous=resolved an open send, inbox=stranded and held).");
m.describe(PUSH_NUDGES, "counter",
"CB-307 push-loop nudges to the primary (delivered|exhausted).");
m.describe(HEARTBEAT_NUDGES, "counter",
"CB-551 idle-lead heartbeat nudges (delivered|failed|exhausted). Quiet-cap exhaustion "
+ "means the lead idled with nothing pending and was told to stand down.");
m.describe(SPAWNS, "counter",
"Worker spawn attempts by peer kind and outcome (ready|timeout|guard_rejected).");
m.describe(HERDR_CALLS, "counter",
@@ -74,7 +69,7 @@ public final class FleetMetrics {
// One gauge per state so a scrape shows the whole census even when a state is empty —
// an absent series and a zero series read very differently on a dashboard.
for (MemberSession.State state : MemberSession.State.values()) {
for (WorkerSession.State state : WorkerSession.State.values()) {
String label = state.name().toLowerCase();
m.gauge(SESSIONS, () -> countIn(sessions, state), "state", label);
}
@@ -83,7 +78,7 @@ public final class FleetMetrics {
// port's non-destructive read — scraping metrics must never ack a reply out of the inbox.
m.collector(INBOX_DEPTH, "target", () -> {
Map<String, Number> depths = new LinkedHashMap<>();
for (MemberSession s : sessions.roster()) {
for (WorkerSession s : sessions.roster()) {
String target = s.terminalId();
if (target == null) {
continue;
@@ -95,7 +90,7 @@ public final class FleetMetrics {
return m;
}
private static long countIn(SessionManager sessions, MemberSession.State state) {
private static long countIn(SessionManager sessions, WorkerSession.State state) {
return sessions.roster().stream().filter(s -> s.state() == state).count();
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.metrics;
package dev.ltms.bridged.metrics;
import java.util.Map;
import java.util.NavigableMap;
@@ -0,0 +1,242 @@
package dev.ltms.bridged.msg;
import com.rabbitmq.client.AMQP;
import com.rabbitmq.client.Channel;
import com.rabbitmq.client.Connection;
import com.rabbitmq.client.ConnectionFactory;
import com.rabbitmq.client.DeliverCallback;
import com.rabbitmq.client.Recoverable;
import com.rabbitmq.client.RecoveryListener;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.concurrent.ConcurrentHashMap;
/**
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
* port {@link InMemoryReplyInbox} implements as soft state.
*
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target has a durable
* queue {@code agent.<target>.inbox}. The gateway that owns the target starts a manual-ack consumer
* ({@link #own}) that pulls persistent messages off that queue into an in-memory <em>held</em> map
* (keyed by {@code msgId}) but does <em>not</em> ack them. {@link #peek} returns that snapshot;
* {@link #ack} acks the broker delivery-tag and drops the entry. Because messages stay unacked until
* the owning gateway actually drains them, a crash (or a {@code java -jar} bounce) before caller-ack
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
* adapter cannot give, with the port contract preserved.
*
* <p><strong>Ownership is explicit.</strong> {@link #own} declares the queue and starts the consumer;
* {@link #release} cancels it. {@link #publish} sends to the queue but does <em>not</em> imply ownership
* and does not attach a consumer. This split is required by CB-308 federation, where one gateway may
* publish to an agent owned by another gateway; in that case the publisher must not compete for
* deliveries.
*
* <p><strong>Dedup.</strong> The consumer keys the held map by {@code msgId}; a redelivered duplicate
* (at-least-once, or a producer double-publish) is acked-and-dropped on arrival, so it never
* double-queues.
*
* <p><strong>Visibility.</strong> Unlike the in-memory adapter, publish → broker → consumer is
* asynchronous, so a {@link #peek} immediately after {@link #publish} may not yet see the message
* (broker delivery latency). Callers that need the reply drained poll (as the primary already does);
* the contract test waits for visibility. This is inherent to broker-backed delivery, not a defect.
*
* <p>The default deploy targets LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1 (URI-only swap),
* so the {@code @Tag("contract")} integration test runs against a RabbitMQ container.
*/
public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
private static final Logger log = LoggerFactory.getLogger(AmqpReplyInbox.class);
private static final String QUEUE_PREFIX = "agent.";
private static final String QUEUE_SUFFIX = ".inbox";
private final Connection connection;
private final Channel channel;
/** All channel operations (publish/declare/ack/cancel) serialize on this — a Channel is not thread-safe. */
private final Object channelLock = new Object();
/** target → (msgId → held delivery). Per-target map is guarded by synchronizing on itself. */
private final ConcurrentHashMap<String, LinkedHashMap<String, Held>> held = new ConcurrentHashMap<>();
/** Targets whose queue is declared and consumer is running, mapped to their broker consumer tag. */
private final ConcurrentHashMap<String, String> consumerTags = new ConcurrentHashMap<>();
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
private record Held(long deliveryTag, InboxMessage message) {}
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) and open the inbox. */
public static AmqpReplyInbox open(String uri) {
try {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"));
} catch (Exception e) {
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
}
}
/** Wrap an already-open connection (injection seam for the contract test). */
AmqpReplyInbox(Connection connection) {
this.connection = connection;
try {
this.channel = connection.createChannel();
} catch (IOException e) {
throw new IllegalStateException("cannot open AMQP channel", e);
}
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
// repopulates it with valid tags (dedup by msgId still prevents any double-queue).
if (connection instanceof Recoverable recoverable) {
recoverable.addRecoveryListener(new RecoveryListener() {
@Override
public void handleRecovery(Recoverable recoverable) {
held.clear();
log.info("AMQP connection recovered; cleared held replies for fresh redelivery");
}
@Override
public void handleRecoveryStarted(Recoverable recoverable) {
// no-op: we act once recovery completes
}
});
}
}
@Override
public void own(String target) {
synchronized (channelLock) {
if (consumerTags.containsKey(target)) {
return; // already owning this target
}
String queue = queueName(target);
try {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
String tag = channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
consumerTags.put(target, tag);
log.debug("AMQP inbox owns queue {} for target {}", queue, target);
} catch (IOException e) {
throw new IllegalStateException("cannot own queue " + queue, e);
}
}
}
@Override
public void release(String target) {
synchronized (channelLock) {
String tag = consumerTags.remove(target);
held.remove(target); // stale delivery tags must not survive release
if (tag == null) {
return;
}
try {
channel.basicCancel(tag);
} catch (IOException e) {
throw new IllegalStateException("cannot cancel consumer for " + target, e);
}
}
}
@Override
public void publish(String target, String msgId, String content) {
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
.messageId(msgId)
.deliveryMode(2) // persistent — survives a broker restart
.contentType("text/plain")
.build();
try {
synchronized (channelLock) {
channel.basicPublish("", queueName(target), props, content.getBytes(StandardCharsets.UTF_8));
}
} catch (IOException e) {
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
}
}
@Override
public List<InboxMessage> peek(String target) {
var perTarget = held.get(target);
if (perTarget == null) {
return List.of();
}
synchronized (perTarget) {
return perTarget.values().stream().map(Held::message).toList();
}
}
@Override
public void ack(String target, String msgId) {
var perTarget = held.get(target);
if (perTarget == null) {
return;
}
Held h;
synchronized (perTarget) {
h = perTarget.remove(msgId);
}
if (h == null) {
return; // never held (or already acked) — no-op
}
try {
synchronized (channelLock) {
channel.basicAck(h.deliveryTag(), false);
}
} catch (IOException e) {
// Ack didn't reach the broker: restore the entry so a later ack (or a redelivery after
// reconnect) can retry. Keeps the at-least-once contract — a reply is never silently lost.
synchronized (perTarget) {
perTarget.putIfAbsent(msgId, h);
}
throw new IllegalStateException("cannot ack reply " + msgId + " on " + queueName(target), e);
}
}
private DeliverCallback deliverCallback(String target) {
return (_, delivery) -> {
String msgId = delivery.getProperties().getMessageId();
long tag = delivery.getEnvelope().getDeliveryTag();
if (msgId == null || msgId.isBlank()) {
msgId = Long.toHexString(tag); // synthesize an id so dedup still has a key
}
String content = new String(delivery.getBody(), StandardCharsets.UTF_8);
var perTarget = held.computeIfAbsent(target, _ -> new LinkedHashMap<>());
boolean duplicate;
synchronized (perTarget) {
if (perTarget.containsKey(msgId)) {
duplicate = true;
} else {
perTarget.put(msgId, new Held(tag, new InboxMessage(msgId, target, content)));
duplicate = false;
}
}
if (duplicate) {
// Redelivered duplicate: ack the new tag and drop it so the broker stops resending.
synchronized (channelLock) {
channel.basicAck(tag, false);
}
}
};
}
private static String queueName(String target) {
return QUEUE_PREFIX + target + QUEUE_SUFFIX;
}
@Override
public void close() {
try {
channel.close();
} catch (Exception e) {
log.debug("AMQP channel close: {}", e.toString());
}
try {
connection.close();
} catch (Exception e) {
log.debug("AMQP connection close: {}", e.toString());
}
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import java.util.LinkedHashMap;
import java.util.List;
@@ -0,0 +1,527 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.UUID;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.CompletionException;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
import java.util.concurrent.atomic.AtomicLong;
import java.util.concurrent.locks.ReentrantLock;
/**
* The blocking delegation feature (CB-104): deliver {@code content} into a worker and block until
* the worker returns a <em>structured reply</em> via {@code bridge_reply} (the {@link Rendezvous}),
* then hand that reply back. Delivery is the {@link Injector}'s job (the background poller sends it
* when the worker is injectable); this service never drives the injector or scrapes the terminal —
* completion is the worker's explicit reply, not a guess about {@code agent_status}.
*
* <p>Sends are serialized per session so exactly one reply can be outstanding per worker, which is
* what lets a reply map unambiguously to its send (no cross-talk between concurrent callers).
*
* <p>If the worker never replies within the timeout, the caller gets a typed "still working" /
* "queued" outcome — the message may still be mid-flight. A finished-but-unreplied turn is caught
* by the CB-106 completion fallback (see {@link Rendezvous#resolveCompletion}).
*
* <p><strong>Async fire-and-poll (CB-107).</strong> A caller's MCP client caps a blocking call at
* ~60s, but a real delegated task runs for minutes. {@link #sendAsync} therefore runs the same
* blocking {@link #send} on a background virtual thread and hands back a <em>ticket</em> the caller
* polls with {@link #poll}. The blocking and async paths share one code path (and the same per-target
* serialization), so async inherits the reply + completion resolution behaviour for free.
*/
public final class MessageService {
private static final Logger log = LoggerFactory.getLogger(MessageService.class);
/**
* The window a fire-and-poll send waits for resolution — generous, since no caller is blocked on
* it; a real delegated task resolves (reply or completion) well within this, and only a genuinely
* hung worker rides it out.
*/
private static final long ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L;
/** How long a finished (terminal) ticket is retained for polling before it is pruned. */
private static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
/** Outcome of a blocking send. */
public enum Outcome {
/** The worker called {@code bridge_reply}; {@code text} holds the structured answer. */
REPLIED,
/**
* The worker's delegated turn finished without a {@code bridge_reply} (CB-106 fallback);
* {@code text} is the scraped transcript tail rather than a structured answer.
*/
COMPLETED_UNREPLIED,
/**
* The worker ran the turn then wedged in an unrecoverable state (CB-109); {@code text} is the
* failure context (e.g. the error screen). Terminal, but not a successful completion.
*/
WORKER_FAILED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with
* {@link #answer(String, String, long)} and the turn resumes.
*/
QUESTION,
/** Timed out after the message was delivered — the worker is still working. */
TIMED_OUT_WORKING,
/** Timed out before delivery — the message is still queued for the worker. */
TIMED_OUT_QUEUED,
/** Another send to this session was in flight for the whole window. */
BUSY,
/**
* An answer ({@link #answer(String, String, long)}) referenced a {@code turnId} that is no
* longer open — the worker's {@code bridge_ask} already timed out or was answered.
*/
STALE_TURN
}
/**
* @param outcome how the send ended (or paused)
* @param text the worker's answer when {@link #completed()} (a structured {@code bridge_reply}
* for {@link Outcome#REPLIED}, a scraped transcript tail for
* {@link Outcome#COMPLETED_UNREPLIED}), or the question for {@link Outcome#QUESTION},
* else {@code null}
* @param turnId correlation id for a {@link Outcome#QUESTION} (answered via
* {@link #answer(String, String, long)}), else {@code null}
*/
public record Reply(Outcome outcome, String text, String turnId) {
/** A reply with no correlation id (the common terminal outcomes). */
public Reply(Outcome outcome, String text) {
this(outcome, text, null);
}
/** Whether the worker's turn actually finished with an answer (replied or scraped). */
public boolean completed() {
return outcome == Outcome.REPLIED || outcome == Outcome.COMPLETED_UNREPLIED;
}
}
/** How a worker's {@code bridge_ask} (CB-205) resolved. */
public enum AskOutcome {
/** The primary answered; {@link AskResult#answer} carries it. */
ANSWERED,
/** No delegation was open to surface the question to — the worker has no one to ask. */
NO_WAITER,
/** The primary did not answer within the window. */
TIMED_OUT
}
/** The outcome of a worker's {@code bridge_ask}: how it resolved and (if answered) the answer. */
public record AskResult(AskOutcome outcome, String answer) {
}
/** Lifecycle phase of an async delegation ticket. */
public enum Phase {
/** Delegated and in flight — queued for the worker or being worked. */
PENDING,
/** The worker's turn finished; {@link TaskView#reply} holds the answer. */
DONE,
/** The delegation could not complete (timed out, worker gone, or busy). */
FAILED
}
/**
* A poll snapshot of an async delegation.
*
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, else {@code null}
* @param replySource {@code "reply"} (structured {@code bridge_reply}) or {@code "transcript"}
* (completion scrape) when {@link Phase#DONE}, else {@code null}
* @param detail a human note (live worker status while pending, or the failure reason)
*/
public record TaskView(String ticket, Phase phase, String reply, String replySource, String detail) {
}
/** An in-flight or finished async delegation, keyed by its ticket. */
private record Task(String target, CompletableFuture<Reply> future, long createdNanos) {
}
private final AgentControl agents;
private final Injector injector;
private final Rendezvous rendezvous;
private final ReplyInbox inbox;
private final ReplyPushLoop pushLoop;
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
Thread.ofVirtual().name("bridge-async-", 0).factory());
/**
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
*
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
* branch ({@link #reply}) so it can nudge the primary to drain the inbox
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop) {
this(agents, injector, rendezvous, inbox, pushLoop, null);
}
/**
* As above, with a metric registry (CB-502). Instrumenting here rather than at the REST and MCP
* edges means both surfaces are counted by one piece of code and cannot drift.
*
* @param metrics nullable — when null, nothing is recorded
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop, Metrics metrics) {
this.agents = agents;
this.injector = injector;
this.rendezvous = rendezvous;
this.inbox = inbox;
this.pushLoop = pushLoop;
this.metrics = metrics;
}
/** Create with an explicit {@link ReplyInbox} and no push loop. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox) {
this(agents, injector, rendezvous, inbox, null);
}
/** Backward-compatible constructor that uses a default {@link InMemoryReplyInbox}. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous) {
this(agents, injector, rendezvous, new InMemoryReplyInbox());
}
/** Current lifecycle status of a worker (the {@code GET /sessions/{id}/status} surface). */
public AgentStatus status(String target) {
return agents.status(target);
}
/**
* Route a worker's explicit {@code bridge_reply}: resolve an open send, or queue it in the
* inbox if no send is currently open. Unlike the bare {@link Rendezvous#resolve}, a no-waiter
* result is <em>not</em> a failure — the reply is held for later drain.
*
* <p><strong>Do NOT use this for mid-turn questions.</strong> {@code bridge_ask} /
* {@link Rendezvous#resolveQuestion} must keep today's {@code NO_WAITER} behaviour — questions
* are interactive and must never be queued.
*
* @return always {@code true} — the reply either resolved a live send or was queued
*/
public boolean reply(String session, String content) {
if (rendezvous.resolve(session, content)) {
count(BridgedMetrics.REPLIES, "path", "rendezvous");
return true; // a live send took it — unchanged fast path
}
inbox.publish(session, UUID.randomUUID().toString(), content);
// A rising inbox share is the signal CB-307 exists to make visible: the worker finished but
// nobody was waiting, so delivery now depends on the push loop and a drain.
count(BridgedMetrics.REPLIES, "path", "inbox");
if (pushLoop != null) {
pushLoop.onReplyQueued(session);
}
return true; // held, not lost
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
if (metrics != null) {
metrics.inc(name, labels);
}
}
/** Count a send's terminal outcome and pass the reply through unchanged. */
private Reply recorded(Reply r) {
String label = sendOutcomeLabel(r.outcome());
if (label != null) {
count(BridgedMetrics.SENDS, "outcome", label);
}
return r;
}
/** Map a terminal send outcome to its metric label, or {@code null} for non-terminal ones. */
private static String sendOutcomeLabel(Outcome o) {
return switch (o) {
case REPLIED -> "replied";
case COMPLETED_UNREPLIED -> "completion_fallback";
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> "timeout";
case WORKER_FAILED -> "failed";
case STALE_TURN, QUESTION -> null; // not a completed delegation
};
}
/**
* Abandon any send still waiting on {@code target} because its session has gone away (CB-516).
*
* <p>Without this, tearing a worker down left its rendezvous waiter open: a blocking
* {@code bridge_send} kept blocking, and an async one kept reporting {@code PENDING} until
* {@link #ASYNC_TIMEOUT_MS} — thirty minutes — even though the worker provably no longer
* existed and the delegation could never complete. Worse, {@code poll} already had the evidence
* (it calls {@code liveStatus} to build its detail string and gets back {@code "unknown"}) and
* reported {@code PENDING} anyway.
*
* <p>Resolving the waiter as a failure — rather than letting it time out — also means the
* outcome is counted, so a torn-down delegation stops being invisible to {@code /metrics}.
*
* @return true if a live waiter was failed
*/
public boolean abandon(String target, String reason) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
return false; // nobody is blocked on this worker — nothing to abandon
}
boolean failed = rendezvous.resolveFailure(waiter, reason);
if (failed) {
log.debug("abandoned send to {}: {}", target, reason);
}
return failed;
}
/**
* Acknowledge a specific reply by {@code msgId} for {@code target}. Removes it from the inbox
* so that a subsequent drain or peek no longer returns it.
*/
public void ackReply(String target, String msgId) {
inbox.ack(target, msgId);
}
/**
* Drain (peek + ack) all pending inbox replies for {@code target}. At-least-once: returns the
* messages and acknowledges them; an in-flight failure between returning and the caller
* processing them re-surfaces them on a subsequent drain (the ack is local).
*
* @return the drained messages, newest last (FIFO); empty list if none
*/
public List<ReplyInbox.InboxMessage> drainReplies(String target) {
var messages = inbox.peek(target);
for (var msg : messages) {
inbox.ack(target, msg.msgId());
}
return messages;
}
/**
* Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
*/
public Reply send(String target, String content, long timeoutMillis) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
}
try {
CompletableFuture<Void> delivered = injector.enqueue(target, content);
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
} finally {
rendezvous.close(target, reply);
}
} finally {
lock.unlock();
}
}
/**
* A worker's mid-turn question (CB-205 reverse rendezvous): surface {@code question} to the
* primary by resolving its open blocking {@code bridge_send}, then block this (worker) call until
* the primary answers via {@link #answer} or {@code timeoutMillis} elapses. Identity is the
* worker's own session — it does not address the primary.
*
* <p>Returns {@link AskOutcome#NO_WAITER} when no delegation is open to surface the question to
* (nothing to answer it), {@link AskOutcome#ANSWERED} with the primary's answer, or
* {@link AskOutcome#TIMED_OUT} if the primary stayed silent. The worker resumes its turn either
* way — an answered ask hands back the answer; an unanswered one leaves it to proceed alone.
*/
public AskResult ask(String workerSession, String question, long timeoutMillis) {
Rendezvous.AskTicket ticket = rendezvous.openAsk(workerSession);
// Only the freshly-opening caller surfaces the question; a coalesced duplicate simply blocks on
// the shared answer future that the fresh owner is already responsible for.
if (ticket.fresh()) {
// Register the reverse waiter first, then surface the question — so the answer, which can
// arrive the instant the primary reacts, always finds an open waiter to resolve.
if (!rendezvous.resolveQuestion(workerSession, question, ticket.turnId())) {
rendezvous.closeAsk(ticket.turnId());
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
}
}
try {
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
return new AskResult(AskOutcome.ANSWERED, answer);
} catch (TimeoutException e) {
log.debug("bridge_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
return new AskResult(AskOutcome.TIMED_OUT, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting the primary's answer for " + workerSession, e);
} finally {
// Only the fresh owner tears down the shared turn; a duplicate must leave it open.
if (ticket.fresh()) {
rendezvous.closeAsk(ticket.turnId());
}
}
}
/**
* The primary's answer to a worker's {@code bridge_ask} (CB-205): resolve the worker's blocked
* question identified by {@code turnId}, then — like a fresh {@link #send} — block for the worker's
* eventual {@code bridge_reply} as it finishes the resumed turn. The worker session is derived from
* {@code turnId}, never a caller argument.
*
* <p>Unlike {@link #send} this does not re-inject through the {@link Injector}: the worker is
* mid-turn (already picked up), so the answer flows back through its own open {@code bridge_ask}
* call, not a new status-gated delivery. The forward waiter is opened <em>before</em> the worker
* is unblocked so a reply that lands the instant it resumes is not lost.
*/
public Reply answer(String turnId, String content, long timeoutMillis) {
String workerSession = rendezvous.askSession(turnId);
if (workerSession == null) {
return new Reply(Outcome.STALE_TURN, null); // the ask lapsed (timed out or already answered)
}
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(workerSession, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null);
}
try {
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession);
if (!rendezvous.answerAsk(turnId, content)) {
rendezvous.close(workerSession, reply);
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
}
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
} catch (TimeoutException e) {
// The worker resumed but hasn't replied yet — no completion fallback arms an answered
// turn (it never re-entered the injector), so a silent worker rides out the window.
return new Reply(Outcome.TIMED_OUT_WORKING, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + workerSession, e);
} finally {
rendezvous.close(workerSession, reply);
}
} finally {
lock.unlock();
}
}
/**
* Fire-and-poll variant of {@link #send}: deliver {@code content} to {@code target} on a
* background virtual thread and return immediately with a ticket to {@link #poll}. This is how a
* long task is delegated without tripping the caller's MCP client call timeout.
*
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content) {
String ticket = "task-" + ticketSeq.incrementAndGet();
CompletableFuture<Reply> future =
CompletableFuture.supplyAsync(() -> send(target, content, ASYNC_TIMEOUT_MS), asyncExecutor);
tasks.put(ticket, new Task(target, future, System.nanoTime()));
pruneTerminalTickets();
log.debug("async send {} -> {}", ticket, target);
return ticket;
}
/**
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket;
* otherwise a {@link Phase#PENDING} view (with the live worker status as detail), a
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason.
*/
public TaskView poll(String ticket) {
Task task = tasks.get(ticket);
if (task == null) {
return null;
}
CompletableFuture<Reply> f = task.future();
if (!f.isDone()) {
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target()));
}
Reply r;
try {
r = f.getNow(null);
} catch (CompletionException | java.util.concurrent.CancellationException e) {
Throwable cause = (e instanceof CompletionException ce && ce.getCause() != null) ? ce.getCause() : e;
return new TaskView(ticket, Phase.FAILED, null, null, cause.getMessage());
}
if (r.completed()) {
String source = r.outcome() == Outcome.REPLIED ? "reply" : "transcript";
return new TaskView(ticket, Phase.DONE, r.text(), source, null);
}
// A wedged worker (CB-109) carries the error context as its reason; the timeout/busy
// outcomes carry none, so fall back to the outcome name.
String detail = r.outcome() == Outcome.WORKER_FAILED && r.text() != null
? r.text()
: "no reply — " + r.outcome().name().toLowerCase();
return new TaskView(ticket, Phase.FAILED, null, null, detail);
}
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
private String liveStatus(String target) {
try {
return agents.status(target).name().toLowerCase();
} catch (RuntimeException e) {
return "unknown";
}
}
/** Drop finished tickets older than the TTL so the registry cannot grow without bound. */
private void pruneTerminalTickets() {
long cutoff = System.nanoTime() - TICKET_TTL_NANOS;
tasks.values().removeIf(t -> t.future().isDone() && t.createdNanos() < cutoff);
}
/** Release the async executor. */
public void close() {
asyncExecutor.shutdown();
}
/** Map a rendezvous {@link Rendezvous.Kind} onto its send {@link Outcome} (shared by send/answer). */
private static Outcome outcomeOf(Rendezvous.Kind kind) {
return switch (kind) {
case REPLY -> Outcome.REPLIED;
case COMPLETION -> Outcome.COMPLETED_UNREPLIED;
case FAILED -> Outcome.WORKER_FAILED;
case QUESTION -> Outcome.QUESTION;
};
}
private static boolean tryLock(ReentrantLock lock, long millis) {
try {
return lock.tryLock(Math.max(0, millis), TimeUnit.MILLISECONDS);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting the session send lock", e);
}
}
private static long remainingMillis(long deadlineNanos) {
return (deadlineNanos - System.nanoTime()) / 1_000_000L;
}
}
@@ -1,13 +1,13 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicLong;
/**
* The reply rendezvous: where a blocking {@code fleet_send} awaits how the worker's delegated turn
* The reply rendezvous: where a blocking {@code bridge_send} awaits how the worker's delegated turn
* ends. The sending (primary) request thread {@link #open}s a waiter; it is resolved either by the
* worker's explicit {@code fleet_reply} ({@link #resolve}, arriving on a different thread via
* worker's explicit {@code bridge_reply} ({@link #resolve}, arriving on a different thread via
* {@code POST /sessions/{id}/reply}) or — the CB-106 fallback — by the injector observing the
* worker's delegated turn return to idle without a reply ({@link #resolveCompletion}).
*
@@ -27,23 +27,16 @@ public final class Rendezvous {
/** How a delegated turn ended (or paused). */
public enum Kind {
/** The worker called {@code fleet_reply} with a structured answer. */
/** The worker called {@code bridge_reply} with a structured answer. */
REPLY,
/** The worker's turn finished without a {@code fleet_reply}; {@code text} is a scrape. */
/** The worker's turn finished without a {@code bridge_reply}; {@code text} is a scrape. */
COMPLETION,
/** The worker ran the turn then wedged (CB-109); {@code text} is the failure context. */
FAILED,
/**
* The turn finished without a {@code fleet_reply}, and the scrape matched the backend's
* configured usage-limit refusal pattern (CB-578 stage A); {@code text} is the reason,
* carrying the matched line. The pane is healthy — only the account is refusing — so this
* is kept separate from a session simply going {@code GONE}.
*/
BACKEND_EXHAUSTED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205 reverse rendezvous);
* {@code text} is the question and {@code turnId} correlates the primary's answer back to
* the worker's blocked {@code fleet_ask}. Not terminal — the turn resumes after the answer.
* the worker's blocked {@code bridge_ask}. Not terminal — the turn resumes after the answer.
*/
QUESTION
}
@@ -75,36 +68,21 @@ public final class Rendezvous {
/** Reverse rendezvous (CB-205): worker questions awaiting the primary's answer, keyed by {@code turnId}. */
private final ConcurrentHashMap<String, AskWaiter> asks = new ConcurrentHashMap<>();
private final AtomicLong askSeq = new AtomicLong();
/** Per-session index of the currently-open ask, so duplicate fleet_ask calls coalesce onto one turn. */
/** Per-session index of the currently-open ask, so duplicate bridge_ask calls coalesce onto one turn. */
private final ConcurrentHashMap<String, String> openAsksBySession = new ConcurrentHashMap<>();
/**
* Register a waiter for {@code session} — the await side of the public {@code resolve*} methods.
* The caller must hold that session's send lock.
*
* <p>Atomic fail-if-present (CB-548): if a waiter is already registered for {@code session}, an
* {@link IllegalStateException} is thrown rather than replacing the first — so any future
* invariant violation fails loudly instead of silently swapping the waiter another send is
* blocked on. {@code MessageService} serializes sends per session (the send lock), so in correct
* code a double open is impossible; this is a tripwire for the day that no longer holds.
*/
public CompletableFuture<Resolution> open(String session) {
CompletableFuture<Resolution> waiter = new CompletableFuture<>();
CompletableFuture<Resolution> existing = waiters.putIfAbsent(session, waiter);
if (existing != null) {
throw new IllegalStateException(
"rendezvous double-open for session " + session + " — a waiter is already registered");
}
waiters.put(session, waiter);
return waiter;
}
/**
* Remove {@code waiter} for {@code session}, only if it is still the registered one. The
* symmetric complement of {@link #open}: a terminal send deregisters its waiter so the next
* send on the session may {@link #open} a fresh one (CB-548 makes double-open an error, so a
* successful {@code open} after a finished turn requires this close to have happened first).
*/
public void close(String session, CompletableFuture<Resolution> waiter) {
/** Remove {@code waiter} for {@code session} (only if it is still the registered one). */
void close(String session, CompletableFuture<Resolution> waiter) {
waiters.remove(session, waiter);
}
@@ -132,7 +110,7 @@ public final class Rendezvous {
return complete(session, new Resolution(Kind.REPLY, content));
}
// --- reverse rendezvous (CB-205 fleet_ask) ------------------------------------------------
// --- reverse rendezvous (CB-205 bridge_ask) ------------------------------------------------
/**
* Open a reverse-rendezvous waiter for a worker's mid-turn question. If {@code session} already has
@@ -166,7 +144,7 @@ public final class Rendezvous {
}
/**
* Surface a worker's mid-turn {@code question} by resolving the primary's open {@code fleet_send}
* Surface a worker's mid-turn {@code question} by resolving the primary's open {@code bridge_send}
* with a {@link Kind#QUESTION} carrying {@code turnId}. Same session-keyed semantics as
* {@link #resolve}: the one outstanding send for {@code session} unblocks with the question.
*
@@ -184,7 +162,7 @@ public final class Rendezvous {
}
/**
* Resolve a worker's blocked {@code fleet_ask} with the primary's {@code answer}, unblocking it
* Resolve a worker's blocked {@code bridge_ask} with the primary's {@code answer}, unblocking it
* to resume its turn.
*
* @return {@code true} if the ask was still open and got the answer; {@code false} if the
@@ -195,7 +173,7 @@ public final class Rendezvous {
return w != null && w.answer().complete(answer);
}
/** Drop a reverse-rendezvous turn once its {@code fleet_ask} has resolved (answered or lapsed). */
/** Drop a reverse-rendezvous turn once its {@code bridge_ask} has resolved (answered or lapsed). */
public void closeAsk(String turnId) {
AskWaiter w = asks.get(turnId);
if (w == null) {
@@ -209,9 +187,9 @@ public final class Rendezvous {
/**
* Resolve a specific captured {@code waiter} as a completion (the delegated turn finished with no
* {@code fleet_reply}); {@code text} is the scraped transcript tail. The waiter is the one
* {@code bridge_reply}); {@code text} is the scraped transcript tail. The waiter is the one
* captured when this turn was delivered, so a late completion for turn N cannot land on turn N+1's
* send (CB-116). A no-op if that waiter was already resolved — a raced {@code fleet_reply} wins.
* send (CB-116). A no-op if that waiter was already resolved — a raced {@code bridge_reply} wins.
*
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
*/
@@ -231,19 +209,6 @@ public final class Rendezvous {
return waiter != null && waiter.complete(new Resolution(Kind.FAILED, reason));
}
/**
* Resolve a specific captured {@code waiter} as {@link Kind#BACKEND_EXHAUSTED} (CB-578 stage A):
* the turn finished with no {@code fleet_reply} and the scrape matched the backend's configured
* usage-limit pattern; {@code reason} carries the matched line. Like
* {@link #resolveCompletion(CompletableFuture, String)} it targets the exact captured send
* (CB-116). A no-op if that waiter was already resolved — first resolution wins.
*
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
*/
public boolean resolveExhausted(CompletableFuture<Resolution> waiter, String reason) {
return waiter != null && waiter.complete(new Resolution(Kind.BACKEND_EXHAUSTED, reason));
}
private boolean complete(String session, Resolution resolution) {
CompletableFuture<Resolution> waiter = waiters.get(session);
return waiter != null && waiter.complete(resolution);
@@ -1,4 +1,4 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import java.util.List;
@@ -0,0 +1,176 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
/**
* Mechanism (b) of CB-307: a dedicated, status-gated push loop that nudges the primary's own
* herdr pane when a worker reply lands with no live {@code bridge_send} to resolve it.
*
* <p>The loop is triggered by {@link #onReplyQueued(String)} (called from
* {@link MessageService#reply} after the durable inbox publish). It checks four conditions
* at each tick via {@link #decide(String, int)}, then either injects a drain nudge,
* waits for the primary to become injectable, or stops reminding.
*
* <p>Bounded: at most {@link #maxReminders} nudges per target, with a configurable backoff
* between them. The reply is never lost — the durable inbox is the backstop.
*/
public final class ReplyPushLoop {
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
private final ReplyInbox inbox;
private final ScheduledExecutorService scheduler;
private final int maxReminders;
private final long backoffMs;
private final Metrics metrics; // CB-512: nullable — no registry in unit tests
/** Track targets that have an active schedule. */
private final ConcurrentHashMap<String, Boolean> activeTargets = new ConcurrentHashMap<>();
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
int maxReminders, long backoffMs) {
this(primaryRegistry, agents, inbox, scheduler, maxReminders, backoffMs, null);
}
/** As above, with a metric registry (CB-512) so push outcomes are counted. */
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
int maxReminders, long backoffMs, Metrics metrics) {
this.primaryRegistry = primaryRegistry;
this.agents = agents;
this.inbox = inbox;
this.scheduler = scheduler;
this.maxReminders = maxReminders;
this.backoffMs = backoffMs;
this.metrics = metrics;
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
if (metrics != null) {
metrics.inc(name, labels);
}
}
// --- decision logic (package-private for unit-testing) -------------------------------------
/** The action the loop should take for a target at the given reminder count. */
enum Action { INJECT, WAIT_BUSY, STOP }
/**
* Pure decision function: examine the current state and return what the loop should do.
*
* @param target the worker session (target terminal id)
* @param reminderCount how many nudges have been sent so far for this target
* @return the action the caller should take
*/
Action decide(String target, int reminderCount) {
if (!primaryRegistry.isKnown()) {
log.debug("push: primary unknown, stopping reminder for {}", target);
return Action.STOP;
}
if (inbox.peek(target).isEmpty()) {
log.debug("push: inbox empty for {}, stopping reminder", target);
return Action.STOP;
}
if (reminderCount >= maxReminders) {
log.debug("push: reminder cap ({}) reached for {}, stopping", maxReminders, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted");
return Action.STOP;
}
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
AgentStatus status;
try {
status = agents.status(primaryTerminal);
} catch (RuntimeException e) {
log.debug("push: status check failed for primary {}, will retry", primaryTerminal, e);
return Action.WAIT_BUSY;
}
if (status.injectable()) {
return Action.INJECT;
}
log.debug("push: primary {} is {} (not injectable), waiting", primaryTerminal, status);
return Action.WAIT_BUSY;
}
// --- public entrypoint ---------------------------------------------------------------------
/**
* Called when a reply is queued for {@code target}. Idempotent per target: a second call while
* a schedule is active is a no-op. The schedule nudges the primary, then schedules a follow-up
* check (reminder on backoff, or re-check on WAIT_BUSY), until the inbox is empty or the cap
* is reached.
*/
public void onReplyQueued(String target) {
if (activeTargets.putIfAbsent(target, Boolean.TRUE) != null) {
log.debug("push: already active for {}, ignoring duplicate trigger", target);
return; // already scheduled
}
log.debug("push: starting reminder loop for {}", target);
scheduleNext(target, 0);
}
/** Execute one loop tick — called on the scheduler thread. */
private void tick(String target, int reminderCount) {
var action = decide(target, reminderCount);
switch (action) {
case INJECT -> {
injectNudge(target, reminderCount);
scheduleNext(target, reminderCount + 1);
}
// Re-check after the configured backoff; the primary may become injectable soon.
case WAIT_BUSY -> scheduleNext(target, reminderCount);
case STOP -> {
activeTargets.remove(target);
log.debug("push: reminder loop ended for {}", target);
}
}
}
/** Send the nudge and log the event. */
private void injectNudge(String target, int reminderCount) {
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
String nudge = NUDGE_FORMAT.formatted(target, target);
try {
agents.send(primaryTerminal, nudge);
log.debug("push: nudge {}/{} sent to primary {} for target {}",
reminderCount + 1, maxReminders, primaryTerminal, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered");
} catch (RuntimeException e) {
log.warn("push: failed to nudge primary {} for target {} (reminder {}/{}): {}",
primaryTerminal, target, reminderCount + 1, maxReminders, e.toString());
}
}
/** Schedule the next tick on the scheduler thread pool. */
private void scheduleNext(String target, int nextReminderCount) {
scheduler.schedule(() -> tick(target, nextReminderCount), backoffMs, TimeUnit.MILLISECONDS);
}
// --- lifecycle -----------------------------------------------------------------------------
/** Shut down the scheduler. Outstanding reminders are cancelled. */
public void stop() {
scheduler.shutdownNow();
activeTargets.clear();
}
/** @see #stop() */
public void close() {
stop();
}
}
@@ -0,0 +1,35 @@
package dev.ltms.bridged.peer;
/**
* Declared capabilities of a {@link PeerLauncher}. The protocol is the union across all
* configured launchers; a verb invoked against a peer that lacks the capability returns a clean
* "unsupported for this peer" rather than a crash. Capabilities keep the protocol honest as peers
* diversify and prevent the core from assuming "every peer is a Claude in a worktree."
*/
public enum Capability {
/**
* The peer supports {@code bridge_ask} rendezvous — pausing its delegated turn to ask
* the primary a question, then resuming once answered. All Claude Code peers support this.
*/
MID_TURN_ASK,
/**
* The peer can open its own PR at the end of an implementation turn (CB-302). Opt-in per
* profile: granted only when the profile carries a git-forge token ({@code gitTokenEnv}).
*/
SELF_PR,
/**
* The peer can run inside a provisioned isolated git worktree. All CLI-based peers support
* this since their cwd is set at spawn time.
*/
WORKTREE,
/**
* The spawner can reconcile orphaned peers on boot — workers that outlived a prior daemon
* process and whose pane ids died with it (CB-117). Claude Code over herdr supports this
* via name-based matching against the herdr agent list.
*/
ORPHAN_REAP
}
@@ -0,0 +1,44 @@
package dev.ltms.bridged.peer;
/**
* An opaque handle returned by {@link PeerLauncher#spawn(SpawnRequest)}. The core routes on
* {@link #id()} (the registry/routing key) and uses {@link #terminalId()} for session tracking;
* launcher-private coordinates beyond these are reachable through the concrete implementation.
*
* <p>A {@link PeerHandle} is returned <em>after</em> the peer process is live — the launcher
* has already completed subscription-guarded env/vfs setup, process start, and placement. The
* handle is a ticket the core exchanges for the running peer, not a lazy/delayed reference.
*/
public interface PeerHandle {
/**
* The registry/routing key — an opaque, launcher-assigned identifier (CB-519). Multiple
* daemon processes may run on one host, so the contract is <em>host-unique</em>, not merely
* process-unique: the herdr-backed launcher mints a fresh UUID per spawn, and a non-herdr
* launcher is likewise expected to return an identifier that cannot collide across processes
* on the same host. This id is the routing key and is deliberately decoupled from any launcher
* transport coordinate (e.g. a herdr pane id), which stays launcher-private. Guaranteed to be
* non-null and unique among live peers on the host.
*/
String id();
/**
* The transport-level session identifier used for message routing and presence tracking.
* For the herdr launcher this is the herdr terminal UUID. A non-herdr launcher may return
* its own analogous identifier, or {@code null} if the concept does not apply.
*/
default String terminalId() {
return null;
}
/**
* The worker profile that spawned this peer, if the launcher resolved one. A launcher that
* performs dynamic profile selection (e.g. CB-518 weighted placement) sets this so the
* session registry records the actual profile rather than the requested/default one.
*
* @return the profile name, or {@code null} when the launcher leaves it unspecified
*/
default String profile() {
return null;
}
}
@@ -0,0 +1,85 @@
package dev.ltms.bridged.peer;
import java.util.List;
import java.util.Set;
/**
* SPI for materializing a connected peer — the only way the bridge core creates or tears down
* a peer process. Every launcher is a first-party, in-tree adapter selected by (future) profile
* config; today's single adapter is the {@code ClaudeCodeLauncher} / Claude Code over herdr.
*
* <p>The core delegates spawn and teardown to this interface without knowing how the peer is set
* up. Environment variables, CLI flags, subscription guards, transport (herdr tab/pane) layout,
* and naming conventions are all adapter-private — the core sees only the returned
* {@link PeerHandle} whose {@code id()} is the registry/routing key.
*
* <p>The interface is a superset of what {@code SessionManager} and {@code Bridged.main} call
* on the concrete launcher today.
*/
public interface PeerLauncher {
/**
* The set of {@link Capability capabilities} this launcher declares. A peer whose profile
* opts into a git-forge token should include {@link Capability#SELF_PR}; the base set for
* the Claude Code herdr adapter is always {@code MID_TURN_ASK, WORKTREE, ORPHAN_REAP}.
*/
Set<Capability> capabilities();
/**
* {@code profileName}/requestedCwd null/blank → default resolution. Returns after the peer
* process is live (env + argv + placement complete). Never returns {@code null}.
*
* @param req the spawn parameters (profile, requested cwd, caller cwd)
* @return a handle whose {@link PeerHandle#id()} is the registry/routing key
* @throws IllegalArgumentException if the profile is unknown and no default is configured
*/
PeerHandle spawn(SpawnRequest req);
/**
* The configured worker profile names — the set of names {@code spawn(profileName)} accepts.
*/
Set<String> profiles();
/**
* The profile a no-argument {@link #spawn(SpawnRequest)} uses, or {@code null} if none is configured.
*/
String defaultProfile();
/**
* Resolve the effective working directory for a spawn {@code req} without actually spawning.
* Resolution order: requestedCwd → profile cwd → callerCwd → daemon cwd.
*
* @return the resolved absolute path, never null/blank
*/
String effectiveCwd(SpawnRequest req);
/**
* The parity-overlay file list for {@code profileName} (default list when unset). Used by
* worktree provisioning to copy config files into the isolated checkout before spawning.
*/
List<String> parityOverlay(String profileName);
/**
* The set of all agents this launcher currently tracks, transport-specific. Each element
* exposes at minimum a pane-like {@code id()} matching this launcher's {@link PeerHandle}
* scheme, plus transport-level status. Callers merge this set with the session registry to
* build a live roster view.
*/
List<?> list();
/**
* Reap orphaned peers left behind by a prior daemon process. Only peers whose naming scheme
* matches this launcher's and whose nonce differs from the current process are eligible.
* Best-effort: a failure to list or to stop any one peer is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
int reapOrphanWorkers();
/**
* Tear a peer down by its registry/routing key ({@link PeerHandle#id()}). Tolerates an
* already-gone peer. Also cleans up launcher-private resources (e.g. empty dedicated tabs)
* when safe to do so.
*/
void stop(String id);
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.peer;
package dev.ltms.bridged.peer;
/**
* Thrown when a {@link PeerLauncher} starts a peer process but the peer
@@ -0,0 +1,13 @@
package dev.ltms.bridged.peer;
/**
* Parameters for a {@link PeerLauncher#spawn(SpawnRequest)} call — the peer-neutral
* aggregation of what the core knows at delegation time: which profile to use, the caller's
* requested working directory, and the caller's own cwd (to inherit when no other cwd is set).
*
* <p>A null or blank {@code profileName} means "use the launcher's default profile."
* A null or blank {@code requestedCwd} means "inherit from config or caller."
* A null {@code callerCwd} means "the request came from the daemon itself (not a primary)."
*/
public record SpawnRequest(String profileName, String requestedCwd, String callerCwd) {
}
@@ -0,0 +1,22 @@
package dev.ltms.bridged.placement;
/**
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps and
* reachability so that a pre-existing config behaves identically after upgrade.
*/
final class FixedPlacementPolicy implements PlacementPolicy {
@Override
public PlacementCandidate select(PlacementContext ctx) {
String d = ctx.defaultProfile();
if (d != null && !d.isBlank()) {
return new PlacementCandidate(d, null, 1.0f, null);
}
if (!ctx.candidates().isEmpty()) {
PlacementCandidate first = ctx.candidates().getFirst();
return new PlacementCandidate(first.profile(), null, first.weight(), first.maxLoad());
}
throw new PlacementException("no worker profiles configured");
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
/**
* A profile (and, in CB-308, a host) that can be chosen by a {@link PlacementPolicy}.
@@ -17,18 +17,4 @@ public record PlacementCandidate(String profile, String host, float weight, Inte
public static PlacementCandidate profile(String profile) {
return new PlacementCandidate(profile, null, 1.0f, null);
}
/**
* True when this candidate carries an explicit {@code weight <= 0} (CB-554) and must be
* skipped by every automatic policy — the same way a quarantined or unreachable candidate is
* skipped. {@code FleetConfig.Profile}'s compact constructor already normalises "absent" to
* {@code 1.0} and "negative" to {@code 0.0}, so this is a plain threshold check here; it does
* not need to distinguish "explicit 0" from "absent" itself.
*
* <p>Exclusion is about <em>automatic</em> selection only — an explicit
* {@code fleet_spawn{profile:"..."}} bypasses placement entirely and is unaffected.
*/
public boolean excluded() {
return weight <= 0.0f;
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.Set;
@@ -11,13 +11,9 @@ import java.util.function.Function;
* @param candidates every configured candidate; the policy filters out those at cap or unreachable
* @param liveCount current live worker count per profile (from the session registry)
* @param unreachable profiles already known to have failed in this spawn attempt
* @param quarantined profiles whose credential is currently quarantined (CB-578 stage B) — a
* {@code BACKEND_EXHAUSTED} classification put it, or a profile it shares a
* credential with, on cooldown. Filtered the same way as {@code unreachable}.
*/
public record PlacementContext(String defaultProfile,
List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable,
Set<String> quarantined) {
Set<String> unreachable) {
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
/**
* Thrown when a {@link PlacementPolicy} has no candidate available. Kept as a distinct type so
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
/**
* Factory for the built-in placement policies.
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
/**
* How {@code bridged} chooses a worker profile when a spawn names none. Implementations are
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
import java.util.ArrayList;
import java.util.List;
@@ -12,16 +12,13 @@ final class PlacementPolicyUtil {
}
/**
* Candidates that are not weight-excluded (CB-554: explicit {@code weight <= 0}, checked
* first because it is a static config choice rather than transient state), not
* known-unreachable, not quarantined (CB-578 stage B), and have not reached their maxLoad.
* Candidates that are not known-unreachable and have not reached their maxLoad.
* A {@code null} maxLoad means unlimited.
*/
static List<PlacementCandidate> available(PlacementContext ctx) {
List<PlacementCandidate> out = new ArrayList<>();
for (PlacementCandidate c : ctx.candidates()) {
if (c.excluded() || ctx.unreachable().contains(c.profile())
|| ctx.quarantined().contains(c.profile())) {
if (ctx.unreachable().contains(c.profile())) {
continue;
}
Integer cap = c.maxLoad();
@@ -37,23 +34,15 @@ final class PlacementPolicyUtil {
}
/**
* Build a clear exception describing why every candidate was dropped: all weight-0, all
* quarantined, all at capacity, all unreachable, or a mix. Each candidate is counted into
* exactly one bucket (weight-excluded takes priority) so a candidate excluded for more than
* one reason is never double-counted.
* Build a clear exception describing why every candidate was dropped: all at capacity,
* all unreachable, or a mix.
*/
static PlacementException emptyException(PlacementContext ctx) {
int weightExcluded = 0;
int atCap = 0;
int unreachable = 0;
int quarantined = 0;
for (PlacementCandidate c : ctx.candidates()) {
Integer cap = c.maxLoad();
if (c.excluded()) {
weightExcluded++;
} else if (ctx.quarantined().contains(c.profile())) {
quarantined++;
} else if (ctx.unreachable().contains(c.profile())) {
if (ctx.unreachable().contains(c.profile())) {
unreachable++;
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
atCap++;
@@ -64,13 +53,6 @@ final class PlacementPolicyUtil {
if (total == 0) {
return new PlacementException("no worker profiles configured");
}
if (weightExcluded == total) {
return new PlacementException(
"all worker profiles have weight 0 (excluded from automatic selection)");
}
if (quarantined == total) {
return new PlacementException("all worker profiles are quarantined (backend exhausted)");
}
if (atCap == total) {
return new PlacementException("all worker profiles are at maxLoad");
}
@@ -78,8 +60,6 @@ final class PlacementPolicyUtil {
return new PlacementException("all worker profiles are unreachable");
}
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
+ unreachable + " unreachable, " + quarantined + " quarantined, "
+ weightExcluded + " weight-0, "
+ (total - atCap - unreachable - quarantined - weightExcluded) + " remaining");
+ unreachable + " unreachable, " + (total - atCap - unreachable) + " remaining");
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.concurrent.atomic.AtomicInteger;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.Map;
@@ -1,26 +1,23 @@
package dev.ltms.fleet.rest;
package dev.ltms.bridged.rest;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.auth.AuditLog;
import dev.ltms.fleet.auth.Authz;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.guard.GuardException;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.placement.PlacementException;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.WorktreeRequest;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.bridged.auth.AuditLog;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.peer.PeerLauncher;
import io.javalin.Javalin;
import io.javalin.http.Context;
import jakarta.servlet.http.HttpServlet;
@@ -42,12 +39,12 @@ import java.util.stream.Collectors;
* <p>Built from injected collaborators so tests supply fakes and run on an ephemeral
* port; {@code main} supplies the real Unix-socket client and worker service.
*/
public final class FleetApp {
public final class BridgedApp {
/** Default blocking window for a message; kept under typical HTTP idle timeouts. */
private static final long DEFAULT_MESSAGE_TIMEOUT_MS = 25_000;
private static final long MAX_MESSAGE_TIMEOUT_MS = 120_000;
/** Blocking window for a worker's fleet_ask (CB-205); the worker's MCP client caps its own call. */
/** Blocking window for a worker's bridge_ask (CB-205); the worker's MCP client caps its own call. */
private static final long DEFAULT_ASK_TIMEOUT_MS = 55_000;
private static final long MAX_ASK_TIMEOUT_MS = 115_000;
@@ -58,7 +55,7 @@ public final class FleetApp {
private final PeerLauncher workers;
private final SessionManager sessions; // CB-301: authoritative session registry
private final MessageService messages;
private final MemberPresence presence; // CB-113: which workers are MCP-connected (available)
private final WorkerPresence presence; // CB-113: which workers are MCP-connected (available)
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
private final CallerResolver auth; // CB-501: null → authz not enforced (legacy behaviour)
private final Metrics metrics; // CB-502: null → /metrics not exposed
@@ -69,8 +66,8 @@ public final class FleetApp {
* behaved before CB-501. Retained so existing acceptance tests keep exercising handler
* behaviour without each needing an auth fixture.
*/
public FleetApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, MemberPresence presence,
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
HttpServlet mcpServlet) {
this(herdr, workers, sessions, messages, presence, mcpServlet, null, null);
}
@@ -81,8 +78,8 @@ public final class FleetApp {
* @param metrics registry to instrument and expose at {@code GET /metrics}; {@code null} omits
* the endpoint
*/
public FleetApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, MemberPresence presence,
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics) {
this.herdr = herdr;
this.workers = workers;
@@ -105,7 +102,7 @@ public final class FleetApp {
}
});
// CB-501: resolve identity once per request, before any handler. /mcp does NOT pass through
// here — it is a raw servlet on Jetty's context handler — so FleetMcp enforces separately
// here — it is a raw servlet on Jetty's context handler — so BridgeMcp enforces separately
// against the same CallerResolver. Any check that lives in only one place is not a control.
if (auth != null) {
app.before(ctx -> ctx.attribute(CALLER,
@@ -118,15 +115,15 @@ public final class FleetApp {
}
app.get("/sessions", this::sessions);
app.get("/agents", this::agents);
app.get("/members", this::listMembers); // CB-304: registry roster + live herdr status
app.get("/profiles", this::profiles); // configured backend profiles
app.post("/members", this::spawnMember); // optional ?role=&profile= or {"role":…,"profile":…}
app.delete("/members/{paneId}", this::stopMember);
app.post("/sessions/{id}/message", this::sendMessage); // fleet_send (primary; blocking, wait:false, or answer via turnId)
app.post("/sessions/{id}/reply", this::replyMessage); // fleet_reply (worker)
app.get("/workers", this::listWorkers); // CB-304: registry roster + live herdr status
app.get("/profiles", this::profiles); // configured worker profiles
app.post("/workers", this::spawnWorker); // optional ?profile= or {"profile":…}
app.delete("/workers/{paneId}", this::stopWorker);
app.post("/sessions/{id}/message", this::sendMessage); // bridge_send (primary; blocking, wait:false, or answer via turnId)
app.post("/sessions/{id}/reply", this::replyMessage); // bridge_reply (worker)
app.get("/sessions/{id}/replies", this::drainReplies); // drain reply inbox (CB-307)
app.post("/sessions/{id}/ask", this::askMessage); // fleet_ask (worker → primary, CB-205)
app.get("/sessions/{id}/status", this::sessionStatus); // fleet_status
app.post("/sessions/{id}/ask", this::askMessage); // bridge_ask (worker → primary, CB-205)
app.get("/sessions/{id}/status", this::sessionStatus); // bridge_status
app.get("/tasks/{ticket}", this::taskStatus); // poll an async (wait:false) send
return app;
}
@@ -167,7 +164,7 @@ public final class FleetApp {
private void countAuthFailure(String reason) {
if (metrics != null) {
metrics.inc(FleetMetrics.AUTH_FAILURES, "reason", reason);
metrics.inc("bridged_auth_failures_total", "reason", reason);
}
}
@@ -220,11 +217,11 @@ public final class FleetApp {
return;
}
ctx.status(200).json(Map.of("agents",
workers.list().stream().map(Agent.class::cast).map(FleetApp::view).toList()));
workers.list().stream().map(Agent.class::cast).map(BridgedApp::view).toList()));
}
/** CB-304: bridge-owned roster merged with live herdr status by paneId. */
private void listMembers(Context ctx) {
private void listWorkers(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
@@ -236,14 +233,7 @@ public final class FleetApp {
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
Map<String, Object> body = new LinkedHashMap<>();
body.put("workers", out);
// CB-586: operator visibility for the refs/wip snapshot store without shelling into the
// repo — how many snapshot refs exist and roughly what they cost. Present only once a
// worktree session has established the repo, so a never-snapshotted fleet reports nothing.
sessions.wipRefs().ifPresent(st -> body.put("wipRefs",
Map.of("count", st.count(), "costBytes", st.costBytes())));
ctx.status(200).json(body);
ctx.status(200).json(Map.of("workers", out));
}
/** The configured worker profiles and which one a no-argument spawn uses. */
@@ -261,11 +251,10 @@ public final class FleetApp {
* body) picks which configured profile; omitted → the default. 403 if the base_url would breach
* the subscription boundary, 400 for an unknown profile.
*/
private void spawnMember(Context ctx) {
private void spawnWorker(Context ctx) {
if (!allow(ctx, Authz.Action.SPAWN, null)) {
return;
}
String role = ctx.queryParam("role");
String profile = ctx.queryParam("profile");
String cwd = ctx.queryParam("cwd");
String worktree = ctx.queryParam("worktree");
@@ -276,7 +265,6 @@ public final class FleetApp {
String body = ctx.body();
if (!body.isBlank()) {
JsonNode b = mapper.readTree(body);
if (role == null || role.isBlank()) role = b.path("role").asText(null);
if (profile == null || profile.isBlank()) profile = b.path("profile").asText(null);
if (cwd == null || cwd.isBlank()) cwd = b.path("cwd").asText(null);
if (worktree == null || worktree.isBlank()) worktree = b.path("worktree").asText(null);
@@ -287,25 +275,12 @@ public final class FleetApp {
}
}
WorktreeRequest wt = worktreeRequest(worktree, ticket);
MemberRole memberRole;
try {
memberRole = (role == null || role.isBlank()) ? MemberRole.DEV : MemberRole.parse(role);
} catch (IllegalArgumentException e) {
ctx.status(400).json(Map.of("error", "unknown_role", "detail", e.getMessage()));
return;
}
try {
// No MCP caller over REST, so callerCwd and ownerTerminal are null.
MemberSession member = sessions.acquire(blankToNull(profile), memberRole,
blankToNull(cwd), null, null, wt);
ctx.status(201).json(view(member));
WorkerSession worker = sessions.acquire(blankToNull(profile), blankToNull(cwd), null, null, wt);
ctx.status(201).json(view(worker));
} catch (GuardException e) {
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
} catch (PlacementException e) {
// CB-599: no candidate had capacity (maxLoad, quarantine, or all-exhausted) — a benign,
// likely-transient refusal, distinct from "profile does not exist" below. 503: the
// request was valid and will likely succeed later.
ctx.status(503).json(Map.of("error", "no_capacity", "detail", e.getMessage()));
} catch (IllegalArgumentException e) {
ctx.status(400).json(Map.of("error", "unknown_profile", "detail", e.getMessage()));
} catch (PeerUnreachableException e) {
@@ -331,7 +306,7 @@ public final class FleetApp {
}
/** Tear a worker down by pane id. */
private void stopMember(Context ctx) {
private void stopWorker(Context ctx) {
String paneId = ctx.pathParam("paneId");
if (!allow(ctx, Authz.Action.STOP, paneId)) {
return;
@@ -342,7 +317,7 @@ public final class FleetApp {
/**
* The blocking delegation call (CB-104): inject {@code content} into the worker via the
* status-gated injector and block until the worker returns a structured {@code fleet_reply}.
* status-gated injector and block until the worker returns a structured {@code bridge_reply}.
* Times out with a typed 202 (working / queued / busy) rather than an error — the message may
* still land.
*/
@@ -371,7 +346,7 @@ public final class FleetApp {
}
timeout = Math.clamp(timeout, 1, MAX_MESSAGE_TIMEOUT_MS);
// Answering a worker's fleet_ask (CB-205): always blocks, and derives the worker from turnId.
// Answering a worker's bridge_ask (CB-205): always blocks, and derives the worker from turnId.
if (turnId != null && !turnId.isBlank()) {
writeReply(ctx, id, messages.answer(turnId, content, timeout), timeout);
return;
@@ -393,7 +368,7 @@ public final class FleetApp {
/**
* Render a {@link MessageService.Reply} onto the response — shared by a normal send and a
* fleet_ask answer. A structured/scraped completion is 200; a worker's mid-turn question a 202
* bridge_ask answer. A structured/scraped completion is 200; a worker's mid-turn question a 202
* (with its {@code turnId}); a stale answer a 409; every other non-terminal outcome a typed 202.
*/
private void writeReply(Context ctx, String id, MessageService.Reply reply, long timeout) {
@@ -405,7 +380,7 @@ public final class FleetApp {
"sessionId", id, "error", "stale_turn",
"detail", "that question is no longer open (timed out or already answered)"));
case REPLIED, COMPLETED_UNREPLIED -> {
// replySource distinguishes a structured fleet_reply from the CB-106 completion
// replySource distinguishes a structured bridge_reply from the CB-106 completion
// fallback (a scrape of the worker's transcript when it finished without replying).
String source = reply.outcome() == MessageService.Outcome.REPLIED ? "reply" : "transcript";
ctx.status(200).json(Map.of("sessionId", id, "reply", reply.text(), "replySource", source));
@@ -417,19 +392,16 @@ public final class FleetApp {
case TIMED_OUT_QUEUED -> "queued";
case BUSY -> "busy";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
default -> "done"; // unreachable (terminal outcomes handled above)
},
"detail", (reply.outcome() == MessageService.Outcome.WORKER_FAILED
|| reply.outcome() == MessageService.Outcome.BACKEND_EXHAUSTED)
&& reply.text() != null
"detail", reply.outcome() == MessageService.Outcome.WORKER_FAILED && reply.text() != null
? reply.text()
: "no reply within " + timeout + "ms; poll status or retry"));
}
}
/**
* A worker's mid-turn question ({@code fleet_ask}, CB-205) — surfaces to the primary's open
* A worker's mid-turn question ({@code bridge_ask}, CB-205) — surfaces to the primary's open
* blocking send and blocks until it answers. 200 with the answer, 409 if no delegation is open,
* 202 if the primary stayed silent.
*/
@@ -466,7 +438,7 @@ public final class FleetApp {
}
/**
* The worker's structured reply ({@code fleet_reply}) — resolves the blocking send awaiting
* The worker's structured reply ({@code bridge_reply}) — resolves the blocking send awaiting
* on this session, or queues the reply in the inbox when no send is open (CB-307).
*/
private void replyMessage(Context ctx) {
@@ -506,7 +478,7 @@ public final class FleetApp {
}
/**
* Live lifecycle status of a worker (MCP `fleet_status` wraps this in CB-105), plus its
* Live lifecycle status of a worker (MCP `bridge_status` wraps this in CB-105), plus its
* <em>readiness</em> (CB-113): {@code ready} is true once the worker's Claude has connected the
* bridge MCP — the reliable "available to receive a task" signal, unlike bare {@code idle}, which
* is also true during boot.
@@ -517,20 +489,10 @@ public final class FleetApp {
return;
}
try {
Map<String, Object> body = new LinkedHashMap<>();
body.put("sessionId", id);
body.put("status", messages.status(id).name().toLowerCase());
body.put("ready", presence.isPresent(id));
// CB-582: a worker paused mid-turn in an async fleet_ask is otherwise invisible to a
// status poll — surface the open question and how to answer it, same as fleet_poll's
// Phase.ASKING view.
MessageService.PendingAsk ask = messages.pendingAsk(id);
if (ask != null) {
body.put("question", ask.question());
body.put("turnId", ask.turnId());
body.put("ticket", ask.ticket());
}
ctx.status(200).json(body);
ctx.status(200).json(Map.of(
"sessionId", id,
"status", messages.status(id).name().toLowerCase(),
"ready", presence.isPresent(id)));
} catch (HerdrException e) {
herdrError(ctx, e);
}
@@ -556,11 +518,6 @@ public final class FleetApp {
if (v.detail() != null) {
body.put("detail", v.detail());
}
// CB-582: Phase.ASKING carries the question in v.reply() (handled above) and its answer-
// correlation id here — a REST caller polling this ticket otherwise has no way to answer it.
if (v.turnId() != null) {
body.put("turnId", v.turnId());
}
ctx.status(200).json(body);
}
@@ -587,7 +544,7 @@ public final class FleetApp {
}
/** CB-301 projection of an authoritative bridge-owned session. */
private static Map<String, Object> view(MemberSession s) {
private static Map<String, Object> view(WorkerSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("terminalId", s.terminalId());
m.put("paneId", s.paneId());
@@ -0,0 +1,217 @@
package dev.ltms.bridged.session;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.BufferedReader;
import java.io.IOException;
import java.io.InputStreamReader;
import java.io.UncheckedIOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.security.SecureRandom;
import java.util.List;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.stream.Collectors;
/**
* Production {@link Worktrees} implementation that shells {@code git} via {@link ProcessBuilder}.
* Non-zero exits become {@link WorktreeException}. Worktree directories live under a configurable
* root (default: a sibling {@code .bridged-worktrees} of the repo root) so they are never nested
* inside the primary working tree.
*/
public final class GitWorktrees implements Worktrees {
private static final Logger log = LoggerFactory.getLogger(GitWorktrees.class);
/** Project-level MCP config. Present in the repo, so every worktree checks the primary's out. */
private static final String MCP_CONFIG = ".mcp.json";
/** What {@link #isolateToolSurface} writes: a valid, explicitly empty server map. */
private static final String NEUTRAL_MCP_CONFIG = "{\n \"mcpServers\": {}\n}\n";
private final String configuredRoot;
private final SecureRandom random = new SecureRandom();
private final AtomicLong seq = new AtomicLong();
/** Default constructor: worktree root is derived per-repo as {@code <repoRoot>/../.bridged-worktrees}. */
public GitWorktrees() {
this(null);
}
/** @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of the repo root. */
public GitWorktrees(String configuredRoot) {
this.configuredRoot = configuredRoot;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
String base = (baseRef == null || baseRef.isBlank()) ? "HEAD" : baseRef;
String nonce = nonce();
Path root = resolveRoot(repoRoot);
Path path = root.resolve(nonce);
try {
Files.createDirectories(root);
} catch (IOException e) {
throw new WorktreeException("cannot create worktree root " + root + ": " + e.getMessage(), e);
}
String wt = path.toAbsolutePath().toString();
log.info("adding worktree branch={} path={} base={}", branch, wt, base);
exec("git", "-C", repoRoot, "worktree", "add", wt, "-b", branch, base);
isolateToolSurface(wt);
return wt;
}
/**
* Neutralize the worktree's project MCP config so a worker inherits only the tools its launcher
* mounts (the bridge, via {@code --mcp-config}) — never the primary's.
*
* <p>This is unconditional, and it is not the same job as the parity overlay. The repo's own
* committed {@code .mcp.json} declares the primary's IDE servers, so a fresh checkout mounts them
* whether or not the overlay copies anything; a worker that inherits them navigates and edits
* through tools bound to the <em>primary's</em> IntelliJ project, which silently hands it absolute
* paths outside its own worktree. That is not hypothetical: a CB-523 worker made all 59 of its
* edits in the primary checkout while compiling its worktree, so every build it ran was of code
* that did not contain its changes.
*
* <p>Writing an empty server map (rather than deleting the file) keeps a project-level
* {@code .mcp.json} present and explicit, and the {@code --skip-worktree} bit keeps the
* neutralized copy from ever showing up as a local modification the worker might commit.
*/
private void isolateToolSurface(String worktreePath) {
Path root = Path.of(worktreePath).toAbsolutePath().normalize();
Path mcp = root.resolve(MCP_CONFIG);
try {
Files.writeString(mcp, NEUTRAL_MCP_CONFIG);
} catch (IOException e) {
throw new WorktreeException("cannot neutralize " + MCP_CONFIG + " in the worktree: "
+ e.getMessage(), e);
}
if (isTracked(root, MCP_CONFIG)) {
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", MCP_CONFIG);
}
log.debug("neutralized {} — worker tool surface is launcher-mounted only", MCP_CONFIG);
}
@Override
public void remove(String repoRoot, String worktreePath) {
Path p = Path.of(worktreePath);
if (!Files.exists(p)) {
log.debug("worktree {} already gone — nothing to remove", worktreePath);
return;
}
log.info("removing worktree {}", worktreePath);
exec("git", "-C", repoRoot, "worktree", "remove", "--force", worktreePath);
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
if (overlay == null || overlay.isEmpty()) {
return;
}
Path srcRoot = Path.of(repoRoot).toAbsolutePath().normalize();
Path dstRoot = Path.of(worktreePath).toAbsolutePath().normalize();
for (String rel : overlay) {
Path src = srcRoot.resolve(rel).normalize();
if (!Files.exists(src)) {
log.debug("parity overlay source missing — skipping {}", rel);
continue;
}
Path dst = dstRoot.resolve(rel).normalize();
try {
Files.createDirectories(dst.getParent());
Files.copy(src, dst, StandardCopyOption.REPLACE_EXISTING, StandardCopyOption.COPY_ATTRIBUTES);
log.debug("copied parity overlay {}", rel);
} catch (IOException e) {
throw new WorktreeException("cannot copy overlay " + rel + ": " + e.getMessage(), e);
}
if (isTracked(dstRoot, rel)) {
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", rel);
log.debug("marked overlay --skip-worktree {}", rel);
}
}
}
@Override
public String repoRoot(String cwd) {
String out = exec("git", "-C", cwd, "rev-parse", "--show-toplevel");
return Path.of(out.trim()).toAbsolutePath().normalize().toString();
}
/** Resolve the directory that will hold per-session worktree checkouts. */
private Path resolveRoot(String repoRoot) {
if (configuredRoot != null && !configuredRoot.isBlank()) {
return Path.of(configuredRoot).toAbsolutePath().normalize();
}
Path repo = Path.of(repoRoot).toAbsolutePath().normalize();
return repo.resolveSibling(".bridged-worktrees");
}
private String nonce() {
return String.format("%06x", random.nextInt(1 << 24)) + "-" + seq.incrementAndGet();
}
private boolean isTracked(Path worktreeRoot, String rel) {
return exitCode("git", "-C", worktreeRoot.toString(), "ls-files", "--error-unmatch", rel) == 0;
}
/**
* Run a command and return its stdout. Non-zero exit → {@link WorktreeException} with both
* stdout and stderr (merged by redirectErrorStream).
*/
private String exec(String... command) {
String out;
int code;
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
try (BufferedReader r = new BufferedReader(new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
out = r.lines().collect(Collectors.joining("\n"));
} catch (IOException e) {
p.destroyForcibly();
throw new UncheckedIOException(e);
}
try {
if (!p.waitFor(30, TimeUnit.SECONDS)) {
p.destroyForcibly();
throw new WorktreeException("command timed out: " + String.join(" ", command) + "\n" + out);
}
code = p.exitValue();
} catch (InterruptedException e) {
p.destroyForcibly();
Thread.currentThread().interrupt();
throw new WorktreeException("interrupted waiting for command: " + String.join(" ", command), e);
}
if (code != 0) {
throw new WorktreeException("exit " + code + " for: " + String.join(" ", command)
+ (out.isBlank() ? "" : "\n" + out));
}
return out;
}
private int exitCode(String... command) {
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
try {
if (!p.waitFor(30, TimeUnit.SECONDS)) {
p.destroyForcibly();
throw new WorktreeException("command timed out: " + String.join(" ", command));
}
return p.exitValue();
} catch (InterruptedException e) {
p.destroyForcibly();
Thread.currentThread().interrupt();
throw new WorktreeException("interrupted waiting for command: " + String.join(" ", command), e);
}
}
}
@@ -0,0 +1,504 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Consumer;
import java.util.function.LongSupplier;
/**
* Authoritative in-daemon registry of the worker sessions this {@code bridged} process spawned.
* Delegates spawn/teardown to a {@link PeerLauncher} (which performs subscription-guarded env
* setup and process/materialization) and adds lifecycle tracking, ownership, and deterministic
* teardown on top.
*
* <p>The state machine is intentionally one-shot / no-reuse: every acquired worker is fresh,
* and a finished or released worker is torn down, never pooled. {@link #recycle} is a convenience
* for {@code release + acquire} with a new distinct pane id.
*
* <p>The manager implements {@link TurnListener} so the injector's turn boundaries drive
* {@code READY → BUSY → DONE} (or {@code FAILED}). It exposes a {@link WorkerPresence} view via
* {@link #asPresence()}: any MCP contact from a worker marks it present and simultaneously
* transitions the session {@code SPAWNING → READY}.
*/
public final class SessionManager implements TurnListener {
private static final Logger log = LoggerFactory.getLogger(SessionManager.class);
private final PeerLauncher launcher;
private final Worktrees worktrees;
private final ConcurrentHashMap<String /*paneId*/, WorkerSession> registry = new ConcurrentHashMap<>();
private final WorkerPresence presence;
private final SecureRandom nonceRandom = new SecureRandom();
private final AtomicLong nonceSeq = new AtomicLong();
private final LongSupplier nowNanos;
private final int contextCap;
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** CB-516: notified with a terminalId on every release; no-op until wired. */
private final List<Consumer<String>> releaseListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
public SessionManager(PeerLauncher launcher) {
this(launcher, new GitWorktrees(), System::nanoTime, 0);
}
/** Backward-compatible constructor with an injectable worktree seam. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees) {
this(launcher, worktrees, System::nanoTime, 0);
}
/** Test constructor with an injectable clock. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos) {
this(launcher, worktrees, nowNanos, 0);
}
/** Production constructor with a configured context turn cap. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, int contextCap) {
this(launcher, worktrees, System::nanoTime, contextCap);
}
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos,
int contextCap) {
this.launcher = launcher;
this.worktrees = worktrees;
this.presence = new PresenceBridge(this);
this.nowNanos = nowNanos;
this.contextCap = contextCap;
}
/**
* The single {@link WorkerPresence} view of this manager: it records availability and forwards
* the signal to the {@code SPAWNING → READY} transition. Pass this to the {@code Injector} and
* {@code BridgeMcp} where they previously accepted a plain {@link WorkerPresence}. The same
* instance is returned every call — presence is shared state, so a fresh bridge per call would
* fragment the {@code present} set and lose signals across callers.
*/
public WorkerPresence asPresence() {
return presence;
}
/**
* Spawn a worker and register it as {@link WorkerSession.State#SPAWNING}. The caller's
* identity is recorded as {@code ownerTerminal} ({@code null} for daemon/anon callers).
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal) {
return acquire(profile, requestedCwd, callerCwd, ownerTerminal, null);
}
/**
* Spawn a worker, optionally inside a fresh git worktree. When {@code wt} is non-null the
* worktree is provisioned, parity-overlaid, and its path becomes the worker's cwd. On any
* failure before registration the worktree is removed so no dangling checkout is left.
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
if (wt == null) {
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd);
PeerHandle handle = launcher.spawn(req);
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd));
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
cwd,
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
null,
null);
registry.put(handle.id(), session);
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
notifyAcquired(session.terminalId());
return session;
}
return acquireWithWorktree(profile, requestedCwd, callerCwd, ownerTerminal, wt);
}
/** Tear a worker down by pane id and remove it from the registry. Idempotent. */
public void release(String paneId) {
WorkerSession removed = registry.remove(paneId);
if (removed != null) {
log.debug("releasing session pane={} terminal={} state={}",
removed.paneId(), removed.terminalId(), removed.state());
// CB-516: a send still waiting on this worker can never be answered now. Tell the
// listener BEFORE the pane is torn down, so a blocked caller fails fast with a real
// reason instead of sitting on a rendezvous nothing will ever resolve.
notifyReleased(removed.terminalId());
}
launcher.stop(paneId);
if (removed != null && removed.worktree() != null) {
worktrees.remove(worktrees.repoRoot(removed.cwd()), removed.worktree());
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is acquired
* (CB-520). This is the hook that lets the reply inbox {@code own} a target's queue.
*/
public void onAcquire(Consumer<String> listener) {
if (listener != null) {
acquireListeners.add(listener);
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is released
* (CB-516). Every teardown path funnels through {@link #release}, so one hook covers the REST
* and MCP stop tools, the idle-TTL reaper, {@code recycle}, and shutdown drain alike.
*
* <p>Added rather than injected because {@code MessageService} — one intended listener — is
* constructed after this manager (it needs the injector and rendezvous, which need the session
* presence view this manager exposes). Wiring it at construction would require breaking that
* cycle for one callback.
*/
public void onRelease(Consumer<String> listener) {
if (listener != null) {
releaseListeners.add(listener);
}
}
/** A listener failure must never prevent the acquisition it is reacting to. */
private void notifyAcquired(String terminalId) {
if (terminalId == null) {
return;
}
for (Consumer<String> listener : acquireListeners) {
try {
listener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("acquire listener failed for terminal {}: {}", terminalId, e.toString());
}
}
}
/** A listener failure must never prevent the teardown it is reacting to. */
private void notifyReleased(String terminalId) {
if (terminalId == null) {
return;
}
for (Consumer<String> listener : releaseListeners) {
try {
listener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
}
}
}
private WorkerSession acquireWithWorktree(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
String preResolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
// nor a caller cwd, so taking the first non-blank of those two yielded null and put
// `git -C null` on the command line — an NPE out of ProcessBuilder, surfacing as HTTP 500.
// The non-worktree path always used this chain; only this branch was missed.
String repoRoot = worktrees.repoRoot(
launcher.effectiveCwd(new SpawnRequest(preResolvedProfile, requestedCwd, callerCwd)));
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
String path = null;
PeerHandle handle;
try {
path = worktrees.add(repoRoot, branch, wt.baseRef());
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd));
} catch (RuntimeException e) {
if (path != null) {
try {
worktrees.remove(repoRoot, path);
} catch (RuntimeException cleanup) {
log.warn("failed to clean up worktree {} after spawn error: {}", path, cleanup.getMessage());
}
}
throw e;
}
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
cwd,
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
path,
branch);
registry.put(handle.id(), session);
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
notifyAcquired(session.terminalId());
return session;
}
private String slug(String raw) {
return raw == null ? "ticket" : raw.toLowerCase().replaceAll("[^a-z0-9]+", "-").replaceAll("^-+|-+$", "");
}
private String nonce() {
return String.format("%06x", nonceRandom.nextInt(1 << 24)) + "-" + nonceSeq.incrementAndGet();
}
/**
* The profile to record for a session. A launcher that performed dynamic selection tells us
* the actual profile via {@link PeerHandle#profile()}; otherwise fall back to what the caller
* requested (or the launcher's default for a no-profile spawn).
*/
private String resolveProfile(PeerHandle handle, String requestedProfile) {
String fromHandle = handle.profile();
if (fromHandle != null && !fromHandle.isBlank()) {
return fromHandle;
}
if (requestedProfile != null && !requestedProfile.isBlank()) {
return requestedProfile;
}
return launcher.defaultProfile();
}
/**
* Release the old session and acquire a fresh one with the same profile and working directory.
* The new session is guaranteed to have a pane id distinct from the old one (no-reuse invariant).
*/
public WorkerSession recycle(String paneId) {
WorkerSession old = registry.get(paneId);
if (old == null) {
throw new IllegalArgumentException("no session for paneId " + paneId);
}
release(paneId);
return acquire(old.profile(), old.cwd(), old.cwd(), old.ownerTerminal());
}
/** The session for {@code paneId}, if it is still registered and not released. */
public Optional<WorkerSession> get(String paneId) {
return Optional.ofNullable(registry.get(paneId));
}
/** Bridge-owned roster: all registered sessions (acquired minus released). */
public List<WorkerSession> roster() {
return List.copyOf(registry.values());
}
/**
* CB-304 merged roster+live view. The registry is authoritative for worktree, branch,
* profile, owner, and state; the optional live agent supplies the herdr-reported status.
*/
public static Map<String, Object> rosterView(WorkerSession session, Agent live) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", session.terminalId());
m.put("paneId", session.paneId());
m.put("profile", session.profile());
m.put("state", session.state().name().toLowerCase());
if (session.worktree() != null) {
m.put("worktree", session.worktree());
}
if (session.branch() != null) {
m.put("branch", session.branch());
}
if (session.ownerTerminal() != null) {
m.put("owner", session.ownerTerminal());
}
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
return m;
}
/** Lifecycle hook: worker became available on the bridge MCP. */
void onReady(String terminalId) {
transitionByTerminal(terminalId, WorkerSession.State.SPAWNING, WorkerSession.State.READY);
}
/**
* Lifecycle hook: a message was delivered into the worker — it is now busy on a turn.
* The turn count is bumped and the activity timestamp is refreshed. A {@code DONE} session
* can be re-delivered for multi-turn reuse until it is released.
*/
@Override
public void onDelivered(String target) {
WorkerSession current = findByTerminal(target);
if (current == null) return;
if (current.state() != WorkerSession.State.READY && current.state() != WorkerSession.State.DONE) {
return;
}
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.BUSY).bumpTurn(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} {} -> BUSY turn={}",
target, current.paneId(), current.state(), updated.turnCount());
}
}
/** Lifecycle hook: the worker's delegated turn completed successfully. */
@Override
public void onTurnComplete(String target) {
WorkerSession current = findByTerminal(target);
if (current == null || current.state() != WorkerSession.State.BUSY) return;
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.DONE).withActivity(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} BUSY -> DONE turn={}",
target, current.paneId(), updated.turnCount());
}
if (contextCap > 0 && updated.turnCount() >= contextCap) {
release(current.paneId());
}
}
/** Lifecycle hook: the worker's delegated turn failed. */
@Override
public void onTurnFailed(String target) {
onFailed(target);
}
/** Lifecycle hook: the worker vanished or was dropped mid-life. */
void onFailed(String target) {
WorkerSession current = findByTerminal(target);
if (current == null) return;
if (current.state() == WorkerSession.State.RELEASED) return;
if (replace(current, current.withState(WorkerSession.State.FAILED))) {
log.debug("session marked failed terminal={} pane={}", target, current.paneId());
}
}
/**
* Best-effort reap of sessions that have been idle longer than {@code idleTtlNanos}. Only
* {@code READY} and {@code DONE} sessions are eligible — never a {@code SPAWNING} or
* {@code BUSY} worker. Returns the number of sessions released.
*/
int reapIdle(long idleTtlNanos) {
long now = nowNanos.getAsLong();
int reaped = 0;
for (WorkerSession s : roster()) {
if (s.state() != WorkerSession.State.READY && s.state() != WorkerSession.State.DONE) {
continue;
}
if (now - s.lastActivityAtNanos() > idleTtlNanos) {
release(s.paneId());
reaped++;
}
}
return reaped;
}
/**
* Gracefully drain all registered sessions. For each session that is {@code BUSY}, poll up to
* {@code timeoutNanos} for it to leave {@code BUSY}, then release it regardless. Non-busy
* sessions are released immediately. A failure releasing one session is logged and does not
* abort the rest.
*/
void drainAll(long timeoutNanos) {
long deadline = System.nanoTime() + timeoutNanos;
for (WorkerSession s : roster()) {
try {
if (s.state() == WorkerSession.State.BUSY) {
while (System.nanoTime() < deadline) {
WorkerSession current = registry.get(s.paneId());
if (current == null || current.state() != WorkerSession.State.BUSY) {
break;
}
try {
long remaining = deadline - System.nanoTime();
Thread.sleep(Math.min(TimeUnit.NANOSECONDS.toMillis(remaining), 50));
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
break;
}
}
}
release(s.paneId());
} catch (RuntimeException e) {
log.warn("drain failed for pane={}; continuing with remaining sessions", s.paneId(), e);
}
}
}
/**
* Close this manager by draining all sessions. The timeout comes from configuration when set,
* otherwise a sensible default.
*/
public void close(Integer drainTimeoutSeconds) {
int seconds = (drainTimeoutSeconds != null && drainTimeoutSeconds > 0) ? drainTimeoutSeconds : 5;
drainAll(TimeUnit.SECONDS.toNanos(seconds));
}
/** Number of sessions currently registered. */
public int size() {
return registry.size();
}
/**
* The registered session owning {@code terminalId}, or {@code null} if none does.
*
* <p>A null {@code terminalId} is a normal input, not a caller bug: every lifecycle hook here is
* fed from the MCP transport, where the <em>primary</em> resolves to a {@link
* dev.ltms.bridged.auth.Principal} with no terminal. {@code BridgeMcp} documents that contact as
* a no-op, and {@link dev.ltms.bridged.inject.WorkerPresence#markPresent} honours it — but
* {@code PresenceBridge} then forwards the same null here. Matching on a null id can never
* succeed anyway (a registered session always has a terminal), so answer "no match" rather than
* throwing: an NPE on this path takes down an unrelated tool call for the primary.
*/
private WorkerSession findByTerminal(String terminalId) {
if (terminalId == null) return null;
for (WorkerSession s : registry.values()) {
if (terminalId.equals(s.terminalId())) return s;
}
return null;
}
private void transitionByTerminal(String terminalId, WorkerSession.State from,
WorkerSession.State to) {
WorkerSession current = findByTerminal(terminalId);
if (current == null || current.state() != from) return;
long now = nowNanos.getAsLong();
if (replace(current, current.withState(to).withActivity(now))) {
log.debug("session transitioned terminal={} pane={} {} -> {}",
terminalId, current.paneId(), from, to);
}
}
private boolean replace(WorkerSession expected, WorkerSession updated) {
return registry.replace(expected.paneId(), expected, updated);
}
/** WorkerPresence bridge that also drives the manager's READY transition. */
private static final class PresenceBridge extends WorkerPresence {
private final SessionManager sessions;
PresenceBridge(SessionManager sessions) {
this.sessions = sessions;
}
@Override
public void markPresent(String terminal) {
if (terminal == null || terminal.isBlank()) {
return; // the primary's contact carries no worker terminal — not a readiness signal
}
super.markPresent(terminal);
sessions.onReady(terminal);
}
}
}
@@ -0,0 +1,70 @@
package dev.ltms.bridged.session;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.TimeUnit;
/**
* Periodic virtual-thread reaper that tears down {@code READY}/{@code DONE} sessions which have
* exceeded their idle TTL. Modeled on {@link dev.ltms.bridged.inject.StatusPoller}: a single
* virtual-thread loop, idempotent start/stop, and no {@code ScheduledExecutorService}.
*/
public final class SessionReaper {
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
private final SessionManager sessions;
private final long idleTtlNanos;
private final long intervalMillis;
private volatile boolean running;
private Thread thread;
/** Construct a reaper with the default 5-second polling interval. */
public SessionReaper(SessionManager sessions, long idleTtlSeconds) {
this(sessions, idleTtlSeconds, DEFAULT_INTERVAL_MILLIS);
}
/** Construct a reaper with an explicit polling interval (useful for tests). */
public SessionReaper(SessionManager sessions, long idleTtlSeconds, long intervalMillis) {
this.sessions = sessions;
this.idleTtlNanos = TimeUnit.SECONDS.toNanos(idleTtlSeconds);
this.intervalMillis = intervalMillis;
}
/** Start the reaper loop on a virtual thread. Idempotent. */
public synchronized void start() {
if (running) return;
running = true;
thread = Thread.ofVirtual().name("session-reaper").start(this::loop);
log.info("session reaper started (idle ttl {}s, interval {}ms)",
TimeUnit.NANOSECONDS.toSeconds(idleTtlNanos), intervalMillis);
}
private void loop() {
while (running) {
try {
sessions.reapIdle(idleTtlNanos);
} catch (RuntimeException e) {
log.warn("session reaper iteration failed; continuing", e);
}
sleep();
}
}
private void sleep() {
try {
Thread.sleep(intervalMillis);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
running = false;
}
}
/** Stop the reaper loop. Idempotent. */
public synchronized void stop() {
running = false;
if (thread != null) thread.interrupt();
}
}
@@ -0,0 +1,61 @@
package dev.ltms.bridged.session;
/**
* A bridge-owned worker session — the authoritative in-daemon record of a worker this
* process spawned. Immutable; state transitions are performed by replacing the record in
* {@link SessionManager}'s registry.
*
* @param paneId the host-unique opaque id (CB-519) — the registry key and the argument to
* teardown. Despite the historical name this is the {@link
* dev.ltms.bridged.peer.PeerHandle#id()}, a UUID, and is distinct from the
* launcher-private herdr pane coordinate.
* @param terminalId herdr terminal handle — the {@code target} for send/read/status
* @param profile the worker profile name that spawned this session
* @param cwd the resolved working directory the worker started in
* @param ownerTerminal the caller that requested this worker ({@code null} = daemon/anon)
* @param spawnedAtNanos {@link System#nanoTime()} when the session was registered
* @param lastActivityAtNanos {@link System#nanoTime()} of the most recent lifecycle event
* @param turnCount number of delegated turns that have been delivered to this session
* @param state current lifecycle state in the one-shot FSM
*/
public record WorkerSession(
String paneId,
String terminalId,
String profile,
String cwd,
String ownerTerminal,
long spawnedAtNanos,
long lastActivityAtNanos,
int turnCount,
State state,
String worktree,
String branch) {
/** One-shot worker lifecycle states. */
public enum State {
SPAWNING,
READY,
BUSY,
DONE,
FAILED,
RELEASED
}
/** Return a copy of this session in {@code state}. */
public WorkerSession withState(State state) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch);
}
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
public WorkerSession withActivity(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount, state, worktree, branch);
}
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
public WorkerSession bumpTurn(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount + 1, state, worktree, branch);
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.session;
package dev.ltms.bridged.session;
/** Non-zero exit or I/O failure from a git worktree operation. */
public final class WorktreeException extends RuntimeException {
@@ -1,4 +1,4 @@
package dev.ltms.fleet.session;
package dev.ltms.bridged.session;
/** Ask {@link SessionManager#acquire} to provision an isolated worktree. null ⇒ run in the shared primary tree. */
public record WorktreeRequest(String ticketSlug, String baseRef) {
@@ -0,0 +1,18 @@
package dev.ltms.bridged.session;
import java.util.List;
/** Seam between {@link SessionManager} and git worktree operations. Tests use a recording fake. */
public interface Worktrees {
/** git -C <repoRoot> worktree add <path> -b <branch> <baseRef|HEAD>. Returns the worktree path. */
String add(String repoRoot, String branch, String baseRef);
/** git -C <repoRoot> worktree remove --force <path>. Idempotent (already-gone tolerated). */
void remove(String repoRoot, String worktreePath);
/** Copy each existing overlay path repoRoot→worktree; mark tracked ones --skip-worktree. */
void overlayParity(String repoRoot, String worktreePath, List<String> overlay);
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
String repoRoot(String cwd);
}
@@ -0,0 +1,195 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>Claude Code</strong> — the safe path from a
* delegation request to a running off-subscription Claude.
*
* <p>Everything transport-related (tab/pane placement, the CB-306 spawn-readiness gate, unique
* naming, CB-117 orphan reap, teardown, listing, cwd resolution) lives in the base. This class
* supplies only the two Claude-specific seams:
* <ul>
* <li>the {@code claude} name prefix (so reap matches {@code claude-*} panes, never another
* adapter's), and</li>
* <li>{@link #buildLaunch}, which encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em> touching
* herdr, and mount the bridge MCP + reply charter as inline launch flags. A worker's base_url
* lives in the env map handed to herdr and nowhere else; {@code bridged}'s own environment is
* never mutated, and nothing is written to the worker's profile.</li>
* </ul>
*/
public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "claude";
private final SubscriptionGuard guard;
/**
* Standing instruction appended to the worker's system prompt so it returns its result via
* {@code bridge_reply}. Injected as a launch flag, so nothing is written to the worker's
* profile — it is guidance, and a worker that never replies is caught by the send's timeout.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet. Every message you "
+ "receive arrives through the bridge, and the ONLY channel back to the sender is the "
+ "bridge_reply MCP tool. Text you write in your terminal is NOT sent anywhere — the "
+ "sender cannot see your screen, so an in-terminal answer is silently discarded. "
+ "Therefore you MUST end EVERY turn by calling bridge_reply with `content` set to your "
+ "complete response. This holds for every message without exception — tasks, questions, "
+ "clarifications, acknowledgements, and ordinary back-and-forth conversation. Call "
+ "bridge_reply exactly once, as the final action of your turn, with your full answer in "
+ "`content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so
* existing deployments and tests keep the legacy non-blocking spawn semantics.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, guard, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300));
}
/**
* Production constructor with spawn-ready gate enabled. The gate polls {@code agents.status()}
* until the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs));
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
* {@code sleeper} is never called when the gate is disabled ({@code spawnReadyTimeoutMs == 0}).
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param guard subscription-boundary guard (checked before spawning)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (e.g. {@code () -> Thread.sleep(pollMs)}); it
* already encodes the poll interval, so the 8th positional argument
* (poll ms) is accepted for API symmetry but otherwise unused here
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.guard = guard;
}
/**
* {@inheritDoc}
*
* <p>The spawn sequence encodes the subscription boundary: assert the profile's base_url is on
* the allowlist <em>before</em> any herdr call, then build the worker env with
* {@code ANTHROPIC_*}, the parity-neutral git-forge grant, and the bridge MCP + reply charter
* mounted as inline launch flags.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
String baseUrl = cfg.baseUrl();
guard.assertWorker(baseUrl); // hard stop before we spawn anything
Map<String, String> workerEnv = baseEnv(cfg);
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithBridge(cfg));
}
/**
* The launch argv, plus — when {@code worker.mcpUrl} is set — inline {@code --mcp-config} for
* the bridge server and {@code --append-system-prompt} for the {@link #REPLY_CHARTER}. Neither
* touches the profile's config; both are pure command-line flags. This inline-flag mount is
* Claude Code specific — other adapters mount MCP and instructions their own way.
*/
private List<String> argvWithBridge(BridgedConfig.Worker cfg) {
if (!cfg.hasMcp()) {
return cfg.argv();
}
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
List<String> argv = mutableArgv(cfg.argv());
argv.add("--mcp-config");
argv.add(mcpJson);
argv.add("--append-system-prompt");
argv.add(REPLY_CHARTER);
return argv;
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
// --- CB-117 reap predicate (Claude prefix), kept for direct unit testing -------------------
/**
* Whether {@code name} is a Claude Code bridge worker started by a <em>different</em> process
* than {@code currentNonce}. A thin {@code claude}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -0,0 +1,48 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import java.nio.file.Path;
/**
* Provisions an isolated {@code CODEX_HOME} for one Codex peer (CB-528).
*
* <p>Codex reads <em>everything</em> from {@code CODEX_HOME} — its config, credentials, sessions,
* skills, plugins, and state. Pointing a peer at the operator's own {@code ~/.codex} would give it
* the operator's tool surface and let it write into the operator's session history, which is the
* same class of failure CB-525 exists to prevent on the Claude side. So every peer gets its own
* directory, and this is the seam that builds it.
*
* <p>It is an interface rather than a method on the launcher for two reasons: provisioning is
* filesystem work with its own failure modes (a missing credential is the most common, and it
* surfaces as an opaque {@code 401} from Codex rather than a spawn error), and keeping it separate
* lets the launcher be tested without touching a real home directory.
*
* <p>Three things the implementation must put in the home, because Codex has no launch flag for
* any of them:
* <ul>
* <li>the bridge MCP server, as {@code [mcp_servers.*]} in {@code config.toml};</li>
* <li>the reply charter, as {@code AGENTS.md} — Codex has no {@code --append-system-prompt},
* so the standing instruction has to reach it as a file;</li>
* <li>credentials, since a freshly created home has none and Codex fails closed.</li>
* </ul>
*/
public interface CodexHome {
/**
* Build a fresh, isolated home for a peer launching under {@code cfg} and return its path,
* suitable for the {@code CODEX_HOME} environment variable.
*
* @param cfg the profile being launched; supplies the MCP URL and any bearer-token variable
* @return the provisioned directory
* @throws RuntimeException if the home cannot be provisioned — including when no credential is
* available, which must fail loudly here rather than as a 401 later
*/
Path provision(BridgedConfig.Worker cfg);
/**
* Remove a home previously returned by {@link #provision}. Idempotent: releasing an already
* released or never provisioned path is not an error, because teardown races teardown.
*/
void release(Path home);
}
@@ -0,0 +1,192 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>OpenAI's {@code codex}</strong> CLI — the third
* peer kind behind the bridge (after Claude Code and opencode). Like {@link OpenCodeLauncher} it
* extends {@link HerdrPeerLauncher} and reuses every line of shared transport (tab/pane placement,
* the CB-306 readiness gate, unique naming + CB-117 reap, teardown, listing, cwd), overriding only
* the launch seams.
*
* <p>The divergences, all confined to {@link #buildLaunch}:
* <ul>
* <li><strong>No subscription boundary.</strong> Codex authenticates through its own
* {@code CODEX_HOME} credentials and has no {@code ANTHROPIC_BASE_URL} seam, so there is no
* {@link dev.ltms.bridged.guard.SubscriptionGuard} — the guard is a Claude-private concern,
* not part of the SPI. This asymmetry is deliberate and safe for the same reason opencode's is:
* the guard exists to stop a worker borrowing the primary's Anthropic subscription, and a
* codex process has no Anthropic credential path at all, so nothing can leak the subscription.
* The bridge injects no {@code ANTHROPIC_*}/{@code CLAUDE_*} variable and reads none.</li>
* <li><strong>Isolated home.</strong> Codex reads everything from {@code CODEX_HOME}. The
* {@link CodexHome} seam provisions a fresh, isolated directory (config, reply charter,
* credentials — none of which Codex can take as a launch flag) and the worker is pointed at it
* with {@code CODEX_HOME}, so a peer never inherits the operator's own {@code ~/.codex}.</li>
* <li><strong>{@code --approve-for-me}</strong> is mandatory: without it, every MCP tool call from
* the peer returns "user cancelled MCP tool call" and the peer can never call
* {@code bridge_reply}. It routes approvals through automatic review while keeping the sandbox
* on — deliberately not {@code --dangerously-bypass-approvals-and-sandbox}.</li>
* <li><strong>Flags, not env/files.</strong> the model ({@code -m}), the bridge MCP override
* ({@code -c mcp_servers.bridged.url=...}), and the bridge bearer-token variable
* ({@code --bearer-token-env-var}) are all command-line.</li>
* <li><strong>{@code codex} name prefix</strong> so reap matches {@code codex-*} panes and never
* another adapter's.</li>
* </ul>
*/
public final class CodexLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "codex";
/** Provisions the isolated {@code CODEX_HOME} each peer runs against (CB-528). */
private final CodexHome codexHome;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so it
* matches the legacy non-blocking spawn semantics.
*/
public CodexLauncher(AgentControl agents, WorkspaceControl spaces, CodexHome codexHome,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, codexHome, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300));
}
/**
* Production constructor with the spawn-ready gate enabled. Polls {@code agents.status()} until
* the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public CodexLauncher(AgentControl agents, WorkspaceControl spaces, CodexHome codexHome,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, codexHome, profiles, defaultProfile, env,
spawnReadyTimeoutMs, System::currentTimeMillis,
() -> sleepUninterruptibly(spawnReadyPollMs));
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply a
* fake clock ({@code nowMillis}), poll-loop wait ({@code sleeper}), and a stub {@link CodexHome}
* whose returned path they inspect as {@code CODEX_HOME}.
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param codexHome seam that provisions the isolated {@code CODEX_HOME} (CB-528)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (never called when the gate is disabled)
*/
public CodexLauncher(AgentControl agents, WorkspaceControl spaces, CodexHome codexHome,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.codexHome = codexHome;
}
/**
* {@inheritDoc}
*
* <p>Builds the codex launch: no {@code ANTHROPIC_*} and no guard (codex reads its own
* {@code CODEX_HOME} credentials); provision an isolated home via {@link #codexHome} and point
* the worker at it with {@code CODEX_HOME}; carry the parity-neutral git-forge grant; and pass
* the model, bridge MCP override, and bearer-token variable as flags — with mandatory
* {@code --approve-for-me}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
Map<String, String> workerEnv = baseEnv(cfg);
workerEnv.put("CODEX_HOME", codexHome.provision(cfg).toString());
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvFor(cfg));
}
/**
* The launch argv: {@code cfg.argv()} as the base command (the profile may override the
* executable), then mandatory {@code --approve-for-me}, then the optional model / bridge-MCP /
* bearer-token flags.
*
* <p>{@code --approve-for-me} is not optional polish — without it Codex auto-rejects every MCP
* tool call ("user cancelled MCP tool call") and the peer can never answer via
* {@code bridge_reply}. It routes approvals through automatic review while keeping the sandbox
* on; it is deliberately not the escape-hatch bypass flag.
*/
private List<String> argvFor(BridgedConfig.Worker cfg) {
List<String> argv = mutableArgv(cfg.argv());
argv.add("--approve-for-me");
if (cfg.model() != null && !cfg.model().isBlank()) {
argv.add("-m");
argv.add(cfg.model());
}
if (cfg.hasMcp()) {
argv.add("-c");
argv.add("mcp_servers.bridged.url=\"" + cfg.mcpUrl() + "\"");
}
if (cfg.tokenEnv() != null && !cfg.tokenEnv().isBlank()) {
argv.add("--bearer-token-env-var");
argv.add(cfg.tokenEnv());
}
return argv;
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
// --- CB-117 reap predicate (codex prefix), kept for direct unit testing --------------------
/**
* Whether {@code name} is a codex bridge worker started by a <em>different</em> process than
* {@code currentNonce}. A thin {@code codex}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -0,0 +1,262 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.PlacementCandidate;
import dev.ltms.bridged.placement.PlacementContext;
import dev.ltms.bridged.placement.PlacementPolicies;
import dev.ltms.bridged.placement.PlacementPolicy;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.Collections;
import java.util.EnumSet;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Function;
/**
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
* router in front of one {@link HerdrPeerLauncher} per peer {@code kind} (Claude Code, opencode, …).
* It owns no transport of its own; it dispatches each SPI call to the delegate that owns the profile
* involved, and fans the fleet-wide queries (list/reap/caps/profiles) across all delegates.
*
* <p>Routing rules:
* <ul>
* <li><strong>By profile</strong> — {@link #spawn}, {@link #effectiveCwd}, {@link #parityOverlay}
* resolve the profile (a null/blank name → the global {@link #defaultProfile}) and delegate to
* the single adapter that declares it. Profiles partition cleanly across adapters: the
* constructor rejects a name claimed by two.</li>
* <li><strong>By pane id</strong> — {@link #stop} routes to the adapter that spawned that pane
* (recorded at spawn time). A pane the composite never spawned (only real for a caller that
* hand-rolls an id) falls back to the first delegate; teardown is pane-id addressed and
* tab cleanup is single-occupant guarded, so it is safe either way.</li>
* <li><strong>Fleet-wide</strong> — {@link #reapOrphanWorkers} and {@link #capabilities} fan out
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
* shares one herdr connection and so reports the same global agent set.</li>
* </ul>
*
* <p>CB-518: an unqualified spawn is routed through a {@link PlacementPolicy}. The default
* {@code fixed} policy reproduces the historical default-profile behaviour; {@code weighted} uses
* smooth weighted round-robin with {@code maxLoad} gating. If a chosen profile fails with
* {@link PeerUnreachableException}, the composite advances to the next available candidate and
* retries, bounded by the number of candidates.
*/
public final class CompositePeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(CompositePeerLauncher.class);
private final List<HerdrPeerLauncher> delegates;
private final Map<String, HerdrPeerLauncher> byProfile;
private final String defaultProfile;
/** paneId → the delegate that spawned it, so {@link #stop} tears down through the right adapter. */
private final Map<String, HerdrPeerLauncher> spawnedBy = new ConcurrentHashMap<>();
private final Map<String, BridgedConfig.Worker> profileConfigs;
private final PlacementPolicy placementPolicy;
private final Function<String, Integer> liveCount;
/**
* Backward-compatible constructor: fixed placement, no live-counting. Use this for tests and
* simple wiring; it preserves the pre-CB-518 behaviour exactly.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to (may be null)
* @throws IllegalArgumentException if {@code delegates} is empty or two adapters claim one profile
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates, String defaultProfile) {
this(delegates, defaultProfile, Map.of(), PlacementPolicies.fixed(), name -> 0);
}
/**
* Production constructor with a placement policy and live-worker counter.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to under {@code fixed} policy
* @param profileConfigs all configured worker profiles (used for candidate weights/caps)
* @param placementPolicy which policy governs unqualified spawns
* @param liveCount live worker count per profile (must never return {@code null})
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Worker> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount) {
if (delegates.isEmpty()) {
throw new IllegalArgumentException("at least one peer adapter must be configured");
}
this.delegates = List.copyOf(delegates);
this.defaultProfile = defaultProfile;
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
// differ from one JVM run to the next.
this.profileConfigs = Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs));
this.placementPolicy = placementPolicy;
this.liveCount = liveCount;
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
HerdrPeerLauncher prev = index.putIfAbsent(profile, d);
if (prev != null) {
throw new IllegalArgumentException(
"worker profile '" + profile + "' is claimed by two peer adapters");
}
}
}
// Order-preserving for the same reason, and because profiles() is user-visible (bridge_profiles).
this.byProfile = Collections.unmodifiableMap(index);
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (resolved == null) {
// No profile and no default configured — hand to the first delegate so it raises the
// same "no default" error it would on its own; keeps the SPI contract single-sourced.
return delegates.getFirst();
}
HerdrPeerLauncher d = byProfile.get(resolved);
if (d == null) {
throw new IllegalArgumentException("unknown worker profile: " + resolved);
}
return d;
}
@Override
public PeerHandle spawn(SpawnRequest req) {
String requestedProfile = req.profileName();
if (requestedProfile != null && !requestedProfile.isBlank()) {
// An explicit profile bypasses the policy entirely.
HerdrPeerLauncher d = route(requestedProfile);
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
}
List<PlacementCandidate> candidates = candidates();
Set<String> unreachable = new HashSet<>();
PlacementContext ctx = new PlacementContext(defaultProfile, candidates, liveCount, unreachable);
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
for (int attempt = 0; attempt < maxAttempts; attempt++) {
PlacementCandidate chosen;
try {
chosen = placementPolicy.select(ctx);
} catch (RuntimeException e) {
// No candidate left (all at cap or all unreachable). The policy already threw a clear
// message; do not wrap it in a generic PeerUnreachableException.
throw e;
}
HerdrPeerLauncher d = byProfile.get(chosen.profile());
if (d == null) {
// A configured profile with no adapter is a wiring bug; fail fast.
unreachable.add(chosen.profile());
continue;
}
SpawnRequest routedReq = new SpawnRequest(chosen.profile(), req.requestedCwd(), req.callerCwd());
try {
PeerHandle handle = d.spawn(routedReq);
spawnedBy.put(handle.id(), d);
return handle;
} catch (PeerUnreachableException e) {
log.warn("spawn on profile {} unreachable, will retry next candidate if any: {}",
chosen.profile(), e.getMessage());
unreachable.add(chosen.profile());
// Update the context for the next selection so the policy excludes this profile.
ctx = new PlacementContext(defaultProfile, candidates, liveCount, unreachable);
}
}
throw new PeerUnreachableException(
"no reachable worker profile available after trying " + unreachable.size()
+ " candidate(s): " + String.join(", ", unreachable));
}
/** Build the candidate list from the configured profiles, in definition order. */
private List<PlacementCandidate> candidates() {
List<PlacementCandidate> out = new ArrayList<>();
for (Map.Entry<String, BridgedConfig.Worker> e : profileConfigs.entrySet()) {
BridgedConfig.Worker w = e.getValue();
out.add(new PlacementCandidate(e.getKey(), null, w.weight(), w.maxLoad()));
}
return out;
}
@Override
public String effectiveCwd(SpawnRequest req) {
return route(req.profileName()).effectiveCwd(req);
}
@Override
public List<String> parityOverlay(String profileName) {
return route(profileName).parityOverlay(profileName);
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.remove(id);
if (d == null) {
log.debug("stop({}) — no recorded owner, routing to the first adapter (pane-addressed)", id);
d = delegates.getFirst();
}
d.stop(id);
}
@Override
public Set<String> profiles() {
return byProfile.keySet();
}
@Override
public String defaultProfile() {
return defaultProfile;
}
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
Map<String, Agent> byPane = new LinkedHashMap<>();
for (HerdrPeerLauncher d : delegates) {
for (Agent a : d.list()) {
if (a.paneId() != null) {
byPane.putIfAbsent(a.paneId(), a);
}
}
}
return List.copyOf(byPane.values());
}
@Override
public int reapOrphanWorkers() {
int reaped = 0;
for (HerdrPeerLauncher d : delegates) {
reaped += d.reapOrphanWorkers();
}
return reaped;
}
/** The union of every adapter's capabilities — a capability any adapter offers, the fleet offers. */
@Override
public Set<Capability> capabilities() {
EnumSet<Capability> caps = EnumSet.noneOf(Capability.class);
for (HerdrPeerLauncher d : delegates) {
caps.addAll(d.capabilities());
}
return Set.copyOf(caps);
}
}
@@ -0,0 +1,600 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.Collection;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ConcurrentMap;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* Abstract base for {@link PeerLauncher} adapters that materialize a peer as a <em>herdr</em>
* agent (a CLI coding agent running in a herdr tab/pane). It owns everything that is the same
* regardless of <em>which</em> coding agent runs: tab/pane placement, the CB-306 spawn-readiness
* gate, unique naming, CB-117 orphan reap, teardown, {@link #list() listing}, and cwd resolution.
*
* <p>Two seams are peer-specific and supplied by the concrete adapter:
* <ul>
* <li>{@code namePrefix} (constructor arg) — the label prefix ({@code claude}, {@code opencode})
* that drives both unique naming and the orphan-reap pattern, so each adapter reaps only its
* own kind of pane and never another's.</li>
* <li>{@link #buildLaunch(BridgedConfig.Worker)} — the peer-specific env map + argv, including any
* subscription/guard check, MCP mount, and instruction injection. The base never sees how the
* peer is configured; it only places and starts the returned {@link Launch}.</li>
* </ul>
*
* <p>Placement: in the default {@code tab} policy a peer lands in its own tab inside a dedicated
* worker space (found-or-created once, then shared), so peers never split or clutter the user's
* real work spaces. Teardown removes the peer's pane <em>and</em> its now-empty tab, tolerating an
* already-gone peer so a repeated DELETE is harmless.
*/
public abstract class HerdrPeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(HerdrPeerLauncher.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
/**
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
*/
private static final int SHELL_READY_RETRIES = 20;
private final String namePrefix; // label prefix: naming + reap scheme
private final AgentControl agents;
private final WorkspaceControl spaces;
private final Map<String, BridgedConfig.Worker> profiles; // profile name → spawn settings
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
/** Host env lookup (injectable for tests); adapters read it in {@link #buildLaunch}. */
protected final Function<String, String> env;
private final AtomicLong nameSeq = new AtomicLong(); // per-peer counter (also the tab #)
private final long spawnReadyTimeoutMs; // 0 = disable gate (legacy non-blocking spawn)
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
private final Runnable sleeper; // sleep/wait hook (injectable for tests; never real-sleep in unit tests)
// Per-process token mixed into each peer name so a fresh process (nameSeq back at 0) cannot
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
// CB-519: PeerHandle.id() is a host-unique opaque UUID, decoupled from the herdr pane id. The
// routing/registry key is the UUID; the herdr pane id is a launcher-private placement/teardown
// coordinate. This map bridges the two so stop(id) can resolve a host-unique key back to the
// exact pane it must tear down. The pane id is launcher-private (never the routing key) — see
// PeerHandle.id().
private final ConcurrentMap<String, String> paneByAgentId = new ConcurrentHashMap<>();
/**
* @param namePrefix label prefix for this peer kind (drives naming and reap)
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param profiles configured peer profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (never called when the gate is disabled); the poll
* interval is baked into this hook, so the base needs no poll field
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
this.namePrefix = namePrefix;
this.agents = agents;
this.spaces = spaces;
this.profiles = Map.copyOf(profiles);
this.defaultProfile = defaultProfile;
this.env = env;
this.spawnReadyTimeoutMs = spawnReadyTimeoutMs;
this.nowMillis = nowMillis;
this.sleeper = sleeper;
}
// --- adapter seams -------------------------------------------------------------------------
/**
* Build the peer-specific launch for {@code cfg}: the environment map and argv handed to herdr.
* Any subscription/guard check, MCP mount, and instruction injection happen here. The env map
* and argv are adapter-private; the base only places and starts what is returned.
*/
protected abstract Launch buildLaunch(BridgedConfig.Worker cfg);
/** A peer-specific launch: the herdr {@code env} map and {@code argv}. */
protected record Launch(Map<String, String> env, List<String> argv) {
}
// --- profile surface -----------------------------------------------------------------------
/** The configured peer profile names (what {@code spawn(profile)} accepts). */
@Override
public Set<String> profiles() {
return profiles.keySet();
}
/** The parity-overlay file list for {@code profileName} (default list when unset). */
@Override
public List<String> parityOverlay(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
return List.of();
}
BridgedConfig.Worker cfg = profiles.get(name);
return cfg == null ? List.of() : cfg.parityOverlay();
}
/** The profile a no-argument spawn uses, or {@code null} if none is configured. */
@Override
public String defaultProfile() {
return defaultProfile;
}
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
protected Collection<BridgedConfig.Worker> profileConfigs() {
return profiles.values();
}
/** Resolve {@code profileName} (null/blank → default) to its config, or throw with the options. */
protected BridgedConfig.Worker requireProfile(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
throw new IllegalArgumentException("no default worker profile is configured — "
+ "pass a profile; configured: " + profiles.keySet());
}
BridgedConfig.Worker cfg = profiles.get(name);
if (cfg == null) {
throw new IllegalArgumentException("unknown worker profile '" + name
+ "' — configured: " + profiles.keySet());
}
return cfg;
}
// --- spawn ---------------------------------------------------------------------------------
/**
* Spawn a peer. {@code profileName} null/blank → the default profile. The working directory
* (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd}, else the
* profile's configured {@code cwd}, else {@code callerCwd} (the primary's cwd, when the spawn
* came from the primary over MCP), else the daemon's cwd — never assumed to be {@code $HOME}.
* The adapter's {@link #buildLaunch} runs before any herdr call.
*/
protected Agent spawnInternal(String profileName, String requestedCwd, String callerCwd) {
BridgedConfig.Worker cfg = requireProfile(profileName);
Launch launch = buildLaunch(cfg);
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
return cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd);
}
/**
* {@inheritDoc}
*
* <p>Delegates to {@link #spawnInternal} and wraps the resulting herdr {@link Agent} in a
* {@link WorkerHandle} whose {@link PeerHandle#id()} is a fresh <em>host-unique</em> opaque
* UUID (CB-519), deliberately decoupled from the herdr pane id: the id is the registry/routing
* key and must never collide across daemon processes on the same host, while the herdr pane id
* stays a launcher-private placement/teardown coordinate, remembered here so {@link #stop}
* can resolve the host-unique key back to its pane. When {@code spawnReadyTimeoutMs > 0},
* blocks until the peer's herdr status is injectable or the timeout elapses; on timeout the
* pane is closed (no orphan) and a {@link PeerUnreachableException} is thrown.
*/
@Override
public PeerHandle spawn(SpawnRequest req) {
Agent agent = spawnInternal(req.profileName(), req.requestedCwd(), req.callerCwd());
String paneId = agent.paneId();
if (spawnReadyTimeoutMs > 0) {
waitUntilInjectableOrThrow(paneId);
}
// CB-519: the handle id is a host-unique UUID; the herdr pane it maps to stays internal.
String id = UUID.randomUUID().toString();
paneByAgentId.put(id, paneId);
return new WorkerHandle(id, agent.terminalId(), requireProfile(req.profileName()).profile());
}
@Override
public String effectiveCwd(SpawnRequest req) {
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
}
/**
* CB-301: the effective working directory a spawn for {@code profileName} would use, without
* actually spawning.
*/
private String effectiveCwd(String profileName, String requestedCwd, String callerCwd) {
return resolveCwd(requestedCwd, requireProfile(profileName), callerCwd);
}
/**
* CB-112 cwd resolution: spawn arg → profile config → the primary's cwd → the daemon's cwd.
* Never returns {@code null}/blank: {@code "."} (the daemon's own working directory) is the
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
*/
private static String resolveCwd(String requestedCwd, BridgedConfig.Worker cfg, String callerCwd) {
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
}
private static String firstNonBlank(String... values) {
for (String v : values) {
if (v != null && !v.isBlank()) return v;
}
return null;
}
/** Dedicated worker space → own tab (carrying cwd+env) → start the peer into the seed pane. */
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId(), cwd, workerEnv);
log.info("spawning {} profile={} space={} tab={} cwd={}",
namePrefix, cfg.profile(), space.workspaceId(), tab.tab().tabId(), cwd);
Started started;
try {
if (tab.rootPaneId() == null) {
// Protocol 19 starts the agent INTO the seed pane — without one there is nowhere
// to start, and a partial tab would be left behind.
throw new IllegalStateException("tab " + tab.tab().tabId()
+ " had no seed pane in the create response — cannot start a peer in it");
}
started = startUniquelyNamed(cfg, argv, tab.rootPaneId());
} catch (RuntimeException e) {
// The peer never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("failed to close orphaned tab {} after spawn error: {}",
tab.tab().tabId(), cleanup.getMessage());
}
throw e;
}
// The peer is LIVE now, in the seed pane itself (no shell pane to drop — protocol 19).
// Labelling is cosmetic: it must not fail the spawn or orphan the running peer — on error
// we log and still return it so the caller gets its paneId and can tear it down.
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
log.info("{} started pane={} tab={} terminal={}",
namePrefix, started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
return started.agent();
}
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
private void tidy(String what, Runnable step) {
try {
step.run();
} catch (RuntimeException e) {
log.warn("post-start step failed ({}) — peer is running regardless: {}", what, e.getMessage());
}
}
/** Legacy placement: split the currently-focused tab; the peer still starts in {@code cwd}. */
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
namePrefix, cfg.profile(), cwd, argv);
String paneId = spaces.splitPane(cwd, workerEnv);
if (paneId == null) {
throw new IllegalStateException("pane.split returned no pane — cannot start a peer");
}
Agent peer = startUniquelyNamed(cfg, argv, paneId).agent();
log.info("{} started pane={} terminal={}", namePrefix, peer.paneId(), peer.terminalId());
return peer;
}
/** A started peer together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
/**
* Start the peer under a unique herdr agent name. herdr requires each running agent's
* {@code name} to be distinct (a 2nd identical {@code name} fails {@code agent_name_taken}) —
* the exact case that makes multiple peers useful. The name is
* {@code <prefix>-<profile>-<nonce>-<seq>}: {@code seq} distinguishes peers within this process,
* and the per-process {@code nonce} keeps a fresh process (whose {@code seq} restarts at 0) from
* colliding with same-profile peers that outlived a restart. The retry is a belt-and-braces
* backstop for the astronomically unlikely nonce+seq clash; the name is a label only — herdr
* detects kind and status from terminal output, not from it.
*/
private Started startUniquelyNamed(BridgedConfig.Worker cfg, List<String> argv, String paneId) {
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
// argv[0] — the configured executable — is dropped and only the extra args are passed.
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(startAwaitingShellPrompt(name, args, paneId), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("peer name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
/** Start the agent into {@code paneId}, waiting out the seed shell's boot with the sleeper. */
private Agent startAwaitingShellPrompt(String name, List<String> args, String paneId) {
HerdrException busy = null;
for (int attempt = 0; attempt < SHELL_READY_RETRIES; attempt++) {
try {
return agents.start(name, namePrefix, args, paneId);
} catch (HerdrException e) {
if (!"agent_pane_busy".equals(e.code())) throw e;
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
busy = e;
sleeper.run();
}
}
throw busy;
}
// --- discovery + reap ----------------------------------------------------------------------
/** All herdr-tracked agents — discovery for "what peers exist". */
@Override
public List<Agent> list() {
return agents.list();
}
/**
* Reap peer panes left behind by an earlier daemon process (CB-117). herdr keeps a peer's pane
* alive across a daemon restart <em>by design</em>, and that pane's id is held only by its
* spawner — so a peer whose owning process exited before issuing the matching teardown leaks
* with nothing tracking it. On boot we scan herdr for agents whose name matches our
* {@code <prefix>-<profile>-<nonce>-<seq>} scheme with a nonce <em>other</em> than this
* process's {@link #nameNonce}, and tear each one down (its pane and, via {@link #stop}, its
* now-empty dedicated tab). A current-nonce peer is ours and live, so it is left running; a
* user's own session carries no such name and is never touched. A peer from a <em>different</em>
* adapter (different prefix) is likewise never touched. Best-effort: a failed listing, or a
* failure to stop any one peer, is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
@Override
public int reapOrphanWorkers() {
List<Agent> all;
try {
all = agents.list();
} catch (RuntimeException e) {
log.warn("orphan-peer reap skipped — agent.list failed: {}", e.getMessage());
return 0;
}
int reaped = 0;
for (Agent a : all) {
if (!isForeignWorker(namePrefix, a.name(), nameNonce)) continue;
try {
stop(a.paneId());
reaped++;
log.info("reaped orphan {} {} (pane={} tab={}) left by a prior daemon",
namePrefix, a.name(), a.paneId(), a.tabId());
} catch (RuntimeException e) {
log.warn("could not reap orphan {} {} (pane={}): {}",
namePrefix, a.name(), a.paneId(), e.getMessage());
}
}
if (reaped > 0) {
log.info("orphan-peer reap complete — {} stale {} peer(s) removed at startup", reaped, namePrefix);
}
return reaped;
}
/** The {@code <prefix>-<profile>-<nonce>-<seq>} name pattern; group 1 captures the 6-hex nonce. */
static Pattern workerNamePattern(String prefix) {
return Pattern.compile(prefix + "-.*-([0-9a-f]{6})-\\d+");
}
/**
* Whether {@code name} is a peer of kind {@code prefix} started by a <em>different</em> process
* than {@code currentNonce} — the reap predicate (CB-117). True only for the prefix's naming
* scheme with a foreign nonce: a non-peer name, a different adapter's name, or our own live
* nonce is excluded. Pure and package-private so the decision is unit-testable without herdr.
*/
static boolean isForeignWorker(String prefix, String name, String currentNonce) {
String nonce = workerNonce(prefix, name);
return nonce != null && !nonce.equals(currentNonce);
}
/** The 6-hex nonce embedded in a {@code prefix} peer name, or {@code null} if not one. */
static String workerNonce(String prefix, String name) {
if (name == null) return null;
Matcher m = workerNamePattern(prefix).matcher(name);
return m.matches() ? m.group(1) : null;
}
/** This process's peer-name nonce (a label component only; exposed for reaper tests). */
String nameNonce() {
return nameNonce;
}
// --- teardown ------------------------------------------------------------------------------
/**
* Tear a peer down: close the pane, and close its tab <em>only</em> when the peer is that tab's
* sole occupant. The single-pane check is what makes this safe regardless of how the peer was
* placed (or a placement-config change across a restart): a pane-placement peer sitting in one
* of the user's shared tabs has siblings, so its tab is never closed — we only ever remove a
* tab we created to hold one peer.
*
* <p>{@code idOrPane} is the {@link PeerHandle#id()} of a peer this launcher spawned (CB-519's
* host-unique opaque UUID), resolved through {@link #paneByAgentId} to the pane it must tear
* down. An argument that is not one of our ids is treated as a raw herdr pane id — the
* {@link #reapOrphanWorkers() orphan-reap} and spawn-gate-timeout paths, plus any caller that
* passes a pane directly, keep working without an owning id.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed peer) is treated as success; any other failure propagates so a
* genuinely failed teardown is not reported as done.
*/
@Override
public void stop(String idOrPane) {
// Teardown knows only the pane, not which profile spawned it. Attempt tab cleanup when any
// profile uses tab placement (so the bridge may have created a dedicated peer tab); the
// single-occupant check below is what actually protects the user's shared tabs.
String paneId = paneByAgentId.remove(idOrPane);
if (paneId == null) {
paneId = idOrPane; // raw-pane fallback (reap, gate timeout, pane-addressed callers)
}
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) throw e;
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
spaces.closeTab(loc.tabId());
} else if (loc != null) {
log.debug("not closing tab {} — it holds {} panes (not a dedicated peer tab)",
loc.tabId(), loc.tabPaneCount());
}
}
/** Whether any configured profile places peers in their own tab (so tabs may need cleanup). */
private boolean usesTabPlacement() {
return profiles.values().stream().anyMatch(BridgedConfig.Worker::tabPlacement);
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
// --- spawn-readiness gate (CB-306) ---------------------------------------------------------
/**
* Poll {@link AgentControl#status} until the pane reports an injectable state or the configured
* timeout elapses. On timeout, close the pane (self-reap) and throw.
*/
private void waitUntilInjectableOrThrow(String paneId) {
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
while (nowMillis.getAsLong() < deadline) {
if (agents.status(paneId).injectable()) {
log.debug("peer pane={} reached injectable state", paneId);
return;
}
sleeper.run();
}
log.warn("peer pane={} did not become injectable within {}ms — closing", paneId, spawnReadyTimeoutMs);
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " did not reach injectable state within "
+ spawnReadyTimeoutMs + "ms");
}
/** A concrete {@link PeerHandle} wrapping herdr agent coordinates and the profile that spawned it. */
private record WorkerHandle(String id, String terminalId, String profile) implements PeerHandle {
}
// --- shared helpers ------------------------------------------------------------------------
/** Put {@code k → v} only when {@code v} is present (non-null, non-blank). */
protected static void putIfPresent(Map<String, String> m, String k, String v) {
if (v != null && !v.isBlank()) {
m.put(k, v);
}
}
/** Host env lookup that tolerates an unconfigured (null/blank) var name — returns null then. */
protected String resolveEnv(String name) {
return (name == null || name.isBlank()) ? null : env.apply(name);
}
/**
* The parity-neutral git-forge token grant (CB-302): when {@code cfg} opts in via
* {@code gitTokenEnv} and the token resolves, inject {@code GITEA_TOKEN} plus its paired
* {@code GITEA_HOST}. Push over SSH is unaffected; the only incremental grant is PR-create.
* Peer-neutral, so every herdr adapter reuses it unchanged.
*/
protected void applyGitToken(Map<String, String> workerEnv, BridgedConfig.Worker cfg) {
if (!cfg.hasGitToken()) {
return;
}
String gitToken = resolveEnv(cfg.gitTokenEnv());
if (gitToken != null) {
workerEnv.put("GITEA_TOKEN", gitToken);
putIfPresent(workerEnv, "GITEA_HOST", resolveEnv(cfg.gitHostEnv()));
}
}
/** A fresh mutable env map — the conventional starting point for {@link #buildLaunch}. */
/**
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
* {@code env:} entries.
*
* <p>Why this exists: bridged passes herdr an explicit env map, and herdr merges it into
* <em>its own</em> process environment. So before this, a worker inherited whatever PATH the
* herdr server happened to be started with — on this host, one from weeks earlier with no JDK
* and no Maven, which left workers unable to run the build they were being asked to run. The
* worker's toolchain must follow from configuration, not from how a long-lived daemon was
* launched.
*
* <p>Adapter-specific variables are layered on top of this by {@code buildLaunch} and therefore
* win. That ordering is deliberate and load-bearing: it stops a profile's {@code env:} from
* overriding {@code ANTHROPIC_BASE_URL} and slipping past {@link
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
* {@code baseUrl} and nothing else.
*/
protected Map<String, String> baseEnv(BridgedConfig.Worker cfg) {
Map<String, String> workerEnv = new LinkedHashMap<>();
String path = env.apply("PATH");
if (path != null && !path.isBlank()) {
workerEnv.put("PATH", path);
}
if (cfg != null && cfg.env() != null) {
workerEnv.putAll(cfg.env());
}
return workerEnv;
}
/** Defensive copy of {@code argv} plus room to append launch flags. */
protected static List<String> mutableArgv(List<String> argv) {
return new ArrayList<>(argv);
}
/**
* Uninterruptible sleep — the production {@link #sleeper}. Tests supply their own no-op /
* fast-faking sleeper so they never real-sleep.
*/
protected static void sleepUninterruptibly(long ms) {
try {
Thread.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
// preserve the interrupt flag but continue — poll loops should not be aborted by an
// interrupt that was not meant for them.
}
}
}
@@ -0,0 +1,316 @@
package dev.ltms.bridged.worker;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.node.ObjectNode;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>opencode</strong> — an open-source,
* provider-agnostic terminal coding agent. Its whole reason for existing is to prove the
* {@code PeerLauncher} SPI is genuinely provider-neutral: opencode shares none of Claude Code's
* private launch seams, yet reuses every line of shared transport in the base (tab/pane placement,
* the CB-306 readiness gate, unique naming + CB-117 reap, teardown, listing, cwd).
*
* <p>The divergences from {@link ClaudeCodeLauncher}, all confined to {@link #buildLaunch}:
* <ul>
* <li><strong>No subscription boundary.</strong> opencode carries no {@code ANTHROPIC_BASE_URL}
* and there is no {@link dev.ltms.bridged.guard.SubscriptionGuard} — the guard is a
* Claude-private concern, not part of the SPI. opencode reads the operator's own provider
* credentials from its global {@code auth.json}; the bridge injects none.</li>
* <li><strong>File-based MCP mount + instructions.</strong> opencode has no inline
* {@code --mcp-config}/{@code --append-system-prompt}. Instead the bridge writes an ephemeral
* {@code opencode.json} that declares the bridge as a {@code remote} MCP server and lists a
* reply-charter file under {@code instructions}, then points the worker at it with
* {@code OPENCODE_CONFIG}. This is the one place the launcher touches disk — Claude never did.</li>
* <li><strong>Model as a flag.</strong> the {@code provider/model} selector is passed as
* {@code -m}, not an env var.</li>
* <li><strong>{@code opencode} name prefix</strong> so reap matches {@code opencode-*} panes and
* never another adapter's.</li>
* </ul>
*/
public final class OpenCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "opencode";
/** Writer for the generated {@code opencode.json}. */
private static final ObjectMapper JSON = new ObjectMapper();
/**
* Standing instruction written to the charter file and mounted via the config's
* {@code instructions} so the worker returns its result through {@code bridge_reply}. Kept on
* disk (not a launch flag) because opencode's {@code instructions} takes file paths, not inline
* text — the file is regenerated per spawn and never touches the worker's own profile.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet, running under opencode. "
+ "Every message you receive arrives through the bridge, and the ONLY channel back to the "
+ "sender is the bridge_reply MCP tool. Text you write in your terminal is NOT sent "
+ "anywhere — the sender cannot see your screen, so an in-terminal answer is silently "
+ "discarded. Therefore you MUST end EVERY turn by calling bridge_reply with `content` set "
+ "to your complete response. This holds for every message without exception — tasks, "
+ "questions, clarifications, acknowledgements, and ordinary back-and-forth conversation. "
+ "Call bridge_reply exactly once, as the final action of your turn, with your full answer "
+ "in `content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/** Root under which per-spawn opencode config dirs are created (injectable for tests). */
private final Path configRoot;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so it
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300),
defaultConfigRoot());
}
/**
* Production constructor with the spawn-ready gate enabled. Polls {@code agents.status()} until
* the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot());
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply a
* fake clock ({@code nowMillis}), poll-loop wait ({@code sleeper}), and a temp {@code configRoot}
* they can inspect the generated {@code opencode.json}/charter under.
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (encodes the poll interval; never called when the
* gate is disabled)
* @param configRoot existing directory under which per-spawn config dirs are created
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper, Path configRoot) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.configRoot = configRoot;
}
private static Path defaultConfigRoot() {
return Path.of(System.getProperty("java.io.tmpdir"));
}
/**
* {@inheritDoc}
*
* <p>Builds the opencode launch: no {@code ANTHROPIC_*} and no guard (opencode reads its own
* provider credentials); when the profile mounts the bridge MCP, generate an ephemeral
* {@code opencode.json} (remote MCP server + reply-charter instructions) and point the worker at
* it via {@code OPENCODE_CONFIG}; carry the parity-neutral git-forge grant; and select the model
* with {@code -m}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
Map<String, String> workerEnv = baseEnv(cfg);
// A config file is needed for the bridge MCP mount, for a pinned endpoint (CB-508), or both.
if (cfg.hasMcp() || hasCustomProvider(cfg)) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg).toString());
}
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithModel(cfg));
}
/**
* True when this profile pins its own OpenAI-compatible endpoint (CB-508) rather than using
* whatever provider opencode resolves by default.
*
* <p>Note this reuses {@code baseUrl}, the same field the Claude adapter injects as
* {@code ANTHROPIC_BASE_URL} — but it does <em>not</em> go through {@code SubscriptionGuard}.
* That asymmetry is deliberate and safe: the guard exists to stop a worker borrowing the
* primary's Anthropic subscription, and an opencode process has no Anthropic credential path
* at all. Pointing it at a local vLLM cannot leak the subscription.
*/
private static boolean hasCustomProvider(BridgedConfig.Worker cfg) {
return cfg.baseUrl() != null && !cfg.baseUrl().isBlank();
}
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
private List<String> argvWithModel(BridgedConfig.Worker cfg) {
List<String> argv = mutableArgv(cfg.argv());
if (cfg.model() != null && !cfg.model().isBlank()) {
argv.add("-m");
argv.add(cfg.model());
}
return argv;
}
/**
* Write an ephemeral {@code opencode.json} (and the reply-charter file it references) into a
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
*/
private Path writeConfig(BridgedConfig.Worker cfg) {
try {
Path dir = Files.createTempDirectory(configRoot, "bridged-opencode-");
dir.toFile().deleteOnExit();
ObjectNode root = JSON.createObjectNode();
root.put("$schema", "https://opencode.ai/config.json");
if (cfg.hasMcp()) {
Path charter = dir.resolve("reply-charter.md");
Files.writeString(charter, REPLY_CHARTER);
charter.toFile().deleteOnExit();
ObjectNode bridge = root.putObject("mcp").putObject("bridge");
bridge.put("type", "remote");
bridge.put("url", cfg.mcpUrl());
bridge.put("enabled", true);
root.putArray("instructions").add(charter.toAbsolutePath().toString());
}
if (hasCustomProvider(cfg)) {
addCustomProvider(root, cfg);
}
Path cfgFile = dir.resolve("opencode.json");
// Built with Jackson rather than string concatenation: the provider block is nested and
// carries operator-supplied values (URL, model id, api key), so escaping must be real.
Files.writeString(cfgFile, JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root));
cfgFile.toFile().deleteOnExit();
return cfgFile;
} catch (IOException e) {
throw new UncheckedIOException(
"cannot write opencode config for profile " + cfg.profile(), e);
}
}
/**
* Declare a custom OpenAI-compatible provider so the worker talks to a pinned endpoint (a local
* vLLM, say) instead of opencode's default gateway (CB-508).
*
* <p>The provider id comes from the {@code provider/model} selector in {@code model:}, so one
* field drives both the declaration and the {@code -m} flag and they cannot drift apart.
*/
private void addCustomProvider(ObjectNode root, BridgedConfig.Worker cfg) {
String[] parts = splitModelSelector(cfg);
String providerId = parts[0];
String modelId = parts[1];
ObjectNode provider = root.putObject("provider").putObject(providerId);
provider.put("npm", "@ai-sdk/openai-compatible");
provider.put("name", providerId + " (bridged)");
ObjectNode options = provider.putObject("options");
options.put("baseURL", openAiBaseUrl(cfg.baseUrl()));
// vLLM and friends usually ignore the key, but the AI SDK still requires a non-empty one.
String token = resolveEnv(cfg.tokenEnv());
options.put("apiKey", (token == null || token.isBlank()) ? "bridged-local-noauth" : token);
provider.putObject("models").putObject(modelId).put("name", modelId);
}
/**
* Split {@code model:} into its {@code provider} and {@code model} halves. A pinned endpoint
* needs both, so a bare model name is rejected loudly rather than silently falling back to the
* default gateway — a worker quietly talking to the wrong endpoint is the failure this avoids.
*/
private static String[] splitModelSelector(BridgedConfig.Worker cfg) {
String model = cfg.model();
int slash = model == null ? -1 : model.indexOf('/');
if (model == null || model.isBlank() || slash <= 0 || slash == model.length() - 1) {
throw new IllegalArgumentException(
"profile " + cfg.profile() + " sets baseUrl (a pinned opencode endpoint) so"
+ " model: must be \"<provider>/<model>\", e.g."
+ " \"local-vllm/deepseek-v4-flash\"; got "
+ (model == null ? "null" : '"' + model + '"'));
}
return new String[]{model.substring(0, slash), model.substring(slash + 1)};
}
/**
* The OpenAI-compatible base URL for {@code baseUrl}. A bare {@code host:port} gets {@code /v1}
* appended (where these servers put the API); a URL that already carries a path is taken as-is,
* so an endpoint mounted somewhere unusual is still reachable.
*/
private static String openAiBaseUrl(String baseUrl) {
String trimmed = baseUrl.trim();
while (trimmed.endsWith("/")) {
trimmed = trimmed.substring(0, trimmed.length() - 1);
}
int schemeEnd = trimmed.indexOf("://");
String afterScheme = schemeEnd < 0 ? trimmed : trimmed.substring(schemeEnd + 3);
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
// --- CB-117 reap predicate (opencode prefix), kept for direct unit testing -----------------
/**
* Whether {@code name} is an opencode bridge worker started by a <em>different</em> process than
* {@code currentNonce}. A thin {@code opencode}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -1,788 +0,0 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.ConfigWatcher;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.LeadTabScanner;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.inject.StatusPoller;
import dev.ltms.fleet.inject.TurnListener;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.health.FleetHealthMonitor;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.mcp.LsofPeerPidLookup;
import dev.ltms.fleet.mcp.LsofProcessCwdLookup;
import dev.ltms.fleet.msg.AmqpReplyInbox;
import dev.ltms.fleet.msg.InMemoryReplyInbox;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
import dev.ltms.fleet.msg.ReplyPushLoop;
import dev.ltms.fleet.rest.FleetApp;
import dev.ltms.fleet.session.GitWorktrees;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.session.SessionReaper;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.member.CompositePeerLauncher;
import dev.ltms.fleet.member.HerdrPeerLauncher;
import dev.ltms.fleet.member.OpenCodeLauncher;
import dev.ltms.fleet.placement.BackendQuarantine;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.regex.Pattern;
import java.util.stream.Collectors;
/**
* {@code bridged} entry point. Wires the real herdr socket client to the REST app and
* starts listening. Before anything else it asserts its own environment is clean —
* {@code bridged} is not a Claude process and must never carry a base_url.
*/
public final class Fleetd {
private static final Logger log = LoggerFactory.getLogger(Fleetd.class);
/** CB-504: how long to wait at startup for herdr's socket before serving degraded. */
private static final long HERDR_WAIT_SECONDS = 30;
private static final long HERDR_WAIT_POLL_MILLIS = 500;
/**
* CB-632: prefer {@code fleetd.yaml} in {@code dir}; fall back to {@code bridged.yaml} when
* the new name is not there. The operator's live file is still named {@code bridged.yaml},
* so the old name keeps working until that file moves.
*/
static Path chooseDefaultConfigFile(Path dir) {
Path fleetd = dir.resolve("fleetd.yaml");
if (Files.exists(fleetd)) {
return fleetd;
}
return dir.resolve("bridged.yaml");
}
static void main(String[] args) {
Path configPath = args.length > 0 ? Path.of(args[0]) : chooseDefaultConfigFile(Path.of(""));
// CB-632: the config file is being renamed bridged.yaml -> fleetd.yaml. Name the file we
// actually loaded, whichever of the two names it carries.
log.info("Using configuration file {}", configPath);
FleetConfig cfg = FleetConfig.load(configPath);
// CB-594: report which secret env vars the config actually needs, by name, before anything
// else can fail on a silently-empty one. A daemon started without a login shell (launchd)
// boots fine either way — this is the only thing that says so out loud.
reportRequiredSecrets(cfg);
// CB-596: an absent (or empty) memberCredentials: block blocks NOTHING — no credential
// name is hardcoded any more to fall back on. Say so loudly, the same way a missing
// secret is reported above, so upgrading past this commit never silently drops CB-592's
// protection.
reportMemberCredentialsGap(cfg);
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
// contract; adding a reader here does not make a key reloadable by itself.
ConfigRef config = new ConfigRef(configPath, cfg);
// The primary/host env that launched bridged must not be tainted.
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
guard.assertPrimaryClean(System.getenv());
// CB-501: refuse to start if the bind is wider than the auth mode can defend. Under
// loopback-trust, "not a known worker" means "the primary" — sound only because the OS
// refuses remote connections to a loopback socket. This throws rather than warns so the
// dangerous configuration cannot be reached by ignoring a log line.
cfg.validateAuthExposure();
cfg.validateLeadTabPrefixes();
// CB-542: a subscription:true profile whose env: reseats ANTHROPIC_BASE_URL/AUTH_TOKEN would
// reach an unguarded endpoint (the launcher skips SubscriptionGuard for it). Refuse at load.
cfg.validateSubscriptionProfiles();
cfg.validateCharters();
// CB-548: every architect slot must name a configured workers: profile — the strong-model
// backend the future spawn lifecycle would read. A stale reference dies here, not later.
cfg.validateMembers();
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
: UnixSocketHerdrClient.defaultSocketPath();
UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect(socket, new com.fasterxml.jackson.databind.ObjectMapper());
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
Map<String, FleetConfig.Profile> claudeProfiles = new LinkedHashMap<>();
Map<String, FleetConfig.Profile> opencodeProfiles = new LinkedHashMap<>();
cfg.profiles().forEach((name, w) -> {
if (w.isOpenCode()) {
opencodeProfiles.put(name, w);
} else {
claudeProfiles.put(name, w);
}
});
List<HerdrPeerLauncher> adapters = new ArrayList<>();
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(),
() -> config.get().memberCredentials()));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(agents, spaces,
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(),
() -> config.get().memberCredentials()));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
// The cooldown is deferred (see FleetConfig#quarantineCooldownSeconds): it is read once
// here, at startup, and a config reload only changes it for a daemon restart.
BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
PeerLauncher workers = new CompositePeerLauncher(
adapters,
cfg.effectiveDefaultProfile(),
config,
profileName -> liveCountRef.get().apply(profileName),
quarantine);
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
// with /healthz reporting "degraded" is strictly more useful than exiting.
boolean herdrUp = awaitHerdr(herdr);
if (herdrUp) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
// with the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
} else {
log.warn("herdr did not answer within {}s — starting anyway; /healthz will report "
+ "degraded until it comes up. Orphaned worker panes (if any) were NOT reaped.",
HERDR_WAIT_SECONDS);
}
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
int contextCap = 0;
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
&& cfg.lifecycle().contextCap() > 0) {
contextCap = cfg.lifecycle().contextCap();
}
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()),
System::nanoTime, contextCap, clearAfterTurn);
liveCountRef.set(profileName -> (int) sessions.roster().stream()
.filter(s -> profileName.equals(s.profile()))
.count());
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
final SessionReaper reaper;
if (cfg.lifecycle() != null
&& cfg.lifecycle().idleTtlSeconds() != null
&& cfg.lifecycle().idleTtlSeconds() > 0) {
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
reaper.start();
} else {
reaper = null;
}
// CB-530: every pane the config names as a lead, merged from `leaders:` and the legacy
// singular pin. PrimaryRegistry below still tracks ONE terminal — it addresses the push
// loop's nudges, which need a single destination — so it keeps the legacy pin.
Map<String, String> leadTerminals = cfg.leaderTerminals();
if (leadTerminals.size() > 1) {
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
}
// CB-531: on top of the legacy primary.terminal pin, discover leads by the tab labels the
// operator writes. CB-557 moved the settings onto the lead they describe, so scanning is on
// whenever a `fleet.leaders:` entry exists — with no leads configured the supplier is a
// constant and never touches herdr, exactly as a missing `leadScan:` block used to behave.
// CB-579: each lead now names its own exact `tab:` label, so one scanner discovers every
// configured lead regardless of how differently their tabs are labelled — the old
// single-shared-tabPrefix limitation (and its warning) is gone.
final Supplier<Map<String, String>> leads;
var leaders = cfg.fleet().leaders();
if (!leaders.isEmpty()) {
Set<String> memberSpaces = cfg.profiles().values().stream()
.map(FleetConfig.Profile::workspace)
.filter(Objects::nonNull)
.collect(Collectors.toSet());
Map<String, String> tabToName = new LinkedHashMap<>();
leaders.forEach((name, leader) -> {
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
tabToName.put(leader.tab(), name);
}
});
// One shared rescan cadence: still taken from the first entry, as before — it is an
// operational cadence, not identity, so there is no correctness reason to give every
// lead its own scanner.
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
leads = new LeadTabScanner(herdr, tabToName, memberSpaces,
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), System::nanoTime);
log.info("lead scan: tabs {} host a lead (rescan every {}s, member spaces {} excluded)",
tabToName.keySet(), scanIntervalSeconds, memberSpaces);
} else {
leads = () -> leadTerminals;
}
// CB-558: start any declared lead that is not already running. After the scanner is built,
// because both read the same tab labels and the ordering makes that dependency visible; and
// only when herdr answered, because the launcher's whole safety property is that it can
// count live leads first — it must never guess and risk a second orchestrator.
if (herdrUp && !leaders.isEmpty()) {
int launched = new LeadLauncher(agents, spaces, cfg).ensureLeads();
if (launched > 0) {
log.info("lead auto-launch: {} lead(s) started", launched);
}
}
// CB-548: config-declared architect slots. Config supplies only the stable name → profile
// map; the terminal → slot binding is owned by the registry and is empty at startup, so no
// pane resolves to an architect until the later spawn lifecycle binds one. The registry is
// what CallerResolver resolves against and what that lifecycle will read profiles from;
// nothing here spawns a slot.
MemberRegistry members = new MemberRegistry(cfg.fleet());
sessions.setMemberLifecycle(members);
if (!members.slots().isEmpty()) {
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
+ "spawn lifecycle binds a live terminal to it)",
members.slots().size(), members.slots().keySet());
}
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
Rendezvous rendezvous = new Rendezvous();
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
// Compiled once at startup, keyed by profile name; a profile with no exhaustedPattern is
// simply absent here, so its workers keep today's completion-fallback behaviour unchanged.
Map<String, Pattern> exhaustedPatternsByProfile = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (profile.hasExhaustedPattern()) {
exhaustedPatternsByProfile.put(name, Pattern.compile(profile.exhaustedPattern()));
}
});
ExhaustedPatternLookup exhaustedPatterns = target -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(session -> exhaustedPatternsByProfile.get(session.profile()))
.orElse(null);
log.info("backend-exhausted classification (CB-578 stage A): {}",
CompletionResolver.coverage(cfg.profiles().keySet(), exhaustedPatternsByProfile.keySet()));
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
// the profile config live off `config`, so a credentialId edit is hot: no restart needed.
ExhaustionSink exhaustionSink = (target, reason) -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(MemberSession::profile)
.map(profileName -> config.get().profiles().get(profileName))
.ifPresent(profile -> {
String credentialId = profile.effectiveCredentialId();
quarantine.quarantine(credentialId);
log.warn("credential '{}' quarantined for {}s (profile '{}' classified "
+ "BACKEND_EXHAUSTED): {}", credentialId,
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
});
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns, exhaustionSink);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
MemberPresence presence = sessions.asPresence();
TurnListener turnListener = new TurnListener() {
@Override
public void onTurnComplete(String target) {
completion.onTurnComplete(target);
sessions.onTurnComplete(target);
}
@Override
public boolean hasPostTurnAction(String target) {
return sessions.hasPostTurnAction(target);
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
completion.resolveBeforePostAction(target);
return sessions.onTurnCompleteWithPostAction(target);
}
@Override
public void onDelivered(String target, dev.ltms.fleet.msg.TurnToken token) {
completion.onDelivered(target, token);
sessions.onDelivered(target, token);
}
@Override
public void onTurnFailed(String target) {
completion.onTurnFailed(target);
sessions.onTurnFailed(target);
}
@Override
public void onTurnFailed(String target, String reason) {
completion.onTurnFailed(target, reason);
sessions.onTurnFailed(target);
}
};
Injector injector = new Injector(agents, turnListener, deliverableTo(presence, leads),
presence::forget);
StatusPoller poller = new StatusPoller(agents, injector, Injector.POLL_INTERVAL_MILLIS);
poller.start();
// CB-307: reply inbox. A broker: block selects the AMQP-backed durable adapter; absent (or
// unusable), bridged stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
// connection, so keep the reference to close it in the ordered shutdown hook.
final ReplyInbox replyInbox = selectReplyInbox(cfg.broker(), System.getenv(), AmqpReplyInbox::open);
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
// The pin also feeds CallerResolver below: a primary running inside a herdr pane would
// otherwise resolve as a worker and be refused every orchestration tool.
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
// CB-532: `primary.terminal` is superseded and no longer needed for either of its jobs —
// identity comes from `leaders:`/`leadScan:`, and reply nudges now follow the delegating
// lead. Say so once at startup rather than leaving a redundant pin to look load-bearing.
if (pinnedPrimaryTerminal != null && !pinnedPrimaryTerminal.isBlank()) {
log.warn("primary.terminal is DEPRECATED (CB-532) and can be deleted: identity now comes "
+ "from leaders:/leadScan:, and reply nudges follow the lead that delegated. "
+ "It still works, and is still the fallback nudge destination when a restart "
+ "has lost the delegation map. Its pushReminders/pushBackoffMs stay valid.");
}
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
// open fleet_send. Uses its own lightweight scheduled executor, separate from the injector.
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
var pushScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-push-").unstarted(r));
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
// CB-512: the push loop takes it too, so nudge outcomes (delivered|exhausted) are counted.
Metrics metrics = FleetMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, agents, replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
// It has its own single-thread scheduler and holds its own scheduler shutdown via close().
final LeadHeartbeatLoop heartbeat;
var heartbeatScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-heartbeat-").unstarted(r));
if (cfg.leadHeartbeat() != null) {
var hb = cfg.leadHeartbeat();
heartbeat = new LeadHeartbeatLoop(primaryRegistry, agents, replyInbox, sessions::roster,
pushLoop, heartbeatScheduler, System::nanoTime,
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
metrics);
heartbeat.start();
} else {
heartbeat = null;
heartbeatScheduler.shutdownNow();
}
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
final FleetHealthMonitor healthMonitor;
var healthScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-health-").unstarted(r));
if (cfg.health() != null && cfg.health().isEnabled()) {
// CB-580: a member found GONE/NEVER_READY must fail whatever ticket is waiting on it,
// through the same idempotent target-wide operation CB-516 already uses on release.
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
System::nanoTime, cfg.health().intervalOrDefault(), messages::abandon);
String coverage = FleetHealthMonitor.coverage(true,
cfg.health().notifications() != null && cfg.health().notifications().configured());
if ("detection-only".equals(coverage)) {
log.warn("fleet health: {} (no notification sink configured)", coverage);
} else {
log.info("fleet health: {}", coverage);
}
healthMonitor.start();
} else {
healthMonitor = null;
healthScheduler.shutdownNow();
}
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
sessions.onAcquire(replyInbox::own);
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
sessions.onRelease(detail -> {
// CB-578 stage C, acceptance criterion 10: a failed ticket's detail should tell a lead
// where to re-dispatch onto the same tree, not just that the worker vanished.
String reason = "the worker session was released before it replied";
if (detail.worktreePath() != null) {
reason += "; worktree=" + detail.worktreePath() + " branch=" + detail.branch()
+ " snapshot=" + (detail.snapshotRef() != null ? detail.snapshotRef() : "none");
}
// CB-584 (issue #65 criterion 5): also name the agent session, so a lead can resume the
// member's conversation instead of only re-dispatching a fresh one onto the same files.
if (detail.agentSessionId() != null) {
reason += " agentSessionId=" + detail.agentSessionId();
}
messages.abandon(detail.terminalId(), reason);
replyInbox.release(detail.terminalId());
primaryRegistry.forgetDelegation(detail.terminalId()); // CB-532: don't leak the lead binding
});
// MCP server face (CB-105): fleet_send/fleet_reply/fleet_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
final CallerResolver callers;
if (cfg.auth().tokenMode()) {
String token = System.getenv(cfg.auth().tokenEnv());
if (token == null || token.isBlank()) {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting bridged");
}
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics, new FleetMcp.CapacitySource(profile -> liveCountRef.get().apply(profile),
profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.maxLoad();
}, () -> config.get().profiles().keySet(), System::nanoTime),
new FleetMcp.HealthCoverageSource(() -> {
var health = config.get().health();
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
health != null && health.notifications() != null && health.notifications().configured());
}),
new FleetMcp.QuarantineSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, quarantine));
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed, so an
// upgraded daemon behaves exactly as before — the file is read once at boot and never again.
final ConfigWatcher configWatcher;
if (cfg.configReload() != null && cfg.configReload().isEnabled()) {
configWatcher = new ConfigWatcher(config, cfg.configReload().intervalSeconds());
configWatcher.start();
} else {
configWatcher = null;
}
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
// last. This replaces the earlier independent hooks that could race and close herdr early.
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
poller.stop();
messages.close();
pushLoop.close();
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
if (healthMonitor != null) healthMonitor.stop();
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
mcp.close();
if (reaper != null) reaper.stop();
// Release the broker connection last among message resources (no-op for the in-memory inbox).
if (replyInbox instanceof AutoCloseable closeable) {
try {
closeable.close();
} catch (Exception e) {
log.debug("reply inbox close: {}", e.toString());
}
}
herdr.close();
}));
Javalin app = new FleetApp(herdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics).build();
app.start(cfg.bind().host(), cfg.bind().port());
log.info("bridged listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
}
/**
* The {@link Injector}'s readiness gate (CB-534): a target is deliverable if it is a spawned
* member whose agent has connected the bridge MCP, <em>or</em> a lead.
*
* <p>The gate exists for one reason — to hold a delivery out of a <em>spawned</em> member's boot
* window, where herdr already reports {@code idle} but the TUI would drop an injected paste. That
* hazard is a property of spawning. A lead is never spawned: the operator started it and named it
* (or labelled its tab) only once it was up, so there is no boot window to guard.
*
* <p>A lead is also never enrolled in {@link MemberPresence} — {@code FleetMcp} marks presence
* for every spawned member (worker and architect), deliberately, since that map doubles as the
* member roster's availability signal and a lead counted there would show up as an available
* member. So without the second disjunct a lead is permanently un-deliverable: every
* lead→lead send sat on the gate for {@code READINESS_GRACE_POLLS} (~60s) and then failed
* having never been typed into the pane.
*
* <p>The lead set is read through the supplier on each call rather than snapshotted, so a lead
* discovered by {@code leadScan} after startup becomes deliverable without a restart.
*/
static Predicate<String> deliverableTo(MemberPresence presence, Supplier<Map<String, String>> leads) {
return target -> presence.isPresent(target) || leads.get().containsKey(target);
}
/** Injection seam for {@link #selectReplyInbox}: production binds {@link AmqpReplyInbox#open}. */
@FunctionalInterface
interface AmqpOpener {
ReplyInbox open(String uri, int prefetch);
}
/**
* CB-151/152: pick the reply inbox. A usable broker — a literal {@code uri}, or a {@code
* uriEnv} whose variable resolves (both read from {@code env}) — selects the durable AMQP inbox.
* Everything else falls back to the in-memory inbox: no broker block, a blank {@code uri}, a
* {@code uriEnv} whose variable is unset or blank, or a broker unreachable at boot. The two
* lossy paths warn <em>loudly</em> — never silently — because what is lost is durable,
* cross-restart reply delivery. Package-private and env-injected so the selection is testable
* without a real broker or a mutable process environment.
*/
static ReplyInbox selectReplyInbox(FleetConfig.Broker broker, Map<String, String> env, AmqpOpener amqp) {
if (broker == null) {
log.info("reply inbox: in-memory (soft-state)");
return new InMemoryReplyInbox();
}
if (broker.hasUriEnv()) {
// uriEnv is authoritative whenever set (CB-151): the operator moved off clear text, so
// it must not quietly fall back onto a stale literal uri.
if (broker.uri() != null && !broker.uri().isBlank()) {
log.info("broker.uri is ignored because broker.uriEnv={} is set", broker.uriEnv());
}
String effectiveUri = broker.effectiveUri(env);
if (effectiveUri == null) {
log.warn("broker.uriEnv={} is unset or blank — durable AMQP reply inbox DISABLED. "
+ "Replies are soft-state and will not survive a restart. Set {} in the "
+ "daemon's environment (see scripts/redeploy-bridged.sh) and restart to "
+ "use the durable broker inbox.",
broker.uriEnv(), broker.uriEnv());
log.info("reply inbox: in-memory (soft-state)");
return new InMemoryReplyInbox();
}
log.info("reply inbox: AMQP broker (durable) via env var {} (prefetch={})",
broker.uriEnv(), broker.prefetchOrDefault());
return openAmqpOrFallback(effectiveUri, broker.prefetchOrDefault(), "uriEnv " + broker.uriEnv(), amqp);
}
// No uriEnv: the literal uri path (existing behaviour).
if (broker.effectiveUri(env) == null) {
log.info("reply inbox: in-memory (soft-state)");
return new InMemoryReplyInbox();
}
log.info("reply inbox: AMQP broker (durable) (prefetch={})", broker.prefetchOrDefault());
return openAmqpOrFallback(broker.uri(), broker.prefetchOrDefault(), "uri", amqp);
}
/**
* Open the AMQP inbox, falling back to the in-memory inbox for this process lifetime if the
* broker cannot be reached at boot (CB-152). Not silent: the warning says durable delivery is
* off, replies are soft-state and will not survive a restart, plus the source that failed and
* the URI <em>with credentials stripped</em>. Never retries in the background — a broker that
* drops <em>after</em> startup already self-heals via the connection factory's automatic
* recovery; only the boot path is changed here.
*/
private static ReplyInbox openAmqpOrFallback(String effectiveUri, int prefetch, String source,
AmqpOpener amqp) {
try {
return amqp.open(effectiveUri, prefetch);
} catch (IllegalStateException e) {
log.warn("cannot reach AMQP broker ({}, {}) — falling back to the in-memory reply inbox "
+ "for this process lifetime. Durable, cross-restart reply delivery is OFF; "
+ "replies are soft-state and will not survive a restart. Reason: {}",
source, stripCredentials(effectiveUri), reasonOf(e));
log.info("reply inbox: in-memory (soft-state)");
return new InMemoryReplyInbox();
}
}
/** An AMQP URI carries {@code user:pass@} inline — show the host/port, never the credentials. */
static String stripCredentials(String uri) {
return uri == null ? null : uri.replaceAll("://[^@/]*@", "://");
}
/** The deepest cause's class and message — the outermost {@code IllegalStateException} echoes the URI (with password). */
private static String reasonOf(Throwable e) {
Throwable t = e;
while (t.getCause() != null && t.getCause() != t) {
t = t.getCause();
}
String msg = t.getMessage();
return t.getClass().getSimpleName() + (msg == null || msg.isBlank() ? "" : ": " + msg);
}
/**
* CB-594: which env vars the loaded config actually needs, and why — every non-{@code
* subscription} profile's {@code tokenEnv} (a subscription profile never reads one, see
* {@link FleetConfig.Profile#isSubscription()}), plus every profile's {@code gitTokenEnv}
* where set (opt-in), plus a configured {@code broker.uriEnv} (CB-151). Derived from the
* config, not hard-coded, so a new profile is covered for free. A var required by more than one
* profile is one entry naming every profile that needs it. Deliberately excludes {@code
* auth.tokenEnv}: that one is already enforced loudly, by a startup throw in {@code main()} —
* about 370 lines <em>below</em> this method's call site
* ({@link #reportRequiredSecrets(FleetConfig)}), not a few lines above it. That throw only
* fires when {@code auth.mode: token} is configured; under the default loopback-trust mode it
* never runs, and {@code auth.tokenEnv} is simply not required.
*
* <p>Package-private and pure (no I/O, no logging) so the derivation is unit-testable without
* capturing log output; {@link #reportRequiredSecrets(FleetConfig)} is the logging caller.
*/
static Map<String, List<String>> requiredSecretEnvVars(FleetConfig cfg) {
Map<String, List<String>> requiredBy = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (!profile.isSubscription()) {
requiredBy.computeIfAbsent(profile.tokenEnv(), _ -> new ArrayList<>())
.add("profile '" + name + "' tokenEnv");
}
if (profile.hasGitToken()) {
requiredBy.computeIfAbsent(profile.gitTokenEnv(), _ -> new ArrayList<>())
.add("profile '" + name + "' gitTokenEnv");
}
});
FleetConfig.Broker broker = cfg.broker();
if (broker != null && broker.hasUriEnv()) {
requiredBy.computeIfAbsent(broker.uriEnv(), _ -> new ArrayList<>())
.add("broker uriEnv");
}
return requiredBy;
}
/**
* CB-594: log, by name only, which required env vars (see {@link #requiredSecretEnvVars}) are
* set in the daemon's own process environment — the environment every profile's {@code
* tokenEnv}/{@code gitTokenEnv} is read from at spawn time (see
* {@code HerdrPeerLauncher.resolveEnv}). Never logs a value, a prefix, or a length.
*
* <p>A missing entry only warns — it must never refuse to start. A daemon that boots and says
* what is wrong is strictly more useful than one that will not boot at all.
*/
private static void reportRequiredSecrets(FleetConfig cfg) {
Map<String, List<String>> requiredBy = requiredSecretEnvVars(cfg);
if (requiredBy.isEmpty()) {
log.info("startup secrets: no profile references a token env var — nothing to check");
return;
}
Map<String, String> env = System.getenv();
requiredBy.forEach((varName, sources) -> {
String value = env.get(varName);
if (value != null && !value.isBlank()) {
log.info("startup secret {}: set ({})", varName, String.join(", ", sources));
} else {
log.warn("startup secret {}: MISSING ({}) — the daemon will start anyway, and this "
+ "failure stays invisible until a worker actually needs it. Fix "
+ "${SHARED_ENV}/tools/secrets.sh and restart bridged from a LOGIN "
+ "shell (see scripts/redeploy-bridged.sh).",
varName, String.join(", ", sources));
}
});
}
/**
* CB-596: {@code known:} empty (block absent entirely, or present but empty) means {@link
* FleetConfig.MemberCredentials#blockedSet()} is empty too — every member pane inherits the
* operator's whole secret store, unblocked, exactly the defect this ticket fixes. Unlike a
* missing token ({@link #reportRequiredSecrets}), there is no name to point at: the point is
* that the block itself is missing. Warn once at startup and say what to add; never refuse to
* start over it — see {@link #reportRequiredSecrets} for why a daemon that boots and says
* what is wrong beats one that will not boot at all.
*
* <p>Package-private so the test can capture the log directly, the same way {@link
* #requiredSecretEnvVars} is exposed for {@link #reportRequiredSecrets}'s own test.
*/
static void reportMemberCredentialsGap(FleetConfig cfg) {
FleetConfig.MemberCredentials creds = cfg.memberCredentials();
if (creds != null && !creds.known().isEmpty()) {
log.info("memberCredentials: policy={}, {} known name(s), {} allowed — blocking {} on "
+ "every spawn{}",
creds.policy(), creds.known().size(), creds.allow().size(), creds.blockedSet().size(),
creds.isAllowList()
? " (allow-list: known/allow are reporting only — the control is the derived ZDOTDIR scrub)"
: "");
return;
}
log.warn("memberCredentials: absent or empty — the daemon will start anyway, and every "
+ "member pane inherits the operator's WHOLE secret store, unblocked (CB-592's "
+ "protection is lost). Add a memberCredentials: block (policy/allow/known) to "
+ "bridged.yaml — see fleetd.example.yaml — and restart.");
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
*
* @return true if herdr answered, false if it never did
*/
private static boolean awaitHerdr(HerdrClient herdr) {
long deadline = System.nanoTime() + HERDR_WAIT_SECONDS * 1_000_000_000L;
boolean waited = false;
while (true) {
try {
herdr.call("ping");
if (waited) {
log.info("herdr is up");
}
return true;
} catch (HerdrException e) {
if (System.nanoTime() >= deadline) {
return false;
}
if (!waited) {
log.info("waiting up to {}s for the herdr socket…", HERDR_WAIT_SECONDS);
waited = true;
}
try {
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
return false;
}
}
}
}
private Fleetd() {
}
}
@@ -1,272 +0,0 @@
package dev.ltms.fleet.auth;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.peer.MemberRole;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
import java.util.Map;
import java.util.function.Function;
import java.util.function.Supplier;
/**
* Resolves every caller to a {@link Principal}, for both entry paths into the core (CB-501).
*
* <p>There are two of them and they are not layered the way the docs suggest: {@code FleetMcp}
* calls the service layer directly and is mounted as a raw servlet (so it never passes through a
* Javalin filter), while the REST routes historically resolved no identity at all. Both now
* delegate here, so the authorization rules are stated once instead of drifting apart.
*
* <p><strong>Resolution order</strong> — connection identity first, token second, nothing third:
* <ol>
* <li>A loopback peer PID that maps to a pane named by {@code leaders:}, by the legacy
* {@code primary.terminal} pin, or by an operator-labelled lead tab (CB-307, CB-530, CB-531)
* ⇒ {@link Role#PRIMARY}, carrying that lead's
* name. The pane mapping is as unforgeable as a worker's, and the config explicitly names
* that pane as a lead's own — without this rule a lead running <em>inside</em> a herdr pane
* is misread as a worker and locked out of orchestration. More than one pane may be named,
* so two leads can work as peers rather than one being demoted.</li>
* <li>A loopback peer PID that maps to a pane bound to a CB-548 architect slot ⇒
* {@link Role#ARCHITECT}, carrying the slot name. Just unforgeable as a worker's, and
* resolved from the <em>live</em> terminal→slot binding (never a request argument), before
* the generic worker fallback.</li>
* <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#WORKER}. This is
* unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured
* regardless of auth mode, so enabling auth never breaks the fleet.</li>
* <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li>
* <li>Otherwise, under {@code loopback-trust}, a loopback caller ⇒ {@link Role#PRIMARY}
* (the historical behaviour, now an explicit configured choice).</li>
* <li>Otherwise {@link Role#ANONYMOUS}.</li>
* </ol>
*/
public final class CallerResolver {
private final ConnectionIdentity identity;
private final boolean tokenMode;
private final byte[] expectedToken; // null unless tokenMode
/**
* terminal_id → lead name; empty when nothing is pinned. CB-530.
*
* <p>A supplier rather than a map because the registry is no longer fixed at startup: CB-531
* discovers leads by scanning herdr for operator-labelled tabs, so a lead that opens its tab
* after the daemon booted must still be recognised. Consulted per resolve; the scanner behind
* it is TTL-cached, so this is a map lookup in the common case.
*/
private final Supplier<Map<String, String>> leadTerminals;
/**
* terminal_id → architect slot name; empty when nothing is configured. CB-548.
*
* <p>Like {@link #leadTerminals}, a supplier rather than a fixed map, so a binding injected
* after startup — when the later spawn lifecycle establishes a live architect session, or an
* operator pins one — takes effect without a restart. Consulted per resolve; today's wiring
* in {@code Fleetd} reads a constant from config, which is the degenerate live case.
*/
private final Supplier<Map<String, String>> architectTerminals;
private final Function<String, MemberRole> memberSlotRoles;
private final Function<String, String> memberSlotNames;
/** Loopback-trust resolver: no token required, historical behaviour. Test-only. */
CallerResolver(ConnectionIdentity identity) {
this(identity, false, null, Map.of());
}
/** As {@link #CallerResolver(ConnectionIdentity, boolean, String, Map)} with no leads pinned. Test-only. */
CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
this(identity, tokenMode, token, Map.of());
}
/**
* Single-pin form for the legacy {@code primary.terminal}-only configuration — one lead, named
* {@code primary}.
*
* <p>A static factory rather than a fourth constructor overload on purpose: {@code String} and
* {@code Map} overloads are ambiguous for a literal {@code null} argument, which is a compile
* error at the call site and exactly the shape "unpinned" is written in.
*
* @param pinnedPrimaryTerminal the primary's own herdr {@code terminal_id}
* ({@code null}/blank = unpinned)
*/
static CallerResolver pinnedTo(ConnectionIdentity identity, boolean tokenMode,
String token, String pinnedPrimaryTerminal) {
return new CallerResolver(identity, tokenMode, token,
pinnedPrimaryTerminal == null || pinnedPrimaryTerminal.isBlank()
? Map.of() : Map.of(pinnedPrimaryTerminal, "primary"));
}
/**
* @param identity connection-based worker identification
* @param tokenMode when true, a non-worker caller must present a valid bearer token
* @param token the expected bearer token; required (non-blank) when {@code tokenMode}
* @param leadTerminals herdr {@code terminal_id} → lead name for every configured lead
* (CB-530). A caller resolving to one of these panes is that lead — a
* {@link Role#PRIMARY} — rather than a worker. Empty = nothing pinned,
* so every pane resolves as a worker.
*/
CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Map<String, String> leadTerminals) {
this(identity, tokenMode, token, fixed(leadTerminals), null);
}
/**
* Live-registry form: {@code leadTerminals} is consulted on every resolve, so leads discovered
* after startup (CB-531's tab scan) take effect without a restart.
*
* <p>A static factory rather than a fourth constructor overload, for the same reason as
* {@link #pinnedTo}: {@code Map} and {@code Supplier} overloads are ambiguous for a literal
* {@code null}.
*/
static CallerResolver withLeads(ConnectionIdentity identity, boolean tokenMode,
String token,
Supplier<Map<String, String>> leadTerminals) {
return new CallerResolver(identity, tokenMode, token, leadTerminals, null);
}
/**
* Live registry form that can confirm a bound slot is an architect slot.
*
* <p>This is the only public construction path. It keeps terminal bindings and slot roles in
* the same {@link MemberRegistry}, so a configured architect can resolve as an architect.
*/
public static CallerResolver withLeadsAndMembers(ConnectionIdentity identity,
boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
MemberRegistry members) {
return new CallerResolver(identity, tokenMode, token, leadTerminals,
members == null ? null : members::snapshot,
members == null ? null : members::roleForSlot,
members == null ? null : members::nameForSlot);
}
private static Supplier<Map<String, String>> fixed(Map<String, String> leadTerminals) {
Map<String, String> snapshot = leadTerminals == null ? Map.of() : Map.copyOf(leadTerminals);
return () -> snapshot;
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals) {
this(identity, tokenMode, token, leadTerminals, architectTerminals, null);
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals,
Function<String, MemberRole> memberSlotRoles) {
this(identity, tokenMode, token, leadTerminals, architectTerminals, memberSlotRoles, null);
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals,
Function<String, MemberRole> memberSlotRoles,
Function<String, String> memberSlotNames) {
if (tokenMode && (token == null || token.isBlank())) {
throw new IllegalArgumentException(
"auth.mode=token requires a non-empty token; check that the env var named by "
+ "auth.tokenEnv is exported to the daemon's environment");
}
this.identity = identity;
this.tokenMode = tokenMode;
this.expectedToken = tokenMode ? token.getBytes(StandardCharsets.UTF_8) : null;
this.leadTerminals = leadTerminals == null ? Map::of : leadTerminals;
this.architectTerminals = architectTerminals == null ? Map::of : architectTerminals;
this.memberSlotRoles = memberSlotRoles == null ? _ -> null : memberSlotRoles;
this.memberSlotNames = memberSlotNames == null ? Function.identity() : memberSlotNames;
}
/**
* The currently-recognised leads, {@code terminal_id → name} (CB-535).
*
* <p>Deliberately read from the same supplier {@link #resolve} consults, rather than from a
* second copy handed to the roster: a lead that is <em>listed</em> but would not <em>resolve</em>
* (or the reverse) is an address a peer cannot actually reach, and the two answers drifting apart
* is precisely the confusion this exists to end. Live, so a lead discovered by the tab scan after
* startup appears without a restart.
*/
public Map<String, String> leads() {
return leadTerminals.get();
}
/**
* The currently-recognised architect slots, {@code terminal_id → slot name} (CB-548).
*
* <p>Read from the same supplier {@link #resolve} consults, so a slot that is <em>listed</em>
* here but would not <em>resolve</em> (or the reverse) cannot drift apart. Live for the same
* reason as {@link #leads()}.
*/
public Map<String, String> members() {
return architectTerminals.get();
}
/**
* Resolve the caller of a request.
*
* @param remoteAddr the connection's remote address
* @param remotePort the connection's remote port (used for the peer-PID lookup)
* @param authorizationHeader the raw {@code Authorization} header, or {@code null}
*/
public Principal resolve(String remoteAddr, int remotePort, String authorizationHeader) {
ConnectionIdentity.Caller c = identity.resolve(remoteAddr, remotePort);
if (c.terminal() != null) {
String lead = leadTerminals.get().get(c.terminal());
if (lead != null) {
// The config names this pane as a lead's own. The pane mapping is exactly as
// unforgeable as a worker's, so it outranks the token path — no credential needed.
// Checked before the architect registry so a pane named in BOTH is still the lead
// (CB-548 preserves every existing leader behaviour).
return Principal.leader(lead, c.terminal(), c.pid());
}
String slot = architectTerminals.get().get(c.terminal());
if (slot != null && memberSlotRoles.apply(slot) == MemberRole.ARCHITECT) {
// The config/live binding names this pane as an architect slot's own. Same
// unforgeable pane mapping; the live binding, never a request argument, decides.
// Check the slot role too: this defence in depth prevents a bad lifecycle bind from
// escalating a dev or reviewer into an architect. Checked before the worker fallback.
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
}
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
}
if (tokenMode) {
return presentedTokenMatches(authorizationHeader)
? Principal.primary(c.pid())
: Principal.anonymous();
}
// loopback-trust: same-host callers that are not workers are the primary. A non-loopback
// caller is anonymous even here — and startup refuses that combination anyway
// (FleetConfig.validateAuthExposure), so this is defence in depth, not the control.
return isLoopback(remoteAddr) ? Principal.primary(c.pid()) : Principal.anonymous();
}
private boolean presentedTokenMatches(String authorizationHeader) {
String presented = bearerValue(authorizationHeader);
if (presented == null) {
return false;
}
// Constant-time: MessageDigest.isEqual does not short-circuit on the first differing byte,
// so a token cannot be recovered a byte at a time by timing the response.
return MessageDigest.isEqual(presented.getBytes(StandardCharsets.UTF_8), expectedToken);
}
/** Extract the credential from {@code Authorization: Bearer <token>}, or {@code null}. */
private static String bearerValue(String header) {
if (header == null) {
return null;
}
String h = header.trim();
if (h.length() < 7 || !h.regionMatches(true, 0, "Bearer ", 0, 7)) {
return null;
}
String token = h.substring(7).trim();
return token.isEmpty() ? null : token;
}
private static boolean isLoopback(String remoteAddr) {
if (remoteAddr == null) {
return false;
}
return remoteAddr.equals("127.0.0.1") || remoteAddr.equals("::1")
|| remoteAddr.equals("0:0:0:0:0:0:0:1") || remoteAddr.startsWith("127.");
}
}
@@ -1,21 +0,0 @@
package dev.ltms.fleet.auth;
import dev.ltms.fleet.peer.MemberRole;
/** Optional session lifecycle hook for live member-slot bindings. */
public interface MemberLifecycle {
MemberLifecycle NONE = new MemberLifecycle() {
@Override
public void acquired(MemberRole role, String profile, String terminal) {
}
@Override
public void released(String terminal) {
}
};
void acquired(MemberRole role, String profile, String terminal);
void released(String terminal);
}
@@ -1,232 +0,0 @@
package dev.ltms.fleet.auth;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.peer.MemberRole;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.HashMap;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
/**
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
* profile it points at, plus the <em>live</em> bindings from a live architect's herdr terminal to
* its slot.
*
* <p>Two halves, split by who owns each:
* <ul>
* <li><b>slots</b> — configured once, keyed by the gateway-local unique name; each carries the
* {@code profile} reference the spawn lifecycle reads when it stands the slot up. A read-only
* snapshot taken at construction.</li>
* <li><b>terminal bindings</b> — owned by this registry and initially <em>empty</em>. Config
* declares no architect terminal, so at startup every slot is idle and nothing resolves to an
* architect; a session only becomes one when the spawn lifecycle {@linkplain #bind(String,
* String) binds} its terminal to a slot. {@link CallerResolver} reads this through
* {@link #snapshot()} to turn a pane into an {@link Role#ARCHITECT}.</li>
* </ul>
*
* <p>Spawning/lifecycle is deliberately a separate unit: this class only owns the bindings and
* exposes the map the resolver resolves against plus the profile lookup lifecycle will call.
* Nothing here creates or manages an architect session.
*/
public final class MemberRegistry implements MemberLifecycle {
private static final Logger log = LoggerFactory.getLogger(MemberRegistry.class);
/**
* One flattened {@code fleet:} entry.
*
* <p>Flattened because a slot name is unique only <em>within</em> its pool — {@code sonnet} may
* legitimately be both a developer and a reviewer — while a terminal binds to exactly one thing.
* The qualified {@link #key()} is what that binding uses.
*
* @param name the slot's key inside its pool
* @param role the pool it came from
* @param profile the backend it runs on
*/
public record Entry(String name, MemberRole role, String profile) {
/** {@code "architect:opus"} — unique across pools, unlike {@link #name()}. */
public String key() {
return role.wireName() + ":" + name;
}
}
private final Map<String, Entry> slots;
/** Live {@code terminal_id → qualified slot key}; guarded by {@code this}. */
private final Map<String, String> terminalToSlot = new HashMap<>();
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
public MemberRegistry(FleetConfig.Fleet fleet) {
Map<String, Entry> flat = new LinkedHashMap<>();
if (fleet != null) {
for (MemberRole role : MemberRole.values()) {
fleet.pool(role).forEach((name, slot) -> {
if (slot != null) {
Entry e = new Entry(name, role, slot.profile());
flat.put(e.key(), e);
}
});
}
}
this.slots = Collections.unmodifiableMap(flat);
}
/** The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot. */
public Map<String, Entry> slots() {
return slots;
}
/** The slots belonging to {@code role}, in definition order. */
public Map<String, Entry> slotsFor(MemberRole role) {
Map<String, Entry> out = new LinkedHashMap<>();
slots.forEach((key, e) -> {
if (e.role() == role) {
out.put(key, e);
}
});
return Collections.unmodifiableMap(out);
}
/**
* An immutable copy of the live {@code terminal_id → slot name} bindings.
*
* <p>Passed to {@link CallerResolver} as the source of architect identity, and what
* {@code fleet_whoami}/the roster will read to say which slot a pane hosts. Empty until the
* spawn lifecycle binds a slot.
*/
public Map<String, String> snapshot() {
synchronized (terminalToSlot) {
return Map.copyOf(terminalToSlot);
}
}
/** The slot a live terminal is bound to, or {@code null} if it is not an architect slot. */
public String slotForTerminal(String terminal) {
if (terminal == null) {
return null;
}
synchronized (terminalToSlot) {
return terminalToSlot.get(terminal);
}
}
/**
* The strong-model profile a slot runs under — what the spawn lifecycle reads.
*
* @return the slot's configured {@code profile}, or {@code null} if the slot is unknown or
* declares none
*/
public String profileForSlot(String slotName) {
Entry e = slots.get(slotName);
return (e == null || e.profile() == null) ? null : e.profile();
}
/** The role a qualified slot key belongs to, or {@code null} when the key is unknown. */
public MemberRole roleForSlot(String slotName) {
Entry e = slots.get(slotName);
return e == null ? null : e.role();
}
/** The unqualified configured name for a slot, or {@code null} if it is unknown. */
public String nameForSlot(String slotName) {
Entry e = slots.get(slotName);
return e == null ? null : e.name();
}
/** True when {@code slotName} is a configured architect slot. */
public boolean isSlot(String slotName) {
return slots.containsKey(slotName);
}
/**
* Bind {@code terminal} to {@code slot} (CB-548).
*
* <p>The spawn lifecycle calls this when it stands a slot up. The bind is atomic and preserves
* the two cardinality invariants: a terminal may occupy at most one slot, and a slot may host at
* most one terminal. Binding the same terminal to the same slot again is a harmless no-op.
*
* @param slot a configured slot name, or the bind is refused
* @param terminal the pane that will act as this architect
* @return {@code true} if the binding is now {@code terminal → slot}; {@code false} if it was
* refused — an unknown slot, a terminal already bound to a different slot, or a slot
* already hosting a different terminal
*/
public boolean bind(String slot, String terminal) {
if (slot == null || terminal == null || terminal.isBlank()) {
return false;
}
synchronized (terminalToSlot) {
if (!isSlot(slot)) {
return false; // unknown slot — nothing to bind to
}
String existingSlot = terminalToSlot.get(terminal);
if (existingSlot != null) {
return slot.equals(existingSlot); // already this slot (idempotent) or a different one
}
if (terminalToSlot.containsValue(slot)) {
return false; // slot already hosts a terminal — no second one
}
terminalToSlot.put(terminal, slot);
return true;
}
}
/**
* Compare-safe unbind of {@code expectedTerminal} from {@code slot} (CB-548).
*
* <p>The spawn lifecycle calls this when it tears a slot down. Only the exact binding
* {@code expectedTerminal → slot} is removed; if that terminal was since rebound to a different
* slot (or the slot to a different terminal), the call is a no-op returning {@code false} — a
* stale unbind must never remove a replacement.
*
* @param slot the slot the caller believes the terminal is bound to
* @param expectedTerminal the terminal it expects to be bound there
* @return {@code true} if {@code expectedTerminal → slot} was removed; {@code false} if nothing
* was (no such binding, or the binding had already moved)
*/
public boolean unbind(String slot, String expectedTerminal) {
if (slot == null || expectedTerminal == null) {
return false;
}
synchronized (terminalToSlot) {
String current = terminalToSlot.get(expectedTerminal);
if (current == null || !slot.equals(current)) {
return false; // absent, or a replacement/moved binding — leave it in place
}
terminalToSlot.remove(expectedTerminal);
return true;
}
}
/**
* Bind only architect sessions to a free slot with the resolved profile.
*
* <p>The role check is lifecycle policy. {@link CallerResolver} repeats it when resolving a
* binding, so a later lifecycle regression cannot turn a worker into an architect.
*/
@Override
public void acquired(MemberRole role, String profile, String terminal) {
if (role != MemberRole.ARCHITECT || terminal == null || terminal.isBlank()) {
return;
}
// slotsFor preserves definition order, so duplicate-profile slots use the first free one.
for (Entry entry : slotsFor(MemberRole.ARCHITECT).values()) {
if (Objects.equals(profile, entry.profile()) && bind(entry.key(), terminal)) {
return;
}
}
log.info("member slot: no free architect slot for profile={}; session remains a worker", profile);
}
/** Unbind a released terminal using the compare-safe registry operation. */
@Override
public void released(String terminal) {
String slot = slotForTerminal(terminal);
if (slot != null) {
unbind(slot, terminal);
}
}
}
@@ -1,127 +0,0 @@
package dev.ltms.fleet.auth;
/**
* A resolved caller: its {@link Role}, and — for a worker — the herdr {@code terminal_id} that
* identifies which worker it is (CB-501).
*
* @param role what this caller is authorized to act as
* @param terminal the herdr {@code terminal_id} of the pane this caller occupies — a worker's, or
* (since CB-532) a named lead's; {@code null} for an unnamed primary resolved off a
* token or loopback trust, and for {@code ANONYMOUS}
* @param pid the connecting process id, or {@code -1} when not resolvable (audit context)
* @param name for a lead resolved from the CB-530 {@code leaders:} registry, which lead it is;
* for an architect resolved from the CB-548 {@code architects:} registry, which
* slot it occupies; {@code null} for every other caller, including an unnamed primary
*/
public record Principal(Role role, String terminal, long pid, String name) {
/**
* Three-arg form for the callers that have no name to carry (workers, anonymous, and the
* token/loopback primary paths). Kept so adding CB-530's {@code name} did not churn every
* construction site — and so a reconstructed principal without a stashed name still works.
*/
public Principal(Role role, String terminal, long pid) {
this(role, terminal, pid, null);
}
/** A caller authenticated as nothing — the default when no check establishes anything else. */
public static Principal anonymous() {
return new Principal(Role.ANONYMOUS, null, -1);
}
/** The orchestrating session, unnamed (token or loopback-trust path). */
public static Principal primary(long pid) {
return new Principal(Role.PRIMARY, null, pid);
}
/**
* A named lead from the {@code leaders:} registry (CB-530).
*
* <p>Carries {@link Role#PRIMARY}: a lead <em>is</em> a primary as far as authorization goes,
* so every existing {@code isPrimary()} gate keeps working unchanged and the role table needed
* no new entry. The name is reporting only — it lets {@code fleet_whoami} say <em>which</em>
* lead is asking once more than one is configured.
*
* <p><strong>CB-532: a lead now carries the terminal it was matched by.</strong> Under CB-530 it
* deliberately did not, because {@code terminal} meant "which worker pane" everywhere and a
* non-null one would have enrolled the lead in the worker presence map. That reading was what
* made a lead unaddressable: {@link #ownsSession} could never be true for it, so
* {@code fleet_reply} was refused and one lead could send to another but never be answered.
* The terminal now means "which pane is this caller", the presence map keys on
* {@link #isSpawnedMember()} instead, and a lead is a peer that can both send and receive.
*/
public static Principal leader(String name, String terminal, long pid) {
return new Principal(Role.PRIMARY, terminal, pid, name);
}
/** A worker peer, identified by its herdr pane. */
public static Principal worker(String terminal, long pid) {
return new Principal(Role.WORKER, terminal, pid);
}
/**
* An architect (CB-548), identified by the slot it occupies and the pane bound to it.
*
* <p>Carries {@link Role#ARCHITECT}. {@code slotName} is reporting only — it lets
* {@code fleet_whoami} say <em>which</em> architect slot is asking, and it is the key the
* (future) spawn lifecycle reads a profile back from. Identity is the {@code terminal}: like a
* worker's it comes from the connection and the live terminal→slot binding, so
* {@code ownsSession} works exactly as it does for a worker — an architect acts as its own
* pane and no other.
*/
public static Principal architect(String slotName, String terminal, long pid) {
return new Principal(Role.ARCHITECT, terminal, pid, slotName);
}
public boolean isPrimary() {
return role == Role.PRIMARY;
}
public boolean isArchitect() {
return role == Role.ARCHITECT;
}
public boolean isWorker() {
return role == Role.WORKER;
}
/**
* Whether this caller is a spawned member with its own pane.
*
* <p>Both workers and architects are spawned members. A lead is excluded because recording it
* as present would count it as an available member in the roster.
*/
public boolean isSpawnedMember() {
return role == Role.WORKER || role == Role.ARCHITECT;
}
public boolean isAnonymous() {
return role == Role.ANONYMOUS;
}
/**
* Whether this caller may act <em>as</em> {@code sessionId} — the "own session only" rule that
* keeps one peer from replying or asking on another's behalf.
*
* <p>The rule is about <em>identity</em>, not rank: a caller may act as the pane it demonstrably
* occupies, and as no other. CB-532 dropped the extra {@code isWorker()} conjunct that used to
* be here. It was not what enforced the rule — {@code terminal.equals(sessionId)} is, and that
* terminal comes from the connection, so it cannot be forged either way. All the conjunct did
* was make a lead permanently unable to answer anyone, since a lead's terminal was null and a
* lead is not a worker. A caller with no terminal at all (an off-host or token-authenticated
* primary) still owns nothing, which is the case the null check covers.
*/
public boolean ownsSession(String sessionId) {
return terminal != null && terminal.equals(sessionId);
}
/** Short, non-sensitive description for audit lines and error details. */
public String describe() {
return switch (role) {
case WORKER -> "worker:" + terminal;
case ARCHITECT -> "architect:" + name;
case PRIMARY -> name == null ? "primary" : "leader:" + name;
case ANONYMOUS -> "anonymous";
};
}
}
@@ -1,294 +0,0 @@
package dev.ltms.fleet.config;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Supplier;
/**
* The daemon's live configuration, re-readable without a restart (CB-559).
*
* <p>Consumers hold this, not a {@link FleetConfig}, and read through {@link #get()} at the point
* of use. A component that captures {@code ref.get()} into a field at construction has opted out of
* reload — which is sometimes right (see <em>deferred</em> below), but it must then be a deliberate
* choice rather than an accident of where the field was initialised.
*
* <h2>Not every key can change under a running daemon</h2>
* Keys fall into three classes, and the difference is about what already exists when the reload
* happens — not about how important the key is.
*
* <ul>
* <li><strong>Hot</strong> — re-read per use, so a reload takes effect on the next spawn:
* {@code fleet:} (every role pool, {@code charters}, and {@code tabLabel}),
* {@code placement:}, and an existing profile's {@code weight} / {@code maxLoad}. Those
* three are read through a supplier on {@code CompositePeerLauncher}, which is what makes
* them hot — not the fact that they are config. <strong>This does NOT include
* {@code fleet.leaders}</strong>: {@code Fleetd.main} reads {@code cfg.fleet().leaders()}
* once at startup to build the {@code LeadTabScanner} and the {@code LeadLauncher}, and
* neither is reconstructed on reload — so a lead added, removed, or re-{@code tab}'d under
* {@code fleet.leaders} needs a restart, the same as any deferred key below.</li>
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code quarantineCooldownSeconds}
* (CB-578 stage B — baked once into the {@code BackendQuarantine} built at startup),
* {@code guard:}, {@code worktreeRoot:}, adding or removing a profile (a new backend needs its own launcher,
* which is constructed once), <em>and an existing profile's launch settings</em> —
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl},
* {@code exhaustedPattern} (CB-578 stage A — compiled once into {@code Fleetd.main}'s
* pattern map at startup), and the rest. {@code credentialId} (CB-578 stage B) is NOT on
* this list — it is read live off the config supplier at every quarantine check and
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so it is hot instead.
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
* each spawn out of that copy, so those never reach a launch until the daemon restarts. A
* reload logs these rather than pretending they applied.</li>
* <li><strong>Cold</strong> — cannot change at all under a running daemon: {@code bind:},
* {@code herdrSocket:}, {@code broker:} and {@code auth:}. The socket is bound, the broker
* connection is open, and the auth mode decides who may reach the port that is already
* listening.</li>
* </ul>
*
* <p><strong>A cold change refuses the whole reload.</strong> Not the hot half applied and the cold
* half warned about: that would leave the running daemon in a state matching no file on disk, which
* is the worst thing a reload can do to an operator debugging one. Refusing keeps the invariant that
* the live config is always some version of the file, and the message names the keys that must
* change through a restart.
*
* <p>A reload that fails to parse or fails validation is also refused, and the previous config keeps
* running. A config file being edited is normally read once mid-save; degrading a working daemon
* because it caught a half-written file would be a bad trade.
*/
public final class ConfigRef implements Supplier<FleetConfig> {
private static final Logger log = LoggerFactory.getLogger(ConfigRef.class);
/** Keys that cannot change under a running daemon — see the class doc. */
private static final Set<String> COLD_KEYS =
Set.of("bind", "herdrSocket", "broker", "auth");
private final Path path;
private final AtomicReference<FleetConfig> current;
public ConfigRef(Path path, FleetConfig initial) {
this.path = path;
this.current = new AtomicReference<>(Objects.requireNonNull(initial, "initial config"));
}
/** A fixed reference that never reloads — for tests and for wiring built from a config in code. */
public static ConfigRef fixed(FleetConfig cfg) {
return new ConfigRef(null, cfg);
}
/** The live configuration. Read this per use; do not cache it in a field. */
@Override
public FleetConfig get() {
return current.get();
}
/** The file this ref reloads from, or {@code null} for a {@link #fixed} ref. */
public Path path() {
return path;
}
/**
* What a reload attempt did.
*
* @param applied true when the new config is now live
* @param coldKeys cold keys whose value changed, which is why an unapplied reload was refused
* @param deferred keys that changed and were accepted, but whose effect waits for a restart
* @param error the parse or validation failure that refused the reload, else {@code null}
*/
public record Outcome(boolean applied, List<String> coldKeys, List<String> deferred,
String error) {
public Outcome {
coldKeys = List.copyOf(coldKeys);
deferred = List.copyOf(deferred);
}
static Outcome refusedCold(List<String> keys) {
return new Outcome(false, keys, List.of(), null);
}
static Outcome failed(String error) {
return new Outcome(false, List.of(), List.of(), error);
}
/** A one-line summary for the operator — the reason, not just the verdict. */
public String summary() {
if (error != null) {
return "config reload refused — " + error;
}
if (!applied) {
return "config reload refused — these keys cannot change under a running daemon: "
+ String.join(", ", coldKeys) + ". Restart bridged to apply them.";
}
if (!deferred.isEmpty()) {
return "config reloaded; these changes need a restart to take effect: "
+ String.join(", ", deferred);
}
return "config reloaded";
}
}
/**
* Re-read the file, validate it, and swap it in when nothing cold changed.
*
* <p>Never throws: a reload is a best-effort operation on a daemon that is already serving, and
* a bad edit must not take it down. Every failure path leaves the previous config live and is
* reported through the returned {@link Outcome}.
*/
public Outcome reload() {
if (path == null) {
return Outcome.failed("this config was built in code and has no file to reload from");
}
FleetConfig old = current.get();
FleetConfig fresh;
try {
fresh = FleetConfig.load(path);
// The same gate startup runs. A config that would have refused to boot must not be able
// to slip in through a reload — that is how a daemon ends up in a state it could never
// have started in, which is the hardest kind to debug.
fresh.validateAuthExposure();
fresh.validateLeadTabPrefixes();
fresh.validateSubscriptionProfiles();
fresh.validateCharters();
fresh.validateMembers();
} catch (RuntimeException e) {
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
log.warn("config reload from {} refused, keeping the running config: {}", path, msg);
return Outcome.failed(msg);
}
List<String> cold = changedColdKeys(old, fresh);
if (!cold.isEmpty()) {
Outcome out = Outcome.refusedCold(cold);
log.warn(out.summary());
return out;
}
List<String> deferred = changedDeferredKeys(old, fresh);
current.set(fresh);
Outcome out = new Outcome(true, List.of(), deferred, null);
log.info(out.summary());
return out;
}
/** Cold keys whose value differs between the running config and the candidate. */
private static List<String> changedColdKeys(FleetConfig old, FleetConfig fresh) {
List<String> changed = new ArrayList<>();
if (!Objects.equals(old.bind(), fresh.bind())) {
changed.add("bind");
}
if (!Objects.equals(old.herdrSocket(), fresh.herdrSocket())) {
changed.add("herdrSocket");
}
if (!Objects.equals(old.broker(), fresh.broker())) {
changed.add("broker");
}
if (!Objects.equals(old.auth(), fresh.auth())) {
changed.add("auth");
}
// Kept in step with COLD_KEYS so the doc and the code cannot drift apart silently.
assert COLD_KEYS.containsAll(changed) : "a cold key was reported that COLD_KEYS omits";
return changed;
}
/** Changed keys that were accepted but whose effect waits for a restart. */
private static List<String> changedDeferredKeys(FleetConfig old, FleetConfig fresh) {
List<String> changed = new ArrayList<>();
if (!Objects.equals(old.lifecycle(), fresh.lifecycle())) {
changed.add("lifecycle");
}
if (!Objects.equals(old.leadHeartbeat(), fresh.leadHeartbeat())) {
changed.add("leadHeartbeat");
}
if (!Objects.equals(old.guard(), fresh.guard())) {
changed.add("guard");
}
if (!Objects.equals(old.worktreeRoot(), fresh.worktreeRoot())) {
changed.add("worktreeRoot");
}
if (!Objects.equals(old.spawnReadyTimeoutMs(), fresh.spawnReadyTimeoutMs())
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
changed.add("spawnReady*");
}
// CB-578 stage B: baked once into the BackendQuarantine built at startup — a running
// quarantine keeps its original cooldown regardless, and a new cooldown only applies to a
// quarantine that starts after a restart.
if (!Objects.equals(old.quarantineCooldownSeconds(), fresh.quarantineCooldownSeconds())) {
changed.add("quarantineCooldownSeconds");
}
Map<String, FleetConfig.Profile> before =
old.profiles() == null ? Map.of() : old.profiles();
Map<String, FleetConfig.Profile> after =
fresh.profiles() == null ? Map.of() : fresh.profiles();
// Adding or removing a profile is deferred: a new backend needs its own launcher, and
// launchers are built once at startup.
if (!before.keySet().equals(after.keySet())) {
Set<String> diff = new LinkedHashSet<>(before.keySet());
diff.addAll(after.keySet());
diff.removeIf(p -> before.containsKey(p) && after.containsKey(p));
changed.add("profiles (added/removed: " + String.join(", ", diff) + ")");
}
// An EXISTING profile's launch settings are deferred too, and this is easy to get wrong:
// `HerdrPeerLauncher` takes `Map.copyOf(profiles)` at construction and `spawn` resolves the
// profile out of that snapshot, so a reloaded model/baseUrl/argv/env never reaches a launch.
// Only weight and maxLoad are genuinely hot, because placement reads them through the
// supplier on the composite rather than from the adapter's copy. Without this check a
// changed model would report "config reloaded" and silently do nothing — the worst outcome
// a reload can produce, because the operator has no reason to doubt it.
List<String> relaunch = new ArrayList<>();
before.forEach((name, was) -> {
FleetConfig.Profile now = after.get(name);
if (now != null && !sameLaunchSettings(was, now)) {
relaunch.add(name);
}
});
if (!relaunch.isEmpty()) {
changed.add("profiles." + String.join("/", relaunch) + " launch settings "
+ "(model, baseUrl, argv, env, …) — the launcher holds a startup snapshot");
}
return changed;
}
/**
* Whether two versions of a profile would launch a peer identically. Compares every component
* the launcher reads at spawn; {@code weight}, {@code maxLoad} and {@code credentialId} are
* excluded because those are read live (by the placement policy and, for credentialId, by
* {@code CompositePeerLauncher}/the CB-578 stage B exhaustion sink) and really do take effect on
* the next spawn.
*/
private static boolean sameLaunchSettings(FleetConfig.Profile a, FleetConfig.Profile b) {
return Objects.equals(a.baseUrl(), b.baseUrl())
&& Objects.equals(a.model(), b.model())
&& Objects.equals(a.configDir(), b.configDir())
&& Objects.equals(a.tokenEnv(), b.tokenEnv())
&& Objects.equals(a.argv(), b.argv())
&& Objects.equals(a.placement(), b.placement())
&& Objects.equals(a.workspace(), b.workspace())
&& Objects.equals(a.tabLabel(), b.tabLabel())
&& Objects.equals(a.mcpUrl(), b.mcpUrl())
// CB-634: the IDE MCP mount is a launch flag, fixed at spawn like mcpUrl — a
// reload changes it only for members spawned after, so a changed value is deferred.
&& Objects.equals(a.ideMcpUrl(), b.ideMcpUrl())
&& Objects.equals(a.cwd(), b.cwd())
&& Objects.equals(a.parityOverlay(), b.parityOverlay())
&& Objects.equals(a.gitTokenEnv(), b.gitTokenEnv())
&& Objects.equals(a.gitHostEnv(), b.gitHostEnv())
&& Objects.equals(a.kind(), b.kind())
&& Objects.equals(a.env(), b.env())
&& Objects.equals(a.subscription(), b.subscription())
// CB-578 stage B: exhaustedPattern is compiled once into Fleetd.main's pattern map
// at startup (see ExhaustedPatternLookup wiring) — a reload never re-reads it, so a
// changed pattern must be reported as deferred, exactly like model/baseUrl/argv.
&& Objects.equals(a.exhaustedPattern(), b.exhaustedPattern());
}
}
@@ -1,96 +0,0 @@
package dev.ltms.fleet.config;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
/**
* Polls {@code bridged.yaml}'s modified time and asks {@link ConfigRef} to reload when it moves
* (CB-559). Opt-in through {@code configReload.enabled}.
*
* <p><strong>Why polling and not a filesystem watch.</strong> {@code WatchService} on macOS has no
* native backend — it falls back to polling internally anyway, at an interval this code does not
* control — and editors save config files in ways that produce a different event mix per editor
* (write-in-place, write-and-rename, write-temp-and-swap). A modified-time check treats all of them
* the same and is a single {@code stat} per tick, which at a ten-second cadence costs nothing worth
* measuring.
*
* <p><strong>A missing or unreadable file is not a reason to act.</strong> Many editors briefly
* unlink the file during a save. Reloading on "it vanished" would mean reloading from a file that no
* longer exists; reporting an error every tick would bury the log. So an unreadable file is skipped
* silently and the next tick tries again — the running config stays live, which is the correct
* outcome either way.
*/
public final class ConfigWatcher {
private static final Logger log = LoggerFactory.getLogger(ConfigWatcher.class);
private final ConfigRef ref;
private final long intervalSeconds;
private final ScheduledExecutorService scheduler;
private volatile long lastSeenMillis;
public ConfigWatcher(ConfigRef ref, long intervalSeconds) {
this.ref = ref;
this.intervalSeconds = intervalSeconds;
this.lastSeenMillis = modifiedMillis(ref.path());
this.scheduler = Executors.newSingleThreadScheduledExecutor(r -> {
Thread t = new Thread(r, "config-watcher");
// A daemon thread: an operator's config watch must never be the reason the JVM refuses
// to exit after everything else has shut down.
t.setDaemon(true);
return t;
});
}
/** Begin watching. A ref with no file (a fixed one) is a no-op rather than an error. */
public void start() {
if (ref.path() == null) {
log.debug("config watch not started — this config has no file behind it");
return;
}
scheduler.scheduleWithFixedDelay(this::tick, intervalSeconds, intervalSeconds,
TimeUnit.SECONDS);
log.info("config watch: {} re-read when it changes (every {}s)", ref.path(), intervalSeconds);
}
/** One poll. Never throws — an exception here would silently cancel the schedule. */
void tick() {
try {
long now = modifiedMillis(ref.path());
if (now == 0 || now == lastSeenMillis) {
return;
}
// Stamp BEFORE reloading. A file whose reload is refused (a bad edit, or a cold key)
// must not be retried every tick — that would log the same refusal forever. The next
// save moves the timestamp again and earns a fresh attempt.
lastSeenMillis = now;
ref.reload();
} catch (RuntimeException e) {
log.warn("config watch tick failed, still watching: {}", e.getMessage());
}
}
private static long modifiedMillis(Path path) {
if (path == null) {
return 0;
}
try {
return Files.getLastModifiedTime(path).toMillis();
} catch (IOException e) {
return 0; // mid-save, or gone: say nothing and try again next tick
}
}
/** Stop polling. Called from the daemon's ordered shutdown hook, alongside the other loops. */
public void stop() {
scheduler.shutdownNow();
}
}
File diff suppressed because it is too large Load Diff
@@ -1,40 +0,0 @@
package dev.ltms.fleet.health;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.session.MemberSession;
/**
* Pure classifier. Collection and repair are deliberately outside this package.
* {@link HealthState#ERROR_ON_SCREEN} is not decided yet because it needs a bounded pane detection
* read and an adapter-specific fatal signature; status facts alone must not guess it.
*/
public final class FleetHealth {
private FleetHealth() { }
public static HealthDecision decide(HealthSnapshot s, HealthPrior prior, long nowNanos) {
if (s.controlLinkDown()) return result(HealthState.CONTROL_LINK_DOWN, false);
if (s.targetNotFound()) return result(HealthState.GONE, false);
if (s.sessionState() == MemberSession.State.SPAWNING && !s.present() && s.readinessGraceElapsed()) {
return result(HealthState.NEVER_READY, false);
}
if (s.orphanedDelegation()) return result(HealthState.DELEGATION_ORPHANED, false);
boolean disagreement = s.sessionState() == MemberSession.State.BUSY && s.acceptedDelivery()
&& (s.liveStatus() == AgentStatus.IDLE || s.liveStatus() == AgentStatus.DONE);
if (disagreement && prior.busyButDone()) return result(HealthState.TURN_BOUNDARY_LOST, true);
if (s.stalled()) return result(HealthState.STALL_SUSPECTED, disagreement);
if (s.replyStranded()) return result(HealthState.REPLY_STRANDED, disagreement);
if (s.queuedDelivery() || s.inboxMessage()) return result(HealthState.WORK_PENDING, disagreement);
if (s.sessionState() == MemberSession.State.SPAWNING) return result(HealthState.STARTING, disagreement);
if (s.acceptedDelivery() && s.liveStatus() == AgentStatus.BLOCKED) {
return result(HealthState.BLOCKED_AMBIGUOUS, disagreement);
}
// An accepted delivery remains bridge work even when herdr is late, unknown, or has already
// reported DONE once. It cannot be IDLE until the delegation has resolved.
if (s.acceptedDelivery()) return result(HealthState.WORKING, disagreement);
return result(HealthState.IDLE, disagreement);
}
private static HealthDecision result(HealthState state, boolean disagreement) {
return new HealthDecision(state, new HealthPrior(disagreement));
}
}
@@ -1,151 +0,0 @@
package dev.ltms.fleet.health;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.session.MemberSession;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.HashMap;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.BiConsumer;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/** Slow whole-fleet evidence collection. It is deliberately separate from the delivery poller. */
public final class FleetHealthMonitor {
private static final Logger log = LoggerFactory.getLogger(FleetHealthMonitor.class);
/** Bounded attempts to run {@link #failTarget} for one transition. Never retried tick-to-tick (CB-580). */
static final int MAX_FAIL_TARGET_ATTEMPTS = 3;
private final AgentControl agents;
private final Supplier<List<MemberSession>> roster;
private final MessageService messages;
private final ScheduledExecutorService scheduler;
private final LongSupplier clock;
private final long intervalSeconds;
private final BiConsumer<String, String> failTarget;
private final Map<String, HealthPrior> priors = new HashMap<>();
private final Map<String, HealthState> states = new HashMap<>();
// These facts need the evidence publishers introduced by later M4 units. They are not negatives.
private static final boolean NOT_YET_OBSERVED = false;
/**
* @param failTarget CB-568's idempotent target-wide failure operation (e.g. {@code messages::abandon}),
* invoked once when a member transitions into a terminal health state. Required —
* there is deliberately no defaulting overload; a caller that does not want the
* fail-tickets-on-terminal-health behavior must pass an explicit inert value (see
* {@code TestTurnTokens.inert} / {@code FleetMcp.CapacitySource.none()} for the pattern).
*/
public FleetHealthMonitor(AgentControl agents, Supplier<List<MemberSession>> roster, MessageService messages,
ScheduledExecutorService scheduler, LongSupplier clock, long intervalSeconds,
BiConsumer<String, String> failTarget) {
this.agents = agents;
this.roster = roster;
this.messages = messages;
this.scheduler = scheduler;
this.clock = clock;
this.intervalSeconds = intervalSeconds;
this.failTarget = Objects.requireNonNull(failTarget, "failTarget");
}
/** Pure per-member decision seam. */
static HealthDecision decide(HealthSnapshot snapshot, HealthPrior prior, long nowNanos) {
return FleetHealth.decide(snapshot, prior, nowNanos);
}
public void start() { scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS); }
public void stop() { scheduler.shutdownNow(); }
// Package-private so tests can run one tick without waiting.
void tick() {
try {
List<Agent> agentsNow = agents.list(); // Exactly one list call for this complete observation.
List<MemberSession> rosterNow = roster.get(); // One in-memory roster snapshot for this tick.
Map<String, Agent> live = new HashMap<>();
for (Agent agent : agentsNow) live.put(agent.terminalId(), agent);
HashSet<String> current = new HashSet<>();
for (MemberSession session : rosterNow) {
current.add(session.terminalId());
Agent agent = live.get(session.terminalId());
AgentStatus status = agent == null ? AgentStatus.UNKNOWN : agent.status();
boolean accepted = messages.hasAcceptedDelivery(session.terminalId());
HealthSnapshot snapshot = new HealthSnapshot(session.state(), status, accepted, NOT_YET_OBSERVED,
messages.hasInboxMessage(session.terminalId()), agent != null, NOT_YET_OBSERVED,
NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED);
HealthDecision decision = decide(snapshot, priors.getOrDefault(session.terminalId(), HealthPrior.NONE),
clock.getAsLong());
priors.put(session.terminalId(), decision.prior());
reportTransition(session.terminalId(), decision.state());
}
priors.keySet().retainAll(current);
states.keySet().retainAll(current);
} catch (Throwable error) {
// A list failure is health evidence, and must never kill the monitor's only scheduler task.
log.warn("fleet health collection failed; will retry next tick", error);
} finally {
if (!scheduler.isShutdown()) {
scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS);
}
}
}
void reportTransition(String target, HealthState next) {
HealthState previous = states.put(target, next);
if (previous == next) return;
if (fault(next)) {
log.warn("fleet health member={} state={} previous={}", target, next, previous);
} else if (previous != null && fault(previous)) {
log.info("fleet health member={} recovered state={} previous={}", target, next, previous);
}
// CB-580: a member entering GONE/NEVER_READY must not leave its waiting tickets pending
// forever. Fire exactly once per transition — never on a tick where the state is unchanged,
// which is what made the rejected commit call abandon() once per tick for as long as a
// member stayed terminal.
if (terminal(next)) {
failTerminalTarget(target, next);
}
}
private void failTerminalTarget(String target, HealthState state) {
String reason = "fleet health: member reached terminal state " + state.name();
RuntimeException last = null;
for (int attempt = 1; attempt <= MAX_FAIL_TARGET_ATTEMPTS; attempt++) {
try {
failTarget.accept(target, reason);
return;
} catch (RuntimeException error) {
last = error;
log.warn("fleet health: failTarget attempt {}/{} failed for member={} state={}",
attempt, MAX_FAIL_TARGET_ATTEMPTS, target, state, error);
}
}
log.warn("fleet health: giving up on failTarget for member={} state={} after {} attempts",
target, state, MAX_FAIL_TARGET_ATTEMPTS, last);
}
private static boolean terminal(HealthState state) {
return state == HealthState.GONE || state == HealthState.NEVER_READY;
}
private static boolean fault(HealthState state) {
return switch (state) {
case NEVER_READY, GONE, TURN_BOUNDARY_LOST, ERROR_ON_SCREEN, STALL_SUSPECTED,
MUTE, REPLY_STRANDED, DELEGATION_ORPHANED, CONTROL_LINK_DOWN -> true;
default -> false;
};
}
public static String coverage(boolean enabled, boolean notificationConfigured) {
return !enabled ? "off" : notificationConfigured ? "full" : "detection-only";
}
}
@@ -1,4 +0,0 @@
package dev.ltms.fleet.health;
/** Classification plus the private fact that the next pure decision needs. */
public record HealthDecision(HealthState state, HealthPrior prior) { }
@@ -1,6 +0,0 @@
package dev.ltms.fleet.health;
/** Private cross-tick observation. It is deliberately not a reported health value. */
public record HealthPrior(boolean busyButDone) {
public static final HealthPrior NONE = new HealthPrior(false);
}

Some files were not shown because too many files have changed in this diff Show More