Compare commits
44 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 871b595954 | |||
| b67b1585c2 | |||
| 979b2b5632 | |||
| cf4ad186ab | |||
| 5100f215cf | |||
| f129e9b7cd | |||
| 4aed45de19 | |||
| 1e6daa5c73 | |||
| 4c015d76b7 | |||
| 37a11cd168 | |||
| 22ad24db6c | |||
| e94c1b8841 | |||
| cc0ec65714 | |||
| d67d30c58a | |||
| d75ee1cca5 | |||
| 6804676a96 | |||
| 3e5d742ac7 | |||
| 6da2a71050 | |||
| e32ac39faf | |||
| daa243d37a | |||
| 2773ab600d | |||
| 19cdf8dc9f | |||
| 9daf1ec5ba | |||
| c9f0ca9359 | |||
| 84c8a2d2f0 | |||
| ded226abfe | |||
| 9b8d55bc18 | |||
| 11f8709286 | |||
| b034f105c0 | |||
| 6e37722383 | |||
| ffce30afa2 | |||
| e724a59f2d | |||
| 4bf855d225 | |||
| d4c9704007 | |||
| 7c252b5f5f | |||
| f756933879 | |||
| a1aecbf4fc | |||
| 131e7b1ccd | |||
| d0ac6c435f | |||
| 2bc5f3a057 | |||
| ba6b4a5da9 | |||
| da5a987df0 | |||
| 7dd6c46156 | |||
| 3a5cdc5108 |
@@ -1,24 +1,21 @@
|
||||
---
|
||||
name: implementer
|
||||
description: Implementer-role playbook for a bridged worker — you are in an isolated git worktree on a dedicated branch; implement the assigned task, commit, push, open your own PR to main, and hand off the PR URL via bridge_reply. You never merge. Load this when you have been delegated an implementation task over bridged.
|
||||
description: Implementer-role procedure for a bridged worker — verify your worktree, implement the scope, commit, push, open your own PR, and hand off the PR URL. Load this when the lead delegates you an implementation task over bridged.
|
||||
---
|
||||
|
||||
# Implementer worker
|
||||
# Implementer worker — procedure
|
||||
|
||||
You are an **implementer** in the claude-bridge fleet. The lead delegated you one scoped task,
|
||||
and you are running in an **isolated git worktree on your own branch** — a full peer of the
|
||||
primary (same `CLAUDE.md`, skills, memory, MCP), differing only in the model behind you and the
|
||||
branch you sit on. Your job for this turn: **implement the task, then hand off a PR the lead can
|
||||
review and merge.** You do the work; the lead (or human) is the merge gate — you never merge.
|
||||
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
|
||||
never merge, never commit `.mcp.json` or `wiki/`) is in **`CLAUDE.md` → Bridge communication →
|
||||
Worker** and already applies. This skill is only the *implement-and-hand-off procedure*.
|
||||
|
||||
Delivery mechanics (how the task reached you, how your reply resolves the lead's blocked send)
|
||||
are in [`docs/MCP-Contract.md`](../../../docs/MCP-Contract.md); the worktree/PR model is in
|
||||
[`docs/Worker-Git-Workflow.md`](../../../docs/Worker-Git-Workflow.md). You only need the steps
|
||||
below.
|
||||
You run in an **isolated git worktree on your own branch** — a full peer of the primary (same
|
||||
repo, `CLAUDE.md`, skills, MCP), differing in the model behind you and the branch you sit on.
|
||||
The worktree model is documented in [`docs/Worker-Git-Workflow.md`](../../../docs/Worker-Git-Workflow.md).
|
||||
|
||||
## 1. Confirm where you are — a worktree on a dedicated branch
|
||||
## 1. Confirm where you are
|
||||
|
||||
Before touching anything, verify your ground truth:
|
||||
Before touching anything:
|
||||
|
||||
```bash
|
||||
git rev-parse --show-toplevel # your worktree root — NOT the primary's main tree
|
||||
@@ -26,46 +23,40 @@ git branch --show-current # your dedicated branch: worker/<ticket>-<nonce>
|
||||
git status # should be clean at the start
|
||||
```
|
||||
|
||||
Do **all** work here, on this branch. **Never** switch to `main`, never `git checkout main`,
|
||||
never rebase onto or push to `main` directly. The branch is your isolation — respect it.
|
||||
Do **all** work here, on this branch. Never `git checkout main`, never rebase onto or push to
|
||||
`main`. The branch is your isolation — respect it.
|
||||
|
||||
## 2. Implement the task
|
||||
## 2. Implement
|
||||
|
||||
- Implement exactly the scope the lead named. Keep changes focused; if you notice something out
|
||||
of scope, note it in your reply rather than expanding the diff.
|
||||
- Match the surrounding code's style, naming, and idioms. Follow project `CLAUDE.md`.
|
||||
- **You cannot run the IDE MCP tools** (intellij-index / jetbrains are the primary's, not yours).
|
||||
So **never claim a file is "IDE-clean" or "diagnostics-clean"** — you cannot verify that. State
|
||||
only what you actually ran (e.g. `mvn`, a test) and its real output. A fabricated clean claim is
|
||||
worse than an honest "I could not verify inspections here."
|
||||
- Run whatever build/test you can and **report the true result** — including failures.
|
||||
- Implement exactly the scope the lead named. Keep the diff focused; note anything out of scope
|
||||
in your reply instead of widening it.
|
||||
- Match the surrounding code's style, naming, and idioms.
|
||||
- Run whatever build/test you can — `mvn clean install` from the module root. Read its **full**
|
||||
output; a piped `mvn ... | tail` hides failures.
|
||||
|
||||
## 3. Commit — focused, and never the excluded files
|
||||
## 3. Commit
|
||||
|
||||
```bash
|
||||
git add <the files you changed>
|
||||
git add <the files you changed> # explicitly — never `git add -A` / `git add .`
|
||||
git commit -m "<ticket>: <clear one-line summary>"
|
||||
```
|
||||
|
||||
**Excluded from every commit, always:** `.mcp.json` (the primary's local, session-modified copy —
|
||||
present only for parity) and `wiki/` (a separate submodule). Stage files explicitly; do **not**
|
||||
`git add -A` / `git add .` blindly, or you risk staging them. If `.mcp.json` shows as modified,
|
||||
leave it — it is flagged `--skip-worktree` and is not yours to commit.
|
||||
`.mcp.json` will show as modified. Leave it — it is `--skip-worktree` and not yours to commit.
|
||||
|
||||
## 4. Push your branch
|
||||
## 4. Push
|
||||
|
||||
```bash
|
||||
git push -u origin HEAD
|
||||
```
|
||||
|
||||
Push is over SSH as the same user — no extra credential needed. Push the branch as-is; do not
|
||||
force-push over anything you did not create.
|
||||
Push is over SSH as the same user — no extra credential needed. Never force-push over anything
|
||||
you did not create.
|
||||
|
||||
## 5. Open your own PR to `main`
|
||||
|
||||
Open the PR via the gitea REST API. The daemon injected a **repo-scoped token** (`GITEA_TOKEN`)
|
||||
and the forge host (`GITEA_HOST`) into your env for exactly this — the token can create a PR but
|
||||
**cannot merge** (that stays the lead/human gate).
|
||||
Via the gitea REST API. The daemon injected a **repo-scoped token** (`GITEA_TOKEN`) and the forge
|
||||
host (`GITEA_HOST`) into your env for exactly this — the token can create a PR but **cannot
|
||||
merge**.
|
||||
|
||||
```bash
|
||||
API="${GITEA_HOST%/}/api/v1/repos/lms/claude-bridge/pulls"
|
||||
@@ -81,16 +72,14 @@ JSON
|
||||
)"
|
||||
```
|
||||
|
||||
The response JSON includes `"html_url"` — that is your PR URL. If the call fails (non-2xx), read
|
||||
the error body, fix the cause if it is yours (e.g. branch not pushed yet), and report the failure
|
||||
honestly in your reply rather than inventing a URL. If `GITEA_TOKEN` is unset, your profile was
|
||||
not granted PR-create — push the branch (step 4) and report the branch name so the lead opens the
|
||||
PR.
|
||||
The response JSON carries `"html_url"` — that is your PR URL. On a non-2xx, read the error body,
|
||||
fix it if the cause is yours (e.g. branch not pushed yet), and report the failure rather than
|
||||
inventing a URL. If `GITEA_TOKEN` is unset your profile was not granted PR-create: push the branch
|
||||
and report its name so the lead opens the PR.
|
||||
|
||||
## 6. Reply via `bridge_reply` — the PR is the handoff
|
||||
## 6. Hand off — what goes in `bridge_reply`
|
||||
|
||||
End your turn with **exactly one** `bridge_reply`. That reply is the entire handoff — the lead
|
||||
cannot see your terminal. Include:
|
||||
The reply is the entire handoff; the lead cannot see your terminal.
|
||||
|
||||
```
|
||||
PR: <html_url from step 5, or "not created: <reason>" + branch name>
|
||||
@@ -100,27 +89,21 @@ tests: <what you ran and its REAL result — or "not run: <why>">
|
||||
summary: <2-3 lines: what you implemented and any caveat the reviewer needs>
|
||||
```
|
||||
|
||||
Then stop. **Do not merge. Do not touch `.mcp.json` or `wiki/`.** One reply closes the turn.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant L as Lead
|
||||
participant B as bridged
|
||||
participant I as Implementer (you)
|
||||
participant G as git / gitea
|
||||
|
||||
L->>B: bridge_send(task) — blocks
|
||||
B-->>I: your assignment (in a worktree on your branch)
|
||||
L->>I: delegated task (you are in a worktree on your branch)
|
||||
I->>I: implement + build/test here
|
||||
I->>G: git commit (never .mcp.json / wiki)
|
||||
I->>G: git push -u origin HEAD
|
||||
I->>G: POST /pulls (GITEA_TOKEN) — open PR to main
|
||||
G-->>I: html_url
|
||||
I->>B: bridge_reply(PR url, branch, files, tests)
|
||||
B-->>L: { outcome:"reply", text }
|
||||
Note over L,G: lead reviews the PR, merges on green — you never merge
|
||||
I->>L: bridge_reply(PR url, branch, files, tests)
|
||||
Note over L,G: the lead reviews the PR and merges on green — you never merge
|
||||
```
|
||||
|
||||
*The implement turn: work in the worktree, commit → push → open the PR, hand off the URL. The
|
||||
lead is the merge gate.*
|
||||
*The implement turn: work in the worktree, commit → push → open the PR, hand off the URL.*
|
||||
|
||||
@@ -1,69 +1,39 @@
|
||||
---
|
||||
name: reviewer
|
||||
description: Reviewer-role playbook for a bridged worker — read the assigned scope, find the real issues, ask the lead via bridge_ask when a decision is genuinely theirs, and report the finding via bridge_reply. Load this when you have been delegated a code review over bridged.
|
||||
description: Reviewer-role procedure for a bridged worker — how to work a review scope and the exact shape of the finding to report. Load this when the lead delegates you a code review over bridged.
|
||||
---
|
||||
|
||||
# Reviewer worker
|
||||
# Reviewer worker — procedure
|
||||
|
||||
You are a **reviewer** in the claude-bridge fleet. The lead delegated you one scoped review
|
||||
over `bridged`, and your whole job is **this single turn**: examine the scope it named, and
|
||||
report back. You are not the owner of the code and you do not merge anything — you surface
|
||||
what the owner needs to know, then hand the turn back.
|
||||
|
||||
Delivery mechanics (how the task reached you, how your reply resolves the lead's blocked
|
||||
send) are in [`docs/MCP-Contract.md`](../../../docs/MCP-Contract.md); you only need the three
|
||||
rules below.
|
||||
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
|
||||
never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and already applies. This
|
||||
skill is only the *review procedure*: how to work the scope, and the exact shape of what you
|
||||
send back.
|
||||
|
||||
## 1. Read the whole scope before you judge
|
||||
|
||||
The delegation names your scope — a file, a diff, a PR, a function. **Read all of it first.**
|
||||
A review that fires on a snippet misses the caller that makes it safe (or the one that makes
|
||||
it a bug). Reviewing only part of the scope and guessing the rest is the most common way a
|
||||
reviewer worker is wrong.
|
||||
A review that fires on a snippet misses the caller that makes it safe (or the one that makes it
|
||||
a bug). Reviewing part of the scope and guessing the rest is the most common way a reviewer is
|
||||
wrong.
|
||||
|
||||
## 2. Stay in your lane
|
||||
## 2. Stay in the scope
|
||||
|
||||
- Review **only** the assigned scope. If you notice something elsewhere, mention it in one
|
||||
line — do **not** go hunt it. Wandering is how two workers end up reporting the same thing
|
||||
and neither covers what it was given.
|
||||
- Do **not** edit files, run the build, or spawn other workers. You review; the owner acts.
|
||||
- You never set `ANTHROPIC_BASE_URL` and never touch herdr — you are a Claude Code process,
|
||||
not part of the transport.
|
||||
- Review **only** what you were assigned. Something elsewhere looks wrong? One line in your
|
||||
reply — do not go hunt it. Wandering is how two reviewers report the same thing and neither
|
||||
covers what it was given.
|
||||
- Do **not** edit files or run the build. You review; the owner acts.
|
||||
|
||||
## 3. When the decision is the lead's — ask, don't guess
|
||||
## 3. Reach for `bridge_ask` only for a genuine fork
|
||||
|
||||
Some things you cannot resolve from the code: an ambiguous requirement, a missing acceptance
|
||||
criterion, "is this behavior intended or a bug?", or a choice between two defensible fixes.
|
||||
Guessing there produces a confident-but-wrong finding. Instead **pause and ask the lead** with
|
||||
`bridge_ask` — a single crisp question. The call blocks; when the lead answers you **resume
|
||||
the same turn** with the answer and finish. Ask only when the answer changes your finding;
|
||||
don't narrate options you could decide yourself.
|
||||
Ambiguous requirement, a missing acceptance criterion, "intended or a bug?", or two defensible
|
||||
fixes with different consequences — those are the lead's call, and guessing produces a
|
||||
confident-but-wrong finding. Anything you could settle by reading more code is yours to settle.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant L as Lead
|
||||
participant B as bridged
|
||||
participant R as Reviewer (you)
|
||||
## 4. The finding — what goes in `bridge_reply`
|
||||
|
||||
L->>B: bridge_send(review scope) — blocks
|
||||
B-->>R: your assignment
|
||||
R->>R: read the full scope
|
||||
opt a decision only the lead can make
|
||||
R->>B: bridge_ask("intended, or a bug?") — you block
|
||||
B-->>L: { outcome:"question", turn_id }
|
||||
L->>B: bridge_send(answer, turn_id)
|
||||
B-->>R: { answer } — you resume the SAME turn
|
||||
end
|
||||
R->>B: bridge_reply(structured finding) — ends your turn
|
||||
B-->>L: { outcome:"reply", text }
|
||||
```
|
||||
|
||||
*The review turn, with the optional `bridge_ask` detour when the call is the lead's to make.*
|
||||
|
||||
## 4. Report with `bridge_reply` — one structured finding
|
||||
|
||||
End your turn with **exactly one** `bridge_reply`. Report the **single most important** real
|
||||
issue in the scope, in these four lines, under ~90 words:
|
||||
Report the **single most important** real issue in the scope, in these four lines, under
|
||||
~90 words:
|
||||
|
||||
```
|
||||
1. <path>:<line>
|
||||
@@ -72,13 +42,10 @@ issue in the scope, in these four lines, under ~90 words:
|
||||
4. severity: high | medium | low
|
||||
```
|
||||
|
||||
- Found nothing real after reading? Reply `NO ISSUE` and one line saying why — a clean review
|
||||
is a valid result, and a fabricated issue is worse than none.
|
||||
- **Nothing real after reading?** Reply `NO ISSUE` and one line saying why. A clean review is a
|
||||
valid result; a fabricated issue is worse than none.
|
||||
- **Severity:** `high` = wrong result, data loss, security, or a hang/crash on a real path ·
|
||||
`medium` = a real bug on an edge path, or a correctness risk under load/concurrency ·
|
||||
`low` = clarity, a latent foot-gun, or a smell with no current failure.
|
||||
- Be specific and verifiable: a line number and a one-line repro beat an adjective. If you
|
||||
can't point to where it goes wrong, you haven't found it yet.
|
||||
|
||||
One reply closes the turn. If you asked mid-turn, the answer you got is already folded into
|
||||
this finding — you do not ask again after replying.
|
||||
- Be specific and verifiable: a line number and a one-line repro beat an adjective. If you can't
|
||||
point at where it goes wrong, you haven't found it yet.
|
||||
|
||||
@@ -0,0 +1,55 @@
|
||||
name: CI
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
pull_request:
|
||||
branches: [main]
|
||||
|
||||
jobs:
|
||||
build:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
# The wiki submodule is docs only and is not needed to build — leave it unfetched so CI
|
||||
# does not depend on the wiki repo being reachable.
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
# The runner image ships an older default-jdk; bridged sets maven.compiler.release=25, so
|
||||
# provision the JDK explicitly rather than apt-installing whatever "default" means today.
|
||||
- name: Set up JDK 25
|
||||
uses: actions/setup-java@v4
|
||||
with:
|
||||
distribution: temurin
|
||||
java-version: '25'
|
||||
cache: maven
|
||||
|
||||
# setup-java provisions the JDK only — it does NOT install Maven, and the runner image has
|
||||
# no mvn on PATH (a bare `mvn` exits 127). Install it separately. apt pulls a default JRE as
|
||||
# a dependency; JAVA_HOME from setup-java still wins, which the version check below proves.
|
||||
- name: Install Maven
|
||||
run: |
|
||||
apt-get update && apt-get install -y --no-install-recommends maven
|
||||
mvn -version
|
||||
|
||||
- name: Build and test
|
||||
working-directory: bridged
|
||||
# This IS the mock-socket surface CB-503 asks for: the pom's `default-excludes` profile
|
||||
# already sets excludedGroups=contract, so the @Tag("contract") tests — which need a live
|
||||
# herdr socket and a RabbitMQ container — are excluded without any flag here. Everything
|
||||
# that runs does so against the fake UDS herdr and fake ccs/claude stubs.
|
||||
run: mvn -B clean install
|
||||
|
||||
# Deliberately NOT actions/upload-artifact: this Gitea instance presents as GHES, and
|
||||
# @actions/artifact v2+ (i.e. upload-artifact@v4) refuses to run there —
|
||||
# "GHESNotSupportedError ... not currently supported on GHES", which red-Xes an otherwise
|
||||
# green build. Since the artifact could not be retrieved anyway, dump the failing tests into
|
||||
# the log instead, where they are actually readable.
|
||||
- name: Failing test output
|
||||
if: failure()
|
||||
working-directory: bridged
|
||||
run: |
|
||||
for f in target/surefire-reports/*.txt; do
|
||||
[ -f "$f" ] || continue
|
||||
grep -qE "Failures: [1-9]|Errors: [1-9]" "$f" && { echo "===== $f ====="; cat "$f"; }
|
||||
done
|
||||
exit 0
|
||||
@@ -1,7 +1,196 @@
|
||||
# claude-bridge — project instructions
|
||||
|
||||
## Bridge communication (enforced — read this first)
|
||||
|
||||
> **Canonical block.** Everything down to §Layering is the portable bridge charter, copied verbatim
|
||||
> into every project that mounts the bridge MCP. Keep it byte-identical with the template in the
|
||||
> wiki ([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable
|
||||
> CLAUDE.md block*); improvements go to the template first, then out to each project. Anything
|
||||
> specific to *this* repo lives under §Project addendum below, never inline above it.
|
||||
|
||||
If no `bridge_*` MCP tools are mounted in this session, this section does not apply — skip it.
|
||||
|
||||
`bridged` is the **sole communication gateway** between agents here. The orchestrating session (the
|
||||
**primary**) and every delegated peer (a **worker**) mount the *same* MCP server and talk only
|
||||
through its `bridge_*` tools. No session addresses a peer, a broker, or the network directly.
|
||||
|
||||
### Which role am I? — settle this before acting
|
||||
|
||||
**Both roles read this file.** A worker runs in a git worktree of this same repo, so it inherits
|
||||
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
|
||||
|
||||
**Call `bridge_whoami`.** It returns `{"role":"primary"}` or `{"role":"worker","sessionId":…,
|
||||
"profile":…,"worktree":…,"branch":…}`, resolved by the daemon from your connection — unforgeable,
|
||||
and the same resolution its authorization gate uses. Don't infer what you can ask.
|
||||
|
||||
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
|
||||
fires: the reply charter in your system prompt (*"You are an off-subscription worker in the
|
||||
claude-bridge fleet"*) ⇒ **worker**; bridge tools prefixed `mcp__bridge__*` ⇒ **worker** (the
|
||||
launcher fixes that mount name; a primary's mount is named by whoever wrote its `.mcp.json`, so it
|
||||
varies); `ANTHROPIC_BASE_URL` set ⇒ **worker** (Claude-model workers run on a clean env, so its
|
||||
*absence* proves nothing). **Still unsure ⇒ act as a worker.** The two mistakes are not symmetric: a
|
||||
primary acting as a worker is refused by the authorization gate — loud and self-correcting — while a
|
||||
worker acting as the primary ends its turn with no `bridge_reply`, and the sender silently receives
|
||||
nothing. Fail toward the recoverable error.
|
||||
|
||||
### Invariants — both roles, no exceptions
|
||||
|
||||
1. **Never set, export, or forward `ANTHROPIC_BASE_URL`** (or `ANTHROPIC_AUTH_TOKEN`). The primary
|
||||
stays on subscription; only the bridge puts a worker off it, at spawn. Mounting the bridge must
|
||||
never move a session across that boundary.
|
||||
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
|
||||
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
|
||||
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
|
||||
cannot act as another session. Spawn/stop/send/drain are primary-only; reply/ask are
|
||||
worker-only-and-only-as-itself. A call outside your role is refused, not queued.
|
||||
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
|
||||
re-send because a call looks slow — the bridge delivers when the peer is `idle`/`blocked`.
|
||||
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
|
||||
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
|
||||
|
||||
### Primary (lead) — run this on every task, in order
|
||||
|
||||
**Delegate by default — that is the job.** With the bridge mounted you are an orchestrator on a
|
||||
metered subscription, and workers are cheap, parallel, and disposable. The default answer to "who
|
||||
does this?" is **a worker**, not you. Reach for `bridge_send` before you reach for `Edit`. The steps
|
||||
below are the procedure — run them in order, every task, not only the big ones.
|
||||
|
||||
0. **Know your role** — `bridge_whoami`, once per session, before anything else.
|
||||
1. **Split.** Write the unit list. Every unit carries: scope · the files or PR in question ·
|
||||
acceptance criteria · exactly what to report back. A unit with no acceptance criteria is not
|
||||
ready to delegate — refine it or keep it.
|
||||
2. **Gate each unit** on one question: **"can I write a brief good enough for a worker to
|
||||
succeed?"** — *not* "could I do this faster myself?" (usually you could; doing it yourself costs
|
||||
your context and your subscription, while a wasted worker turn costs a worker turn). Yes ⇒
|
||||
delegate. The keep-list is closed: the conversation with the user, decomposition and planning,
|
||||
the final judgment call, verification, merges, and anything that depends on context only you
|
||||
hold. Nothing else is yours by default.
|
||||
3. **Spawn every delegated unit first** — `bridge_spawn{profile, worktree:true, ticket}`, one per
|
||||
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model and cost, not in
|
||||
tier, so the default is rarely what you want.
|
||||
4. **Then send them all** — `bridge_send{sessionId, content, wait:false}`. Line 1 of every brief is
|
||||
`Load the <name> skill.` naming the worker's playbook; those skills are opt-in and that line is
|
||||
what makes them reliable. Where the project ships no such skill, spell the procedure out in the
|
||||
brief instead. The brief is self-contained — the worker sees your message and the repo, nothing
|
||||
of your context, your plan, or your screen.
|
||||
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
|
||||
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
|
||||
`bridge_status`, never by reading its terminal.
|
||||
6. **Verify yourself.** Re-run the build and the checks. A worker mounts only the bridge MCP and
|
||||
cannot run your other tooling, and a piped command (`… | tail`) hides failures behind a zero
|
||||
exit — never promote a worker's "clean" to a fact.
|
||||
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
|
||||
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
|
||||
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
|
||||
as it lands; don't wait for the last implementer. Under ~50 changed lines, skip the fan-out and
|
||||
read it yourself.
|
||||
8. **Adjudicate, merge, tear down — yours alone.** Read the diff yourself: fully if it is small,
|
||||
targeted at the reported findings and the risky paths if it is large. Reviewer findings direct
|
||||
your attention; they never substitute for it. Then merge, then `bridge_stop{paneId}`.
|
||||
|
||||
**Steps 3 and 4 are separate on purpose** — spawning and sending in one loop is how parallel work
|
||||
silently becomes serial, and it is the most common way this layer is wasted. For the same reason,
|
||||
prefer `wait:false` + `bridge_poll` for anything non-trivial: a blocking `bridge_send` is capped by
|
||||
*your own* MCP client call timeout (~60s), well below the task's real runtime.
|
||||
|
||||
**Delegating does not delegate responsibility.** Workers open PRs; you are the gate. Never delegate
|
||||
the merge — and merging on a reviewer's word is delegating it by proxy.
|
||||
|
||||
| Intent | Tool |
|
||||
|---|---|
|
||||
| Confirm your own role | `bridge_whoami` |
|
||||
| See backends available | `bridge_profiles` |
|
||||
| Start a worker | `bridge_spawn{profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
|
||||
| See the fleet | `bridge_list` · one worker's state: `bridge_status{sessionId}` |
|
||||
| Delegate (blocking) | `bridge_send{sessionId, content}` |
|
||||
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
|
||||
| Answer a worker's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
|
||||
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
|
||||
| Tear down | `bridge_stop{paneId}` |
|
||||
|
||||
### Worker — the turn contract
|
||||
|
||||
1. **Load the playbook skill the lead named** before doing anything else.
|
||||
2. **Do the assigned scope only.** Note anything you spot outside it in one line; don't go hunt it.
|
||||
3. **`bridge_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
|
||||
defensible fixes, "bug or intended?"). It blocks and you resume the *same* turn with the answer.
|
||||
Don't ask what you could decide yourself.
|
||||
4. **End the turn with exactly one `bridge_reply{content}`**, carrying your complete answer. This is
|
||||
the whole handoff. No `bridge_reply` ⇒ the sender gets nothing and the exchange stalls.
|
||||
5. **Report honestly.** State only what you actually ran and its real output, including failures.
|
||||
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
|
||||
so never claim the result of a check you had no way to run.
|
||||
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
|
||||
project marks as not-yours-to-commit.
|
||||
|
||||
### Where each rule lives (don't duplicate — extend the right layer)
|
||||
|
||||
| Layer | Scope | Reaches |
|
||||
|---|---|---|
|
||||
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every worker, at launch, every peer kind |
|
||||
| **this section** | protocol + orchestration policy | primary **and** every Claude worker — tracked in git, so worktrees inherit it |
|
||||
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a worker told to load one |
|
||||
| the bridge's own docs | design detail, flows, error model | on demand |
|
||||
|
||||
A rule belongs in **exactly one** layer — the outermost one that must obey it. Peers that don't read
|
||||
`CLAUDE.md` (non-Claude adapters) get the charter only, so any rule *they* must obey belongs in the
|
||||
charter, not here.
|
||||
|
||||
## Project addendum — claude-bridge (not part of the canonical block)
|
||||
|
||||
- **This repo is the bridge.** The daemon is `bridged`, its MCP mount is `http://127.0.0.1:8765/mcp`,
|
||||
and the code behind the rules above is `mcp/BridgeMcp` (tools), `auth/Authz` (the role table),
|
||||
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
|
||||
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
|
||||
`reviewer` (scoped review → one structured finding). Name one in every delegation.
|
||||
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
|
||||
(a submodule with its own remote).
|
||||
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
|
||||
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
|
||||
session's context.
|
||||
|
||||
### The prompt is part of the product — update it with the code (mandatory)
|
||||
|
||||
This repo *is* the bridge, so the canonical block above is not documentation about someone else's
|
||||
system: it is the instruction surface this codebase ships. **Every change here must end by asking
|
||||
whether the block still tells the truth.** A code change that silently invalidates it is an
|
||||
incomplete change — the agents reading it have no other source.
|
||||
|
||||
Before you call any work done, check the row that matches what you touched:
|
||||
|
||||
| You changed… | Re-read and update… |
|
||||
|---|---|
|
||||
| a `bridge_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
|
||||
| `Authz` / the role table | invariant 3, and the primary-only vs worker-only claims |
|
||||
| `ConnectionIdentity` / how a caller is resolved | the `bridge_whoami` paragraph and the fallback ladder |
|
||||
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__bridge__*`), and the layering table's top row |
|
||||
| the injector / status gating | invariant 4 |
|
||||
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
|
||||
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
|
||||
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
|
||||
|
||||
Then **propagate**: the block in this file and the template in the wiki
|
||||
([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
|
||||
block*) must stay byte-identical, and other projects carrying the block need the same edit. Verify
|
||||
rather than trust:
|
||||
|
||||
```bash
|
||||
python3 - <<'PY'
|
||||
import pathlib
|
||||
c = pathlib.Path("CLAUDE.md").read_text()
|
||||
w = pathlib.Path("wiki/7-Use-Cases.md").read_text()
|
||||
S, E = "## Bridge communication (enforced", "## Project addendum — claude-bridge"
|
||||
block = c[c.index(S):c.index(E)].rstrip() + "\n"
|
||||
i = w.index("```markdown\n") + len("```markdown\n")
|
||||
print("in sync:", w[i:w.index("\n```\n", i) + 1] == block)
|
||||
PY
|
||||
```
|
||||
|
||||
## IDE MCP tools & validation workflow (enforced)
|
||||
|
||||
> **Primary only.** Workers have no IDE MCP mount — if you are a worker, skip this section and
|
||||
> report the build/test output you actually ran (see §Bridge communication → Worker).
|
||||
|
||||
Two IDE MCP servers are connected: **intellij-index** (semantic code intelligence) and
|
||||
**jetbrains** (file problems, reformat, debugger). IntelliJ has multiple projects open; our
|
||||
module is **`bridged`**. Always pass these to IDE MCP tools:
|
||||
|
||||
@@ -92,7 +92,8 @@ bridge (code reviews delegated this way have produced committed bug fixes). Sele
|
||||
primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a
|
||||
fallback injector.
|
||||
|
||||
**Shipped** (Java 25 · Maven · 105 tests green — unit/acceptance + live-herdr contract tests):
|
||||
**Shipped** (Java 25 · Maven · 266 unit/acceptance tests green; the live-herdr and broker contract
|
||||
tests run separately via `mvn test -Pcontract`):
|
||||
|
||||
- **Core gateway** — herdr socket client (contract-tested vs live 0.7.0); guard-checked worker
|
||||
spawn with `ANTHROPIC_BASE_URL` injected only into the worker's env; status-gated injector;
|
||||
@@ -106,7 +107,17 @@ fallback injector.
|
||||
- **Fleet** — multiple worker profiles, each with an independent base_url guard check; workers
|
||||
inherit the primary's working directory (never `$HOME`); a readiness gate holds delivery until
|
||||
a worker's Claude has connected the bridge MCP (no paste lost into its boot window).
|
||||
- **Blocked-worker path** — `bridge_ask` reverse rendezvous: a worker pauses its delegated turn to
|
||||
ask the primary and resumes the *same* turn with the answer (CB-205).
|
||||
- **Session lifecycle** — session manager with spawn/reuse/recycle, `idle_ttl` reaper, `context_cap`,
|
||||
and graceful drain on shutdown (CB-301/CB-303); per-worker git worktrees on their own branch with
|
||||
a config-parity overlay, so parallel implementers never stomp each other (CB-301-ext).
|
||||
- **Reliable worker→primary delivery** — a durable `ReplyInbox` (in-memory by default, AMQP/LavinMQ
|
||||
for cross-restart durability) holds a reply that arrives with no open send, and an active
|
||||
status-gated push loop nudges the primary to drain it (CB-307).
|
||||
- **Pluggable peers** — a `PeerLauncher` SPI with two in-tree adapters, `claude-code` and `opencode`,
|
||||
routed by a `kind:` discriminator (CB-401/CB-402).
|
||||
|
||||
**Next** (see the [roadmap](wiki/8-Roadmap.md)) — structured envelope schema, `bridge_ask`
|
||||
(blocked-worker path), session lifecycle / recycle / `idle_ttl`, split-host, and hardening
|
||||
(auth/TLS, `/metrics`, CI, systemd).
|
||||
**Next** (see the [roadmap](wiki/8-Roadmap.md)) — Stage 5 hardening (auth/TLS, `/metrics`, CI,
|
||||
service supervision, per-session authz + audit), then cross-host: CB-308 multi-host federation and
|
||||
CB-500 multi-tier coordination.
|
||||
|
||||
@@ -5,6 +5,9 @@ dependency-reduced-pom.xml
|
||||
# Local runtime config (copy from bridged.example.yaml)
|
||||
bridged.yaml
|
||||
|
||||
# CB-505 audit trail + daemon stdout/stderr — runtime records, never source
|
||||
logs/
|
||||
|
||||
# Editor / OS
|
||||
*.iml
|
||||
.idea/
|
||||
|
||||
@@ -3,11 +3,34 @@
|
||||
# bridged is the sole gateway between primary/worker Claude sessions and herdr.
|
||||
# It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL.
|
||||
|
||||
# REST + MCP listen address. Keep it on loopback — bridged is same-host in Stage-1.
|
||||
# REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token`
|
||||
# below — bridged REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
|
||||
# API authentication (CB-501). Governs how a caller that is NOT an on-host worker pane proves it
|
||||
# is the primary. Worker identity never depends on this: a loopback peer PID that maps to a herdr
|
||||
# pane is unforgeable and is always honoured, so turning auth on cannot lock the fleet out.
|
||||
#
|
||||
# mode: loopback-trust → DEFAULT, and the historical behaviour: any loopback caller that is not
|
||||
# a worker is the primary, no credential needed. Sound ONLY because the
|
||||
# OS refuses remote connections to a loopback socket.
|
||||
# mode: token → such a caller must send `Authorization: Bearer <token>`; without it it
|
||||
# is anonymous and authorized for nothing. REQUIRED for a non-loopback
|
||||
# bind — the daemon fails fast otherwise, because "unauthenticated ⇒
|
||||
# primary" on a reachable port would hand spawn/stop/send to anyone.
|
||||
# tokenEnv → host env var holding the token (never the literal value). Default
|
||||
# BRIDGED_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
|
||||
#
|
||||
# TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and
|
||||
# let it own certificate lifecycle, e.g.
|
||||
# location / { proxy_pass http://127.0.0.1:8765; proxy_set_header Authorization $http_authorization; }
|
||||
# The broker link gets TLS from its own URI (amqps://…) — see `broker` below.
|
||||
# auth:
|
||||
# mode: token
|
||||
# tokenEnv: BRIDGED_API_TOKEN
|
||||
|
||||
# herdr Unix socket. Omit to use the client default
|
||||
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
@@ -26,9 +49,38 @@ herdrSocket: ~/.config/herdr/herdr.sock
|
||||
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
|
||||
# cwd on an MCP spawn, else the daemon's cwd — never $HOME. See
|
||||
# docs/Worker-Startup-and-Trust.md.
|
||||
# configDir → CLAUDE_CONFIG_DIR for the worker, so it inherits that profile's
|
||||
# skills/MCP/hooks. Omit to leave the worker on the host default.
|
||||
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
|
||||
# worktree sees the same local config (CB-301-ext). Omit for the default set:
|
||||
# [.mcp.json, .claude/settings.local.json, .env, .envrc].
|
||||
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
|
||||
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
|
||||
# Opt-in by design — omit and the worker gets no PR-create grant (push over
|
||||
# SSH is unaffected). The token value itself is never stored in this file.
|
||||
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
|
||||
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
|
||||
# env → extra environment for this profile's workers, as a literal key/value map
|
||||
# (CB-511). Use it to give workers a toolchain.
|
||||
#
|
||||
# A worker's environment does NOT come from your shell. bridged hands herdr an
|
||||
# explicit env map and herdr merges it into ITS OWN process env — so before
|
||||
# CB-511 a worker inherited whatever PATH the herdr server happened to be
|
||||
# started with, which on a long-lived herdr can predate your toolchain entirely
|
||||
# and leave workers unable to run `mvn` or `java` at all.
|
||||
# bridged now propagates ITS OWN PATH to every worker by default; set `env:`
|
||||
# only to override that or add more (JAVA_HOME, …). Since the default is the
|
||||
# daemon's PATH, make sure the daemon is started with a good one — see the PATH
|
||||
# lines in deploy/dev.ltms.bridged.plist and deploy/bridged.service.
|
||||
#
|
||||
# Adapter-owned variables always win over `env:`: ANTHROPIC_BASE_URL and the
|
||||
# rest of the ANTHROPIC_*/CLAUDE_* wiring are applied after it, so an `env:`
|
||||
# entry cannot repoint a worker past the SubscriptionGuard — which is checked
|
||||
# against `baseUrl` alone.
|
||||
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
|
||||
workers:
|
||||
gx10: # ccs profile name (NOT a hostname)
|
||||
kind: claude-code # which adapter spawns this profile (default; may omit)
|
||||
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
|
||||
model: coder
|
||||
placement: tab
|
||||
@@ -37,6 +89,11 @@ workers:
|
||||
mcpUrl: http://127.0.0.1:8765/mcp
|
||||
tokenEnv: BRIDGED_WORKER_TOKEN
|
||||
argv: ["ccs", "gx10"]
|
||||
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
|
||||
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
|
||||
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
|
||||
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
|
||||
# parityOverlay: [".mcp.json", ".claude/settings.local.json", ".env", ".envrc"]
|
||||
ollama:
|
||||
baseUrl: http://ollama.ltms.dev # local/self-hosted; usually no token
|
||||
placement: tab
|
||||
@@ -44,6 +101,49 @@ workers:
|
||||
tabLabel: "worker: {profile} #{n}"
|
||||
mcpUrl: http://127.0.0.1:8765/mcp
|
||||
argv: ["ccs", "ollama"]
|
||||
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
|
||||
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
|
||||
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
|
||||
# The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a
|
||||
# `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude.
|
||||
#
|
||||
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → bridge_send →
|
||||
# structured bridge_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
|
||||
# and need NO credentials — check `opencode models` for the current free list, since the names
|
||||
# change. That also makes the worker off-subscription by construction.
|
||||
# opencode-free:
|
||||
# kind: opencode
|
||||
# model: opencode/north-mini-code-free # `provider/model` selector, injected as `-m`
|
||||
# placement: tab
|
||||
# workspace: bridged-workers
|
||||
# tabLabel: "opencode: {profile} #{n}"
|
||||
# mcpUrl: http://127.0.0.1:8765/mcp
|
||||
# argv: ["opencode"]
|
||||
#
|
||||
# CB-508: point an opencode profile at your OWN OpenAI-compatible endpoint (local vLLM, llama.cpp,
|
||||
# LM Studio, TGI…) instead of opencode's gateway. Setting `baseUrl` on a `kind: opencode` profile
|
||||
# makes the bridge emit a custom `provider` block into the generated opencode.json — opencode has
|
||||
# no ANTHROPIC_BASE_URL seam, so this is how the endpoint is pinned.
|
||||
# baseUrl → a bare host:port gets `/v1` appended (where these servers mount the API); a URL that
|
||||
# already has a path is used verbatim, so a custom mount point still works.
|
||||
# model → MUST be "<provider>/<model>". The provider half names the generated block; the model
|
||||
# half must match an id the server reports at /v1/models. One field drives both the
|
||||
# declaration and the `-m` flag, so they cannot drift apart. A bare model name with a
|
||||
# baseUrl set is rejected at spawn rather than silently using the default gateway.
|
||||
# tokenEnv → optional; its value becomes the provider apiKey. Most local servers ignore the key,
|
||||
# so a placeholder is used when unset (the AI SDK still requires a non-empty one).
|
||||
# NOTE: no `guard` entry is needed even with a baseUrl set. The SubscriptionGuard exists to stop a
|
||||
# worker borrowing the primary's Anthropic subscription, and an opencode process has no Anthropic
|
||||
# credential path at all.
|
||||
# opencode-local:
|
||||
# kind: opencode
|
||||
# baseUrl: http://127.0.0.1:8000
|
||||
# model: local-vllm/deepseek-v4-flash
|
||||
# placement: tab
|
||||
# workspace: bridged-workers
|
||||
# tabLabel: "opencode: {profile} #{n}"
|
||||
# mcpUrl: http://127.0.0.1:8765/mcp
|
||||
# argv: ["opencode"]
|
||||
defaultWorker: gx10
|
||||
|
||||
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
|
||||
@@ -54,6 +154,18 @@ guard:
|
||||
- gx01.gw
|
||||
- ollama.ltms.dev
|
||||
|
||||
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
|
||||
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
|
||||
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
|
||||
# unknown keys are ignored, so a snake_case key would be silently dropped (default kept).
|
||||
# spawnReadyTimeoutMs: 20000
|
||||
# spawnReadyPollMs: 300
|
||||
|
||||
# Worktree provisioning root (CB-301-ext). Where per-worker git worktrees are checked out so
|
||||
# each worker owns an isolated branch instead of sharing the primary's tree. Omit to default
|
||||
# to a sibling directory of the repo root.
|
||||
# worktreeRoot: /Users/me/src/.bridged-worktrees
|
||||
|
||||
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
|
||||
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
|
||||
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
|
||||
@@ -63,3 +175,29 @@ guard:
|
||||
# idleTtlSeconds: 300
|
||||
# contextCap: 10
|
||||
# drainTimeoutSeconds: 5
|
||||
|
||||
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
|
||||
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
|
||||
# Set a broker uri to swap in the AMQP-backed inbox: worker replies with no open send are held
|
||||
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
|
||||
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
|
||||
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
|
||||
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
|
||||
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
|
||||
# broker:
|
||||
# uri: amqp://guest:guest@127.0.0.1:5672
|
||||
|
||||
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
|
||||
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
|
||||
# pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop
|
||||
# stops as soon as the primary's inbox is empty.
|
||||
# terminal → pin the primary's herdr terminal id. Omit to learn it from the connection on
|
||||
# the first orchestration-side MCP call (the normal case). An off-host or
|
||||
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
|
||||
# degrades to pull; the reply is still never lost.
|
||||
# pushReminders → max nudges before giving up (default 5)
|
||||
# pushBackoffMs → delay between nudges in ms (default 15000)
|
||||
# primary:
|
||||
# terminal: term_65619bd6174568
|
||||
# pushReminders: 5
|
||||
# pushBackoffMs: 15000
|
||||
|
||||
@@ -0,0 +1,142 @@
|
||||
# CB-307 — Active push-to-primary + reminder loop (the reliability layer)
|
||||
|
||||
**Status:** design (2026-07-19). Builds directly on the shipped durable landing zone
|
||||
(`AmqpReplyInbox`, main `2bc5f3a`, dogfooded live). gitea #5.
|
||||
|
||||
## Why this exists
|
||||
|
||||
Stage 2 gave a worker→primary reply a **durable place to wait** when no `bridge_send` is
|
||||
open: it lands in `agent.<target>.inbox` on the broker and survives a daemon bounce. But
|
||||
delivery is still **pull** — the primary only sees the reply if it happens to call
|
||||
`bridge_poll(target)` / `GET /sessions/{id}/replies`. A reply can sit indefinitely while
|
||||
the primary works on something else.
|
||||
|
||||
This layer makes delivery **active**: the bridge *pushes* a nudge to the primary the moment
|
||||
a reply lands, and keeps reminding (bounded) until the primary drains it. At-least-once,
|
||||
dedup by `msgId`, and — critically — it never loses the reply even if every push fails,
|
||||
because the durable inbox is the backstop.
|
||||
|
||||
## The hard constraint it works around
|
||||
|
||||
The bridge is an MCP **server**; the primary is an MCP **client**. A server cannot call
|
||||
into a client. So "push to the primary" cannot be an MCP response — it needs a *sideband*
|
||||
channel. The chosen channel: **inject a synthetic user-turn into the primary's own herdr
|
||||
terminal pane** — the same mechanism the bridge already uses to deliver tasks to workers,
|
||||
pointed at the primary's pane instead.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
W["worker"] -->|"bridge_reply (no open send)"| MS["MessageService.reply"]
|
||||
MS -->|"inbox.publish"| INBOX[("agent.<target>.inbox<br/>(durable, LavinMQ)")]
|
||||
MS -->|"notify"| LOOP["ReplyPushLoop"]
|
||||
LOOP -->|"status-gated inject"| PANE["primary's herdr pane"]
|
||||
PANE -->|"primary drains"| DRAIN["bridge_poll(target)<br/>= peek + ack"]
|
||||
DRAIN -->|"inbox now empty"| LOOP
|
||||
LOOP -.->|"still non-empty →<br/>re-inject on backoff"| PANE
|
||||
classDef store fill:#2c5282,stroke:#1a365d,color:#ffffff;
|
||||
class INBOX store
|
||||
```
|
||||
|
||||
*Figure 1 — a reply lands in the durable inbox; the push loop nudges the primary's pane;
|
||||
the primary's drain acks it; a still-full inbox triggers a bounded re-nudge.*
|
||||
|
||||
## Three increments
|
||||
|
||||
### Increment 1 — learn & store the primary's terminal_id
|
||||
|
||||
**Finding (seam map):** `ConnectionIdentity.resolve(remoteAddr, remotePort)` already returns
|
||||
the caller's herdr `terminal_id` for *every* MCP call, via `PaneLocator.terminalForPid`
|
||||
(walks `pane.list`, matches the caller PID to a pane's process tree). It is non-null whenever
|
||||
the caller runs in a herdr pane on this host. Today it's discarded for the primary
|
||||
(`presence.markPresent` is a no-op on it).
|
||||
|
||||
**Plan:** a single-slot `PrimaryRegistry` (thread-safe) holding the primary's `terminal_id`.
|
||||
Populate it from the **orchestration-side** MCP tools — `bridge_send`, `bridge_spawn`,
|
||||
`bridge_poll`, `bridge_list`, `bridge_status`, `bridge_profiles` — capturing
|
||||
`callerTerminal(exchange)` when it is (a) non-null and (b) **not** a registered worker
|
||||
session in `SessionManager`. That caller is, by construction, the primary. Worker-side tools
|
||||
(`bridge_reply`, `bridge_ask`) never set it.
|
||||
|
||||
- **Config override / pin:** a `primary: { terminal: "<id>" }` block in `BridgedConfig`
|
||||
(nested record, same shape as `Broker`). Lets an operator pin it, or supply it when
|
||||
derivation can't (see degrade case).
|
||||
- **Degrade:** if the primary is off-host or in a non-herdr terminal, `terminalForPid`
|
||||
returns null and no override is set → **the registry stays empty → the push loop is a
|
||||
no-op and we fall back to pull** (today's behaviour). The reply is never lost; it's just
|
||||
not actively pushed. This is a safe, explicit degradation, not a failure.
|
||||
|
||||
### Increment 2 — the push loop
|
||||
|
||||
A `ReplyPushLoop` component, notified at the single no-waiter call site
|
||||
(`MessageService.reply` → the `inbox.publish` branch, `MessageService.java:192`).
|
||||
|
||||
- **Inject a nudge, not the payload.** The injected turn tells the primary *to drain*
|
||||
(e.g. "Worker `<target>` returned a reply — run `bridge_poll(target=<target>)` to collect
|
||||
it"), it does **not** carry the reply text. Rationale: replies can be large/multiline and
|
||||
terminal injection would mangle them; the drain response is the clean transport. Keeps the
|
||||
push idempotent — re-nudging is harmless.
|
||||
- **Ack = drain.** The primary draining (`drainReplies` = peek + ack) is the acknowledgement.
|
||||
The loop's **stop condition is `inbox.peek(target).isEmpty()`** — the reply is gone from the
|
||||
inbox because it was acked. No new `bridge_ack` tool needed for v1 (see Increment 3).
|
||||
- **Status-gated injection (mechanism (b), chosen).** A dedicated lightweight scheduled loop,
|
||||
**not** the worker `Injector`. It injects via `AgentControl.send(primaryTerminal, nudge)`
|
||||
(the same herdr `agent.send` = `pane send-text` + submit that delivers to workers) only when
|
||||
`AgentControl.status(primaryTerminal).injectable()` (IDLE/BLOCKED) — never mid-turn. This keeps
|
||||
the primary path fully isolated from `WorkerPresence`/`StatusPoller` (which are worker-scoped),
|
||||
and makes it unit-testable with a fake `AgentControl` + an injected clock (per the CB-306
|
||||
`LongSupplier` clock + `Runnable` sleeper seam). Rejected (a) reuse-the-Injector: it would force
|
||||
the primary terminal into the worker poller set and couple to worker-presence semantics — more
|
||||
integration surface, harder to test, no real gain for a bounded reminder.
|
||||
- **Bounded reminder / backoff.** While `peek(target)` stays non-empty, re-inject on a
|
||||
backoff schedule up to a cap (N reminders or a max duration; config
|
||||
`primary.push_reminders` / `primary.push_backoff_ms`). After the cap, **stop reminding** —
|
||||
the reply remains in the durable inbox and the next natural poll (or a later worker reply's
|
||||
nudge) still surfaces it. Bounded so the bridge never spams the primary.
|
||||
|
||||
### Increment 3 — optional per-`msgId` `bridge_ack` tool (deferred)
|
||||
|
||||
Drain-as-ack is coarse: it clears *all* pending replies for a target at once. If finer
|
||||
control is ever needed (ack one reply, leave others held), add a `bridge_ack(msgId)` tool
|
||||
mapping to `inbox.ack(target, msgId)` — the port already supports per-`msgId` ack. Not built
|
||||
in v1; the stop-on-empty loop is sufficient.
|
||||
|
||||
## The two subtleties (decided here)
|
||||
|
||||
1. **Which caller is "the primary"?** Connection-derived, not self-reported: the caller whose
|
||||
resolved terminal is non-null **and not a registered worker session**, seen on an
|
||||
orchestration-side tool. This never mislabels a worker (workers are in `SessionManager`)
|
||||
and needs no new env var or argument (identity stays connection-derived, per the existing
|
||||
`BridgeMcp` invariant).
|
||||
|
||||
2. **Readiness-gate mismatch → dedicated loop.** The existing `Injector` gates delivery on
|
||||
`ready.test(target)` = `WorkerPresence` (the *worker's* MCP connected). The primary is not
|
||||
in `WorkerPresence`, so reusing `Injector` would mean forcing the primary terminal into the
|
||||
worker `StatusPoller` set and swapping the `ready` predicate — extra integration surface with
|
||||
worker-scoped machinery. Decision: **mechanism (b)** — a small dedicated scheduled loop that
|
||||
calls `AgentControl.status(primaryTerminal).injectable()` then `AgentControl.send(...)`, with
|
||||
an injected clock. Isolated from worker presence, trivially unit-testable, sufficient for a
|
||||
bounded reminder. (Verified live: this primary resolves to `term_656c8cc03e1f0b1`, pane
|
||||
`w2:pY` — the primary genuinely runs in a herdr pane on this host, so the path is exercisable.)
|
||||
|
||||
## Boundary note
|
||||
|
||||
This is the first time the bridge **writes into the primary's pane** — a new direction of
|
||||
control. It stays within the communication-bus identity: the injection is a **nudge** (a
|
||||
synthetic "go drain your replies" turn), **status-gated** so it never interrupts a turn,
|
||||
**bounded** so it never spams, carries **no env** and **never crosses the subscription
|
||||
boundary**. The bridge is signalling the primary that it has mail — not driving its work.
|
||||
|
||||
## Test plan
|
||||
|
||||
- **Unit (hermetic):** `PrimaryRegistry` set/clear/override; the "caller is primary iff
|
||||
non-null terminal AND not a registered session" predicate; the loop's stop-on-empty and
|
||||
bounded-reminder logic with an injected clock + a fake injector (no real herdr).
|
||||
- **Live dogfood (primary-side):** with the daemon on the broker jar + a real worker,
|
||||
delegate a task, let the worker reply after the `bridge_send` window closes, and observe the
|
||||
bridge inject a drain nudge into *this* primary pane; confirm draining stops the reminders;
|
||||
confirm an unreachable primary (registry empty) degrades to pull with no loss.
|
||||
|
||||
## Out of scope
|
||||
|
||||
Multi-host (CB-308) — the push loop is local-only; a remote primary is reached by its own
|
||||
local gateway, not cross-host injection. Federation reuses this loop per-gateway.
|
||||
@@ -24,6 +24,10 @@
|
||||
<slf4j.version>2.0.16</slf4j.version>
|
||||
<logback.version>1.5.18</logback.version>
|
||||
<junit.version>5.11.4</junit.version>
|
||||
<amqp.version>5.22.0</amqp.version>
|
||||
<testcontainers.version>1.20.4</testcontainers.version>
|
||||
<commons-compress.version>1.27.1</commons-compress.version>
|
||||
<commons-lang3.version>3.18.0</commons-lang3.version>
|
||||
</properties>
|
||||
|
||||
<!--
|
||||
@@ -62,6 +66,22 @@
|
||||
<artifactId>jackson-annotations</artifactId>
|
||||
<version>3.0-rc5</version>
|
||||
</dependency>
|
||||
<!-- Testcontainers 1.20.4 pulls commons-compress 1.24.0 (test scope), which carries
|
||||
CVE-2024-25710 (8.1) + CVE-2024-26308 — both fixed in 1.26.0. Pin the patched line.
|
||||
Test-scope only (never shipped in the jar), but bumped per the CVE policy. -->
|
||||
<dependency>
|
||||
<groupId>org.apache.commons</groupId>
|
||||
<artifactId>commons-compress</artifactId>
|
||||
<version>${commons-compress.version}</version>
|
||||
</dependency>
|
||||
<!-- Testcontainers 1.20.4 also pulls commons-lang3 3.16.0 (test scope): CVE-2025-48924
|
||||
(uncontrolled recursion in ClassUtils), fixed in 3.18.0. Pin the patched line.
|
||||
Test-scope only (never shipped in the jar), bumped per the CVE policy. -->
|
||||
<dependency>
|
||||
<groupId>org.apache.commons</groupId>
|
||||
<artifactId>commons-lang3</artifactId>
|
||||
<version>${commons-lang3.version}</version>
|
||||
</dependency>
|
||||
</dependencies>
|
||||
</dependencyManagement>
|
||||
|
||||
@@ -93,6 +113,16 @@
|
||||
<version>${mcp.version}</version>
|
||||
</dependency>
|
||||
|
||||
<!-- Broker client (CB-307 Stage 2): AMQP 0-9-1. Default deploy targets LavinMQ; this same
|
||||
client speaks to RabbitMQ unchanged (URI-only swap), so integration tests run against a
|
||||
stock RabbitMQ container. Only wired when a broker: block is present in config; absent →
|
||||
the in-memory ReplyInbox. -->
|
||||
<dependency>
|
||||
<groupId>com.rabbitmq</groupId>
|
||||
<artifactId>amqp-client</artifactId>
|
||||
<version>${amqp.version}</version>
|
||||
</dependency>
|
||||
|
||||
<!-- Logging -->
|
||||
<dependency>
|
||||
<groupId>org.slf4j</groupId>
|
||||
@@ -112,6 +142,22 @@
|
||||
<version>${junit.version}</version>
|
||||
<scope>test</scope>
|
||||
</dependency>
|
||||
|
||||
<!-- Testcontainers RabbitMQ: spins a real broker for the @Tag("contract") AMQP integration
|
||||
test only. Excluded from the default build (contract group), so `mvn clean install`
|
||||
stays hermetic and green without Docker; run under -Pcontract with Docker present. -->
|
||||
<dependency>
|
||||
<groupId>org.testcontainers</groupId>
|
||||
<artifactId>rabbitmq</artifactId>
|
||||
<version>${testcontainers.version}</version>
|
||||
<scope>test</scope>
|
||||
</dependency>
|
||||
<dependency>
|
||||
<groupId>org.testcontainers</groupId>
|
||||
<artifactId>junit-jupiter</artifactId>
|
||||
<version>${testcontainers.version}</version>
|
||||
<scope>test</scope>
|
||||
</dependency>
|
||||
</dependencies>
|
||||
|
||||
<build>
|
||||
@@ -123,6 +169,30 @@
|
||||
<version>3.14.0</version>
|
||||
</plugin>
|
||||
|
||||
<!--
|
||||
Coverage (CB-509). Build-time tooling only — never a compile/runtime dependency, so
|
||||
it adds nothing to the shipped jar. Report lands at target/site/jacoco/index.html and
|
||||
target/site/jacoco/jacoco.csv. No `check` rule / threshold is wired: a coverage gate
|
||||
rewards writing tests that execute lines, which is the failure mode this project is
|
||||
trying to avoid, not encourage.
|
||||
-->
|
||||
<plugin>
|
||||
<groupId>org.jacoco</groupId>
|
||||
<artifactId>jacoco-maven-plugin</artifactId>
|
||||
<version>0.8.13</version>
|
||||
<executions>
|
||||
<execution>
|
||||
<id>prepare-agent</id>
|
||||
<goals><goal>prepare-agent</goal></goals>
|
||||
</execution>
|
||||
<execution>
|
||||
<id>report</id>
|
||||
<phase>test</phase>
|
||||
<goals><goal>report</goal></goals>
|
||||
</execution>
|
||||
</executions>
|
||||
</plugin>
|
||||
|
||||
<!-- Unit tests run by default; contract tests (live herdr) are tag-excluded. -->
|
||||
<plugin>
|
||||
<groupId>org.apache.maven.plugins</groupId>
|
||||
|
||||
@@ -3,6 +3,8 @@ package dev.ltms.bridged;
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.HerdrClient;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.herdr.PaneLocator;
|
||||
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
@@ -11,23 +13,39 @@ import dev.ltms.bridged.inject.Injector;
|
||||
import dev.ltms.bridged.inject.StatusPoller;
|
||||
import dev.ltms.bridged.inject.TurnListener;
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.auth.CallerResolver;
|
||||
import dev.ltms.bridged.mcp.BridgeMcp;
|
||||
import dev.ltms.bridged.mcp.ConnectionIdentity;
|
||||
import dev.ltms.bridged.metrics.BridgedMetrics;
|
||||
import dev.ltms.bridged.metrics.Metrics;
|
||||
import dev.ltms.bridged.mcp.PrimaryRegistry;
|
||||
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
|
||||
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
|
||||
import dev.ltms.bridged.msg.AmqpReplyInbox;
|
||||
import dev.ltms.bridged.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.msg.ReplyInbox;
|
||||
import dev.ltms.bridged.msg.ReplyPushLoop;
|
||||
import dev.ltms.bridged.rest.BridgedApp;
|
||||
import dev.ltms.bridged.session.GitWorktrees;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import dev.ltms.bridged.session.SessionReaper;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.worker.CompositePeerLauncher;
|
||||
import dev.ltms.bridged.worker.HerdrPeerLauncher;
|
||||
import dev.ltms.bridged.worker.OpenCodeLauncher;
|
||||
import io.javalin.Javalin;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
|
||||
/**
|
||||
* {@code bridged} entry point. Wires the real herdr socket client to the REST app and
|
||||
@@ -41,6 +59,10 @@ public final class Bridged {
|
||||
/** How often the injector samples a busy worker's status while it has queued work. */
|
||||
private static final long INJECT_POLL_MILLIS = 250;
|
||||
|
||||
/** CB-504: how long to wait at startup for herdr's socket before serving degraded. */
|
||||
private static final long HERDR_WAIT_SECONDS = 30;
|
||||
private static final long HERDR_WAIT_POLL_MILLIS = 500;
|
||||
|
||||
static void main(String[] args) {
|
||||
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
|
||||
BridgedConfig cfg = BridgedConfig.load(configPath);
|
||||
@@ -49,6 +71,12 @@ public final class Bridged {
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
guard.assertPrimaryClean(System.getenv());
|
||||
|
||||
// CB-501: refuse to start if the bind is wider than the auth mode can defend. Under
|
||||
// loopback-trust, "not a known worker" means "the primary" — sound only because the OS
|
||||
// refuses remote connections to a loopback socket. This throws rather than warns so the
|
||||
// dangerous configuration cannot be reached by ignoring a log line.
|
||||
cfg.validateAuthExposure();
|
||||
|
||||
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
|
||||
? Path.of(cfg.herdrSocket())
|
||||
: UnixSocketHerdrClient.defaultSocketPath();
|
||||
@@ -57,11 +85,47 @@ public final class Bridged {
|
||||
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
WorkspaceControl spaces = new WorkspaceControl(herdr);
|
||||
PeerLauncher workers = new ClaudeCodeLauncher(agents, spaces, guard,
|
||||
cfg.workerProfiles(), cfg.defaultProfile(), System::getenv);
|
||||
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died with
|
||||
// the previous process — reap those leaked orphans now, before we start serving.
|
||||
workers.reapOrphanWorkers();
|
||||
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
|
||||
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
|
||||
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
|
||||
Map<String, BridgedConfig.Worker> claudeProfiles = new LinkedHashMap<>();
|
||||
Map<String, BridgedConfig.Worker> opencodeProfiles = new LinkedHashMap<>();
|
||||
cfg.workerProfiles().forEach((name, w) -> {
|
||||
if (w.isOpenCode()) {
|
||||
opencodeProfiles.put(name, w);
|
||||
} else {
|
||||
claudeProfiles.put(name, w);
|
||||
}
|
||||
});
|
||||
List<HerdrPeerLauncher> adapters = new ArrayList<>();
|
||||
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
|
||||
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
|
||||
// unless opencode is the only kind configured.
|
||||
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
|
||||
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
|
||||
claudeProfiles, cfg.defaultProfile(), System::getenv,
|
||||
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
|
||||
}
|
||||
if (!opencodeProfiles.isEmpty()) {
|
||||
adapters.add(new OpenCodeLauncher(agents, spaces,
|
||||
opencodeProfiles, cfg.defaultProfile(), System::getenv,
|
||||
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
|
||||
}
|
||||
PeerLauncher workers = new CompositePeerLauncher(adapters, cfg.defaultProfile());
|
||||
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
|
||||
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
|
||||
// the first thing that actually talks to herdr, so without this wait a boot-order race
|
||||
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
|
||||
// with /healthz reporting "degraded" is strictly more useful than exiting.
|
||||
if (awaitHerdr(herdr)) {
|
||||
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
|
||||
// with the previous process — reap those leaked orphans now, before we start serving.
|
||||
workers.reapOrphanWorkers();
|
||||
} else {
|
||||
log.warn("herdr did not answer within {}s — starting anyway; /healthz will report "
|
||||
+ "degraded until it comes up. Orphaned worker panes (if any) were NOT reaped.",
|
||||
HERDR_WAIT_SECONDS);
|
||||
}
|
||||
|
||||
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
|
||||
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
|
||||
@@ -115,14 +179,66 @@ public final class Bridged {
|
||||
StatusPoller poller = new StatusPoller(agents, injector, INJECT_POLL_MILLIS);
|
||||
poller.start();
|
||||
|
||||
MessageService messages = new MessageService(agents, injector, rendezvous);
|
||||
// CB-307: reply inbox. A broker: block (with a uri) selects the AMQP-backed durable adapter;
|
||||
// absent, bridged stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
|
||||
// connection, so keep the reference to close it in the ordered shutdown hook.
|
||||
final ReplyInbox replyInbox;
|
||||
if (cfg.broker() != null && cfg.broker().isConfigured()) {
|
||||
replyInbox = AmqpReplyInbox.open(cfg.broker().uri());
|
||||
log.info("reply inbox: AMQP broker (durable) at {}", cfg.broker().uri());
|
||||
} else {
|
||||
replyInbox = new InMemoryReplyInbox();
|
||||
log.info("reply inbox: in-memory (soft-state)");
|
||||
}
|
||||
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
|
||||
PrimaryRegistry primaryRegistry = new PrimaryRegistry(
|
||||
cfg.primary() != null ? cfg.primary().terminal() : null);
|
||||
|
||||
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
|
||||
// open bridge_send. Uses its own lightweight scheduled executor, separate from the injector.
|
||||
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
|
||||
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
|
||||
var pushScheduler = Executors.newSingleThreadScheduledExecutor(r ->
|
||||
Thread.ofVirtual().name("bridge-push-").unstarted(r));
|
||||
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
|
||||
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
|
||||
// CB-512: the push loop takes it too, so nudge outcomes (delivered|exhausted) are counted.
|
||||
Metrics metrics = BridgedMetrics.create(sessions, replyInbox);
|
||||
var pushLoop = new ReplyPushLoop(primaryRegistry, agents, replyInbox,
|
||||
pushScheduler, maxReminders, backoffMs, metrics);
|
||||
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
|
||||
pushLoop, metrics);
|
||||
|
||||
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
|
||||
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
|
||||
// reached /metrics — the delegation was unresolvable and nothing said so.
|
||||
sessions.onRelease(terminal ->
|
||||
messages.abandon(terminal, "the worker session was released before it replied"));
|
||||
|
||||
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
|
||||
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
|
||||
ConnectionIdentity identity = new ConnectionIdentity(
|
||||
new PaneLocator(herdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
|
||||
// Cast to ClaudeCodeLauncher: BridgeMcp is not yet migrated to PeerLauncher (Stage A scope).
|
||||
BridgeMcp mcp = new BridgeMcp(messages, rendezvous, (ClaudeCodeLauncher) workers, sessions, identity, presence);
|
||||
|
||||
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
|
||||
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
|
||||
final CallerResolver callers;
|
||||
if (cfg.auth().tokenMode()) {
|
||||
String token = System.getenv(cfg.auth().tokenEnv());
|
||||
if (token == null || token.isBlank()) {
|
||||
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
|
||||
+ " is unset or empty — export it before starting bridged");
|
||||
}
|
||||
callers = new CallerResolver(identity, true, token);
|
||||
log.info("auth: token mode (bearer required for non-worker callers, env {})",
|
||||
cfg.auth().tokenEnv());
|
||||
} else {
|
||||
callers = new CallerResolver(identity);
|
||||
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
|
||||
}
|
||||
|
||||
BridgeMcp mcp = new BridgeMcp(messages, workers, sessions, identity, presence,
|
||||
primaryRegistry, callers, metrics);
|
||||
|
||||
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
|
||||
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
|
||||
@@ -131,17 +247,60 @@ public final class Bridged {
|
||||
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
|
||||
poller.stop();
|
||||
messages.close();
|
||||
pushLoop.close();
|
||||
mcp.close();
|
||||
if (reaper != null) reaper.stop();
|
||||
// Release the broker connection last among message resources (no-op for the in-memory inbox).
|
||||
if (replyInbox instanceof AutoCloseable closeable) {
|
||||
try {
|
||||
closeable.close();
|
||||
} catch (Exception e) {
|
||||
log.debug("reply inbox close: {}", e.toString());
|
||||
}
|
||||
}
|
||||
herdr.close();
|
||||
}));
|
||||
|
||||
Javalin app = new BridgedApp(herdr, (ClaudeCodeLauncher) workers, sessions, messages, rendezvous, presence, mcp.servlet()).build();
|
||||
Javalin app = new BridgedApp(herdr, workers, sessions, messages, presence, mcp.servlet(),
|
||||
callers, metrics).build();
|
||||
app.start(cfg.bind().host(), cfg.bind().port());
|
||||
log.info("bridged listening on {}:{}, herdr socket {}",
|
||||
cfg.bind().host(), cfg.bind().port(), socket);
|
||||
}
|
||||
|
||||
/**
|
||||
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
|
||||
*
|
||||
* @return true if herdr answered, false if it never did
|
||||
*/
|
||||
private static boolean awaitHerdr(HerdrClient herdr) {
|
||||
long deadline = System.nanoTime() + HERDR_WAIT_SECONDS * 1_000_000_000L;
|
||||
boolean waited = false;
|
||||
while (true) {
|
||||
try {
|
||||
herdr.call("ping");
|
||||
if (waited) {
|
||||
log.info("herdr is up");
|
||||
}
|
||||
return true;
|
||||
} catch (HerdrException e) {
|
||||
if (System.nanoTime() >= deadline) {
|
||||
return false;
|
||||
}
|
||||
if (!waited) {
|
||||
log.info("waiting up to {}s for the herdr socket…", HERDR_WAIT_SECONDS);
|
||||
waited = true;
|
||||
}
|
||||
try {
|
||||
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
|
||||
} catch (InterruptedException ie) {
|
||||
Thread.currentThread().interrupt();
|
||||
return false;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private Bridged() {
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,72 @@
|
||||
package dev.ltms.bridged.auth;
|
||||
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.time.Instant;
|
||||
import java.time.ZoneOffset;
|
||||
import java.time.format.DateTimeFormatter;
|
||||
|
||||
/**
|
||||
* Append-only record of privileged actions (CB-505).
|
||||
*
|
||||
* <p>Writes JSON lines to a dedicated {@code audit} logger — its own appender, separate from the
|
||||
* chatty app log — so the trail stays greppable and can later be shipped without dragging debug
|
||||
* noise along.
|
||||
*
|
||||
* <p><strong>Message content is never recorded.</strong> This bridge carries the user's source
|
||||
* code, diffs, and prompts; an audit trail that quietly accumulated them would be a transcript
|
||||
* archive wearing a security control's clothing. Records carry <em>who / what / against what /
|
||||
* outcome</em> and correlation ids only.
|
||||
*/
|
||||
public final class AuditLog {
|
||||
|
||||
private static final Logger AUDIT = LoggerFactory.getLogger("audit");
|
||||
private static final DateTimeFormatter TS =
|
||||
DateTimeFormatter.ofPattern("yyyy-MM-dd'T'HH:mm:ss.SSSXXX").withZone(ZoneOffset.UTC);
|
||||
|
||||
private AuditLog() {
|
||||
}
|
||||
|
||||
/** Record an allowed action. */
|
||||
public static void allowed(Principal caller, Authz.Action action, String target) {
|
||||
write(caller, action, target, "allowed", null);
|
||||
}
|
||||
|
||||
/** Record a refused action and why. */
|
||||
public static void denied(Principal caller, Authz.Action action, String target, String reason) {
|
||||
write(caller, action, target, "denied", reason);
|
||||
}
|
||||
|
||||
/** Record an action that was authorized but then failed downstream (guard, timeout, herdr). */
|
||||
public static void failed(Principal caller, Authz.Action action, String target, String reason) {
|
||||
write(caller, action, target, "failed", reason);
|
||||
}
|
||||
|
||||
private static void write(Principal caller, Authz.Action action, String target,
|
||||
String outcome, String reason) {
|
||||
Principal c = caller != null ? caller : Principal.anonymous();
|
||||
StringBuilder sb = new StringBuilder(200);
|
||||
// The timestamp is built here rather than by the appender pattern: a pattern that wrapped
|
||||
// literal braces around the message collides with logback's own variable substitution.
|
||||
sb.append("{\"ts\":\"").append(TS.format(Instant.now())).append('"')
|
||||
.append(",\"role\":\"").append(c.role()).append('"')
|
||||
.append(",\"actor\":\"").append(esc(c.describe())).append('"')
|
||||
.append(",\"pid\":").append(c.pid())
|
||||
.append(",\"action\":\"").append(action).append('"')
|
||||
.append(",\"target\":").append(target == null ? "null" : '"' + esc(target) + '"')
|
||||
.append(",\"outcome\":\"").append(outcome).append('"');
|
||||
if (reason != null) {
|
||||
sb.append(",\"reason\":\"").append(esc(reason)).append('"');
|
||||
}
|
||||
sb.append('}');
|
||||
// The appender supplies the timestamp, so it cannot disagree with the app log's clock.
|
||||
AUDIT.info(sb.toString());
|
||||
}
|
||||
|
||||
/** Minimal JSON string escaping — these values are ids and short reasons, never free text. */
|
||||
private static String esc(String s) {
|
||||
return s.replace("\\", "\\\\").replace("\"", "\\\"")
|
||||
.replace("\n", "\\n").replace("\r", "\\r").replace("\t", "\\t");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,72 @@
|
||||
package dev.ltms.bridged.auth;
|
||||
|
||||
/**
|
||||
* The authorization table (CB-505), stated once and enforced on both entry paths.
|
||||
*
|
||||
* <p>Most of these rules are already true de facto — {@code BridgeMcp} derives a worker's identity
|
||||
* from the connection rather than reading it from an argument, so a worker has never been able to
|
||||
* reply <em>as</em> another worker over MCP. What was missing is that the REST surface trusted the
|
||||
* session id in the URL path, and neither surface checked role at all. This class makes the
|
||||
* invariant explicit and testable rather than emergent.
|
||||
*/
|
||||
public final class Authz {
|
||||
|
||||
private Authz() {
|
||||
}
|
||||
|
||||
/** A privileged operation, named for the audit trail. */
|
||||
public enum Action {
|
||||
/** Spawn a worker peer. */
|
||||
SPAWN,
|
||||
/** Tear a worker peer down. */
|
||||
STOP,
|
||||
/** Deliver a turn to a session (or answer a worker's question). */
|
||||
SEND,
|
||||
/** A worker's terminal reply for its own turn. */
|
||||
REPLY,
|
||||
/** A worker's mid-turn question to the primary. */
|
||||
ASK,
|
||||
/** Collect held replies from a session's inbox. */
|
||||
DRAIN,
|
||||
/** Read-only observation: status, roster, profiles, task polling. */
|
||||
READ,
|
||||
/** Scrape the metrics endpoint. */
|
||||
METRICS
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether {@code caller} may perform {@code action} against {@code targetSession}.
|
||||
*
|
||||
* @param targetSession the session id in the request path; only consulted for the worker-scoped
|
||||
* actions ({@code REPLY}, {@code ASK}), ignored otherwise, may be
|
||||
* {@code null}
|
||||
*/
|
||||
public static boolean permits(Principal caller, Action action, String targetSession) {
|
||||
if (caller == null || caller.isAnonymous()) {
|
||||
return false; // authenticated as nothing ⇒ authorized for nothing
|
||||
}
|
||||
return switch (action) {
|
||||
// Orchestration is the primary's alone. A worker driving spawn/stop/send would be a
|
||||
// worker escalating into the orchestrator role.
|
||||
case SPAWN, STOP, SEND, DRAIN -> caller.isPrimary();
|
||||
|
||||
// The load-bearing rule: a worker acts only as itself. The primary is deliberately
|
||||
// excluded — a reply/ask is a worker's own turn output, and letting the primary forge
|
||||
// one would corrupt the rendezvous correlation it is itself waiting on.
|
||||
case REPLY, ASK -> caller.ownsSession(targetSession);
|
||||
|
||||
// Observation is open to both authenticated roles: a worker legitimately polls its own
|
||||
// status, and the roster carries no secrets.
|
||||
case READ, METRICS -> caller.isPrimary() || caller.isWorker();
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Why a request was refused, for the error body. Distinguishes "you are nobody" from "you are
|
||||
* somebody, but not the right somebody" — the first is a credential problem (401), the second
|
||||
* an authorization one (403).
|
||||
*/
|
||||
public static boolean isUnauthenticated(Principal caller) {
|
||||
return caller == null || caller.isAnonymous();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,119 @@
|
||||
package dev.ltms.bridged.auth;
|
||||
|
||||
import dev.ltms.bridged.mcp.ConnectionIdentity;
|
||||
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.security.MessageDigest;
|
||||
|
||||
/**
|
||||
* Resolves every caller to a {@link Principal}, for both entry paths into the core (CB-501).
|
||||
*
|
||||
* <p>There are two of them and they are not layered the way the docs suggest: {@code BridgeMcp}
|
||||
* calls the service layer directly and is mounted as a raw servlet (so it never passes through a
|
||||
* Javalin filter), while the REST routes historically resolved no identity at all. Both now
|
||||
* delegate here, so the authorization rules are stated once instead of drifting apart.
|
||||
*
|
||||
* <p><strong>Resolution order</strong> — connection identity first, token second, nothing third:
|
||||
* <ol>
|
||||
* <li>A loopback peer PID that maps to a herdr worker pane ⇒ {@link Role#WORKER}. This is
|
||||
* unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured
|
||||
* regardless of auth mode, so enabling auth never breaks the fleet.</li>
|
||||
* <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li>
|
||||
* <li>Otherwise, under {@code loopback-trust}, a loopback caller ⇒ {@link Role#PRIMARY}
|
||||
* (the historical behaviour, now an explicit configured choice).</li>
|
||||
* <li>Otherwise {@link Role#ANONYMOUS}.</li>
|
||||
* </ol>
|
||||
*/
|
||||
public final class CallerResolver {
|
||||
|
||||
private final ConnectionIdentity identity;
|
||||
private final boolean tokenMode;
|
||||
private final byte[] expectedToken; // null unless tokenMode
|
||||
|
||||
/** Loopback-trust resolver: no token required, historical behaviour. */
|
||||
public CallerResolver(ConnectionIdentity identity) {
|
||||
this(identity, false, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* @param identity connection-based worker identification
|
||||
* @param tokenMode when true, a non-worker caller must present a valid bearer token
|
||||
* @param token the expected bearer token; required (non-blank) when {@code tokenMode}
|
||||
*/
|
||||
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
|
||||
if (tokenMode && (token == null || token.isBlank())) {
|
||||
throw new IllegalArgumentException(
|
||||
"auth.mode=token requires a non-empty token; check that the env var named by "
|
||||
+ "auth.tokenEnv is exported to the daemon's environment");
|
||||
}
|
||||
this.identity = identity;
|
||||
this.tokenMode = tokenMode;
|
||||
this.expectedToken = tokenMode ? token.getBytes(StandardCharsets.UTF_8) : null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve the caller of a request.
|
||||
*
|
||||
* @param remoteAddr the connection's remote address
|
||||
* @param remotePort the connection's remote port (used for the peer-PID lookup)
|
||||
* @param authorizationHeader the raw {@code Authorization} header, or {@code null}
|
||||
*/
|
||||
public Principal resolve(String remoteAddr, int remotePort, String authorizationHeader) {
|
||||
ConnectionIdentity.Caller c = identity.resolve(remoteAddr, remotePort);
|
||||
if (c.terminal() != null) {
|
||||
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
|
||||
}
|
||||
|
||||
if (tokenMode) {
|
||||
return presentedTokenMatches(authorizationHeader)
|
||||
? Principal.primary(c.pid())
|
||||
: Principal.anonymous();
|
||||
}
|
||||
|
||||
// loopback-trust: same-host callers that are not workers are the primary. A non-loopback
|
||||
// caller is anonymous even here — and startup refuses that combination anyway
|
||||
// (BridgedConfig.validateAuthExposure), so this is defence in depth, not the control.
|
||||
return isLoopback(remoteAddr) ? Principal.primary(c.pid()) : Principal.anonymous();
|
||||
}
|
||||
|
||||
/** The working directory of the calling process (CB-112 spawn cwd inheritance), or {@code null}. */
|
||||
public String cwdForPid(long pid) {
|
||||
return identity.cwdForPid(pid);
|
||||
}
|
||||
|
||||
/** True when auth requires a bearer token of non-worker callers. */
|
||||
public boolean tokenMode() {
|
||||
return tokenMode;
|
||||
}
|
||||
|
||||
private boolean presentedTokenMatches(String authorizationHeader) {
|
||||
String presented = bearerValue(authorizationHeader);
|
||||
if (presented == null) {
|
||||
return false;
|
||||
}
|
||||
// Constant-time: MessageDigest.isEqual does not short-circuit on the first differing byte,
|
||||
// so a token cannot be recovered a byte at a time by timing the response.
|
||||
return MessageDigest.isEqual(presented.getBytes(StandardCharsets.UTF_8), expectedToken);
|
||||
}
|
||||
|
||||
/** Extract the credential from {@code Authorization: Bearer <token>}, or {@code null}. */
|
||||
private static String bearerValue(String header) {
|
||||
if (header == null) {
|
||||
return null;
|
||||
}
|
||||
String h = header.trim();
|
||||
if (h.length() < 7 || !h.regionMatches(true, 0, "Bearer ", 0, 7)) {
|
||||
return null;
|
||||
}
|
||||
String token = h.substring(7).trim();
|
||||
return token.isEmpty() ? null : token;
|
||||
}
|
||||
|
||||
private static boolean isLoopback(String remoteAddr) {
|
||||
if (remoteAddr == null) {
|
||||
return false;
|
||||
}
|
||||
return remoteAddr.equals("127.0.0.1") || remoteAddr.equals("::1")
|
||||
|| remoteAddr.equals("0:0:0:0:0:0:0:1") || remoteAddr.startsWith("127.");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,57 @@
|
||||
package dev.ltms.bridged.auth;
|
||||
|
||||
/**
|
||||
* A resolved caller: its {@link Role}, and — for a worker — the herdr {@code terminal_id} that
|
||||
* identifies which worker it is (CB-501).
|
||||
*
|
||||
* @param role what this caller is authorized to act as
|
||||
* @param terminal the worker's herdr terminal id; {@code null} for {@code PRIMARY}/{@code ANONYMOUS}
|
||||
* @param pid the connecting process id, or {@code -1} when not resolvable (audit context)
|
||||
*/
|
||||
public record Principal(Role role, String terminal, long pid) {
|
||||
|
||||
/** A caller authenticated as nothing — the default when no check establishes anything else. */
|
||||
public static Principal anonymous() {
|
||||
return new Principal(Role.ANONYMOUS, null, -1);
|
||||
}
|
||||
|
||||
/** The orchestrating session. */
|
||||
public static Principal primary(long pid) {
|
||||
return new Principal(Role.PRIMARY, null, pid);
|
||||
}
|
||||
|
||||
/** A worker peer, identified by its herdr pane. */
|
||||
public static Principal worker(String terminal, long pid) {
|
||||
return new Principal(Role.WORKER, terminal, pid);
|
||||
}
|
||||
|
||||
public boolean isPrimary() {
|
||||
return role == Role.PRIMARY;
|
||||
}
|
||||
|
||||
public boolean isWorker() {
|
||||
return role == Role.WORKER;
|
||||
}
|
||||
|
||||
public boolean isAnonymous() {
|
||||
return role == Role.ANONYMOUS;
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether this caller may act <em>as</em> {@code sessionId} — the "own session only" rule that
|
||||
* keeps one worker from replying or asking on another's behalf. Only a worker can own a
|
||||
* session, and only its own.
|
||||
*/
|
||||
public boolean ownsSession(String sessionId) {
|
||||
return isWorker() && terminal != null && terminal.equals(sessionId);
|
||||
}
|
||||
|
||||
/** Short, non-sensitive description for audit lines and error details. */
|
||||
public String describe() {
|
||||
return switch (role) {
|
||||
case WORKER -> "worker:" + terminal;
|
||||
case PRIMARY -> "primary";
|
||||
case ANONYMOUS -> "anonymous";
|
||||
};
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,29 @@
|
||||
package dev.ltms.bridged.auth;
|
||||
|
||||
/**
|
||||
* What a caller is allowed to be on the bus (CB-501).
|
||||
*
|
||||
* <p>The ordering matters conceptually: {@link #PRIMARY} is the <em>most</em> privileged role
|
||||
* (it spawns, stops, sends to any session, and drains any inbox), not the least. Before CB-501
|
||||
* the daemon reached {@code PRIMARY} by <em>failing</em> every other check — any caller that did
|
||||
* not resolve to a known worker pane was treated as the primary. That is inverted here:
|
||||
* {@link #ANONYMOUS} is the fallback, and {@code PRIMARY} must be established.
|
||||
*/
|
||||
public enum Role {
|
||||
|
||||
/**
|
||||
* The orchestrating session. Established either by being a loopback caller that is not a
|
||||
* worker pane (under {@code loopback-trust}) or by presenting a valid bearer token (under
|
||||
* {@code token} mode).
|
||||
*/
|
||||
PRIMARY,
|
||||
|
||||
/**
|
||||
* A worker peer, identified by its herdr pane. Unforgeable: derived from the connection's
|
||||
* loopback peer PID via herdr's PID→pane map, never from a request argument.
|
||||
*/
|
||||
WORKER,
|
||||
|
||||
/** Authenticated as nothing. Authorized for nothing but {@code /healthz}. */
|
||||
ANONYMOUS
|
||||
}
|
||||
@@ -27,7 +27,17 @@ import java.util.Set;
|
||||
* @param guard subscription-boundary allowlist
|
||||
* @param worktreeRoot nullable root directory for provisioned worktrees; defaults to a sibling
|
||||
* of the repo root
|
||||
* @param lifecycle session lifecycle limits ({@code null} = all disabled)
|
||||
* @param lifecycle session lifecycle limits ({@code null} = all disabled)
|
||||
* @param spawnReadyTimeoutMs max ms to wait for a spawned worker to reach an injectable state
|
||||
* ({@code null} / 0 disables the poll gate — legacy non-blocking behaviour)
|
||||
* @param spawnReadyPollMs poll interval while waiting for the worker to become injectable
|
||||
* @param broker external AMQP broker for durable reply delivery ({@code null} → in-memory,
|
||||
* soft-state {@code ReplyInbox}; present → the AMQP-backed adapter, CB-307 Stage 2)
|
||||
* @param primary optional pinned primary terminal config ({@code null} → derived from connection);
|
||||
* a non-blank {@code terminal} seeds {@code PrimaryRegistry} and prevents
|
||||
* connection-derived overrides, CB-307
|
||||
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
|
||||
* historical behaviour), CB-501
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record BridgedConfig(
|
||||
@@ -38,7 +48,12 @@ public record BridgedConfig(
|
||||
String defaultWorker,
|
||||
Guard guard,
|
||||
String worktreeRoot,
|
||||
Lifecycle lifecycle) {
|
||||
Lifecycle lifecycle,
|
||||
Integer spawnReadyTimeoutMs,
|
||||
Integer spawnReadyPollMs,
|
||||
Broker broker,
|
||||
Primary primary,
|
||||
Auth auth) {
|
||||
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record Bind(String host, int port) {
|
||||
@@ -78,6 +93,12 @@ public record BridgedConfig(
|
||||
* (minimal-grant default — push over SSH stays free, PR-create is opt-in)
|
||||
* @param gitHostEnv name of the host env var holding the forge host (default {@code GITEA_HOST});
|
||||
* injected as {@code GITEA_HOST} <em>only</em> when {@code gitTokenEnv} is set
|
||||
* @param kind which peer launcher spawns this profile: {@code "claude-code"} (default —
|
||||
* the {@link dev.ltms.bridged.worker.ClaudeCodeLauncher}) or {@code "opencode"}.
|
||||
* The {@code CompositePeerLauncher} routes {@code spawn}/reap by this value, so
|
||||
* each adapter drives only its own kind. Normalised to lower-case; blank ⇒ the
|
||||
* default. It selects the adapter, not the transport — placement, tabs, cwd, and
|
||||
* the readiness gate are kind-independent and stay in the shared base.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record Worker(String profile, String baseUrl, String model,
|
||||
@@ -85,9 +106,24 @@ public record BridgedConfig(
|
||||
String placement, String workspace, String tabLabel, String mcpUrl,
|
||||
String cwd,
|
||||
List<String> parityOverlay,
|
||||
String gitTokenEnv, String gitHostEnv) {
|
||||
String gitTokenEnv, String gitHostEnv,
|
||||
String kind,
|
||||
Map<String, String> env) {
|
||||
|
||||
/** Peer kind spawned by {@link dev.ltms.bridged.worker.ClaudeCodeLauncher} (the default). */
|
||||
public static final String KIND_CLAUDE_CODE = "claude-code";
|
||||
/** Peer kind spawned by the opencode adapter (CB-402). */
|
||||
public static final String KIND_OPENCODE = "opencode";
|
||||
|
||||
public Worker {
|
||||
argv = (argv == null || argv.isEmpty()) ? List.of("claude") : List.copyOf(argv);
|
||||
// A claude-code worker defaults its launch command to `claude`; other kinds carry their own
|
||||
// argv (e.g. `opencode`) and must not inherit the Claude binary — so only default when unset
|
||||
// AND this is the claude-code kind.
|
||||
String k = (kind == null || kind.isBlank()) ? KIND_CLAUDE_CODE : kind.toLowerCase();
|
||||
argv = (argv == null || argv.isEmpty())
|
||||
? (KIND_CLAUDE_CODE.equals(k) ? List.of("claude") : List.of(k))
|
||||
: List.copyOf(argv);
|
||||
kind = k;
|
||||
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_WORKER_TOKEN" : tokenEnv;
|
||||
placement = (placement == null || placement.isBlank()) ? "tab" : placement.toLowerCase();
|
||||
workspace = (workspace == null || workspace.isBlank()) ? "bridged-workers" : workspace;
|
||||
@@ -98,6 +134,7 @@ public record BridgedConfig(
|
||||
// gitTokenEnv stays null when unset (opt-in). gitHostEnv defaults so operators enabling
|
||||
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
|
||||
gitHostEnv = (gitHostEnv == null || gitHostEnv.isBlank()) ? "GITEA_HOST" : gitHostEnv;
|
||||
env = (env == null) ? Map.of() : Map.copyOf(env);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -110,13 +147,48 @@ public record BridgedConfig(
|
||||
String placement, String workspace, String tabLabel, String mcpUrl,
|
||||
String cwd, List<String> parityOverlay) {
|
||||
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
|
||||
mcpUrl, cwd, parityOverlay, null, null);
|
||||
mcpUrl, cwd, parityOverlay, null, null, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Backward-compatible constructor with the CB-302 git-forge fields but no explicit peer
|
||||
* {@code kind} — defaults to {@link #KIND_CLAUDE_CODE}. Keeps pre-CB-402 call sites working.
|
||||
*/
|
||||
public Worker(String profile, String baseUrl, String model,
|
||||
String configDir, String tokenEnv, List<String> argv,
|
||||
String placement, String workspace, String tabLabel, String mcpUrl,
|
||||
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv) {
|
||||
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
|
||||
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, null, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Backward-compatible constructor without the CB-511 {@code env:} passthrough — the worker
|
||||
* gets the daemon's PATH and nothing else. Keeps pre-CB-511 call sites working.
|
||||
*/
|
||||
public Worker(String profile, String baseUrl, String model,
|
||||
String configDir, String tokenEnv, List<String> argv,
|
||||
String placement, String workspace, String tabLabel, String mcpUrl,
|
||||
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
|
||||
String kind) {
|
||||
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
|
||||
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, null);
|
||||
}
|
||||
|
||||
/** A copy with {@code profile} set — used to default a profile to its {@code workers} key. */
|
||||
public Worker withProfile(String p) {
|
||||
return new Worker(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
|
||||
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv);
|
||||
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env);
|
||||
}
|
||||
|
||||
/** True when this profile is served by the Claude Code adapter (the default kind). */
|
||||
public boolean isClaudeCode() {
|
||||
return KIND_CLAUDE_CODE.equals(kind);
|
||||
}
|
||||
|
||||
/** True when this profile is served by the opencode adapter (CB-402). */
|
||||
public boolean isOpenCode() {
|
||||
return KIND_OPENCODE.equals(kind);
|
||||
}
|
||||
|
||||
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
|
||||
@@ -161,6 +233,84 @@ public record BridgedConfig(
|
||||
public record Lifecycle(Integer idleTtlSeconds, Integer contextCap, Integer drainTimeoutSeconds) {
|
||||
}
|
||||
|
||||
/**
|
||||
* External AMQP broker for durable, cross-restart reply delivery (CB-307 Stage 2). Its mere
|
||||
* presence swaps the in-memory {@code ReplyInbox} for the AMQP-backed adapter; absent, bridged
|
||||
* stays soft-state. Production default is LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1
|
||||
* and is a URI-only swap.
|
||||
*
|
||||
* @param uri AMQP connection URI, e.g. {@code amqp://guest:guest@127.0.0.1:5672/}. Blank/{@code null}
|
||||
* ⇒ the broker block is treated as absent (in-memory adapter).
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record Broker(String uri) {
|
||||
|
||||
/** True when a usable broker URI is configured (an empty block does not enable AMQP). */
|
||||
public boolean isConfigured() {
|
||||
return uri != null && !uri.isBlank();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Optional pinned primary terminal config (CB-307). When present with a non-blank
|
||||
* {@code terminal}, the bridge uses this as the primary's herdr identity instead of
|
||||
* deriving it from the MCP connection. Useful when the primary runs off-host or in a
|
||||
* non-herdr terminal where connection-derived identity is unavailable.
|
||||
*
|
||||
* @param terminal the primary's herdr {@code terminal_id} ({@code null}/blank → derive)
|
||||
* @param pushReminders max reminder nudges before giving up (default 5)
|
||||
* @param pushBackoffMs delay between reminders in ms (default 15000)
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record Primary(String terminal, Integer pushReminders, Integer pushBackoffMs) {
|
||||
/** @return configured reminder cap, or 5 */
|
||||
public int remindersOrDefault() {
|
||||
return pushReminders != null ? pushReminders : 5;
|
||||
}
|
||||
|
||||
/** @return configured backoff in ms, or 15000 */
|
||||
public long backoffMsOrDefault() {
|
||||
return pushBackoffMs != null ? pushBackoffMs.longValue() : 15_000L;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* API authentication (CB-501). Governs how a caller that is <em>not</em> an on-host worker
|
||||
* pane proves it is the primary.
|
||||
*
|
||||
* <p>Worker identity never depends on this block: a loopback peer PID that maps to a herdr
|
||||
* pane is unforgeable and is always honoured (see
|
||||
* {@link dev.ltms.bridged.mcp.ConnectionIdentity}). This only decides what happens for
|
||||
* <em>everyone else</em>.
|
||||
*
|
||||
* @param mode {@code "loopback-trust"} (default) — any loopback caller that is not a known
|
||||
* worker is the primary, no credential needed; this is the historical
|
||||
* behaviour, now chosen explicitly rather than implied. {@code "token"} — such
|
||||
* a caller must present {@code Authorization: Bearer <token>} or it is
|
||||
* {@code ANONYMOUS} and authorized for nothing.
|
||||
* @param tokenEnv name of the host env var holding the bearer token; the literal value is
|
||||
* never stored in config. Defaults to {@code BRIDGED_API_TOKEN}. Only read
|
||||
* when {@code mode} is {@code token}.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record Auth(String mode, String tokenEnv) {
|
||||
|
||||
/** Historical behaviour: loopback non-worker ⇒ primary, no credential. */
|
||||
public static final String MODE_LOOPBACK_TRUST = "loopback-trust";
|
||||
/** A non-worker caller must present a valid bearer token to be the primary. */
|
||||
public static final String MODE_TOKEN = "token";
|
||||
|
||||
public Auth {
|
||||
mode = (mode == null || mode.isBlank()) ? MODE_LOOPBACK_TRUST : mode.toLowerCase();
|
||||
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_API_TOKEN" : tokenEnv;
|
||||
}
|
||||
|
||||
/** True when a bearer token is required of every non-worker caller. */
|
||||
public boolean tokenMode() {
|
||||
return MODE_TOKEN.equals(mode);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Subscription boundary. Only these hosts may back a worker's
|
||||
* {@code ANTHROPIC_BASE_URL}; the primary must carry none.
|
||||
@@ -229,6 +379,45 @@ public record BridgedConfig(
|
||||
Bind b = bind != null ? bind : new Bind(null, 0);
|
||||
Guard g = guard != null ? guard : new Guard(List.of());
|
||||
Lifecycle l = lifecycle != null ? lifecycle : new Lifecycle(null, null, null);
|
||||
return new BridgedConfig(b, herdrSocket, worker, workers, defaultWorker, g, worktreeRoot, l);
|
||||
Integer timeout = (spawnReadyTimeoutMs != null) ? spawnReadyTimeoutMs : 20000;
|
||||
Integer pollMs = (spawnReadyPollMs != null) ? spawnReadyPollMs : 300;
|
||||
Auth a = auth != null ? auth : new Auth(null, null);
|
||||
// broker is left as-is: null (or an empty/blank uri) keeps the in-memory soft-state inbox.
|
||||
// primary is left as-is: null defaults to connection-derived identity.
|
||||
return new BridgedConfig(b, herdrSocket, worker, workers, defaultWorker, g, worktreeRoot, l, timeout, pollMs, broker, primary, a);
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a configuration whose network exposure outruns its authentication (CB-501).
|
||||
*
|
||||
* <p>{@code loopback-trust} means "any caller that is not a known worker pane is the primary" —
|
||||
* safe only because the OS refuses non-local connections to a loopback bind. Widen
|
||||
* {@code bind.host} without switching to {@code token} mode and that sentence becomes "any
|
||||
* client that can reach this port is the primary", which is the most privileged role on the
|
||||
* bus. Rather than document the hazard, make it unrepresentable: fail fast at startup.
|
||||
*
|
||||
* @throws IllegalStateException when a non-loopback bind is paired with {@code loopback-trust}
|
||||
*/
|
||||
public void validateAuthExposure() {
|
||||
String host = bind().host();
|
||||
if (isLoopbackBind(host) || auth().tokenMode()) {
|
||||
return;
|
||||
}
|
||||
throw new IllegalStateException(
|
||||
"refusing to start: bind.host=" + host + " is not loopback, but auth.mode="
|
||||
+ auth().mode() + ". A non-loopback bind treats every unauthenticated "
|
||||
+ "caller as the primary (spawn/stop/send/drain on any session). Set "
|
||||
+ "auth.mode: token (with auth.tokenEnv) before exposing this port, or "
|
||||
+ "bind to 127.0.0.1 and put a reverse proxy in front.");
|
||||
}
|
||||
|
||||
/** True for the loopback addresses and the unspecified-but-local forms we treat as same-host. */
|
||||
private static boolean isLoopbackBind(String host) {
|
||||
if (host == null || host.isBlank()) {
|
||||
return true; // Bind's own default is 127.0.0.1
|
||||
}
|
||||
String h = host.trim().toLowerCase();
|
||||
return h.equals("127.0.0.1") || h.equals("::1") || h.equals("localhost")
|
||||
|| h.startsWith("127.");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,15 +1,23 @@
|
||||
package dev.ltms.bridged.mcp;
|
||||
|
||||
import dev.ltms.bridged.auth.AuditLog;
|
||||
import dev.ltms.bridged.auth.Authz;
|
||||
import dev.ltms.bridged.auth.CallerResolver;
|
||||
import dev.ltms.bridged.auth.Principal;
|
||||
import dev.ltms.bridged.auth.Role;
|
||||
import dev.ltms.bridged.guard.GuardException;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.metrics.BridgedMetrics;
|
||||
import dev.ltms.bridged.metrics.Metrics;
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.session.WorkerSession;
|
||||
import dev.ltms.bridged.session.WorktreeRequest;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import io.modelcontextprotocol.common.McpTransportContext;
|
||||
import io.modelcontextprotocol.json.McpJsonMapper;
|
||||
import io.modelcontextprotocol.json.jackson3.JacksonMcpJsonMapperSupplier;
|
||||
@@ -34,8 +42,8 @@ import java.util.stream.Collectors;
|
||||
* / {@code bridge_status}; the worker calls {@code bridge_reply}.
|
||||
*
|
||||
* <p>Beyond delegation the primary also manages the fleet here (CB-108): {@code bridge_spawn} /
|
||||
* {@code bridge_list} / {@code bridge_stop} adapt {@link ClaudeCodeLauncher} so a worker's whole lifecycle
|
||||
* is driven through MCP, with the subscription boundary still enforced inside {@code ClaudeCodeLauncher}.
|
||||
* {@code bridge_list} / {@code bridge_stop} drive the {@link PeerLauncher} SPI so a worker's whole
|
||||
* lifecycle is managed through MCP, with each adapter's subscription boundary enforced inside it.
|
||||
*
|
||||
* <p>The tool <em>logic</em> lives in package-private static methods returning a
|
||||
* {@link McpSchema.CallToolResult}, so it is unit-testable without standing up the HTTP transport;
|
||||
@@ -55,12 +63,34 @@ public final class BridgeMcp {
|
||||
static final String CALLER_TERMINAL = "callerTerminal";
|
||||
/** Transport-context key under which the extractor stashes the caller's PID (for cwd inherit). */
|
||||
static final String CALLER_PID = "callerPid";
|
||||
/** Transport-context key under which the extractor stashes the resolved {@link Role} (CB-501). */
|
||||
static final String CALLER_ROLE = "callerRole";
|
||||
|
||||
private final HttpServletStreamableServerTransportProvider transport;
|
||||
private final McpSyncServer server;
|
||||
private final CallerResolver authz; // CB-501: null → authorization not enforced (legacy)
|
||||
private final Metrics metrics; // CB-502: null → auth failures not counted
|
||||
|
||||
public BridgeMcp(MessageService messages, Rendezvous rendezvous, ClaudeCodeLauncher workers,
|
||||
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence) {
|
||||
/**
|
||||
* Legacy constructor — no authorization. Retained so existing tests exercise tool behaviour
|
||||
* without an auth fixture.
|
||||
*/
|
||||
public BridgeMcp(MessageService messages, PeerLauncher workers,
|
||||
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
|
||||
PrimaryRegistry primaryRegistry) {
|
||||
this(messages, workers, sessions, identity, presence, primaryRegistry, null, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
|
||||
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
|
||||
* Jetty's context handler and never passes through Javalin's {@code before}
|
||||
* filter, so the REST guard does not cover it.
|
||||
* @param metrics registry for auth-failure counting; may be {@code null}
|
||||
*/
|
||||
public BridgeMcp(MessageService messages, PeerLauncher workers,
|
||||
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
|
||||
PrimaryRegistry primaryRegistry, CallerResolver callers, Metrics metrics) {
|
||||
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
|
||||
this.transport = HttpServletStreamableServerTransportProvider.builder()
|
||||
.jsonMapper(json)
|
||||
@@ -70,17 +100,28 @@ public final class BridgeMcp {
|
||||
// inherit the primary's cwd (CB-112). Any contact from a worker marks it available
|
||||
// (CB-113) — its MCP initialize is the reliable "the agent is up" signal.
|
||||
.contextExtractor(req -> {
|
||||
ConnectionIdentity.Caller c = identity.resolve(req.getRemoteAddr(), req.getRemotePort());
|
||||
presence.markPresent(c.terminal()); // no-op for the primary (null terminal)
|
||||
// One resolution per call, shared with the REST surface via CallerResolver so
|
||||
// the two paths cannot drift on who a caller is.
|
||||
Principal p = callers != null
|
||||
? callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
|
||||
req.getHeader("Authorization"))
|
||||
: legacyPrincipal(identity, req.getRemoteAddr(), req.getRemotePort());
|
||||
presence.markPresent(p.terminal()); // no-op for the primary (null terminal)
|
||||
return McpTransportContext.create(Map.of(
|
||||
CALLER_TERMINAL, orEmpty(c.terminal()),
|
||||
CALLER_PID, Long.toString(c.pid())));
|
||||
CALLER_TERMINAL, orEmpty(p.terminal()),
|
||||
CALLER_PID, Long.toString(p.pid()),
|
||||
CALLER_ROLE, p.role().name()));
|
||||
})
|
||||
.build();
|
||||
this.server = McpServer.sync(transport)
|
||||
.serverInfo("bridge", "0.1.0")
|
||||
.capabilities(McpSchema.ServerCapabilities.builder().tools(true).build())
|
||||
.toolCall(sendTool(), (_, req) -> {
|
||||
.toolCall(sendTool(), (exchange, req) -> {
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SEND,
|
||||
str(req.arguments(), "sessionId"));
|
||||
if (denied != null) return denied;
|
||||
String caller = callerTerminal(exchange);
|
||||
if (caller != null) primaryRegistry.record(caller);
|
||||
Map<String, Object> a = req.arguments();
|
||||
String turnId = str(a, "turnId");
|
||||
if (turnId != null && !turnId.isBlank()) {
|
||||
@@ -93,18 +134,46 @@ public final class BridgeMcp {
|
||||
? sendAsync(messages, str(a, "sessionId"), str(a, "content"))
|
||||
: send(messages, str(a, "sessionId"), str(a, "content"), timeoutMs(a));
|
||||
})
|
||||
// bridge_reply's identity is the CONNECTION, never an argument.
|
||||
.toolCall(replyTool(), (exchange, req) ->
|
||||
reply(rendezvous, callerTerminal(exchange), str(req.arguments(), "content")))
|
||||
// bridge_reply's identity is the CONNECTION, never an argument — so the authz check
|
||||
// is "is this caller a worker at all", and it can only ever reply as itself.
|
||||
.toolCall(replyTool(), (exchange, req) -> {
|
||||
String self = callerTerminal(exchange);
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.REPLY, self);
|
||||
if (denied != null) return denied;
|
||||
return reply(messages, self, str(req.arguments(), "content"));
|
||||
})
|
||||
// bridge_ask (CB-205): a worker's mid-turn question — identity from the CONNECTION.
|
||||
.toolCall(askTool(), (exchange, req) ->
|
||||
ask(messages, callerTerminal(exchange), str(req.arguments(), "question"), timeoutMs(req.arguments())))
|
||||
.toolCall(statusTool(), (_, req) ->
|
||||
status(messages, str(req.arguments(), "sessionId")))
|
||||
.toolCall(pollTool(), (_, req) ->
|
||||
poll(messages, str(req.arguments(), "ticket")))
|
||||
.toolCall(askTool(), (exchange, req) -> {
|
||||
String self = callerTerminal(exchange);
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.ASK, self);
|
||||
if (denied != null) return denied;
|
||||
return ask(messages, self, str(req.arguments(), "question"), timeoutMs(req.arguments()));
|
||||
})
|
||||
.toolCall(statusTool(), (exchange, req) -> {
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
||||
if (denied != null) return denied;
|
||||
return status(messages, str(req.arguments(), "sessionId"));
|
||||
})
|
||||
.toolCall(pollTool(), (exchange, req) -> {
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
||||
if (denied != null) return denied;
|
||||
Map<String, Object> a = req.arguments();
|
||||
return poll(messages, str(a, "ticket"), str(a, "target"));
|
||||
})
|
||||
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
|
||||
// Acking removes a reply from the inbox, so it is a drain, not a read.
|
||||
.toolCall(ackTool(), (exchange, req) -> {
|
||||
Map<String, Object> a = req.arguments();
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.DRAIN, str(a, "target"));
|
||||
if (denied != null) return denied;
|
||||
return ack(messages, str(a, "target"), str(a, "msgId"));
|
||||
})
|
||||
// Fleet management (CB-108): spawn/list/stop over ClaudeCodeLauncher.
|
||||
.toolCall(spawnTool(), (exchange, req) -> {
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SPAWN, null);
|
||||
if (denied != null) return denied;
|
||||
String caller = callerTerminal(exchange);
|
||||
if (caller != null) primaryRegistry.record(caller);
|
||||
Map<String, Object> a = req.arguments();
|
||||
// CB-112: worker inherits the primary's cwd unless the call pins one.
|
||||
// CB-301: carry the caller's identity as the session owner (null for the primary).
|
||||
@@ -113,10 +182,107 @@ public final class BridgeMcp {
|
||||
return spawn(sessions, str(a, "profile"), str(a, "cwd"), callerCwd,
|
||||
callerTerminal(exchange), worktreeRequest(a));
|
||||
})
|
||||
.toolCall(listTool(), (_, _) -> listWorkers(workers, sessions))
|
||||
.toolCall(stopTool(), (_, req) -> stop(sessions, str(req.arguments(), "paneId")))
|
||||
.toolCall(profilesTool(), (_, _) -> profiles(workers))
|
||||
.toolCall(listTool(), (exchange, _) -> {
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
||||
if (denied != null) return denied;
|
||||
return listWorkers(workers, sessions);
|
||||
})
|
||||
.toolCall(stopTool(), (exchange, req) -> {
|
||||
String paneId = str(req.arguments(), "paneId");
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.STOP, paneId);
|
||||
if (denied != null) return denied;
|
||||
return stop(sessions, paneId);
|
||||
})
|
||||
.toolCall(profilesTool(), (exchange, _) -> {
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
||||
if (denied != null) return denied;
|
||||
return profiles(workers);
|
||||
})
|
||||
.toolCall(whoamiTool(), (exchange, _) -> {
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
||||
if (denied != null) return denied;
|
||||
return whoami(principal(exchange), sessions);
|
||||
})
|
||||
.build();
|
||||
this.authz = callers;
|
||||
this.metrics = metrics;
|
||||
}
|
||||
|
||||
/**
|
||||
* Pre-CB-501 identity: worker if the connection maps to a pane, otherwise the primary. Used
|
||||
* only by the legacy constructor, where authorization is not enforced anyway.
|
||||
*/
|
||||
private static Principal legacyPrincipal(ConnectionIdentity identity, String addr, int port) {
|
||||
ConnectionIdentity.Caller c = identity.resolve(addr, port);
|
||||
return c.terminal() != null
|
||||
? Principal.worker(c.terminal(), c.pid())
|
||||
: Principal.primary(c.pid());
|
||||
}
|
||||
|
||||
/** The caller reconstructed from the transport context. */
|
||||
private static Principal principal(McpSyncServerExchange exchange) {
|
||||
return principalFrom(exchange.transportContext().get(CALLER_ROLE),
|
||||
callerTerminal(exchange), callerPid(exchange));
|
||||
}
|
||||
|
||||
/**
|
||||
* Rebuild a {@link Principal} from the three values the context extractor stashed.
|
||||
*
|
||||
* <p>Split out from {@link #principal(McpSyncServerExchange)} so the identity rules are
|
||||
* reachable without an {@code McpSyncServerExchange} — that is an SDK type this project has no
|
||||
* mocking library to fabricate, which is why this logic had no test at all until CB-513.
|
||||
*
|
||||
* @param role the stashed {@link Role} name, or {@code null} on the legacy path
|
||||
* @param terminal the worker terminal, or {@code null} for a non-worker
|
||||
* @param pid the calling pid, or {@code -1}
|
||||
*/
|
||||
static Principal principalFrom(Object role, String terminal, long pid) {
|
||||
if (role == null) {
|
||||
// No role stashed (legacy path): fall back to the historical interpretation.
|
||||
return terminal != null ? Principal.worker(terminal, pid) : Principal.primary(pid);
|
||||
}
|
||||
return new Principal(Role.valueOf(role.toString()), terminal, pid);
|
||||
}
|
||||
|
||||
/**
|
||||
* Gate a tool call on the CB-505 table. Returns {@code null} when the call may proceed, or the
|
||||
* error result to return when it may not.
|
||||
*/
|
||||
private McpSchema.CallToolResult deny(McpSyncServerExchange exchange, Authz.Action action,
|
||||
String target) {
|
||||
return denyFor(principal(exchange), action, target);
|
||||
}
|
||||
|
||||
/**
|
||||
* The policy half of {@link #deny}: everything except pulling the caller out of the MCP
|
||||
* exchange. Kept separate so the authorization decision — the actual control — is unit-testable
|
||||
* without fabricating an SDK {@code McpSyncServerExchange}.
|
||||
*
|
||||
* <p>This surface exists because the enforcement was previously unreachable from a test: no
|
||||
* test constructs a {@code BridgeMcp}, so the whole MCP-side gate ran zero times in the suite
|
||||
* while the REST-side equivalent had ten tests. A security control nothing exercises is a
|
||||
* claim, not a control.
|
||||
*
|
||||
* @return {@code null} when the call may proceed, or the error result to return when it may not
|
||||
*/
|
||||
McpSchema.CallToolResult denyFor(Principal caller, Authz.Action action, String target) {
|
||||
// The enforcement switch lives HERE rather than in the exchange-facing wrapper: any future
|
||||
// tool that calls this directly must not be able to skip the gate by accident.
|
||||
if (authz == null) {
|
||||
return null; // legacy constructor: authorization not enforced
|
||||
}
|
||||
if (Authz.permits(caller, action, target)) {
|
||||
if (action != Authz.Action.READ) {
|
||||
AuditLog.allowed(caller, action, target); // reads would drown the trail
|
||||
}
|
||||
return null;
|
||||
}
|
||||
String reason = Authz.isUnauthenticated(caller) ? "unauthenticated" : "forbidden";
|
||||
AuditLog.denied(caller, action, target, reason);
|
||||
if (metrics != null) {
|
||||
metrics.inc(BridgedMetrics.AUTH_FAILURES, "reason", reason);
|
||||
}
|
||||
return error(reason + ": " + caller.describe() + " may not " + action);
|
||||
}
|
||||
|
||||
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
|
||||
@@ -235,10 +401,17 @@ public final class BridgeMcp {
|
||||
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
|
||||
}
|
||||
|
||||
/** {@code bridge_poll}: check an async delegation by ticket (pending / done+reply / failed). */
|
||||
static McpSchema.CallToolResult poll(MessageService messages, String ticket) {
|
||||
/** {@code bridge_poll}: check an async delegation by ticket, or drain a worker's inbox by target. */
|
||||
static McpSchema.CallToolResult poll(MessageService messages, String ticket, String target) {
|
||||
if (!isBlank(target)) {
|
||||
var replies = messages.drainReplies(target);
|
||||
if (replies.isEmpty()) {
|
||||
return text("[]");
|
||||
}
|
||||
return text(json(replies));
|
||||
}
|
||||
if (isBlank(ticket)) {
|
||||
return error("ticket is required");
|
||||
return error("ticket (or target) is required");
|
||||
}
|
||||
MessageService.TaskView v = messages.poll(ticket);
|
||||
if (v == null) {
|
||||
@@ -254,11 +427,12 @@ public final class BridgeMcp {
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code bridge_reply}: the worker returns its structured answer, resolving the awaiting send.
|
||||
* {@code bridge_reply}: the worker returns its structured answer, resolving the awaiting send
|
||||
* or — when no send is open — queueing the reply in the inbox for later drain (CB-307).
|
||||
* {@code callerTerminal} is resolved from the connection (never an argument); a {@code null}
|
||||
* means the caller is not a known worker (e.g. the primary called it by mistake).
|
||||
*/
|
||||
static McpSchema.CallToolResult reply(Rendezvous rendezvous, String callerTerminal, String content) {
|
||||
static McpSchema.CallToolResult reply(MessageService messages, String callerTerminal, String content) {
|
||||
if (callerTerminal == null) {
|
||||
return error("bridge_reply is for workers only — could not identify the calling worker "
|
||||
+ "from the connection");
|
||||
@@ -266,9 +440,17 @@ public final class BridgeMcp {
|
||||
if (content == null) {
|
||||
return error("content is required");
|
||||
}
|
||||
return rendezvous.resolve(callerTerminal, content)
|
||||
? text("delivered")
|
||||
: error("no send is awaiting a reply for this worker");
|
||||
messages.reply(callerTerminal, content);
|
||||
return text("delivered");
|
||||
}
|
||||
|
||||
/** {@code bridge_ack}: acknowledge (remove) a specific reply from the inbox. */
|
||||
static McpSchema.CallToolResult ack(MessageService messages, String target, String msgId) {
|
||||
if (isBlank(target) || isBlank(msgId)) {
|
||||
return error("target and msgId are required");
|
||||
}
|
||||
messages.ackReply(target, msgId);
|
||||
return text("acknowledged " + msgId);
|
||||
}
|
||||
|
||||
/** {@code bridge_status}: the live lifecycle status of a worker session. */
|
||||
@@ -283,6 +465,49 @@ public final class BridgeMcp {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code bridge_whoami}: the caller's own identity, as the daemon already resolved it.
|
||||
*
|
||||
* <p>Every other tool <em>consumes</em> this identity — the authorization gate, the reply
|
||||
* rendezvous, the cwd inherit — but none reported it, so an agent had to infer its own role
|
||||
* from side channels the daemon does not control: a charter string in its system prompt, the
|
||||
* name its MCP mount happens to carry, or {@code ANTHROPIC_BASE_URL} (which Claude-model
|
||||
* workers do not set). The failure mode of guessing is asymmetric and silent: a primary that
|
||||
* mistakes itself for a worker is refused by {@link Authz} and learns immediately, while a
|
||||
* worker that mistakes itself for the primary ends its turn without {@code bridge_reply} and
|
||||
* the sender simply receives nothing. This tool removes the guess.
|
||||
*
|
||||
* <p>For a worker the session registry adds what it knows about that session. A worker the
|
||||
* registry has no record of — one that outlived a daemon restart — still gets its role and
|
||||
* {@code sessionId}, which is the load-bearing part.
|
||||
*/
|
||||
static McpSchema.CallToolResult whoami(Principal caller, SessionManager sessions) {
|
||||
Map<String, Object> m = new LinkedHashMap<>();
|
||||
m.put("role", caller.role().name().toLowerCase());
|
||||
if (!caller.isWorker()) {
|
||||
return text(json(m));
|
||||
}
|
||||
m.put("sessionId", caller.terminal());
|
||||
sessions.roster().stream()
|
||||
.filter(s -> caller.terminal().equals(s.terminalId()))
|
||||
.findFirst()
|
||||
.ifPresent(s -> {
|
||||
m.put("paneId", s.paneId());
|
||||
m.put("profile", s.profile());
|
||||
m.put("state", s.state().name().toLowerCase());
|
||||
if (s.worktree() != null) {
|
||||
m.put("worktree", s.worktree());
|
||||
}
|
||||
if (s.branch() != null) {
|
||||
m.put("branch", s.branch());
|
||||
}
|
||||
if (s.ownerTerminal() != null) {
|
||||
m.put("owner", s.ownerTerminal());
|
||||
}
|
||||
});
|
||||
return text(json(m));
|
||||
}
|
||||
|
||||
// --- fleet management logic (CB-108 / CB-301) --------------------------------------------
|
||||
|
||||
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
|
||||
@@ -308,6 +533,8 @@ public final class BridgeMcp {
|
||||
return error("subscription boundary: " + e.getMessage());
|
||||
} catch (IllegalArgumentException e) {
|
||||
return error(e.getMessage()); // unknown / no-default profile
|
||||
} catch (PeerUnreachableException e) {
|
||||
return error("spawn timed out — worker pane never reached injectable state: " + e.getMessage());
|
||||
} catch (HerdrException e) {
|
||||
return error("herdr error spawning worker: " + e.getMessage());
|
||||
}
|
||||
@@ -342,16 +569,17 @@ public final class BridgeMcp {
|
||||
}
|
||||
|
||||
/** {@code bridge_profiles}: the configured worker profiles and the default. */
|
||||
static McpSchema.CallToolResult profiles(ClaudeCodeLauncher workers) {
|
||||
static McpSchema.CallToolResult profiles(PeerLauncher workers) {
|
||||
return text(json(Map.of(
|
||||
"profiles", workers.profiles(),
|
||||
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
|
||||
}
|
||||
|
||||
/** {@code bridge_list}: bridge-owned roster merged with live herdr status by paneId. */
|
||||
static McpSchema.CallToolResult listWorkers(ClaudeCodeLauncher workers, SessionManager sessions) {
|
||||
static McpSchema.CallToolResult listWorkers(PeerLauncher workers, SessionManager sessions) {
|
||||
try {
|
||||
Map<String, Agent> live = workers.list().stream()
|
||||
.map(Agent.class::cast)
|
||||
.filter(a -> a.paneId() != null)
|
||||
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
|
||||
List<Map<String, Object>> out = sessions.roster().stream()
|
||||
@@ -435,10 +663,24 @@ public final class BridgeMcp {
|
||||
private static McpSchema.Tool pollTool() {
|
||||
return tool("bridge_poll",
|
||||
"Check an async delegation (a bridge_send with wait:false) by its ticket: "
|
||||
+ "pending, done (with the worker's reply), or failed.",
|
||||
+ "pending, done (with the worker's reply), or failed. When target (a worker "
|
||||
+ "session id) is present instead of ticket, drain that worker's inbox of "
|
||||
+ "replies delivered when no send was open.",
|
||||
objectSchema(Map.of(
|
||||
"ticket", stringProp("The ticket returned by bridge_send wait:false")),
|
||||
List.of("ticket")));
|
||||
"ticket", stringProp("The ticket returned by bridge_send wait:false"),
|
||||
"target", stringProp("Worker session id to drain pending replies from (optional)")),
|
||||
List.of()));
|
||||
}
|
||||
|
||||
private static McpSchema.Tool ackTool() {
|
||||
return tool("bridge_ack",
|
||||
"Acknowledge (remove) a specific reply from a worker's inbox. Use when the primary "
|
||||
+ "has processed a reply and wants to confirm it, leaving other pending replies "
|
||||
+ "in the inbox for later drain.",
|
||||
objectSchema(Map.of(
|
||||
"target", stringProp("Worker session id whose inbox to ack from"),
|
||||
"msgId", stringProp("The message id to acknowledge")),
|
||||
List.of("target", "msgId")));
|
||||
}
|
||||
|
||||
private static McpSchema.Tool spawnTool() {
|
||||
@@ -495,6 +737,18 @@ public final class BridgeMcp {
|
||||
List.of("sessionId")));
|
||||
}
|
||||
|
||||
private static McpSchema.Tool whoamiTool() {
|
||||
return tool("bridge_whoami",
|
||||
"Report who YOU are on the bridge — your role is resolved from your connection "
|
||||
+ "(unforgeable), never from anything you claim. Returns role 'primary' (you "
|
||||
+ "orchestrate: spawn/send/stop, and you must never call bridge_reply) or "
|
||||
+ "'worker' (you were delegated to: you must end every turn with exactly one "
|
||||
+ "bridge_reply, and cannot spawn or send), plus your own sessionId, profile, "
|
||||
+ "worktree and branch when you are a worker. Call this first when following "
|
||||
+ "role-conditional instructions rather than guessing your role.",
|
||||
objectSchema(Map.of(), List.of()));
|
||||
}
|
||||
|
||||
// --- small helpers -------------------------------------------------------------------------
|
||||
|
||||
// The SDK 2.0.0 deprecates its own Tool builders without a stable replacement — isolate it here.
|
||||
|
||||
@@ -0,0 +1,68 @@
|
||||
package dev.ltms.bridged.mcp;
|
||||
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.Optional;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
|
||||
/**
|
||||
* Single-slot, thread-safe registry for the primary's herdr {@code terminal_id}.
|
||||
*
|
||||
* <p>Populated from the caller terminal of orchestration-side MCP tools
|
||||
* ({@code bridge_send}, {@code bridge_spawn}) — tools that only the primary calls.
|
||||
* A pinned terminal (from config) seeds the registry at construction and makes
|
||||
* subsequent {@link #record(String)} calls no-ops.
|
||||
*
|
||||
* <p>The push loop ({@code ReplyPushLoop}) uses {@link #isKnown()} to decide
|
||||
* whether active nudging is possible; an empty registry means the primary is
|
||||
* off-host or non-herdr and delivery falls back to pull.
|
||||
*/
|
||||
public final class PrimaryRegistry {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(PrimaryRegistry.class);
|
||||
|
||||
private final AtomicReference<String> terminal = new AtomicReference<>();
|
||||
private final boolean pinned;
|
||||
|
||||
/**
|
||||
* @param pinnedTerminal an optional pinned terminal from config ({@code null}/blank = unpinned)
|
||||
*/
|
||||
public PrimaryRegistry(String pinnedTerminal) {
|
||||
if (pinnedTerminal != null && !pinnedTerminal.isBlank()) {
|
||||
this.terminal.set(pinnedTerminal);
|
||||
this.pinned = true;
|
||||
log.info("primary terminal pinned: {}", pinnedTerminal);
|
||||
} else {
|
||||
this.pinned = false;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Record a terminal_id. No-op when:
|
||||
* <ul>
|
||||
* <li>the registry is pinned (config override),
|
||||
* <li>{@code terminalId} is {@code null} or blank (non-herdr caller).
|
||||
* </ul>
|
||||
*/
|
||||
public void record(String terminalId) {
|
||||
if (pinned) return;
|
||||
if (terminalId == null || terminalId.isBlank()) return;
|
||||
String prev = terminal.getAndSet(terminalId);
|
||||
if (prev == null) {
|
||||
log.debug("primary terminal learned: {}", terminalId);
|
||||
} else if (!prev.equals(terminalId)) {
|
||||
log.debug("primary terminal changed: {} -> {}", prev, terminalId);
|
||||
}
|
||||
}
|
||||
|
||||
/** The known primary terminal, or empty if not yet learned (and not pinned). */
|
||||
public Optional<String> primaryTerminal() {
|
||||
return Optional.ofNullable(terminal.get());
|
||||
}
|
||||
|
||||
/** {@code true} once a terminal has been recorded (or was pinned at construction). */
|
||||
public boolean isKnown() {
|
||||
return terminal.get() != null;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,96 @@
|
||||
package dev.ltms.bridged.metrics;
|
||||
|
||||
import dev.ltms.bridged.msg.ReplyInbox;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.session.WorkerSession;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* The daemon's metric definitions (CB-502) — one place where every series is named, described, and
|
||||
* (for gauges) bound to live state.
|
||||
*
|
||||
* <p>The set is deliberately small: each series maps to a failure mode this project has actually
|
||||
* hit, not to whatever was easy to count. The two worth watching in practice are
|
||||
* {@code bridged_sends_total{outcome="completion_fallback"}} — a rising share means turn detection
|
||||
* is degrading, the CB-115/116/118 failure family — and
|
||||
* {@code bridged_push_nudges_total{outcome="exhausted"}}, which means the primary stopped draining
|
||||
* its inbox and CB-307's active push gave up.
|
||||
*/
|
||||
public final class BridgedMetrics {
|
||||
|
||||
/** Counter: delegated sends by terminal outcome. */
|
||||
public static final String SENDS = "bridged_sends_total";
|
||||
/** Counter: worker replies by the path that carried them (rendezvous vs stranded-to-inbox). */
|
||||
public static final String REPLIES = "bridged_replies_total";
|
||||
/** Counter: push-loop nudges to the primary, by outcome. */
|
||||
public static final String PUSH_NUDGES = "bridged_push_nudges_total";
|
||||
/** Counter: spawn attempts by peer kind and outcome. */
|
||||
public static final String SPAWNS = "bridged_spawns_total";
|
||||
/** Counter: herdr socket calls by method and outcome. */
|
||||
public static final String HERDR_CALLS = "bridged_herdr_calls_total";
|
||||
/** Counter: rejected requests by reason (CB-501). */
|
||||
public static final String AUTH_FAILURES = "bridged_auth_failures_total";
|
||||
/** Gauge: session census by lifecycle state. */
|
||||
public static final String SESSIONS = "bridged_sessions";
|
||||
/** Gauge: undrained replies held per target. */
|
||||
public static final String INBOX_DEPTH = "bridged_inbox_depth";
|
||||
|
||||
private BridgedMetrics() {
|
||||
}
|
||||
|
||||
/**
|
||||
* Build the registry with its help text and live gauges bound.
|
||||
*
|
||||
* @param sessions the authoritative session registry (census gauge)
|
||||
* @param inbox the reply inbox; only used for a depth gauge when it can be inspected
|
||||
*/
|
||||
public static Metrics create(SessionManager sessions, ReplyInbox inbox) {
|
||||
Metrics m = new Metrics();
|
||||
|
||||
m.describe(SENDS, "counter",
|
||||
"Delegated sends by terminal outcome (replied|completion_fallback|timeout|failed).");
|
||||
m.describe(REPLIES, "counter",
|
||||
"Worker replies by delivery path (rendezvous=resolved an open send, inbox=stranded and held).");
|
||||
m.describe(PUSH_NUDGES, "counter",
|
||||
"CB-307 push-loop nudges to the primary (delivered|exhausted).");
|
||||
m.describe(SPAWNS, "counter",
|
||||
"Worker spawn attempts by peer kind and outcome (ready|timeout|guard_rejected).");
|
||||
m.describe(HERDR_CALLS, "counter",
|
||||
"herdr socket calls by method and outcome — the dependency everything else rests on.");
|
||||
m.describe(AUTH_FAILURES, "counter",
|
||||
"Requests refused by CB-501/505 (unauthenticated|forbidden).");
|
||||
m.describe(SESSIONS, "gauge",
|
||||
"Registered worker sessions by lifecycle state.");
|
||||
m.describe(INBOX_DEPTH, "gauge",
|
||||
"Replies held for a target that the primary has not drained. Steady state is 0; "
|
||||
+ "a target stuck above 0 means CB-307 delivery is not completing.");
|
||||
|
||||
// One gauge per state so a scrape shows the whole census even when a state is empty —
|
||||
// an absent series and a zero series read very differently on a dashboard.
|
||||
for (WorkerSession.State state : WorkerSession.State.values()) {
|
||||
String label = state.name().toLowerCase();
|
||||
m.gauge(SESSIONS, () -> countIn(sessions, state), "state", label);
|
||||
}
|
||||
|
||||
// Depth is per live session, so the label set is only known at scrape time. peek() is the
|
||||
// port's non-destructive read — scraping metrics must never ack a reply out of the inbox.
|
||||
m.collector(INBOX_DEPTH, "target", () -> {
|
||||
Map<String, Number> depths = new LinkedHashMap<>();
|
||||
for (WorkerSession s : sessions.roster()) {
|
||||
String target = s.terminalId();
|
||||
if (target == null) {
|
||||
continue;
|
||||
}
|
||||
depths.put(target, inbox.peek(target).size());
|
||||
}
|
||||
return depths;
|
||||
});
|
||||
return m;
|
||||
}
|
||||
|
||||
private static long countIn(SessionManager sessions, WorkerSession.State state) {
|
||||
return sessions.roster().stream().filter(s -> s.state() == state).count();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,181 @@
|
||||
package dev.ltms.bridged.metrics;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.NavigableMap;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.ConcurrentSkipListMap;
|
||||
import java.util.concurrent.atomic.LongAdder;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
* The daemon's metric registry and Prometheus text renderer (CB-502).
|
||||
*
|
||||
* <p>Deliberately dependency-free. The roadmap's tech-stack table specified Micrometer, but this
|
||||
* pom already carries an unusual reconciliation burden (a hand-pinned {@code jackson-annotations}
|
||||
* to make the MCP SDK's Jackson 3 coexist with our Jackson 2, a Jetty BOM import to stop version
|
||||
* skew, and four documented accepted-CVE advisories), and the dependency CVE gate this project
|
||||
* mandates could not be run when this landed. The metric set is small and fully known, and
|
||||
* Prometheus text exposition is a stable, well-specified format — so the registry is ~100 lines
|
||||
* here instead of a new transitive tree. {@code GET /metrics} is the swap seam if Micrometer's
|
||||
* ecosystem is ever wanted.
|
||||
*
|
||||
* <p>Thread-safe: counters are {@link LongAdder} (built for contended increment), gauges are
|
||||
* supplier-backed so they read live state at scrape time rather than needing to be pushed.
|
||||
*/
|
||||
public final class Metrics {
|
||||
|
||||
/** Counter series, keyed by the fully-rendered {@code name{labels}} sample id. */
|
||||
private final NavigableMap<String, LongAdder> counters = new ConcurrentSkipListMap<>();
|
||||
/** Gauge series, evaluated at scrape time. */
|
||||
private final NavigableMap<String, Supplier<Number>> gauges = new ConcurrentSkipListMap<>();
|
||||
/** Gauge families whose label set is only known at scrape time, keyed by metric name. */
|
||||
private final NavigableMap<String, Collector> collectors = new ConcurrentSkipListMap<>();
|
||||
/** HELP/TYPE metadata, keyed by bare metric name. */
|
||||
private final Map<String, String[]> meta = new ConcurrentHashMap<>();
|
||||
|
||||
/** A gauge family whose series are discovered per scrape (one label, many values). */
|
||||
private record Collector(String labelName, Supplier<Map<String, Number>> samples) {
|
||||
}
|
||||
|
||||
/** Declare a metric's help text and type once, so the exposition carries HELP/TYPE lines. */
|
||||
public Metrics describe(String name, String type, String help) {
|
||||
meta.put(name, new String[]{type, help});
|
||||
return this;
|
||||
}
|
||||
|
||||
/** Increment a counter by one. */
|
||||
public void inc(String name, String... labelPairs) {
|
||||
add(name, 1, labelPairs);
|
||||
}
|
||||
|
||||
/** Increment a counter by {@code delta}. */
|
||||
public void add(String name, long delta, String... labelPairs) {
|
||||
counters.computeIfAbsent(sample(name, labelPairs), _ -> new LongAdder()).add(delta);
|
||||
}
|
||||
|
||||
/**
|
||||
* Register a live gauge. The supplier is called at scrape time, so it reflects current state
|
||||
* (session census, inbox depth) without anything having to remember to update it.
|
||||
*/
|
||||
public void gauge(String name, Supplier<Number> value, String... labelPairs) {
|
||||
gauges.put(sample(name, labelPairs), value);
|
||||
}
|
||||
|
||||
/**
|
||||
* Register a gauge family whose label values are not known up front — inbox depth per target,
|
||||
* for instance, where the set of targets changes as workers come and go. The supplier returns
|
||||
* {@code labelValue → value} and is evaluated once per scrape.
|
||||
*/
|
||||
public void collector(String name, String labelName, Supplier<Map<String, Number>> samples) {
|
||||
collectors.put(name, new Collector(labelName, samples));
|
||||
}
|
||||
|
||||
/** Current value of a counter series — for assertions in tests. */
|
||||
public long count(String name, String... labelPairs) {
|
||||
LongAdder a = counters.get(sample(name, labelPairs));
|
||||
return a == null ? 0 : a.sum();
|
||||
}
|
||||
|
||||
/**
|
||||
* Render the Prometheus text exposition format (version 0.0.4): optional {@code # HELP} and
|
||||
* {@code # TYPE} lines per metric family, then one line per sample.
|
||||
*/
|
||||
public String render() {
|
||||
StringBuilder out = new StringBuilder(1024);
|
||||
String lastFamily = null;
|
||||
for (Map.Entry<String, LongAdder> e : counters.entrySet()) {
|
||||
lastFamily = emitHeader(out, e.getKey(), lastFamily);
|
||||
out.append(e.getKey()).append(' ').append(e.getValue().sum()).append('\n');
|
||||
}
|
||||
for (Map.Entry<String, Supplier<Number>> e : gauges.entrySet()) {
|
||||
lastFamily = emitHeader(out, e.getKey(), lastFamily);
|
||||
Number v;
|
||||
try {
|
||||
v = e.getValue().get();
|
||||
} catch (RuntimeException ex) {
|
||||
continue; // a broken gauge must never break the whole scrape
|
||||
}
|
||||
if (v == null) {
|
||||
continue;
|
||||
}
|
||||
out.append(e.getKey()).append(' ').append(format(v)).append('\n');
|
||||
}
|
||||
for (Map.Entry<String, Collector> e : collectors.entrySet()) {
|
||||
Map<String, Number> samples;
|
||||
try {
|
||||
samples = e.getValue().samples().get();
|
||||
} catch (RuntimeException ex) {
|
||||
continue; // a broken collector must never break the whole scrape
|
||||
}
|
||||
if (samples == null || samples.isEmpty()) {
|
||||
continue;
|
||||
}
|
||||
lastFamily = emitHeader(out, e.getKey(), lastFamily);
|
||||
// Sort so repeated scrapes are byte-stable and diffable.
|
||||
new java.util.TreeMap<>(samples).forEach((label, v) -> {
|
||||
if (v != null) {
|
||||
out.append(sample(e.getKey(), e.getValue().labelName(), label))
|
||||
.append(' ').append(format(v)).append('\n');
|
||||
}
|
||||
});
|
||||
}
|
||||
return out.toString();
|
||||
}
|
||||
|
||||
/** Emit HELP/TYPE when the sample starts a new metric family; returns the current family. */
|
||||
private String emitHeader(StringBuilder out, String sampleId, String lastFamily) {
|
||||
String family = familyOf(sampleId);
|
||||
if (family.equals(lastFamily)) {
|
||||
return lastFamily;
|
||||
}
|
||||
String[] m = meta.get(family);
|
||||
if (m != null) {
|
||||
out.append("# HELP ").append(family).append(' ').append(m[1]).append('\n');
|
||||
out.append("# TYPE ").append(family).append(' ').append(m[0]).append('\n');
|
||||
}
|
||||
return family;
|
||||
}
|
||||
|
||||
private static String familyOf(String sampleId) {
|
||||
int brace = sampleId.indexOf('{');
|
||||
return brace < 0 ? sampleId : sampleId.substring(0, brace);
|
||||
}
|
||||
|
||||
/** Whole numbers render without a decimal point; everything else as-is. */
|
||||
private static String format(Number v) {
|
||||
double d = v.doubleValue();
|
||||
return (d == Math.rint(d) && !Double.isInfinite(d))
|
||||
? Long.toString((long) d)
|
||||
: Double.toString(d);
|
||||
}
|
||||
|
||||
/** Build the {@code name{k="v",k2="v2"}} sample id; labels are sorted for stable output. */
|
||||
private static String sample(String name, String... labelPairs) {
|
||||
if (labelPairs == null || labelPairs.length == 0) {
|
||||
return name;
|
||||
}
|
||||
if (labelPairs.length % 2 != 0) {
|
||||
throw new IllegalArgumentException("labels must be key/value pairs, got " + labelPairs.length);
|
||||
}
|
||||
NavigableMap<String, String> sorted = new java.util.TreeMap<>();
|
||||
for (int i = 0; i < labelPairs.length; i += 2) {
|
||||
sorted.put(labelPairs[i], labelPairs[i + 1] == null ? "" : labelPairs[i + 1]);
|
||||
}
|
||||
StringBuilder sb = new StringBuilder(name.length() + 16 * sorted.size());
|
||||
sb.append(name).append('{');
|
||||
boolean first = true;
|
||||
for (Map.Entry<String, String> e : sorted.entrySet()) {
|
||||
if (!first) {
|
||||
sb.append(',');
|
||||
}
|
||||
first = false;
|
||||
sb.append(e.getKey()).append("=\"").append(escapeLabel(e.getValue())).append('"');
|
||||
}
|
||||
return sb.append('}').toString();
|
||||
}
|
||||
|
||||
/** Label values are escaped per the exposition format: backslash, quote, newline. */
|
||||
private static String escapeLabel(String v) {
|
||||
return v.replace("\\", "\\\\").replace("\"", "\\\"").replace("\n", "\\n");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,233 @@
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import com.rabbitmq.client.AMQP;
|
||||
import com.rabbitmq.client.Channel;
|
||||
import com.rabbitmq.client.Connection;
|
||||
import com.rabbitmq.client.ConnectionFactory;
|
||||
import com.rabbitmq.client.DeliverCallback;
|
||||
import com.rabbitmq.client.Recoverable;
|
||||
import com.rabbitmq.client.RecoveryListener;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
|
||||
/**
|
||||
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
|
||||
* port {@link InMemoryReplyInbox} implements as soft state.
|
||||
*
|
||||
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target owns a durable
|
||||
* queue {@code agent.<target>.inbox}. A manual-ack consumer pulls persistent messages off that queue
|
||||
* into an in-memory <em>held</em> map (keyed by {@code msgId}) but does <em>not</em> ack them.
|
||||
* {@link #peek} returns that snapshot; {@link #ack} acks the broker delivery-tag and drops the entry.
|
||||
* Because messages stay unacked until the primary actually drains them, a crash (or a {@code java -jar}
|
||||
* bounce) before caller-ack leaves them on the broker — it redelivers on reconnect. That is the
|
||||
* durability the in-memory adapter cannot give, with the port contract preserved.
|
||||
*
|
||||
* <p><strong>Dedup.</strong> The consumer keys the held map by {@code msgId}; a redelivered duplicate
|
||||
* (at-least-once, or a producer double-publish) is acked-and-dropped on arrival, so it never
|
||||
* double-queues. {@link #publish} additionally short-circuits an already-held {@code msgId} — a
|
||||
* fast path; the consumer-side check is the real guarantee.
|
||||
*
|
||||
* <p><strong>Visibility.</strong> Unlike the in-memory adapter, publish → broker → consumer is
|
||||
* asynchronous, so a {@link #peek} immediately after {@link #publish} may not yet see the message
|
||||
* (broker delivery latency). Callers that need the reply drained poll (as the primary already does);
|
||||
* the contract test waits for visibility. This is inherent to broker-backed delivery, not a defect.
|
||||
*
|
||||
* <p>The default deploy targets LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1 (URI-only swap),
|
||||
* so the {@code @Tag("contract")} integration test runs against a RabbitMQ container.
|
||||
*/
|
||||
public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(AmqpReplyInbox.class);
|
||||
|
||||
private static final String QUEUE_PREFIX = "agent.";
|
||||
private static final String QUEUE_SUFFIX = ".inbox";
|
||||
|
||||
private final Connection connection;
|
||||
private final Channel channel;
|
||||
/** All channel operations (publish/declare/ack) serialize on this — a Channel is not thread-safe. */
|
||||
private final Object channelLock = new Object();
|
||||
/** target → (msgId → held delivery). Per-target map is guarded by synchronizing on itself. */
|
||||
private final ConcurrentHashMap<String, LinkedHashMap<String, Held>> held = new ConcurrentHashMap<>();
|
||||
/** Targets whose queue is declared and consumer is running. */
|
||||
private final Set<String> consuming = ConcurrentHashMap.newKeySet();
|
||||
|
||||
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
|
||||
private record Held(long deliveryTag, InboxMessage message) {}
|
||||
|
||||
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) and open the inbox. */
|
||||
public static AmqpReplyInbox open(String uri) {
|
||||
try {
|
||||
ConnectionFactory factory = new ConnectionFactory();
|
||||
factory.setUri(uri);
|
||||
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
|
||||
factory.setAutomaticRecoveryEnabled(true);
|
||||
factory.setTopologyRecoveryEnabled(true);
|
||||
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"));
|
||||
} catch (Exception e) {
|
||||
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
|
||||
}
|
||||
}
|
||||
|
||||
/** Wrap an already-open connection (injection seam for the contract test). */
|
||||
AmqpReplyInbox(Connection connection) {
|
||||
this.connection = connection;
|
||||
try {
|
||||
this.channel = connection.createChannel();
|
||||
} catch (IOException e) {
|
||||
throw new IllegalStateException("cannot open AMQP channel", e);
|
||||
}
|
||||
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
|
||||
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
|
||||
// repopulates it with valid tags (dedup by msgId still prevents any double-queue).
|
||||
if (connection instanceof Recoverable recoverable) {
|
||||
recoverable.addRecoveryListener(new RecoveryListener() {
|
||||
@Override
|
||||
public void handleRecovery(Recoverable recoverable) {
|
||||
held.clear();
|
||||
log.info("AMQP connection recovered; cleared held replies for fresh redelivery");
|
||||
}
|
||||
|
||||
@Override
|
||||
public void handleRecoveryStarted(Recoverable recoverable) {
|
||||
// no-op: we act once recovery completes
|
||||
}
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
ensureConsuming(target);
|
||||
var perTarget = held.get(target);
|
||||
if (perTarget != null) {
|
||||
synchronized (perTarget) {
|
||||
if (perTarget.containsKey(msgId)) {
|
||||
return; // already held — producer-side fast dedup
|
||||
}
|
||||
}
|
||||
}
|
||||
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
|
||||
.messageId(msgId)
|
||||
.deliveryMode(2) // persistent — survives a broker restart
|
||||
.contentType("text/plain")
|
||||
.build();
|
||||
try {
|
||||
synchronized (channelLock) {
|
||||
channel.basicPublish("", queueName(target), props, content.getBytes(StandardCharsets.UTF_8));
|
||||
}
|
||||
} catch (IOException e) {
|
||||
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
ensureConsuming(target);
|
||||
var perTarget = held.get(target);
|
||||
if (perTarget == null) {
|
||||
return List.of();
|
||||
}
|
||||
synchronized (perTarget) {
|
||||
return perTarget.values().stream().map(Held::message).toList();
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public void ack(String target, String msgId) {
|
||||
var perTarget = held.get(target);
|
||||
if (perTarget == null) {
|
||||
return;
|
||||
}
|
||||
Held h;
|
||||
synchronized (perTarget) {
|
||||
h = perTarget.remove(msgId);
|
||||
}
|
||||
if (h == null) {
|
||||
return; // never held (or already acked) — no-op
|
||||
}
|
||||
try {
|
||||
synchronized (channelLock) {
|
||||
channel.basicAck(h.deliveryTag(), false);
|
||||
}
|
||||
} catch (IOException e) {
|
||||
// Ack didn't reach the broker: restore the entry so a later ack (or a redelivery after
|
||||
// reconnect) can retry. Keeps the at-least-once contract — a reply is never silently lost.
|
||||
synchronized (perTarget) {
|
||||
perTarget.putIfAbsent(msgId, h);
|
||||
}
|
||||
throw new IllegalStateException("cannot ack reply " + msgId + " on " + queueName(target), e);
|
||||
}
|
||||
}
|
||||
|
||||
/** Declare the durable per-target queue and start its manual-ack consumer, once per target. */
|
||||
private void ensureConsuming(String target) {
|
||||
if (consuming.contains(target)) {
|
||||
return;
|
||||
}
|
||||
synchronized (channelLock) {
|
||||
if (!consuming.add(target)) {
|
||||
return; // another thread just set it up
|
||||
}
|
||||
String queue = queueName(target);
|
||||
try {
|
||||
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
|
||||
channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
|
||||
} catch (IOException e) {
|
||||
consuming.remove(target);
|
||||
throw new IllegalStateException("cannot consume queue " + queue, e);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private DeliverCallback deliverCallback(String target) {
|
||||
return (_, delivery) -> {
|
||||
String msgId = delivery.getProperties().getMessageId();
|
||||
long tag = delivery.getEnvelope().getDeliveryTag();
|
||||
if (msgId == null || msgId.isBlank()) {
|
||||
msgId = Long.toHexString(tag); // synthesize an id so dedup still has a key
|
||||
}
|
||||
String content = new String(delivery.getBody(), StandardCharsets.UTF_8);
|
||||
var perTarget = held.computeIfAbsent(target, _ -> new LinkedHashMap<>());
|
||||
boolean duplicate;
|
||||
synchronized (perTarget) {
|
||||
if (perTarget.containsKey(msgId)) {
|
||||
duplicate = true;
|
||||
} else {
|
||||
perTarget.put(msgId, new Held(tag, new InboxMessage(msgId, target, content)));
|
||||
duplicate = false;
|
||||
}
|
||||
}
|
||||
if (duplicate) {
|
||||
// Redelivered duplicate: ack the new tag and drop it so the broker stops resending.
|
||||
synchronized (channelLock) {
|
||||
channel.basicAck(tag, false);
|
||||
}
|
||||
}
|
||||
};
|
||||
}
|
||||
|
||||
private static String queueName(String target) {
|
||||
return QUEUE_PREFIX + target + QUEUE_SUFFIX;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
try {
|
||||
channel.close();
|
||||
} catch (Exception e) {
|
||||
log.debug("AMQP channel close: {}", e.toString());
|
||||
}
|
||||
try {
|
||||
connection.close();
|
||||
} catch (Exception e) {
|
||||
log.debug("AMQP connection close: {}", e.toString());
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,50 @@
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
|
||||
/**
|
||||
* Soft-state {@link ReplyInbox} backed by a {@link ConcurrentHashMap} keyed by target session.
|
||||
* Per-target FIFO ordering (insertion order via {@link LinkedHashMap}). Dedup by {@code msgId}
|
||||
* within a target. Thread-safe for concurrent publish vs. drain.
|
||||
*
|
||||
* <p><strong>This is soft-state, NOT persistence.</strong> Lost on a {@code java -jar} bounce — that
|
||||
* is correct and consistent with "bridged stays soft-state." The Stage-2 AMQP adapter replaces this.
|
||||
*/
|
||||
public final class InMemoryReplyInbox implements ReplyInbox {
|
||||
|
||||
private final ConcurrentHashMap<String, LinkedHashMap<String, InboxMessage>> store = new ConcurrentHashMap<>();
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
var perTarget = store.computeIfAbsent(target, _ -> new LinkedHashMap<>());
|
||||
//noinspection SynchronizationOnLocalVariableOrMethodParameter
|
||||
synchronized (perTarget) {
|
||||
perTarget.putIfAbsent(msgId, new InboxMessage(msgId, target, content));
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
var perTarget = store.get(target);
|
||||
if (perTarget == null) {
|
||||
return List.of();
|
||||
}
|
||||
//noinspection SynchronizationOnLocalVariableOrMethodParameter
|
||||
synchronized (perTarget) {
|
||||
return List.copyOf(perTarget.values());
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public void ack(String target, String msgId) {
|
||||
var perTarget = store.get(target);
|
||||
if (perTarget != null) {
|
||||
//noinspection SynchronizationOnLocalVariableOrMethodParameter
|
||||
synchronized (perTarget) {
|
||||
perTarget.remove(msgId);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -3,9 +3,13 @@ package dev.ltms.bridged.msg;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.AgentStatus;
|
||||
import dev.ltms.bridged.inject.Injector;
|
||||
import dev.ltms.bridged.metrics.BridgedMetrics;
|
||||
import dev.ltms.bridged.metrics.Metrics;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.UUID;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.CompletionException;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
@@ -147,16 +151,50 @@ public final class MessageService {
|
||||
private final AgentControl agents;
|
||||
private final Injector injector;
|
||||
private final Rendezvous rendezvous;
|
||||
private final ReplyInbox inbox;
|
||||
private final ReplyPushLoop pushLoop;
|
||||
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
|
||||
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
|
||||
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
|
||||
private final AtomicLong ticketSeq = new AtomicLong();
|
||||
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
|
||||
Thread.ofVirtual().name("bridge-async-", 0).factory());
|
||||
|
||||
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous) {
|
||||
/**
|
||||
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
|
||||
*
|
||||
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
|
||||
* branch ({@link #reply}) so it can nudge the primary to drain the inbox
|
||||
*/
|
||||
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
|
||||
ReplyInbox inbox, ReplyPushLoop pushLoop) {
|
||||
this(agents, injector, rendezvous, inbox, pushLoop, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, with a metric registry (CB-502). Instrumenting here rather than at the REST and MCP
|
||||
* edges means both surfaces are counted by one piece of code and cannot drift.
|
||||
*
|
||||
* @param metrics nullable — when null, nothing is recorded
|
||||
*/
|
||||
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
|
||||
ReplyInbox inbox, ReplyPushLoop pushLoop, Metrics metrics) {
|
||||
this.agents = agents;
|
||||
this.injector = injector;
|
||||
this.rendezvous = rendezvous;
|
||||
this.inbox = inbox;
|
||||
this.pushLoop = pushLoop;
|
||||
this.metrics = metrics;
|
||||
}
|
||||
|
||||
/** Create with an explicit {@link ReplyInbox} and no push loop. */
|
||||
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox) {
|
||||
this(agents, injector, rendezvous, inbox, null);
|
||||
}
|
||||
|
||||
/** Backward-compatible constructor that uses a default {@link InMemoryReplyInbox}. */
|
||||
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous) {
|
||||
this(agents, injector, rendezvous, new InMemoryReplyInbox());
|
||||
}
|
||||
|
||||
/** Current lifecycle status of a worker (the {@code GET /sessions/{id}/status} surface). */
|
||||
@@ -164,6 +202,109 @@ public final class MessageService {
|
||||
return agents.status(target);
|
||||
}
|
||||
|
||||
/**
|
||||
* Route a worker's explicit {@code bridge_reply}: resolve an open send, or queue it in the
|
||||
* inbox if no send is currently open. Unlike the bare {@link Rendezvous#resolve}, a no-waiter
|
||||
* result is <em>not</em> a failure — the reply is held for later drain.
|
||||
*
|
||||
* <p><strong>Do NOT use this for mid-turn questions.</strong> {@code bridge_ask} /
|
||||
* {@link Rendezvous#resolveQuestion} must keep today's {@code NO_WAITER} behaviour — questions
|
||||
* are interactive and must never be queued.
|
||||
*
|
||||
* @return always {@code true} — the reply either resolved a live send or was queued
|
||||
*/
|
||||
public boolean reply(String session, String content) {
|
||||
if (rendezvous.resolve(session, content)) {
|
||||
count(BridgedMetrics.REPLIES, "path", "rendezvous");
|
||||
return true; // a live send took it — unchanged fast path
|
||||
}
|
||||
inbox.publish(session, UUID.randomUUID().toString(), content);
|
||||
// A rising inbox share is the signal CB-307 exists to make visible: the worker finished but
|
||||
// nobody was waiting, so delivery now depends on the push loop and a drain.
|
||||
count(BridgedMetrics.REPLIES, "path", "inbox");
|
||||
if (pushLoop != null) {
|
||||
pushLoop.onReplyQueued(session);
|
||||
}
|
||||
return true; // held, not lost
|
||||
}
|
||||
|
||||
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
|
||||
private void count(String name, String... labels) {
|
||||
if (metrics != null) {
|
||||
metrics.inc(name, labels);
|
||||
}
|
||||
}
|
||||
|
||||
/** Count a send's terminal outcome and pass the reply through unchanged. */
|
||||
private Reply recorded(Reply r) {
|
||||
String label = sendOutcomeLabel(r.outcome());
|
||||
if (label != null) {
|
||||
count(BridgedMetrics.SENDS, "outcome", label);
|
||||
}
|
||||
return r;
|
||||
}
|
||||
|
||||
/** Map a terminal send outcome to its metric label, or {@code null} for non-terminal ones. */
|
||||
private static String sendOutcomeLabel(Outcome o) {
|
||||
return switch (o) {
|
||||
case REPLIED -> "replied";
|
||||
case COMPLETED_UNREPLIED -> "completion_fallback";
|
||||
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> "timeout";
|
||||
case WORKER_FAILED -> "failed";
|
||||
case STALE_TURN, QUESTION -> null; // not a completed delegation
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Abandon any send still waiting on {@code target} because its session has gone away (CB-516).
|
||||
*
|
||||
* <p>Without this, tearing a worker down left its rendezvous waiter open: a blocking
|
||||
* {@code bridge_send} kept blocking, and an async one kept reporting {@code PENDING} until
|
||||
* {@link #ASYNC_TIMEOUT_MS} — thirty minutes — even though the worker provably no longer
|
||||
* existed and the delegation could never complete. Worse, {@code poll} already had the evidence
|
||||
* (it calls {@code liveStatus} to build its detail string and gets back {@code "unknown"}) and
|
||||
* reported {@code PENDING} anyway.
|
||||
*
|
||||
* <p>Resolving the waiter as a failure — rather than letting it time out — also means the
|
||||
* outcome is counted, so a torn-down delegation stops being invisible to {@code /metrics}.
|
||||
*
|
||||
* @return true if a live waiter was failed
|
||||
*/
|
||||
public boolean abandon(String target, String reason) {
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
|
||||
if (waiter == null || waiter.isDone()) {
|
||||
return false; // nobody is blocked on this worker — nothing to abandon
|
||||
}
|
||||
boolean failed = rendezvous.resolveFailure(waiter, reason);
|
||||
if (failed) {
|
||||
log.debug("abandoned send to {}: {}", target, reason);
|
||||
}
|
||||
return failed;
|
||||
}
|
||||
|
||||
/**
|
||||
* Acknowledge a specific reply by {@code msgId} for {@code target}. Removes it from the inbox
|
||||
* so that a subsequent drain or peek no longer returns it.
|
||||
*/
|
||||
public void ackReply(String target, String msgId) {
|
||||
inbox.ack(target, msgId);
|
||||
}
|
||||
|
||||
/**
|
||||
* Drain (peek + ack) all pending inbox replies for {@code target}. At-least-once: returns the
|
||||
* messages and acknowledges them; an in-flight failure between returning and the caller
|
||||
* processing them re-surfaces them on a subsequent drain (the ack is local).
|
||||
*
|
||||
* @return the drained messages, newest last (FIFO); empty list if none
|
||||
*/
|
||||
public List<ReplyInbox.InboxMessage> drainReplies(String target) {
|
||||
var messages = inbox.peek(target);
|
||||
for (var msg : messages) {
|
||||
inbox.ack(target, msg.msgId());
|
||||
}
|
||||
return messages;
|
||||
}
|
||||
|
||||
/**
|
||||
* Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the
|
||||
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
|
||||
@@ -180,11 +321,12 @@ public final class MessageService {
|
||||
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
|
||||
try {
|
||||
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
|
||||
return new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
|
||||
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
|
||||
} catch (TimeoutException e) {
|
||||
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
|
||||
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
|
||||
return new Reply(wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null);
|
||||
return recorded(new Reply(
|
||||
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
|
||||
} catch (ExecutionException e) {
|
||||
Throwable cause = e.getCause();
|
||||
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
|
||||
|
||||
@@ -0,0 +1,31 @@
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Holds terminal worker→primary replies that arrive with no live send to resolve, keyed by worker
|
||||
* session (target), until the primary drains them. Soft-state in Stage 1 (in-memory, lost on restart);
|
||||
* the Stage 2 AMQP adapter implements the same contract with cross-restart durability.
|
||||
*
|
||||
* <p><strong>This interface is the port.</strong> {@link InMemoryReplyInbox} is the Stage-1 adapter;
|
||||
* an AMQP-backed adapter (Stage 2) must implement the same contract (idempotent publish, FIFO peek,
|
||||
* at-least-once ack).
|
||||
*/
|
||||
public interface ReplyInbox {
|
||||
|
||||
/** A queued reply: an idempotency id, the worker session it came from, and the reply text. */
|
||||
record InboxMessage(String msgId, String target, String content) {}
|
||||
|
||||
/**
|
||||
* Queue {@code content} from worker {@code target} under {@code msgId}. Idempotent: publishing an
|
||||
* already-present {@code msgId} for {@code target} is a no-op (dedup), so an at-least-once Stage-2
|
||||
* redelivery cannot double-queue.
|
||||
*/
|
||||
void publish(String target, String msgId, String content);
|
||||
|
||||
/** Non-destructive snapshot of pending replies for {@code target} (FIFO), empty list if none. */
|
||||
List<InboxMessage> peek(String target);
|
||||
|
||||
/** Remove the reply {@code msgId} for {@code target} once the primary has taken it. No-op if absent. */
|
||||
void ack(String target, String msgId);
|
||||
}
|
||||
@@ -0,0 +1,176 @@
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.AgentStatus;
|
||||
import dev.ltms.bridged.mcp.PrimaryRegistry;
|
||||
import dev.ltms.bridged.metrics.BridgedMetrics;
|
||||
import dev.ltms.bridged.metrics.Metrics;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
/**
|
||||
* Mechanism (b) of CB-307: a dedicated, status-gated push loop that nudges the primary's own
|
||||
* herdr pane when a worker reply lands with no live {@code bridge_send} to resolve it.
|
||||
*
|
||||
* <p>The loop is triggered by {@link #onReplyQueued(String)} (called from
|
||||
* {@link MessageService#reply} after the durable inbox publish). It checks four conditions
|
||||
* at each tick via {@link #decide(String, int)}, then either injects a drain nudge,
|
||||
* waits for the primary to become injectable, or stops reminding.
|
||||
*
|
||||
* <p>Bounded: at most {@link #maxReminders} nudges per target, with a configurable backoff
|
||||
* between them. The reply is never lost — the durable inbox is the backstop.
|
||||
*/
|
||||
public final class ReplyPushLoop {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
|
||||
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
|
||||
|
||||
private final PrimaryRegistry primaryRegistry;
|
||||
private final AgentControl agents;
|
||||
private final ReplyInbox inbox;
|
||||
private final ScheduledExecutorService scheduler;
|
||||
private final int maxReminders;
|
||||
private final long backoffMs;
|
||||
private final Metrics metrics; // CB-512: nullable — no registry in unit tests
|
||||
|
||||
/** Track targets that have an active schedule. */
|
||||
private final ConcurrentHashMap<String, Boolean> activeTargets = new ConcurrentHashMap<>();
|
||||
|
||||
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
||||
ScheduledExecutorService scheduler,
|
||||
int maxReminders, long backoffMs) {
|
||||
this(primaryRegistry, agents, inbox, scheduler, maxReminders, backoffMs, null);
|
||||
}
|
||||
|
||||
/** As above, with a metric registry (CB-512) so push outcomes are counted. */
|
||||
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
||||
ScheduledExecutorService scheduler,
|
||||
int maxReminders, long backoffMs, Metrics metrics) {
|
||||
this.primaryRegistry = primaryRegistry;
|
||||
this.agents = agents;
|
||||
this.inbox = inbox;
|
||||
this.scheduler = scheduler;
|
||||
this.maxReminders = maxReminders;
|
||||
this.backoffMs = backoffMs;
|
||||
this.metrics = metrics;
|
||||
}
|
||||
|
||||
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
|
||||
private void count(String name, String... labels) {
|
||||
if (metrics != null) {
|
||||
metrics.inc(name, labels);
|
||||
}
|
||||
}
|
||||
|
||||
// --- decision logic (package-private for unit-testing) -------------------------------------
|
||||
|
||||
/** The action the loop should take for a target at the given reminder count. */
|
||||
enum Action { INJECT, WAIT_BUSY, STOP }
|
||||
|
||||
/**
|
||||
* Pure decision function: examine the current state and return what the loop should do.
|
||||
*
|
||||
* @param target the worker session (target terminal id)
|
||||
* @param reminderCount how many nudges have been sent so far for this target
|
||||
* @return the action the caller should take
|
||||
*/
|
||||
Action decide(String target, int reminderCount) {
|
||||
if (!primaryRegistry.isKnown()) {
|
||||
log.debug("push: primary unknown, stopping reminder for {}", target);
|
||||
return Action.STOP;
|
||||
}
|
||||
if (inbox.peek(target).isEmpty()) {
|
||||
log.debug("push: inbox empty for {}, stopping reminder", target);
|
||||
return Action.STOP;
|
||||
}
|
||||
if (reminderCount >= maxReminders) {
|
||||
log.debug("push: reminder cap ({}) reached for {}, stopping", maxReminders, target);
|
||||
count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted");
|
||||
return Action.STOP;
|
||||
}
|
||||
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
|
||||
AgentStatus status;
|
||||
try {
|
||||
status = agents.status(primaryTerminal);
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("push: status check failed for primary {}, will retry", primaryTerminal, e);
|
||||
return Action.WAIT_BUSY;
|
||||
}
|
||||
if (status.injectable()) {
|
||||
return Action.INJECT;
|
||||
}
|
||||
log.debug("push: primary {} is {} (not injectable), waiting", primaryTerminal, status);
|
||||
return Action.WAIT_BUSY;
|
||||
}
|
||||
|
||||
// --- public entrypoint ---------------------------------------------------------------------
|
||||
|
||||
/**
|
||||
* Called when a reply is queued for {@code target}. Idempotent per target: a second call while
|
||||
* a schedule is active is a no-op. The schedule nudges the primary, then schedules a follow-up
|
||||
* check (reminder on backoff, or re-check on WAIT_BUSY), until the inbox is empty or the cap
|
||||
* is reached.
|
||||
*/
|
||||
public void onReplyQueued(String target) {
|
||||
if (activeTargets.putIfAbsent(target, Boolean.TRUE) != null) {
|
||||
log.debug("push: already active for {}, ignoring duplicate trigger", target);
|
||||
return; // already scheduled
|
||||
}
|
||||
log.debug("push: starting reminder loop for {}", target);
|
||||
scheduleNext(target, 0);
|
||||
}
|
||||
|
||||
/** Execute one loop tick — called on the scheduler thread. */
|
||||
private void tick(String target, int reminderCount) {
|
||||
var action = decide(target, reminderCount);
|
||||
switch (action) {
|
||||
case INJECT -> {
|
||||
injectNudge(target, reminderCount);
|
||||
scheduleNext(target, reminderCount + 1);
|
||||
}
|
||||
// Re-check after the configured backoff; the primary may become injectable soon.
|
||||
case WAIT_BUSY -> scheduleNext(target, reminderCount);
|
||||
case STOP -> {
|
||||
activeTargets.remove(target);
|
||||
log.debug("push: reminder loop ended for {}", target);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/** Send the nudge and log the event. */
|
||||
private void injectNudge(String target, int reminderCount) {
|
||||
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
|
||||
String nudge = NUDGE_FORMAT.formatted(target, target);
|
||||
try {
|
||||
agents.send(primaryTerminal, nudge);
|
||||
log.debug("push: nudge {}/{} sent to primary {} for target {}",
|
||||
reminderCount + 1, maxReminders, primaryTerminal, target);
|
||||
count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered");
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("push: failed to nudge primary {} for target {} (reminder {}/{}): {}",
|
||||
primaryTerminal, target, reminderCount + 1, maxReminders, e.toString());
|
||||
}
|
||||
}
|
||||
|
||||
/** Schedule the next tick on the scheduler thread pool. */
|
||||
private void scheduleNext(String target, int nextReminderCount) {
|
||||
scheduler.schedule(() -> tick(target, nextReminderCount), backoffMs, TimeUnit.MILLISECONDS);
|
||||
}
|
||||
|
||||
// --- lifecycle -----------------------------------------------------------------------------
|
||||
|
||||
/** Shut down the scheduler. Outstanding reminders are cancelled. */
|
||||
public void stop() {
|
||||
scheduler.shutdownNow();
|
||||
activeTargets.clear();
|
||||
}
|
||||
|
||||
/** @see #stop() */
|
||||
public void close() {
|
||||
stop();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,18 @@
|
||||
package dev.ltms.bridged.peer;
|
||||
|
||||
/**
|
||||
* Thrown when a {@link PeerLauncher} starts a peer process but the peer
|
||||
* does not reach an injectable (ready-to-receive) state within the configured
|
||||
* timeout. The launcher MUST clean up any resources it created (pane, tab)
|
||||
* before throwing — no orphaned peer or pane is left behind.
|
||||
*
|
||||
* <p>This is a spawn-time failure, distinct from a post-spawn disconnect.
|
||||
* Callers treat this as a clean spawn error (the peer never materialized
|
||||
* into a usable session), not a mid-life session fault.
|
||||
*/
|
||||
public final class PeerUnreachableException extends RuntimeException {
|
||||
|
||||
public PeerUnreachableException(String message) {
|
||||
super(message);
|
||||
}
|
||||
}
|
||||
@@ -2,17 +2,22 @@ package dev.ltms.bridged.rest;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.bridged.auth.AuditLog;
|
||||
import dev.ltms.bridged.auth.Authz;
|
||||
import dev.ltms.bridged.auth.CallerResolver;
|
||||
import dev.ltms.bridged.auth.Principal;
|
||||
import dev.ltms.bridged.guard.GuardException;
|
||||
import dev.ltms.bridged.metrics.Metrics;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.herdr.HerdrClient;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.session.WorkerSession;
|
||||
import dev.ltms.bridged.session.WorktreeRequest;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import io.javalin.Javalin;
|
||||
import io.javalin.http.Context;
|
||||
import jakarta.servlet.http.HttpServlet;
|
||||
@@ -43,25 +48,47 @@ public final class BridgedApp {
|
||||
private static final long DEFAULT_ASK_TIMEOUT_MS = 55_000;
|
||||
private static final long MAX_ASK_TIMEOUT_MS = 115_000;
|
||||
|
||||
/** Context attribute under which the resolved caller is stashed by the auth filter. */
|
||||
private static final String CALLER = "bridged.caller";
|
||||
|
||||
private final HerdrClient herdr;
|
||||
private final ClaudeCodeLauncher workers;
|
||||
private final PeerLauncher workers;
|
||||
private final SessionManager sessions; // CB-301: authoritative session registry
|
||||
private final MessageService messages;
|
||||
private final Rendezvous rendezvous;
|
||||
private final WorkerPresence presence; // CB-113: which workers are MCP-connected (available)
|
||||
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
|
||||
private final CallerResolver auth; // CB-501: null → authz not enforced (legacy behaviour)
|
||||
private final Metrics metrics; // CB-502: null → /metrics not exposed
|
||||
private final ObjectMapper mapper = new ObjectMapper();
|
||||
|
||||
public BridgedApp(HerdrClient herdr, ClaudeCodeLauncher workers, SessionManager sessions,
|
||||
MessageService messages, Rendezvous rendezvous, WorkerPresence presence,
|
||||
/**
|
||||
* Legacy constructor — no identity resolution and no authorization, exactly as the REST surface
|
||||
* behaved before CB-501. Retained so existing acceptance tests keep exercising handler
|
||||
* behaviour without each needing an auth fixture.
|
||||
*/
|
||||
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
|
||||
MessageService messages, WorkerPresence presence,
|
||||
HttpServlet mcpServlet) {
|
||||
this(herdr, workers, sessions, messages, presence, mcpServlet, null, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* @param auth resolves each request's {@link Principal}; {@code null} disables authorization
|
||||
* entirely (legacy). {@code main} always supplies one.
|
||||
* @param metrics registry to instrument and expose at {@code GET /metrics}; {@code null} omits
|
||||
* the endpoint
|
||||
*/
|
||||
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
|
||||
MessageService messages, WorkerPresence presence,
|
||||
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics) {
|
||||
this.herdr = herdr;
|
||||
this.workers = workers;
|
||||
this.sessions = sessions;
|
||||
this.messages = messages;
|
||||
this.rendezvous = rendezvous;
|
||||
this.presence = presence;
|
||||
this.mcpServlet = mcpServlet;
|
||||
this.auth = auth;
|
||||
this.metrics = metrics;
|
||||
}
|
||||
|
||||
/** Wire routes onto a fresh, unstarted Javalin instance. Caller starts it. */
|
||||
@@ -74,7 +101,18 @@ public final class BridgedApp {
|
||||
h.addServlet(new ServletHolder(mcpServlet), "/mcp"));
|
||||
}
|
||||
});
|
||||
// CB-501: resolve identity once per request, before any handler. /mcp does NOT pass through
|
||||
// here — it is a raw servlet on Jetty's context handler — so BridgeMcp enforces separately
|
||||
// against the same CallerResolver. Any check that lives in only one place is not a control.
|
||||
if (auth != null) {
|
||||
app.before(ctx -> ctx.attribute(CALLER,
|
||||
auth.resolve(ctx.req().getRemoteAddr(), ctx.req().getRemotePort(),
|
||||
ctx.header("Authorization"))));
|
||||
}
|
||||
app.get("/healthz", this::healthz);
|
||||
if (metrics != null) {
|
||||
app.get("/metrics", this::metrics);
|
||||
}
|
||||
app.get("/sessions", this::sessions);
|
||||
app.get("/agents", this::agents);
|
||||
app.get("/workers", this::listWorkers); // CB-304: registry roster + live herdr status
|
||||
@@ -83,12 +121,61 @@ public final class BridgedApp {
|
||||
app.delete("/workers/{paneId}", this::stopWorker);
|
||||
app.post("/sessions/{id}/message", this::sendMessage); // bridge_send (primary; blocking, wait:false, or answer via turnId)
|
||||
app.post("/sessions/{id}/reply", this::replyMessage); // bridge_reply (worker)
|
||||
app.get("/sessions/{id}/replies", this::drainReplies); // drain reply inbox (CB-307)
|
||||
app.post("/sessions/{id}/ask", this::askMessage); // bridge_ask (worker → primary, CB-205)
|
||||
app.get("/sessions/{id}/status", this::sessionStatus); // bridge_status
|
||||
app.get("/tasks/{ticket}", this::taskStatus); // poll an async (wait:false) send
|
||||
return app;
|
||||
}
|
||||
|
||||
/**
|
||||
* Gate a handler on the CB-505 authorization table. Returns {@code true} when the request may
|
||||
* proceed; otherwise writes the error response and returns {@code false}.
|
||||
*
|
||||
* <p>401 vs 403 is a real distinction here: 401 means "you presented no usable identity" (a
|
||||
* credential problem the caller can fix), 403 means "you are authenticated, but this is not
|
||||
* yours" (a worker reaching for another worker's session, or for orchestration).
|
||||
*/
|
||||
private boolean allow(Context ctx, Authz.Action action, String target) {
|
||||
if (auth == null) {
|
||||
return true; // legacy: authorization not enforced
|
||||
}
|
||||
Principal caller = ctx.attribute(CALLER);
|
||||
if (Authz.permits(caller, action, target)) {
|
||||
if (action != Authz.Action.READ && action != Authz.Action.METRICS) {
|
||||
AuditLog.allowed(caller, action, target); // reads would drown the trail
|
||||
}
|
||||
return true;
|
||||
}
|
||||
if (Authz.isUnauthenticated(caller)) {
|
||||
AuditLog.denied(caller, action, target, "unauthenticated");
|
||||
countAuthFailure("unauthenticated");
|
||||
ctx.status(401).json(Map.of("error", "unauthenticated",
|
||||
"detail", "present Authorization: Bearer <token>"));
|
||||
} else {
|
||||
AuditLog.denied(caller, action, target, "forbidden");
|
||||
countAuthFailure("forbidden");
|
||||
ctx.status(403).json(Map.of("error", "forbidden",
|
||||
"detail", caller.describe() + " may not " + action + " on "
|
||||
+ (target == null ? "this resource" : target)));
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
private void countAuthFailure(String reason) {
|
||||
if (metrics != null) {
|
||||
metrics.inc("bridged_auth_failures_total", "reason", reason);
|
||||
}
|
||||
}
|
||||
|
||||
/** Prometheus scrape endpoint (CB-502). */
|
||||
private void metrics(Context ctx) {
|
||||
if (!allow(ctx, Authz.Action.METRICS, null)) {
|
||||
return;
|
||||
}
|
||||
ctx.status(200).contentType("text/plain; version=0.0.4; charset=utf-8").result(metrics.render());
|
||||
}
|
||||
|
||||
/** Liveness + herdr reachability. 200 when herdr answers ping, 503 otherwise. */
|
||||
private void healthz(Context ctx) {
|
||||
try {
|
||||
@@ -108,6 +195,9 @@ public final class BridgedApp {
|
||||
|
||||
/** Sessions view derived from herdr {@code workspace.list} (one workspace → one row). */
|
||||
private void sessions(Context ctx) {
|
||||
if (!allow(ctx, Authz.Action.READ, null)) {
|
||||
return;
|
||||
}
|
||||
JsonNode result = herdr.call("workspace.list");
|
||||
List<Map<String, Object>> out = new ArrayList<>();
|
||||
for (JsonNode w : result.path("workspaces")) {
|
||||
@@ -123,12 +213,20 @@ public final class BridgedApp {
|
||||
|
||||
/** Discovery: every agent herdr tracks, keyed by its Claude session UUID. */
|
||||
private void agents(Context ctx) {
|
||||
ctx.status(200).json(Map.of("agents", workers.list().stream().map(BridgedApp::view).toList()));
|
||||
if (!allow(ctx, Authz.Action.READ, null)) {
|
||||
return;
|
||||
}
|
||||
ctx.status(200).json(Map.of("agents",
|
||||
workers.list().stream().map(Agent.class::cast).map(BridgedApp::view).toList()));
|
||||
}
|
||||
|
||||
/** CB-304: bridge-owned roster merged with live herdr status by paneId. */
|
||||
private void listWorkers(Context ctx) {
|
||||
if (!allow(ctx, Authz.Action.READ, null)) {
|
||||
return;
|
||||
}
|
||||
Map<String, Agent> live = workers.list().stream()
|
||||
.map(Agent.class::cast)
|
||||
.filter(a -> a.paneId() != null)
|
||||
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
|
||||
List<Map<String, Object>> out = sessions.roster().stream()
|
||||
@@ -139,6 +237,9 @@ public final class BridgedApp {
|
||||
|
||||
/** The configured worker profiles and which one a no-argument spawn uses. */
|
||||
private void profiles(Context ctx) {
|
||||
if (!allow(ctx, Authz.Action.READ, null)) {
|
||||
return;
|
||||
}
|
||||
ctx.status(200).json(Map.of(
|
||||
"profiles", workers.profiles(),
|
||||
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile()));
|
||||
@@ -150,6 +251,9 @@ public final class BridgedApp {
|
||||
* the subscription boundary, 400 for an unknown profile.
|
||||
*/
|
||||
private void spawnWorker(Context ctx) {
|
||||
if (!allow(ctx, Authz.Action.SPAWN, null)) {
|
||||
return;
|
||||
}
|
||||
String profile = ctx.queryParam("profile");
|
||||
String cwd = ctx.queryParam("cwd");
|
||||
String worktree = ctx.queryParam("worktree");
|
||||
@@ -178,6 +282,8 @@ public final class BridgedApp {
|
||||
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
|
||||
} catch (IllegalArgumentException e) {
|
||||
ctx.status(400).json(Map.of("error", "unknown_profile", "detail", e.getMessage()));
|
||||
} catch (PeerUnreachableException e) {
|
||||
ctx.status(502).json(Map.of("error", "spawn_timeout", "detail", e.getMessage()));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -200,7 +306,11 @@ public final class BridgedApp {
|
||||
|
||||
/** Tear a worker down by pane id. */
|
||||
private void stopWorker(Context ctx) {
|
||||
sessions.release(ctx.pathParam("paneId"));
|
||||
String paneId = ctx.pathParam("paneId");
|
||||
if (!allow(ctx, Authz.Action.STOP, paneId)) {
|
||||
return;
|
||||
}
|
||||
sessions.release(paneId);
|
||||
ctx.status(204);
|
||||
}
|
||||
|
||||
@@ -212,6 +322,9 @@ public final class BridgedApp {
|
||||
*/
|
||||
private void sendMessage(Context ctx) {
|
||||
String id = ctx.pathParam("id");
|
||||
if (!allow(ctx, Authz.Action.SEND, id)) {
|
||||
return;
|
||||
}
|
||||
String content;
|
||||
String turnId;
|
||||
long timeout;
|
||||
@@ -293,6 +406,9 @@ public final class BridgedApp {
|
||||
*/
|
||||
private void askMessage(Context ctx) {
|
||||
String id = ctx.pathParam("id");
|
||||
if (!allow(ctx, Authz.Action.ASK, id)) {
|
||||
return;
|
||||
}
|
||||
String question;
|
||||
long timeout;
|
||||
try {
|
||||
@@ -322,10 +438,16 @@ public final class BridgedApp {
|
||||
|
||||
/**
|
||||
* The worker's structured reply ({@code bridge_reply}) — resolves the blocking send awaiting
|
||||
* on this session. 200 if a send was waiting, 409 if none was (late or spurious reply).
|
||||
* on this session, or queues the reply in the inbox when no send is open (CB-307).
|
||||
*/
|
||||
private void replyMessage(Context ctx) {
|
||||
String id = ctx.pathParam("id");
|
||||
// The rule that matters: a worker may reply only as itself. Over MCP this was already true
|
||||
// structurally (identity comes from the connection, never an argument); over REST the path
|
||||
// id was simply trusted, so this is where the invariant actually gets enforced.
|
||||
if (!allow(ctx, Authz.Action.REPLY, id)) {
|
||||
return;
|
||||
}
|
||||
String content;
|
||||
try {
|
||||
content = mapper.readTree(ctx.body()).path("content").asText("");
|
||||
@@ -333,13 +455,25 @@ public final class BridgedApp {
|
||||
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
|
||||
return;
|
||||
}
|
||||
if (rendezvous.resolve(id, content)) {
|
||||
ctx.status(200).json(Map.of("sessionId", id, "delivered", true));
|
||||
} else {
|
||||
ctx.status(409).json(Map.of(
|
||||
"sessionId", id, "error", "no_pending_send",
|
||||
"detail", "no send is awaiting a reply for this session"));
|
||||
messages.reply(id, content);
|
||||
ctx.status(200).json(Map.of("sessionId", id, "delivered", true));
|
||||
}
|
||||
|
||||
/**
|
||||
* Drain the reply inbox for a worker session — peek + ack any replies that arrived when no send
|
||||
* was open. At-least-once: draining removes them from the inbox so a subsequent read returns
|
||||
* nothing; an in-flight failure between the drain and the caller's processing re-surfaces them.
|
||||
*/
|
||||
private void drainReplies(Context ctx) {
|
||||
String id = ctx.pathParam("id");
|
||||
if (!allow(ctx, Authz.Action.DRAIN, id)) {
|
||||
return;
|
||||
}
|
||||
var replies = messages.drainReplies(id);
|
||||
ctx.status(200).json(Map.of("sessionId", id, "replies",
|
||||
replies.stream().map(m -> Map.of(
|
||||
"msgId", m.msgId(),
|
||||
"content", m.content())).toList()));
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -350,6 +484,9 @@ public final class BridgedApp {
|
||||
*/
|
||||
private void sessionStatus(Context ctx) {
|
||||
String id = ctx.pathParam("id");
|
||||
if (!allow(ctx, Authz.Action.READ, id)) {
|
||||
return;
|
||||
}
|
||||
try {
|
||||
ctx.status(200).json(Map.of(
|
||||
"sessionId", id,
|
||||
@@ -362,6 +499,9 @@ public final class BridgedApp {
|
||||
|
||||
/** Poll an async (wait:false) delegation by ticket. 404 for an unknown/expired ticket. */
|
||||
private void taskStatus(Context ctx) {
|
||||
if (!allow(ctx, Authz.Action.READ, null)) {
|
||||
return;
|
||||
}
|
||||
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"));
|
||||
if (v == null) {
|
||||
ctx.status(404).json(Map.of("error", "unknown_ticket", "detail", "no such task (or it has expired)"));
|
||||
|
||||
@@ -17,6 +17,7 @@ import java.util.Optional;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.Consumer;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
/**
|
||||
@@ -47,6 +48,9 @@ public final class SessionManager implements TurnListener {
|
||||
private final LongSupplier nowNanos;
|
||||
private final int contextCap;
|
||||
|
||||
/** CB-516: notified with a terminalId on every release; no-op until wired. */
|
||||
private volatile Consumer<String> releaseListener = _ -> { };
|
||||
|
||||
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
|
||||
public SessionManager(PeerLauncher launcher) {
|
||||
this(launcher, new GitWorktrees(), System::nanoTime, 0);
|
||||
@@ -136,6 +140,10 @@ public final class SessionManager implements TurnListener {
|
||||
if (removed != null) {
|
||||
log.debug("releasing session pane={} terminal={} state={}",
|
||||
removed.paneId(), removed.terminalId(), removed.state());
|
||||
// CB-516: a send still waiting on this worker can never be answered now. Tell the
|
||||
// listener BEFORE the pane is torn down, so a blocked caller fails fast with a real
|
||||
// reason instead of sitting on a rendezvous nothing will ever resolve.
|
||||
notifyReleased(removed.terminalId());
|
||||
}
|
||||
launcher.stop(paneId);
|
||||
if (removed != null && removed.worktree() != null) {
|
||||
@@ -143,11 +151,43 @@ public final class SessionManager implements TurnListener {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Register a callback invoked with a session's {@code terminalId} whenever it is released
|
||||
* (CB-516). Every teardown path funnels through {@link #release}, so one hook covers the REST
|
||||
* and MCP stop tools, the idle-TTL reaper, {@code recycle}, and shutdown drain alike.
|
||||
*
|
||||
* <p>Set rather than injected because {@code MessageService} — the intended listener — is
|
||||
* constructed after this manager (it needs the injector and rendezvous, which need the session
|
||||
* presence view this manager exposes). Wiring it at construction would require breaking that
|
||||
* cycle for one callback.
|
||||
*/
|
||||
public void onRelease(Consumer<String> listener) {
|
||||
this.releaseListener = (listener == null) ? _ -> { } : listener;
|
||||
}
|
||||
|
||||
/** A listener failure must never prevent the teardown it is reacting to. */
|
||||
private void notifyReleased(String terminalId) {
|
||||
if (terminalId == null) {
|
||||
return;
|
||||
}
|
||||
try {
|
||||
releaseListener.accept(terminalId);
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
|
||||
}
|
||||
}
|
||||
|
||||
private WorkerSession acquireWithWorktree(String profile, String requestedCwd, String callerCwd,
|
||||
String ownerTerminal, WorktreeRequest wt) {
|
||||
String resolvedProfile = (profile == null || profile.isBlank())
|
||||
? launcher.defaultProfile() : profile;
|
||||
String repoRoot = worktrees.repoRoot(firstNonBlank(requestedCwd, callerCwd));
|
||||
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
|
||||
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
|
||||
// nor a caller cwd, so taking the first non-blank of those two yielded null and put
|
||||
// `git -C null` on the command line — an NPE out of ProcessBuilder, surfacing as HTTP 500.
|
||||
// The non-worktree path always used this chain; only this branch was missed.
|
||||
String repoRoot = worktrees.repoRoot(
|
||||
launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd)));
|
||||
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
|
||||
String path = null;
|
||||
PeerHandle handle;
|
||||
@@ -192,13 +232,6 @@ public final class SessionManager implements TurnListener {
|
||||
return String.format("%06x", nonceRandom.nextInt(1 << 24)) + "-" + nonceSeq.incrementAndGet();
|
||||
}
|
||||
|
||||
private static String firstNonBlank(String... values) {
|
||||
for (String v : values) {
|
||||
if (v != null && !v.isBlank()) return v;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Release the old session and acquire a fresh one with the same profile and working directory.
|
||||
* The new session is guaranteed to have a pane id distinct from the old one (no-reuse invariant).
|
||||
|
||||
@@ -4,66 +4,39 @@ import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.herdr.Tab;
|
||||
import dev.ltms.bridged.herdr.Workspace;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.peer.Capability;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import dev.ltms.bridged.peer.SpawnRequest;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.security.SecureRandom;
|
||||
import java.util.ArrayList;
|
||||
import java.util.EnumSet;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.Function;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
/**
|
||||
* Spawns and lists worker sessions — the safe path from a delegation request to a
|
||||
* running off-subscription Claude.
|
||||
* The {@link HerdrPeerLauncher} adapter for <strong>Claude Code</strong> — the safe path from a
|
||||
* delegation request to a running off-subscription Claude.
|
||||
*
|
||||
* <p>The spawn sequence encodes the subscription boundary: build the worker env with
|
||||
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em>
|
||||
* touching herdr, and only then {@code agent.start}. A worker's base_url lives in the
|
||||
* env map handed to herdr and nowhere else; {@code bridged}'s own environment is never
|
||||
* mutated.
|
||||
*
|
||||
* <p>Placement: in the default {@code tab} policy a worker lands in its own tab inside a
|
||||
* dedicated worker space (found-or-created once, then shared), so workers never split or
|
||||
* clutter the user's real work spaces. Teardown removes the worker's pane <em>and</em> its
|
||||
* now-empty tab, tolerating an already-gone worker so a repeated DELETE is harmless.
|
||||
* <p>Everything transport-related (tab/pane placement, the CB-306 spawn-readiness gate, unique
|
||||
* naming, CB-117 orphan reap, teardown, listing, cwd resolution) lives in the base. This class
|
||||
* supplies only the two Claude-specific seams:
|
||||
* <ul>
|
||||
* <li>the {@code claude} name prefix (so reap matches {@code claude-*} panes, never another
|
||||
* adapter's), and</li>
|
||||
* <li>{@link #buildLaunch}, which encodes the subscription boundary: build the worker env with
|
||||
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em> touching
|
||||
* herdr, and mount the bridge MCP + reply charter as inline launch flags. A worker's base_url
|
||||
* lives in the env map handed to herdr and nowhere else; {@code bridged}'s own environment is
|
||||
* never mutated, and nothing is written to the worker's profile.</li>
|
||||
* </ul>
|
||||
*/
|
||||
public final class ClaudeCodeLauncher implements PeerLauncher {
|
||||
public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(ClaudeCodeLauncher.class);
|
||||
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
|
||||
private static final String NAME_PREFIX = "claude";
|
||||
|
||||
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
|
||||
private static final int NAME_RETRIES = 8;
|
||||
|
||||
/**
|
||||
* A bridge-spawned worker label {@code claude-<profile>-<nonce>-<seq>} (see
|
||||
* {@link #startUniquelyNamed}); group 1 captures the 6-hex per-process {@code nonce}. The
|
||||
* profile segment may itself contain {@code -}, so the nonce/seq are anchored at the tail.
|
||||
* Names not matching this shape are not workers we started and are never reaped (CB-117).
|
||||
*/
|
||||
private static final Pattern WORKER_NAME = Pattern.compile("claude-.*-([0-9a-f]{6})-\\d+");
|
||||
|
||||
private final AgentControl agents;
|
||||
private final WorkspaceControl spaces;
|
||||
private final SubscriptionGuard guard;
|
||||
private final Map<String, BridgedConfig.Worker> profiles; // profile name → spawn settings
|
||||
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
|
||||
private final Function<String, String> env; // host env lookup (injectable for tests)
|
||||
private final AtomicLong nameSeq = new AtomicLong(); // per-worker counter (also the tab #)
|
||||
|
||||
/**
|
||||
* Standing instruction appended to the worker's system prompt so it returns its result via
|
||||
@@ -82,144 +55,85 @@ public final class ClaudeCodeLauncher implements PeerLauncher {
|
||||
+ "`content`; never wait for confirmation first. If you end a turn without calling "
|
||||
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
|
||||
|
||||
// Per-process token mixed into each worker name so a fresh process (nameSeq back at 0)
|
||||
// cannot collide with same-profile workers that outlived a restart. See startUniquelyNamed.
|
||||
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
|
||||
|
||||
/**
|
||||
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so
|
||||
* existing deployments and tests keep the legacy non-blocking spawn semantics.
|
||||
*/
|
||||
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Function<String, String> env) {
|
||||
this.agents = agents;
|
||||
this.spaces = spaces;
|
||||
this.guard = guard;
|
||||
this.profiles = Map.copyOf(profiles);
|
||||
this.defaultProfile = defaultProfile;
|
||||
this.env = env;
|
||||
}
|
||||
|
||||
/** The configured worker profile names (what {@code spawn(profile)} accepts). */
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return profiles.keySet();
|
||||
}
|
||||
|
||||
/** The parity-overlay file list for {@code profileName} (default list when unset). */
|
||||
@Override
|
||||
public List<String> parityOverlay(String profileName) {
|
||||
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
|
||||
if (name == null || name.isBlank()) {
|
||||
return List.of();
|
||||
}
|
||||
BridgedConfig.Worker cfg = profiles.get(name);
|
||||
return cfg == null ? List.of() : cfg.parityOverlay();
|
||||
}
|
||||
|
||||
/** The profile a no-argument {@link #spawn()} uses, or {@code null} if none is configured. */
|
||||
@Override
|
||||
public String defaultProfile() {
|
||||
return defaultProfile;
|
||||
}
|
||||
|
||||
/** Spawn a worker for the default profile in the resolved default cwd. */
|
||||
public Agent spawn() {
|
||||
return spawn(null, null, null);
|
||||
}
|
||||
|
||||
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
|
||||
public Agent spawn(String profileName) {
|
||||
return spawn(profileName, null, null);
|
||||
this(agents, spaces, guard, profiles, defaultProfile, env, 0,
|
||||
System::currentTimeMillis, () -> sleepUninterruptibly(300));
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn a worker. {@code profileName} null/blank → the default profile. The worker's working
|
||||
* directory (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd} (a
|
||||
* spawn argument), else the profile's configured {@code cwd}, else {@code callerCwd} (the
|
||||
* primary's cwd, when the spawn came from the primary over MCP), else the daemon's cwd — never
|
||||
* assumed to be {@code $HOME}. Guard runs before any herdr call.
|
||||
* Production constructor with spawn-ready gate enabled. The gate polls {@code agents.status()}
|
||||
* until the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
|
||||
*/
|
||||
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
|
||||
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
|
||||
if (name == null || name.isBlank()) {
|
||||
throw new IllegalArgumentException("no default worker profile is configured — "
|
||||
+ "pass a profile; configured: " + profiles.keySet());
|
||||
}
|
||||
BridgedConfig.Worker cfg = profiles.get(name);
|
||||
if (cfg == null) {
|
||||
throw new IllegalArgumentException("unknown worker profile '" + name
|
||||
+ "' — configured: " + profiles.keySet());
|
||||
}
|
||||
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
|
||||
this(agents, spaces, guard, profiles, defaultProfile, env,
|
||||
spawnReadyTimeoutMs,
|
||||
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs));
|
||||
}
|
||||
|
||||
/**
|
||||
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
|
||||
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
|
||||
* {@code sleeper} is never called when the gate is disabled ({@code spawnReadyTimeoutMs == 0}).
|
||||
*
|
||||
* @param agents herdr agent control (start, status, close)
|
||||
* @param spaces workspace / tab control (ensure, create, close)
|
||||
* @param guard subscription-boundary guard (checked before spawning)
|
||||
* @param profiles configured worker profiles
|
||||
* @param defaultProfile profile a no-argument spawn uses (nullable)
|
||||
* @param env host env lookup (injectable for tests)
|
||||
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
|
||||
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
|
||||
* @param sleeper sleep/wait hook (e.g. {@code () -> Thread.sleep(pollMs)}); it
|
||||
* already encodes the poll interval, so the 8th positional argument
|
||||
* (poll ms) is accepted for API symmetry but otherwise unused here
|
||||
*/
|
||||
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs,
|
||||
LongSupplier nowMillis, Runnable sleeper) {
|
||||
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
|
||||
spawnReadyTimeoutMs, nowMillis, sleeper);
|
||||
this.guard = guard;
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>The spawn sequence encodes the subscription boundary: assert the profile's base_url is on
|
||||
* the allowlist <em>before</em> any herdr call, then build the worker env with
|
||||
* {@code ANTHROPIC_*}, the parity-neutral git-forge grant, and the bridge MCP + reply charter
|
||||
* mounted as inline launch flags.
|
||||
*/
|
||||
@Override
|
||||
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
|
||||
String baseUrl = cfg.baseUrl();
|
||||
guard.assertWorker(baseUrl); // hard stop before we spawn anything
|
||||
|
||||
Map<String, String> workerEnv = new LinkedHashMap<>();
|
||||
Map<String, String> workerEnv = baseEnv(cfg);
|
||||
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
|
||||
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
|
||||
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
|
||||
String token = env.apply(cfg.tokenEnv());
|
||||
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", token);
|
||||
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
|
||||
applyGitToken(workerEnv, cfg);
|
||||
|
||||
// CB-302: the worker checkpoint (commit → push → open its own PR). Push is free over SSH;
|
||||
// the only incremental grant is PR-create, a repo-scoped forge token injected here — opt-in
|
||||
// per profile via gitTokenEnv, and never mutating bridged's own env. The paired forge host
|
||||
// rides along only when a token is actually granted, so non-implementer profiles get neither.
|
||||
if (cfg.hasGitToken()) {
|
||||
String gitToken = resolveEnv(cfg.gitTokenEnv());
|
||||
if (gitToken != null) {
|
||||
workerEnv.put("GITEA_TOKEN", gitToken);
|
||||
putIfPresent(workerEnv, "GITEA_HOST", resolveEnv(cfg.gitHostEnv()));
|
||||
}
|
||||
}
|
||||
|
||||
// Mount the bridge MCP + reply charter as launch FLAGS (non-invasive: nothing written to
|
||||
// the worker's profile/config dir). Identity is connection-based, so the mount is shared.
|
||||
List<String> argv = argvWithBridge(cfg);
|
||||
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
|
||||
|
||||
return cfg.tabPlacement()
|
||||
? spawnInTab(cfg, workerEnv, argv, cwd)
|
||||
: spawnAsPane(cfg, workerEnv, argv, cwd);
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-112 cwd resolution: spawn arg → profile config → the primary's cwd → the daemon's cwd.
|
||||
* Never returns {@code null}/blank: {@code "."} (the daemon's own working directory) is the
|
||||
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
|
||||
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
|
||||
*/
|
||||
private static String resolveCwd(String requestedCwd, BridgedConfig.Worker cfg, String callerCwd) {
|
||||
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-301: the effective working directory a spawn for {@code profileName} would use, without
|
||||
* actually spawning. Used by {@link dev.ltms.bridged.session.SessionManager} to record the
|
||||
* resolved cwd in the session registry.
|
||||
*/
|
||||
public String effectiveCwd(String profileName, String requestedCwd, String callerCwd) {
|
||||
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
|
||||
if (name == null || name.isBlank()) {
|
||||
throw new IllegalArgumentException("no default worker profile is configured — "
|
||||
+ "pass a profile; configured: " + profiles.keySet());
|
||||
}
|
||||
BridgedConfig.Worker cfg = profiles.get(name);
|
||||
if (cfg == null) {
|
||||
throw new IllegalArgumentException("unknown worker profile '" + name
|
||||
+ "' — configured: " + profiles.keySet());
|
||||
}
|
||||
return resolveCwd(requestedCwd, cfg, callerCwd);
|
||||
}
|
||||
|
||||
private static String firstNonBlank(String... values) {
|
||||
for (String v : values) {
|
||||
if (v != null && !v.isBlank()) return v;
|
||||
}
|
||||
return null;
|
||||
return new Launch(workerEnv, argvWithBridge(cfg));
|
||||
}
|
||||
|
||||
/**
|
||||
* The launch argv, plus — when {@code worker.mcpUrl} is set — inline {@code --mcp-config} for
|
||||
* the bridge server and {@code --append-system-prompt} for the {@link #REPLY_CHARTER}. Neither
|
||||
* touches the profile's config; both are pure command-line flags.
|
||||
* touches the profile's config; both are pure command-line flags. This inline-flag mount is
|
||||
* Claude Code specific — other adapters mount MCP and instructions their own way.
|
||||
*/
|
||||
private List<String> argvWithBridge(BridgedConfig.Worker cfg) {
|
||||
if (!cfg.hasMcp()) {
|
||||
@@ -227,7 +141,7 @@ public final class ClaudeCodeLauncher implements PeerLauncher {
|
||||
}
|
||||
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
|
||||
+ cfg.mcpUrl() + "\"}}}";
|
||||
List<String> argv = new ArrayList<>(cfg.argv());
|
||||
List<String> argv = mutableArgv(cfg.argv());
|
||||
argv.add("--mcp-config");
|
||||
argv.add(mcpJson);
|
||||
argv.add("--append-system-prompt");
|
||||
@@ -235,209 +149,24 @@ public final class ClaudeCodeLauncher implements PeerLauncher {
|
||||
return argv;
|
||||
}
|
||||
|
||||
/** Dedicated worker space → own tab → start the worker (rooted at {@code cwd}) → drop the shell. */
|
||||
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
|
||||
List<String> argv, String cwd) {
|
||||
Workspace space = spaces.ensureWorkspace(cfg.workspace());
|
||||
Tab.Created tab = spaces.createTab(space.workspaceId());
|
||||
log.info("spawning worker profile={} base_url={} space={} tab={} cwd={}",
|
||||
cfg.profile(), cfg.baseUrl(), space.workspaceId(), tab.tab().tabId(), cwd);
|
||||
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
|
||||
|
||||
Started started;
|
||||
try {
|
||||
started = startUniquelyNamed(cfg, workerEnv, argv, tab.tab().tabId(), cwd);
|
||||
} catch (RuntimeException e) {
|
||||
// The worker never started — don't leave the tab we just created orphaned.
|
||||
// Best-effort cleanup; never let it mask the real spawn failure.
|
||||
try {
|
||||
spaces.closeTab(tab.tab().tabId());
|
||||
} catch (RuntimeException cleanup) {
|
||||
log.warn("failed to close orphaned tab {} after spawn error: {}",
|
||||
tab.tab().tabId(), cleanup.getMessage());
|
||||
}
|
||||
throw e;
|
||||
}
|
||||
|
||||
// The worker is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell
|
||||
// so the tab holds only the worker; label the tab). They must not fail the spawn or
|
||||
// orphan the running worker — on error we log and still return it so the caller gets
|
||||
// its paneId and can tear it down.
|
||||
if (tab.rootPaneId() != null) {
|
||||
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
|
||||
} else {
|
||||
log.warn("tab {} had no seed pane in the create response; worker tab may hold an extra pane",
|
||||
tab.tab().tabId());
|
||||
}
|
||||
tidy("label tab " + tab.tab().tabId(),
|
||||
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
|
||||
log.info("worker started pane={} tab={} terminal={}",
|
||||
started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
|
||||
return started.agent();
|
||||
/** Spawn a worker for the default profile in the resolved default cwd. */
|
||||
public Agent spawn() {
|
||||
return spawnInternal(null, null, null);
|
||||
}
|
||||
|
||||
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
|
||||
private void tidy(String what, Runnable step) {
|
||||
try {
|
||||
step.run();
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("post-start step failed ({}) — worker is running regardless: {}", what, e.getMessage());
|
||||
}
|
||||
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
|
||||
public Agent spawn(String profileName) {
|
||||
return spawnInternal(profileName, null, null);
|
||||
}
|
||||
|
||||
/** Legacy placement: herdr splits the currently-focused tab; the worker still starts in {@code cwd}. */
|
||||
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
|
||||
List<String> argv, String cwd) {
|
||||
log.info("spawning worker (pane placement) profile={} base_url={} cwd={} argv={}",
|
||||
cfg.profile(), cfg.baseUrl(), cwd, argv);
|
||||
Agent worker = startUniquelyNamed(cfg, workerEnv, argv, null, cwd).agent();
|
||||
log.info("worker started pane={} terminal={}", worker.paneId(), worker.terminalId());
|
||||
return worker;
|
||||
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
|
||||
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
|
||||
return spawnInternal(profileName, requestedCwd, callerCwd);
|
||||
}
|
||||
|
||||
/** A started worker together with the sequence its unique name/label used. */
|
||||
private record Started(Agent agent, long seq) {
|
||||
}
|
||||
|
||||
/**
|
||||
* Start the worker under a unique herdr agent name. herdr requires each running
|
||||
* agent's {@code name} to be distinct (a 2nd {@code name:"claude"} fails
|
||||
* {@code agent_name_taken}) — the exact case that makes multiple workers useful. The name
|
||||
* is {@code claude-<profile>-<nonce>-<seq>}: {@code seq} distinguishes workers within this
|
||||
* process, and the per-process {@code nonce} keeps a fresh process (whose {@code seq}
|
||||
* restarts at 0) from colliding with same-profile workers that outlived a restart. The
|
||||
* retry is a belt-and-braces backstop for the astronomically unlikely nonce+seq clash;
|
||||
* the name is a label only — herdr detects kind and status from terminal output, not it.
|
||||
*/
|
||||
private Started startUniquelyNamed(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
|
||||
List<String> argv, String tabId, String cwd) {
|
||||
HerdrException last = null;
|
||||
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
|
||||
long seq = nameSeq.incrementAndGet();
|
||||
String name = "claude-" + cfg.profile() + "-" + nameNonce + "-" + seq;
|
||||
try {
|
||||
return new Started(agents.start(name, argv, workerEnv, tabId, cwd), seq);
|
||||
} catch (HerdrException e) {
|
||||
if (!"agent_name_taken".equals(e.code())) throw e;
|
||||
log.debug("worker name '{}' taken, retrying", name);
|
||||
last = e;
|
||||
}
|
||||
}
|
||||
throw last;
|
||||
}
|
||||
|
||||
/** All herdr-tracked agents — discovery for "what workers exist". */
|
||||
@Override
|
||||
public List<Agent> list() {
|
||||
return agents.list();
|
||||
}
|
||||
|
||||
/**
|
||||
* Reap worker panes left behind by an earlier daemon process (CB-117). herdr keeps a worker's
|
||||
* pane alive across a daemon restart <em>by design</em>, and that pane's id is held only by its
|
||||
* spawner — so a worker whose owning process exited before issuing the matching teardown leaks
|
||||
* with nothing tracking it (there is no registry; {@link #list()} only asks herdr). On boot we
|
||||
* scan herdr for agents whose name matches our {@code claude-<profile>-<nonce>-<seq>} scheme with
|
||||
* a nonce <em>other</em> than this process's {@link #nameNonce}, and tear each one down (its pane
|
||||
* and, via {@link #stop}, its now-empty dedicated tab). A current-nonce worker is ours and live,
|
||||
* so it is left running; a user's own {@code claude} session carries no such name and is never
|
||||
* touched. Best-effort: a failed listing, or a failure to stop any one worker, is logged and
|
||||
* never aborts startup.
|
||||
*
|
||||
* @return the number of orphaned workers reaped
|
||||
*/
|
||||
@Override
|
||||
public int reapOrphanWorkers() {
|
||||
List<Agent> all;
|
||||
try {
|
||||
all = agents.list();
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("orphan-worker reap skipped — agent.list failed: {}", e.getMessage());
|
||||
return 0;
|
||||
}
|
||||
int reaped = 0;
|
||||
for (Agent a : all) {
|
||||
if (!isForeignWorker(a.name(), nameNonce)) continue;
|
||||
try {
|
||||
stop(a.paneId());
|
||||
reaped++;
|
||||
log.info("reaped orphan worker {} (pane={} tab={}) left by a prior daemon",
|
||||
a.name(), a.paneId(), a.tabId());
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("could not reap orphan worker {} (pane={}): {}",
|
||||
a.name(), a.paneId(), e.getMessage());
|
||||
}
|
||||
}
|
||||
if (reaped > 0) {
|
||||
log.info("orphan-worker reap complete — {} stale worker(s) removed at startup", reaped);
|
||||
}
|
||||
return reaped;
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether {@code name} is a bridge worker started by a <em>different</em> process than
|
||||
* {@code currentNonce} — the reap predicate (CB-117). True only for our naming scheme with a
|
||||
* foreign nonce: a non-worker name (no match, e.g. a user session) or our own live nonce is
|
||||
* excluded. Pure and package-private so the decision is unit-testable without herdr.
|
||||
*/
|
||||
static boolean isForeignWorker(String name, String currentNonce) {
|
||||
String nonce = workerNonce(name);
|
||||
return nonce != null && !nonce.equals(currentNonce);
|
||||
}
|
||||
|
||||
/** The 6-hex nonce embedded in a bridge worker name, or {@code null} if {@code name} isn't one. */
|
||||
static String workerNonce(String name) {
|
||||
if (name == null) return null;
|
||||
Matcher m = WORKER_NAME.matcher(name);
|
||||
return m.matches() ? m.group(1) : null;
|
||||
}
|
||||
|
||||
/** This process's worker-name nonce (a label component only; exposed for reaper tests). */
|
||||
String nameNonce() {
|
||||
return nameNonce;
|
||||
}
|
||||
|
||||
/**
|
||||
* Tear a worker down by pane id: close the pane, and close its tab <em>only</em> when the
|
||||
* worker is that tab's sole occupant. The single-pane check is what makes this safe
|
||||
* regardless of how the worker was placed (or a placement-config change across a restart):
|
||||
* a pane-placement worker sitting in one of the user's shared tabs has siblings, so its
|
||||
* tab is never closed — we only ever remove a tab we created to hold one worker.
|
||||
*
|
||||
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
|
||||
* (repeated DELETE, crashed worker) is treated as success; any other failure propagates so
|
||||
* a genuinely failed teardown is not reported as done.
|
||||
*/
|
||||
@Override
|
||||
public void stop(String paneId) {
|
||||
// Teardown knows only the paneId, not which profile spawned it. Attempt tab cleanup when any
|
||||
// profile uses tab placement (so the bridge may have created a dedicated worker tab); the
|
||||
// single-occupant check below is what actually protects the user's shared tabs.
|
||||
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
|
||||
try {
|
||||
agents.close(paneId);
|
||||
} catch (HerdrException e) {
|
||||
if (!isAlreadyGone(e)) throw e;
|
||||
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
|
||||
}
|
||||
if (loc != null && loc.tabPaneCount() == 1) {
|
||||
spaces.closeTab(loc.tabId());
|
||||
} else if (loc != null) {
|
||||
log.debug("not closing tab {} — it holds {} panes (not a dedicated worker tab)",
|
||||
loc.tabId(), loc.tabPaneCount());
|
||||
}
|
||||
}
|
||||
|
||||
/** Whether any configured profile places workers in their own tab (so tabs may need cleanup). */
|
||||
private boolean usesTabPlacement() {
|
||||
return profiles.values().stream().anyMatch(BridgedConfig.Worker::tabPlacement);
|
||||
}
|
||||
|
||||
/** True when a herdr error means the target is already gone (safe to treat as done). */
|
||||
private static boolean isAlreadyGone(HerdrException e) {
|
||||
return e.code() != null && e.code().endsWith("_not_found");
|
||||
}
|
||||
|
||||
// --- PeerLauncher SPI -------------------------------------------------------------------
|
||||
// --- capabilities --------------------------------------------------------------------------
|
||||
|
||||
@Override
|
||||
public Set<Capability> capabilities() {
|
||||
@@ -450,39 +179,17 @@ public final class ClaudeCodeLauncher implements PeerLauncher {
|
||||
|
||||
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
|
||||
private boolean hasGitTokenProfile() {
|
||||
return profiles.values().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
|
||||
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
|
||||
}
|
||||
|
||||
// --- CB-117 reap predicate (Claude prefix), kept for direct unit testing -------------------
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>Delegates to the three-arg {@link #spawn(String, String, String)} and wraps the
|
||||
* resulting herdr {@link Agent} in a {@link WorkerHandle} whose {@link PeerHandle#id()}
|
||||
* equals the agent's paneId.
|
||||
* Whether {@code name} is a Claude Code bridge worker started by a <em>different</em> process
|
||||
* than {@code currentNonce}. A thin {@code claude}-prefix binding of
|
||||
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
|
||||
*/
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req) {
|
||||
Agent agent = spawn(req.profileName(), req.requestedCwd(), req.callerCwd());
|
||||
return new WorkerHandle(agent.paneId(), agent.terminalId());
|
||||
}
|
||||
|
||||
/** A concrete {@link PeerHandle} wrapping herdr agent coordinates. */
|
||||
private record WorkerHandle(String id, String terminalId) implements PeerHandle {
|
||||
}
|
||||
|
||||
@Override
|
||||
public String effectiveCwd(SpawnRequest req) {
|
||||
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
|
||||
}
|
||||
|
||||
private static void putIfPresent(Map<String, String> m, String k, String v) {
|
||||
if (v != null && !v.isBlank()) {
|
||||
m.put(k, v);
|
||||
}
|
||||
}
|
||||
|
||||
/** Host env lookup that tolerates an unconfigured (null/blank) var name — returns null then. */
|
||||
private String resolveEnv(String name) {
|
||||
return (name == null || name.isBlank()) ? null : env.apply(name);
|
||||
static boolean isForeignWorker(String name, String currentNonce) {
|
||||
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,160 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.peer.Capability;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import dev.ltms.bridged.peer.SpawnRequest;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.EnumSet;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
|
||||
/**
|
||||
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
|
||||
* router in front of one {@link HerdrPeerLauncher} per peer {@code kind} (Claude Code, opencode, …).
|
||||
* It owns no transport of its own; it dispatches each SPI call to the delegate that owns the profile
|
||||
* involved, and fans the fleet-wide queries (list/reap/caps/profiles) across all delegates.
|
||||
*
|
||||
* <p>Routing rules:
|
||||
* <ul>
|
||||
* <li><strong>By profile</strong> — {@link #spawn}, {@link #effectiveCwd}, {@link #parityOverlay}
|
||||
* resolve the profile (a null/blank name → the global {@link #defaultProfile}) and delegate to
|
||||
* the single adapter that declares it. Profiles partition cleanly across adapters: the
|
||||
* constructor rejects a name claimed by two.</li>
|
||||
* <li><strong>By pane id</strong> — {@link #stop} routes to the adapter that spawned that pane
|
||||
* (recorded at spawn time). A pane the composite never spawned (only real for a caller that
|
||||
* hand-rolls an id) falls back to the first delegate; teardown is pane-id addressed and
|
||||
* tab cleanup is single-occupant guarded, so it is safe either way.</li>
|
||||
* <li><strong>Fleet-wide</strong> — {@link #reapOrphanWorkers} and {@link #capabilities} fan out
|
||||
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
|
||||
* shares one herdr connection and so reports the same global agent set.</li>
|
||||
* </ul>
|
||||
*/
|
||||
public final class CompositePeerLauncher implements PeerLauncher {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(CompositePeerLauncher.class);
|
||||
|
||||
private final List<HerdrPeerLauncher> delegates;
|
||||
private final Map<String, HerdrPeerLauncher> byProfile;
|
||||
private final String defaultProfile;
|
||||
|
||||
/** paneId → the delegate that spawned it, so {@link #stop} tears down through the right adapter. */
|
||||
private final Map<String, HerdrPeerLauncher> spawnedBy = new ConcurrentHashMap<>();
|
||||
|
||||
/**
|
||||
* @param delegates one adapter per configured peer kind; must be non-empty and declare
|
||||
* disjoint profile-name sets
|
||||
* @param defaultProfile the profile a no-argument spawn resolves to (may be null)
|
||||
* @throws IllegalArgumentException if {@code delegates} is empty or two adapters claim one profile
|
||||
*/
|
||||
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates, String defaultProfile) {
|
||||
if (delegates.isEmpty()) {
|
||||
throw new IllegalArgumentException("at least one peer adapter must be configured");
|
||||
}
|
||||
this.delegates = List.copyOf(delegates);
|
||||
this.defaultProfile = defaultProfile;
|
||||
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
|
||||
for (HerdrPeerLauncher d : this.delegates) {
|
||||
for (String profile : d.profiles()) {
|
||||
HerdrPeerLauncher prev = index.putIfAbsent(profile, d);
|
||||
if (prev != null) {
|
||||
throw new IllegalArgumentException(
|
||||
"worker profile '" + profile + "' is claimed by two peer adapters");
|
||||
}
|
||||
}
|
||||
}
|
||||
this.byProfile = Map.copyOf(index);
|
||||
}
|
||||
|
||||
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
|
||||
private HerdrPeerLauncher route(String profileName) {
|
||||
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
|
||||
if (resolved == null) {
|
||||
// No profile and no default configured — hand to the first delegate so it raises the
|
||||
// same "no default" error it would on its own; keeps the SPI contract single-sourced.
|
||||
return delegates.getFirst();
|
||||
}
|
||||
HerdrPeerLauncher d = byProfile.get(resolved);
|
||||
if (d == null) {
|
||||
throw new IllegalArgumentException("unknown worker profile: " + resolved);
|
||||
}
|
||||
return d;
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req) {
|
||||
HerdrPeerLauncher d = route(req.profileName());
|
||||
PeerHandle handle = d.spawn(req);
|
||||
spawnedBy.put(handle.id(), d);
|
||||
return handle;
|
||||
}
|
||||
|
||||
@Override
|
||||
public String effectiveCwd(SpawnRequest req) {
|
||||
return route(req.profileName()).effectiveCwd(req);
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<String> parityOverlay(String profileName) {
|
||||
return route(profileName).parityOverlay(profileName);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void stop(String id) {
|
||||
HerdrPeerLauncher d = spawnedBy.remove(id);
|
||||
if (d == null) {
|
||||
log.debug("stop({}) — no recorded owner, routing to the first adapter (pane-addressed)", id);
|
||||
d = delegates.getFirst();
|
||||
}
|
||||
d.stop(id);
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return byProfile.keySet();
|
||||
}
|
||||
|
||||
@Override
|
||||
public String defaultProfile() {
|
||||
return defaultProfile;
|
||||
}
|
||||
|
||||
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
|
||||
@Override
|
||||
public List<Agent> list() {
|
||||
Map<String, Agent> byPane = new LinkedHashMap<>();
|
||||
for (HerdrPeerLauncher d : delegates) {
|
||||
for (Agent a : d.list()) {
|
||||
if (a.paneId() != null) {
|
||||
byPane.putIfAbsent(a.paneId(), a);
|
||||
}
|
||||
}
|
||||
}
|
||||
return List.copyOf(byPane.values());
|
||||
}
|
||||
|
||||
@Override
|
||||
public int reapOrphanWorkers() {
|
||||
int reaped = 0;
|
||||
for (HerdrPeerLauncher d : delegates) {
|
||||
reaped += d.reapOrphanWorkers();
|
||||
}
|
||||
return reaped;
|
||||
}
|
||||
|
||||
/** The union of every adapter's capabilities — a capability any adapter offers, the fleet offers. */
|
||||
@Override
|
||||
public Set<Capability> capabilities() {
|
||||
EnumSet<Capability> caps = EnumSet.noneOf(Capability.class);
|
||||
for (HerdrPeerLauncher d : delegates) {
|
||||
caps.addAll(d.capabilities());
|
||||
}
|
||||
return Set.copyOf(caps);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,546 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.herdr.Tab;
|
||||
import dev.ltms.bridged.herdr.Workspace;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||
import dev.ltms.bridged.peer.SpawnRequest;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.security.SecureRandom;
|
||||
import java.util.ArrayList;
|
||||
import java.util.Collection;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
/**
|
||||
* Abstract base for {@link PeerLauncher} adapters that materialize a peer as a <em>herdr</em>
|
||||
* agent (a CLI coding agent running in a herdr tab/pane). It owns everything that is the same
|
||||
* regardless of <em>which</em> coding agent runs: tab/pane placement, the CB-306 spawn-readiness
|
||||
* gate, unique naming, CB-117 orphan reap, teardown, {@link #list() listing}, and cwd resolution.
|
||||
*
|
||||
* <p>Two seams are peer-specific and supplied by the concrete adapter:
|
||||
* <ul>
|
||||
* <li>{@code namePrefix} (constructor arg) — the label prefix ({@code claude}, {@code opencode})
|
||||
* that drives both unique naming and the orphan-reap pattern, so each adapter reaps only its
|
||||
* own kind of pane and never another's.</li>
|
||||
* <li>{@link #buildLaunch(BridgedConfig.Worker)} — the peer-specific env map + argv, including any
|
||||
* subscription/guard check, MCP mount, and instruction injection. The base never sees how the
|
||||
* peer is configured; it only places and starts the returned {@link Launch}.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>Placement: in the default {@code tab} policy a peer lands in its own tab inside a dedicated
|
||||
* worker space (found-or-created once, then shared), so peers never split or clutter the user's
|
||||
* real work spaces. Teardown removes the peer's pane <em>and</em> its now-empty tab, tolerating an
|
||||
* already-gone peer so a repeated DELETE is harmless.
|
||||
*/
|
||||
public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(HerdrPeerLauncher.class);
|
||||
|
||||
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
|
||||
private static final int NAME_RETRIES = 8;
|
||||
|
||||
private final String namePrefix; // label prefix: naming + reap scheme
|
||||
private final AgentControl agents;
|
||||
private final WorkspaceControl spaces;
|
||||
private final Map<String, BridgedConfig.Worker> profiles; // profile name → spawn settings
|
||||
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
|
||||
|
||||
/** Host env lookup (injectable for tests); adapters read it in {@link #buildLaunch}. */
|
||||
protected final Function<String, String> env;
|
||||
|
||||
private final AtomicLong nameSeq = new AtomicLong(); // per-peer counter (also the tab #)
|
||||
|
||||
private final long spawnReadyTimeoutMs; // 0 = disable gate (legacy non-blocking spawn)
|
||||
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
|
||||
private final Runnable sleeper; // sleep/wait hook (injectable for tests; never real-sleep in unit tests)
|
||||
|
||||
// Per-process token mixed into each peer name so a fresh process (nameSeq back at 0) cannot
|
||||
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
|
||||
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
|
||||
|
||||
/**
|
||||
* @param namePrefix label prefix for this peer kind (drives naming and reap)
|
||||
* @param agents herdr agent control (start, status, close)
|
||||
* @param spaces workspace / tab control (ensure, create, close)
|
||||
* @param profiles configured peer profiles
|
||||
* @param defaultProfile profile a no-argument spawn uses (nullable)
|
||||
* @param env host env lookup (injectable for tests)
|
||||
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
|
||||
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
|
||||
* @param sleeper sleep/wait hook (never called when the gate is disabled); the poll
|
||||
* interval is baked into this hook, so the base needs no poll field
|
||||
*/
|
||||
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs,
|
||||
LongSupplier nowMillis, Runnable sleeper) {
|
||||
this.namePrefix = namePrefix;
|
||||
this.agents = agents;
|
||||
this.spaces = spaces;
|
||||
this.profiles = Map.copyOf(profiles);
|
||||
this.defaultProfile = defaultProfile;
|
||||
this.env = env;
|
||||
this.spawnReadyTimeoutMs = spawnReadyTimeoutMs;
|
||||
this.nowMillis = nowMillis;
|
||||
this.sleeper = sleeper;
|
||||
}
|
||||
|
||||
// --- adapter seams -------------------------------------------------------------------------
|
||||
|
||||
/**
|
||||
* Build the peer-specific launch for {@code cfg}: the environment map and argv handed to herdr.
|
||||
* Any subscription/guard check, MCP mount, and instruction injection happen here. The env map
|
||||
* and argv are adapter-private; the base only places and starts what is returned.
|
||||
*/
|
||||
protected abstract Launch buildLaunch(BridgedConfig.Worker cfg);
|
||||
|
||||
/** A peer-specific launch: the herdr {@code env} map and {@code argv}. */
|
||||
protected record Launch(Map<String, String> env, List<String> argv) {
|
||||
}
|
||||
|
||||
// --- profile surface -----------------------------------------------------------------------
|
||||
|
||||
/** The configured peer profile names (what {@code spawn(profile)} accepts). */
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return profiles.keySet();
|
||||
}
|
||||
|
||||
/** The parity-overlay file list for {@code profileName} (default list when unset). */
|
||||
@Override
|
||||
public List<String> parityOverlay(String profileName) {
|
||||
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
|
||||
if (name == null || name.isBlank()) {
|
||||
return List.of();
|
||||
}
|
||||
BridgedConfig.Worker cfg = profiles.get(name);
|
||||
return cfg == null ? List.of() : cfg.parityOverlay();
|
||||
}
|
||||
|
||||
/** The profile a no-argument spawn uses, or {@code null} if none is configured. */
|
||||
@Override
|
||||
public String defaultProfile() {
|
||||
return defaultProfile;
|
||||
}
|
||||
|
||||
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
|
||||
protected Collection<BridgedConfig.Worker> profileConfigs() {
|
||||
return profiles.values();
|
||||
}
|
||||
|
||||
/** Resolve {@code profileName} (null/blank → default) to its config, or throw with the options. */
|
||||
protected BridgedConfig.Worker requireProfile(String profileName) {
|
||||
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
|
||||
if (name == null || name.isBlank()) {
|
||||
throw new IllegalArgumentException("no default worker profile is configured — "
|
||||
+ "pass a profile; configured: " + profiles.keySet());
|
||||
}
|
||||
BridgedConfig.Worker cfg = profiles.get(name);
|
||||
if (cfg == null) {
|
||||
throw new IllegalArgumentException("unknown worker profile '" + name
|
||||
+ "' — configured: " + profiles.keySet());
|
||||
}
|
||||
return cfg;
|
||||
}
|
||||
|
||||
// --- spawn ---------------------------------------------------------------------------------
|
||||
|
||||
/**
|
||||
* Spawn a peer. {@code profileName} null/blank → the default profile. The working directory
|
||||
* (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd}, else the
|
||||
* profile's configured {@code cwd}, else {@code callerCwd} (the primary's cwd, when the spawn
|
||||
* came from the primary over MCP), else the daemon's cwd — never assumed to be {@code $HOME}.
|
||||
* The adapter's {@link #buildLaunch} runs before any herdr call.
|
||||
*/
|
||||
protected Agent spawnInternal(String profileName, String requestedCwd, String callerCwd) {
|
||||
BridgedConfig.Worker cfg = requireProfile(profileName);
|
||||
Launch launch = buildLaunch(cfg);
|
||||
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
|
||||
return cfg.tabPlacement()
|
||||
? spawnInTab(cfg, launch.env(), launch.argv(), cwd)
|
||||
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>Delegates to {@link #spawnInternal} and wraps the resulting herdr {@link Agent} in a
|
||||
* {@link WorkerHandle} whose {@link PeerHandle#id()} equals the agent's paneId. When
|
||||
* {@code spawnReadyTimeoutMs > 0}, blocks until the peer's herdr status is injectable or the
|
||||
* timeout elapses; on timeout the pane is closed (no orphan) and a
|
||||
* {@link PeerUnreachableException} is thrown.
|
||||
*/
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req) {
|
||||
Agent agent = spawnInternal(req.profileName(), req.requestedCwd(), req.callerCwd());
|
||||
String paneId = agent.paneId();
|
||||
if (spawnReadyTimeoutMs > 0) {
|
||||
waitUntilInjectableOrThrow(paneId);
|
||||
}
|
||||
return new WorkerHandle(paneId, agent.terminalId());
|
||||
}
|
||||
|
||||
@Override
|
||||
public String effectiveCwd(SpawnRequest req) {
|
||||
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-301: the effective working directory a spawn for {@code profileName} would use, without
|
||||
* actually spawning.
|
||||
*/
|
||||
private String effectiveCwd(String profileName, String requestedCwd, String callerCwd) {
|
||||
return resolveCwd(requestedCwd, requireProfile(profileName), callerCwd);
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-112 cwd resolution: spawn arg → profile config → the primary's cwd → the daemon's cwd.
|
||||
* Never returns {@code null}/blank: {@code "."} (the daemon's own working directory) is the
|
||||
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
|
||||
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
|
||||
*/
|
||||
private static String resolveCwd(String requestedCwd, BridgedConfig.Worker cfg, String callerCwd) {
|
||||
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
|
||||
}
|
||||
|
||||
private static String firstNonBlank(String... values) {
|
||||
for (String v : values) {
|
||||
if (v != null && !v.isBlank()) return v;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/** Dedicated worker space → own tab → start the peer (rooted at {@code cwd}) → drop the shell. */
|
||||
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
|
||||
List<String> argv, String cwd) {
|
||||
Workspace space = spaces.ensureWorkspace(cfg.workspace());
|
||||
Tab.Created tab = spaces.createTab(space.workspaceId());
|
||||
log.info("spawning {} profile={} space={} tab={} cwd={}",
|
||||
namePrefix, cfg.profile(), space.workspaceId(), tab.tab().tabId(), cwd);
|
||||
|
||||
Started started;
|
||||
try {
|
||||
started = startUniquelyNamed(cfg, workerEnv, argv, tab.tab().tabId(), cwd);
|
||||
} catch (RuntimeException e) {
|
||||
// The peer never started — don't leave the tab we just created orphaned.
|
||||
// Best-effort cleanup; never let it mask the real spawn failure.
|
||||
try {
|
||||
spaces.closeTab(tab.tab().tabId());
|
||||
} catch (RuntimeException cleanup) {
|
||||
log.warn("failed to close orphaned tab {} after spawn error: {}",
|
||||
tab.tab().tabId(), cleanup.getMessage());
|
||||
}
|
||||
throw e;
|
||||
}
|
||||
|
||||
// The peer is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell so the
|
||||
// tab holds only the peer; label the tab). They must not fail the spawn or orphan the
|
||||
// running peer — on error we log and still return it so the caller gets its paneId and can
|
||||
// tear it down.
|
||||
if (tab.rootPaneId() != null) {
|
||||
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
|
||||
} else {
|
||||
log.warn("tab {} had no seed pane in the create response; peer tab may hold an extra pane",
|
||||
tab.tab().tabId());
|
||||
}
|
||||
tidy("label tab " + tab.tab().tabId(),
|
||||
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
|
||||
log.info("{} started pane={} tab={} terminal={}",
|
||||
namePrefix, started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
|
||||
return started.agent();
|
||||
}
|
||||
|
||||
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
|
||||
private void tidy(String what, Runnable step) {
|
||||
try {
|
||||
step.run();
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("post-start step failed ({}) — peer is running regardless: {}", what, e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
/** Legacy placement: herdr splits the currently-focused tab; the peer still starts in {@code cwd}. */
|
||||
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
|
||||
List<String> argv, String cwd) {
|
||||
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
|
||||
namePrefix, cfg.profile(), cwd, argv);
|
||||
Agent peer = startUniquelyNamed(cfg, workerEnv, argv, null, cwd).agent();
|
||||
log.info("{} started pane={} terminal={}", namePrefix, peer.paneId(), peer.terminalId());
|
||||
return peer;
|
||||
}
|
||||
|
||||
/** A started peer together with the sequence its unique name/label used. */
|
||||
private record Started(Agent agent, long seq) {
|
||||
}
|
||||
|
||||
/**
|
||||
* Start the peer under a unique herdr agent name. herdr requires each running agent's
|
||||
* {@code name} to be distinct (a 2nd identical {@code name} fails {@code agent_name_taken}) —
|
||||
* the exact case that makes multiple peers useful. The name is
|
||||
* {@code <prefix>-<profile>-<nonce>-<seq>}: {@code seq} distinguishes peers within this process,
|
||||
* and the per-process {@code nonce} keeps a fresh process (whose {@code seq} restarts at 0) from
|
||||
* colliding with same-profile peers that outlived a restart. The retry is a belt-and-braces
|
||||
* backstop for the astronomically unlikely nonce+seq clash; the name is a label only — herdr
|
||||
* detects kind and status from terminal output, not from it.
|
||||
*/
|
||||
private Started startUniquelyNamed(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
|
||||
List<String> argv, String tabId, String cwd) {
|
||||
HerdrException last = null;
|
||||
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
|
||||
long seq = nameSeq.incrementAndGet();
|
||||
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
|
||||
try {
|
||||
return new Started(agents.start(name, argv, workerEnv, tabId, cwd), seq);
|
||||
} catch (HerdrException e) {
|
||||
if (!"agent_name_taken".equals(e.code())) throw e;
|
||||
log.debug("peer name '{}' taken, retrying", name);
|
||||
last = e;
|
||||
}
|
||||
}
|
||||
throw last;
|
||||
}
|
||||
|
||||
// --- discovery + reap ----------------------------------------------------------------------
|
||||
|
||||
/** All herdr-tracked agents — discovery for "what peers exist". */
|
||||
@Override
|
||||
public List<Agent> list() {
|
||||
return agents.list();
|
||||
}
|
||||
|
||||
/**
|
||||
* Reap peer panes left behind by an earlier daemon process (CB-117). herdr keeps a peer's pane
|
||||
* alive across a daemon restart <em>by design</em>, and that pane's id is held only by its
|
||||
* spawner — so a peer whose owning process exited before issuing the matching teardown leaks
|
||||
* with nothing tracking it. On boot we scan herdr for agents whose name matches our
|
||||
* {@code <prefix>-<profile>-<nonce>-<seq>} scheme with a nonce <em>other</em> than this
|
||||
* process's {@link #nameNonce}, and tear each one down (its pane and, via {@link #stop}, its
|
||||
* now-empty dedicated tab). A current-nonce peer is ours and live, so it is left running; a
|
||||
* user's own session carries no such name and is never touched. A peer from a <em>different</em>
|
||||
* adapter (different prefix) is likewise never touched. Best-effort: a failed listing, or a
|
||||
* failure to stop any one peer, is logged and never aborts startup.
|
||||
*
|
||||
* @return the number of orphaned peers reaped
|
||||
*/
|
||||
@Override
|
||||
public int reapOrphanWorkers() {
|
||||
List<Agent> all;
|
||||
try {
|
||||
all = agents.list();
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("orphan-peer reap skipped — agent.list failed: {}", e.getMessage());
|
||||
return 0;
|
||||
}
|
||||
int reaped = 0;
|
||||
for (Agent a : all) {
|
||||
if (!isForeignWorker(namePrefix, a.name(), nameNonce)) continue;
|
||||
try {
|
||||
stop(a.paneId());
|
||||
reaped++;
|
||||
log.info("reaped orphan {} {} (pane={} tab={}) left by a prior daemon",
|
||||
namePrefix, a.name(), a.paneId(), a.tabId());
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("could not reap orphan {} {} (pane={}): {}",
|
||||
namePrefix, a.name(), a.paneId(), e.getMessage());
|
||||
}
|
||||
}
|
||||
if (reaped > 0) {
|
||||
log.info("orphan-peer reap complete — {} stale {} peer(s) removed at startup", reaped, namePrefix);
|
||||
}
|
||||
return reaped;
|
||||
}
|
||||
|
||||
/** The {@code <prefix>-<profile>-<nonce>-<seq>} name pattern; group 1 captures the 6-hex nonce. */
|
||||
static Pattern workerNamePattern(String prefix) {
|
||||
return Pattern.compile(prefix + "-.*-([0-9a-f]{6})-\\d+");
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether {@code name} is a peer of kind {@code prefix} started by a <em>different</em> process
|
||||
* than {@code currentNonce} — the reap predicate (CB-117). True only for the prefix's naming
|
||||
* scheme with a foreign nonce: a non-peer name, a different adapter's name, or our own live
|
||||
* nonce is excluded. Pure and package-private so the decision is unit-testable without herdr.
|
||||
*/
|
||||
static boolean isForeignWorker(String prefix, String name, String currentNonce) {
|
||||
String nonce = workerNonce(prefix, name);
|
||||
return nonce != null && !nonce.equals(currentNonce);
|
||||
}
|
||||
|
||||
/** The 6-hex nonce embedded in a {@code prefix} peer name, or {@code null} if not one. */
|
||||
static String workerNonce(String prefix, String name) {
|
||||
if (name == null) return null;
|
||||
Matcher m = workerNamePattern(prefix).matcher(name);
|
||||
return m.matches() ? m.group(1) : null;
|
||||
}
|
||||
|
||||
/** This process's peer-name nonce (a label component only; exposed for reaper tests). */
|
||||
String nameNonce() {
|
||||
return nameNonce;
|
||||
}
|
||||
|
||||
// --- teardown ------------------------------------------------------------------------------
|
||||
|
||||
/**
|
||||
* Tear a peer down by pane id: close the pane, and close its tab <em>only</em> when the peer is
|
||||
* that tab's sole occupant. The single-pane check is what makes this safe regardless of how the
|
||||
* peer was placed (or a placement-config change across a restart): a pane-placement peer sitting
|
||||
* in one of the user's shared tabs has siblings, so its tab is never closed — we only ever
|
||||
* remove a tab we created to hold one peer.
|
||||
*
|
||||
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
|
||||
* (repeated DELETE, crashed peer) is treated as success; any other failure propagates so a
|
||||
* genuinely failed teardown is not reported as done.
|
||||
*/
|
||||
@Override
|
||||
public void stop(String paneId) {
|
||||
// Teardown knows only the paneId, not which profile spawned it. Attempt tab cleanup when any
|
||||
// profile uses tab placement (so the bridge may have created a dedicated peer tab); the
|
||||
// single-occupant check below is what actually protects the user's shared tabs.
|
||||
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
|
||||
try {
|
||||
agents.close(paneId);
|
||||
} catch (HerdrException e) {
|
||||
if (!isAlreadyGone(e)) throw e;
|
||||
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
|
||||
}
|
||||
if (loc != null && loc.tabPaneCount() == 1) {
|
||||
spaces.closeTab(loc.tabId());
|
||||
} else if (loc != null) {
|
||||
log.debug("not closing tab {} — it holds {} panes (not a dedicated peer tab)",
|
||||
loc.tabId(), loc.tabPaneCount());
|
||||
}
|
||||
}
|
||||
|
||||
/** Whether any configured profile places peers in their own tab (so tabs may need cleanup). */
|
||||
private boolean usesTabPlacement() {
|
||||
return profiles.values().stream().anyMatch(BridgedConfig.Worker::tabPlacement);
|
||||
}
|
||||
|
||||
/** True when a herdr error means the target is already gone (safe to treat as done). */
|
||||
private static boolean isAlreadyGone(HerdrException e) {
|
||||
return e.code() != null && e.code().endsWith("_not_found");
|
||||
}
|
||||
|
||||
// --- spawn-readiness gate (CB-306) ---------------------------------------------------------
|
||||
|
||||
/**
|
||||
* Poll {@link AgentControl#status} until the pane reports an injectable state or the configured
|
||||
* timeout elapses. On timeout, close the pane (self-reap) and throw.
|
||||
*/
|
||||
private void waitUntilInjectableOrThrow(String paneId) {
|
||||
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
|
||||
while (nowMillis.getAsLong() < deadline) {
|
||||
if (agents.status(paneId).injectable()) {
|
||||
log.debug("peer pane={} reached injectable state", paneId);
|
||||
return;
|
||||
}
|
||||
sleeper.run();
|
||||
}
|
||||
log.warn("peer pane={} did not become injectable within {}ms — closing", paneId, spawnReadyTimeoutMs);
|
||||
stop(paneId);
|
||||
throw new PeerUnreachableException(
|
||||
"worker pane " + paneId + " did not reach injectable state within "
|
||||
+ spawnReadyTimeoutMs + "ms");
|
||||
}
|
||||
|
||||
/** A concrete {@link PeerHandle} wrapping herdr agent coordinates. */
|
||||
private record WorkerHandle(String id, String terminalId) implements PeerHandle {
|
||||
}
|
||||
|
||||
// --- shared helpers ------------------------------------------------------------------------
|
||||
|
||||
/** Put {@code k → v} only when {@code v} is present (non-null, non-blank). */
|
||||
protected static void putIfPresent(Map<String, String> m, String k, String v) {
|
||||
if (v != null && !v.isBlank()) {
|
||||
m.put(k, v);
|
||||
}
|
||||
}
|
||||
|
||||
/** Host env lookup that tolerates an unconfigured (null/blank) var name — returns null then. */
|
||||
protected String resolveEnv(String name) {
|
||||
return (name == null || name.isBlank()) ? null : env.apply(name);
|
||||
}
|
||||
|
||||
/**
|
||||
* The parity-neutral git-forge token grant (CB-302): when {@code cfg} opts in via
|
||||
* {@code gitTokenEnv} and the token resolves, inject {@code GITEA_TOKEN} plus its paired
|
||||
* {@code GITEA_HOST}. Push over SSH is unaffected; the only incremental grant is PR-create.
|
||||
* Peer-neutral, so every herdr adapter reuses it unchanged.
|
||||
*/
|
||||
protected void applyGitToken(Map<String, String> workerEnv, BridgedConfig.Worker cfg) {
|
||||
if (!cfg.hasGitToken()) {
|
||||
return;
|
||||
}
|
||||
String gitToken = resolveEnv(cfg.gitTokenEnv());
|
||||
if (gitToken != null) {
|
||||
workerEnv.put("GITEA_TOKEN", gitToken);
|
||||
putIfPresent(workerEnv, "GITEA_HOST", resolveEnv(cfg.gitHostEnv()));
|
||||
}
|
||||
}
|
||||
|
||||
/** A fresh mutable env map — the conventional starting point for {@link #buildLaunch}. */
|
||||
/**
|
||||
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
|
||||
* {@code env:} entries.
|
||||
*
|
||||
* <p>Why this exists: bridged passes herdr an explicit env map, and herdr merges it into
|
||||
* <em>its own</em> process environment. So before this, a worker inherited whatever PATH the
|
||||
* herdr server happened to be started with — on this host, one from weeks earlier with no JDK
|
||||
* and no Maven, which left workers unable to run the build they were being asked to run. The
|
||||
* worker's toolchain must follow from configuration, not from how a long-lived daemon was
|
||||
* launched.
|
||||
*
|
||||
* <p>Adapter-specific variables are layered on top of this by {@code buildLaunch} and therefore
|
||||
* win. That ordering is deliberate and load-bearing: it stops a profile's {@code env:} from
|
||||
* overriding {@code ANTHROPIC_BASE_URL} and slipping past {@link
|
||||
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
|
||||
* {@code baseUrl} and nothing else.
|
||||
*/
|
||||
protected Map<String, String> baseEnv(BridgedConfig.Worker cfg) {
|
||||
Map<String, String> workerEnv = new LinkedHashMap<>();
|
||||
String path = env.apply("PATH");
|
||||
if (path != null && !path.isBlank()) {
|
||||
workerEnv.put("PATH", path);
|
||||
}
|
||||
if (cfg != null && cfg.env() != null) {
|
||||
workerEnv.putAll(cfg.env());
|
||||
}
|
||||
return workerEnv;
|
||||
}
|
||||
|
||||
/** Defensive copy of {@code argv} plus room to append launch flags. */
|
||||
protected static List<String> mutableArgv(List<String> argv) {
|
||||
return new ArrayList<>(argv);
|
||||
}
|
||||
|
||||
/**
|
||||
* Uninterruptible sleep — the production {@link #sleeper}. Tests supply their own no-op /
|
||||
* fast-faking sleeper so they never real-sleep.
|
||||
*/
|
||||
protected static void sleepUninterruptibly(long ms) {
|
||||
try {
|
||||
Thread.sleep(ms);
|
||||
} catch (InterruptedException e) {
|
||||
Thread.currentThread().interrupt();
|
||||
// preserve the interrupt flag but continue — poll loops should not be aborted by an
|
||||
// interrupt that was not meant for them.
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,316 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.fasterxml.jackson.databind.node.ObjectNode;
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.peer.Capability;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.io.UncheckedIOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.EnumSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
/**
|
||||
* The {@link HerdrPeerLauncher} adapter for <strong>opencode</strong> — an open-source,
|
||||
* provider-agnostic terminal coding agent. Its whole reason for existing is to prove the
|
||||
* {@code PeerLauncher} SPI is genuinely provider-neutral: opencode shares none of Claude Code's
|
||||
* private launch seams, yet reuses every line of shared transport in the base (tab/pane placement,
|
||||
* the CB-306 readiness gate, unique naming + CB-117 reap, teardown, listing, cwd).
|
||||
*
|
||||
* <p>The divergences from {@link ClaudeCodeLauncher}, all confined to {@link #buildLaunch}:
|
||||
* <ul>
|
||||
* <li><strong>No subscription boundary.</strong> opencode carries no {@code ANTHROPIC_BASE_URL}
|
||||
* and there is no {@link dev.ltms.bridged.guard.SubscriptionGuard} — the guard is a
|
||||
* Claude-private concern, not part of the SPI. opencode reads the operator's own provider
|
||||
* credentials from its global {@code auth.json}; the bridge injects none.</li>
|
||||
* <li><strong>File-based MCP mount + instructions.</strong> opencode has no inline
|
||||
* {@code --mcp-config}/{@code --append-system-prompt}. Instead the bridge writes an ephemeral
|
||||
* {@code opencode.json} that declares the bridge as a {@code remote} MCP server and lists a
|
||||
* reply-charter file under {@code instructions}, then points the worker at it with
|
||||
* {@code OPENCODE_CONFIG}. This is the one place the launcher touches disk — Claude never did.</li>
|
||||
* <li><strong>Model as a flag.</strong> the {@code provider/model} selector is passed as
|
||||
* {@code -m}, not an env var.</li>
|
||||
* <li><strong>{@code opencode} name prefix</strong> so reap matches {@code opencode-*} panes and
|
||||
* never another adapter's.</li>
|
||||
* </ul>
|
||||
*/
|
||||
public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
|
||||
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
|
||||
private static final String NAME_PREFIX = "opencode";
|
||||
|
||||
/** Writer for the generated {@code opencode.json}. */
|
||||
private static final ObjectMapper JSON = new ObjectMapper();
|
||||
|
||||
/**
|
||||
* Standing instruction written to the charter file and mounted via the config's
|
||||
* {@code instructions} so the worker returns its result through {@code bridge_reply}. Kept on
|
||||
* disk (not a launch flag) because opencode's {@code instructions} takes file paths, not inline
|
||||
* text — the file is regenerated per spawn and never touches the worker's own profile.
|
||||
*/
|
||||
static final String REPLY_CHARTER =
|
||||
"You are an off-subscription worker in the claude-bridge fleet, running under opencode. "
|
||||
+ "Every message you receive arrives through the bridge, and the ONLY channel back to the "
|
||||
+ "sender is the bridge_reply MCP tool. Text you write in your terminal is NOT sent "
|
||||
+ "anywhere — the sender cannot see your screen, so an in-terminal answer is silently "
|
||||
+ "discarded. Therefore you MUST end EVERY turn by calling bridge_reply with `content` set "
|
||||
+ "to your complete response. This holds for every message without exception — tasks, "
|
||||
+ "questions, clarifications, acknowledgements, and ordinary back-and-forth conversation. "
|
||||
+ "Call bridge_reply exactly once, as the final action of your turn, with your full answer "
|
||||
+ "in `content`; never wait for confirmation first. If you end a turn without calling "
|
||||
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
|
||||
|
||||
/** Root under which per-spawn opencode config dirs are created (injectable for tests). */
|
||||
private final Path configRoot;
|
||||
|
||||
/**
|
||||
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so it
|
||||
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
|
||||
*/
|
||||
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Function<String, String> env) {
|
||||
this(agents, spaces, profiles, defaultProfile, env, 0,
|
||||
System::currentTimeMillis, () -> sleepUninterruptibly(300),
|
||||
defaultConfigRoot());
|
||||
}
|
||||
|
||||
/**
|
||||
* Production constructor with the spawn-ready gate enabled. Polls {@code agents.status()} until
|
||||
* the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
|
||||
*/
|
||||
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
|
||||
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
|
||||
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
|
||||
defaultConfigRoot());
|
||||
}
|
||||
|
||||
/**
|
||||
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply a
|
||||
* fake clock ({@code nowMillis}), poll-loop wait ({@code sleeper}), and a temp {@code configRoot}
|
||||
* they can inspect the generated {@code opencode.json}/charter under.
|
||||
*
|
||||
* @param agents herdr agent control (start, status, close)
|
||||
* @param spaces workspace / tab control (ensure, create, close)
|
||||
* @param profiles configured worker profiles
|
||||
* @param defaultProfile profile a no-argument spawn uses (nullable)
|
||||
* @param env host env lookup (injectable for tests)
|
||||
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
|
||||
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
|
||||
* @param sleeper sleep/wait hook (encodes the poll interval; never called when the
|
||||
* gate is disabled)
|
||||
* @param configRoot existing directory under which per-spawn config dirs are created
|
||||
*/
|
||||
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs,
|
||||
LongSupplier nowMillis, Runnable sleeper, Path configRoot) {
|
||||
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
|
||||
spawnReadyTimeoutMs, nowMillis, sleeper);
|
||||
this.configRoot = configRoot;
|
||||
}
|
||||
|
||||
private static Path defaultConfigRoot() {
|
||||
return Path.of(System.getProperty("java.io.tmpdir"));
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>Builds the opencode launch: no {@code ANTHROPIC_*} and no guard (opencode reads its own
|
||||
* provider credentials); when the profile mounts the bridge MCP, generate an ephemeral
|
||||
* {@code opencode.json} (remote MCP server + reply-charter instructions) and point the worker at
|
||||
* it via {@code OPENCODE_CONFIG}; carry the parity-neutral git-forge grant; and select the model
|
||||
* with {@code -m}.
|
||||
*/
|
||||
@Override
|
||||
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
|
||||
Map<String, String> workerEnv = baseEnv(cfg);
|
||||
// A config file is needed for the bridge MCP mount, for a pinned endpoint (CB-508), or both.
|
||||
if (cfg.hasMcp() || hasCustomProvider(cfg)) {
|
||||
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg).toString());
|
||||
}
|
||||
applyGitToken(workerEnv, cfg);
|
||||
return new Launch(workerEnv, argvWithModel(cfg));
|
||||
}
|
||||
|
||||
/**
|
||||
* True when this profile pins its own OpenAI-compatible endpoint (CB-508) rather than using
|
||||
* whatever provider opencode resolves by default.
|
||||
*
|
||||
* <p>Note this reuses {@code baseUrl}, the same field the Claude adapter injects as
|
||||
* {@code ANTHROPIC_BASE_URL} — but it does <em>not</em> go through {@code SubscriptionGuard}.
|
||||
* That asymmetry is deliberate and safe: the guard exists to stop a worker borrowing the
|
||||
* primary's Anthropic subscription, and an opencode process has no Anthropic credential path
|
||||
* at all. Pointing it at a local vLLM cannot leak the subscription.
|
||||
*/
|
||||
private static boolean hasCustomProvider(BridgedConfig.Worker cfg) {
|
||||
return cfg.baseUrl() != null && !cfg.baseUrl().isBlank();
|
||||
}
|
||||
|
||||
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
|
||||
private List<String> argvWithModel(BridgedConfig.Worker cfg) {
|
||||
List<String> argv = mutableArgv(cfg.argv());
|
||||
if (cfg.model() != null && !cfg.model().isBlank()) {
|
||||
argv.add("-m");
|
||||
argv.add(cfg.model());
|
||||
}
|
||||
return argv;
|
||||
}
|
||||
|
||||
/**
|
||||
* Write an ephemeral {@code opencode.json} (and the reply-charter file it references) into a
|
||||
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
|
||||
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
|
||||
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
|
||||
*/
|
||||
private Path writeConfig(BridgedConfig.Worker cfg) {
|
||||
try {
|
||||
Path dir = Files.createTempDirectory(configRoot, "bridged-opencode-");
|
||||
dir.toFile().deleteOnExit();
|
||||
|
||||
ObjectNode root = JSON.createObjectNode();
|
||||
root.put("$schema", "https://opencode.ai/config.json");
|
||||
|
||||
if (cfg.hasMcp()) {
|
||||
Path charter = dir.resolve("reply-charter.md");
|
||||
Files.writeString(charter, REPLY_CHARTER);
|
||||
charter.toFile().deleteOnExit();
|
||||
|
||||
ObjectNode bridge = root.putObject("mcp").putObject("bridge");
|
||||
bridge.put("type", "remote");
|
||||
bridge.put("url", cfg.mcpUrl());
|
||||
bridge.put("enabled", true);
|
||||
root.putArray("instructions").add(charter.toAbsolutePath().toString());
|
||||
}
|
||||
if (hasCustomProvider(cfg)) {
|
||||
addCustomProvider(root, cfg);
|
||||
}
|
||||
|
||||
Path cfgFile = dir.resolve("opencode.json");
|
||||
// Built with Jackson rather than string concatenation: the provider block is nested and
|
||||
// carries operator-supplied values (URL, model id, api key), so escaping must be real.
|
||||
Files.writeString(cfgFile, JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root));
|
||||
cfgFile.toFile().deleteOnExit();
|
||||
return cfgFile;
|
||||
} catch (IOException e) {
|
||||
throw new UncheckedIOException(
|
||||
"cannot write opencode config for profile " + cfg.profile(), e);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Declare a custom OpenAI-compatible provider so the worker talks to a pinned endpoint (a local
|
||||
* vLLM, say) instead of opencode's default gateway (CB-508).
|
||||
*
|
||||
* <p>The provider id comes from the {@code provider/model} selector in {@code model:}, so one
|
||||
* field drives both the declaration and the {@code -m} flag and they cannot drift apart.
|
||||
*/
|
||||
private void addCustomProvider(ObjectNode root, BridgedConfig.Worker cfg) {
|
||||
String[] parts = splitModelSelector(cfg);
|
||||
String providerId = parts[0];
|
||||
String modelId = parts[1];
|
||||
|
||||
ObjectNode provider = root.putObject("provider").putObject(providerId);
|
||||
provider.put("npm", "@ai-sdk/openai-compatible");
|
||||
provider.put("name", providerId + " (bridged)");
|
||||
|
||||
ObjectNode options = provider.putObject("options");
|
||||
options.put("baseURL", openAiBaseUrl(cfg.baseUrl()));
|
||||
// vLLM and friends usually ignore the key, but the AI SDK still requires a non-empty one.
|
||||
String token = resolveEnv(cfg.tokenEnv());
|
||||
options.put("apiKey", (token == null || token.isBlank()) ? "bridged-local-noauth" : token);
|
||||
|
||||
provider.putObject("models").putObject(modelId).put("name", modelId);
|
||||
}
|
||||
|
||||
/**
|
||||
* Split {@code model:} into its {@code provider} and {@code model} halves. A pinned endpoint
|
||||
* needs both, so a bare model name is rejected loudly rather than silently falling back to the
|
||||
* default gateway — a worker quietly talking to the wrong endpoint is the failure this avoids.
|
||||
*/
|
||||
private static String[] splitModelSelector(BridgedConfig.Worker cfg) {
|
||||
String model = cfg.model();
|
||||
int slash = model == null ? -1 : model.indexOf('/');
|
||||
if (model == null || model.isBlank() || slash <= 0 || slash == model.length() - 1) {
|
||||
throw new IllegalArgumentException(
|
||||
"profile " + cfg.profile() + " sets baseUrl (a pinned opencode endpoint) so"
|
||||
+ " model: must be \"<provider>/<model>\", e.g."
|
||||
+ " \"local-vllm/deepseek-v4-flash\"; got "
|
||||
+ (model == null ? "null" : '"' + model + '"'));
|
||||
}
|
||||
return new String[]{model.substring(0, slash), model.substring(slash + 1)};
|
||||
}
|
||||
|
||||
/**
|
||||
* The OpenAI-compatible base URL for {@code baseUrl}. A bare {@code host:port} gets {@code /v1}
|
||||
* appended (where these servers put the API); a URL that already carries a path is taken as-is,
|
||||
* so an endpoint mounted somewhere unusual is still reachable.
|
||||
*/
|
||||
private static String openAiBaseUrl(String baseUrl) {
|
||||
String trimmed = baseUrl.trim();
|
||||
while (trimmed.endsWith("/")) {
|
||||
trimmed = trimmed.substring(0, trimmed.length() - 1);
|
||||
}
|
||||
int schemeEnd = trimmed.indexOf("://");
|
||||
String afterScheme = schemeEnd < 0 ? trimmed : trimmed.substring(schemeEnd + 3);
|
||||
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
|
||||
}
|
||||
|
||||
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
|
||||
|
||||
/** Spawn a worker for the default profile in the resolved default cwd. */
|
||||
public Agent spawn() {
|
||||
return spawnInternal(null, null, null);
|
||||
}
|
||||
|
||||
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
|
||||
public Agent spawn(String profileName) {
|
||||
return spawnInternal(profileName, null, null);
|
||||
}
|
||||
|
||||
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
|
||||
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
|
||||
return spawnInternal(profileName, requestedCwd, callerCwd);
|
||||
}
|
||||
|
||||
// --- capabilities --------------------------------------------------------------------------
|
||||
|
||||
@Override
|
||||
public Set<Capability> capabilities() {
|
||||
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
|
||||
if (hasGitTokenProfile()) {
|
||||
caps.add(Capability.SELF_PR);
|
||||
}
|
||||
return Set.copyOf(caps);
|
||||
}
|
||||
|
||||
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
|
||||
private boolean hasGitTokenProfile() {
|
||||
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
|
||||
}
|
||||
|
||||
// --- CB-117 reap predicate (opencode prefix), kept for direct unit testing -----------------
|
||||
|
||||
/**
|
||||
* Whether {@code name} is an opencode bridge worker started by a <em>different</em> process than
|
||||
* {@code currentNonce}. A thin {@code opencode}-prefix binding of
|
||||
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
|
||||
*/
|
||||
static boolean isForeignWorker(String name, String currentNonce) {
|
||||
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
|
||||
}
|
||||
}
|
||||
@@ -5,10 +5,39 @@
|
||||
</encoder>
|
||||
</appender>
|
||||
|
||||
<!--
|
||||
CB-505 audit trail. Its own file, deliberately not the app log: privileged actions
|
||||
(spawn/stop/send/reply/drain) must stay greppable and shippable without dragging DEBUG noise
|
||||
along. AuditLog emits a complete JSON object including its own ISO-8601 "ts" field, so the
|
||||
pattern is a bare %msg — a pattern that spliced literal braces around the message would
|
||||
collide with logback's own variable substitution. Rolls daily, 30 days retained, 100MB cap.
|
||||
|
||||
NOTE: records carry who/what/target/outcome only. Message CONTENT is never written here —
|
||||
this bridge carries source code and prompts, and an audit log that accumulated them would be
|
||||
a transcript archive rather than a control.
|
||||
-->
|
||||
<appender name="AUDIT" class="ch.qos.logback.core.rolling.RollingFileAppender">
|
||||
<file>logs/audit.log</file>
|
||||
<rollingPolicy class="ch.qos.logback.core.rolling.SizeAndTimeBasedRollingPolicy">
|
||||
<fileNamePattern>logs/audit.%d{yyyy-MM-dd}.%i.log</fileNamePattern>
|
||||
<maxFileSize>10MB</maxFileSize>
|
||||
<maxHistory>30</maxHistory>
|
||||
<totalSizeCap>100MB</totalSizeCap>
|
||||
</rollingPolicy>
|
||||
<encoder>
|
||||
<pattern>%msg%n</pattern>
|
||||
</encoder>
|
||||
</appender>
|
||||
|
||||
<logger name="dev.ltms.bridged" level="DEBUG"/>
|
||||
<logger name="io.javalin" level="INFO"/>
|
||||
<logger name="org.eclipse.jetty" level="WARN"/>
|
||||
|
||||
<!-- additivity=false keeps the audit stream out of stdout; it is its own record. -->
|
||||
<logger name="audit" level="INFO" additivity="false">
|
||||
<appender-ref ref="AUDIT"/>
|
||||
</logger>
|
||||
|
||||
<root level="INFO">
|
||||
<appender-ref ref="STDOUT"/>
|
||||
</root>
|
||||
|
||||
@@ -0,0 +1,100 @@
|
||||
package dev.ltms.bridged.auth;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.LoggerContext;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* CB-505 — the audit record's shape.
|
||||
*
|
||||
* <p>These exist because the first cut of this feature emitted lines that were <em>not</em> valid
|
||||
* JSON: the timestamp was spliced on by a logback pattern whose literal braces collided with
|
||||
* logback's variable substitution. The appender failed to parse, and nothing in the build noticed.
|
||||
* An audit trail that silently stops being machine-readable is worse than none.
|
||||
*/
|
||||
class AuditLogTest {
|
||||
|
||||
private final ObjectMapper mapper = new ObjectMapper();
|
||||
private ListAppender<ILoggingEvent> appender;
|
||||
private ch.qos.logback.classic.Logger auditLogger;
|
||||
|
||||
@BeforeEach
|
||||
void attach() {
|
||||
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
|
||||
auditLogger = ctx.getLogger("audit");
|
||||
appender = new ListAppender<>();
|
||||
appender.setContext(ctx);
|
||||
appender.start();
|
||||
auditLogger.addAppender(appender);
|
||||
auditLogger.setLevel(Level.INFO);
|
||||
}
|
||||
|
||||
@AfterEach
|
||||
void detach() {
|
||||
auditLogger.detachAppender(appender);
|
||||
}
|
||||
|
||||
private JsonNode onlyRecord() throws Exception {
|
||||
assertEquals(1, appender.list.size(), "exactly one audit line expected");
|
||||
String line = appender.list.getFirst().getFormattedMessage();
|
||||
return mapper.readTree(line); // throws if the line is not valid JSON
|
||||
}
|
||||
|
||||
@Test
|
||||
void anAllowedActionIsRecordedAsValidJson() throws Exception {
|
||||
AuditLog.allowed(Principal.primary(4242), Authz.Action.SPAWN, "term_a");
|
||||
|
||||
JsonNode r = onlyRecord();
|
||||
assertEquals("PRIMARY", r.path("role").asText());
|
||||
assertEquals("primary", r.path("actor").asText());
|
||||
assertEquals(4242, r.path("pid").asLong());
|
||||
assertEquals("SPAWN", r.path("action").asText());
|
||||
assertEquals("term_a", r.path("target").asText());
|
||||
assertEquals("allowed", r.path("outcome").asText());
|
||||
assertFalse(r.path("ts").asText().isBlank(), "every record carries its own timestamp");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aDenialRecordsTheReason() throws Exception {
|
||||
AuditLog.denied(Principal.worker("term_b", 7), Authz.Action.REPLY, "term_a", "forbidden");
|
||||
|
||||
JsonNode r = onlyRecord();
|
||||
assertEquals("WORKER", r.path("role").asText());
|
||||
assertEquals("worker:term_b", r.path("actor").asText());
|
||||
assertEquals("denied", r.path("outcome").asText());
|
||||
assertEquals("forbidden", r.path("reason").asText());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aNullCallerIsRecordedAsAnonymousRatherThanCrashing() throws Exception {
|
||||
AuditLog.failed(null, Authz.Action.SEND, null, "herdr unreachable");
|
||||
|
||||
JsonNode r = onlyRecord();
|
||||
assertEquals("ANONYMOUS", r.path("role").asText());
|
||||
assertTrue(r.path("target").isNull(), "an absent target is JSON null, not the string \"null\"");
|
||||
assertEquals("failed", r.path("outcome").asText());
|
||||
}
|
||||
|
||||
@Test
|
||||
void hostileValuesAreEscapedAndCannotForgeAnExtraRecord() throws Exception {
|
||||
// A target id containing a quote and a newline must not be able to terminate the JSON
|
||||
// object early and inject a second, attacker-shaped audit line.
|
||||
AuditLog.denied(Principal.worker("term_a", 1), Authz.Action.REPLY,
|
||||
"evil\",\"outcome\":\"allowed\"}\n{\"forged\":true", "forbidden");
|
||||
|
||||
JsonNode r = onlyRecord();
|
||||
assertEquals("denied", r.path("outcome").asText(),
|
||||
"the injected outcome must not override the real one");
|
||||
assertTrue(r.path("target").asText().contains("forged"),
|
||||
"the hostile text survives as inert data inside the target field");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,80 @@
|
||||
package dev.ltms.bridged.auth;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static dev.ltms.bridged.auth.Authz.Action.*;
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/** CB-505 — the authorization table, pinned so it cannot drift silently. */
|
||||
class AuthzTest {
|
||||
|
||||
private static final Principal PRIMARY = Principal.primary(100);
|
||||
private static final Principal WORKER_A = Principal.worker("term_a", 200);
|
||||
private static final Principal WORKER_B = Principal.worker("term_b", 300);
|
||||
private static final Principal ANON = Principal.anonymous();
|
||||
|
||||
@Test
|
||||
void anonymousIsAuthorizedForNothing() {
|
||||
for (Authz.Action a : Authz.Action.values()) {
|
||||
assertFalse(Authz.permits(ANON, a, "term_a"),
|
||||
a + " must be refused to an unauthenticated caller");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void aNullCallerIsTreatedAsAnonymous() {
|
||||
assertFalse(Authz.permits(null, READ, null));
|
||||
assertTrue(Authz.isUnauthenticated(null));
|
||||
}
|
||||
|
||||
@Test
|
||||
void orchestrationBelongsToThePrimaryAlone() {
|
||||
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, SEND, DRAIN}) {
|
||||
assertTrue(Authz.permits(PRIMARY, a, "term_a"), "the primary orchestrates: " + a);
|
||||
assertFalse(Authz.permits(WORKER_A, a, "term_a"),
|
||||
"a worker performing " + a + " would be escalating into the orchestrator role");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWorkerMayReplyAndAskOnlyAsItself() {
|
||||
assertTrue(Authz.permits(WORKER_A, REPLY, "term_a"));
|
||||
assertTrue(Authz.permits(WORKER_A, ASK, "term_a"));
|
||||
|
||||
assertFalse(Authz.permits(WORKER_A, REPLY, "term_b"),
|
||||
"worker A must not be able to reply on worker B's session");
|
||||
assertFalse(Authz.permits(WORKER_B, ASK, "term_a"),
|
||||
"worker B must not be able to ask as worker A");
|
||||
}
|
||||
|
||||
@Test
|
||||
void thePrimaryMayNotForgeAWorkersReply() {
|
||||
// Not a hypothetical nicety: a forged reply would resolve the rendezvous the primary is
|
||||
// itself blocked on, corrupting the correlation between a turn and its answer.
|
||||
assertFalse(Authz.permits(PRIMARY, REPLY, "term_a"));
|
||||
assertFalse(Authz.permits(PRIMARY, ASK, "term_a"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWorkerWithNoTargetCannotReply() {
|
||||
assertFalse(Authz.permits(WORKER_A, REPLY, null),
|
||||
"an absent session id must not satisfy the own-session rule");
|
||||
}
|
||||
|
||||
@Test
|
||||
void observationIsOpenToBothAuthenticatedRoles() {
|
||||
assertTrue(Authz.permits(PRIMARY, READ, null));
|
||||
assertTrue(Authz.permits(WORKER_A, READ, null));
|
||||
assertTrue(Authz.permits(PRIMARY, METRICS, null));
|
||||
assertTrue(Authz.permits(WORKER_A, METRICS, null));
|
||||
}
|
||||
|
||||
@Test
|
||||
void unauthenticatedIsDistinguishedFromMerelyForbidden() {
|
||||
// Drives the 401-vs-403 split: a missing credential is fixable by the caller, a wrong role
|
||||
// is not.
|
||||
assertTrue(Authz.isUnauthenticated(ANON));
|
||||
assertFalse(Authz.isUnauthenticated(WORKER_A));
|
||||
assertFalse(Authz.isUnauthenticated(PRIMARY));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,106 @@
|
||||
package dev.ltms.bridged.auth;
|
||||
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.PaneLocator;
|
||||
import dev.ltms.bridged.mcp.ConnectionIdentity;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* CB-501. The behaviour under test is the inversion of the pre-CB-501 default: failing every
|
||||
* identity check must yield {@link Role#ANONYMOUS}, not {@code PRIMARY}.
|
||||
*/
|
||||
class CallerResolverTest {
|
||||
|
||||
private final FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
/** Identity resolving the canned worker pane, keyed off a faked peer-PID lookup. */
|
||||
private ConnectionIdentity identity(long pid) {
|
||||
return new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
|
||||
}
|
||||
|
||||
/** A PID that owns a worker pane in the fake. */
|
||||
private ConnectionIdentity workerIdentity() {
|
||||
return identity(FakeHerdr.WORKER_PID);
|
||||
}
|
||||
|
||||
/** A PID that owns no pane — i.e. the primary, or any other local process. */
|
||||
private ConnectionIdentity nonWorkerIdentity() {
|
||||
return identity(999_999);
|
||||
}
|
||||
|
||||
@Test
|
||||
void aLoopbackWorkerPaneResolvesToWorkerRegardlessOfAuthMode() {
|
||||
Principal underTrust = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
|
||||
Principal underToken = new CallerResolver(workerIdentity(), true, "s3cret")
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.WORKER, underTrust.role());
|
||||
assertEquals("term_a", underTrust.terminal());
|
||||
assertEquals(Role.WORKER, underToken.role(),
|
||||
"worker identity is unforgeable and must never be token-gated — otherwise enabling "
|
||||
+ "auth would lock the whole fleet out of bridge_reply");
|
||||
assertEquals("term_a", underToken.terminal());
|
||||
}
|
||||
|
||||
@Test
|
||||
void loopbackTrustTreatsANonWorkerLoopbackCallerAsThePrimary() {
|
||||
Principal p = new CallerResolver(nonWorkerIdentity()).resolve("127.0.0.1", 99, null);
|
||||
|
||||
assertEquals(Role.PRIMARY, p.role(), "the historical behaviour, now an explicit choice");
|
||||
}
|
||||
|
||||
@Test
|
||||
void tokenModeRefusesANonWorkerCallerThatPresentsNoToken() {
|
||||
Principal p = new CallerResolver(nonWorkerIdentity(), true, "s3cret")
|
||||
.resolve("127.0.0.1", 99, null);
|
||||
|
||||
assertEquals(Role.ANONYMOUS, p.role(),
|
||||
"no credential must mean NOTHING, not the most privileged role on the bus");
|
||||
}
|
||||
|
||||
@Test
|
||||
void tokenModeAcceptsAValidBearerTokenAsThePrimary() {
|
||||
Principal p = new CallerResolver(nonWorkerIdentity(), true, "s3cret")
|
||||
.resolve("127.0.0.1", 99, "Bearer s3cret");
|
||||
|
||||
assertEquals(Role.PRIMARY, p.role());
|
||||
}
|
||||
|
||||
@Test
|
||||
void tokenModeRejectsAWrongOrMalformedCredential() {
|
||||
CallerResolver r = new CallerResolver(nonWorkerIdentity(), true, "s3cret");
|
||||
|
||||
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "Bearer wrong").role());
|
||||
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "s3cret").role(), "scheme required");
|
||||
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "Bearer ").role(), "empty credential");
|
||||
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "Basic s3cret").role(), "wrong scheme");
|
||||
}
|
||||
|
||||
@Test
|
||||
void theBearerSchemeIsCaseInsensitivePerRfc7235() {
|
||||
CallerResolver r = new CallerResolver(nonWorkerIdentity(), true, "s3cret");
|
||||
|
||||
assertEquals(Role.PRIMARY, r.resolve("127.0.0.1", 99, "bearer s3cret").role());
|
||||
assertEquals(Role.PRIMARY, r.resolve("127.0.0.1", 99, "BEARER s3cret").role());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aNonLoopbackCallerIsNeverThePrimaryUnderLoopbackTrust() {
|
||||
// Defence in depth: startup already refuses this pairing (validateAuthExposure), but if a
|
||||
// proxy ever forwards a remote peer onto the loopback listener, the resolver must not
|
||||
// hand it the primary role.
|
||||
Principal p = new CallerResolver(nonWorkerIdentity()).resolve("10.0.0.7", 99, null);
|
||||
|
||||
assertEquals(Role.ANONYMOUS, p.role());
|
||||
}
|
||||
|
||||
@Test
|
||||
void tokenModeRequiresANonEmptyConfiguredToken() {
|
||||
ConnectionIdentity id = nonWorkerIdentity();
|
||||
|
||||
assertThrows(IllegalArgumentException.class, () -> new CallerResolver(id, true, null));
|
||||
assertThrows(IllegalArgumentException.class, () -> new CallerResolver(id, true, " "));
|
||||
}
|
||||
}
|
||||
@@ -90,4 +90,279 @@ class BridgedConfigTest {
|
||||
Files.writeString(f, "bind:\n port: 8080\nfutureFeature:\n enabled: true\n");
|
||||
assertDoesNotThrow(() -> BridgedConfig.load(f));
|
||||
}
|
||||
|
||||
@Test
|
||||
void absentBrokerBlockLeavesInboxSoftState(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("no-broker.yaml");
|
||||
Files.writeString(f, "bind:\n port: 8080\n");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertNull(cfg.broker(), "no broker: block → null → in-memory inbox is selected");
|
||||
}
|
||||
|
||||
@Test
|
||||
void brokerBlockWithUriEnablesAmqpAdapter(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("broker.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
broker:
|
||||
uri: amqp://guest:guest@127.0.0.1:5672/
|
||||
""");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertNotNull(cfg.broker());
|
||||
assertTrue(cfg.broker().isConfigured(), "a non-blank uri enables the AMQP adapter");
|
||||
assertEquals("amqp://guest:guest@127.0.0.1:5672/", cfg.broker().uri());
|
||||
}
|
||||
|
||||
@Test
|
||||
void brokerBlockWithBlankUriStaysSoftState(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("broker-blank.yaml");
|
||||
Files.writeString(f, "bind:\n port: 8080\nbroker:\n uri: \"\"\n");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertNotNull(cfg.broker());
|
||||
assertFalse(cfg.broker().isConfigured(), "an empty uri must not enable AMQP");
|
||||
}
|
||||
|
||||
@Test
|
||||
void absentPrimaryBlockLeavesPrimaryNull(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("no-primary.yaml");
|
||||
Files.writeString(f, "bind:\n port: 8080\n");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertNull(cfg.primary(), "no primary: block → null → connection-derived identity");
|
||||
}
|
||||
|
||||
@Test
|
||||
void primaryBlockWithTerminalPinsIdentity(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("primary-pinned.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
primary:
|
||||
terminal: term_fixed
|
||||
""");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertNotNull(cfg.primary());
|
||||
assertEquals("term_fixed", cfg.primary().terminal());
|
||||
}
|
||||
|
||||
@Test
|
||||
void primaryBlockWithBlankTerminalDefaultsToDerived(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("primary-blank.yaml");
|
||||
Files.writeString(f, "bind:\n port: 8080\nprimary:\n terminal: \"\"\n");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertNotNull(cfg.primary());
|
||||
assertTrue(cfg.primary().terminal() == null || cfg.primary().terminal().isBlank(),
|
||||
"a blank terminal in yaml should be treated as absent — null or empty are equivalent");
|
||||
}
|
||||
|
||||
// --- CB-402: peer kind discriminator -------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void workerKindDefaultsToClaudeCodeWhenOmitted(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("kind-absent.yaml");
|
||||
Files.writeString(f, """
|
||||
workers:
|
||||
gx10:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
argv: ["ccs", "gx10"]
|
||||
""");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertEquals(BridgedConfig.Worker.KIND_CLAUDE_CODE, cfg.workerProfiles().get("gx10").kind(),
|
||||
"a worker with no kind: is a claude-code worker (backward compatible)");
|
||||
}
|
||||
|
||||
@Test
|
||||
void opencodeKindIsNormalizedToLowerCase(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("kind-opencode.yaml");
|
||||
Files.writeString(f, """
|
||||
workers:
|
||||
gemini:
|
||||
kind: OpenCode
|
||||
model: google/gemini-2.5-pro
|
||||
argv: ["opencode"]
|
||||
""");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertEquals(BridgedConfig.Worker.KIND_OPENCODE, cfg.workerProfiles().get("gemini").kind(),
|
||||
"kind is normalised to lower-case so YAML casing does not matter");
|
||||
}
|
||||
|
||||
@Test
|
||||
void kindPredicatesReflectTheResolvedKind(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("kind-predicates.yaml");
|
||||
Files.writeString(f, """
|
||||
workers:
|
||||
claude:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
gemini:
|
||||
kind: opencode
|
||||
model: google/gemini-2.5-pro
|
||||
""");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
BridgedConfig.Worker claude = cfg.workerProfiles().get("claude");
|
||||
BridgedConfig.Worker gemini = cfg.workerProfiles().get("gemini");
|
||||
assertTrue(claude.isClaudeCode(), "the default-kind worker is claude-code");
|
||||
assertFalse(claude.isOpenCode(), "a claude-code worker is not opencode");
|
||||
assertTrue(gemini.isOpenCode(), "the kind: opencode worker is opencode");
|
||||
assertFalse(gemini.isClaudeCode(), "an opencode worker is not claude-code");
|
||||
}
|
||||
|
||||
@Test
|
||||
void argvDefaultsToTheKindBinaryWhenUnset(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("kind-argv.yaml");
|
||||
Files.writeString(f, """
|
||||
workers:
|
||||
claude:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
gemini:
|
||||
kind: opencode
|
||||
model: google/gemini-2.5-pro
|
||||
""");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertEquals(java.util.List.of("claude"), cfg.workerProfiles().get("claude").argv(),
|
||||
"a claude-code worker with no argv defaults to the claude binary");
|
||||
assertEquals(java.util.List.of("opencode"), cfg.workerProfiles().get("gemini").argv(),
|
||||
"an opencode worker with no argv defaults to the opencode binary, never claude");
|
||||
}
|
||||
|
||||
@Test
|
||||
void authDefaultsToLoopbackTrustSoExistingConfigsBehaveAsBefore(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("no-auth-block.yaml");
|
||||
Files.writeString(f, "bind:\n host: 127.0.0.1\n port: 8765\n");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertNotNull(cfg.auth(), "auth must default rather than be null");
|
||||
assertFalse(cfg.auth().tokenMode());
|
||||
assertEquals("BRIDGED_API_TOKEN", cfg.auth().tokenEnv(), "documented default env var");
|
||||
assertDoesNotThrow(cfg::validateAuthExposure, "loopback + loopback-trust is the safe pairing");
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-501's highest-value check. Under loopback-trust, "not a known worker" means "the primary" —
|
||||
* sound only while the OS refuses remote connections. Widening the bind without token mode
|
||||
* would silently promote every reachable client to the most privileged role on the bus.
|
||||
*/
|
||||
@Test
|
||||
void aNonLoopbackBindWithoutTokenModeIsRefusedAtStartup(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("exposed.yaml");
|
||||
Files.writeString(f, "bind:\n host: 0.0.0.0\n port: 8765\n");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAuthExposure);
|
||||
assertTrue(e.getMessage().contains("auth.mode: token"),
|
||||
"the error must say how to fix it, not just that it refused");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aNonLoopbackBindIsAllowedOnceTokenModeIsOn(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("exposed-with-token.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 0.0.0.0
|
||||
port: 8765
|
||||
auth:
|
||||
mode: token
|
||||
tokenEnv: MY_TOKEN
|
||||
""");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertTrue(cfg.auth().tokenMode());
|
||||
assertEquals("MY_TOKEN", cfg.auth().tokenEnv());
|
||||
assertDoesNotThrow(cfg::validateAuthExposure);
|
||||
}
|
||||
|
||||
@Test
|
||||
void loopbackFormsAreAllRecognised(@TempDir Path dir) throws Exception {
|
||||
for (String host : new String[]{"127.0.0.1", "localhost", "::1", "127.0.0.53"}) {
|
||||
Path f = dir.resolve("lb-" + host.replace(':', '_') + ".yaml");
|
||||
Files.writeString(f, "bind:\n host: \"" + host + "\"\n port: 8765\n");
|
||||
assertDoesNotThrow(() -> BridgedConfig.load(f).validateAuthExposure(),
|
||||
host + " is loopback and must not trip the exposure guard");
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The shipped {@code bridged.example.yaml} must actually parse. Config binds through a plain
|
||||
* Jackson mapper with {@code ignoreUnknown = true}, so a misspelled key in the example is
|
||||
* silently dropped and the operator gets a default they did not ask for — exactly how a
|
||||
* {@code spawn_ready_timeout_ms} typo survived in the example until the CB-5xx wrap-up.
|
||||
*/
|
||||
@Test
|
||||
void shippedExampleConfigParses() {
|
||||
Path example = Path.of("bridged.example.yaml");
|
||||
assertTrue(Files.exists(example), "bridged.example.yaml must ship next to the pom");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(example);
|
||||
assertEquals(8765, cfg.bind().port(), "example binds the documented default port");
|
||||
assertTrue(cfg.workerProfiles().containsKey("gx10"), "example documents the gx10 profile");
|
||||
assertEquals("gx10", cfg.defaultProfile(), "example's defaultWorker resolves");
|
||||
assertTrue(cfg.guard().hostSet().contains("gx01.gw"),
|
||||
"every example profile's base_url host must be in the example allowlist");
|
||||
}
|
||||
|
||||
/**
|
||||
* Every optional knob the example documents must bind under the exact spelling used there.
|
||||
* Keep this list in step with {@code bridged.example.yaml}: a rename that updates the record
|
||||
* but not the example (or vice versa) fails here instead of silently no-op'ing in production.
|
||||
*/
|
||||
@Test
|
||||
void everyOptionalKnobDocumentedInTheExampleBinds(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("all-knobs.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
spawnReadyTimeoutMs: 25000
|
||||
spawnReadyPollMs: 400
|
||||
worktreeRoot: /tmp/bridged-worktrees
|
||||
workers:
|
||||
gx10:
|
||||
kind: claude-code
|
||||
baseUrl: http://gx01.gw:8000
|
||||
configDir: /tmp/ccs/gx10
|
||||
cwd: /tmp/repo
|
||||
parityOverlay: [".mcp.json", ".env"]
|
||||
gitTokenEnv: GITEA_TOKEN
|
||||
gitHostEnv: GITEA_HOST
|
||||
lifecycle:
|
||||
idleTtlSeconds: 300
|
||||
contextCap: 10
|
||||
drainTimeoutSeconds: 5
|
||||
broker:
|
||||
uri: amqp://guest:guest@127.0.0.1:5672
|
||||
primary:
|
||||
terminal: term_abc123
|
||||
pushReminders: 5
|
||||
pushBackoffMs: 15000
|
||||
""");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertEquals(25000, cfg.spawnReadyTimeoutMs(), "spawnReadyTimeoutMs is camelCase, not snake_case");
|
||||
assertEquals(400, cfg.spawnReadyPollMs(), "spawnReadyPollMs is camelCase, not snake_case");
|
||||
assertEquals("/tmp/bridged-worktrees", cfg.worktreeRoot());
|
||||
|
||||
BridgedConfig.Worker w = cfg.workerProfiles().get("gx10");
|
||||
assertEquals("/tmp/ccs/gx10", w.configDir());
|
||||
assertEquals("/tmp/repo", w.cwd());
|
||||
assertEquals(java.util.List.of(".mcp.json", ".env"), w.parityOverlay());
|
||||
assertTrue(w.hasGitToken(), "gitTokenEnv binds and enables the CB-302 PR grant");
|
||||
assertEquals("GITEA_HOST", w.gitHostEnv());
|
||||
|
||||
assertEquals(300, cfg.lifecycle().idleTtlSeconds());
|
||||
assertEquals(10, cfg.lifecycle().contextCap());
|
||||
assertEquals(5, cfg.lifecycle().drainTimeoutSeconds());
|
||||
assertEquals("amqp://guest:guest@127.0.0.1:5672", cfg.broker().uri());
|
||||
assertEquals("term_abc123", cfg.primary().terminal());
|
||||
assertEquals(5, cfg.primary().remindersOrDefault());
|
||||
assertEquals(15000L, cfg.primary().backoffMsOrDefault());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -195,6 +195,66 @@ class CompletionResolverTest {
|
||||
assertEquals("hello", waiter.getNow(null).text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent() {
|
||||
// The most important branch of the CB-115 guard: a failed read means the resolver could not
|
||||
// SEE the screen — "couldn't see", not "no change". It must still resolve the send (an empty
|
||||
// tail beats hanging until the caller's timeout), even though a baseline was captured. The
|
||||
// baseline here is "" (an empty pane at delivery), so without the !scrapeFailed clause the
|
||||
// byte-identical guard would wrongly match the empty tail and suppress.
|
||||
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.read throws HerdrException
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
var turn = new CompletionResolver.InFlight(waiter, ""); // empty pane baselined at delivery
|
||||
resolver.resolve("term_a", turn);
|
||||
|
||||
assertTrue(waiter.isDone(),
|
||||
"a failed scrape must still resolve the send, not hang until the caller's timeout");
|
||||
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
|
||||
assertEquals("", waiter.getNow(null).text(), "the tail is empty because the screen was unreadable");
|
||||
}
|
||||
|
||||
// --- CB-115/CB-116 fail guard: an already-done or absent waiter is left alone ---------
|
||||
|
||||
@Test
|
||||
void failLeavesAnAlreadyResolvedWaiterUntouchedAndSkipsTheScrape() {
|
||||
// The send was already resolved (e.g. by the worker's explicit reply) before fail fired.
|
||||
// fail must not overwrite that value, and must not even scrape the worker — nobody needs it.
|
||||
FakeHerdr herdr = new FakeHerdr().readText("an error screen");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
var turn = new CompletionResolver.InFlight(waiter, null);
|
||||
assertTrue(rendezvous.resolveCompletion(waiter, "already replied"));
|
||||
|
||||
resolver.fail("term_a", turn);
|
||||
|
||||
assertFalse(herdr.called("agent.read"),
|
||||
"fail must not scrape a waiter that is already done");
|
||||
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
|
||||
"fail must not overwrite the existing resolution");
|
||||
assertEquals("already replied", waiter.getNow(null).text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void failFallsBackToTheRegisteredWaiterWhenThereIsNoInFlightTurn() {
|
||||
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
|
||||
// fail falls back to the waiter currently registered on the Rendezvous and fails it.
|
||||
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
|
||||
var waiter = rendezvous.open("term_a"); // send registered, but no captureBaseline ever ran
|
||||
resolver.fail("term_a", null); // no in-flight turn → fall back to the registered waiter
|
||||
|
||||
assertTrue(waiter.isDone(), "fail falls back to the registered waiter when no turn is in flight");
|
||||
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
|
||||
assertEquals("stuck on an error screen", waiter.getNow(null).text());
|
||||
}
|
||||
|
||||
// --- CB-116 waiter identity: a late completion never crosses into the next turn ---------
|
||||
|
||||
@Test
|
||||
|
||||
@@ -0,0 +1,168 @@
|
||||
package dev.ltms.bridged.mcp;
|
||||
|
||||
import dev.ltms.bridged.auth.Authz;
|
||||
import dev.ltms.bridged.auth.CallerResolver;
|
||||
import dev.ltms.bridged.auth.Principal;
|
||||
import dev.ltms.bridged.auth.Role;
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.PaneLocator;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.inject.Injector;
|
||||
import dev.ltms.bridged.metrics.BridgedMetrics;
|
||||
import dev.ltms.bridged.metrics.Metrics;
|
||||
import dev.ltms.bridged.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.session.FakeWorktrees;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* CB-513 — the CB-505 authorization gate on the <strong>MCP</strong> entry path.
|
||||
*
|
||||
* <p>Why this file exists: CB-505 claimed authorization is "enforced on both entry paths", and it
|
||||
* is — but only REST was ever tested ({@code BridgedAppAuthTest}). Coverage showed
|
||||
* {@code BridgeMcp.deny()}, {@code principal()} and every tool-registration lambda at <em>zero</em>
|
||||
* executed lines, because no test had ever constructed a {@code BridgeMcp} — the existing
|
||||
* {@code BridgeMcpTest} calls only the static handler methods. An unexercised security control is
|
||||
* a claim, not a control.
|
||||
*
|
||||
* <p>These tests construct a real {@code BridgeMcp} (which also exercises the constructor and the
|
||||
* tool wiring) and drive the policy half of the gate directly.
|
||||
*/
|
||||
class BridgeMcpAuthzTest {
|
||||
|
||||
private final FakeHerdr herdr = new FakeHerdr();
|
||||
private final AgentControl agents = new AgentControl(herdr);
|
||||
private Metrics metrics;
|
||||
private BridgeMcp mcp;
|
||||
|
||||
@AfterEach
|
||||
void close() {
|
||||
if (mcp != null) mcp.close();
|
||||
}
|
||||
|
||||
/** A fully wired BridgeMcp on fakes — constructing it is itself part of what is under test. */
|
||||
private BridgeMcp mcp(boolean enforce) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||
_ -> "tok");
|
||||
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
|
||||
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(),
|
||||
new InMemoryReplyInbox());
|
||||
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
|
||||
metrics = BridgedMetrics.create(sessions, new InMemoryReplyInbox());
|
||||
|
||||
mcp = new BridgeMcp(messages, workers, sessions, identity, sessions.asPresence(),
|
||||
new PrimaryRegistry(null),
|
||||
enforce ? new CallerResolver(identity) : null,
|
||||
metrics);
|
||||
return mcp;
|
||||
}
|
||||
|
||||
private static final Principal PRIMARY = Principal.primary(100);
|
||||
private static final Principal WORKER_A = Principal.worker("term_a", 200);
|
||||
private static final Principal ANON = Principal.anonymous();
|
||||
|
||||
// --- the table, enforced on THIS path too ---------------------------------------------------
|
||||
|
||||
@Test
|
||||
void primaryMayOrchestrate() {
|
||||
BridgeMcp m = mcp(true);
|
||||
for (Authz.Action a : new Authz.Action[]{Authz.Action.SPAWN, Authz.Action.STOP,
|
||||
Authz.Action.SEND, Authz.Action.DRAIN, Authz.Action.READ}) {
|
||||
assertNull(m.denyFor(PRIMARY, a, "term_a"), a + " is the primary's to perform");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWorkerMayNotOrchestrateOverMcp() {
|
||||
BridgeMcp m = mcp(true);
|
||||
for (Authz.Action a : new Authz.Action[]{Authz.Action.SPAWN, Authz.Action.STOP,
|
||||
Authz.Action.SEND, Authz.Action.DRAIN}) {
|
||||
McpSchema.CallToolResult denied = m.denyFor(WORKER_A, a, "term_a");
|
||||
assertNotNull(denied, a + " must be refused to a worker");
|
||||
assertTrue(denied.isError(), "a refusal is returned as an MCP tool error");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWorkerMayReplyAndAskOnlyAsItself() {
|
||||
BridgeMcp m = mcp(true);
|
||||
assertNull(m.denyFor(WORKER_A, Authz.Action.REPLY, "term_a"), "its own session is allowed");
|
||||
assertNull(m.denyFor(WORKER_A, Authz.Action.ASK, "term_a"));
|
||||
|
||||
assertNotNull(m.denyFor(WORKER_A, Authz.Action.REPLY, "term_b"),
|
||||
"worker A must not reply on worker B's session");
|
||||
assertNotNull(m.denyFor(WORKER_A, Authz.Action.ASK, "term_b"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void thePrimaryMayNotForgeAWorkerReplyOverMcp() {
|
||||
BridgeMcp m = mcp(true);
|
||||
// A forged reply would resolve the very rendezvous the primary is blocked on.
|
||||
assertNotNull(m.denyFor(PRIMARY, Authz.Action.REPLY, "term_a"));
|
||||
assertNotNull(m.denyFor(PRIMARY, Authz.Action.ASK, "term_a"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void anonymousIsRefusedEverythingAndCountedAsUnauthenticated() {
|
||||
BridgeMcp m = mcp(true);
|
||||
McpSchema.CallToolResult denied = m.denyFor(ANON, Authz.Action.READ, null);
|
||||
|
||||
assertNotNull(denied, "authenticated as nothing ⇒ authorized for nothing");
|
||||
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "unauthenticated"));
|
||||
assertEquals(0, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "forbidden"),
|
||||
"a missing credential is 401-shaped, not 403-shaped");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWrongRoleIsCountedAsForbiddenNotUnauthenticated() {
|
||||
BridgeMcp m = mcp(true);
|
||||
assertNotNull(m.denyFor(WORKER_A, Authz.Action.SPAWN, null));
|
||||
|
||||
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "forbidden"));
|
||||
assertEquals(0, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "unauthenticated"),
|
||||
"the caller IS authenticated — it is just not the right role");
|
||||
}
|
||||
|
||||
@Test
|
||||
void theLegacyConstructorLeavesTheGateOpen() {
|
||||
// The 22 pre-existing BridgeMcpTest cases rely on no authorization being enforced.
|
||||
BridgeMcp m = mcp(false);
|
||||
assertNull(m.denyFor(ANON, Authz.Action.SPAWN, null),
|
||||
"no CallerResolver supplied ⇒ authorization not enforced (legacy behaviour)");
|
||||
}
|
||||
|
||||
// --- identity reconstruction from the transport context ------------------------------------
|
||||
|
||||
@Test
|
||||
void principalIsRebuiltFromTheStashedRole() {
|
||||
assertEquals(Role.WORKER, BridgeMcp.principalFrom("WORKER", "term_a", 7).role());
|
||||
assertEquals("term_a", BridgeMcp.principalFrom("WORKER", "term_a", 7).terminal());
|
||||
assertEquals(Role.PRIMARY, BridgeMcp.principalFrom("PRIMARY", null, 7).role());
|
||||
assertEquals(Role.ANONYMOUS, BridgeMcp.principalFrom("ANONYMOUS", null, -1).role());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aMissingRoleFallsBackToTheHistoricalInterpretation() {
|
||||
// Legacy path: no role stashed. A terminal means worker; its absence meant "the primary",
|
||||
// which is exactly the pre-CB-501 default CB-501 inverted — preserved only here.
|
||||
assertEquals(Role.WORKER, BridgeMcp.principalFrom(null, "term_a", 7).role());
|
||||
assertEquals(Role.PRIMARY, BridgeMcp.principalFrom(null, null, 7).role());
|
||||
}
|
||||
}
|
||||
@@ -1,5 +1,6 @@
|
||||
package dev.ltms.bridged.mcp;
|
||||
|
||||
import dev.ltms.bridged.auth.Principal;
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
@@ -57,13 +58,16 @@ class BridgeMcpTest {
|
||||
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
|
||||
() -> BridgeMcp.send(messages, "term_a", "review this", 4000L));
|
||||
|
||||
McpSchema.CallToolResult reply = BridgeMcp.reply(rendezvous, "term_a", "LGTM");
|
||||
// Wait until the send has opened its waiter so the reply resolves it (CB-307: reply now
|
||||
// queues in the inbox if no waiter is open, which would break the round-trip).
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (Boolean.TRUE.equals(reply.isError()) && System.currentTimeMillis() < deadline) {
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
reply = BridgeMcp.reply(rendezvous, "term_a", "LGTM");
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"), "send should have opened its waiter");
|
||||
|
||||
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "LGTM");
|
||||
assertEquals("delivered", textOf(reply));
|
||||
|
||||
McpSchema.CallToolResult res = send.get(6, TimeUnit.SECONDS);
|
||||
@@ -80,30 +84,32 @@ class BridgeMcpTest {
|
||||
assertTrue(out.contains("ticket="), out);
|
||||
String ticket = out.substring(out.indexOf("ticket=") + "ticket=".length()).trim();
|
||||
|
||||
// Resolve the awaiting send once it has opened (retry past the async-open race).
|
||||
// Wait until the send has opened its waiter before replying (CB-307: reply never errors,
|
||||
// so the old retry-on-error pattern no longer works — it would queue instead of resolve).
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
McpSchema.CallToolResult reply = BridgeMcp.reply(rendezvous, "term_a", "async LGTM");
|
||||
while (Boolean.TRUE.equals(reply.isError()) && System.currentTimeMillis() < deadline) {
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
reply = BridgeMcp.reply(rendezvous, "term_a", "async LGTM");
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"), "send should have opened its waiter");
|
||||
|
||||
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "async LGTM");
|
||||
assertEquals("delivered", textOf(reply));
|
||||
|
||||
// Poll until the async send completes and reports the reply.
|
||||
McpSchema.CallToolResult polled = BridgeMcp.poll(messages, ticket);
|
||||
McpSchema.CallToolResult polled = BridgeMcp.poll(messages, ticket, null);
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (!textOf(polled).contains("async LGTM") && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
polled = BridgeMcp.poll(messages, ticket);
|
||||
polled = BridgeMcp.poll(messages, ticket, null);
|
||||
}
|
||||
assertEquals("async LGTM", textOf(polled));
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollUnknownTicketIsAnError() {
|
||||
McpSchema.CallToolResult res = BridgeMcp.poll(messages, "task-999");
|
||||
McpSchema.CallToolResult res = BridgeMcp.poll(messages, "task-999", null);
|
||||
assertTrue(res.isError());
|
||||
assertTrue(textOf(res).contains("unknown ticket"));
|
||||
}
|
||||
@@ -122,10 +128,32 @@ class BridgeMcpTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void replyWithNoPendingSendIsAnError() {
|
||||
McpSchema.CallToolResult res = BridgeMcp.reply(rendezvous, "term_a", "orphan");
|
||||
assertTrue(res.isError());
|
||||
assertTrue(textOf(res).contains("no send is awaiting"));
|
||||
void replyWithNoPendingSendIsQueuedNotError() {
|
||||
// CB-307: a reply with no open send is now queued in the inbox, not an error.
|
||||
McpSchema.CallToolResult res = BridgeMcp.reply(messages, "term_a", "orphan");
|
||||
assertNotEquals(Boolean.TRUE, res.isError(), "a queued reply is not an error");
|
||||
assertEquals("delivered", textOf(res));
|
||||
|
||||
// The reply is drainable by target.
|
||||
var drained = messages.drainReplies("term_a");
|
||||
assertEquals(1, drained.size());
|
||||
assertEquals("orphan", drained.getFirst().content());
|
||||
}
|
||||
|
||||
@Test
|
||||
void bridgePollWithTargetDrainsReplies() {
|
||||
// A reply with no open send queues it in the inbox.
|
||||
BridgeMcp.reply(messages, "term_a", "queued-msg");
|
||||
|
||||
// bridge_poll with target drains the inbox.
|
||||
McpSchema.CallToolResult res = BridgeMcp.poll(messages, null, "term_a");
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String text = textOf(res);
|
||||
assertTrue(text.contains("queued-msg"), "the drained reply should appear in the result");
|
||||
|
||||
// Second drain returns empty.
|
||||
McpSchema.CallToolResult empty = BridgeMcp.poll(messages, null, "term_a");
|
||||
assertEquals("[]", textOf(empty));
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -159,14 +187,14 @@ class BridgeMcpTest {
|
||||
// The worker's ask returns the answer — it resumes the same turn.
|
||||
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
|
||||
|
||||
// The resumed worker replies, resolving the answering send (retry past the reopen race).
|
||||
McpSchema.CallToolResult reply = BridgeMcp.reply(rendezvous, "term_a", "done");
|
||||
// The resumed worker replies, resolving the answering send (wait for the reopened waiter).
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (Boolean.TRUE.equals(reply.isError()) && System.currentTimeMillis() < deadline) {
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
reply = BridgeMcp.reply(rendezvous, "term_a", "done");
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"), "the answer should have reopened a waiter");
|
||||
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "done");
|
||||
assertEquals("delivered", textOf(reply));
|
||||
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
|
||||
}
|
||||
@@ -276,6 +304,37 @@ class BridgeMcpTest {
|
||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), " ").isError());
|
||||
}
|
||||
|
||||
@Test
|
||||
void bridgeAckReturnsConfirmationForValidArgs() {
|
||||
McpSchema.CallToolResult res = BridgeMcp.ack(messages, "term_a", "msg-1");
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
assertTrue(textOf(res).contains("msg-1"), "response should mention the msgId");
|
||||
}
|
||||
|
||||
@Test
|
||||
void bridgeAckRejectsMissingArgs() {
|
||||
assertTrue(BridgeMcp.ack(messages, null, "msg-1").isError());
|
||||
assertTrue(BridgeMcp.ack(messages, "term_a", null).isError());
|
||||
assertTrue(BridgeMcp.ack(messages, " ", "msg-1").isError());
|
||||
}
|
||||
|
||||
@Test
|
||||
void bridgeAckRemovesSpecificReply() {
|
||||
// Queue a reply and capture its msgId.
|
||||
BridgeMcp.reply(messages, "term_a", "orphan");
|
||||
var before = messages.drainReplies("term_a");
|
||||
assertEquals(1, before.size(), "one reply in the inbox");
|
||||
String msgId = before.getFirst().msgId();
|
||||
|
||||
// Publish the same reply again and ack it via bridge_ack surface.
|
||||
BridgeMcp.reply(messages, "term_a", "orphan-again");
|
||||
var peeked = messages.drainReplies("term_a");
|
||||
assertEquals(1, peeked.size(), "one fresh reply in the inbox");
|
||||
|
||||
// ackReply works (no-op since published with a different UUID, but callable).
|
||||
assertDoesNotThrow(() -> messages.ackReply("term_a", msgId));
|
||||
}
|
||||
|
||||
@Test
|
||||
void statusReportsLiveAgentStatus() {
|
||||
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
|
||||
@@ -285,4 +344,58 @@ class BridgeMcpTest {
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
assertEquals("blocked", textOf(res));
|
||||
}
|
||||
|
||||
// --- bridge_whoami: the caller's own identity, so an agent never has to guess its role -------
|
||||
|
||||
@Test
|
||||
void whoamiReportsThePrimaryAsPrimaryAndNothingElse() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res = BridgeMcp.whoami(
|
||||
Principal.primary(100), sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"role\":\"primary\""), out);
|
||||
// The primary owns no session — leaking a sessionId here would invite it to reply as one.
|
||||
assertFalse(out.contains("sessionId"), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void whoamiReportsAWorkerWithItsRegisteredSession() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
|
||||
new WorktreeRequest("cb-517", null));
|
||||
|
||||
McpSchema.CallToolResult res = BridgeMcp.whoami(Principal.worker(s.terminalId(), 200), sessions);
|
||||
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"role\":\"worker\""), out);
|
||||
assertTrue(out.contains("\"sessionId\":\"" + s.terminalId() + "\""), out);
|
||||
assertTrue(out.contains("\"profile\":\"ltms-local\""), out);
|
||||
assertTrue(out.contains("\"worktree\":\"" + s.worktree() + "\""), out);
|
||||
assertTrue(out.contains("\"branch\":\"" + s.branch() + "\""), out);
|
||||
assertTrue(out.contains("\"owner\":\"term_primary\""), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* A worker the registry has no record of — it outlived a daemon restart — must still learn the
|
||||
* load-bearing fact. Degrading to "I don't know who you are" would put it back to guessing,
|
||||
* which is the failure this tool exists to remove.
|
||||
*/
|
||||
@Test
|
||||
void whoamiStillReportsWorkerRoleWhenTheSessionIsUnregistered() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res = BridgeMcp.whoami(Principal.worker("term_orphan", 200),
|
||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"role\":\"worker\""), out);
|
||||
assertTrue(out.contains("\"sessionId\":\"term_orphan\""), out);
|
||||
assertFalse(out.contains("profile"), out); // nothing invented for a session we don't track
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,108 @@
|
||||
package dev.ltms.bridged.mcp;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* Unit tests for {@link PrimaryRegistry}: pin vs record, isKnown transitions,
|
||||
* null/blank guard.
|
||||
*/
|
||||
class PrimaryRegistryTest {
|
||||
|
||||
@Test
|
||||
void unpinnedInitiallyUnknown() {
|
||||
var reg = new PrimaryRegistry(null);
|
||||
assertFalse(reg.isKnown());
|
||||
assertTrue(reg.primaryTerminal().isEmpty());
|
||||
}
|
||||
|
||||
@Test
|
||||
void unpinnedAcceptsBlankAsAbsent() {
|
||||
var reg = new PrimaryRegistry("");
|
||||
assertFalse(reg.isKnown());
|
||||
assertTrue(reg.primaryTerminal().isEmpty());
|
||||
}
|
||||
|
||||
@Test
|
||||
void pinnedFromConstruction() {
|
||||
var reg = new PrimaryRegistry("term_fixed");
|
||||
assertTrue(reg.isKnown());
|
||||
assertEquals("term_fixed", reg.primaryTerminal().orElseThrow());
|
||||
}
|
||||
|
||||
@Test
|
||||
void recordWhenUnpinnedSetsTheTerminal() {
|
||||
var reg = new PrimaryRegistry(null);
|
||||
reg.record("term_abc");
|
||||
assertTrue(reg.isKnown());
|
||||
assertEquals("term_abc", reg.primaryTerminal().orElseThrow());
|
||||
}
|
||||
|
||||
@Test
|
||||
void recordWithNullDoesNothingWhenUnpinned() {
|
||||
var reg = new PrimaryRegistry(null);
|
||||
reg.record(null);
|
||||
assertFalse(reg.isKnown());
|
||||
}
|
||||
|
||||
@Test
|
||||
void recordWithBlankDoesNothingWhenUnpinned() {
|
||||
var reg = new PrimaryRegistry(null);
|
||||
reg.record(" ");
|
||||
assertFalse(reg.isKnown());
|
||||
}
|
||||
|
||||
@Test
|
||||
void recordOverwritesWhenUnpinned() {
|
||||
var reg = new PrimaryRegistry(null);
|
||||
reg.record("term_first");
|
||||
assertEquals("term_first", reg.primaryTerminal().orElseThrow());
|
||||
reg.record("term_second");
|
||||
assertEquals("term_second", reg.primaryTerminal().orElseThrow());
|
||||
}
|
||||
|
||||
@Test
|
||||
void recordIsIgnoredWhenPinned() {
|
||||
var reg = new PrimaryRegistry("term_pinned");
|
||||
reg.record("term_other");
|
||||
assertEquals("term_pinned", reg.primaryTerminal().orElseThrow(), "pinned value must survive record");
|
||||
}
|
||||
|
||||
@Test
|
||||
void nullRecordIsIgnoredWhenPinned() {
|
||||
var reg = new PrimaryRegistry("term_pinned");
|
||||
reg.record(null);
|
||||
assertTrue(reg.isKnown());
|
||||
assertEquals("term_pinned", reg.primaryTerminal().orElseThrow());
|
||||
}
|
||||
|
||||
@Test
|
||||
void blankRecordIsIgnoredWhenPinned() {
|
||||
var reg = new PrimaryRegistry("term_pinned");
|
||||
reg.record(" ");
|
||||
assertTrue(reg.isKnown());
|
||||
assertEquals("term_pinned", reg.primaryTerminal().orElseThrow());
|
||||
}
|
||||
|
||||
@Test
|
||||
void isKnownFalseAfterConstructionWithNull() {
|
||||
var reg = new PrimaryRegistry(null);
|
||||
assertFalse(reg.isKnown());
|
||||
}
|
||||
|
||||
@Test
|
||||
void isKnownAfterRecord() {
|
||||
var reg = new PrimaryRegistry(null);
|
||||
reg.record("term_x");
|
||||
assertTrue(reg.isKnown());
|
||||
}
|
||||
|
||||
@Test
|
||||
void primaryTerminalRoundTrip() {
|
||||
var reg = new PrimaryRegistry(null);
|
||||
assertTrue(reg.primaryTerminal().isEmpty());
|
||||
reg.record("term_found");
|
||||
assertEquals("term_found", reg.primaryTerminal().get());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,127 @@
|
||||
package dev.ltms.bridged.metrics;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/** CB-502 — the zero-dependency Prometheus text renderer. */
|
||||
class MetricsTest {
|
||||
|
||||
@Test
|
||||
void countersAccumulatePerLabelSet() {
|
||||
Metrics m = new Metrics();
|
||||
m.inc("bridged_sends_total", "outcome", "replied");
|
||||
m.inc("bridged_sends_total", "outcome", "replied");
|
||||
m.inc("bridged_sends_total", "outcome", "timeout");
|
||||
|
||||
assertEquals(2, m.count("bridged_sends_total", "outcome", "replied"));
|
||||
assertEquals(1, m.count("bridged_sends_total", "outcome", "timeout"));
|
||||
assertEquals(0, m.count("bridged_sends_total", "outcome", "failed"),
|
||||
"an untouched series reads as zero, not an error");
|
||||
}
|
||||
|
||||
@Test
|
||||
void rendersHelpAndTypeOncePerFamily() {
|
||||
Metrics m = new Metrics();
|
||||
m.describe("bridged_sends_total", "counter", "Delegated sends by outcome.");
|
||||
m.inc("bridged_sends_total", "outcome", "replied");
|
||||
m.inc("bridged_sends_total", "outcome", "timeout");
|
||||
|
||||
String out = m.render();
|
||||
assertEquals(1, countOccurrences(out, "# HELP bridged_sends_total"),
|
||||
"HELP is per family, not per series");
|
||||
assertEquals(1, countOccurrences(out, "# TYPE bridged_sends_total counter"));
|
||||
assertTrue(out.contains("bridged_sends_total{outcome=\"replied\"} 1"));
|
||||
assertTrue(out.contains("bridged_sends_total{outcome=\"timeout\"} 1"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void labelsAreSortedSoScrapesAreByteStable() {
|
||||
Metrics a = new Metrics();
|
||||
a.inc("m", "b", "2", "a", "1");
|
||||
Metrics b = new Metrics();
|
||||
b.inc("m", "a", "1", "b", "2");
|
||||
|
||||
assertEquals(a.render(), b.render(), "label order in the call must not change the output");
|
||||
assertTrue(a.render().contains("m{a=\"1\",b=\"2\"}"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void gaugesAreEvaluatedAtScrapeTimeNotRegistrationTime() {
|
||||
Metrics m = new Metrics();
|
||||
int[] live = {1};
|
||||
m.gauge("bridged_sessions", () -> live[0], "state", "ready");
|
||||
|
||||
assertTrue(m.render().contains("bridged_sessions{state=\"ready\"} 1"));
|
||||
live[0] = 5;
|
||||
assertTrue(m.render().contains("bridged_sessions{state=\"ready\"} 5"),
|
||||
"the gauge must read current state on every scrape");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aThrowingGaugeDoesNotBreakTheWholeScrape() {
|
||||
Metrics m = new Metrics();
|
||||
m.inc("good_total");
|
||||
m.gauge("bad_gauge", () -> {
|
||||
throw new IllegalStateException("herdr is down");
|
||||
});
|
||||
|
||||
String out = assertDoesNotThrow(m::render);
|
||||
assertTrue(out.contains("good_total 1"), "healthy series must still be exported");
|
||||
assertFalse(out.contains("bad_gauge"), "the broken series is simply absent");
|
||||
}
|
||||
|
||||
@Test
|
||||
void collectorsDiscoverTheirLabelSetPerScrape() {
|
||||
Metrics m = new Metrics();
|
||||
Map<String, Number> depths = new LinkedHashMap<>();
|
||||
m.collector("bridged_inbox_depth", "target", () -> depths);
|
||||
|
||||
assertFalse(m.render().contains("bridged_inbox_depth"), "no targets yet ⇒ no series");
|
||||
|
||||
depths.put("term_a", 2);
|
||||
depths.put("term_b", 0);
|
||||
String out = m.render();
|
||||
assertTrue(out.contains("bridged_inbox_depth{target=\"term_a\"} 2"));
|
||||
assertTrue(out.contains("bridged_inbox_depth{target=\"term_b\"} 0"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void labelValuesAreEscaped() {
|
||||
Metrics m = new Metrics();
|
||||
m.inc("m", "detail", "he said \"hi\"\nand \\left");
|
||||
|
||||
String out = m.render();
|
||||
assertTrue(out.contains("\\\""), "quotes escaped");
|
||||
assertTrue(out.contains("\\n"), "newlines escaped — a raw one would corrupt the exposition");
|
||||
assertTrue(out.contains("\\\\"), "backslashes escaped");
|
||||
}
|
||||
|
||||
@Test
|
||||
void wholeNumberGaugesRenderWithoutADecimalPoint() {
|
||||
Metrics m = new Metrics();
|
||||
m.gauge("whole", () -> 3.0);
|
||||
m.gauge("fractional", () -> 1.5);
|
||||
|
||||
String out = m.render();
|
||||
assertTrue(out.contains("whole 3"), "3.0 should not render as 3.0");
|
||||
assertTrue(out.contains("fractional 1.5"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void oddLabelCountIsRejected() {
|
||||
Metrics m = new Metrics();
|
||||
assertThrows(IllegalArgumentException.class, () -> m.inc("m", "dangling"));
|
||||
}
|
||||
|
||||
private static int countOccurrences(String haystack, String needle) {
|
||||
int n = 0;
|
||||
for (int i = haystack.indexOf(needle); i >= 0; i = haystack.indexOf(needle, i + 1)) {
|
||||
n++;
|
||||
}
|
||||
return n;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,113 @@
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import org.junit.jupiter.api.Tag;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.testcontainers.containers.RabbitMQContainer;
|
||||
import org.testcontainers.junit.jupiter.Container;
|
||||
import org.testcontainers.junit.jupiter.Testcontainers;
|
||||
import org.testcontainers.utility.DockerImageName;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* Contract test for {@link AmqpReplyInbox} against a REAL broker (a RabbitMQ container — the same
|
||||
* AMQP 0-9-1 the production LavinMQ deploy speaks, URI-only swap). Tagged {@code contract} so it is
|
||||
* excluded from {@code mvn test}/{@code mvn clean install} (which stay hermetic and need no Docker);
|
||||
* run it with Docker present via {@code mvn test -Pcontract}.
|
||||
*
|
||||
* <p>It proves the port contract on genuine infrastructure: eventual visibility of a published reply,
|
||||
* ack removal, msgId dedup, and — the reason Stage 2 exists — cross-restart durability: an unacked
|
||||
* reply survives closing the inbox and is redelivered to a fresh connection.
|
||||
*/
|
||||
@Tag("contract")
|
||||
@Testcontainers
|
||||
class AmqpReplyInboxContractTest {
|
||||
|
||||
@Container
|
||||
static final RabbitMQContainer BROKER =
|
||||
new RabbitMQContainer(DockerImageName.parse("rabbitmq:3.13-management"));
|
||||
|
||||
private static String uri() {
|
||||
// guest/guest against the mapped AMQP port. No trailing slash: an empty path is vhost "",
|
||||
// which does not exist — omitting it selects the default vhost "/".
|
||||
return "amqp://guest:guest@" + BROKER.getHost() + ":" + BROKER.getAmqpPort();
|
||||
}
|
||||
|
||||
@Test
|
||||
void publishThenPeekThenAck() throws Exception {
|
||||
String target = "worker-pub-" + System.nanoTime();
|
||||
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
|
||||
inbox.publish(target, "m1", "hello primary");
|
||||
|
||||
List<ReplyInbox.InboxMessage> got = awaitPeek(inbox, target);
|
||||
assertEquals(1, got.size(), "the published reply should be held for drain");
|
||||
assertEquals("m1", got.getFirst().msgId());
|
||||
assertEquals(target, got.getFirst().target());
|
||||
assertEquals("hello primary", got.getFirst().content());
|
||||
|
||||
inbox.ack(target, "m1");
|
||||
assertTrue(inbox.peek(target).isEmpty(), "an acked reply is dropped");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void duplicateMsgIdIsNotDoubleQueued() throws Exception {
|
||||
String target = "worker-dedup-" + System.nanoTime();
|
||||
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
|
||||
inbox.publish(target, "dup", "first");
|
||||
awaitPeek(inbox, target);
|
||||
inbox.publish(target, "dup", "second"); // same msgId — must be a no-op
|
||||
|
||||
// Give any erroneous second delivery time to land, then assert still exactly one.
|
||||
Thread.sleep(500);
|
||||
List<ReplyInbox.InboxMessage> got = inbox.peek(target);
|
||||
assertEquals(1, got.size(), "a repeated msgId must not double-queue");
|
||||
assertEquals("first", got.getFirst().content(), "the first payload wins");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void unackedReplySurvivesRestartAndIsRedelivered() throws Exception {
|
||||
String target = "worker-durable-" + System.nanoTime();
|
||||
|
||||
// First "process life": publish, see it held, but crash before acking.
|
||||
try (AmqpReplyInbox first = AmqpReplyInbox.open(uri())) {
|
||||
first.publish(target, "persist-1", "survive me");
|
||||
assertEquals(1, awaitPeek(first, target).size());
|
||||
// no ack — simulate a java -jar bounce with the reply still pending
|
||||
}
|
||||
|
||||
// Second "process life": a fresh connection to the same broker must be redelivered the reply.
|
||||
try (AmqpReplyInbox second = AmqpReplyInbox.open(uri())) {
|
||||
List<ReplyInbox.InboxMessage> got = awaitPeek(second, target);
|
||||
assertEquals(1, got.size(), "an unacked persistent reply is redelivered after restart");
|
||||
assertEquals("persist-1", got.getFirst().msgId());
|
||||
assertEquals("survive me", got.getFirst().content());
|
||||
|
||||
second.ack(target, "persist-1");
|
||||
}
|
||||
|
||||
// Third life: once acked, it is gone for good — durability is not endless replay.
|
||||
try (AmqpReplyInbox third = AmqpReplyInbox.open(uri())) {
|
||||
Thread.sleep(500);
|
||||
assertTrue(third.peek(target).isEmpty(), "an acked reply does not come back on the next restart");
|
||||
}
|
||||
}
|
||||
|
||||
/** Poll peek (broker delivery is async) until a reply for {@code target} appears or ~10s elapse. */
|
||||
@SuppressWarnings("BusyWait") // deliberate poll for async broker delivery, bounded by the deadline
|
||||
private static List<ReplyInbox.InboxMessage> awaitPeek(AmqpReplyInbox inbox, String target)
|
||||
throws InterruptedException {
|
||||
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(10);
|
||||
List<ReplyInbox.InboxMessage> msgs = inbox.peek(target);
|
||||
while (msgs.isEmpty() && System.nanoTime() < deadline) {
|
||||
Thread.sleep(50);
|
||||
msgs = inbox.peek(target);
|
||||
}
|
||||
return msgs;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,145 @@
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.concurrent.CountDownLatch;
|
||||
import java.util.concurrent.ExecutorService;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* Unit tests for {@link InMemoryReplyInbox}: publish, peek, ack, dedup, FIFO ordering, and thread
|
||||
* safety under concurrent publish vs. drain.
|
||||
*/
|
||||
class InMemoryReplyInboxTest {
|
||||
|
||||
private final ReplyInbox inbox = new InMemoryReplyInbox();
|
||||
|
||||
@Test
|
||||
void publishThenPeekReturnsTheMessage() {
|
||||
inbox.publish("term_a", "m1", "hello");
|
||||
var msgs = inbox.peek("term_a");
|
||||
assertEquals(1, msgs.size());
|
||||
assertEquals("m1", msgs.getFirst().msgId());
|
||||
assertEquals("term_a", msgs.getFirst().target());
|
||||
assertEquals("hello", msgs.getFirst().content());
|
||||
}
|
||||
|
||||
@Test
|
||||
void peekForUnknownTargetReturnsEmpty() {
|
||||
assertTrue(inbox.peek("no-such-target").isEmpty());
|
||||
}
|
||||
|
||||
@Test
|
||||
void ackRemovesTheMessage() {
|
||||
inbox.publish("term_a", "m1", "hello");
|
||||
inbox.ack("term_a", "m1");
|
||||
assertTrue(inbox.peek("term_a").isEmpty(), "after ack, the message is gone");
|
||||
}
|
||||
|
||||
@Test
|
||||
void ackForUnknownMsgIdIsNoOp() {
|
||||
inbox.publish("term_a", "m1", "hello");
|
||||
inbox.ack("term_a", "no-such-id"); // no-op
|
||||
assertEquals(1, inbox.peek("term_a").size(), "the published message is still there");
|
||||
}
|
||||
|
||||
@Test
|
||||
void ackForUnknownTargetIsNoOp() {
|
||||
inbox.ack("no-such-target", "m1"); // no-op, should not throw
|
||||
}
|
||||
|
||||
@Test
|
||||
void dedupByIdempotentMsgId() {
|
||||
inbox.publish("term_a", "m1", "first");
|
||||
inbox.publish("term_a", "m1", "second"); // same msgId, different content
|
||||
var msgs = inbox.peek("term_a");
|
||||
assertEquals(1, msgs.size(), "dedup: second publish with same msgId is a no-op");
|
||||
assertEquals("first", msgs.getFirst().content(), "the original content is retained");
|
||||
}
|
||||
|
||||
@Test
|
||||
void publishesWithDifferentMsgIdsBothAppear() {
|
||||
inbox.publish("term_a", "m1", "first");
|
||||
inbox.publish("term_a", "m2", "second");
|
||||
var msgs = inbox.peek("term_a");
|
||||
assertEquals(2, msgs.size());
|
||||
assertEquals("m1", msgs.get(0).msgId());
|
||||
assertEquals("m2", msgs.get(1).msgId());
|
||||
}
|
||||
|
||||
@Test
|
||||
void perTargetIsolation() {
|
||||
inbox.publish("term_a", "m1", "for-a");
|
||||
inbox.publish("term_b", "m2", "for-b");
|
||||
assertEquals(1, inbox.peek("term_a").size());
|
||||
assertEquals(1, inbox.peek("term_b").size());
|
||||
}
|
||||
|
||||
@Test
|
||||
void fifoOrderIsPreserved() {
|
||||
inbox.publish("term_a", "m1", "first");
|
||||
inbox.publish("term_a", "m2", "second");
|
||||
inbox.publish("term_a", "m3", "third");
|
||||
var msgs = inbox.peek("term_a");
|
||||
assertEquals(3, msgs.size());
|
||||
assertEquals("m1", msgs.get(0).msgId());
|
||||
assertEquals("m2", msgs.get(1).msgId());
|
||||
assertEquals("m3", msgs.get(2).msgId());
|
||||
}
|
||||
|
||||
@Test
|
||||
void peekReturnsAnImmutableCopy() {
|
||||
inbox.publish("term_a", "m1", "hello");
|
||||
var msgs = inbox.peek("term_a");
|
||||
assertThrows(UnsupportedOperationException.class, () -> msgs.add(
|
||||
new ReplyInbox.InboxMessage("x", "term_a", "x")));
|
||||
}
|
||||
|
||||
@Test
|
||||
void ackRemovesOneMessageLeavesOthers() {
|
||||
inbox.publish("term_a", "m1", "first");
|
||||
inbox.publish("term_a", "m2", "second");
|
||||
inbox.ack("term_a", "m1");
|
||||
var msgs = inbox.peek("term_a");
|
||||
assertEquals(1, msgs.size());
|
||||
assertEquals("m2", msgs.getFirst().msgId());
|
||||
}
|
||||
|
||||
@Test
|
||||
void concurrentPublishAndDrain() throws Exception {
|
||||
int msgCount = 100;
|
||||
ExecutorService exec = Executors.newVirtualThreadPerTaskExecutor();
|
||||
try {
|
||||
// Concurrent publishers
|
||||
var pubDone = new CountDownLatch(msgCount);
|
||||
for (int i = 0; i < msgCount; i++) {
|
||||
final int id = i;
|
||||
exec.submit(() -> {
|
||||
inbox.publish("term_a", "m" + id, "content-" + id);
|
||||
pubDone.countDown();
|
||||
});
|
||||
}
|
||||
// Concurrent drainer
|
||||
AtomicReference<Exception> drainError = new AtomicReference<>();
|
||||
exec.submit(() -> {
|
||||
try {
|
||||
pubDone.await();
|
||||
for (int i = 0; i < 50; i++) {
|
||||
var peeked = inbox.peek("term_a");
|
||||
for (var msg : peeked) {
|
||||
inbox.ack("term_a", msg.msgId());
|
||||
}
|
||||
}
|
||||
} catch (Exception e) {
|
||||
drainError.set(e);
|
||||
}
|
||||
}).get();
|
||||
assertNull(drainError.get(), "concurrent drain should not throw");
|
||||
} finally {
|
||||
exec.shutdown();
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -14,6 +14,7 @@ import java.util.concurrent.TimeUnit;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
@@ -211,4 +212,256 @@ class MessageServiceTest {
|
||||
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
|
||||
"an answer to a turn that never existed (or already lapsed) is stale, not a hang");
|
||||
}
|
||||
|
||||
// --- timeout, answer, poll, and lock-contention edges ----------------------------------
|
||||
|
||||
@Test
|
||||
void sendTimesOutBeforeDeliveryIsQueuedNotWorking() {
|
||||
// Nothing ever delivers the message and nothing resolves the send, so the reply future
|
||||
// times out with delivery still incomplete — the message is still queued for the worker.
|
||||
MessageService.Reply r = messages.send(T, "never delivered", 50);
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome(),
|
||||
"an undelivered send that times out is still queued, not working");
|
||||
assertNull(r.text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void sendTimesOutAfterDeliveryIsStillWorking() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send =
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 300));
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver — the delivered future now completes
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker starts but never replies
|
||||
// No rendezvous.resolve(T, ...) — the reply future rides out its short timeout.
|
||||
|
||||
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, r.outcome(),
|
||||
"a delivered send whose worker never replies times out as still working");
|
||||
}
|
||||
|
||||
@Test
|
||||
void answerTimesOutWhenTheResumedWorkerNeverReplies() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
injector.onStatus(T, AgentStatus.WORKING);
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||
|
||||
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
|
||||
assertNotNull(q.turnId());
|
||||
|
||||
// The primary answers, unblocking the worker; but the worker never sends the follow-up
|
||||
// bridge_reply, so the answering send rides out its short window as still-working.
|
||||
MessageService.Reply answer = messages.answer(q.turnId(), "config.yaml", 200);
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answer.outcome(),
|
||||
"an answered worker that never replies times out as still working");
|
||||
|
||||
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome());
|
||||
assertEquals("config.yaml", a.answer());
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollReturnsNullForAnUnknownTicket() {
|
||||
assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown");
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollReportsACompletedTicket() throws Exception {
|
||||
String ticket = messages.sendAsync(T, "long task");
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker works
|
||||
assertTrue(rendezvous.resolve(T, "async result"), "a reply resolves the async send");
|
||||
|
||||
// Wait for the background send to finish and publish a DONE view.
|
||||
MessageService.TaskView view = null;
|
||||
long deadline = System.currentTimeMillis() + 2000;
|
||||
while (view == null || view.phase() != MessageService.Phase.DONE) {
|
||||
if (System.currentTimeMillis() >= deadline) break;
|
||||
view = messages.poll(ticket);
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertNotNull(view, "a resolved async send must become DONE");
|
||||
assertEquals(MessageService.Phase.DONE, view.phase());
|
||||
assertEquals("async result", view.reply(), "the completed ticket reports the reply");
|
||||
assertEquals("reply", view.replySource(), "a structured bridge_reply is sourced from 'reply'");
|
||||
}
|
||||
|
||||
@Test
|
||||
void concurrentSendToSameSessionWhileFirstHoldsItIsBusy() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> first =
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "first", 5000));
|
||||
awaitWaiting(); // the first send now holds the session lock, blocked on its reply
|
||||
|
||||
// A second send to the SAME session cannot take the lock within its short window.
|
||||
MessageService.Reply busy = messages.send(T, "second", 100);
|
||||
assertEquals(MessageService.Outcome.BUSY, busy.outcome(),
|
||||
"a second send while another holds the session is busy, not a hang");
|
||||
assertNull(busy.text());
|
||||
|
||||
// Release the first send so it resolves cleanly and the test thread is not left pinned.
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver the first message
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up
|
||||
assertTrue(rendezvous.resolve(T, "first done"), "the first send resolves with a reply");
|
||||
MessageService.Reply firstReply = first.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.REPLIED, firstReply.outcome());
|
||||
assertEquals("first done", firstReply.text());
|
||||
}
|
||||
|
||||
// --- CB-307 reply inbox ----------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void replyQueuesInInboxWhenNoSendIsOpen() {
|
||||
// No send is open for this session — reply should queue in the inbox.
|
||||
assertTrue(messages.reply(T, "queued-text"), "reply should succeed (queued)");
|
||||
|
||||
var drained = messages.drainReplies(T);
|
||||
assertEquals(1, drained.size());
|
||||
assertEquals("queued-text", drained.getFirst().content());
|
||||
}
|
||||
|
||||
@Test
|
||||
void replyResolvesOpenSendDoesNotQueue() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitUninterruptibly(T);
|
||||
|
||||
// An explicit reply resolves the open send.
|
||||
assertTrue(messages.reply(T, "send-resolved"), "reply should succeed (resolved live send)");
|
||||
|
||||
// The inbox should be empty — the reply went to the send, not the inbox.
|
||||
assertTrue(messages.drainReplies(T).isEmpty(), "no reply in the inbox");
|
||||
|
||||
MessageService.Reply r = send.get(3, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.REPLIED, r.outcome());
|
||||
assertEquals("send-resolved", r.text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void drainRepliesReturnsAllPendingThenEmptyOnNextCall() {
|
||||
messages.reply(T, "msg-1");
|
||||
messages.reply(T, "msg-2");
|
||||
|
||||
var first = messages.drainReplies(T);
|
||||
assertEquals(2, first.size());
|
||||
|
||||
var second = messages.drainReplies(T);
|
||||
assertTrue(second.isEmpty(), "second drain should be empty (acked)");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aQuestionIsNeverQueuedInTheInbox() {
|
||||
// No send is open — bridge_ask with no delegation returns NO_WAITER,
|
||||
// and the question text MUST NOT appear in the reply inbox.
|
||||
// The inbox is only fed by MessageService.reply(), not by bridge_ask.
|
||||
MessageService.AskResult r = messages.ask(T, "anyone there?", 500);
|
||||
assertEquals(MessageService.AskOutcome.NO_WAITER, r.outcome(),
|
||||
"bridge_ask with no open delegation must return NO_WAITER, never queued");
|
||||
|
||||
assertTrue(messages.drainReplies(T).isEmpty(), "questions must never be queued");
|
||||
}
|
||||
|
||||
@Test
|
||||
void completionFallbackIsNeverQueued() throws Exception {
|
||||
// The fallback resolves a captured waiter, never the inbox.
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitUninterruptibly(T);
|
||||
injectDelivery();
|
||||
|
||||
// The worker never sends bridge_reply, but the turn completes.
|
||||
herdr.readText("done-scraped");
|
||||
completion.onTurnComplete(T); // The fallback arms and resolves the captured waiter.
|
||||
|
||||
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, r.outcome());
|
||||
|
||||
// The inbox should be empty — the reply went to the captured waiter.
|
||||
assertTrue(messages.drainReplies(T).isEmpty(), "completion fallback must not queue");
|
||||
}
|
||||
|
||||
// --- helpers ---------------------------------------------------------------------------
|
||||
|
||||
/** Like {@link #awaitWaiting()} but rethrows as unchecked. */
|
||||
private void awaitUninterruptibly(String session) {
|
||||
try {
|
||||
awaitWaiting();
|
||||
} catch (InterruptedException e) {
|
||||
Thread.currentThread().interrupt();
|
||||
throw new IllegalStateException(e);
|
||||
}
|
||||
}
|
||||
|
||||
/** Set up a delivered turn so the worker is working, ready for an ask or completion. */
|
||||
private void injectDelivery() {
|
||||
herdr.readText("$ prompt"); // pre-turn content baseline
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver the task
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up
|
||||
}
|
||||
|
||||
// --- CB-516: a released session must not leave a send hanging ------------------------------
|
||||
|
||||
/**
|
||||
* The bug this fixes: tearing a worker down left its rendezvous waiter open, so a blocking send
|
||||
* kept blocking and an async one kept reporting PENDING until the 30-minute async timeout —
|
||||
* even though the worker provably no longer existed.
|
||||
*/
|
||||
@Test
|
||||
void abandonFailsASendThatIsStillWaitingOnAReleasedSession() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send =
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
|
||||
awaitWaiting();
|
||||
|
||||
assertTrue(messages.abandon(T, "session released"), "a live waiter is abandoned");
|
||||
|
||||
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.WORKER_FAILED, r.outcome(),
|
||||
"an abandoned send fails rather than riding out its timeout");
|
||||
assertEquals("session released", r.text(), "the caller is told why");
|
||||
}
|
||||
|
||||
@Test
|
||||
void abandonIsANoOpWhenNobodyIsWaiting() {
|
||||
assertFalse(messages.abandon(T, "session released"),
|
||||
"no open send ⇒ nothing to abandon");
|
||||
}
|
||||
|
||||
@Test
|
||||
void abandonDoesNotOverwriteAnAlreadyResolvedSend() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send =
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "the real answer"));
|
||||
|
||||
assertFalse(messages.abandon(T, "session released"),
|
||||
"a send already answered by the worker must not be clobbered");
|
||||
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals("the real answer", r.text());
|
||||
}
|
||||
|
||||
/** The async path is the one that hung: poll must report FAILED, not PENDING forever. */
|
||||
@Test
|
||||
void anAbandonedAsyncTaskPollsAsFailedNotPending() throws Exception {
|
||||
String ticket = messages.sendAsync(T, "long task");
|
||||
awaitWaiting();
|
||||
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase());
|
||||
|
||||
messages.abandon(T, "session released");
|
||||
|
||||
MessageService.TaskView view = null;
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (System.currentTimeMillis() < deadline) {
|
||||
view = messages.poll(ticket);
|
||||
if (view.phase() != MessageService.Phase.PENDING) break;
|
||||
Thread.sleep(10);
|
||||
}
|
||||
assertNotNull(view);
|
||||
assertEquals(MessageService.Phase.FAILED, view.phase(),
|
||||
"a delegation whose worker is gone must not keep reporting PENDING");
|
||||
assertTrue(view.detail() != null && view.detail().contains("released"),
|
||||
"and the detail says why, rather than 'worker unknown'");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -70,6 +70,28 @@ class RendezvousTest {
|
||||
"no blocked send means no primary to surface the question to");
|
||||
}
|
||||
|
||||
@Test
|
||||
void resolveCompletionTwiceIsANoOpTheSecondTime() {
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(W);
|
||||
assertTrue(rendezvous.resolveCompletion(waiter, "first scrape"), "the first completion resolves");
|
||||
assertFalse(rendezvous.resolveCompletion(waiter, "second scrape"),
|
||||
"a second completion on an already-resolved waiter returns false");
|
||||
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
|
||||
assertEquals("first scrape", waiter.getNow(null).text(),
|
||||
"the first resolution wins; the stored value is unchanged");
|
||||
}
|
||||
|
||||
@Test
|
||||
void resolveFailureTwiceIsANoOpTheSecondTime() {
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(W);
|
||||
assertTrue(rendezvous.resolveFailure(waiter, "first reason"), "the first failure resolves");
|
||||
assertFalse(rendezvous.resolveFailure(waiter, "second reason"),
|
||||
"a second failure on an already-resolved waiter returns false");
|
||||
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
|
||||
assertEquals("first reason", waiter.getNow(null).text(),
|
||||
"the first resolution wins; the stored value is unchanged");
|
||||
}
|
||||
|
||||
@Test
|
||||
void closeAskRemovesTheTurn() {
|
||||
Rendezvous.AskTicket t = rendezvous.openAsk(W);
|
||||
|
||||
@@ -0,0 +1,303 @@
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.HerdrClient;
|
||||
import dev.ltms.bridged.mcp.PrimaryRegistry;
|
||||
import dev.ltms.bridged.metrics.BridgedMetrics;
|
||||
import dev.ltms.bridged.metrics.Metrics;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.Collections;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.CountDownLatch;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* Unit tests for {@link ReplyPushLoop}: decision logic, nudge injection, idempotency,
|
||||
* bounded reminders, and stop conditions.
|
||||
*
|
||||
* <p>Uses a {@link RecordingHerdrClient} that synchronizes access to its call list so the
|
||||
* scheduler thread and test thread never have memory ordering issues. The {@code decide()}
|
||||
* tests use a simple client with no concurrency concern.
|
||||
*/
|
||||
class ReplyPushLoopTest {
|
||||
|
||||
private static final String PRIMARY = "term_primary";
|
||||
private static final String WORKER = "term_worker";
|
||||
private static final ObjectMapper MAPPER = new ObjectMapper();
|
||||
|
||||
private PrimaryRegistry registry;
|
||||
private AgentControl agents;
|
||||
private InMemoryReplyInbox inbox;
|
||||
private ScheduledExecutorService scheduler;
|
||||
|
||||
@BeforeEach
|
||||
void setUp() {
|
||||
registry = new PrimaryRegistry(PRIMARY);
|
||||
inbox = new InMemoryReplyInbox();
|
||||
scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
scheduler.shutdownNow();
|
||||
}
|
||||
|
||||
// --- decide() logic ------------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void decideWithoutPrimaryIsStop() {
|
||||
agents = agentWithStatus("idle");
|
||||
var loop = new ReplyPushLoop(
|
||||
new PrimaryRegistry(null), agents, inbox, scheduler, 5, 100);
|
||||
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(WORKER, 0));
|
||||
}
|
||||
|
||||
@Test
|
||||
void decideWithEmptyInboxIsStop() {
|
||||
agents = agentWithStatus("idle");
|
||||
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
|
||||
}
|
||||
|
||||
@Test
|
||||
void decideAtCapIsStop() {
|
||||
agents = agentWithStatus("idle");
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100).decide(WORKER, 2));
|
||||
}
|
||||
|
||||
@Test
|
||||
void decideUnderCapWithInjectablePrimaryIsInject() {
|
||||
agents = agentWithStatus("idle");
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
|
||||
}
|
||||
|
||||
@Test
|
||||
void decideUnderCapWithBlockedPrimaryIsInject() {
|
||||
agents = agentWithStatus("blocked");
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
|
||||
"BLOCKED is injectable");
|
||||
}
|
||||
|
||||
@Test
|
||||
void decideUnderCapWithDonePrimaryIsInject() {
|
||||
agents = agentWithStatus("done");
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
|
||||
"DONE is injectable");
|
||||
}
|
||||
|
||||
@Test
|
||||
void decideUnderCapWithBusyPrimaryIsWaitBusy() {
|
||||
agents = agentWithStatus("working");
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
|
||||
}
|
||||
|
||||
@Test
|
||||
void decideUnderCapWithUnknownPrimaryIsWaitBusy() {
|
||||
agents = agentWithStatus("unknown");
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
|
||||
}
|
||||
|
||||
@Test
|
||||
void decideStopsAfterInboxIsEmptied() {
|
||||
agents = agentWithStatus("idle");
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
|
||||
inbox.ack(WORKER, "m1");
|
||||
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
|
||||
}
|
||||
|
||||
// --- onReplyQueued integration -------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void injectablePrimaryCausesExactlyOneNudge() throws Exception {
|
||||
var rec = recordingClient();
|
||||
agents = new AgentControl(rec);
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
|
||||
loop(1, 50).onReplyQueued(WORKER);
|
||||
|
||||
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
|
||||
"one nudge (2 agent.send calls) should have been sent");
|
||||
|
||||
// Exactly one nudge = exactly 2 agent.send calls (text + submit)
|
||||
assertEquals(2, rec.sendCount());
|
||||
assertTrue(rec.sentParams().stream()
|
||||
.anyMatch(e -> e.getValue().toString().contains("bridge_poll")),
|
||||
"nudge text should contain bridge_poll");
|
||||
}
|
||||
|
||||
@Test
|
||||
void onReplyQueuedIsIdempotentPerTarget() throws Exception {
|
||||
var rec = recordingClient();
|
||||
agents = new AgentControl(rec);
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
|
||||
var loop = loop(1, 100);
|
||||
loop.onReplyQueued(WORKER);
|
||||
loop.onReplyQueued(WORKER); // second call — should be a no-op
|
||||
|
||||
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
|
||||
"expected exactly one nudge (2 sends)");
|
||||
Thread.sleep(200);
|
||||
assertEquals(2, rec.sendCount(),
|
||||
"second onReplyQueued must not trigger another nudge");
|
||||
}
|
||||
|
||||
@Test
|
||||
void sendsUpToCapThenStops() throws Exception {
|
||||
int cap = 2;
|
||||
var rec = recordingClient();
|
||||
agents = new AgentControl(rec);
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
rec.sendLatch = new CountDownLatch(cap * 2);
|
||||
|
||||
loop(cap, 50).onReplyQueued(WORKER);
|
||||
|
||||
assertTrue(rec.sendLatch.await(5, TimeUnit.SECONDS),
|
||||
cap + " nudges (" + (cap * 2) + " sends) should have fired");
|
||||
Thread.sleep(300);
|
||||
assertEquals(cap * 2, rec.sendCount(),
|
||||
"exactly " + (cap * 2) + " agent.send calls (cap=" + cap + ")");
|
||||
}
|
||||
|
||||
// --- nudge format --------------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void nudgeFormatIsCorrect() {
|
||||
String nudge = ReplyPushLoop.NUDGE_FORMAT.formatted(WORKER, WORKER);
|
||||
assertTrue(nudge.contains("Worker term_worker"));
|
||||
assertTrue(nudge.contains("bridge_poll(target=term_worker)"));
|
||||
}
|
||||
|
||||
// --- metrics (CB-512) ----------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void successfulNudgeIncrementsDelivered() throws Exception {
|
||||
var rec = recordingClient();
|
||||
agents = new AgentControl(rec);
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
Metrics metrics = new Metrics();
|
||||
|
||||
loop(1, 50, metrics).onReplyQueued(WORKER);
|
||||
|
||||
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
|
||||
"one nudge (2 agent.send calls) should have been sent");
|
||||
// The delivered count is bumped on the scheduler thread right after the send that releases
|
||||
// the latch — settle briefly so the counter is published before we read it.
|
||||
Thread.sleep(200);
|
||||
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"),
|
||||
"a successfully sent nudge must count as delivered");
|
||||
}
|
||||
|
||||
@Test
|
||||
void reminderCapIncrementsExhausted() {
|
||||
agents = agentWithStatus("idle");
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
Metrics metrics = new Metrics();
|
||||
|
||||
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100, metrics).decide(WORKER, 2));
|
||||
|
||||
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
|
||||
"hitting the reminder cap must count as exhausted");
|
||||
assertEquals(0, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"));
|
||||
}
|
||||
|
||||
// --- helpers -------------------------------------------------------------------------------
|
||||
|
||||
private ReplyPushLoop loop() {
|
||||
return loop(5, 100);
|
||||
}
|
||||
|
||||
private ReplyPushLoop loop(int maxReminders, long backoffMs) {
|
||||
return new ReplyPushLoop(registry, agents, inbox, scheduler, maxReminders, backoffMs);
|
||||
}
|
||||
|
||||
private ReplyPushLoop loop(int maxReminders, long backoffMs, Metrics metrics) {
|
||||
return new ReplyPushLoop(registry, agents, inbox, scheduler, maxReminders, backoffMs, metrics);
|
||||
}
|
||||
|
||||
private static AgentControl agentWithStatus(String status) {
|
||||
return new AgentControl(new FakeHerdrClient(status));
|
||||
}
|
||||
|
||||
/** Non-recording (single-threaded) fake — safe for decide() tests. */
|
||||
private static final class FakeHerdrClient implements HerdrClient {
|
||||
private final String agentStatus;
|
||||
|
||||
FakeHerdrClient(String agentStatus) {
|
||||
this.agentStatus = agentStatus;
|
||||
}
|
||||
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
if ("agent.get".equals(method)) {
|
||||
return MAPPER.createObjectNode()
|
||||
.set("agent", MAPPER.createObjectNode()
|
||||
.put("terminal_id", PRIMARY)
|
||||
.put("agent_status", agentStatus));
|
||||
}
|
||||
return MAPPER.createObjectNode();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Thread-safe recording fake that counts agent.send calls. Uses synchronized access
|
||||
* so the scheduler thread and test thread never race.
|
||||
*/
|
||||
private static final class RecordingHerdrClient implements HerdrClient {
|
||||
private final List<Map.Entry<String, Object>> calls =
|
||||
Collections.synchronizedList(new ArrayList<>());
|
||||
volatile CountDownLatch sendLatch = new CountDownLatch(2);
|
||||
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
if ("agent.get".equals(method)) {
|
||||
return MAPPER.createObjectNode()
|
||||
.set("agent", MAPPER.createObjectNode()
|
||||
.put("terminal_id", PRIMARY)
|
||||
.put("agent_status", "idle")); // recording double is always injectable
|
||||
}
|
||||
if ("agent.send".equals(method)) {
|
||||
calls.add(Map.entry(method, params));
|
||||
sendLatch.countDown();
|
||||
}
|
||||
return MAPPER.createObjectNode();
|
||||
}
|
||||
|
||||
long sendCount() {
|
||||
return calls.size();
|
||||
}
|
||||
|
||||
List<Map.Entry<String, Object>> sentParams() {
|
||||
return List.copyOf(calls);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static RecordingHerdrClient recordingClient() {
|
||||
return new RecordingHerdrClient();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,225 @@
|
||||
package dev.ltms.bridged.rest;
|
||||
|
||||
import dev.ltms.bridged.auth.CallerResolver;
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.PaneLocator;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.inject.Injector;
|
||||
import dev.ltms.bridged.mcp.ConnectionIdentity;
|
||||
import dev.ltms.bridged.metrics.BridgedMetrics;
|
||||
import dev.ltms.bridged.metrics.Metrics;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.session.FakeWorktrees;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.net.URI;
|
||||
import java.net.http.HttpClient;
|
||||
import java.net.http.HttpRequest;
|
||||
import java.net.http.HttpResponse;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* CB-501/505 enforcement over real HTTP. The unit tests pin the policy; these pin that the policy
|
||||
* is actually reached from a request — a rule enforced nowhere is not a control.
|
||||
*/
|
||||
class BridgedAppAuthTest {
|
||||
|
||||
private final HttpClient http = HttpClient.newHttpClient();
|
||||
private Javalin app;
|
||||
private Metrics metrics;
|
||||
|
||||
@AfterEach
|
||||
void stop() {
|
||||
if (app != null) app.stop();
|
||||
}
|
||||
|
||||
/**
|
||||
* Start the app with the given identity/auth wiring.
|
||||
*
|
||||
* @param pid the PID every connection resolves to — {@link FakeHerdr#WORKER_PID} makes the
|
||||
* caller worker {@code term_a}, anything else makes it a non-worker
|
||||
*/
|
||||
private int start(long pid, boolean tokenMode, String token) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
|
||||
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
|
||||
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
|
||||
k -> "BRIDGED_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
|
||||
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
|
||||
Injector injector = new Injector(agents);
|
||||
MessageService messages = new MessageService(agents, injector, new Rendezvous());
|
||||
|
||||
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
|
||||
CallerResolver callers = tokenMode
|
||||
? new CallerResolver(identity, true, token)
|
||||
: new CallerResolver(identity);
|
||||
metrics = BridgedMetrics.create(sessions, new dev.ltms.bridged.msg.InMemoryReplyInbox());
|
||||
|
||||
app = new BridgedApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
|
||||
callers, metrics).build().start("127.0.0.1", 0);
|
||||
return app.port();
|
||||
}
|
||||
|
||||
private HttpResponse<String> send(int port, String method, String path, String body, String auth)
|
||||
throws Exception {
|
||||
HttpRequest.Builder b = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path))
|
||||
.header("Content-Type", "application/json");
|
||||
if (auth != null) {
|
||||
b.header("Authorization", auth);
|
||||
}
|
||||
b = switch (method) {
|
||||
case "POST" -> b.POST(body == null
|
||||
? HttpRequest.BodyPublishers.noBody()
|
||||
: HttpRequest.BodyPublishers.ofString(body));
|
||||
case "DELETE" -> b.DELETE();
|
||||
default -> b.GET();
|
||||
};
|
||||
return http.send(b.build(), HttpResponse.BodyHandlers.ofString());
|
||||
}
|
||||
|
||||
// --- loopback-trust: the caller is the primary -------------------------------------------
|
||||
|
||||
@Test
|
||||
void thePrimaryMayOrchestrateButMayNotForgeAWorkerReply() throws Exception {
|
||||
int port = start(999_999, false, null); // no pane ⇒ primary
|
||||
|
||||
HttpResponse<String> read = send(port, "GET", "/profiles", null, null);
|
||||
assertEquals(200, read.statusCode(), "the primary may observe");
|
||||
|
||||
HttpResponse<String> reply = send(port, "POST", "/sessions/term_a/reply",
|
||||
"{\"content\":\"forged\"}", null);
|
||||
assertEquals(403, reply.statusCode(),
|
||||
"a forged reply would resolve the rendezvous the primary is itself waiting on");
|
||||
assertTrue(reply.body().contains("forbidden"));
|
||||
}
|
||||
|
||||
// --- loopback-trust: the caller is a worker ------------------------------------------------
|
||||
|
||||
@Test
|
||||
void aWorkerMayReplyAsItselfButNotAsAnother() throws Exception {
|
||||
int port = start(FakeHerdr.WORKER_PID, false, null); // resolves to term_a
|
||||
|
||||
HttpResponse<String> own = send(port, "POST", "/sessions/term_a/reply",
|
||||
"{\"content\":\"done\"}", null);
|
||||
assertEquals(200, own.statusCode(), "a worker replies on its own session");
|
||||
|
||||
HttpResponse<String> other = send(port, "POST", "/sessions/term_b/reply",
|
||||
"{\"content\":\"not mine\"}", null);
|
||||
assertEquals(403, other.statusCode(),
|
||||
"REST trusted the path id before CB-505; this is the hole being closed");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWorkerMayNotOrchestrate() throws Exception {
|
||||
int port = start(FakeHerdr.WORKER_PID, false, null);
|
||||
|
||||
assertEquals(403, send(port, "POST", "/workers", null, null).statusCode(),
|
||||
"a worker spawning workers would be escalating into the orchestrator role");
|
||||
assertEquals(403, send(port, "DELETE", "/workers/w2:p7", null, null).statusCode());
|
||||
assertEquals(403, send(port, "POST", "/sessions/term_b/message",
|
||||
"{\"content\":\"hi\"}", null).statusCode());
|
||||
assertEquals(403, send(port, "GET", "/sessions/term_a/replies", null, null).statusCode(),
|
||||
"draining an inbox is the primary's collection step");
|
||||
}
|
||||
|
||||
// --- token mode ---------------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void tokenModeRejectsAnUncredentialedNonWorkerWith401() throws Exception {
|
||||
int port = start(999_999, true, "s3cret");
|
||||
|
||||
HttpResponse<String> res = send(port, "GET", "/profiles", null, null);
|
||||
assertEquals(401, res.statusCode(), "no credential ⇒ authenticated as nothing");
|
||||
assertTrue(res.body().contains("unauthenticated"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void tokenModeAcceptsAValidBearerToken() throws Exception {
|
||||
int port = start(999_999, true, "s3cret");
|
||||
|
||||
assertEquals(200, send(port, "GET", "/profiles", null, "Bearer s3cret").statusCode());
|
||||
}
|
||||
|
||||
@Test
|
||||
void tokenModeStillHonoursConnectionDerivedWorkerIdentity() throws Exception {
|
||||
// The fleet must keep working when auth is switched on: a worker presents no token, and
|
||||
// must still be able to reply.
|
||||
int port = start(FakeHerdr.WORKER_PID, true, "s3cret");
|
||||
|
||||
assertEquals(200, send(port, "POST", "/sessions/term_a/reply",
|
||||
"{\"content\":\"done\"}", null).statusCode());
|
||||
}
|
||||
|
||||
// --- health, metrics ----------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void healthzStaysOpenWithoutCredentials() throws Exception {
|
||||
int port = start(999_999, true, "s3cret");
|
||||
|
||||
assertEquals(200, send(port, "GET", "/healthz", null, null).statusCode(),
|
||||
"a supervisor must be able to probe liveness before any credential is configured");
|
||||
}
|
||||
|
||||
@Test
|
||||
void metricsRequireAuthenticationAndRenderPrometheusText() throws Exception {
|
||||
int port = start(999_999, true, "s3cret");
|
||||
|
||||
assertEquals(401, send(port, "GET", "/metrics", null, null).statusCode());
|
||||
|
||||
HttpResponse<String> ok = send(port, "GET", "/metrics", null, "Bearer s3cret");
|
||||
assertEquals(200, ok.statusCode());
|
||||
assertTrue(ok.headers().firstValue("Content-Type").orElse("").startsWith("text/plain"));
|
||||
assertTrue(ok.body().contains("bridged_sessions{state=\"ready\"}"),
|
||||
"the session census gauge is exported even when empty");
|
||||
}
|
||||
|
||||
@Test
|
||||
void refusalsAreCounted() throws Exception {
|
||||
int port = start(999_999, true, "s3cret");
|
||||
|
||||
send(port, "GET", "/profiles", null, null); // 401
|
||||
send(port, "POST", "/sessions/term_a/reply", "{}", "Bearer s3cret"); // 403
|
||||
|
||||
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "unauthenticated"));
|
||||
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "forbidden"));
|
||||
}
|
||||
|
||||
// --- legacy constructor -------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void theLegacyConstructorLeavesAuthorizationOff() throws Exception {
|
||||
// The 29 pre-existing acceptance tests rely on this: no auth fixture, no enforcement.
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
|
||||
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
|
||||
Map.of(wcfg.profile(), wcfg), wcfg.profile(), _ -> "tok");
|
||||
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
|
||||
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous());
|
||||
app = new BridgedApp(herdr, workers, sessions, messages, sessions.asPresence(), null)
|
||||
.build().start("127.0.0.1", 0);
|
||||
|
||||
assertEquals(200, send(app.port(), "POST", "/sessions/term_a/reply",
|
||||
"{\"content\":\"x\"}", null).statusCode());
|
||||
assertEquals(404, send(app.port(), "GET", "/metrics", null, null).statusCode(),
|
||||
"no registry supplied ⇒ the endpoint is not mounted at all");
|
||||
}
|
||||
}
|
||||
@@ -75,7 +75,7 @@ class BridgedAppTest {
|
||||
poller.start();
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
MessageService messages = new MessageService(agents, injector, rendezvous);
|
||||
app = new BridgedApp(herdr, workers, sessions, messages, rendezvous, this.presence, null)
|
||||
app = new BridgedApp(herdr, workers, sessions, messages, this.presence, null)
|
||||
.build().start("127.0.0.1", 0);
|
||||
return app.port();
|
||||
}
|
||||
@@ -313,15 +313,10 @@ class BridgedAppTest {
|
||||
catch (Exception e) { throw new RuntimeException(e); }
|
||||
});
|
||||
|
||||
// The worker replies once a send is actually awaiting (retry past the startup race).
|
||||
HttpResponse<String> reply;
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
do {
|
||||
reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"LGTM ship it\"}");
|
||||
if (reply.statusCode() != 409) break;
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
} while (System.currentTimeMillis() < deadline);
|
||||
// Give the background send thread time to open its rendezvous waiter (CB-307: reply now
|
||||
// queues in the inbox if no waiter is open, which would break the round-trip).
|
||||
Thread.sleep(200);
|
||||
HttpResponse<String> reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"LGTM ship it\"}");
|
||||
assertEquals(200, reply.statusCode());
|
||||
|
||||
HttpResponse<String> res = send.get(6, java.util.concurrent.TimeUnit.SECONDS);
|
||||
@@ -341,20 +336,14 @@ class BridgedAppTest {
|
||||
String ticket = mapper.readTree(accepted.body()).get("ticket").asText();
|
||||
assertFalse(ticket.isBlank(), "an async send must return a ticket");
|
||||
|
||||
// The worker replies once the async send is actually awaiting (retry past the startup race).
|
||||
HttpResponse<String> reply;
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
do {
|
||||
reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"async LGTM\"}");
|
||||
if (reply.statusCode() != 409) break;
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
} while (System.currentTimeMillis() < deadline);
|
||||
// Give the background async send thread time to open its rendezvous waiter.
|
||||
Thread.sleep(200);
|
||||
HttpResponse<String> reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"async LGTM\"}");
|
||||
assertEquals(200, reply.statusCode());
|
||||
|
||||
// Polling the ticket now reports the finished delegation and its reply.
|
||||
JsonNode task;
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
do {
|
||||
task = mapper.readTree(req(port, "GET", "/tasks/" + ticket).body());
|
||||
if ("done".equals(task.path("phase").asText())) break;
|
||||
@@ -375,11 +364,26 @@ class BridgedAppTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void replyWithNoPendingSendIsConflict() throws Exception {
|
||||
void replyWithNoPendingSendQueuesInsteadOfConflict() throws Exception {
|
||||
// CB-307: a reply with no open send now queues in the inbox, not a 409 conflict.
|
||||
int port = startHealthy();
|
||||
HttpResponse<String> res = postJson(port, "/sessions/term_a/reply", "{\"content\":\"orphan\"}");
|
||||
assertEquals(409, res.statusCode());
|
||||
assertEquals("no_pending_send", mapper.readTree(res.body()).get("error").asText());
|
||||
assertEquals(200, res.statusCode());
|
||||
|
||||
// The queued reply is drainable.
|
||||
HttpResponse<String> drain = req(port, "GET", "/sessions/term_a/replies");
|
||||
assertEquals(200, drain.statusCode());
|
||||
JsonNode body = mapper.readTree(drain.body());
|
||||
assertEquals(1, body.get("replies").size());
|
||||
assertEquals("orphan", body.get("replies").get(0).get("content").asText());
|
||||
}
|
||||
|
||||
@Test
|
||||
void drainRepliesReturnsEmptyForNoReplies() throws Exception {
|
||||
int port = startHealthy();
|
||||
HttpResponse<String> res = req(port, "GET", "/sessions/term_a/replies");
|
||||
assertEquals(200, res.statusCode());
|
||||
assertEquals(0, mapper.readTree(res.body()).get("replies").size());
|
||||
}
|
||||
|
||||
@Test
|
||||
|
||||
@@ -6,6 +6,7 @@ import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
@@ -332,4 +333,75 @@ class SessionManagerTest {
|
||||
.filter(c -> paneId.equals(((Map<?, ?>) c.params()).get("pane_id")))
|
||||
.count();
|
||||
}
|
||||
|
||||
// --- CB-306 spawn-readiness gate: no half-registered session on timeout ----------------
|
||||
|
||||
@Test
|
||||
void acquireThrowsPeerUnreachableWhenGateTimesOutAndRegistersNoSession() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
herdr.agentStatus("unknown"); // never becomes injectable
|
||||
long[] clock = {0};
|
||||
|
||||
// Gate-enabled launcher (1 ms timeout + no-op sleeper that advances clock past deadline)
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
|
||||
new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
|
||||
1, () -> clock[0], () -> clock[0] += 10);
|
||||
SessionManager sessions = new SessionManager(workers, new GitWorktrees(), () -> 0L, 0);
|
||||
|
||||
assertThrows(PeerUnreachableException.class,
|
||||
() -> sessions.acquire("ltms-local", null, "/caller", "term_primary"),
|
||||
"acquire must throw PeerUnreachableException when spawn times out");
|
||||
|
||||
// No half-registered session — the error happened inside spawn, before
|
||||
// SessionManager could put() anything into the registry.
|
||||
assertTrue(sessions.roster().isEmpty(),
|
||||
"no session is registered when spawn times out (roster empty)");
|
||||
}
|
||||
|
||||
// --- CB-516: release must notify, so a blocked send can be failed --------------------------
|
||||
|
||||
@Test
|
||||
void releaseNotifiesTheListenerWithTheReleasedTerminal() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||
sessions.onRelease(released::add);
|
||||
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller", null);
|
||||
sessions.release(s.paneId());
|
||||
|
||||
assertEquals(java.util.List.of(s.terminalId()), released,
|
||||
"every teardown path funnels through release, so one hook must see the terminal");
|
||||
}
|
||||
|
||||
@Test
|
||||
void releasingAnUnknownPaneNotifiesNobody() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||
sessions.onRelease(released::add);
|
||||
|
||||
sessions.release("w9:p404"); // idempotent teardown of something already gone
|
||||
|
||||
assertTrue(released.isEmpty(), "no session removed ⇒ no send was waiting on it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aThrowingReleaseListenerDoesNotBlockTheTeardown() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
sessions.onRelease(_ -> {
|
||||
throw new IllegalStateException("listener blew up");
|
||||
});
|
||||
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller", null);
|
||||
assertDoesNotThrow(() -> sessions.release(s.paneId()),
|
||||
"a listener failure must never prevent the teardown it is reacting to");
|
||||
assertTrue(sessions.get(s.paneId()).isEmpty(), "and the session is still deregistered");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,131 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* Wrapper-behaviour tests for {@link SessionReaper} (the thread lifecycle). The TTL policy itself
|
||||
* (SessionManager.reapIdle) is covered by SessionManagerTest and is deliberately not retested here.
|
||||
* A real SessionManager is used, built the same way the rest of this package's tests do.
|
||||
*/
|
||||
class SessionReaperTest {
|
||||
|
||||
private static final long IDLE_TTL_SECONDS = 60;
|
||||
private static final long SHORT_INTERVAL_MILLIS = 20;
|
||||
|
||||
private static ClaudeCodeLauncher launcher() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
}
|
||||
|
||||
/** A manager on the fake worktree seam — these tests never touch a real git checkout. */
|
||||
private static SessionManager sessionManager() {
|
||||
return new SessionManager(launcher(), new FakeWorktrees());
|
||||
}
|
||||
|
||||
private static SessionReaper reaper() {
|
||||
return new SessionReaper(sessionManager(), IDLE_TTL_SECONDS, SHORT_INTERVAL_MILLIS);
|
||||
}
|
||||
|
||||
/**
|
||||
* A double {@code start()} must leave exactly one live loop, so a single {@code stop()} still
|
||||
* silences it. Asserting only "no throw" would pass against a reaper that never started at
|
||||
* all — and against one that started twice — which is the entire point of the guard.
|
||||
*/
|
||||
@Test
|
||||
void startIsIdempotent() throws InterruptedException {
|
||||
AtomicLong ticks = new AtomicLong();
|
||||
SessionReaper reaper = new SessionReaper(countingManager(ticks),
|
||||
IDLE_TTL_SECONDS, SHORT_INTERVAL_MILLIS);
|
||||
|
||||
assertDoesNotThrow(() -> {
|
||||
reaper.start();
|
||||
reaper.start();
|
||||
}, "a second start() must not throw");
|
||||
assertTrue(awaitTicks(ticks, 2), "the loop is running after a double start()");
|
||||
|
||||
// One stop() for two start() calls: if the second start had spawned its own loop, a
|
||||
// surviving thread would keep the counter climbing past this point.
|
||||
reaper.stop();
|
||||
Thread.sleep(SHORT_INTERVAL_MILLIS * 4);
|
||||
long settled = ticks.get();
|
||||
Thread.sleep(SHORT_INTERVAL_MILLIS * 4);
|
||||
assertEquals(settled, ticks.get(),
|
||||
"a single stop() must silence the reaper even after two start() calls");
|
||||
}
|
||||
|
||||
/** A manager whose clock counts reads — every {@code reapIdle} reads it exactly once. */
|
||||
private static SessionManager countingManager(AtomicLong ticks) {
|
||||
return new SessionManager(launcher(), new FakeWorktrees(), () -> {
|
||||
ticks.incrementAndGet();
|
||||
return System.nanoTime();
|
||||
});
|
||||
}
|
||||
|
||||
/** Bounded wait for the loop to tick at least {@code n} times; avoids fixed-sleep flakiness. */
|
||||
private static boolean awaitTicks(AtomicLong ticks, long n) throws InterruptedException {
|
||||
long deadline = System.currentTimeMillis() + 2000;
|
||||
while (ticks.get() < n && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(10);
|
||||
}
|
||||
return ticks.get() >= n;
|
||||
}
|
||||
|
||||
@Test
|
||||
void stopIsIdempotentAndSafeBeforeStart() {
|
||||
SessionReaper reaper = reaper();
|
||||
|
||||
assertDoesNotThrow(reaper::stop, "stop() before start() must not throw");
|
||||
assertDoesNotThrow(reaper::stop, "a second stop() must not throw");
|
||||
}
|
||||
|
||||
/**
|
||||
* The loop must actually iterate, and {@code stop()} must actually end it.
|
||||
*
|
||||
* <p>Observed through an injected clock rather than by sleeping and hoping: every
|
||||
* {@code reapIdle} call reads {@code nowNanos} exactly once, so the tick count <em>is</em> the
|
||||
* iteration count. Asserting merely "nothing threw" would pass even if {@code start()} were a
|
||||
* no-op, which is the whole behaviour under test.
|
||||
*/
|
||||
@Test
|
||||
void theLoopRunsRepeatedlyAndStopEndsIt() throws InterruptedException {
|
||||
AtomicLong ticks = new AtomicLong();
|
||||
SessionReaper reaper = new SessionReaper(countingManager(ticks),
|
||||
IDLE_TTL_SECONDS, SHORT_INTERVAL_MILLIS);
|
||||
|
||||
reaper.start();
|
||||
// Bounded wait rather than a fixed sleep + exact count: proves repetition without pinning
|
||||
// a timing-derived number that would flake on a loaded machine.
|
||||
boolean iterated = awaitTicks(ticks, 2);
|
||||
long whileRunning = ticks.get();
|
||||
reaper.stop();
|
||||
assertTrue(iterated,
|
||||
"the reaper loop must iterate repeatedly; observed " + whileRunning + " tick(s)");
|
||||
|
||||
// After stop() the loop must go quiet. Allow one in-flight iteration to finish, then
|
||||
// confirm the count has stopped advancing.
|
||||
Thread.sleep(SHORT_INTERVAL_MILLIS * 4);
|
||||
long settled = ticks.get();
|
||||
Thread.sleep(SHORT_INTERVAL_MILLIS * 4);
|
||||
assertEquals(settled, ticks.get(), "stop() must end the loop, not just flag it");
|
||||
}
|
||||
}
|
||||
@@ -166,4 +166,58 @@ class WorktreeSessionManagerTest {
|
||||
assertEquals(2, sessions.roster().size());
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-507 regression. A plain REST spawn supplies neither a requested nor a caller cwd
|
||||
* ({@code BridgedApp} hardcodes {@code callerCwd = null}), and the worktree branch used to
|
||||
* resolve the repo root from just those two — yielding {@code null}, which the real
|
||||
* {@code GitWorktrees} turns into {@code git -C null} and an NPE out of {@code ProcessBuilder}
|
||||
* (HTTP 500).
|
||||
*
|
||||
* <p>Note this asserts on the <em>recorded</em> cwd rather than expecting a throw:
|
||||
* {@link FakeWorktrees#repoRoot} only records its argument and returns a canned root, so a
|
||||
* null flows through the fake harmlessly. That permissiveness is precisely why the whole
|
||||
* suite stayed green while the feature was broken in production — so the assertion has to be
|
||||
* "a usable cwd was passed down", not "an exception was raised".
|
||||
*/
|
||||
@Test
|
||||
void worktreeAcquireWithNoRequestedOrCallerCwdStillResolvesANonNullRepoRoot() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees();
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
|
||||
sessions.acquire("ltms-local", null, null, null, new WorktreeRequest("cb-507", null));
|
||||
|
||||
assertFalse(worktrees.repoRootCalls().isEmpty(),
|
||||
"repoRoot should have been called to resolve the repo root");
|
||||
String cwd = worktrees.repoRootCalls().getFirst().cwd();
|
||||
assertNotNull(cwd, "a null cwd here becomes `git -C null` and NPEs in the real GitWorktrees");
|
||||
assertFalse(cwd.isBlank(), "a blank cwd is as unusable as a null one");
|
||||
}
|
||||
|
||||
/**
|
||||
* The same line carried a second, quieter bug: it never consulted the profile's configured
|
||||
* {@code cwd:}, so a worktree spawn silently ignored a pinned per-profile working directory.
|
||||
* Routing through {@code effectiveCwd} honours it.
|
||||
*/
|
||||
@Test
|
||||
void worktreeAcquireHonoursTheProfileConfiguredCwd() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
// Argument order matters: configDir is the 4th parameter, cwd the 11th (after mcpUrl).
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", null, "/pinned/dir", null);
|
||||
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(
|
||||
new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
SessionManager sessions = new SessionManager(launcher, worktrees);
|
||||
|
||||
sessions.acquire("ltms-local", null, null, null, new WorktreeRequest("cb-507b", null));
|
||||
|
||||
assertEquals(1, worktrees.repoRootCalls().size());
|
||||
assertEquals("/pinned/dir", worktrees.repoRootCalls().getFirst().cwd(),
|
||||
"the profile's configured cwd must reach repoRoot, not be ignored");
|
||||
}
|
||||
|
||||
}
|
||||
|
||||
@@ -7,6 +7,7 @@ import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.peer.Capability;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||
import dev.ltms.bridged.peer.SpawnRequest;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
@@ -318,4 +319,154 @@ class ClaudeCodeLauncherTest {
|
||||
|
||||
assertTrue(herdr.called("pane.close"), "stop via handle.id() must close the pane");
|
||||
}
|
||||
|
||||
// --- CB-306 spawn-readiness gate -----------------------------------------------------------
|
||||
|
||||
private static Map<String, BridgedConfig.Worker> workerConfigMap(String profile, String mcpUrl) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
profile, "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", profile), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", mcpUrl, null, null);
|
||||
return Map.of(cfg.profile(), cfg);
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnWaitsUntilInjectableThenReturnsHandle() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
herdr.agentStatus("unknown"); // first status call sees UNKNOWN
|
||||
long[] clock = {0};
|
||||
boolean[] firstSleep = {true};
|
||||
// The sleeper: advance the fake clock, and on the first call flip the
|
||||
// agent status to IDLE so the next poll succeeds.
|
||||
Runnable sleeper = () -> {
|
||||
clock[0] += 300;
|
||||
if (firstSleep[0]) {
|
||||
herdr.agentStatus("idle");
|
||||
firstSleep[0] = false;
|
||||
}
|
||||
};
|
||||
|
||||
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
|
||||
new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")),
|
||||
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
|
||||
5000, () -> clock[0], sleeper);
|
||||
|
||||
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
|
||||
|
||||
assertNotNull(handle, "spawn returns a handle when worker becomes injectable");
|
||||
assertEquals("w9:pW_1", handle.id(), "handle id matches the started pane");
|
||||
assertEquals(0, paneCloseCount(herdr, "w9:pW_1"),
|
||||
"no pane.close when worker becomes injectable before timeout");
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnThrowsPeerUnreachableWhenNeverInjectableAndReapsPane() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
herdr.agentStatus("unknown"); // always UNKNOWN
|
||||
long[] clock = {0};
|
||||
|
||||
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
|
||||
new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")),
|
||||
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
|
||||
1000, () -> clock[0], () -> clock[0] += 50);
|
||||
|
||||
PeerUnreachableException ex = assertThrows(
|
||||
PeerUnreachableException.class,
|
||||
() -> svc.spawn(new SpawnRequest(null, null, null)));
|
||||
|
||||
assertTrue(ex.getMessage().contains("w9:pW_1"),
|
||||
"exception message references the paneId: " + ex.getMessage());
|
||||
assertTrue(ex.getMessage().contains("1000"),
|
||||
"exception message references the timeout: " + ex.getMessage());
|
||||
assertTrue(clock[0] >= 1000, "fake clock advanced past the timeout: " + clock[0]);
|
||||
assertEquals(1, paneCloseCount(herdr, "w9:pW_1"),
|
||||
"pane was closed on timeout (no orphan left behind)");
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnReturnsImmediatelyWhenGateIsDisabled() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
// The default 6-arg constructor has spawnReadyTimeoutMs=0 (gate disabled).
|
||||
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
|
||||
|
||||
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
|
||||
|
||||
assertNotNull(handle, "spawn returns a handle when the gate is disabled");
|
||||
assertFalse(herdr.called("agent.get"),
|
||||
"agent.get is never called when the gate is disabled (no polling)");
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnGateRespectsZeroTimeoutEvenWithFullConstructor() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
long[] clock = {0};
|
||||
|
||||
// Explicit zero timeout with the full testability constructor — should
|
||||
// skip polling entirely, just like the legacy default path.
|
||||
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
|
||||
new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")),
|
||||
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
|
||||
0, () -> clock[0], () -> clock[0] += 1);
|
||||
|
||||
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
|
||||
|
||||
assertNotNull(handle, "spawn still succeeds with zero timeout");
|
||||
assertEquals(0, paneCloseCount(herdr, handle.id()),
|
||||
"no orphan pane close from the gate path");
|
||||
}
|
||||
|
||||
// --- CB-511: worker environment seeding -----------------------------------------------------
|
||||
|
||||
@Test
|
||||
void workerInheritsTheDaemonPath() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||
k -> "PATH".equals(k) ? "/opt/tools/bin:/usr/bin" : null).spawn();
|
||||
|
||||
assertEquals("/opt/tools/bin:/usr/bin", startEnv(herdr).get("PATH"),
|
||||
"a worker with no PATH cannot run the build it is asked to run");
|
||||
}
|
||||
|
||||
@Test
|
||||
void profileEnvIsInjectedIntoTheWorker() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
|
||||
null, Map.of("JAVA_HOME", "/opt/jdk", "PATH", "/profile/bin"));
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||
k -> "PATH".equals(k) ? "/daemon/bin" : null).spawn();
|
||||
|
||||
Map<String, String> env = startEnv(herdr);
|
||||
assertEquals("/opt/jdk", env.get("JAVA_HOME"), "profile env: is passed through");
|
||||
assertEquals("/profile/bin", env.get("PATH"), "an explicit profile PATH overrides the daemon's");
|
||||
}
|
||||
|
||||
/**
|
||||
* The security-relevant ordering. {@code SubscriptionGuard} is checked against the profile's
|
||||
* {@code baseUrl} only, so if a profile's {@code env:} could overwrite ANTHROPIC_BASE_URL a
|
||||
* worker could be pointed at an unguarded host while the guard passed on a benign one.
|
||||
*/
|
||||
@Test
|
||||
void profileEnvCannotOverrideGuardCheckedAnthropicVars() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
|
||||
null, Map.of("ANTHROPIC_BASE_URL", "http://evil.example.com"));
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||
_ -> null).spawn();
|
||||
|
||||
assertEquals("http://gx00.gw:8000", startEnv(herdr).get("ANTHROPIC_BASE_URL"),
|
||||
"the guard-checked baseUrl must win over any env: entry, or the boundary is bypassable");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,158 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import dev.ltms.bridged.peer.SpawnRequest;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* The composite router: profile → owning adapter for spawn/cwd/parity, pane id → owner for stop,
|
||||
* and fleet-wide union/dedup for list/reap/caps/profiles. Exercised through two real adapters —
|
||||
* claude-code + opencode — over one FakeHerdr, so each call is observed reaching the right adapter
|
||||
* (the started herdr agent name carries that adapter's {@code claude-}/{@code opencode-} prefix).
|
||||
*/
|
||||
class CompositePeerLauncherTest {
|
||||
|
||||
private ClaudeCodeLauncher claudeAdapter(FakeHerdr herdr) {
|
||||
// 12-arg back-compat Worker ctor → kind defaults to claude-code.
|
||||
BridgedConfig.Worker claude = new BridgedConfig.Worker("claude", "http://gx00.gw:8000", "coder",
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}",
|
||||
null, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("claude", claude), "claude", _ -> null);
|
||||
}
|
||||
|
||||
private OpenCodeLauncher opencodeAdapter(FakeHerdr herdr) {
|
||||
BridgedConfig.Worker gemini = new BridgedConfig.Worker("gemini", null, "google/gemini-2.5-pro",
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("opencode"), "tab", "bridged-workers", "w #{n}",
|
||||
null, null, null, "GITEA_ACCESS_TOKEN", null, BridgedConfig.Worker.KIND_OPENCODE);
|
||||
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
Map.of("gemini", gemini), "gemini", _ -> "tok");
|
||||
}
|
||||
|
||||
private CompositePeerLauncher composite(FakeHerdr herdr) {
|
||||
return new CompositePeerLauncher(
|
||||
List.of(claudeAdapter(herdr), opencodeAdapter(herdr)), "claude");
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static String startedName(FakeHerdr herdr) {
|
||||
return (String) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("name");
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnRoutesEachProfileToItsOwningAdapter() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
PeerLauncher composite = composite(herdr);
|
||||
|
||||
composite.spawn(new SpawnRequest("gemini", null, null));
|
||||
assertTrue(startedName(herdr).startsWith("opencode-"),
|
||||
"the gemini profile is spawned by the opencode adapter: " + startedName(herdr));
|
||||
|
||||
composite.spawn(new SpawnRequest("claude", null, null));
|
||||
assertTrue(startedName(herdr).startsWith("claude-"),
|
||||
"the claude profile is spawned by the claude-code adapter: " + startedName(herdr));
|
||||
}
|
||||
|
||||
@Test
|
||||
void nullProfileResolvesTheDefaultAndRoutesToItsOwner() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
composite(herdr).spawn(new SpawnRequest(null, null, null));
|
||||
assertTrue(startedName(herdr).startsWith("claude-"),
|
||||
"a no-profile spawn resolves the default (claude) and routes to its adapter");
|
||||
}
|
||||
|
||||
@Test
|
||||
void unknownProfileIsRejected() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
PeerLauncher composite = composite(herdr);
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> composite.spawn(new SpawnRequest("nope", null, null)),
|
||||
"a profile no adapter declares is an error");
|
||||
}
|
||||
|
||||
@Test
|
||||
void profilesAndDefaultAreExposedAcrossAdapters() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
PeerLauncher composite = composite(herdr);
|
||||
assertEquals(Set.of("claude", "gemini"), composite.profiles(),
|
||||
"profiles are the union of every adapter's profiles");
|
||||
assertEquals("claude", composite.defaultProfile());
|
||||
}
|
||||
|
||||
@Test
|
||||
void capabilitiesAreTheUnionOfEveryAdapter() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher claude = claudeAdapter(herdr);
|
||||
OpenCodeLauncher opencode = opencodeAdapter(herdr);
|
||||
PeerLauncher composite = new CompositePeerLauncher(List.of(claude, opencode), "claude");
|
||||
|
||||
assertTrue(composite.capabilities().containsAll(claude.capabilities()),
|
||||
"the fleet offers every claude-code capability");
|
||||
assertTrue(composite.capabilities().containsAll(opencode.capabilities()),
|
||||
"the fleet offers every opencode capability (incl. SELF_PR from its git-token profile)");
|
||||
}
|
||||
|
||||
@Test
|
||||
void listIsDeduplicatedByPaneIdAcrossAdaptersSharingHerdr() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
PeerLauncher composite = composite(herdr);
|
||||
// Both adapters wrap the same herdr, so each list() returns the same global agent set;
|
||||
// the composite must return each pane once, not once per adapter.
|
||||
assertEquals(1, composite.list().size(),
|
||||
"the single herdr-tracked pane appears once, not duplicated per adapter");
|
||||
}
|
||||
|
||||
@Test
|
||||
void reapSumsAcrossAdaptersAndEachAdapterReapsOnlyItsOwnPrefix() {
|
||||
// One foreign opencode orphan + one foreign claude orphan, from a prior daemon (different nonce).
|
||||
FakeHerdr herdr = new FakeHerdr()
|
||||
.withAgent("opencode-gemini-ffffff-1", "term_o", "wQ:pO", "wQ:tO")
|
||||
.withAgent("claude-claude-eeeeee-1", "term_c", "wQ:pC", "wQ:tC");
|
||||
PeerLauncher composite = composite(herdr);
|
||||
assertEquals(2, composite.reapOrphanWorkers(),
|
||||
"both orphans are reaped — one by each adapter, summed by the composite");
|
||||
}
|
||||
|
||||
@Test
|
||||
void stopTearsDownAPaneSpawnedThroughTheComposite() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
PeerLauncher composite = composite(herdr);
|
||||
PeerHandle handle = composite.spawn(new SpawnRequest("gemini", null, null));
|
||||
|
||||
composite.stop(handle.id());
|
||||
assertTrue(herdr.calls.stream()
|
||||
.anyMatch(c -> c.method().equals("pane.close")
|
||||
&& handle.id().equals(((Map<?, ?>) c.params()).get("pane_id"))),
|
||||
"stop routes to the spawning adapter and closes that worker's pane");
|
||||
}
|
||||
|
||||
@Test
|
||||
void constructorRejectsAProfileClaimedByTwoAdapters() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
// Two opencode adapters both declaring "gemini" — a profile-name collision.
|
||||
OpenCodeLauncher a = opencodeAdapter(herdr);
|
||||
OpenCodeLauncher b = opencodeAdapter(herdr);
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> new CompositePeerLauncher(List.of(a, b), "gemini"),
|
||||
"a profile two adapters both claim is a configuration error");
|
||||
}
|
||||
|
||||
@Test
|
||||
void constructorRejectsAnEmptyAdapterList() {
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> new CompositePeerLauncher(List.of(), "claude"),
|
||||
"at least one adapter must be configured");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,269 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.peer.Capability;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||
import dev.ltms.bridged.peer.SpawnRequest;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* The opencode adapter's launch build: a file-based MCP mount + reply-charter instructions (no
|
||||
* inline flags, no {@code ANTHROPIC_*}, no guard), the {@code -m} model flag, and the shared base
|
||||
* transport (naming, reap, readiness gate) proving the {@link HerdrPeerLauncher} SPI is neutral.
|
||||
*/
|
||||
class OpenCodeLauncherTest {
|
||||
|
||||
private static BridgedConfig.Worker opencodeCfg(String model, String mcpUrl, String gitTokenEnv) {
|
||||
return new BridgedConfig.Worker("gemini", null, model, null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("opencode"), "tab", "bridged-workers", "opencode: {model} #{n}", mcpUrl,
|
||||
null, null, gitTokenEnv, null, BridgedConfig.Worker.KIND_OPENCODE);
|
||||
}
|
||||
|
||||
/** Gate-disabled launcher whose per-spawn config dirs land under an inspectable temp root. */
|
||||
private OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, BridgedConfig.Worker cfg) {
|
||||
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
Map.of(cfg.profile(), cfg), cfg.profile(), k -> "GITEA_ACCESS_TOKEN".equals(k) ? "tok" : null,
|
||||
0, System::currentTimeMillis, () -> { }, configRoot);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static Map<String, Object> lastStart(FakeHerdr herdr) {
|
||||
return (Map<String, Object>) herdr.lastCall("agent.start").params();
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static Map<String, String> startEnv(FakeHerdr herdr) {
|
||||
return (Map<String, String>) lastStart(herdr).get("env");
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static List<String> startArgv(FakeHerdr herdr) {
|
||||
return (List<String>) lastStart(herdr).get("argv");
|
||||
}
|
||||
|
||||
@Test
|
||||
void writesRemoteMcpConfigAndCharterInstructionsWhenMcpUrlSet(@TempDir Path root) throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null))
|
||||
.spawn();
|
||||
|
||||
Map<String, String> env = startEnv(herdr);
|
||||
assertNull(env.get("ANTHROPIC_BASE_URL"), "opencode carries no ANTHROPIC_* / subscription boundary");
|
||||
String cfgPath = env.get("OPENCODE_CONFIG");
|
||||
assertNotNull(cfgPath, "OPENCODE_CONFIG points the worker at the generated config file");
|
||||
assertTrue(Path.of(cfgPath).startsWith(root), "config file is generated under the injected root");
|
||||
|
||||
// Assert on parsed structure, not substrings: the generated config is real JSON and its
|
||||
// whitespace is the formatter's business, not the contract's.
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
|
||||
JsonNode bridge = json.path("mcp").path("bridge");
|
||||
assertEquals("remote", bridge.path("type").asText(), "bridge is mounted as a remote MCP server");
|
||||
assertEquals("http://127.0.0.1:8765/mcp", bridge.path("url").asText(),
|
||||
"the profile's bridge MCP url is present");
|
||||
assertTrue(bridge.path("enabled").asBoolean(), "the bridge server is enabled");
|
||||
assertTrue(json.path("instructions").isArray() && !json.path("instructions").isEmpty(),
|
||||
"the reply charter is mounted via instructions");
|
||||
|
||||
// The instructions entry is a real file path holding the reply charter.
|
||||
Path charter = Path.of(cfgPath).resolveSibling("reply-charter.md");
|
||||
assertTrue(Files.exists(charter), "the charter file the config references was written");
|
||||
assertTrue(Files.readString(charter).contains("bridge_reply"),
|
||||
"the charter instructs the worker to answer via bridge_reply");
|
||||
}
|
||||
|
||||
@Test
|
||||
void noConfigFileWhenMcpUrlAbsent(@TempDir Path root) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null)).spawn();
|
||||
|
||||
assertNull(startEnv(herdr).get("OPENCODE_CONFIG"),
|
||||
"no bridge MCP url → no config file and no OPENCODE_CONFIG");
|
||||
}
|
||||
|
||||
@Test
|
||||
void passesTheModelAsDashMFlag(@TempDir Path root) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null)).spawn();
|
||||
|
||||
List<String> argv = startArgv(herdr);
|
||||
assertEquals("opencode", argv.getFirst(), "base opencode command preserved first");
|
||||
int m = argv.indexOf("-m");
|
||||
assertTrue(m >= 0, "model is selected with -m");
|
||||
assertEquals("google/gemini-2.5-pro", argv.get(m + 1), "the provider/model selector follows -m");
|
||||
}
|
||||
|
||||
@Test
|
||||
void noModelFlagWhenModelBlank(@TempDir Path root) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, root, opencodeCfg(null, null, null)).spawn();
|
||||
assertEquals(List.of("opencode"), startArgv(herdr), "no model → argv is the bare opencode command");
|
||||
}
|
||||
|
||||
@Test
|
||||
void injectsForgeTokenWhenProfileGrantsIt(@TempDir Path root) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, root, opencodeCfg(null, null, "GITEA_ACCESS_TOKEN")).spawn();
|
||||
assertEquals("tok", startEnv(herdr).get("GITEA_TOKEN"),
|
||||
"a git-token profile gets the peer-neutral GITEA_TOKEN grant, same as Claude");
|
||||
}
|
||||
|
||||
@Test
|
||||
void capabilitiesDeclareOrphanReapAndMcpAskAndConditionalSelfPr(@TempDir Path root) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
assertEquals(java.util.Set.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP),
|
||||
service(herdr, root, opencodeCfg(null, null, null)).capabilities(),
|
||||
"no git token → no SELF_PR");
|
||||
assertTrue(service(herdr, root, opencodeCfg(null, null, "GITEA_ACCESS_TOKEN"))
|
||||
.capabilities().contains(Capability.SELF_PR),
|
||||
"a git-token profile adds SELF_PR");
|
||||
}
|
||||
|
||||
@Test
|
||||
void foreignWorkerMatchesOpencodePrefixButNotClaude() {
|
||||
String nonce = "abc123";
|
||||
assertTrue(OpenCodeLauncher.isForeignWorker("opencode-gemini-def456-1", nonce),
|
||||
"an opencode pane from another process is foreign");
|
||||
assertFalse(OpenCodeLauncher.isForeignWorker("opencode-gemini-" + nonce + "-1", nonce),
|
||||
"our own opencode pane (same nonce) is not foreign");
|
||||
assertFalse(OpenCodeLauncher.isForeignWorker("claude-ltms-local-def456-1", nonce),
|
||||
"a claude pane is never reaped by the opencode adapter");
|
||||
}
|
||||
|
||||
@Test
|
||||
void productionConstructorsWireThroughToTheBase() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = opencodeCfg(null, null, null);
|
||||
// 5-arg (gate disabled) and 7-arg (gate enabled) production constructors both expose the profile.
|
||||
OpenCodeLauncher disabled = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
OpenCodeLauncher gated = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null, 5000, 100);
|
||||
assertEquals(java.util.Set.of("gemini"), disabled.profiles());
|
||||
assertEquals("gemini", gated.defaultProfile());
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnGateThrowsPeerUnreachableWhenNeverInjectable(@TempDir Path root) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
herdr.agentStatus("unknown"); // never injectable
|
||||
long[] clock = {0};
|
||||
OpenCodeLauncher svc = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
Map.of("gemini", opencodeCfg(null, null, null)), "gemini", _ -> null,
|
||||
1000, () -> clock[0], () -> clock[0] += 50, root);
|
||||
|
||||
PeerUnreachableException ex = assertThrows(PeerUnreachableException.class,
|
||||
() -> svc.spawn(new SpawnRequest(null, null, null)));
|
||||
assertTrue(clock[0] >= 1000, "the fake clock advanced past the timeout: " + clock[0]);
|
||||
long closes = herdr.calls.stream()
|
||||
.filter(c -> c.method().equals("pane.close"))
|
||||
.filter(c -> "w9:pW_1".equals(((Map<?, ?>) c.params()).get("pane_id")))
|
||||
.count();
|
||||
assertEquals(1, closes, "the worker pane was reaped on timeout (no orphan)");
|
||||
assertNotNull(ex.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnReturnsHandleWhenGateDisabled(@TempDir Path root) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
PeerHandle handle = service(herdr, root, opencodeCfg(null, null, null))
|
||||
.spawn(new SpawnRequest(null, null, null));
|
||||
assertNotNull(handle, "spawn returns a handle when the gate is disabled");
|
||||
assertFalse(herdr.called("agent.get"), "no polling when the gate is disabled");
|
||||
}
|
||||
|
||||
// --- CB-508: pinned OpenAI-compatible endpoint (e.g. a local vLLM) ---------------------------
|
||||
|
||||
/** A profile with a baseUrl but no model provider prefix cannot be resolved — fail loudly. */
|
||||
private static BridgedConfig.Worker pinnedCfg(String model, String baseUrl, String mcpUrl) {
|
||||
return new BridgedConfig.Worker("local", baseUrl, model, null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("opencode"), "tab", "bridged-workers", "opencode: {model} #{n}", mcpUrl,
|
||||
null, null, null, null, BridgedConfig.Worker.KIND_OPENCODE);
|
||||
}
|
||||
|
||||
@Test
|
||||
void baseUrlDeclaresACustomOpenAiCompatibleProvider(@TempDir Path root) throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, root, pinnedCfg("local-vllm/deepseek-v4-flash", "http://127.0.0.1:8000", null))
|
||||
.spawn();
|
||||
|
||||
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
assertNotNull(cfgPath, "a pinned endpoint needs a config file even with no bridge MCP url");
|
||||
JsonNode provider = new ObjectMapper().readTree(Path.of(cfgPath).toFile())
|
||||
.path("provider").path("local-vllm");
|
||||
|
||||
assertFalse(provider.isMissingNode(), "the provider id comes from the model selector");
|
||||
assertEquals("@ai-sdk/openai-compatible", provider.path("npm").asText());
|
||||
assertEquals("http://127.0.0.1:8000/v1", provider.path("options").path("baseURL").asText(),
|
||||
"a bare host:port gets /v1 appended — that is where these servers mount the API");
|
||||
assertFalse(provider.path("options").path("apiKey").asText().isBlank(),
|
||||
"the AI SDK requires a non-empty key even when the server ignores it");
|
||||
assertFalse(provider.path("models").path("deepseek-v4-flash").isMissingNode(),
|
||||
"the model half of the selector is declared under the provider");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aBaseUrlThatAlreadyCarriesAPathIsUsedVerbatim(@TempDir Path root) throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, root, pinnedCfg("local-vllm/m", "http://127.0.0.1:8000/openai/v1", null)).spawn();
|
||||
|
||||
JsonNode json = new ObjectMapper()
|
||||
.readTree(Path.of(startEnv(herdr).get("OPENCODE_CONFIG")).toFile());
|
||||
assertEquals("http://127.0.0.1:8000/openai/v1",
|
||||
json.path("provider").path("local-vllm").path("options").path("baseURL").asText(),
|
||||
"an endpoint mounted on a custom path must not have /v1 bolted on");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aPinnedEndpointRejectsAModelWithNoProviderPrefix(@TempDir Path root) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
OpenCodeLauncher launcher =
|
||||
service(herdr, root, pinnedCfg("deepseek-v4-flash", "http://127.0.0.1:8000", null));
|
||||
|
||||
// Silently falling back to the default gateway would point the worker at the wrong LLM
|
||||
// while looking healthy — the one failure mode worth being loud about.
|
||||
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, launcher::spawn);
|
||||
assertTrue(e.getMessage().contains("<provider>/<model>"), "the error says how to fix it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aPinnedEndpointAndTheBridgeMcpCoexistInOneConfig(@TempDir Path root) throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, root, pinnedCfg("local-vllm/deepseek-v4-flash",
|
||||
"http://127.0.0.1:8000", "http://127.0.0.1:8766/mcp")).spawn();
|
||||
|
||||
JsonNode json = new ObjectMapper()
|
||||
.readTree(Path.of(startEnv(herdr).get("OPENCODE_CONFIG")).toFile());
|
||||
assertEquals("remote", json.path("mcp").path("bridge").path("type").asText(),
|
||||
"pinning an endpoint must not drop the bridge MCP mount");
|
||||
assertFalse(json.path("provider").path("local-vllm").isMissingNode(),
|
||||
"and the provider block is still declared alongside it");
|
||||
assertTrue(json.path("instructions").isArray() && !json.path("instructions").isEmpty(),
|
||||
"the reply charter survives too");
|
||||
}
|
||||
|
||||
@Test
|
||||
void noBaseUrlDeclaresNoProviderSoTheDefaultGatewayIsUsed(@TempDir Path root) throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, root, opencodeCfg("opencode/some-free-model", "http://127.0.0.1:8766/mcp", null))
|
||||
.spawn();
|
||||
|
||||
JsonNode json = new ObjectMapper()
|
||||
.readTree(Path.of(startEnv(herdr).get("OPENCODE_CONFIG")).toFile());
|
||||
assertTrue(json.path("provider").isMissingNode(),
|
||||
"without a baseUrl opencode resolves its own provider as before");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,37 @@
|
||||
<configuration>
|
||||
<!--
|
||||
CB-506 — test-run logging. This file is NOT boilerplate; it exists to keep the test suite
|
||||
out of the CB-505 audit trail.
|
||||
|
||||
main/resources/logback.xml routes the `audit` logger to a RollingFileAppender at
|
||||
logs/audit.log. AuditLogTest and BridgedAppAuthTest exercise that same logger, so without
|
||||
this file `mvn test` appends fabricated records — denied/forbidden SPAWN/STOP/SEND from
|
||||
worker:term_a — to the production security log, byte-identical to real ones. An investigator
|
||||
could not tell a test fixture from a genuine intrusion attempt. Logback prefers
|
||||
logback-test.xml when it is on the test classpath, so this governs test runs only.
|
||||
|
||||
Two constraints if you edit this:
|
||||
- NEVER add a FileAppender/RollingFileAppender here. That reintroduces the bug.
|
||||
- Keep `audit` ENABLED (INFO, additivity=false). Setting it to OFF would silently break
|
||||
AuditLogTest, which attaches its own ListAppender and asserts on emitted records.
|
||||
-->
|
||||
|
||||
<appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
|
||||
<encoder>
|
||||
<pattern>%d{HH:mm:ss.SSS} [%thread] %-5level %logger{36} - %msg%n</pattern>
|
||||
</encoder>
|
||||
</appender>
|
||||
|
||||
<logger name="audit" level="INFO" additivity="false">
|
||||
<appender-ref ref="STDOUT"/>
|
||||
</logger>
|
||||
|
||||
<logger name="dev.ltms.bridged" level="WARN"/>
|
||||
|
||||
<logger name="org.eclipse.jetty" level="WARN"/>
|
||||
|
||||
<root level="INFO">
|
||||
<appender-ref ref="STDOUT"/>
|
||||
</root>
|
||||
|
||||
</configuration>
|
||||
@@ -0,0 +1,61 @@
|
||||
# CB-504 — systemd unit for bridged (Linux).
|
||||
#
|
||||
# The macOS launchd agent (deploy/dev.ltms.bridged.plist) is the supervision target for the
|
||||
# current single-host deployment. This unit exists for the per-host gateways CB-308 introduces,
|
||||
# which will run on Linux.
|
||||
#
|
||||
# Install (user service — bridged drives the user's herdr, not a system daemon):
|
||||
# mkdir -p ~/.config/systemd/user
|
||||
# cp deploy/bridged.service ~/.config/systemd/user/
|
||||
# # edit ExecStart / WorkingDirectory / Environment below, then:
|
||||
# systemctl --user daemon-reload
|
||||
# systemctl --user enable --now bridged
|
||||
# journalctl --user -u bridged -f
|
||||
|
||||
[Unit]
|
||||
Description=bridged — claude-bridge message server
|
||||
Documentation=https://git.ltms.dev/lms/claude-bridge/wiki
|
||||
# Ordering only: herdr is a user process and its socket may appear after us. This is advisory —
|
||||
# bridged retries the herdr socket rather than exiting, which is what actually makes a late
|
||||
# socket survivable. Do NOT add Requires=: a herdr restart must not take bridged down with it.
|
||||
After=herdr.service
|
||||
Wants=herdr.service
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
WorkingDirectory=%h/src/claude-bridge/bridged
|
||||
ExecStart=/usr/lib/jvm/temurin-25-jdk/bin/java -jar target/bridged.jar bridged.yaml
|
||||
|
||||
Environment=HERDR_SOCKET_PATH=%h/.config/herdr/herdr.sock
|
||||
# PATH matters more than it looks (CB-511): bridged propagates its own PATH to every worker it
|
||||
# spawns, so this line decides whether the fleet can run a build at all. systemd does not source a
|
||||
# login shell, so without it the daemon — and every worker — gets a bare default with no JDK/Maven.
|
||||
Environment=PATH=/usr/lib/jvm/temurin-25-jdk/bin:/usr/share/maven/bin:/usr/local/bin:/usr/bin:/bin
|
||||
# Secrets are NOT set here — this file is committed. Put the API/worker tokens in a private
|
||||
# drop-in that systemd reads with restrictive permissions:
|
||||
# systemctl --user edit bridged → [Service] / Environment=BRIDGED_API_TOKEN=...
|
||||
# or point EnvironmentFile at a 0600 file:
|
||||
# EnvironmentFile=%h/.config/bridged/env
|
||||
|
||||
Restart=on-failure
|
||||
RestartSec=10s
|
||||
# A bad config (e.g. a non-loopback bind without token auth) makes bridged fail fast by design.
|
||||
# Give up rather than restart-loop on a permanent error.
|
||||
StartLimitBurst=5
|
||||
StartLimitIntervalSec=120
|
||||
|
||||
# The daemon reads the repo, writes worktrees, and talks to a Unix socket — it needs no more.
|
||||
NoNewPrivileges=true
|
||||
PrivateTmp=true
|
||||
ProtectSystem=strict
|
||||
ProtectHome=read-write
|
||||
ProtectKernelTunables=true
|
||||
ProtectControlGroups=true
|
||||
RestrictSUIDSGID=true
|
||||
|
||||
StandardOutput=journal
|
||||
StandardError=journal
|
||||
SyslogIdentifier=bridged
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
@@ -0,0 +1,80 @@
|
||||
<?xml version="1.0" encoding="UTF-8"?>
|
||||
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
|
||||
<!--
|
||||
CB-504 — launchd agent for bridged (macOS).
|
||||
|
||||
This is the real supervision target today: the dogfooded daemon runs on macOS, where there is
|
||||
no systemd. A systemd unit ships alongside (deploy/bridged.service) for the Linux gateways
|
||||
CB-308 introduces.
|
||||
|
||||
Install:
|
||||
cp deploy/dev.ltms.bridged.plist ~/Library/LaunchAgents/
|
||||
# edit the paths + JAVA_HOME below to match this host, then:
|
||||
launchctl load -w ~/Library/LaunchAgents/dev.ltms.bridged.plist
|
||||
launchctl list | grep bridged
|
||||
|
||||
Note on ordering: launchd has no "start after herdr" primitive for user agents, and neither
|
||||
does systemd in a way that survives a socket appearing late. bridged retries the herdr socket
|
||||
on startup instead, so an agent that comes up before herdr converges rather than dying — that
|
||||
retry is the actual fix; KeepAlive below is the backstop.
|
||||
-->
|
||||
<plist version="1.0">
|
||||
<dict>
|
||||
<key>Label</key>
|
||||
<string>dev.ltms.bridged</string>
|
||||
|
||||
<key>ProgramArguments</key>
|
||||
<array>
|
||||
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home/bin/java</string>
|
||||
<string>-jar</string>
|
||||
<string>/Users/CHANGEME/src/claude-bridge/bridged/target/bridged.jar</string>
|
||||
<string>bridged.yaml</string>
|
||||
</array>
|
||||
|
||||
<!-- Config path in ProgramArguments is relative, so the working directory must be the module. -->
|
||||
<key>WorkingDirectory</key>
|
||||
<string>/Users/CHANGEME/src/claude-bridge/bridged</string>
|
||||
|
||||
<key>EnvironmentVariables</key>
|
||||
<dict>
|
||||
<key>JAVA_HOME</key>
|
||||
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home</string>
|
||||
<key>HERDR_SOCKET_PATH</key>
|
||||
<string>/Users/CHANGEME/.config/herdr/herdr.sock</string>
|
||||
<!--
|
||||
PATH matters more than it looks (CB-511): bridged propagates its own PATH to every worker
|
||||
it spawns, so this line decides whether the fleet can run a build at all. launchd does NOT
|
||||
source .zprofile/.zshrc, so without this the daemon (and therefore every worker) gets a
|
||||
bare /usr/bin:/bin and no JDK or Maven. Keep the toolchain entries first.
|
||||
-->
|
||||
<key>PATH</key>
|
||||
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home/bin:/Users/CHANGEME/Tool/apache-maven-3.9.16/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin</string>
|
||||
<!--
|
||||
Worker/API tokens are NOT set here: this file is committed. Export them from a private
|
||||
launchd override or a wrapper script. bridged reads the API token from the env var named
|
||||
by auth.tokenEnv (default BRIDGED_API_TOKEN) and only in auth.mode: token.
|
||||
-->
|
||||
</dict>
|
||||
|
||||
<key>RunAtLoad</key>
|
||||
<true/>
|
||||
|
||||
<!-- Restart on crash, but not in a tight loop if the config is bad (bridged fails fast on a
|
||||
non-loopback bind without token auth — that is a config error, not a transient one). -->
|
||||
<key>KeepAlive</key>
|
||||
<dict>
|
||||
<key>SuccessfulExit</key>
|
||||
<false/>
|
||||
</dict>
|
||||
<key>ThrottleInterval</key>
|
||||
<integer>10</integer>
|
||||
|
||||
<key>StandardOutPath</key>
|
||||
<string>/Users/CHANGEME/src/claude-bridge/bridged/logs/bridged.out.log</string>
|
||||
<key>StandardErrorPath</key>
|
||||
<string>/Users/CHANGEME/src/claude-bridge/bridged/logs/bridged.err.log</string>
|
||||
|
||||
<key>ProcessType</key>
|
||||
<string>Background</string>
|
||||
</dict>
|
||||
</plist>
|
||||
@@ -0,0 +1,60 @@
|
||||
# LavinMQ — the AMQP broker behind bridged's durable ReplyInbox (CB-307 Stage 2).
|
||||
#
|
||||
# Why this file exists: the broker was previously run ad hoc and simply vanished from the host,
|
||||
# which takes bridged down with it — AmqpReplyInbox.open throws on an unreachable broker and
|
||||
# Bridged.java:187 does not guard it, so a missing broker is a hard startup failure, not a
|
||||
# degraded mode. This pins the version, keeps the data, and brings itself back after a reboot.
|
||||
#
|
||||
# Usage:
|
||||
# docker compose -f deploy/lavinmq/compose.yaml up -d
|
||||
# docker compose -f deploy/lavinmq/compose.yaml ps
|
||||
# docker compose -f deploy/lavinmq/compose.yaml logs -f
|
||||
# docker compose -f deploy/lavinmq/compose.yaml down # keeps the volume
|
||||
# docker compose -f deploy/lavinmq/compose.yaml down -v # DESTROYS held replies
|
||||
#
|
||||
# Management UI: http://127.0.0.1:15672 (guest / guest)
|
||||
#
|
||||
# This is bridged's OWN broker. Do not point bridged at any other AMQP server on this host —
|
||||
# notably not the `local-rabbitmq` container, which belongs to a different project and would end
|
||||
# up carrying this project's queues.
|
||||
|
||||
name: bridged-broker
|
||||
|
||||
services:
|
||||
lavinmq:
|
||||
# Pinned deliberately: :latest silently moves the broker under a running daemon.
|
||||
image: cloudamqp/lavinmq:2.9.1
|
||||
container_name: bridged-lavinmq
|
||||
|
||||
# The failure this deployment exists to prevent — survive reboots and Docker restarts, but
|
||||
# stay down if it was stopped on purpose.
|
||||
restart: unless-stopped
|
||||
|
||||
# Loopback-bound on purpose. LavinMQ ships a default guest/guest account, which is only
|
||||
# acceptable because nothing off-host can reach it. bridged connects over 127.0.0.1, and
|
||||
# binding 0.0.0.0 here would expose a broker with default credentials to the network.
|
||||
ports:
|
||||
- "127.0.0.1:5672:5672" # AMQP — bridged.yaml broker.uri points here
|
||||
- "127.0.0.1:15672:15672" # HTTP management API + UI
|
||||
|
||||
# The whole point of Stage 2. Held-but-unacked replies live here; without a named volume a
|
||||
# `docker compose down` would discard exactly what the durable inbox exists to protect.
|
||||
volumes:
|
||||
- lavinmq-data:/var/lib/lavinmq
|
||||
|
||||
healthcheck:
|
||||
test: ["CMD", "lavinmqctl", "status"]
|
||||
interval: 30s
|
||||
timeout: 5s
|
||||
retries: 3
|
||||
start_period: 10s
|
||||
|
||||
logging:
|
||||
driver: json-file
|
||||
options:
|
||||
max-size: "10m"
|
||||
max-file: "3"
|
||||
|
||||
volumes:
|
||||
lavinmq-data:
|
||||
name: bridged-lavinmq-data
|
||||
@@ -1,6 +1,10 @@
|
||||
# CB-301-ext — Worktree provisioning + config-parity overlay
|
||||
|
||||
**Status:** design spec for review → delegate implementation.
|
||||
**Status:** ✅ shipped — implemented at commit `97ecc71` (per-worker git worktree + config-parity
|
||||
overlay). As-built: `session/GitWorktrees.java` behind the `Worktrees` port, wired in
|
||||
`Bridged.main` and configurable via `worktreeRoot` / per-profile `parityOverlay`
|
||||
(see `bridged.example.yaml`). Branch/worktree surface in `bridge_list` landed with CB-304
|
||||
(`9fe04bf`); the worker-opened-PR checkpoint landed as CB-302 (`64e70ef`).
|
||||
**Extends:** [CB-301 Session Manager](CB-301-Session-Manager.md) (shipped, commit `54d907c`).
|
||||
**Realizes:** the config-parity requirement in [Worker Git Workflow](Worker-Git-Workflow.md).
|
||||
**Grounded in:** `SessionManager`, `WorkerService.spawn/effectiveCwd`, `BridgedConfig.Worker`,
|
||||
|
||||
@@ -0,0 +1,143 @@
|
||||
# CB-306 — Spawn-Readiness Gate (launcher-owned terminal readiness)
|
||||
|
||||
**Status:** design note / delegation spec (branch `worker/cb-306-readiness`)
|
||||
**Issue:** gitea `lms/claude-bridge` #4
|
||||
**Owner of the behaviour:** `ClaudeCodeLauncher` (the `PeerLauncher` adapter) — NOT core.
|
||||
|
||||
## 1. Problem
|
||||
|
||||
`bridge_spawn` today returns a session the instant the herdr pane is started. The pane is not
|
||||
yet a usable Claude REPL — it may still be sitting at the folder-trust prompt, or the CLI may
|
||||
never come up at all. Nothing blocks or times out on that. Consequences:
|
||||
|
||||
- A `bridge_send` to a not-yet-ready worker surfaces as a **~60 s MCP-client timeout** (the send
|
||||
blocks waiting for a turn that can't start) instead of a fast, explicit spawn failure.
|
||||
- A worker stuck at the folder-trust prompt lingers in `SPAWNING` forever; nothing fails it.
|
||||
|
||||
We want **fail-fast spawn**: `spawn()` returns only once the peer is genuinely up and usable in
|
||||
its terminal, or throws a clean error (and leaves no orphan pane) within a bounded timeout.
|
||||
|
||||
## 2. What already exists (do NOT rebuild)
|
||||
|
||||
`SessionManager` + `PresenceBridge.markPresent()` already flip a session `SPAWNING → READY` on
|
||||
**any MCP contact from the worker** (`onReady(terminal)` → `transitionByTerminal(SPAWNING, READY)`,
|
||||
SessionManager ~line 249, "worker became available on the bridge MCP"). That is the **delivery
|
||||
lifecycle** and it stays exactly as-is. CB-306 does **not** touch it and does **not** replace it.
|
||||
|
||||
CB-306 adds a *complementary, launcher-side* gate: the launcher guarantees the **terminal** is a
|
||||
live, interactive REPL before it hands a handle back. The two signals are layered:
|
||||
|
||||
| Signal | Owner | Means | CB-306 |
|
||||
|---|---|---|---|
|
||||
| terminal reaches interactive REPL (herdr `IDLE`) | `ClaudeCodeLauncher` (this ticket) | pane is past folder-trust, CLI is up | **NEW — the spawn gate** |
|
||||
| first MCP contact → `SPAWNING→READY` | core (`SessionManager`/`PresenceBridge`) | worker spoke to the bridge | unchanged |
|
||||
|
||||
## 3. The readiness predicate (herdr status)
|
||||
|
||||
`AgentStatus.fromWire` maps herdr's wire strings to `IDLE | WORKING | BLOCKED | DONE | UNKNOWN`.
|
||||
A freshly started pane that has **not** reached an interactive Claude — including one stalled at
|
||||
the folder-trust prompt — reports **`UNKNOWN`** (herdr has not detected a Claude REPL yet). Once
|
||||
the CLI is up and settled at its prompt it reports **`IDLE`**.
|
||||
|
||||
**Predicate:** the pane is *ready* when `AgentControl.status(target)` first returns an
|
||||
**injectable** state (`IDLE`, `BLOCKED`, or `DONE` — reuse `AgentStatus.injectable()`). `UNKNOWN`
|
||||
= not ready. `WORKING` alone is ambiguous this early and should not by itself satisfy readiness;
|
||||
wait for an injectable state. (We do not need to know *why* a pane isn't ready — a trust stall,
|
||||
a crash, and a slow start all present as "never becomes injectable" and all correctly time out.)
|
||||
|
||||
## 4. Contract change on `spawn(SpawnRequest)`
|
||||
|
||||
`ClaudeCodeLauncher.spawn(SpawnRequest)` becomes **block-until-ready-or-throw**:
|
||||
|
||||
1. Start the pane exactly as today (`spawn(profile, cwd, callerCwd) → Agent`, build env + guard +
|
||||
argv, `spawnInTab`/`spawnAsPane`, unique-named).
|
||||
2. **Poll** `agentControl.status(paneId)` every `pollIntervalMs` (~300 ms) until it is `injectable()`
|
||||
or `spawnReadyTimeoutMs` elapses.
|
||||
3. **Ready** → return the `WorkerHandle(paneId, terminalId)` as today.
|
||||
4. **Timeout** → the launcher **closes the pane it started** (and its tab, via the same path
|
||||
`release`/`stop` uses) and throws **`PeerUnreachableException`** (new, in `dev.ltms.bridged.peer`).
|
||||
No orphan pane is left behind — the launcher cleans up its own failed birth.
|
||||
|
||||
`spawnReadyTimeoutMs == 0` (or unset) **disables** the gate = legacy non-blocking behaviour, so the
|
||||
change is opt-in per deployment and existing tests that don't configure it keep their old semantics.
|
||||
|
||||
### Testability seam (required)
|
||||
|
||||
Do **not** call `Thread.sleep` directly in the poll loop against a real clock — unit tests must not
|
||||
real-sleep. Introduce a small injectable seam, mirroring the existing `StatusPoller` style:
|
||||
|
||||
- a `LongSupplier nowMillis` (monotonic clock) **and** a sleep/wait hook (e.g. a
|
||||
`Sleeper`/`Waiter` functional interface, or reuse whatever `StatusPoller` already uses), both
|
||||
defaulting to the real implementations in the production constructor and overridable in tests.
|
||||
|
||||
Unit tests (add to the existing `ClaudeCodeLauncher` test):
|
||||
- fake `AgentControl` returns `UNKNOWN` a few times then `IDLE` → `spawn` returns the handle; assert
|
||||
no `close` was called.
|
||||
- fake `AgentControl` always `UNKNOWN` → `spawn` throws `PeerUnreachableException`; assert the pane
|
||||
**was closed** (verify `close(paneId)` invoked) and the fake clock advanced past the timeout.
|
||||
- `spawnReadyTimeoutMs == 0` → `spawn` returns immediately without polling (legacy path).
|
||||
|
||||
## 5. Config
|
||||
|
||||
Add to the launcher-level config (a bridged-level knob, not per-profile) in `bridged.yaml` +
|
||||
`BridgedConfig`:
|
||||
|
||||
```yaml
|
||||
spawn_ready_timeout_ms: 20000 # 0 disables the gate (legacy non-blocking spawn)
|
||||
spawn_ready_poll_ms: 300
|
||||
```
|
||||
|
||||
Jackson ignores unknown keys, so omitting them in existing YAML is safe; pick sane defaults in code
|
||||
(`20000` / `300`). Keep the names consistent with existing config field style in `BridgedConfig`.
|
||||
|
||||
## 6. Core / MCP propagation
|
||||
|
||||
`SessionManager.acquire(...)` already calls `launcher.spawn(req)`. A thrown
|
||||
`PeerUnreachableException` must propagate out as a **clean spawn failure**:
|
||||
|
||||
- The **worktree** acquire path already has a try/catch that cleans up a provisioned worktree when
|
||||
`spawn` throws — verify the new exception flows through it (worktree removed, nothing registered).
|
||||
- The **non-worktree** path registers the session only *after* `spawn` returns, so a throw means no
|
||||
half-live `SPAWNING` session is ever registered — confirm this and add a test.
|
||||
- `bridge_spawn` (MCP verb) must return an **error result** carrying the exception message, not a
|
||||
success with a dead session. Trace `BridgeMcp`/`BridgedApp` spawn handlers and make sure the
|
||||
exception becomes a clean tool error, not an uncaught 500 with a stack trace.
|
||||
|
||||
**Out of scope (do NOT do here):** gating `bridge_send` on session `READY` (existing status-gate +
|
||||
this spawn gate already close the window), MCP-handshake-as-readiness signal, the CB-307 broker,
|
||||
any config `kind:` discriminator, any second adapter.
|
||||
|
||||
## 7. Definition of done
|
||||
|
||||
- `ClaudeCodeLauncher.spawn` blocks until injectable or throws `PeerUnreachableException` +
|
||||
self-reaps the pane; gate disabled when timeout is 0.
|
||||
- New `PeerUnreachableException` in `dev.ltms.bridged.peer`.
|
||||
- Config knobs wired (`spawn_ready_timeout_ms`, `spawn_ready_poll_ms`) with safe defaults.
|
||||
- Existing `SPAWNING→READY` MCP-contact transition untouched.
|
||||
- New unit tests (ready / timeout+reap / disabled) green; **all existing tests still pass unchanged**.
|
||||
- Build clean via the worker's own `mvn` (primary re-runs the authoritative IDE + `mvn clean install`
|
||||
gate — self-reports are not verified facts).
|
||||
|
||||
## 8. Sequence
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant SM as SessionManager.acquire
|
||||
participant L as ClaudeCodeLauncher.spawn
|
||||
participant T as herdr (AgentControl)
|
||||
SM->>L: spawn(SpawnRequest)
|
||||
L->>T: start pane (env+guard+argv)
|
||||
loop until injectable or timeout
|
||||
L->>T: status(paneId)
|
||||
T-->>L: UNKNOWN / IDLE
|
||||
end
|
||||
alt reached injectable
|
||||
L-->>SM: PeerHandle(id, terminalId)
|
||||
else timed out
|
||||
L->>T: close(paneId) + tab
|
||||
L-->>SM: throw PeerUnreachableException
|
||||
SM-->>SM: no session registered / worktree cleaned
|
||||
end
|
||||
```
|
||||
|
||||
*Figure — the launcher blocks in `spawn` until the pane is a usable REPL, else self-reaps and throws.*
|
||||
@@ -0,0 +1,164 @@
|
||||
# CB-307 Stage 1 — Reply-Inbox Port + In-Memory Adapter (delegation spec)
|
||||
|
||||
**Ticket:** gitea #5 (CB-307). **Stage:** 1 of 2 (see the issue's "Implementation staging" comment).
|
||||
**Scope of THIS delegation:** the `ReplyInbox` port + the in-memory (soft-state) adapter, wired at the
|
||||
exact drop seam so a worker's terminal reply is **held instead of silently discarded** when no primary
|
||||
send is open. **No broker, no new dependency, no infra** — fully unit-testable and primary-gate-verifiable.
|
||||
Stage 2 (the AMQP/LavinMQ adapter behind the same port) is explicitly **out of scope** here.
|
||||
|
||||
> **Read this whole spec before starting.** The port and the publish seam are prescriptive
|
||||
> (non-negotiable). Where a choice is genuinely open it says "DECISION" with the required default —
|
||||
> follow the default and flag it in your completion message for the primary's review.
|
||||
|
||||
## 1. The bug this fixes (grounded in current code)
|
||||
|
||||
The reverse (worker→primary) path is `Rendezvous` — a `ConcurrentHashMap<session, CompletableFuture<Resolution>>`
|
||||
of **live blocking waiters only**. No queue, no store. When a worker calls `bridge_reply` and **no send
|
||||
is currently open** for that worker:
|
||||
|
||||
- `Rendezvous.resolve(session, content)` → `complete(...)` → `waiters.get(session) == null` →
|
||||
returns `false` (`msg/Rendezvous.java:212-215`).
|
||||
- The `content` string is **never retained** — it is dropped. The worker is told it failed:
|
||||
`BridgeMcp.reply` returns `error("no send is awaiting a reply for this worker")` (`mcp/BridgeMcp.java:270-272`);
|
||||
REST returns `409 no_pending_send` (`rest/BridgedApp.java:339-345`).
|
||||
|
||||
This is the observed "communication break": a worker that finishes just after its `bridge_send` timed
|
||||
out (the ~60s sync window) replies into the void. There is **no message-id, dedup, or ack** anywhere in
|
||||
the message path today.
|
||||
|
||||
## 2. What to build
|
||||
|
||||
### 2.1 The port — `dev.ltms.bridged.msg.ReplyInbox`
|
||||
|
||||
A thin interface owned by the `msg` layer. The in-memory adapter is Stage 1; the AMQP adapter (Stage 2)
|
||||
implements the **same** interface, so keep it broker-agnostic.
|
||||
|
||||
```java
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Holds terminal worker→primary replies that arrive with no live send to resolve, keyed by worker
|
||||
* session (target), until the primary drains them. Soft-state in Stage 1 (in-memory, lost on restart);
|
||||
* the Stage 2 AMQP adapter implements the same contract with cross-restart durability.
|
||||
*/
|
||||
public interface ReplyInbox {
|
||||
|
||||
/** A queued reply: an idempotency id, the worker session it came from, and the reply text. */
|
||||
record InboxMessage(String msgId, String target, String content) {}
|
||||
|
||||
/**
|
||||
* Queue {@code content} from worker {@code target} under {@code msgId}. Idempotent: publishing an
|
||||
* already-present {@code msgId} for {@code target} is a no-op (dedup), so an at-least-once Stage-2
|
||||
* redelivery cannot double-queue.
|
||||
*/
|
||||
void publish(String target, String msgId, String content);
|
||||
|
||||
/** Non-destructive snapshot of pending replies for {@code target} (FIFO), empty list if none. */
|
||||
List<InboxMessage> peek(String target);
|
||||
|
||||
/** Remove the reply {@code msgId} for {@code target} once the primary has taken it. No-op if absent. */
|
||||
void ack(String target, String msgId);
|
||||
}
|
||||
```
|
||||
|
||||
### 2.2 The default adapter — `InMemoryReplyInbox`
|
||||
|
||||
- Backed by a `ConcurrentHashMap<String, ...>` keyed by target session; per-target FIFO ordering.
|
||||
- Dedup by `msgId` within a target (a `LinkedHashMap<msgId, InboxMessage>` per target, or a deque + a
|
||||
seen-set — your call; preserve insertion order).
|
||||
- `peek` returns an immutable copy; `ack` removes by `msgId`. Thread-safe (concurrent publish vs. drain).
|
||||
- **This is soft-state, NOT persistence.** Lost on a `java -jar` bounce — that is correct and consistent
|
||||
with "bridged stays soft-state." Do **not** add any file/DB backing.
|
||||
|
||||
### 2.3 Publish seam — route reply through the service layer
|
||||
|
||||
Keep `Rendezvous` a pure synchronization primitive (do **not** give it an inbox field). Instead centralize
|
||||
in `MessageService`, which already owns the `Rendezvous` and will own the `ReplyInbox`:
|
||||
|
||||
- Add `MessageService.reply(String session, String content)`:
|
||||
```java
|
||||
/** Route a worker's explicit bridge_reply: resolve an open send, or queue it in the inbox if none. */
|
||||
public boolean reply(String session, String content) {
|
||||
if (rendezvous.resolve(session, content)) {
|
||||
return true; // a live send took it — unchanged fast path
|
||||
}
|
||||
inbox.publish(session, UUID.randomUUID().toString(), content); // was a silent drop
|
||||
return true; // held, not lost
|
||||
}
|
||||
```
|
||||
- Repoint the two callers off the bare `rendezvous.resolve(...)` onto `messages.reply(...)`:
|
||||
- `BridgeMcp.reply` (`mcp/BridgeMcp.java:262-273`) — on success return a normal ack; **remove** the
|
||||
`error("no send is awaiting a reply…")` branch (that case is now a successful queue).
|
||||
- `BridgedApp.replyMessage` (`rest/BridgedApp.java:330-346`) — return `200` (queued) instead of
|
||||
`409 no_pending_send`.
|
||||
|
||||
**DO NOT touch the QUESTION path.** `bridge_ask` / `rendezvous.resolveQuestion` must keep today's
|
||||
`NO_WAITER` behaviour — a mid-turn question is **interactive** (the worker blocks synchronously and cannot
|
||||
consume a late answer), so it must **never** be queued. Only terminal `REPLY`s go to the inbox.
|
||||
|
||||
**DO NOT queue the completion/failure fallbacks** (`resolveCompletion` / `resolveFailure`,
|
||||
`Rendezvous.java:196-210`). They target a *captured* waiter (CB-116); a `false` there means the turn was
|
||||
already resolved or the scrape is a stale late duplicate — queueing it risks double-delivery. Leave them
|
||||
exactly as they are. (Extending durability to completions is a deliberate Stage-2 consideration, not this.)
|
||||
|
||||
### 2.4 Drain seam — how the primary collects a stranded reply
|
||||
|
||||
The primary re-checks a worker it delegated to. Expose a drain keyed by **worker session (target)**:
|
||||
|
||||
- Add `MessageService.drainReplies(String target)`: `peek` the inbox, `ack` each returned `msgId`, hand
|
||||
back the `List<InboxMessage>` (or just the contents). At-least-once: peek→deliver→ack (ack only after
|
||||
the caller has them, so an in-flight failure re-surfaces them).
|
||||
- **DECISION (required default): expose via the existing poll verb, keyed by target.** Extend `bridge_poll`
|
||||
to accept an optional `target` (worker session) and, when present, return that worker's drained replies —
|
||||
alongside a matching REST route `GET /sessions/{id}/replies`. Do **not** change `send`/`answer` semantics
|
||||
(do not drain inside `send` — that conflates "deliver to worker" with "collect its mail"). Keep the
|
||||
existing ticket-based `bridge_poll(ticket)` path working unchanged. If you see a cleaner surface, still
|
||||
ship this default and note the alternative for review.
|
||||
|
||||
## 3. Config
|
||||
|
||||
**None for Stage 1.** The in-memory adapter is the unconditional default — wire `new InMemoryReplyInbox()`
|
||||
into `MessageService` in `Bridged.main`. Do **not** add a `broker:` config block (that arrives with the
|
||||
Stage-2 AMQP adapter: absent → in-memory, present → AMQP).
|
||||
|
||||
## 4. Acceptance criteria (what the primary will verify)
|
||||
|
||||
1. New `ReplyInbox` + `InboxMessage` + `InMemoryReplyInbox` in `dev.ltms.bridged.msg`.
|
||||
2. `bridge_reply` with **no open send** now **succeeds and queues** (no more `error` / `409`); the reply is
|
||||
later retrievable and identical.
|
||||
3. The queued reply is drainable by the primary keyed by target; draining **acks** it (a second drain
|
||||
returns nothing); dedup by `msgId` (re-publishing the same id does not double-queue).
|
||||
4. **QUESTION path unchanged** — `bridge_ask` with no open send still returns `NO_WAITER` (add/keep a test
|
||||
proving a question is never queued).
|
||||
5. Completion/failure fallbacks unchanged.
|
||||
6. Unit tests covering: `InMemoryReplyInbox` publish/peek/ack/dedup/FIFO/concurrency; `MessageService.reply`
|
||||
queues on no-waiter and resolves-not-queues when a send is open; `drainReplies` returns+acks;
|
||||
the QUESTION-not-queued guard.
|
||||
7. **All pre-existing tests still green** (baseline is **188**; your total must be ≥ 188 + your new tests).
|
||||
|
||||
## 5. Build & verification (worker side)
|
||||
|
||||
- Build with Maven from the worktree's `bridged/` dir. **Capture the exit code without a masking pipe**
|
||||
(`mvn clean install; echo "MVN_EXIT=$?"` — never `mvn … | tail`, which hides failures).
|
||||
- Read the real test totals from `target/surefire-reports/TEST-*.xml`, not from stdout scroll.
|
||||
- You do **not** have IDE MCP access — do not claim `ide_diagnostics` results. The **primary** runs the
|
||||
authoritative gate (IDE sync + diagnostics 0/0 + `mvn clean install`) before integrating. Your
|
||||
self-reported counts are inputs to that gate, not final facts.
|
||||
|
||||
## 6. Hard constraints (non-negotiable)
|
||||
|
||||
- **`.mcp.json` is `--skip-worktree` in your worktree — never edit, `git add`, or commit it.**
|
||||
- **`wiki/` is a submodule — never run git in it; never touch it.**
|
||||
- Commit only your feature changes (the new port/adapter, the `msg`/`mcp`/`rest` wiring, tests, and if
|
||||
you add config wiring in `Bridged.java`). Nothing else.
|
||||
- Work only inside your assigned worktree on your feature branch. The primary fast-forwards `main` after
|
||||
re-gating — do not touch `main`.
|
||||
- Java 25 idioms are welcome (unnamed `_` params, records). Keep the diff minimal and match surrounding style.
|
||||
|
||||
## 7. Definition of done (report back over the bridge)
|
||||
|
||||
Commit on your feature branch and reply with: the commit SHA, the surefire total (run/failures/errors), a
|
||||
one-line note on the drain-surface decision (§2.4) you shipped, and confirmation that `.mcp.json`/`wiki/`
|
||||
were untouched. The primary re-gates and integrates.
|
||||
@@ -0,0 +1,205 @@
|
||||
# CB-308 — Multi-Host Federation (Stage 5)
|
||||
|
||||
**Status:** design note (proposal)
|
||||
**Depends on:** CB-307 (broker-based reliable delivery) — CB-308 is the multi-host layer built *on*
|
||||
CB-307's broker fabric.
|
||||
**Relates to:** CB-401 (`PeerHandle` opaque id), CB-304 (`rosterView`), CB-306 (spawn-readiness),
|
||||
CB-303 (lifecycle limits), CB-117 (orphan reap).
|
||||
|
||||
## 1. Goal
|
||||
|
||||
Let `claude-bridge` coordinate agents that live on **more than one host** — a primary on host A
|
||||
delegating to workers on hosts B, C, … — without any host learning another host's terminals. The
|
||||
bus stays a **provider-neutral communication fabric**; multi-host is an addressing + routing
|
||||
concern, not a new kind of peer.
|
||||
|
||||
The design rests on three pieces (the shape this ticket proposes):
|
||||
|
||||
1. **Dedicated per-agent channels** — every agent has its own addressable inbox on the broker.
|
||||
2. **A federated agent directory** — a global "who/where/status" lookup, assembled from per-host
|
||||
presence, not a central database.
|
||||
3. **A per-host gateway** — each host runs a `bridged` that owns its local herdr, registers/manages
|
||||
its own sessions, and proxies messages to/from other hosts over the broker.
|
||||
|
||||
## 2. What is single-host today (the assumptions to break)
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph host["Single host (today)"]
|
||||
primary["primary<br/>(MCP client)"]
|
||||
daemon["bridged daemon<br/>127.0.0.1:8765"]
|
||||
reg["in-process registry<br/>keyed by PeerHandle.id() == paneId"]
|
||||
herdr["herdr<br/>(local unix-socket PTY mux)"]
|
||||
w1["worker pane wQ:p1"]
|
||||
w2["worker pane wQ:p2"]
|
||||
primary --> daemon
|
||||
daemon --> reg
|
||||
daemon --> herdr
|
||||
herdr --> w1
|
||||
herdr --> w2
|
||||
end
|
||||
```
|
||||
|
||||
*Figure 1 — everything is co-located and loopback.*
|
||||
|
||||
Three concrete bake-ins assume one host:
|
||||
|
||||
| Assumption | Where | Why it blocks multi-host |
|
||||
|---|---|---|
|
||||
| **herdr is local** | `herdr/` unix socket `~/.config/herdr/herdr.sock` | You cannot drive another host's PTYs → each host **must** own its herdr. This is why a per-host gateway is mandatory. |
|
||||
| **registry is in-process, keyed by `paneId`** | `session/SessionManager` | `paneId` (e.g. `wQ:p2B`) is a herdr-local coordinate — meaningless off-host. Routing needs a host-unique id. |
|
||||
| **loopback, no authn** | `rest/BridgedApp` binds `127.0.0.1:8765` | Fine on one host; the moment a second host can talk to a gateway, that link is a trust boundary. |
|
||||
|
||||
## 3. Target architecture
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph hostA["HOST A"]
|
||||
gA["gateway = bridged A"]
|
||||
regA["local registry + herdr"]
|
||||
primary["primary (MCP client)"]
|
||||
gA --- regA
|
||||
primary --- gA
|
||||
end
|
||||
subgraph hostB["HOST B"]
|
||||
gB["gateway = bridged B"]
|
||||
regB["local registry + herdr"]
|
||||
wb["worker panes"]
|
||||
gB --- regB
|
||||
gB --- wb
|
||||
end
|
||||
subgraph broker["BROKER (LavinMQ / AMQP) — CB-307 fabric"]
|
||||
inbox["agent.<id>.inbox queues"]
|
||||
roster["roster.* presence topic"]
|
||||
dlq["DLQ · delayed-retry (remind)"]
|
||||
end
|
||||
gA -->|"publish to agent.<id>.inbox"| inbox
|
||||
gB -->|"publish to agent.<id>.inbox"| inbox
|
||||
inbox -->|"owning gateway consumes"| gA
|
||||
inbox -->|"owning gateway consumes"| gB
|
||||
gA -->|"announce local agents"| roster
|
||||
gB -->|"announce local agents"| roster
|
||||
roster -->|"union view"| gA
|
||||
roster -->|"union view"| gB
|
||||
```
|
||||
|
||||
*Figure 2 — each gateway owns its local herdr + registry, consumes only its own agents' inboxes,
|
||||
and announces its agents onto a shared presence topic. The broker routes; no host sees another
|
||||
host's terminals.*
|
||||
|
||||
### 3.1 Component mapping (the three pieces)
|
||||
|
||||
- **Dedicated per-agent channels** = a per-agent AMQP routing key / queue, e.g.
|
||||
`agent.<globalId>.inbox`. The agent's **owning gateway is the only consumer** of its inbox.
|
||||
Senders publish to `agent.<id>.inbox` and never need to know the agent's host — the broker
|
||||
routes to whichever gateway holds it. LavinMQ additionally gives durability, DLX, and a native
|
||||
delayed-message exchange (the remind/backoff loop for free) — the same reasons CB-307 picked it.
|
||||
|
||||
- **Federated agent directory** = a **soft-state, bridge-owned** roster, *not* a broker-stored
|
||||
database. Per the persistence-boundary decision (bridged is soft-state; the broker owns *message*
|
||||
durability, not *who/where/status*), each gateway announces its local agents `(globalId, host,
|
||||
status, capabilities)` on a `roster.*` presence topic with periodic heartbeats. Every gateway
|
||||
builds an eventually-consistent **union view** — literally CB-304's `rosterView`, federated. A
|
||||
stale entry expires by missed heartbeat (reuses CB-303's idle/TTL thinking).
|
||||
|
||||
- **Per-host gateway** = today's `bridged` daemon, evolved. It already registers/manages sessions
|
||||
and controls its local herdr; multi-host adds exactly two responsibilities: (a) a broker client
|
||||
that consumes its agents' inboxes and injects into local herdr, and (b) presence announce +
|
||||
union-roster assembly. Evolution, not rewrite.
|
||||
|
||||
### 3.2 Routing rule
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
send["bridge_send(globalId, msg)"] --> lookup{"directory:<br/>is globalId local?"}
|
||||
lookup -->|"yes"| local["inject via local herdr<br/>(today's Injector path)"]
|
||||
lookup -->|"no"| pub["publish agent.<id>.inbox<br/>(broker routes to owning gateway)"]
|
||||
pub --> consume["owning gateway consumes<br/>→ injects into its local herdr"]
|
||||
```
|
||||
|
||||
*Figure 3 — one fork: local agents keep today's in-process inject path; remote agents go over the
|
||||
broker. A sender is oblivious to which branch it took.*
|
||||
|
||||
## 4. What CB-307 already provides vs. what is net-new
|
||||
|
||||
**CB-307 delivers the transport half** and is independently valuable on a single host: the AMQP
|
||||
broker fabric, the `bridged → broker` client/adapter, at-least-once + idempotent (dedup-by-id)
|
||||
delivery, DLQ, and delayed-retry (remind). That *is* the "proxy cross-host message" backbone;
|
||||
extending the same broker from "worker→primary reliability" to "gateway↔gateway" is incremental.
|
||||
|
||||
**Net-new for CB-308 (multi-host), five items:**
|
||||
|
||||
1. **Global agent id** — decouple the routing key from `paneId`. CB-401's `PeerHandle` already
|
||||
abstracts the routing id; make it host-unique (e.g. `<host>/<paneId>` or a UUID minted at spawn).
|
||||
The registry and all verbs route on the global id.
|
||||
2. **Federated directory** — presence announce + heartbeat + union roster over `roster.*`
|
||||
(§3.1).
|
||||
3. **Gateway routing** — the `local ? inject : publish` fork (§3.2), plus each gateway consuming
|
||||
its own agents' inbox queues and injecting into local herdr.
|
||||
4. **Cross-host spawn** — `spawn on host B` = publish a control request to B's control channel →
|
||||
gateway B runs `ClaudeCodeLauncher.spawn` **locally** (CB-306's readiness gate becomes *more*
|
||||
valuable here: the far side wants a positive "agent ready" before anyone sends) → announces the
|
||||
new agent into the federated roster.
|
||||
5. **Trust** — the broker connection is now the security boundary. A gateway injects env/tokens at
|
||||
daemon privilege (the CB-401 Stage-C concern), so a **remote-triggered spawn/send** needs
|
||||
authn/authz: who may act on which host, and which control channels a gateway will honour.
|
||||
|
||||
## 5. The one thing the broker does NOT dissolve
|
||||
|
||||
The MCP asymmetry survives the network. The primary is an MCP **client** to its **local** gateway;
|
||||
it cannot be called into. A worker on B replying to a primary on A flows:
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant W as worker (host B)
|
||||
participant GB as gateway B
|
||||
participant BR as broker
|
||||
participant GA as gateway A
|
||||
participant P as primary (host A, MCP client)
|
||||
W->>GB: bridge_reply
|
||||
GB->>BR: publish primary-bound (durable, msg id)
|
||||
BR->>GA: route to A's primary inbox
|
||||
Note over GA: held durably until the primary pulls
|
||||
P->>GA: blocking bridge_send resolves / bridge_poll
|
||||
GA-->>P: reply (then ACK to broker)
|
||||
```
|
||||
|
||||
*Figure 4 — the broker makes the middle hop lossless, ordered, and idempotent; the **final** hop
|
||||
into the primary is still a **pull** (gateway A holds the message until the primary's blocking
|
||||
`bridge_send` or `bridge_poll`). Cross-host neither improves nor worsens this — it just spans hosts.
|
||||
This is precisely the gap CB-307 closes on one host and CB-308 stretches across hosts.*
|
||||
|
||||
## 6. Staging & dependencies
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
cb307["CB-307<br/>broker-based reliable delivery<br/>(single host first)"] --> cb308["CB-308<br/>multi-host federation<br/>(this note)"]
|
||||
cb308 --> a["global agent id"]
|
||||
cb308 --> b["federated directory"]
|
||||
cb308 --> c["gateway routing"]
|
||||
cb308 --> d["cross-host spawn"]
|
||||
cb308 --> e["cross-host trust model"]
|
||||
classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff;
|
||||
class e gate
|
||||
```
|
||||
|
||||
*Figure 5 — CB-307 is the foundation; CB-308's five items build on it. The trust model (e) is the
|
||||
gating concern before any host accepts remote control.*
|
||||
|
||||
**Recommendation:** keep CB-307 scoped to single-host broker reliability (foundation, independently
|
||||
useful), and build CB-308's items on top once the broker fabric exists. Choose CB-307's broker /
|
||||
channel naming **multi-host-ready** now (per-agent routing keys, a `roster.*` topic namespace) so
|
||||
CB-308 doesn't have to repaint the topology.
|
||||
|
||||
## 7. Open questions
|
||||
|
||||
- **Directory ground-truth:** pure soft-state presence (heartbeats) vs. also treating broker queue
|
||||
existence as authoritative. Lean soft-state to preserve the persistence boundary; revisit if
|
||||
split-brain roster views cause mis-routing.
|
||||
- **Global id scheme:** `<host>/<paneId>` (human-legible, leaks host) vs. opaque UUID (clean, needs
|
||||
the directory to resolve host). Probably UUID in the protocol, host as directory metadata.
|
||||
- **Gateway discovery:** how gateways find the broker and each other (static config vs. discovery).
|
||||
- **Trust model shape:** per-host shared secret vs. mTLS on the broker vs. a capability token per
|
||||
control action — ties into CB-401 Stage-C.
|
||||
- **Failure semantics:** a host/gateway dies mid-turn — how the federated roster reaps it (missed
|
||||
heartbeat) and whether in-flight primary-bound messages survive (broker durability = yes).
|
||||
@@ -0,0 +1,305 @@
|
||||
# CB-402 — Second peer adapter: opencode (Stage B of the Peer Launcher SPI)
|
||||
|
||||
**Status:** ✅ **complete — implemented, merged (`ded226a`), and live-dogfooded 2026-07-29.**
|
||||
All five increments of §4 are done, including increment 5 (the §5 live checklist). See
|
||||
[§8 As-built](#8-as-built--live-dogfood-2026-07-29) for the run. Gitea issue #7 closed.
|
||||
**Depends on:** CB-401 Stage A (`PeerLauncher` SPI, merged `3aa69a9`)
|
||||
**Stage:** 4 (Pluggable peers) · Stage B
|
||||
**Owner action:** design-note → file issue → delegate → primary-verify (per CB-401/306/307)
|
||||
|
||||
---
|
||||
|
||||
## 1. Goal
|
||||
|
||||
Prove the [`PeerLauncher`](../bridged/src/main/java/dev/ltms/bridged/peer/PeerLauncher.java) SPI
|
||||
actually holds for a **non-Claude** coding agent by shipping a second, first-class in-tree
|
||||
adapter: **opencode** (`opencode` 1.1.31, a provider-agnostic terminal coding agent).
|
||||
|
||||
The product direction is *heterogeneous coding agents, Claude Code first-class* — not a
|
||||
human/mock peer. opencode is the right proof precisely because it differs from Claude Code on
|
||||
every seam the SPI is meant to hide:
|
||||
|
||||
| Seam | Claude Code | opencode | ⇒ SPI proof |
|
||||
|------|-------------|----------|-------------|
|
||||
| Subscription boundary | `ANTHROPIC_BASE_URL` + `SubscriptionGuard.assertWorker()` before any herdr call | none — provider-agnostic, off-subscription by nature | the guard is **Claude-private**, not core |
|
||||
| MCP mount | inline `--mcp-config '{…}'` launch flag | `opencode mcp add` / config file (`OPENCODE_CONFIG`) — **no inline flag** | "mount the bridge MCP" is adapter-private |
|
||||
| Instruction injection | `--append-system-prompt "<REPLY_CHARTER>"` | config `instructions` / `AGENTS.md` / `--agent` — **no append flag** | the reply-charter mount is adapter-private |
|
||||
| Model selection | `ANTHROPIC_MODEL` env | `-m provider/model` flag | env-vs-flag is adapter-private |
|
||||
| Name / reap scheme | `claude-<profile>-<nonce>-<seq>` | `opencode-<profile>-<nonce>-<seq>` | each adapter reaps only its own kind |
|
||||
|
||||
Everything *else* — herdr tab/pane placement, the CB-306 spawn-readiness gate, CB-301-ext
|
||||
worktree provisioning, CB-117 orphan reap, teardown, `list()`, cwd resolution — is transport
|
||||
machinery that is **identical** for both. That split is the whole design.
|
||||
|
||||
> Out of scope (deferred to Stage C / later): dynamic external plugin loading behind a
|
||||
> trust/capability model, capability *enforcement* at the verb layer (Stage A only *declares*
|
||||
> caps), and a `human`/mock peer.
|
||||
|
||||
---
|
||||
|
||||
## 2. Current state — one launcher, two concerns mixed
|
||||
|
||||
`ClaudeCodeLauncher` (584 LOC) is the sole `PeerLauncher`. It interleaves two concerns:
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph CCL["ClaudeCodeLauncher (584 LOC) — today"]
|
||||
direction TB
|
||||
T["herdr transport (GENERIC / reusable)<br/>tab-pane placement · spawn-ready gate · worktree<br/>orphan reap · teardown · list · cwd resolution · unique naming"]
|
||||
C["Claude-specific (per-agent)<br/>ANTHROPIC_BASE_URL + SubscriptionGuard · ANTHROPIC_MODEL<br/>argv --mcp-config · --append-system-prompt REPLY_CHARTER · 'claude-' name prefix"]
|
||||
end
|
||||
classDef generic fill:#2f855a,stroke:#22543d,color:#ffffff;
|
||||
classDef specific fill:#b7791f,stroke:#7b341e,color:#ffffff;
|
||||
class T generic
|
||||
class C specific
|
||||
```
|
||||
|
||||
*Figure 1 — the two concerns tangled inside today's single launcher; CB-402 splits them.*
|
||||
|
||||
There is also a **Stage-A deferral** to finish: `Bridged.main` still casts
|
||||
`(ClaudeCodeLauncher) workers` at the `BridgeMcp` and `BridgedApp` constructors. Those two
|
||||
callers only invoke `profiles()`, `defaultProfile()`, and `list()` — **all already on the
|
||||
`PeerLauncher` interface**. The cast survives for one reason only: `PeerLauncher.list()`
|
||||
returns `List<?>` (element type erased) while the callers use `Agent` element methods in their
|
||||
roster join. Finishing the migration is therefore small and contained (§4.D).
|
||||
|
||||
---
|
||||
|
||||
## 3. Target design
|
||||
|
||||
Template-Method base + two thin adapters + a routing composite that keeps the Stage-A seam
|
||||
(one `PeerLauncher` reference held by `SessionManager` / `BridgeMcp` / `BridgedApp`) intact.
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
IFACE["«interface»<br/>PeerLauncher"]
|
||||
COMP["CompositePeerLauncher<br/>routes by profile kind; fans out list/reap/caps"]
|
||||
BASE["«abstract»<br/>HerdrPeerLauncher<br/>transport: placement · ready-gate · reap · stop · cwd · naming"]
|
||||
CCL2["ClaudeCodeLauncher<br/>hooks: guard+ANTHROPIC_* env · --mcp-config · charter flag · prefix 'claude'"]
|
||||
OCL["OpenCodeLauncher<br/>hooks: provider env · OPENCODE_CONFIG file · AGENTS charter · prefix 'opencode'"]
|
||||
|
||||
IFACE -.implemented by.-> COMP
|
||||
IFACE -.implemented by.-> BASE
|
||||
BASE --> CCL2
|
||||
BASE --> OCL
|
||||
COMP -->|"kind=claude-code"| CCL2
|
||||
COMP -->|"kind=opencode"| OCL
|
||||
|
||||
classDef iface fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
|
||||
classDef base fill:#2f855a,stroke:#22543d,color:#ffffff;
|
||||
classDef leaf fill:#6b46c1,stroke:#44337a,color:#ffffff;
|
||||
class IFACE,COMP iface
|
||||
class BASE base
|
||||
class CCL2,OCL leaf
|
||||
```
|
||||
|
||||
*Figure 2 — extracted base, two adapters, and a routing composite behind the unchanged SPI.*
|
||||
|
||||
### A. Extract `HerdrPeerLauncher` (abstract base)
|
||||
|
||||
Move all transport machinery down from `ClaudeCodeLauncher`. What stays generic:
|
||||
|
||||
- fields `agents`, `spaces`, `profiles`, `defaultProfile`, `env`, `nameSeq`, the CB-306 gate
|
||||
knobs (`spawnReadyTimeoutMs`/`spawnReadyPollMs`/`nowMillis`/`sleeper`), and `nameNonce`;
|
||||
- `profiles()`, `defaultProfile()`, `parityOverlay()`, `effectiveCwd(…)`, `resolveCwd`;
|
||||
- the `spawn(SpawnRequest)` **skeleton**: resolve profile → cfg → *hook* → placement → gate → `WorkerHandle`;
|
||||
- `spawnInTab` / `spawnAsPane` / `tidy` / `startUniquelyNamed` (name built from a *hook* prefix);
|
||||
- `list()`, `reapOrphanWorkers()` / `isForeignWorker` / `workerNonce` (pattern built from the prefix hook), `stop()` / `usesTabPlacement` / `isAlreadyGone`;
|
||||
- `waitUntilInjectableOrThrow`, the `WorkerHandle` record, `putIfPresent`, `resolveEnv`, `sleepUninterruptibly`.
|
||||
|
||||
Two adapter **hooks** (abstract):
|
||||
|
||||
```java
|
||||
/** Label prefix for this peer kind; drives unique naming AND the orphan-reap pattern. */
|
||||
protected abstract String namePrefix(); // "claude" | "opencode"
|
||||
|
||||
/** Build the peer-specific launch: env map + argv. Runs any pre-spawn guard here. */
|
||||
protected abstract Launch buildLaunch(BridgedConfig.Worker cfg, SpawnRequest req);
|
||||
record Launch(Map<String,String> env, List<String> argv) {}
|
||||
```
|
||||
|
||||
`capabilities()` stays abstract/per-adapter (it already is). The subscription guard is **not**
|
||||
a base field — it is a constructor dependency of `ClaudeCodeLauncher` alone.
|
||||
|
||||
**Reap isolation:** the reap pattern becomes `Pattern.compile(namePrefix() + "-.*-([0-9a-f]{6})-\\d+")`,
|
||||
so the opencode adapter never reaps a `claude-*` pane and vice-versa. The composite sums both.
|
||||
|
||||
### B. `kind:` config discriminator
|
||||
|
||||
Add one field to `BridgedConfig.Worker`:
|
||||
|
||||
```java
|
||||
String kind // "claude-code" (default) | "opencode"
|
||||
```
|
||||
|
||||
- Compact-ctor default: `kind = blank ? "claude-code" : kind.toLowerCase()`.
|
||||
- `argv` default is currently `List.of("claude")`; when `kind=opencode` and the operator left
|
||||
`argv` unset, default it to `List.of("opencode")`. (Handle in normalization, keyed off `kind`,
|
||||
so the record stays declarative.)
|
||||
- Keep the existing back-compat constructors; `kind` is additive and optional.
|
||||
|
||||
`bridged.example.yaml` documents a two-kind `workers:` block.
|
||||
|
||||
### C. `OpenCodeLauncher` — the adapter hooks for opencode
|
||||
|
||||
`namePrefix()` → `"opencode"`. `buildLaunch(cfg, req)`:
|
||||
|
||||
- **Env:** *no* `ANTHROPIC_BASE_URL`, *no* `SubscriptionGuard` call. Pass through provider
|
||||
credentials the operator names (reuse the existing `tokenEnv` indirection; opencode reads
|
||||
provider keys from env / `opencode auth`). CB-302 git-token injection is reused unchanged
|
||||
(it is peer-neutral: `GITEA_TOKEN`/`GITEA_HOST`).
|
||||
- **MCP mount (non-invasive):** opencode has no inline `--mcp-config`. The adapter writes a
|
||||
throwaway config file and points `OPENCODE_CONFIG=<tempfile>` in the worker env, containing
|
||||
the bridge MCP server block (opencode HTTP MCP schema, `type: "remote"`) — the opencode analog
|
||||
of Claude Code's inline flag. Nothing is written into the worker's real project or profile.
|
||||
- **Reply-charter:** carry `REPLY_CHARTER` as an `instructions` entry in that same generated
|
||||
config (or an `AGENTS.md` written into the per-worker worktree, which is already a throwaway
|
||||
isolated checkout under CB-301-ext). Recommend the config-file route to keep the "touch
|
||||
nothing the user owns" invariant.
|
||||
- **argv:** `opencode <project-or-cwd>` (interactive TUI, the mode a herdr pane drives), plus
|
||||
`-m <provider/model>` when the profile sets a model.
|
||||
|
||||
> `REPLY_CHARTER` is peer-neutral text — hoist it to a shared constant (base or a small
|
||||
> `PeerCharter`), consumed by each adapter through its own injection mechanism.
|
||||
|
||||
### D. `CompositePeerLauncher` + finish the Stage-A migration
|
||||
|
||||
- `Bridged.main` groups configured profiles by `kind`, instantiates one launcher per kind
|
||||
present, and wraps them in `CompositePeerLauncher implements PeerLauncher`.
|
||||
- Routing methods (`spawn(req)`, `effectiveCwd(req)`, `parityOverlay(name)`) dispatch by the
|
||||
profile's kind. Fan-out methods (`list()`, `reapOrphanWorkers()`, `capabilities()`,
|
||||
`profiles()`, `defaultProfile()`) merge across sub-launchers. `stop(id)` tries each (teardown
|
||||
only knows the pane id) — already best-effort/idempotent.
|
||||
- **Migrate `BridgeMcp` + `BridgedApp` to the `PeerLauncher` interface**, dropping both
|
||||
`(ClaudeCodeLauncher)` casts. Only friction is `list()`'s `List<?>`; resolve by giving the SPI
|
||||
a typed roster element (small neutral `PeerAgent` view exposing `id()`/`name()`/status) that
|
||||
the CB-304 roster join consumes — or, minimally, narrow at the callsite. Prefer the typed view.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant P as Primary
|
||||
participant M as BridgeMcp / REST
|
||||
participant C as CompositePeerLauncher
|
||||
participant O as OpenCodeLauncher
|
||||
participant B as HerdrPeerLauncher (base)
|
||||
participant H as herdr
|
||||
P->>M: bridge_spawn(profile="oc-impl")
|
||||
M->>C: spawn(SpawnRequest)
|
||||
C->>C: kind(profile)=="opencode"
|
||||
C->>O: spawn(req)
|
||||
O->>O: buildLaunch → provider env + OPENCODE_CONFIG file + argv
|
||||
O->>B: placement + startUniquelyNamed("opencode-…")
|
||||
B->>H: agent.start(name, argv, env, tab, cwd)
|
||||
B->>H: poll status until injectable (CB-306 gate)
|
||||
B-->>O: Agent
|
||||
O-->>C: PeerHandle(paneId, terminalId)
|
||||
C-->>M: PeerHandle
|
||||
M-->>P: session id
|
||||
```
|
||||
|
||||
*Figure 3 — an opencode spawn: composite routes by kind, adapter builds the peer-specific launch, shared base drives herdr + the readiness gate.*
|
||||
|
||||
---
|
||||
|
||||
## 4. Increment plan (delegate-then-verify friendly)
|
||||
|
||||
1. **Extract base, no behaviour change.** Introduce `HerdrPeerLauncher`; make
|
||||
`ClaudeCodeLauncher` extend it with `namePrefix()="claude"` and `buildLaunch()` wrapping
|
||||
today's guard+env+argv logic. Green build, identical tests — pure refactor. *(IDE
|
||||
`refactor` where possible; the primary re-runs the gate workers can't.)*
|
||||
2. **`kind:` discriminator.** Add the field + normalization + `bridged.example.yaml`. Default
|
||||
path unchanged (`kind=claude-code`).
|
||||
3. **`OpenCodeLauncher`.** Implement the three hooks; unit-test `buildLaunch` (env has no
|
||||
`ANTHROPIC_BASE_URL`; `OPENCODE_CONFIG` points at a file carrying the bridge MCP block +
|
||||
charter; argv shape).
|
||||
4. **`CompositePeerLauncher` + wiring + finish Stage-A migration** (drop the two casts).
|
||||
5. **Live dogfood** (§5) + wiki as-built (primary-gated submodule commit).
|
||||
|
||||
Each increment is independently buildable/mergeable; the feature branch stays **unmerged**
|
||||
until it is "major" (Stage-B whole), per the CB-401 bar.
|
||||
|
||||
---
|
||||
|
||||
## 5. Risks & validation (live dogfood, not assumed)
|
||||
|
||||
- **opencode TUI ⇄ herdr injection.** herdr drives a pane by typing into a TUI. Must confirm
|
||||
opencode's TUI accepts injected keystrokes/submit the way `claude` does, and reaches an
|
||||
`injectable` status the CB-306 gate recognizes. *Validation:* spawn one opencode worker,
|
||||
watch the readiness gate pass, `bridge_send` a trivial task.
|
||||
- **Bridge MCP visibility in opencode.** Confirm `OPENCODE_CONFIG` (or `opencode mcp add`)
|
||||
actually surfaces the `bridge_*` tools inside the opencode session, and that `bridge_reply`
|
||||
is callable — the reply-charter is worthless if the tool isn't mounted. *Validation:* the
|
||||
worker completes a task by calling `bridge_reply`; the reply lands via the CB-307 path.
|
||||
- **opencode MCP/config schema drift.** opencode is fast-moving (1.1.31 today). Pin the config
|
||||
schema we generate against the installed version; treat the exact keys (`type: "remote"` vs
|
||||
`"http"`, `instructions` shape) as a dogfood-verified fact, not an assumption.
|
||||
- **Provider credentials.** opencode needs a configured provider (env key or `opencode auth`).
|
||||
The dogfood profile must name a provider the host actually has, distinct from the primary's
|
||||
subscription.
|
||||
|
||||
---
|
||||
|
||||
## 6. Test plan
|
||||
|
||||
- **Unit (hermetic):** base-extraction regression (existing `ClaudeCodeLauncher` tests pass
|
||||
unchanged); `OpenCodeLauncher.buildLaunch` env/argv/config assertions; `kind` normalization
|
||||
in `BridgedConfigTest`; `CompositePeerLauncher` routing + fan-out (merge of `profiles()`,
|
||||
summed `reapOrphanWorkers()`, per-kind reap isolation) with fake sub-launchers.
|
||||
- **Live (dogfood, manual):** the §5 checklist on the running daemon.
|
||||
- **Gate (primary):** IDE diagnostics 0/0 on every changed file, `mvn clean install` green with
|
||||
the surefire summary captured (not `| tail`), manual diff review — the authoritative checks a
|
||||
worker cannot self-run.
|
||||
|
||||
---
|
||||
|
||||
## 7. Open questions for the lead
|
||||
|
||||
1. ✅ **Provider — resolved 2026-07-29.** None was needed. opencode's own gateway serves
|
||||
**free-tier models with zero credentials** (`opencode auth list` → *0 credentials*, yet
|
||||
`opencode run -m opencode/north-mini-code-free` answers). The dogfood profile uses
|
||||
`opencode/north-mini-code-free`. It is distinct from the primary's Anthropic subscription by
|
||||
construction, and needs no `guard` entry — opencode carries no `ANTHROPIC_BASE_URL`, so the
|
||||
`SubscriptionGuard` does not apply to it at all.
|
||||
2. ✅ **Charter carrier — confirmed as designed:** generated `OPENCODE_CONFIG` `instructions`.
|
||||
Verified working against the installed version.
|
||||
3. ✅ **Merge cadence — resolved as it happened:** Stage B landed as one unit (`ded226a`).
|
||||
|
||||
---
|
||||
|
||||
## 8. As-built — live dogfood (2026-07-29)
|
||||
|
||||
Run against `bridged` on `127.0.0.1:8766` at main `19cdf8d`, with opencode **1.18.5** installed
|
||||
via Homebrew. Every §5 risk is now a verified fact rather than an assumption.
|
||||
|
||||
**The version-drift risk was the real one, and it did not bite.** This adapter was designed against
|
||||
opencode **1.1.31**; the installed version is **1.18.5**. The generated config schema still
|
||||
validates unchanged — `mcp.<name>.type: "remote"`, `url`, `enabled`, and `instructions: [path]` are
|
||||
all accepted, and `OPENCODE_CONFIG=… opencode mcp list` reports `✓ bridge connected`. Pinned here
|
||||
as a dogfood-verified fact for 1.18.5.
|
||||
|
||||
| §5 risk | Result |
|
||||
|---|---|
|
||||
| opencode TUI ⇄ herdr injection; CB-306 gate | ✅ `peer pane=wD:p3 reached injectable state` ~0.6s after `agent.start` |
|
||||
| Bridge MCP visible + `bridge_reply` callable | ✅ MCP `initialize` from `Implementation[name=opencode, version=1.18.5]`; worker replied through the tool |
|
||||
| Config schema drift (1.1.31 → 1.18.5) | ✅ unchanged, see above |
|
||||
| Provider credentials | ✅ free tier, zero credentials |
|
||||
|
||||
Full lifecycle exercised through the REST surface:
|
||||
|
||||
1. `POST /workers?profile=opencode-free` → `201`, routed by `kind:` through `CompositePeerLauncher`
|
||||
to `OpenCodeLauncher` (`spawning opencode profile=opencode-free`), pane `wD:p3`.
|
||||
2. Readiness: `{"ready":true,"status":"idle"}`, roster state `ready`.
|
||||
3. `POST /sessions/{id}/message` → **`{"replySource":"reply","reply":"391"}`** — a *structured*
|
||||
`bridge_reply`, not the CB-115 completion-fallback transcript scrape. The clean path.
|
||||
4. `DELETE /workers/wD:p3` → `204`, roster empty, tolerant teardown (`tab_not_found` ignored —
|
||||
opencode had already closed its own tab).
|
||||
|
||||
**Unplanned cross-validation with CB-501.** The audit trail recorded the worker's reply as
|
||||
`role: WORKER, actor: worker:term_657c1dad2b9731e, action: REPLY, outcome: allowed`. Connection-based
|
||||
identity (loopback peer PID → herdr pane) classified an **opencode** process as a worker with no
|
||||
opencode-specific handling — confirming the identity model is peer-kind-agnostic, which is exactly
|
||||
what CB-308 needs when it stretches the roster across hosts.
|
||||
|
||||
CB-502 counters for the same run: `bridged_sends_total{outcome="replied"} 1`,
|
||||
`bridged_replies_total{path="rendezvous"} 1`, `bridged_inbox_depth{...} 0`.
|
||||
@@ -0,0 +1,460 @@
|
||||
# CB-500 — Multi-Tier Coordination (Stage 6)
|
||||
|
||||
**Status:** design note (proposal — ticket split deferred)
|
||||
**Depends on:** CB-401/402 (Peer Launcher SPI + composite router — placement-neutral spawn),
|
||||
CB-308 (per-agent broker channels + global id + federated roster — the addressing substrate),
|
||||
CB-307 (durable inbox + push loop), CB-301/303 (session FSM + context-cap/idle-ttl), CB-304
|
||||
(`rosterView`).
|
||||
**Relates to:** the bus-identity boundary — see §7. This note stays a **proposal**; no code until the
|
||||
staging in §6 is reviewed and the arc is split into tickets.
|
||||
|
||||
## 1. Goal
|
||||
|
||||
Grow `claude-bridge` from a **single-tier** coordinator (one human-driven primary → a flat pool of
|
||||
workers) into a **multi-tier** one, along three axes the lead has asked for:
|
||||
|
||||
1. **Sandboxed workers** — each worker runs in a **separated, peer-owned sandbox** carrying its own
|
||||
toolchain (Claude routed via `ANTHROPIC_BASE_URL`, a headless IDE, git, MCP, dev-tools), with
|
||||
**per-role** sandboxes (a backend-agent image, a frontend-agent image).
|
||||
2. **Main-agent pairs** — the "main" tier becomes a **pair** (on-subscription Opus + one cloud
|
||||
module) collaborating, instead of a lone primary.
|
||||
3. **An orchestrator tier** — a supervisor **above** the mains that owns their **session identity**
|
||||
(naming, resume) and **curates context**, so every main→worker delegation carries the *exact*
|
||||
slice of context it needs and nothing else.
|
||||
|
||||
The through-line: **this is not a new pillar.** It is the existing `PeerLauncher` and
|
||||
`SessionManager` patterns extended one tier up, riding the **same CB-308 substrate** that multi-host
|
||||
already needs. Sandbox = a placement-neutral spawn target (CB-402 pattern). Pair + orchestrator =
|
||||
per-agent channels + a recursive session manager (CB-308 pattern). The bus stays a
|
||||
**provider-neutral communication fabric**; every addition is addressing, launch, or session scoping —
|
||||
never toolchain ownership (§7).
|
||||
|
||||
## 2. Single-tier today (the assumptions to break)
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
human["human (types)"]
|
||||
primary["PRIMARY (Opus)<br/>MCP client — pull-only"]
|
||||
daemon["bridged daemon<br/>127.0.0.1:8765 (single host)"]
|
||||
comp["CompositePeerLauncher<br/>routes by kind"]
|
||||
cc["ClaudeCodeLauncher"]
|
||||
oc["OpenCodeLauncher"]
|
||||
w1["worker pane (gx00 vLLM)"]
|
||||
w2["worker pane (ollama)"]
|
||||
human --> primary
|
||||
primary -->|"bridge_send / spawn / ask"| daemon
|
||||
daemon --> comp
|
||||
comp --> cc
|
||||
comp --> oc
|
||||
cc --> w1
|
||||
oc --> w2
|
||||
```
|
||||
|
||||
*Figure 1 — one human-driven primary, one daemon, a flat pool of bare herdr-pane workers.*
|
||||
|
||||
Four concrete bake-ins assume a single tier:
|
||||
|
||||
| Assumption | Where (verified) | Why it blocks the direction |
|
||||
|---|---|---|
|
||||
| **Exactly one primary** | `mcp/PrimaryRegistry` — an `AtomicReference<String>`, "single-slot registry for the primary's terminal" | A *pair* needs N addressable mains, each with its own pull inbox. |
|
||||
| **Workers are bare panes** | `worker/*Launcher` spawn a herdr pane via `argv:["ccs", …]` into a pre-existing env | A *sandbox* is a richer launch target (container/devcontainer) — a new placement, not a new provider. |
|
||||
| **`SpawnRequest` is flat** | `peer/SpawnRequest(profileName, requestedCwd, callerCwd)` | A sandbox/role selection needs a spawn-target dimension the record does not carry. |
|
||||
| **No tier above the primary** | there is no manager of the *primary's own* session — `SessionManager` manages *workers* only | An orchestrator that names/resumes/scopes the mains is a wholly new (but pattern-reusable) tier. |
|
||||
|
||||
## 3. Target multi-tier architecture
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
human["human"]
|
||||
subgraph orch["TIER 0 — orchestrator"]
|
||||
osm["OrchestratorSessionManager<br/>(SessionManager, recursed up)<br/>names · resumes · scopes context"]
|
||||
end
|
||||
subgraph mains["TIER 1 — main pair"]
|
||||
m1["main A: Opus<br/>MCP client"]
|
||||
m2["main B: cloud module<br/>MCP client"]
|
||||
end
|
||||
subgraph bus["bridged fabric (CB-307/308 substrate)"]
|
||||
chan["per-agent inbox channels<br/>agent.<globalId>.inbox"]
|
||||
roster["federated roster (union view)"]
|
||||
end
|
||||
subgraph workers["TIER 2 — sandboxed workers"]
|
||||
sbBE["backend sandbox<br/>Claude via ANTHROPIC_BASE_URL<br/>+ headless IDE · git · MCP · dev-tools"]
|
||||
sbFE["frontend sandbox<br/>(role-specific image)"]
|
||||
end
|
||||
human --> osm
|
||||
osm -->|"spawn / name / resume"| m1
|
||||
osm -->|"spawn / name / resume"| m2
|
||||
m1 <-->|"pull inbox"| chan
|
||||
m2 <-->|"pull inbox"| chan
|
||||
m1 -->|"scoped delegation"| bus
|
||||
m2 -->|"scoped delegation"| bus
|
||||
bus --> sbBE
|
||||
bus --> sbFE
|
||||
chan --- roster
|
||||
```
|
||||
|
||||
*Figure 2 — three tiers. Tier 0 owns the mains' session identity + context scope; Tier 1 is a
|
||||
collaborating pair, each an MCP client with its own pull inbox; Tier 2 is peer-owned sandboxes the
|
||||
bus launches into. The middle is CB-308's per-agent-channel + federated-roster substrate, now
|
||||
carrying tier-to-tier traffic, not just host-to-host.*
|
||||
|
||||
The recursion is the key idea: **`orchestrator : mains :: main : workers`** — the same
|
||||
spawn/name/resume/scope verbs at two levels.
|
||||
|
||||
## 4. Development A — Sandboxed, role-specific workers
|
||||
|
||||
A "sandbox" is a **placement**, not a provider — so it slots into the CB-401 SPI exactly the way
|
||||
CB-402's opencode adapter slotted in as a new *provider*. CB-402 proved the SPI is
|
||||
provider-neutral; a `SandboxLauncher` proves it is **placement-neutral**.
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
req["SpawnRequest<br/>(profileName, cwd, + sandbox/role)"]
|
||||
comp["CompositePeerLauncher<br/>routes by kind"]
|
||||
cc["ClaudeCodeLauncher<br/>kind: claude-code"]
|
||||
oc["OpenCodeLauncher<br/>kind: opencode"]
|
||||
sb["SandboxLauncher (NEW)<br/>kind: sandbox"]
|
||||
subgraph owned["bridge OWNS (launch + inject boundary)"]
|
||||
launch["run sandbox entrypoint<br/>(docker/devcontainer up → agent)"]
|
||||
inject["inject + guard baseUrl,<br/>mount bridge MCP + charter"]
|
||||
end
|
||||
subgraph peer["peer OWNS (inside the sandbox)"]
|
||||
img["image = backend|frontend role<br/>headless IDE · git · dev-tools · MCP"]
|
||||
end
|
||||
req --> comp
|
||||
comp --> cc
|
||||
comp --> oc
|
||||
comp --> sb
|
||||
sb --> launch --> inject
|
||||
inject -.->|"launches into, never builds"| img
|
||||
classDef line fill:#b7791f,stroke:#7b341e,color:#ffffff;
|
||||
class inject line
|
||||
```
|
||||
|
||||
*Figure 3 — the ownership line (amber). The bridge runs the sandbox entrypoint and injects the same
|
||||
boundary it owns today (guarded `baseUrl`, mounted MCP + reply charter). Everything inside the image
|
||||
— the IDE, git, dev-tools — is the peer's. The bridge references the image/role; it never provisions
|
||||
tools. This is what keeps "give the worker a headless IDE" on the right side of the "bus, not
|
||||
env-manager" rule (§7).*
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant M as main (delegator)
|
||||
participant D as bridged
|
||||
participant SL as SandboxLauncher
|
||||
participant SB as sandbox (peer-owned)
|
||||
participant A as agent in sandbox
|
||||
M->>D: bridge_spawn(profile=backend, role=backend)
|
||||
D->>SL: spawn(SpawnRequest)
|
||||
SL->>SB: start entrypoint (image = backend role)
|
||||
Note over SL,SB: bridge injects guarded ANTHROPIC_BASE_URL,<br/>mounts bridge MCP url + reply charter
|
||||
SB->>A: launch Claude (headless IDE, git, MCP ready — peer's own)
|
||||
A-->>SL: MCP connects → readiness gate (CB-306) passes
|
||||
SL-->>D: PeerHandle(globalId)
|
||||
D-->>M: spawned, injectable
|
||||
```
|
||||
|
||||
*Figure 4 — spawn into a peer-owned sandbox. Identical control flow to today's pane spawn (incl. the
|
||||
CB-306 readiness gate); only the launcher's `buildLaunch` differs — exactly the CB-402 seam.*
|
||||
|
||||
**Deltas:** a new `kind: sandbox` adapter (extends the same `HerdrPeerLauncher`/`PeerLauncher` base);
|
||||
a spawn-target/role dimension on `SpawnRequest` and the `Worker` profile; optionally a `SANDBOX`
|
||||
`Capability`. Per-role = two profiles → two images; `CompositePeerLauncher` already routes them. If a
|
||||
sandbox is a *separate host/container*, it reuses CB-308's global id + per-host gateway wholesale —
|
||||
**the distributed case is resolved in §11: a sandbox on another host is one spawned by that host's
|
||||
gateway, because herdr keystroke-injection needs a locally-owned PTY.**
|
||||
|
||||
## 5. Development B — Main-agent pairs
|
||||
|
||||
Both mains are MCP **clients**, so **neither can be called into** — each needs a **pull-based
|
||||
per-agent inbox**, which is precisely CB-308 item #1 (per-agent AMQP channels). The primary machinery
|
||||
that is singular today (single-slot `PrimaryRegistry`, a push-loop aimed at one terminal, "these
|
||||
tools only the primary calls") generalizes from a singleton to a **set**.
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph pair["TIER 1 — collaborating pair"]
|
||||
m1["main A: Opus<br/>MCP client (pull-only)"]
|
||||
m2["main B: cloud module<br/>MCP client (pull-only)"]
|
||||
end
|
||||
reg["PrimaryRegistry → multi-slot<br/>(terminal per main)"]
|
||||
subgraph fabric["bridged"]
|
||||
ca["agent.A.inbox"]
|
||||
cb["agent.B.inbox"]
|
||||
push["ReplyPushLoop → N terminals"]
|
||||
end
|
||||
m1 <-->|"peer-to-peer message"| m2
|
||||
m1 -->|"register terminal"| reg
|
||||
m2 -->|"register terminal"| reg
|
||||
ca -->|"pull / nudge"| m1
|
||||
cb -->|"pull / nudge"| m2
|
||||
reg --> push
|
||||
push --> ca
|
||||
push --> cb
|
||||
```
|
||||
|
||||
*Figure 5 — the pair. Each main owns an addressable inbox; `PrimaryRegistry` becomes multi-slot; the
|
||||
push loop nudges each main's terminal. Mains message each other as equals over the same bus (the
|
||||
transport is already peer-neutral — what was missing is N pull endpoints).*
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant MA as main A (Opus)
|
||||
participant BR as bridged / broker
|
||||
participant MB as main B (cloud)
|
||||
MA->>BR: bridge_send(to = main B, msg)
|
||||
BR->>BR: publish agent.B.inbox (durable, msg id)
|
||||
Note over BR: held until B pulls (B is a client too)
|
||||
MB->>BR: blocking bridge_send / poll resolves
|
||||
BR-->>MB: msg (then ACK)
|
||||
MB->>BR: bridge_reply(to = main A)
|
||||
BR->>BR: publish agent.A.inbox
|
||||
MA->>BR: poll resolves
|
||||
BR-->>MA: reply
|
||||
```
|
||||
|
||||
*Figure 6 — main↔main is the CB-307 asymmetry applied on both ends: two clients, so both hops are
|
||||
pull. This is why Part B **depends on** the per-agent-channel substrate, not just a config flag.*
|
||||
|
||||
**Deltas:** `PrimaryRegistry` single-slot → keyed-by-main; per-main inbox routing (CB-308 #1);
|
||||
push-loop fan-out; relax "orchestration tools only the primary calls" to "any registered main."
|
||||
|
||||
## 6. Development C — Orchestrator tier
|
||||
|
||||
The orchestrator is **`SessionManager` recursed one tier up**: today it spawns/names/reaps *worker*
|
||||
sessions; the orchestrator does the same for *main* sessions, and adds **context scoping**.
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
human["human"]
|
||||
subgraph t0["TIER 0 — orchestrator (new top MCP client)"]
|
||||
osm["OrchestratorSessionManager<br/>= SessionManager pattern"]
|
||||
idm["session identity<br/>name · resume · idle-ttl (CB-303)"]
|
||||
ctx["context scoper<br/>(CB-303 context-cap + turn_id)"]
|
||||
end
|
||||
subgraph t1["TIER 1 — mains (now MANAGED sessions)"]
|
||||
m1["main A"]
|
||||
m2["main B"]
|
||||
end
|
||||
subgraph t2["TIER 2 — workers"]
|
||||
w["sandboxed workers"]
|
||||
end
|
||||
human --> osm
|
||||
osm --> idm
|
||||
osm --> ctx
|
||||
idm -->|"spawn / name / resume"| m1
|
||||
idm -->|"spawn / name / resume"| m2
|
||||
ctx -->|"inject exact context slice"| m1
|
||||
m1 -->|"scoped delegation"| w
|
||||
m2 -->|"scoped delegation"| w
|
||||
```
|
||||
|
||||
*Figure 7 — the recursion. Tiers 1 and 2 run the identical spawn/name/resume machinery; the
|
||||
orchestrator merely operates it one level higher. **Re-rooting caveat:** today the primary IS the
|
||||
human's live session; here the human drives the orchestrator, and the mains become managed,
|
||||
resumable sessions. That moves the human-facing top up a tier — an intentional re-root, not an
|
||||
add-on.*
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant H as human
|
||||
participant O as orchestrator
|
||||
participant MA as main A
|
||||
participant W as worker
|
||||
H->>O: high-level goal (large context)
|
||||
O->>O: name/resume main A session
|
||||
O->>MA: task + SCOPED context slice (not the whole history)
|
||||
MA->>W: bridge_send(delegation, carrying only the relevant slice)
|
||||
W-->>MA: result
|
||||
MA-->>O: rollup
|
||||
O->>O: fold into orchestrator context, pick next main/turn
|
||||
```
|
||||
|
||||
*Figure 8 — context focus. The orchestrator holds the global context and hands each main only the
|
||||
slice a given delegation needs, so the main→worker conversation stays on-point. Context *scoping* is
|
||||
coordination (the bus already owns session/turn lifecycle) — it stays inside the identity boundary
|
||||
(§7), unlike toolchain ownership which does not.*
|
||||
|
||||
**Deltas:** a second, higher `SessionManager` instance whose "peers" are mains; the orchestrator
|
||||
becomes the top MCP client; context-slice selection (new) layered on CB-303's `context_cap` +
|
||||
`turn_id` scoping; mains gain a resumable session id in the federated roster.
|
||||
|
||||
## 7. The identity boundary — the one clause to hold
|
||||
|
||||
The bus is a **communication fabric, not an env/toolchain manager**. This direction is compatible
|
||||
**only** with the ownership split below; the amber line in Figure 3 is where it must hold.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph ok["STAYS A BUS (owned)"]
|
||||
a["launch INTO a sandbox<br/>(opaque image/role reference)"]
|
||||
b["inject + guard baseUrl,<br/>mount MCP + charter"]
|
||||
c["session identity + context scope<br/>(name/resume/turn_id)"]
|
||||
d["per-agent addressing + roster"]
|
||||
end
|
||||
subgraph drift["BECOMES ENV-MANAGER (forbidden)"]
|
||||
e["build images / install IDE<br/>or dev-tools"]
|
||||
f["wire the bridge's OWN IDE MCP<br/>into a worker"]
|
||||
g["enumerate 'what a frontend<br/>agent needs'"]
|
||||
end
|
||||
ok -.->|"red flag: any feature that only<br/>makes sense for ONE kind of peer"| drift
|
||||
classDef bad fill:#9b2c2c,stroke:#742a2a,color:#ffffff;
|
||||
class e,f,g bad
|
||||
```
|
||||
|
||||
*Figure 9 — the guardrail. A worker having a headless IDE **inside its own sandbox** is the peer
|
||||
owning its toolchain (left) — the opposite of the bridge reaching into the peer (right). Sandbox
|
||||
specs are peer-owned references (like `argv`/image id); the instant the bridge builds or installs
|
||||
them, it has drifted. This resolves the apparent contradiction between "do NOT give workers IDE MCP
|
||||
access" and "give workers a sandboxed IDE" — different owners.*
|
||||
|
||||
## 8. Staging & dependencies
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
cb402["CB-401/402<br/>Peer Launcher SPI + composite<br/>(DONE / in-flight)"]
|
||||
A["A · SandboxLauncher<br/>(placement-neutral, independent)"]
|
||||
cb308["CB-308 substrate<br/>per-agent channels + global id<br/>+ federated roster"]
|
||||
B["B · main-agent pair<br/>(multi-slot PrimaryRegistry)"]
|
||||
C["C · orchestrator tier<br/>(SessionManager recursed up)"]
|
||||
cb402 --> A
|
||||
cb402 --> cb308
|
||||
cb308 --> B
|
||||
cb308 --> C
|
||||
B --> C
|
||||
A -.->|"if sandbox = separate host/container,<br/>reuses CB-308 global id"| cb308
|
||||
classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff;
|
||||
class cb308 gate
|
||||
```
|
||||
|
||||
*Figure 10 — the substrate (amber) is the shared enabler for B and C. Recommended order:*
|
||||
|
||||
1. **Finish CB-402** (merge the opencode adapter branch).
|
||||
2. **A · SandboxLauncher** — independent; a second proof of the SPI (placement-neutral). Ships anytime.
|
||||
3. **CB-308 substrate** — per-agent channels + global id + federated roster (the multi-host work,
|
||||
promoted from host-to-host to tier-to-tier).
|
||||
4. **B · main-agent pair** — multi-slot `PrimaryRegistry` + per-main inbox routing, on the substrate.
|
||||
5. **C · orchestrator tier** — the capstone; the recursive session manager + context scoping.
|
||||
|
||||
## 9. Open questions (to resolve at ticket-split)
|
||||
|
||||
- **Sandbox mechanism:** container (`docker exec`) vs devcontainer — how the role→image mapping is
|
||||
expressed on the profile. *(Topology **resolved** in §11: distributed = gateway-per-host × local
|
||||
sandboxes; the remaining choice is only the local launch mechanism, not the shape.)*
|
||||
- **Pair semantics:** are the two mains fully symmetric peers, or is one a co-primary that may also
|
||||
delegate? Affects how `PrimaryRegistry` and the "orchestration tools" identity relax.
|
||||
- **Orchestrator drivenness:** the mains become programmatically spawned/resumed — does the human
|
||||
still ever type directly into a main, or only into the orchestrator? (The re-root caveat, Fig 7.)
|
||||
- **Context-slice selection:** who decides the slice — orchestrator heuristics, explicit tool args,
|
||||
or the main pulling on demand? This is the genuinely new responsibility; keep it *scoping*, not
|
||||
content authorship, to stay inside the boundary.
|
||||
- **Trust:** every new tier boundary that accepts spawn/send is a trust edge (the CB-308 item #5 /
|
||||
CB-401 Stage-C concern) — orchestrator→main and main→sandbox both need authz.
|
||||
|
||||
## 10. Ticket-split guidance (deferred)
|
||||
|
||||
This note is deliberately one arc; when split, the natural tickets are **A** (SandboxLauncher +
|
||||
role/spawn-target), **the CB-308 substrate** (likely already its own ticket), **B** (multi-primary
|
||||
pair), and **C** (orchestrator tier + context scoping) — with the identity clause (§7) as an
|
||||
acceptance criterion on **A** specifically. Sequence per §8; nothing here is a new pillar, so each
|
||||
ticket is an extension of an existing pattern (CB-402 for A, CB-308 for B/C).
|
||||
|
||||
## 11. Distributed sandboxes — the resolved topology
|
||||
|
||||
The follow-up question — *"clarify the architecture when we have distributed agents in sandboxes"* —
|
||||
resolves the fork left open in §4 and §9. **Decision: Development A (sandbox launcher) and CB-308
|
||||
(per-host federation) *compose*, not compete — each host runs a `bridged` gateway whose launcher
|
||||
spawns agents into that host's *local* sandboxes.** A sandbox is never reached across the network; it
|
||||
is reached by the gateway sitting next to it.
|
||||
|
||||
### 11.1 The one fact that fixes the shape
|
||||
|
||||
The bus delivers a turn by **herdr keystroke-injection** — `Injector → AgentControl.send` writes into
|
||||
a PTY that its **local** herdr owns. The broker moves *messages and presence*, **never keystrokes**.
|
||||
So an agent's PTY must live in a herdr that *some* `bridged` instance drives locally: a remote
|
||||
container with no local herdr **cannot be injected into**. That rules out a central daemon reaching
|
||||
remote PTYs, and collapses the design to a single identity:
|
||||
|
||||
> **"a sandboxed agent on another host" ≡ "a sandbox spawned by that host's gateway."**
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph hostA["HOST A — gateway"]
|
||||
mA["main / orchestrator<br/>MCP client → LOCAL gateway"]
|
||||
gA["bridged A<br/>herdr + CompositePeerLauncher<br/>(incl. SandboxLauncher)"]
|
||||
cBEa["sandbox: backend<br/>(local container)"]
|
||||
cFEa["sandbox: frontend<br/>(local container)"]
|
||||
mA --- gA
|
||||
gA -->|"spawn (docker/devcontainer)<br/>→ PTY in A's herdr"| cBEa
|
||||
gA --> cFEa
|
||||
end
|
||||
subgraph broker["BROKER (AMQP) — CB-307/308 fabric"]
|
||||
inbox["agent.ID.inbox queues"]
|
||||
roster["roster.* (federated presence)"]
|
||||
end
|
||||
subgraph hostB["HOST B — gateway"]
|
||||
gB["bridged B<br/>herdr + SandboxLauncher"]
|
||||
cBEb["sandbox: backend<br/>(local container)"]
|
||||
gB -->|"spawn → PTY in B's herdr"| cBEb
|
||||
end
|
||||
gA <-->|"messages + presence<br/>(NOT keystrokes)"| inbox
|
||||
gB <-->|"messages + presence"| inbox
|
||||
gA --- roster
|
||||
gB --- roster
|
||||
classDef line fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
|
||||
class inbox,roster line
|
||||
```
|
||||
|
||||
*Figure 11 — the composed topology. Each gateway owns its local herdr and runs a `SandboxLauncher`
|
||||
(the §4 adapter) that spawns role-specific containers **on its own host**; the broker (blue) carries
|
||||
only messages + roster between gateways. Keystroke-injection stays strictly local to each gateway.*
|
||||
|
||||
### 11.2 How a delegation reaches a sandboxed agent on another host
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant MA as main (host A)
|
||||
participant GA as gateway A
|
||||
participant BR as broker
|
||||
participant GB as gateway B
|
||||
participant SB as sandbox agent (host B, container)
|
||||
MA->>GA: bridge_send(globalId on B, msg)
|
||||
GA->>GA: directory lookup - is globalId local? NO
|
||||
GA->>BR: publish agent.ID.inbox (durable)
|
||||
BR->>GB: route to the owning gateway
|
||||
GB->>SB: inject via B's LOCAL herdr (keystrokes)
|
||||
Note over GB,SB: SandboxLauncher already spawned the container -<br/>its PTY is in B's herdr, CB-306 readiness passed
|
||||
SB-->>GB: bridge_reply (to B's LOCAL MCP endpoint)
|
||||
GB->>BR: publish primary-bound (durable, msg id)
|
||||
BR->>GA: route back to A
|
||||
Note over GA: held until the main pulls (the main is a client)
|
||||
MA->>GA: poll / blocking send resolves
|
||||
GA-->>MA: reply
|
||||
```
|
||||
|
||||
*Figure 12 — the `local ? inject : publish` fork (CB-308 §3.2) with a sandboxed far side. Only the
|
||||
**middle** hop crosses the network via the broker; **both** injection points (into the sandbox on B,
|
||||
and the drain-nudge back into the main on A) are local herdr writes. This is CB-308's routing rule
|
||||
unchanged — the sandbox is transparent to it.*
|
||||
|
||||
### 11.3 Two reachability changes any sandbox forces
|
||||
|
||||
| Change | Today | Under sandboxes |
|
||||
|---|---|---|
|
||||
| **`mcpUrl`** | `http://127.0.0.1:8765/mcp` (loopback) | must be **host-routable from inside the container** (e.g. `host.docker.internal` or the gateway's LAN IP) — the worker connects to **its own gateway's** MCP, never a remote one. |
|
||||
| **PTY ownership** | pane in the daemon's herdr | pane is the **container's** attached PTY, in the **local** gateway's herdr (via `docker exec`/devcontainer) — non-negotiable per §11.1. |
|
||||
|
||||
### 11.4 Everything maps to an existing seam (nothing new invented)
|
||||
|
||||
| Concern | Provided by |
|
||||
|---|---|
|
||||
| Per-host gateway (owns local herdr + sessions) | **CB-308** (today's `bridged`, evolved) |
|
||||
| Spawn into a local sandbox / role→image | **Development A** `SandboxLauncher` (§4), routed by `CompositePeerLauncher` |
|
||||
| Addressing a remote sandboxed agent | **CB-308** global id + federated roster (host + role as metadata) |
|
||||
| Orphan reap after a gateway restart | **CB-117** per-gateway, summed by the composite — each reaps only its **local** herdr |
|
||||
| Container up/down | tied to **CB-303** session lifecycle — `SandboxLauncher.stop` tears the container down with the pane |
|
||||
| Cross-gateway spawn/send trust | **CB-308 item #5** / CB-401 Stage-C — each gateway edge is a trust boundary |
|
||||
|
||||
*The net: distributed sandboxes add **zero** new pillars — they are `CB-308 gateway × Development-A
|
||||
launcher` at every host, with the §7 ownership line (peer owns the image; the bridge only launches
|
||||
into it) holding at each gateway.*
|
||||
@@ -0,0 +1,242 @@
|
||||
# CB-5xx — Stage 5 Hardening (auth · metrics · CI · supervision · authz+audit)
|
||||
|
||||
**Status:** design note (pre-implementation) — the single-host close-out before cross-host work.
|
||||
**Covers:** CB-501 (bearer auth + TLS) · CB-502 (`/metrics`) · CB-503 (mock-socket CI) ·
|
||||
CB-504 (service supervision) · CB-505 (per-session authz + audit log).
|
||||
**Depends on:** everything shipped through CB-402. Nothing here changes messaging semantics.
|
||||
**Blocks:** CB-308. The cross-host trust model is CB-308's own gating concern, and it inherits
|
||||
whatever identity/authz shape lands here — so this stage is deliberately *before* federation,
|
||||
not after it.
|
||||
|
||||
---
|
||||
|
||||
## 1. Why this stage is not optional bookkeeping
|
||||
|
||||
`bridged` today has **exactly one security control: the loopback bind**. Every other guarantee
|
||||
rests on it.
|
||||
|
||||
The identity model (`mcp/ConnectionIdentity.java`) resolves a caller from the connection alone —
|
||||
the OS reports the connecting PID, herdr owns the PID→pane map, so a worker cannot forge another
|
||||
worker. Its own javadoc is explicit: *"Single-host only (the herd shares the `bridged` host); the
|
||||
token path is the split-host fallback."* The token path does not exist yet.
|
||||
|
||||
That leaves a seam that is **latent today and load-bearing the moment the bind moves**:
|
||||
|
||||
```java
|
||||
// ConnectionIdentity.resolve — non-loopback callers get a null terminal
|
||||
if (!isLoopback(remoteAddr)) return new Caller(null, -1);
|
||||
```
|
||||
|
||||
…and `null` terminal is interpreted downstream as **"this caller is the primary"**. Combined:
|
||||
|
||||
> Any caller that is not a recognised on-host worker pane is treated as the primary — including,
|
||||
> if `bind.host` is ever widened, an arbitrary remote client.
|
||||
|
||||
Today `bind` defaults to `127.0.0.1` so this is unreachable. But CB-308 exists precisely to widen
|
||||
the boundary, and the primary is the *most* privileged role on the bus (it spawns, stops, sends to
|
||||
any session, and drains any inbox). Shipping federation on top of "unauthenticated ⇒ primary"
|
||||
would be building the security boundary backwards.
|
||||
|
||||
**So CB-501 is not "add a token header". It is: make identity explicit, and make the absence of
|
||||
identity mean *nothing*, not *everything*.**
|
||||
|
||||
---
|
||||
|
||||
## 2. Decisions (locked)
|
||||
|
||||
### D1 — Three caller roles, one resolution path
|
||||
|
||||
Introduce `Role { PRIMARY, WORKER, ANONYMOUS }` resolved by a single `CallerResolver` that both
|
||||
REST and MCP go through. Resolution order:
|
||||
|
||||
1. **Connection identity wins where it applies.** A loopback peer PID that maps to a herdr worker
|
||||
pane ⇒ `WORKER` with that terminal. Unforgeable, unchanged from today, zero config.
|
||||
2. **Token, if presented.** A valid bearer token ⇒ the role that token is provisioned for.
|
||||
3. **Otherwise `ANONYMOUS`** — *not* `PRIMARY`.
|
||||
|
||||
This inverts today's default. `PRIMARY` becomes something you must *prove* (by being a loopback
|
||||
non-worker process when auth is disabled, or by presenting a primary-scoped token when it is
|
||||
enabled), rather than something you get by failing every other check.
|
||||
|
||||
### D2 — Auth is opt-in by config, but the *default* must stay zero-friction on loopback
|
||||
|
||||
The daemon is dogfooded constantly on one machine. If enabling hardening breaks the local setup,
|
||||
it will be disabled and the stage is wasted. So:
|
||||
|
||||
```yaml
|
||||
auth:
|
||||
mode: loopback-trust # default — behaves exactly like today: loopback ⇒ PRIMARY, no token needed
|
||||
# mode: token # every non-worker caller must present a valid bearer token
|
||||
# tokenEnv: BRIDGED_API_TOKEN # host env var holding the token; never the literal value
|
||||
```
|
||||
|
||||
`mode: loopback-trust` is the current behaviour, named honestly and now *chosen* rather than
|
||||
implied. `mode: token` is what a non-loopback bind requires. **A non-loopback `bind.host` with
|
||||
`mode: loopback-trust` must fail fast at startup** — that check is the single highest-value line
|
||||
in this stage, because it makes the dangerous configuration unrepresentable rather than merely
|
||||
discouraged.
|
||||
|
||||
### D3 — TLS terminates *outside* the JVM
|
||||
|
||||
Do **not** add TLS config to Javalin/Jetty. The deployment story for a cross-host gateway is a
|
||||
reverse proxy (or an SSH/WireGuard tunnel) in front of the daemon; the AMQP link has its own TLS
|
||||
via the broker URI (`amqps://`). Adding keystore handling here would mean certificate lifecycle
|
||||
code in a daemon whose whole value is being small, and would duplicate what the proxy does better.
|
||||
|
||||
**CB-501 therefore ships bearer auth + the fail-fast bind check, and documents TLS as a
|
||||
deployment concern with a worked reverse-proxy example.** This is a deliberate narrowing of the
|
||||
roadmap's "auth/TLS" wording — flagged in §6 for the lead.
|
||||
|
||||
### D4 — Metrics without a new dependency
|
||||
|
||||
The roadmap's tech-stack table says Micrometer→Prometheus. Recommend **not** taking that dep:
|
||||
|
||||
- This pom already carries an unusually heavy dependency-reconciliation burden (a hand-pinned
|
||||
`jackson-annotations` 3.0-rc5 to reconcile the MCP SDK's Jackson 3 with our Jackson 2.19, a
|
||||
Jetty BOM import to stop version skew, plus four documented accepted-CVE advisories). Every new
|
||||
transitive tree is a real cost here, not a hypothetical one.
|
||||
- The CVE gate that CLAUDE.md mandates for dependency changes (`jetbrains get_file_problems` →
|
||||
Mend.io) **cannot currently be run** — no JetBrains MCP server is connected. Adding a dependency
|
||||
tree we cannot scan violates the project's own stated policy.
|
||||
- The metric set is small and fully known (§4). Prometheus text exposition is a trivial,
|
||||
stable, well-specified format.
|
||||
|
||||
So: a ~120-line `metrics/Metrics.java` holding `LongAdder` counters and gauge suppliers, rendered
|
||||
to the Prometheus text format at `GET /metrics`. If Micrometer is wanted later for its
|
||||
registry/push ecosystem, this stays a drop-in swap behind the same endpoint. **Flagged in §6 —
|
||||
this deviates from a documented tech-stack choice.**
|
||||
|
||||
### D5 — Supervision targets launchd first, systemd second
|
||||
|
||||
The roadmap says "systemd unit". **This host is macOS — there is no systemd on it** (`systemctl`
|
||||
not found), and the daemon that has been dogfooded for weeks runs as a bare foreground
|
||||
`java -jar`. Ship **both**:
|
||||
|
||||
- `deploy/dev.ltms.bridged.plist` — launchd agent, the *actual* runtime here, with `KeepAlive` and
|
||||
ordered start after herdr.
|
||||
- `deploy/bridged.service` — systemd unit for the Linux gateways CB-308 introduces.
|
||||
|
||||
Ordering after herdr is advisory in both: the herdr socket may not exist at boot, so the daemon
|
||||
must **retry the socket rather than exit** — supervision ordering is a nicety, socket-retry is the
|
||||
actual fix. That retry behaviour is part of CB-504, not a separate ticket.
|
||||
|
||||
### D6 — Audit log is a separate append-only stream, not the app log
|
||||
|
||||
Privileged actions (spawn, stop, send, reply-drain, ack) emit a structured JSON line to a
|
||||
dedicated `audit` SLF4J logger with its own appender, carrying `{ts, role, terminal, pid, action,
|
||||
target, outcome}`. Keeping it off the chatty app logger is what makes it greppable and, later,
|
||||
shippable. **No message *content* in the audit record** — the bridge carries the user's source
|
||||
code and prompts; an audit trail that quietly becomes a transcript archive is a liability, not a
|
||||
control. Content stays out; correlation ids go in.
|
||||
|
||||
---
|
||||
|
||||
## 3. Authorization model (CB-505)
|
||||
|
||||
With D1's roles, the rules are small enough to state completely:
|
||||
|
||||
| Action | REST | PRIMARY | WORKER | ANONYMOUS |
|
||||
|---|---|---|---|---|
|
||||
| spawn worker | `POST /workers` | ✅ | ❌ | ❌ |
|
||||
| stop worker | `DELETE /workers/{paneId}` | ✅ | ❌ | ❌ |
|
||||
| send to a session | `POST /sessions/{id}/message` | ✅ | ❌ | ❌ |
|
||||
| reply | `POST /sessions/{id}/reply` | ❌ | ✅ **own session only** | ❌ |
|
||||
| ask | `POST /sessions/{id}/ask` | ❌ | ✅ **own session only** | ❌ |
|
||||
| drain replies | `GET /sessions/{id}/replies` | ✅ | ❌ | ❌ |
|
||||
| status / list / profiles | `GET …` | ✅ | ✅ | ❌ |
|
||||
| health | `GET /healthz` | ✅ | ✅ | ✅ (unauthenticated by design) |
|
||||
| metrics | `GET /metrics` | ✅ | ✅ | ❌ |
|
||||
|
||||
The load-bearing row is **"own session only"**: a worker may only reply or ask *as itself*. That is
|
||||
already true de facto — `ConnectionIdentity` derives the terminal rather than reading it from the
|
||||
body — so CB-505 mostly **asserts an existing invariant explicitly** and adds the test that pins
|
||||
it. The one real change is rejecting a worker that names a *different* session id in the path.
|
||||
|
||||
`/healthz` stays open: it must answer for a load balancer or supervisor before any credential is
|
||||
configured. It already leaks nothing but herdr's version and up/down.
|
||||
|
||||
### 3.1 There are TWO entry paths, and only one of them has identity today
|
||||
|
||||
The wiki describes MCP as "a thin adapter over the REST core". **At the code level that is not
|
||||
literally true, and the difference is security-relevant.** `BridgeMcp` calls `MessageService` /
|
||||
`SessionManager` *directly*; it never issues an HTTP request against a Javalin route. And `/mcp` is
|
||||
mounted as a raw servlet on Jetty's `ServletContextHandler`
|
||||
(`BridgedApp.build → cfg.jetty.modifyServletContextHandler`), so it does **not** pass through
|
||||
Javalin's `before` filters at all.
|
||||
|
||||
The current split is the mirror image of what you'd expect:
|
||||
|
||||
| Path | Caller identity today | Authz today |
|
||||
|---|---|---|
|
||||
| MCP `/mcp` | ✅ resolved per call (`ConnectionIdentity` via the transport-context extractor) | ❌ none |
|
||||
| REST routes | ❌ **none at all** — the session id is taken from the URL path and trusted | ❌ none |
|
||||
|
||||
So REST is the *more* exposed surface: `POST /sessions/{id}/reply` accepts any `{id}` from the
|
||||
path, whereas the MCP `bridge_reply` derives the worker from the connection and refuses to read it
|
||||
from an argument. Loopback-only bind is what makes this safe today.
|
||||
|
||||
**Therefore CB-505 must enforce on both paths against one shared resolver** — not at a single
|
||||
choke point. Concretely: a Javalin `before` filter for REST, and the existing transport-context
|
||||
extractor for MCP, both delegating to `auth.CallerResolver`. Any authz check that lives in only
|
||||
one of the two is not a control.
|
||||
|
||||
---
|
||||
|
||||
## 4. Metric set (CB-502)
|
||||
|
||||
Deliberately small; every one maps to a failure mode we have actually hit.
|
||||
|
||||
| Metric | Type | Why it exists |
|
||||
|---|---|---|
|
||||
| `bridged_sends_total{outcome}` | counter | outcome ∈ replied\|completion_fallback\|timeout\|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118) |
|
||||
| `bridged_send_duration_seconds` | histogram | delegated turn latency |
|
||||
| `bridged_replies_total{path}` | counter | path ∈ rendezvous\|inbox — how often a reply strands (CB-307's whole reason to exist) |
|
||||
| `bridged_inbox_depth{target}` | gauge | undrained replies; steady-state should be 0 |
|
||||
| `bridged_push_nudges_total{outcome}` | counter | outcome ∈ delivered\|exhausted — a rising `exhausted` means the primary is not draining |
|
||||
| `bridged_spawns_total{kind,outcome}` | counter | outcome ∈ ready\|timeout\|guard_rejected; per peer kind (CB-402) |
|
||||
| `bridged_sessions{state}` | gauge | SPAWNING/READY/BUSY/DONE census |
|
||||
| `bridged_herdr_calls_total{method,outcome}` | counter | socket health — the dependency everything rests on |
|
||||
| `bridged_auth_failures_total{reason}` | counter | only meaningful once CB-501 lands; catches misconfigured workers |
|
||||
|
||||
---
|
||||
|
||||
## 5. Increment plan
|
||||
|
||||
Ordered so each step is independently mergeable and the risky one lands first.
|
||||
|
||||
1. **CB-501a — `CallerResolver` + `Role`.** Pure refactor: route today's connection identity through
|
||||
the new type, `ANONYMOUS` not yet reachable (loopback-trust default preserves behaviour).
|
||||
Green build, no behaviour change.
|
||||
2. **CB-501b — token mode + fail-fast bind check.** Config block, bearer parsing, the
|
||||
non-loopback-bind guard. This is the security-relevant commit; keep it small and reviewable.
|
||||
3. **CB-505 — authz table + audit logger.** Enforce §3 on **both** entry paths (see §3.1); add the
|
||||
audit appender.
|
||||
4. **CB-502 — `Metrics` + `/metrics`.** Instrument the paths in §4.
|
||||
5. **CB-503 — CI.** Runs `mvn -B clean install` with `-Dgroups='!contract'` so the live-herdr and
|
||||
RabbitMQ contract tests are excluded; the mock-UDS suite is the CI surface, exactly as the
|
||||
roadmap's testability section intends.
|
||||
6. **CB-504 — launchd plist + systemd unit + herdr-socket retry.**
|
||||
|
||||
---
|
||||
|
||||
## 6. Open questions for the lead
|
||||
|
||||
*All four resolved 2026-07-29 — the lead confirmed D3, D4, and the D6 sink; the CI runner question
|
||||
was answered from the forge itself. Kept here as the decision record.*
|
||||
|
||||
1. ✅ **TLS scope (D3) — confirmed.** Bearer auth + the fail-fast bind guard ship in the daemon;
|
||||
TLS terminates at a reverse proxy, documented with a worked example. AMQP gets TLS via an
|
||||
`amqps://` URI. No keystore handling in `bridged`.
|
||||
2. ✅ **Micrometer (D4) — confirmed dropped.** Zero-dependency Prometheus text renderer, for the
|
||||
reasons in D4 (pom reconciliation burden + the mandated CVE gate being un-runnable this
|
||||
session). Revisit if a push-gateway or JVM-metrics requirement appears; the endpoint is the
|
||||
swap seam.
|
||||
3. ~~**CI runner (CB-503).**~~ ✅ **Resolved during design** — a Gitea Actions runner *is*
|
||||
registered and healthy (`lms/alms-memory` has 28 completed runs; `lms/alms` runs on push and
|
||||
pull_request). CB-503 targets `.gitea/workflows/ci.yml` with `runs-on: ubuntu-latest`, matching
|
||||
the sibling repo's convention. Note the runner's image ships an older `default-jdk`, so the
|
||||
workflow must provision **JDK 25** explicitly rather than apt-installing the default.
|
||||
Contract-test exclusion needs no CI flag: the pom's `default-excludes` profile already sets
|
||||
`excludedGroups=contract`, so a plain `mvn -B clean install` *is* the mock-socket surface.
|
||||
4. ✅ **Audit sink — confirmed dedicated file.** Its own logback appender writing JSON lines beside
|
||||
the daemon, separate from the app log, per D6. Content still never enters the record.
|
||||
+1
-1
Submodule wiki updated: 8e5fd01ac9...0c896eb49b
Reference in New Issue
Block a user