Compare commits
50 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 3aa69a9e32 | |||
| e056c7e1fa | |||
| d63273d082 | |||
| 0efb65ca0a | |||
| 9fe04bfb08 | |||
| 09d3948acf | |||
| 954351a80b | |||
| 8d51066ddd | |||
| 84102baab4 | |||
| 64e70efdf1 | |||
| 97ecc7136e | |||
| f9073e2320 | |||
| 54d907c314 | |||
| 82c7d6553a | |||
| 19b10e3216 | |||
| 0f79e6bed5 | |||
| aa0cf814ff | |||
| 4aa9d03da2 | |||
| 37abfdfd6e | |||
| d5c3ede215 | |||
| 426855e378 | |||
| 358c6970b5 | |||
| 55ebd5b949 | |||
| 2a61fe69f1 | |||
| ff6aacdc78 | |||
| 5f0ec034d9 | |||
| 31e34d177b | |||
| b2d85af78b | |||
| 968a5c68b6 | |||
| 3c05823f19 | |||
| 2fb46f670f | |||
| 37f7ad4185 | |||
| 926724a279 | |||
| 959f04bc96 | |||
| 7f1b6b3a0d | |||
| ab771ea24e | |||
| a629a7ee73 | |||
| 2052929768 | |||
| a97c287aee | |||
| 8ed2370fbe | |||
| 5b26caca0c | |||
| 6988bfe88f | |||
| 07722a6007 | |||
| fa570ab32a | |||
| 46f87b9cd3 | |||
| 8ffbfd1f06 | |||
| 597ac2e562 | |||
| bc04637694 | |||
| c7b58f2195 | |||
| 01cfda7965 |
@@ -0,0 +1,126 @@
|
||||
---
|
||||
name: implementer
|
||||
description: Implementer-role playbook for a bridged worker — you are in an isolated git worktree on a dedicated branch; implement the assigned task, commit, push, open your own PR to main, and hand off the PR URL via bridge_reply. You never merge. Load this when you have been delegated an implementation task over bridged.
|
||||
---
|
||||
|
||||
# Implementer worker
|
||||
|
||||
You are an **implementer** in the claude-bridge fleet. The lead delegated you one scoped task,
|
||||
and you are running in an **isolated git worktree on your own branch** — a full peer of the
|
||||
primary (same `CLAUDE.md`, skills, memory, MCP), differing only in the model behind you and the
|
||||
branch you sit on. Your job for this turn: **implement the task, then hand off a PR the lead can
|
||||
review and merge.** You do the work; the lead (or human) is the merge gate — you never merge.
|
||||
|
||||
Delivery mechanics (how the task reached you, how your reply resolves the lead's blocked send)
|
||||
are in [`docs/MCP-Contract.md`](../../../docs/MCP-Contract.md); the worktree/PR model is in
|
||||
[`docs/Worker-Git-Workflow.md`](../../../docs/Worker-Git-Workflow.md). You only need the steps
|
||||
below.
|
||||
|
||||
## 1. Confirm where you are — a worktree on a dedicated branch
|
||||
|
||||
Before touching anything, verify your ground truth:
|
||||
|
||||
```bash
|
||||
git rev-parse --show-toplevel # your worktree root — NOT the primary's main tree
|
||||
git branch --show-current # your dedicated branch: worker/<ticket>-<nonce>
|
||||
git status # should be clean at the start
|
||||
```
|
||||
|
||||
Do **all** work here, on this branch. **Never** switch to `main`, never `git checkout main`,
|
||||
never rebase onto or push to `main` directly. The branch is your isolation — respect it.
|
||||
|
||||
## 2. Implement the task
|
||||
|
||||
- Implement exactly the scope the lead named. Keep changes focused; if you notice something out
|
||||
of scope, note it in your reply rather than expanding the diff.
|
||||
- Match the surrounding code's style, naming, and idioms. Follow project `CLAUDE.md`.
|
||||
- **You cannot run the IDE MCP tools** (intellij-index / jetbrains are the primary's, not yours).
|
||||
So **never claim a file is "IDE-clean" or "diagnostics-clean"** — you cannot verify that. State
|
||||
only what you actually ran (e.g. `mvn`, a test) and its real output. A fabricated clean claim is
|
||||
worse than an honest "I could not verify inspections here."
|
||||
- Run whatever build/test you can and **report the true result** — including failures.
|
||||
|
||||
## 3. Commit — focused, and never the excluded files
|
||||
|
||||
```bash
|
||||
git add <the files you changed>
|
||||
git commit -m "<ticket>: <clear one-line summary>"
|
||||
```
|
||||
|
||||
**Excluded from every commit, always:** `.mcp.json` (the primary's local, session-modified copy —
|
||||
present only for parity) and `wiki/` (a separate submodule). Stage files explicitly; do **not**
|
||||
`git add -A` / `git add .` blindly, or you risk staging them. If `.mcp.json` shows as modified,
|
||||
leave it — it is flagged `--skip-worktree` and is not yours to commit.
|
||||
|
||||
## 4. Push your branch
|
||||
|
||||
```bash
|
||||
git push -u origin HEAD
|
||||
```
|
||||
|
||||
Push is over SSH as the same user — no extra credential needed. Push the branch as-is; do not
|
||||
force-push over anything you did not create.
|
||||
|
||||
## 5. Open your own PR to `main`
|
||||
|
||||
Open the PR via the gitea REST API. The daemon injected a **repo-scoped token** (`GITEA_TOKEN`)
|
||||
and the forge host (`GITEA_HOST`) into your env for exactly this — the token can create a PR but
|
||||
**cannot merge** (that stays the lead/human gate).
|
||||
|
||||
```bash
|
||||
API="${GITEA_HOST%/}/api/v1/repos/lms/claude-bridge/pulls"
|
||||
BRANCH="$(git branch --show-current)"
|
||||
curl -sS -X POST "$API" \
|
||||
-H "Authorization: token ${GITEA_TOKEN}" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$(cat <<JSON
|
||||
{"head": "${BRANCH}", "base": "main",
|
||||
"title": "<ticket>: <concise change summary>",
|
||||
"body": "<what changed and why; reference the ticket; note tests run and their result>"}
|
||||
JSON
|
||||
)"
|
||||
```
|
||||
|
||||
The response JSON includes `"html_url"` — that is your PR URL. If the call fails (non-2xx), read
|
||||
the error body, fix the cause if it is yours (e.g. branch not pushed yet), and report the failure
|
||||
honestly in your reply rather than inventing a URL. If `GITEA_TOKEN` is unset, your profile was
|
||||
not granted PR-create — push the branch (step 4) and report the branch name so the lead opens the
|
||||
PR.
|
||||
|
||||
## 6. Reply via `bridge_reply` — the PR is the handoff
|
||||
|
||||
End your turn with **exactly one** `bridge_reply`. That reply is the entire handoff — the lead
|
||||
cannot see your terminal. Include:
|
||||
|
||||
```
|
||||
PR: <html_url from step 5, or "not created: <reason>" + branch name>
|
||||
branch: <your branch>
|
||||
files: <the files you changed>
|
||||
tests: <what you ran and its REAL result — or "not run: <why>">
|
||||
summary: <2-3 lines: what you implemented and any caveat the reviewer needs>
|
||||
```
|
||||
|
||||
Then stop. **Do not merge. Do not touch `.mcp.json` or `wiki/`.** One reply closes the turn.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant L as Lead
|
||||
participant B as bridged
|
||||
participant I as Implementer (you)
|
||||
participant G as git / gitea
|
||||
|
||||
L->>B: bridge_send(task) — blocks
|
||||
B-->>I: your assignment (in a worktree on your branch)
|
||||
I->>I: implement + build/test here
|
||||
I->>G: git commit (never .mcp.json / wiki)
|
||||
I->>G: git push -u origin HEAD
|
||||
I->>G: POST /pulls (GITEA_TOKEN) — open PR to main
|
||||
G-->>I: html_url
|
||||
I->>B: bridge_reply(PR url, branch, files, tests)
|
||||
B-->>L: { outcome:"reply", text }
|
||||
Note over L,G: lead reviews the PR, merges on green — you never merge
|
||||
```
|
||||
|
||||
*The implement turn: work in the worktree, commit → push → open the PR, hand off the URL. The
|
||||
lead is the merge gate.*
|
||||
@@ -0,0 +1,84 @@
|
||||
---
|
||||
name: reviewer
|
||||
description: Reviewer-role playbook for a bridged worker — read the assigned scope, find the real issues, ask the lead via bridge_ask when a decision is genuinely theirs, and report the finding via bridge_reply. Load this when you have been delegated a code review over bridged.
|
||||
---
|
||||
|
||||
# Reviewer worker
|
||||
|
||||
You are a **reviewer** in the claude-bridge fleet. The lead delegated you one scoped review
|
||||
over `bridged`, and your whole job is **this single turn**: examine the scope it named, and
|
||||
report back. You are not the owner of the code and you do not merge anything — you surface
|
||||
what the owner needs to know, then hand the turn back.
|
||||
|
||||
Delivery mechanics (how the task reached you, how your reply resolves the lead's blocked
|
||||
send) are in [`docs/MCP-Contract.md`](../../../docs/MCP-Contract.md); you only need the three
|
||||
rules below.
|
||||
|
||||
## 1. Read the whole scope before you judge
|
||||
|
||||
The delegation names your scope — a file, a diff, a PR, a function. **Read all of it first.**
|
||||
A review that fires on a snippet misses the caller that makes it safe (or the one that makes
|
||||
it a bug). Reviewing only part of the scope and guessing the rest is the most common way a
|
||||
reviewer worker is wrong.
|
||||
|
||||
## 2. Stay in your lane
|
||||
|
||||
- Review **only** the assigned scope. If you notice something elsewhere, mention it in one
|
||||
line — do **not** go hunt it. Wandering is how two workers end up reporting the same thing
|
||||
and neither covers what it was given.
|
||||
- Do **not** edit files, run the build, or spawn other workers. You review; the owner acts.
|
||||
- You never set `ANTHROPIC_BASE_URL` and never touch herdr — you are a Claude Code process,
|
||||
not part of the transport.
|
||||
|
||||
## 3. When the decision is the lead's — ask, don't guess
|
||||
|
||||
Some things you cannot resolve from the code: an ambiguous requirement, a missing acceptance
|
||||
criterion, "is this behavior intended or a bug?", or a choice between two defensible fixes.
|
||||
Guessing there produces a confident-but-wrong finding. Instead **pause and ask the lead** with
|
||||
`bridge_ask` — a single crisp question. The call blocks; when the lead answers you **resume
|
||||
the same turn** with the answer and finish. Ask only when the answer changes your finding;
|
||||
don't narrate options you could decide yourself.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant L as Lead
|
||||
participant B as bridged
|
||||
participant R as Reviewer (you)
|
||||
|
||||
L->>B: bridge_send(review scope) — blocks
|
||||
B-->>R: your assignment
|
||||
R->>R: read the full scope
|
||||
opt a decision only the lead can make
|
||||
R->>B: bridge_ask("intended, or a bug?") — you block
|
||||
B-->>L: { outcome:"question", turn_id }
|
||||
L->>B: bridge_send(answer, turn_id)
|
||||
B-->>R: { answer } — you resume the SAME turn
|
||||
end
|
||||
R->>B: bridge_reply(structured finding) — ends your turn
|
||||
B-->>L: { outcome:"reply", text }
|
||||
```
|
||||
|
||||
*The review turn, with the optional `bridge_ask` detour when the call is the lead's to make.*
|
||||
|
||||
## 4. Report with `bridge_reply` — one structured finding
|
||||
|
||||
End your turn with **exactly one** `bridge_reply`. Report the **single most important** real
|
||||
issue in the scope, in these four lines, under ~90 words:
|
||||
|
||||
```
|
||||
1. <path>:<line>
|
||||
2. issue: <one sentence — what is wrong and why it matters>
|
||||
3. fix: <one line — the concrete change>
|
||||
4. severity: high | medium | low
|
||||
```
|
||||
|
||||
- Found nothing real after reading? Reply `NO ISSUE` and one line saying why — a clean review
|
||||
is a valid result, and a fabricated issue is worse than none.
|
||||
- **Severity:** `high` = wrong result, data loss, security, or a hang/crash on a real path ·
|
||||
`medium` = a real bug on an edge path, or a correctness risk under load/concurrency ·
|
||||
`low` = clarity, a latent foot-gun, or a smell with no current failure.
|
||||
- Be specific and verifiable: a line number and a one-line repro beat an adjective. If you
|
||||
can't point to where it goes wrong, you haven't found it yet.
|
||||
|
||||
One reply closes the turn. If you asked mid-turn, the answer you got is already folded into
|
||||
this finding — you do not ask again after replying.
|
||||
@@ -20,6 +20,15 @@ module is **`bridged`**. Always pass these to IDE MCP tools:
|
||||
test run). A per-file-clean file can still break the build or another module. This is the
|
||||
whole-project gate before declaring work done or committing.
|
||||
|
||||
**Whenever dependencies change (or a `pom.xml` edit), validate CVEs with
|
||||
`jetbrains get_file_problems{filePath: "bridged/pom.xml"}`** — its Mend.io check reflects the
|
||||
dependencies on disk. (Note: `ide_diagnostics` / intellij-index does NOT re-resolve dependencies
|
||||
after a pom edit without a full Maven reimport, so it reports stale CVE results — don't trust it
|
||||
for this.) Treat a CVE warning like any other: bump to a patched version and confirm
|
||||
`mvn clean install` still passes. If the latest available version is still flagged (EOL line,
|
||||
"insufficient information", or config-file-only advisories), document it as accepted in the pom
|
||||
rather than chasing a fix that doesn't exist.
|
||||
|
||||
### Use IDE MCP tools for navigation, refactoring, and diagnostics only
|
||||
|
||||
- **Navigate (prefer over Grep/Read for symbols):** `ide_find_definition`, `ide_find_class`,
|
||||
|
||||
@@ -50,8 +50,9 @@ flowchart LR
|
||||
plain daemon (no Anthropic quota), so it may poll/subscribe freely.
|
||||
- **One gateway (unified MCP setup):** `bridged` is the **sole communication path** for every
|
||||
Claude session. Primary and workers each mount it as an MCP server (one `claude mcp add`
|
||||
line, same on both) and talk over MCP tools — `bridge_send` / `bridge_reply` / `bridge_ask` /
|
||||
`bridge_status`. **No Claude session ever addresses a broker, a peer, or the network
|
||||
line, same on both) and talk over MCP tools — `bridge_send` / `bridge_reply` /
|
||||
`bridge_status` (with `bridge_ask` planned for the blocked-worker path). **No Claude session
|
||||
ever addresses a broker, a peer, or the network
|
||||
directly**; any queue is `bridged`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so
|
||||
mounting the bridge is subscription-safe by construction.
|
||||
- **How the primary consumes a reply:** a single **blocking MCP call** (`bridge_send`);
|
||||
@@ -85,6 +86,27 @@ Gitea wiki.
|
||||
|
||||
## Status
|
||||
|
||||
🟢 Design — herdr-centric **`bridged`** message server selected as the primary approach
|
||||
(2026-07-11), superseding the AgentAPI plan (2026-07-08). AgentAPI retained as fallback
|
||||
injector.
|
||||
🟢 **Implemented & dogfooded** — the herdr-centric **`bridged`** message server is built and in
|
||||
real use: an Opus primary delegates tasks to off-subscription workers that reply through the
|
||||
bridge (code reviews delegated this way have produced committed bug fixes). Selected as the
|
||||
primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a
|
||||
fallback injector.
|
||||
|
||||
**Shipped** (Java 25 · Maven · 105 tests green — unit/acceptance + live-herdr contract tests):
|
||||
|
||||
- **Core gateway** — herdr socket client (contract-tested vs live 0.7.0); guard-checked worker
|
||||
spawn with `ANTHROPIC_BASE_URL` injected only into the worker's env; status-gated injector;
|
||||
blocking `bridge_send` with reply rendezvous; MCP server as a thin adapter over the REST core.
|
||||
- **MCP tools** — `bridge_send` / `bridge_reply` / `bridge_status` (messaging) and `bridge_spawn`
|
||||
/ `bridge_list` / `bridge_stop` / `bridge_profiles` / `bridge_poll` (fleet). Caller identity is
|
||||
connection-based (loopback peer PID → herdr pane), so the same mount serves primary and workers.
|
||||
- **Delivery reliability** — completion fallback (a confirmed `working→idle` turn resolves a
|
||||
send); async fire-and-poll (beats the caller's MCP call timeout for long tasks); and failure
|
||||
detection for wedged (`unknown`), vanished, and never-ready workers so a send never hangs.
|
||||
- **Fleet** — multiple worker profiles, each with an independent base_url guard check; workers
|
||||
inherit the primary's working directory (never `$HOME`); a readiness gate holds delivery until
|
||||
a worker's Claude has connected the bridge MCP (no paste lost into its boot window).
|
||||
|
||||
**Next** (see the [roadmap](wiki/8-Roadmap.md)) — structured envelope schema, `bridge_ask`
|
||||
(blocked-worker path), session lifecycle / recycle / `idle_ttl`, split-host, and hardening
|
||||
(auth/TLS, `/metrics`, CI, systemd).
|
||||
|
||||
@@ -6,28 +6,60 @@
|
||||
# REST + MCP listen address. Keep it on loopback — bridged is same-host in Stage-1.
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8080
|
||||
port: 8765
|
||||
|
||||
# herdr Unix socket. Omit to use the client default
|
||||
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
|
||||
# How a worker session is spawned. Stage-1 uses the existing ccs `ltms-local`
|
||||
# profile, whose .claude.json routes to the gx00 vLLM below.
|
||||
worker:
|
||||
profile: ltms-local
|
||||
baseUrl: http://gx00.gw:8000 # the gx00 vLLM (models: coder / deepseek-v4-flash)
|
||||
model: coder
|
||||
# Placement: each worker lands in its OWN tab inside a dedicated worker space, so it
|
||||
# never splits or clutters your real work spaces. Use `pane` for the legacy behaviour
|
||||
# (split the currently-focused tab).
|
||||
placement: tab # tab | pane
|
||||
workspace: bridged-workers # the dedicated worker space (found-or-created, shared)
|
||||
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
|
||||
# How worker sessions are spawned. Define one or more named profiles (backends) under
|
||||
# `workers`; each key is the profile name (also the ccs profile). `defaultWorker` picks
|
||||
# which one a no-argument spawn uses (bridge_spawn with no profile / POST /workers).
|
||||
#
|
||||
# Shared knobs (placement/workspace/tabLabel) can be repeated per profile; they usually match.
|
||||
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
|
||||
# Use `pane` for the legacy behaviour (split the focused tab).
|
||||
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
|
||||
# (--append-system-prompt) as launch flags; nothing is written to the profile.
|
||||
# tokenEnv → host env var holding the worker's auth token (value never stored in config);
|
||||
# omit for a backend that needs no token (e.g. a local ollama).
|
||||
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
|
||||
# cwd on an MCP spawn, else the daemon's cwd — never $HOME. See
|
||||
# docs/Worker-Startup-and-Trust.md.
|
||||
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
|
||||
workers:
|
||||
gx10: # ccs profile name (NOT a hostname)
|
||||
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
|
||||
model: coder
|
||||
placement: tab
|
||||
workspace: bridged-workers
|
||||
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
|
||||
mcpUrl: http://127.0.0.1:8765/mcp
|
||||
tokenEnv: BRIDGED_WORKER_TOKEN
|
||||
argv: ["ccs", "gx10"]
|
||||
ollama:
|
||||
baseUrl: http://ollama.ltms.dev # local/self-hosted; usually no token
|
||||
placement: tab
|
||||
workspace: bridged-workers
|
||||
tabLabel: "worker: {profile} #{n}"
|
||||
mcpUrl: http://127.0.0.1:8765/mcp
|
||||
argv: ["ccs", "ollama"]
|
||||
defaultWorker: gx10
|
||||
|
||||
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
|
||||
# must carry none. Grounded in ltms-local's real endpoints.
|
||||
# must carry none. Every profile above must have its host listed here.
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
- gx01.gw
|
||||
- ollama.ltms.dev
|
||||
|
||||
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
|
||||
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
|
||||
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
|
||||
# contextCap → force-release a session after this many delegated turns
|
||||
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
|
||||
# lifecycle:
|
||||
# idleTtlSeconds: 300
|
||||
# contextCap: 10
|
||||
# drainTimeoutSeconds: 5
|
||||
|
||||
+53
-4
@@ -17,13 +17,54 @@
|
||||
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
|
||||
<mainClass>dev.ltms.bridged.Bridged</mainClass>
|
||||
|
||||
<jackson.version>2.18.2</jackson.version>
|
||||
<javalin.version>6.3.0</javalin.version>
|
||||
<jackson.version>2.19.0</jackson.version>
|
||||
<javalin.version>6.7.0</javalin.version>
|
||||
<jetty.version>11.0.25</jetty.version>
|
||||
<mcp.version>2.0.0</mcp.version>
|
||||
<slf4j.version>2.0.16</slf4j.version>
|
||||
<logback.version>1.5.15</logback.version>
|
||||
<logback.version>1.5.18</logback.version>
|
||||
<junit.version>5.11.4</junit.version>
|
||||
</properties>
|
||||
|
||||
<!--
|
||||
Dependency security (validate with the JetBrains analyzer's Mend.io check on this pom).
|
||||
Deps are pinned to the latest available versions. Residual advisories with NO upstream fix,
|
||||
accepted for this loopback-bound daemon that processes no untrusted config:
|
||||
- jetty-http 11.0.25 (via Javalin): CVE-2026-2332, CVE-2025-11143 — Jetty 11 is EOL;
|
||||
fixed only in Jetty 12, which needs a Javalin major (6.x rides Jetty 11).
|
||||
- logback-core 1.5.18: CVE-2025-11226, CVE-2026-1225 — both require a MALICIOUS
|
||||
logback.xml (attacker with config write already has code execution); ours is trusted.
|
||||
- jackson-core 2.19.0: WS-2026-0003 — "insufficient information", no fixed version published.
|
||||
- tools.jackson.core (Jackson 3) 3.0.3 via the MCP SDK: CVE-2026-29062 (nesting-depth
|
||||
resource exhaustion). The SDK 2.0.0 is pinned to Jackson 3.0.3 + jackson-annotations
|
||||
3.0-rc5; bumping Jackson 3 to the patched 3.2.x breaks the SDK (annotation mismatch).
|
||||
Only the loopback /mcp endpoint parses this JSON, from trusted local Claude clients.
|
||||
The 11.0.23 -> 11.0.25 bump did clear jetty CVE-2024-8184 (5.9) and CVE-2024-6763.
|
||||
-->
|
||||
|
||||
<!-- Force the latest patched Jetty 11.x across all Javalin-pulled Jetty modules (no version
|
||||
skew). Javalin 6.x rides Jetty 11; a move to Jetty 12 needs a Javalin major. -->
|
||||
<dependencyManagement>
|
||||
<dependencies>
|
||||
<dependency>
|
||||
<groupId>org.eclipse.jetty</groupId>
|
||||
<artifactId>jetty-bom</artifactId>
|
||||
<version>${jetty.version}</version>
|
||||
<type>pom</type>
|
||||
<scope>import</scope>
|
||||
</dependency>
|
||||
<!-- The MCP SDK (Jackson 3) needs jackson-annotations with JsonFormat.Shape.POJO
|
||||
(the 3.0 line); it shares the com.fasterxml.jackson.annotation package with our
|
||||
Jackson 2.19 databind, so both must resolve to the same jar. 3.0 is built to work
|
||||
with Jackson 2.19 databind too — pin it to reconcile the two. -->
|
||||
<dependency>
|
||||
<groupId>com.fasterxml.jackson.core</groupId>
|
||||
<artifactId>jackson-annotations</artifactId>
|
||||
<version>3.0-rc5</version>
|
||||
</dependency>
|
||||
</dependencies>
|
||||
</dependencyManagement>
|
||||
|
||||
<dependencies>
|
||||
<!-- JSON + YAML (config, herdr wire format, REST bodies) -->
|
||||
<dependency>
|
||||
@@ -44,6 +85,14 @@
|
||||
<version>${javalin.version}</version>
|
||||
</dependency>
|
||||
|
||||
<!-- MCP server: the SERVER face. Streamable-HTTP servlet mounted on Javalin's Jetty at
|
||||
/mcp, exposing bridge_send/bridge_reply/bridge_status as thin adapters over REST. -->
|
||||
<dependency>
|
||||
<groupId>io.modelcontextprotocol.sdk</groupId>
|
||||
<artifactId>mcp</artifactId>
|
||||
<version>${mcp.version}</version>
|
||||
</dependency>
|
||||
|
||||
<!-- Logging -->
|
||||
<dependency>
|
||||
<groupId>org.slf4j</groupId>
|
||||
@@ -116,7 +165,7 @@
|
||||
</profile>
|
||||
<profile>
|
||||
<id>contract</id>
|
||||
<properties><excludedGroups></excludedGroups></properties>
|
||||
<properties><excludedGroups/></properties>
|
||||
</profile>
|
||||
</profiles>
|
||||
</project>
|
||||
|
||||
@@ -3,12 +3,26 @@ package dev.ltms.bridged;
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.PaneLocator;
|
||||
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.inject.CompletionResolver;
|
||||
import dev.ltms.bridged.inject.Injector;
|
||||
import dev.ltms.bridged.inject.StatusPoller;
|
||||
import dev.ltms.bridged.inject.TurnListener;
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.mcp.BridgeMcp;
|
||||
import dev.ltms.bridged.mcp.ConnectionIdentity;
|
||||
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
|
||||
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.rest.BridgedApp;
|
||||
import dev.ltms.bridged.worker.WorkerService;
|
||||
import dev.ltms.bridged.session.GitWorktrees;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import dev.ltms.bridged.session.SessionReaper;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import io.javalin.Javalin;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
@@ -40,20 +54,89 @@ public final class Bridged {
|
||||
: UnixSocketHerdrClient.defaultSocketPath();
|
||||
|
||||
UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect(socket, new com.fasterxml.jackson.databind.ObjectMapper());
|
||||
Runtime.getRuntime().addShutdownHook(new Thread(herdr::close));
|
||||
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
WorkspaceControl spaces = new WorkspaceControl(herdr);
|
||||
WorkerService workers = new WorkerService(agents, spaces, guard, cfg.worker(), System::getenv);
|
||||
PeerLauncher workers = new ClaudeCodeLauncher(agents, spaces, guard,
|
||||
cfg.workerProfiles(), cfg.defaultProfile(), System::getenv);
|
||||
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died with
|
||||
// the previous process — reap those leaked orphans now, before we start serving.
|
||||
workers.reapOrphanWorkers();
|
||||
|
||||
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
|
||||
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
|
||||
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
|
||||
int contextCap = 0;
|
||||
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
|
||||
&& cfg.lifecycle().contextCap() > 0) {
|
||||
contextCap = cfg.lifecycle().contextCap();
|
||||
}
|
||||
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()), contextCap);
|
||||
|
||||
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
|
||||
final SessionReaper reaper;
|
||||
if (cfg.lifecycle() != null
|
||||
&& cfg.lifecycle().idleTtlSeconds() != null
|
||||
&& cfg.lifecycle().idleTtlSeconds() > 0) {
|
||||
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
|
||||
reaper.start();
|
||||
} else {
|
||||
reaper = null;
|
||||
}
|
||||
|
||||
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
|
||||
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
|
||||
Injector injector = new Injector(agents);
|
||||
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver completion = new CompletionResolver(agents, rendezvous);
|
||||
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
|
||||
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
|
||||
WorkerPresence presence = sessions.asPresence();
|
||||
TurnListener turnListener = new TurnListener() {
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
completion.onTurnComplete(target);
|
||||
sessions.onTurnComplete(target);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onDelivered(String target) {
|
||||
completion.onDelivered(target);
|
||||
sessions.onDelivered(target);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnFailed(String target) {
|
||||
completion.onTurnFailed(target);
|
||||
sessions.onTurnFailed(target);
|
||||
}
|
||||
};
|
||||
Injector injector = new Injector(agents, turnListener, presence::isPresent, presence::forget);
|
||||
StatusPoller poller = new StatusPoller(agents, injector, INJECT_POLL_MILLIS);
|
||||
poller.start();
|
||||
Runtime.getRuntime().addShutdownHook(new Thread(poller::stop));
|
||||
|
||||
Javalin app = new BridgedApp(herdr, workers).build();
|
||||
MessageService messages = new MessageService(agents, injector, rendezvous);
|
||||
|
||||
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
|
||||
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
|
||||
ConnectionIdentity identity = new ConnectionIdentity(
|
||||
new PaneLocator(herdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
|
||||
// Cast to ClaudeCodeLauncher: BridgeMcp is not yet migrated to PeerLauncher (Stage A scope).
|
||||
BridgeMcp mcp = new BridgeMcp(messages, rendezvous, (ClaudeCodeLauncher) workers, sessions, identity, presence);
|
||||
|
||||
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
|
||||
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
|
||||
// last. This replaces the earlier independent hooks that could race and close herdr early.
|
||||
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
|
||||
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
|
||||
poller.stop();
|
||||
messages.close();
|
||||
mcp.close();
|
||||
if (reaper != null) reaper.stop();
|
||||
herdr.close();
|
||||
}));
|
||||
|
||||
Javalin app = new BridgedApp(herdr, (ClaudeCodeLauncher) workers, sessions, messages, rendezvous, presence, mcp.servlet()).build();
|
||||
app.start(cfg.bind().host(), cfg.bind().port());
|
||||
log.info("bridged listening on {}:{}, herdr socket {}",
|
||||
cfg.bind().host(), cfg.bind().port(), socket);
|
||||
|
||||
@@ -8,7 +8,9 @@ import java.io.IOException;
|
||||
import java.io.UncheckedIOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
/**
|
||||
@@ -16,23 +18,33 @@ import java.util.Set;
|
||||
* {@code bridged.example.yaml}). Unknown keys are ignored so config can grow ahead
|
||||
* of the code.
|
||||
*
|
||||
* @param bind REST/MCP listen host:port
|
||||
* @param herdrSocket path to herdr's Unix socket ({@code null} → client default)
|
||||
* @param worker worker-spawn settings
|
||||
* @param guard subscription-boundary allowlist
|
||||
* @param bind REST/MCP listen host:port
|
||||
* @param herdrSocket path to herdr's Unix socket ({@code null} → client default)
|
||||
* @param worker single worker profile (legacy; superseded by {@code workers})
|
||||
* @param workers named worker profiles, keyed by profile name (multi-backend fleet)
|
||||
* @param defaultWorker which {@code workers} key a no-argument spawn uses ({@code null} → the
|
||||
* single {@code worker}, or the sole/first profile)
|
||||
* @param guard subscription-boundary allowlist
|
||||
* @param worktreeRoot nullable root directory for provisioned worktrees; defaults to a sibling
|
||||
* of the repo root
|
||||
* @param lifecycle session lifecycle limits ({@code null} = all disabled)
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record BridgedConfig(
|
||||
Bind bind,
|
||||
String herdrSocket,
|
||||
Worker worker,
|
||||
Guard guard) {
|
||||
Map<String, Worker> workers,
|
||||
String defaultWorker,
|
||||
Guard guard,
|
||||
String worktreeRoot,
|
||||
Lifecycle lifecycle) {
|
||||
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record Bind(String host, int port) {
|
||||
public Bind {
|
||||
if (host == null || host.isBlank()) host = "127.0.0.1";
|
||||
if (port <= 0) port = 8080;
|
||||
if (port <= 0) port = 8765;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -53,17 +65,63 @@ public record BridgedConfig(
|
||||
* @param tabLabel template for a worker tab's label; {@code {profile}}/{@code {model}}
|
||||
* and {@code {n}} (per-worker number, to keep sibling tabs distinct)
|
||||
* are substituted (default {@code "worker: {profile} #{n}"})
|
||||
* @param mcpUrl bridge MCP URL to provision into the worker's {@code configDir} so it
|
||||
* can call {@code bridge_reply} ({@code null}/blank → no provisioning; the
|
||||
* worker won't reply, only the fallback/timeout resolves the send)
|
||||
* @param cwd fixed working directory for this profile's workers (CB-112 "told otherwise");
|
||||
* {@code null}/blank → inherit the primary's cwd, else the daemon's
|
||||
* @param parityOverlay repo-relative paths copied primary→worktree for config parity; null/empty
|
||||
* defaults to a sensible set of local config files
|
||||
* @param gitTokenEnv name of the host env var holding the git-forge API token; when set, its
|
||||
* value is injected as {@code GITEA_TOKEN} so the worker can open its own PR
|
||||
* at checkpoint (CB-302). {@code null}/blank ⇒ no token is injected
|
||||
* (minimal-grant default — push over SSH stays free, PR-create is opt-in)
|
||||
* @param gitHostEnv name of the host env var holding the forge host (default {@code GITEA_HOST});
|
||||
* injected as {@code GITEA_HOST} <em>only</em> when {@code gitTokenEnv} is set
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record Worker(String profile, String baseUrl, String model,
|
||||
String configDir, String tokenEnv, List<String> argv,
|
||||
String placement, String workspace, String tabLabel) {
|
||||
String placement, String workspace, String tabLabel, String mcpUrl,
|
||||
String cwd,
|
||||
List<String> parityOverlay,
|
||||
String gitTokenEnv, String gitHostEnv) {
|
||||
public Worker {
|
||||
argv = (argv == null || argv.isEmpty()) ? List.of("claude") : List.copyOf(argv);
|
||||
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_WORKER_TOKEN" : tokenEnv;
|
||||
placement = (placement == null || placement.isBlank()) ? "tab" : placement.toLowerCase();
|
||||
workspace = (workspace == null || workspace.isBlank()) ? "bridged-workers" : workspace;
|
||||
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? "worker: {profile} #{n}" : tabLabel;
|
||||
parityOverlay = (parityOverlay == null || parityOverlay.isEmpty())
|
||||
? List.of(".mcp.json", ".claude/settings.local.json", ".env", ".envrc")
|
||||
: List.copyOf(parityOverlay);
|
||||
// gitTokenEnv stays null when unset (opt-in). gitHostEnv defaults so operators enabling
|
||||
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
|
||||
gitHostEnv = (gitHostEnv == null || gitHostEnv.isBlank()) ? "GITEA_HOST" : gitHostEnv;
|
||||
}
|
||||
|
||||
/**
|
||||
* Backward-compatible constructor without the CB-302 git-forge fields — the worker is
|
||||
* granted no PR-create token (push over SSH is unaffected). Keeps pre-CB-302 call sites
|
||||
* (and any {@code workers:} YAML that omits the git keys) working unchanged.
|
||||
*/
|
||||
public Worker(String profile, String baseUrl, String model,
|
||||
String configDir, String tokenEnv, List<String> argv,
|
||||
String placement, String workspace, String tabLabel, String mcpUrl,
|
||||
String cwd, List<String> parityOverlay) {
|
||||
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
|
||||
mcpUrl, cwd, parityOverlay, null, null);
|
||||
}
|
||||
|
||||
/** A copy with {@code profile} set — used to default a profile to its {@code workers} key. */
|
||||
public Worker withProfile(String p) {
|
||||
return new Worker(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
|
||||
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv);
|
||||
}
|
||||
|
||||
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
|
||||
public boolean hasGitToken() {
|
||||
return gitTokenEnv != null && !gitTokenEnv.isBlank();
|
||||
}
|
||||
|
||||
/** True when workers should land in their own tab in the worker space. */
|
||||
@@ -71,6 +129,11 @@ public record BridgedConfig(
|
||||
return "tab".equals(placement);
|
||||
}
|
||||
|
||||
/** True when the bridge MCP should be mounted into a spawned worker (via launch flags). */
|
||||
public boolean hasMcp() {
|
||||
return mcpUrl != null && !mcpUrl.isBlank();
|
||||
}
|
||||
|
||||
/**
|
||||
* Render {@link #tabLabel} for the {@code n}-th worker (substitutes
|
||||
* {@code {profile}}/{@code {model}}/{@code {n}}), so sibling worker tabs are distinct.
|
||||
@@ -83,6 +146,21 @@ public record BridgedConfig(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Session lifecycle limits. All knobs are opt-in: {@code null} or {@code 0} disables the
|
||||
* feature so existing configs keep the previous behaviour.
|
||||
*
|
||||
* @param idleTtlSeconds max seconds a {@code READY}/{@code DONE} session may sit idle
|
||||
* before it is reaped ({@code null} → disabled)
|
||||
* @param contextCap max delegated turns a session serves before force-release
|
||||
* ({@code null} → disabled)
|
||||
* @param drainTimeoutSeconds seconds to wait for {@code BUSY} sessions to finish before
|
||||
* forced teardown on shutdown (default 5 when unset)
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record Lifecycle(Integer idleTtlSeconds, Integer contextCap, Integer drainTimeoutSeconds) {
|
||||
}
|
||||
|
||||
/**
|
||||
* Subscription boundary. Only these hosts may back a worker's
|
||||
* {@code ANTHROPIC_BASE_URL}; the primary must carry none.
|
||||
@@ -100,6 +178,40 @@ public record BridgedConfig(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The effective worker profiles, keyed by profile name. Prefers the {@code workers} map (each
|
||||
* value's {@code profile} defaulted to its key); falls back to the legacy singular {@code worker}
|
||||
* (keyed by its own profile). Empty if neither is configured.
|
||||
*/
|
||||
public Map<String, Worker> workerProfiles() {
|
||||
if (workers != null && !workers.isEmpty()) {
|
||||
Map<String, Worker> out = new LinkedHashMap<>();
|
||||
workers.forEach((name, w) -> out.put(name,
|
||||
(w.profile() == null || w.profile().isBlank()) ? w.withProfile(name) : w));
|
||||
return Map.copyOf(out);
|
||||
}
|
||||
if (worker != null) {
|
||||
String name = (worker.profile() == null || worker.profile().isBlank()) ? "default" : worker.profile();
|
||||
return Map.of(name, worker);
|
||||
}
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
/**
|
||||
* The profile a no-argument spawn uses: {@code defaultWorker} if set, else the legacy single
|
||||
* {@code worker}'s profile, else the sole/first configured profile, else {@code null}.
|
||||
*/
|
||||
public String defaultProfile() {
|
||||
if (defaultWorker != null && !defaultWorker.isBlank()) {
|
||||
return defaultWorker;
|
||||
}
|
||||
if (worker != null && worker.profile() != null && !worker.profile().isBlank()) {
|
||||
return worker.profile();
|
||||
}
|
||||
Map<String, Worker> p = workerProfiles();
|
||||
return p.isEmpty() ? null : p.keySet().iterator().next();
|
||||
}
|
||||
|
||||
private static final ObjectMapper YAML = new ObjectMapper(new YAMLFactory());
|
||||
|
||||
/** Load and validate config from {@code path}. */
|
||||
@@ -116,6 +228,7 @@ public record BridgedConfig(
|
||||
public BridgedConfig withDefaults() {
|
||||
Bind b = bind != null ? bind : new Bind(null, 0);
|
||||
Guard g = guard != null ? guard : new Guard(List.of());
|
||||
return new BridgedConfig(b, herdrSocket, worker, g);
|
||||
Lifecycle l = lifecycle != null ? lifecycle : new Lifecycle(null, null, null);
|
||||
return new BridgedConfig(b, herdrSocket, worker, workers, defaultWorker, g, worktreeRoot, l);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -15,6 +15,10 @@ import com.fasterxml.jackson.databind.JsonNode;
|
||||
* @param agentType detected agent kind, e.g. {@code "claude"}, or {@code null} before herdr
|
||||
* has detected it (the start-time shape)
|
||||
* @param status current lifecycle state
|
||||
* @param name the unique label the agent was started with — for a bridge worker this is
|
||||
* {@code claude-<profile>-<nonce>-<seq>} (CB-117 keys orphan reaping on the
|
||||
* nonce); {@code null} for agents the bridge did not start, e.g. a user's own
|
||||
* Claude session
|
||||
*/
|
||||
public record Agent(
|
||||
String terminalId,
|
||||
@@ -23,7 +27,8 @@ public record Agent(
|
||||
String tabId,
|
||||
String sessionId,
|
||||
String agentType,
|
||||
AgentStatus status) {
|
||||
AgentStatus status,
|
||||
String name) {
|
||||
|
||||
/** Project a herdr {@code agent} node. Tolerates the start-time shape (no session yet). */
|
||||
public static Agent from(JsonNode a) {
|
||||
@@ -42,6 +47,7 @@ public record Agent(
|
||||
a.path("tab_id").asText(null),
|
||||
sessionId,
|
||||
type,
|
||||
AgentStatus.fromWire(a.path("agent_status").asText(null)));
|
||||
AgentStatus.fromWire(a.path("agent_status").asText(null)),
|
||||
a.path("name").asText(null));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -18,6 +18,17 @@ import java.util.Map;
|
||||
*/
|
||||
public final class AgentControl {
|
||||
|
||||
/**
|
||||
* The keystroke that submits a prompt in the Claude Code TUI: a carriage return (Enter).
|
||||
* It must be delivered as its <em>own</em> {@code agent.send} call — herdr delivers a message's
|
||||
* text as a bracketed paste, and a {@code "\r"} appended to that same text is swallowed as
|
||||
* literal newline content, not a submit. Sent as a separate keystroke event it lands outside
|
||||
* the paste and submits. (A bare {@code "\n"} inserts a newline either way.) Verified live
|
||||
* against Claude Code v2.1.210: an injected task stayed unsubmitted with {@code "text\r"} in
|
||||
* one call, and submitted the instant a standalone {@code "\r"} was sent.
|
||||
*/
|
||||
static final String SUBMIT_KEY = "\r";
|
||||
|
||||
private final HerdrClient herdr;
|
||||
|
||||
public AgentControl(HerdrClient herdr) {
|
||||
@@ -36,12 +47,20 @@ public final class AgentControl {
|
||||
return start(name, argv, env, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn an agent into a specific tab. With a non-null {@code tabId} the worker lands
|
||||
* in that tab (the placement policy's dedicated worker tab); with {@code null} herdr
|
||||
* splits the currently-focused tab (legacy pane placement).
|
||||
*/
|
||||
/** Spawn an agent into {@code tabId} at herdr's default cwd. */
|
||||
public Agent start(String name, List<String> argv, Map<String, String> env, String tabId) {
|
||||
return start(name, argv, env, tabId, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn an agent. With a non-null {@code tabId} the worker lands in that tab (the placement
|
||||
* policy's dedicated worker tab); with {@code null} herdr splits the currently-focused tab
|
||||
* (legacy pane placement). A non-blank {@code cwd} sets the worker process's working directory —
|
||||
* {@code agent.start} honours {@code cwd} directly (an agent pane does <em>not</em> inherit the
|
||||
* tab's or workspace's cwd, so this is the only way to root a worker in the primary's directory;
|
||||
* CB-112).
|
||||
*/
|
||||
public Agent start(String name, List<String> argv, Map<String, String> env, String tabId, String cwd) {
|
||||
Map<String, Object> params = new LinkedHashMap<>();
|
||||
params.put("name", name);
|
||||
params.put("argv", argv);
|
||||
@@ -49,13 +68,31 @@ public final class AgentControl {
|
||||
if (tabId != null) {
|
||||
params.put("tab_id", tabId);
|
||||
}
|
||||
if (cwd != null && !cwd.isBlank()) {
|
||||
params.put("cwd", cwd);
|
||||
}
|
||||
JsonNode result = herdr.call("agent.start", params);
|
||||
return Agent.from(result.get("agent"));
|
||||
}
|
||||
|
||||
/** Deliver {@code text} to an agent (its next prompt input). */
|
||||
/**
|
||||
* Deliver {@code text} to an agent as its next prompt <em>and submit it</em> — two keystroke
|
||||
* events: the message (a bracketed paste, so any embedded newlines are preserved verbatim),
|
||||
* then a standalone {@link #SUBMIT_KEY} (Enter) that actually submits it. Without the second
|
||||
* event the text just sits in the worker's input box, never processed (see {@link #SUBMIT_KEY}).
|
||||
*/
|
||||
public void send(String target, String text) {
|
||||
herdr.call("agent.send", Map.of("target", target, "text", text));
|
||||
herdr.call("agent.send", Map.of("target", target, "text", SUBMIT_KEY));
|
||||
}
|
||||
|
||||
/**
|
||||
* Re-send the submit keystroke (Enter) to {@code target}. The Enter that accompanies a delivery
|
||||
* can race the paste — especially right as the worker's TUI becomes interactive — leaving the
|
||||
* text unsubmitted; the injector nudges it with this until the worker actually picks up (CB-113).
|
||||
*/
|
||||
public void submit(String target) {
|
||||
herdr.call("agent.send", Map.of("target", target, "text", SUBMIT_KEY));
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -2,28 +2,38 @@ package dev.ltms.bridged.herdr;
|
||||
|
||||
/**
|
||||
* A herdr agent's lifecycle state, as reported by {@code agent_status}. Drives the
|
||||
* status-gated injector: a worker is safe to inject into only when {@link #IDLE} or
|
||||
* {@link #BLOCKED}, never mid-turn ({@link #WORKING}).
|
||||
* status-gated injector: a worker is safe to inject into only when {@link #IDLE},
|
||||
* {@link #BLOCKED}, or {@link #DONE}, never mid-turn ({@link #WORKING}).
|
||||
*/
|
||||
public enum AgentStatus {
|
||||
IDLE,
|
||||
WORKING,
|
||||
BLOCKED,
|
||||
/**
|
||||
* The worker has finished its turn and is settled at an idle prompt. herdr emits this
|
||||
* (observed live alongside {@code idle}) as a turn-complete marker; earlier code mapped the
|
||||
* unrecognized string to {@link #UNKNOWN}, which both wedged delivery (not {@link #injectable})
|
||||
* and mis-fired the CB-109 stall-failure on a worker that had actually answered. It is a
|
||||
* turn-boundary equivalent to {@link #IDLE}: injectable, and a {@code working → done} edge is a
|
||||
* real completion.
|
||||
*/
|
||||
DONE,
|
||||
UNKNOWN;
|
||||
|
||||
/** Map herdr's wire string ({@code idle|working|blocked|unknown}) to the enum. */
|
||||
/** Map herdr's wire string ({@code idle|working|blocked|done|unknown}) to the enum. */
|
||||
public static AgentStatus fromWire(String s) {
|
||||
if (s == null) return UNKNOWN;
|
||||
return switch (s.toLowerCase()) {
|
||||
case "idle" -> IDLE;
|
||||
case "working" -> WORKING;
|
||||
case "blocked" -> BLOCKED;
|
||||
case "done" -> DONE;
|
||||
default -> UNKNOWN;
|
||||
};
|
||||
}
|
||||
|
||||
/** Whether {@code bridged} may inject a message now without stepping on a live turn. */
|
||||
public boolean injectable() {
|
||||
return this == IDLE || this == BLOCKED;
|
||||
return this == IDLE || this == BLOCKED || this == DONE;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,59 @@
|
||||
package dev.ltms.bridged.herdr;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Resolves which herdr pane a process belongs to — the herdr half of connection-based MCP
|
||||
* identity (CB-105). Given the PID that opened an MCP connection, {@link #terminalForPid} finds
|
||||
* the agent pane whose process tree contains it, so {@code bridged} can tell <em>which worker</em>
|
||||
* is calling without the worker sending anything spoofable.
|
||||
*
|
||||
* <p>herdr owns the PID→pane truth: {@code pane.process_info} reports each pane's {@code shell_pid}
|
||||
* and foreground process PIDs. This scans agent panes; a spawn-time {@code pid→terminal} cache is
|
||||
* the obvious optimization once wired into {@code ClaudeCodeLauncher}.
|
||||
*/
|
||||
public final class PaneLocator {
|
||||
|
||||
private final HerdrClient herdr;
|
||||
|
||||
public PaneLocator(HerdrClient herdr) {
|
||||
this.herdr = herdr;
|
||||
}
|
||||
|
||||
/**
|
||||
* The {@code terminal_id} of the agent pane whose process tree contains {@code pid}, or
|
||||
* {@code null} if no agent pane owns it (e.g. the caller is the primary, or off-host).
|
||||
*/
|
||||
public String terminalForPid(long pid) {
|
||||
if (pid <= 0) {
|
||||
return null;
|
||||
}
|
||||
for (JsonNode pane : herdr.call("pane.list", Map.of()).path("panes")) {
|
||||
String paneId = pane.path("pane_id").asText(null);
|
||||
if (paneId != null && paneOwnsPid(paneId, pid)) {
|
||||
return pane.path("terminal_id").asText(null);
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private boolean paneOwnsPid(String paneId, long pid) {
|
||||
JsonNode info;
|
||||
try {
|
||||
info = herdr.call("pane.process_info", Map.of("pane_id", paneId)).path("process_info");
|
||||
} catch (HerdrException e) {
|
||||
return false; // pane vanished mid-scan — just skip it
|
||||
}
|
||||
if (info.path("shell_pid").asLong(-1) == pid) {
|
||||
return true;
|
||||
}
|
||||
for (JsonNode p : info.path("foreground_processes")) {
|
||||
if (p.path("pid").asLong(-1) == pid) {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
}
|
||||
@@ -65,13 +65,13 @@ public final class WorkspaceControl {
|
||||
}
|
||||
|
||||
/**
|
||||
* A brand-new tab in {@code workspaceId} plus the placeholder shell pane herdr seeds
|
||||
* it with. Start the worker into the tab, then {@code pane.close} the root pane so the
|
||||
* tab holds only the worker.
|
||||
* A brand-new tab in {@code workspaceId} plus the placeholder shell pane herdr seeds it with.
|
||||
* Start the worker into the tab, then {@code pane.close} the root pane so the tab holds only the
|
||||
* worker. (The worker's own cwd is set on {@code agent.start}, not here — an {@code agent.start}
|
||||
* pane does not inherit the tab's cwd; see {@code AgentControl.start}.)
|
||||
*/
|
||||
public Tab.Created createTab(String workspaceId) {
|
||||
JsonNode result = herdr.call("tab.create", Map.of("workspace_id", workspaceId));
|
||||
return Tab.Created.from(result);
|
||||
return Tab.Created.from(herdr.call("tab.create", Map.of("workspace_id", workspaceId)));
|
||||
}
|
||||
|
||||
/** Give a worker's tab a human label in the tab bar. */
|
||||
|
||||
@@ -0,0 +1,256 @@
|
||||
package dev.ltms.bridged.inject;
|
||||
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
|
||||
/**
|
||||
* The CB-106 completion fallback: bridges the {@link Injector}'s turn-completion signal to the
|
||||
* {@link Rendezvous} so a blocking {@code bridge_send} resolves even when the worker finishes its
|
||||
* task without ever calling {@code bridge_reply} — the common case for a real delegated coding task.
|
||||
*
|
||||
* <p>On a confirmed {@code working → idle} boundary it scrapes the worker's recent transcript and
|
||||
* resolves the awaiting send with that tail (a {@link Rendezvous.Kind#COMPLETION} resolution, so the
|
||||
* caller can tell a scrape from a structured reply). It scrapes only when a send is actually waiting
|
||||
* — a fleet worker's own turns, or a send that already timed out, cost no herdr traffic. An explicit
|
||||
* {@code bridge_reply} that raced in first wins; {@link Rendezvous#resolveCompletion} is then a no-op.
|
||||
*
|
||||
* <p>It also handles the CB-109 stall signal ({@link #onTurnFailed}): a worker that ran a turn then
|
||||
* wedged in an {@code unknown} state resolves the send as a failure (with the error screen as
|
||||
* context) rather than leaving it to time out.
|
||||
*
|
||||
* <p>The scrape is cleaned to the last {@code ⏺} assistant block (stripping TUI chrome) and guarded
|
||||
* against misattribution (CB-115): the pane content is baselined on delivery ({@link #onDelivered}),
|
||||
* and a completion whose scrape is unchanged from that baseline — the previous turn's wind-down
|
||||
* sampled as this turn's boundary on a rapid back-to-back send — is suppressed rather than resolving
|
||||
* the send with a stale answer.
|
||||
*
|
||||
* <p><strong>Waiter-specific resolution (CB-116).</strong> On delivery we also capture the exact
|
||||
* {@link Rendezvous} waiter this turn belongs to, and the completion/failure fallbacks resolve
|
||||
* <em>that</em> waiter — never "whatever send is waiting now". A completion fallback runs on a virtual
|
||||
* thread and can land after the worker's {@code bridge_reply} already resolved the turn and the
|
||||
* <em>next</em> send opened its own waiter on the same session; resolving the current waiter would
|
||||
* then deliver turn N's stale scrape as turn N+1's answer. Targeting the captured waiter makes a late
|
||||
* completion a harmless no-op (its waiter is already done) instead of a cross-turn stale reply.
|
||||
*
|
||||
* <p>Wired as the {@link Injector}'s {@link TurnListener}; the handlers hand off to a virtual thread
|
||||
* so the scrape's herdr round-trip never stalls the status poller. The captured waiter is read on the
|
||||
* poller thread (before any next-turn delivery can overwrite it) and passed into the virtual thread.
|
||||
*/
|
||||
public final class CompletionResolver implements TurnListener {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(CompletionResolver.class);
|
||||
|
||||
/**
|
||||
* herdr {@code agent.read} source for the completion scrape. {@code recent} returns the tail of
|
||||
* the transcript (the worker's last output), which is what a delegator wants when the worker
|
||||
* didn't structure a reply.
|
||||
*/
|
||||
static final String SCRAPE_SOURCE = "recent";
|
||||
|
||||
/** Cap the scraped tail so a long transcript can't return an unbounded blob. */
|
||||
static final int MAX_SCRAPE_CHARS = 4000;
|
||||
|
||||
private final AgentControl agents;
|
||||
private final Rendezvous rendezvous;
|
||||
|
||||
/**
|
||||
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
|
||||
* delivering send opened, plus the assistant block present when it was delivered.
|
||||
*
|
||||
* <p>The {@code waiter} is what makes a late fallback safe (CB-116): we resolve it, not "whoever
|
||||
* is waiting now", so a completion that fires after the next send has opened its own waiter is a
|
||||
* no-op rather than a cross-turn stale reply. The {@code baseline} is the CB-115 staleness
|
||||
* reference: a completion scrape equal to it means the worker produced no new output (the previous
|
||||
* turn's wind-down sampled as this boundary), so it is suppressed. Overwritten on each delivery;
|
||||
* cleared when the turn resolves. Package-private so tests can capture and replay a specific turn.
|
||||
*/
|
||||
record InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline) {
|
||||
}
|
||||
|
||||
private final ConcurrentHashMap<String, InFlight> inFlight = new ConcurrentHashMap<>();
|
||||
|
||||
public CompletionResolver(AgentControl agents, Rendezvous rendezvous) {
|
||||
this.agents = agents;
|
||||
this.rendezvous = rendezvous;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onDelivered(String target) {
|
||||
// Capture the exact waiter this turn belongs to (CB-116) and snapshot the pane's pre-turn
|
||||
// content — what it shows *before* the just-delivered turn produces output — as the staleness
|
||||
// reference (CB-115). Done synchronously (like the delivering send itself) so both are in
|
||||
// place before this turn's completion can fire.
|
||||
captureBaseline(target);
|
||||
}
|
||||
|
||||
/** Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link #onDelivered}). */
|
||||
void captureBaseline(String target) {
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
|
||||
if (waiter == null) {
|
||||
inFlight.remove(target); // no send is waiting on this delivery — nothing to resolve later
|
||||
return;
|
||||
}
|
||||
String baseline;
|
||||
try {
|
||||
// Clip to the same cap resolve() applies to the tail (line ~134): the CB-115 misattribution
|
||||
// guard compares baseline.equals(tail), so both sides must be the same capped representation.
|
||||
// An unclipped baseline vs a clipped tail would never match for a >MAX_SCRAPE_CHARS block,
|
||||
// defeating the guard and letting a stale completion resolve the send.
|
||||
baseline = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
|
||||
} catch (RuntimeException e) {
|
||||
baseline = null; // fail open: no baseline ⇒ no suppression
|
||||
log.debug("delivery baseline for {} failed: {}", target, e.getMessage());
|
||||
}
|
||||
inFlight.put(target, new InFlight(waiter, baseline));
|
||||
}
|
||||
|
||||
/** The turn currently baselined for {@code target}, or {@code null} — a test hook for the captureBaseline path. */
|
||||
InFlight inFlight(String target) {
|
||||
return inFlight.get(target);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
// Read the in-flight turn on the poller thread — before any next-turn delivery can overwrite
|
||||
// it — then off-load the scrape (a herdr round-trip we must not block polling on) to a vthread.
|
||||
InFlight turn = inFlight.get(target);
|
||||
Thread.ofVirtual().name("completion-" + target).start(() -> resolve(target, turn));
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnFailed(String target) {
|
||||
InFlight turn = inFlight.get(target);
|
||||
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn));
|
||||
}
|
||||
|
||||
/** Synchronous resolve (the unit-testable core of {@link #onTurnComplete}). */
|
||||
void resolve(String target, InFlight turn) {
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = turn == null ? null : turn.waiter();
|
||||
if (waiter == null || waiter.isDone()) {
|
||||
// Nobody is blocked on THIS turn (it had no send, or its bridge_reply already won). Skip
|
||||
// the scrape; resolving the current waiter here would be the CB-116 cross-turn stale reply.
|
||||
inFlight.remove(target, turn);
|
||||
return;
|
||||
}
|
||||
String tail;
|
||||
boolean scrapeFailed = false;
|
||||
try {
|
||||
tail = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
|
||||
} catch (RuntimeException e) {
|
||||
// The worker finished but we couldn't read its screen — still resolve the send so the
|
||||
// caller unblocks; an empty tail beats hanging until the caller's timeout.
|
||||
log.warn("completion scrape for {} failed; resolving with an empty tail: {}",
|
||||
target, e.getMessage());
|
||||
tail = "";
|
||||
scrapeFailed = true;
|
||||
}
|
||||
// Misattribution guard (CB-115): if the scrape is byte-identical to the pane content at
|
||||
// delivery, this turn produced no new output — the boundary belongs to the previous turn's
|
||||
// wind-down (common on rapid back-to-back sends). Suppress rather than resolve the send with
|
||||
// a stale answer; the real bridge_reply (or a later genuine completion) resolves it instead.
|
||||
// A scrape that failed to read is exempt — an empty tail there is "couldn't see", not "no change".
|
||||
String baseline = turn.baseline();
|
||||
if (!scrapeFailed && baseline != null && baseline.equals(tail)) {
|
||||
log.debug("suppressing misattributed completion for {} (no output change since delivery)",
|
||||
target);
|
||||
return; // keep the in-flight record: a later genuine completion still needs it
|
||||
}
|
||||
if (rendezvous.resolveCompletion(waiter, tail)) {
|
||||
inFlight.remove(target, turn);
|
||||
log.debug("resolved send to {} via turn-completion fallback ({} chars scraped)",
|
||||
target, tail.length());
|
||||
}
|
||||
}
|
||||
|
||||
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
|
||||
void fail(String target, InFlight turn) {
|
||||
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
|
||||
// fall back to the currently-registered waiter (unambiguous — that send never completed, so
|
||||
// no next turn exists to confuse it with).
|
||||
CompletableFuture<Rendezvous.Resolution> waiter =
|
||||
turn != null ? turn.waiter() : rendezvous.currentWaiter(target);
|
||||
if (waiter == null || waiter.isDone()) {
|
||||
inFlight.remove(target, turn); // nobody blocked on this worker — nothing to fail
|
||||
return;
|
||||
}
|
||||
String reason;
|
||||
try {
|
||||
reason = clip(agents.read(target, SCRAPE_SOURCE));
|
||||
} catch (RuntimeException e) {
|
||||
reason = "";
|
||||
}
|
||||
if (reason.isBlank()) {
|
||||
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
|
||||
reason = "worker did not reply; its turn ended in an unrecoverable state "
|
||||
+ "(worker unreachable or stuck)";
|
||||
}
|
||||
if (rendezvous.resolveFailure(waiter, reason)) {
|
||||
inFlight.remove(target, turn);
|
||||
log.debug("failed send to {} via turn-stall fallback", target);
|
||||
}
|
||||
}
|
||||
|
||||
private static String clip(String s) {
|
||||
if (s == null) return "";
|
||||
String trimmed = s.strip();
|
||||
return trimmed.length() <= MAX_SCRAPE_CHARS
|
||||
? trimmed
|
||||
: trimmed.substring(trimmed.length() - MAX_SCRAPE_CHARS);
|
||||
}
|
||||
|
||||
/**
|
||||
* Extract the last assistant message from a raw Claude Code pane scrape (CB-115). Claude Code
|
||||
* prefixes each assistant turn with {@code ⏺}; the delegator wants that answer, not the TUI
|
||||
* chrome around it. Take everything from the final {@code ⏺} onward and stop at the <em>first</em>
|
||||
* hard interface boundary below it — the spinner/status line, input box, {@code ❯} prompt (which
|
||||
* may echo the <em>next</em> turn's text), footer, or tips/warnings. Stopping at the first
|
||||
* boundary (rather than trimming only trailing chrome) is what keeps a following turn's echoed
|
||||
* prompt out of this reply. Blank lines are not boundaries, so a multi-paragraph answer survives;
|
||||
* trailing blanks are trimmed at the end. With no {@code ⏺} marker (an unusual render) the whole
|
||||
* text is scanned the same way, so we never lose the reply.
|
||||
*
|
||||
* <p>Package-private and pure so it is unit-testable without herdr.
|
||||
*/
|
||||
static String lastAssistantBlock(String raw) {
|
||||
if (raw == null || raw.isBlank()) return "";
|
||||
int marker = raw.lastIndexOf('⏺');
|
||||
String block = marker >= 0 ? raw.substring(marker + 1) : raw;
|
||||
StringBuilder out = new StringBuilder();
|
||||
int kept = 0;
|
||||
for (String line : block.split("\n", -1)) {
|
||||
if (isBoundary(line)) break; // first TUI boundary ends the assistant message
|
||||
if (kept++ > 0) out.append('\n');
|
||||
out.append(line);
|
||||
}
|
||||
return out.toString().strip();
|
||||
}
|
||||
|
||||
/**
|
||||
* A hard TUI boundary line that marks the end of an assistant message and the start of interface
|
||||
* chrome (input box, prompt, spinner, footer, tips/warnings). Blank lines are <em>not</em>
|
||||
* boundaries — an answer may contain them — so they are kept and trimmed only if trailing.
|
||||
*/
|
||||
private static boolean isBoundary(String line) {
|
||||
String t = line.strip();
|
||||
if (t.isEmpty()) return false;
|
||||
// A horizontal rule / all box-drawing separators (e.g. "──────").
|
||||
if (t.chars().allMatch(c -> c == '─' || c == '—' || c == '━' || c == '═' || c == '-')) {
|
||||
return true;
|
||||
}
|
||||
String lower = t.toLowerCase();
|
||||
return t.startsWith("╭") || t.startsWith("│") || t.startsWith("╰") || t.startsWith("┌")
|
||||
|| t.startsWith("└") || t.startsWith("❯") || t.startsWith("⏵")
|
||||
|| t.startsWith("⎿") || t.startsWith("⚠")
|
||||
// Status/spinner lines Claude Code renders below a settled or in-flight turn,
|
||||
// e.g. "✻ Baked for 21s", "✶ Forming…".
|
||||
|| t.startsWith("✻") || t.startsWith("✳") || t.startsWith("✽") || t.startsWith("·")
|
||||
|| t.startsWith("●") || t.startsWith("◐") || t.startsWith("✢") || t.startsWith("✶")
|
||||
|| lower.contains("auto mode") || lower.contains("for shortcuts")
|
||||
|| lower.contains("esc to interrupt") || lower.contains("bypass permissions");
|
||||
}
|
||||
}
|
||||
@@ -12,6 +12,8 @@ import java.util.List;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.function.Consumer;
|
||||
import java.util.function.Predicate;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
@@ -32,6 +34,14 @@ import java.util.stream.Collectors;
|
||||
* a transient {@code unknown} — counts as a real pickup, so a detection glitch can't prematurely
|
||||
* release the latch. Perfectly reliable turn boundaries require a herdr {@code events.subscribe}
|
||||
* stream; that is the intended upgrade and would replace only the sampling, not this queue.
|
||||
*
|
||||
* <p><strong>Turn completion (CB-106).</strong> Beyond delivery, the injector reports when a
|
||||
* delegated turn <em>finishes</em>: after a delivery is picked up (a real {@code working} sample),
|
||||
* the next injectable sample is a confirmed {@code working → idle} boundary and fires
|
||||
* {@link TurnListener#onTurnComplete}. Completion is only ever synthesized from a <em>confirmed</em>
|
||||
* turn — the pickup-grace path (a turn too fast to sample) unwedges the queue but does not fire
|
||||
* completion, since without a sampled {@code working} there is no trustworthy "the worker just
|
||||
* finished the task" signal to act on.
|
||||
*/
|
||||
public final class Injector {
|
||||
|
||||
@@ -45,11 +55,68 @@ public final class Injector {
|
||||
*/
|
||||
private static final int PICKUP_GRACE_POLLS = 8;
|
||||
|
||||
/**
|
||||
* How many consecutive {@code unknown} samples while a delegation is outstanding before we
|
||||
* declare it stalled and fire {@link TurnListener#onTurnFailed} (CB-109). A worker wedged in a
|
||||
* state herdr can't classify (e.g. an API-error screen) stays {@code unknown} indefinitely and
|
||||
* would otherwise never resolve; any {@code working}/{@code idle} sample resets the streak, so a
|
||||
* transient detection glitch cannot trip it. At the 250ms poll interval this is ~30s — far longer
|
||||
* than any real detection blip, and still vastly better than the async send's timeout.
|
||||
*/
|
||||
private static final int TURN_STALL_GRACE_POLLS = 120;
|
||||
|
||||
/**
|
||||
* How many consecutive injectable samples a queued-but-undelivered message may wait on the
|
||||
* {@link #ready} gate before we give up and fail it (CB-114). The gate holds a message out of a
|
||||
* worker's boot window (herdr reports {@code idle} while its Claude is still starting), but a
|
||||
* worker whose Claude crashes during boot — or never connects the bridge MCP — stays "idle and
|
||||
* not ready" forever: {@link #ready} never accepts it, the message is never delivered, and the
|
||||
* target would be polled indefinitely with its caller's future never completing. After this
|
||||
* grace the queued messages are failed and the target released. At the 250ms poll interval this
|
||||
* is ~60s — deliberately longer than {@link #TURN_STALL_GRACE_POLLS}, since a first boot (spawn
|
||||
* + model load + MCP connect) legitimately takes longer than an in-turn detection blip.
|
||||
*/
|
||||
private static final int READINESS_GRACE_POLLS = 240;
|
||||
|
||||
private final AgentControl agents;
|
||||
private final TurnListener turnListener;
|
||||
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
|
||||
private final Consumer<String> forget; // CB-114: clear a gone worker's readiness/presence
|
||||
private final ConcurrentHashMap<String, Target> targets = new ConcurrentHashMap<>();
|
||||
|
||||
/** Delivery only; completion signalling is a no-op and every target is treated as available. */
|
||||
public Injector(AgentControl agents) {
|
||||
this(agents, TurnListener.NOOP);
|
||||
}
|
||||
|
||||
/** Delivery plus turn-completion signalling (CB-106); every target is treated as available. */
|
||||
public Injector(AgentControl agents, TurnListener turnListener) {
|
||||
this(agents, turnListener, _ -> true);
|
||||
}
|
||||
|
||||
/**
|
||||
* Delivery, completion signalling (CB-106), and a readiness gate (CB-113): a message is delivered
|
||||
* only when {@code ready} accepts the target — i.e. the worker's Claude has connected the bridge
|
||||
* MCP. This holds the first delivery out of the worker's boot window, where herdr already reports
|
||||
* {@code idle} but the TUI would drop an injected paste.
|
||||
*/
|
||||
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready) {
|
||||
this(agents, turnListener, ready, _ -> {
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Delivery, completion signalling (CB-106), a readiness gate (CB-113), and readiness cleanup
|
||||
* (CB-114): {@code forget} is invoked with a target when its worker is gone — dropped
|
||||
* (pane crash) or timed out on the readiness gate — so its stale presence/readiness is cleared
|
||||
* and does not linger past the worker's life.
|
||||
*/
|
||||
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
|
||||
Consumer<String> forget) {
|
||||
this.agents = agents;
|
||||
this.turnListener = turnListener;
|
||||
this.ready = ready;
|
||||
this.forget = forget;
|
||||
}
|
||||
|
||||
/** A pending message and the future that completes when it has been delivered. */
|
||||
@@ -59,8 +126,12 @@ public final class Injector {
|
||||
/** Per-worker delivery state, guarded by its own monitor (single writer per worker). */
|
||||
private static final class Target {
|
||||
final Deque<Pending> queue = new ArrayDeque<>();
|
||||
boolean awaitingPickup; // sent a message, waiting for the worker to pick it up
|
||||
boolean awaitingPickup; // sent a message, waiting for the worker to pick it up
|
||||
int injectableSincePickup; // consecutive injectable samples while awaitingPickup
|
||||
boolean awaitingCompletion; // a delivered message's turn is not yet known-complete
|
||||
boolean turnObserved; // saw a real `working` sample since that delivery (turn ran)
|
||||
int unknownSinceTurn; // consecutive `unknown` samples while a delegation is outstanding (CB-109)
|
||||
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
|
||||
|
||||
synchronized void add(Pending p) {
|
||||
queue.add(p);
|
||||
@@ -99,63 +170,154 @@ public final class Injector {
|
||||
|
||||
Pending sent = null;
|
||||
RuntimeException sendError = null;
|
||||
boolean turnCompleted = false;
|
||||
boolean turnFailed = false;
|
||||
boolean resubmit = false;
|
||||
List<Pending> notReady = null; // queued messages failed because the worker never became ready
|
||||
synchronized (t) {
|
||||
if (status == AgentStatus.WORKING) {
|
||||
// Definitive pickup: the worker is busy on our last message.
|
||||
// Definitive pickup: the worker is busy on our last message, and (if a delivery is
|
||||
// outstanding) a real turn is now confirmed to be running.
|
||||
t.awaitingPickup = false;
|
||||
t.injectableSincePickup = 0;
|
||||
t.unknownSinceTurn = 0;
|
||||
t.notReadySincePoll = 0;
|
||||
if (t.awaitingCompletion) t.turnObserved = true;
|
||||
} else if (status.injectable()) { // IDLE or BLOCKED
|
||||
if (t.awaitingPickup && ++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
|
||||
// Pickup edge was never sampled (turn faster than the poll, or status lag).
|
||||
// The worker has plainly moved on — release the latch rather than wedge.
|
||||
t.awaitingPickup = false;
|
||||
t.injectableSincePickup = 0;
|
||||
t.unknownSinceTurn = 0;
|
||||
if (t.awaitingPickup) {
|
||||
if (++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
|
||||
// Pickup edge was never sampled (turn faster than the poll, or status lag).
|
||||
// Release the latch rather than wedge — and give up on synthesizing a
|
||||
// completion for this message, since without a confirmed `working` we cannot
|
||||
// trust that a task-processing turn actually ran.
|
||||
t.awaitingPickup = false;
|
||||
t.injectableSincePickup = 0;
|
||||
t.awaitingCompletion = false;
|
||||
t.turnObserved = false;
|
||||
} else {
|
||||
// Delivered but still idle → the worker hasn't picked it up; the submit
|
||||
// keystroke likely raced the paste (esp. right as the TUI became ready).
|
||||
// Re-nudge Enter (CB-113) until the worker starts (WORKING) or the grace ends.
|
||||
resubmit = true;
|
||||
}
|
||||
}
|
||||
if (!t.awaitingPickup) {
|
||||
Pending p = t.queue.peek();
|
||||
if (p != null) {
|
||||
try {
|
||||
agents.send(target, p.text());
|
||||
t.queue.poll();
|
||||
t.awaitingPickup = true;
|
||||
t.injectableSincePickup = 0;
|
||||
sent = p;
|
||||
} catch (RuntimeException e) {
|
||||
// Delivery failed at herdr; drop the poisoned message and surface it
|
||||
// rather than blocking the queue behind it.
|
||||
t.queue.poll();
|
||||
sent = p;
|
||||
sendError = e;
|
||||
// A confirmed turn (a `working` sample was seen) that has now returned to idle is
|
||||
// a trustworthy `working → idle` completion boundary.
|
||||
if (t.awaitingCompletion && t.turnObserved) {
|
||||
t.awaitingCompletion = false;
|
||||
t.turnObserved = false;
|
||||
turnCompleted = true;
|
||||
}
|
||||
// Deliver the next queued message only once the prior turn is fully settled, so a
|
||||
// completion is never confused with the pickup of the following message — and only
|
||||
// once the worker is available (CB-113), so we never paste into its boot window.
|
||||
if (!t.awaitingCompletion) {
|
||||
Pending p = t.queue.peek();
|
||||
if (p != null && ready.test(target)) {
|
||||
t.notReadySincePoll = 0;
|
||||
try {
|
||||
agents.send(target, p.text());
|
||||
t.queue.poll();
|
||||
t.awaitingPickup = true;
|
||||
t.awaitingCompletion = true;
|
||||
t.turnObserved = false;
|
||||
t.injectableSincePickup = 0;
|
||||
sent = p;
|
||||
} catch (RuntimeException e) {
|
||||
// Delivery failed at herdr; drop the poisoned message and surface it
|
||||
// rather than blocking the queue behind it.
|
||||
t.queue.poll();
|
||||
sent = p;
|
||||
sendError = e;
|
||||
}
|
||||
} else if (p != null && ++t.notReadySincePoll >= READINESS_GRACE_POLLS) {
|
||||
// The worker has been idle-but-not-ready for the whole grace: its Claude
|
||||
// never connected the bridge MCP (crashed during boot, or wedged on a
|
||||
// startup prompt). The readiness gate would hold this message forever, so
|
||||
// fail every queued message and release the target (CB-114) instead of
|
||||
// polling it indefinitely with the caller's future never completing.
|
||||
notReady = new ArrayList<>(t.queue);
|
||||
t.queue.clear();
|
||||
t.notReadySincePoll = 0;
|
||||
}
|
||||
}
|
||||
}
|
||||
} else {
|
||||
// UNKNOWN (or any other non-injectable, non-working): not a safe window nor a
|
||||
// reliable pickup signal, so we never deliver or release the pickup latch here. But
|
||||
// an outstanding delegation whose worker has gone unresponsive — stuck in a state
|
||||
// herdr can't classify (CB-109) — will never yield a working→idle boundary. After a
|
||||
// sustained streak, declare it failed so the awaiting send resolves rather than
|
||||
// riding out the async timeout. (This also frees a delivery that wedged before it
|
||||
// was ever picked up, which the injectable-only pickup grace could never release.)
|
||||
if (t.awaitingCompletion && ++t.unknownSinceTurn >= TURN_STALL_GRACE_POLLS) {
|
||||
t.awaitingPickup = false;
|
||||
t.awaitingCompletion = false;
|
||||
t.turnObserved = false;
|
||||
t.unknownSinceTurn = 0;
|
||||
turnFailed = true;
|
||||
}
|
||||
}
|
||||
// UNKNOWN (and any other non-injectable, non-working): do nothing — neither a safe
|
||||
// window nor a reliable pickup signal, so we must not deliver or release the latch.
|
||||
|
||||
// Reclaim the entry once the worker is fully quiescent, so the map cannot grow without
|
||||
// bound across many short-lived workers.
|
||||
if (t.queue.isEmpty() && !t.awaitingPickup) {
|
||||
// Reclaim the entry once the worker is fully quiescent (nothing queued, no pickup or
|
||||
// completion awaited), so the map cannot grow without bound across short-lived workers.
|
||||
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion) {
|
||||
targets.remove(target, t);
|
||||
}
|
||||
}
|
||||
|
||||
// Fire listeners / herdr calls after releasing the monitor so nothing runs on the poller
|
||||
// thread while it holds the target lock.
|
||||
if (resubmit) {
|
||||
try {
|
||||
agents.submit(target); // nudge a raced Enter so the pending paste submits
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("resubmit to {} failed (will retry next poll): {}", target, e.getMessage());
|
||||
}
|
||||
}
|
||||
if (notReady != null) {
|
||||
// Worker never became available: forget its (never-set) readiness, unblock every queued
|
||||
// caller, and route the awaiting send through the same failure path as a stalled turn so
|
||||
// a blocking or async waiter resolves WORKER_FAILED rather than riding out the timeout.
|
||||
forget.accept(target);
|
||||
RuntimeException cause = new IllegalStateException(
|
||||
target + " never became available (no bridge MCP connection within the boot window)");
|
||||
for (Pending p : notReady) {
|
||||
p.delivered().completeExceptionally(cause);
|
||||
}
|
||||
turnListener.onTurnFailed(target);
|
||||
}
|
||||
if (turnCompleted) {
|
||||
turnListener.onTurnComplete(target);
|
||||
}
|
||||
if (turnFailed) {
|
||||
turnListener.onTurnFailed(target);
|
||||
}
|
||||
if (sent != null) {
|
||||
if (sendError != null) {
|
||||
log.warn("inject to {} failed, dropped message: {}", target, sendError.getMessage());
|
||||
sent.delivered().completeExceptionally(sendError);
|
||||
} else {
|
||||
// Baseline the pane's pre-turn content so a misattributed completion (no new output)
|
||||
// can't resolve this send with the previous turn's stale answer (CB-115).
|
||||
turnListener.onDelivered(target);
|
||||
sent.delivered().complete(null);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/** Targets the poller must keep sampling: those with a queued message or an awaited pickup. */
|
||||
/**
|
||||
* Targets the poller must keep sampling: those with a queued message, an awaited pickup, or an
|
||||
* awaited turn completion (so the {@code working → idle} boundary is observed).
|
||||
*/
|
||||
public Set<String> activeTargets() {
|
||||
return targets.entrySet().stream()
|
||||
.filter(e -> {
|
||||
synchronized (e.getValue()) {
|
||||
return !e.getValue().queue.isEmpty() || e.getValue().awaitingPickup;
|
||||
Target t = e.getValue();
|
||||
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion;
|
||||
}
|
||||
})
|
||||
.map(java.util.Map.Entry::getKey)
|
||||
@@ -163,20 +325,31 @@ public final class Injector {
|
||||
}
|
||||
|
||||
/**
|
||||
* Forget a target whose worker is gone, failing every still-queued message so awaiting
|
||||
* callers unblock instead of hanging forever. Futures are completed after the monitor is
|
||||
* released.
|
||||
* Forget a target whose worker is gone, failing every still-queued message so awaiting callers
|
||||
* unblock instead of hanging forever. If a message had already been <em>delivered</em> but its
|
||||
* turn was not yet resolved (CB-110 — the worker vanished mid-turn, e.g. its pane crashed), fire
|
||||
* {@link TurnListener#onTurnFailed} for it: a delivered message is no longer in the queue, so
|
||||
* failing queued waiters alone would leave that send's rendezvous hanging until the async
|
||||
* timeout. Futures and listeners are completed after the monitor is released.
|
||||
*/
|
||||
public void drop(String target, Throwable cause) {
|
||||
Target t = targets.remove(target);
|
||||
if (t == null) return;
|
||||
List<Pending> pending;
|
||||
boolean hadDeliveredTurn;
|
||||
synchronized (t) {
|
||||
pending = new ArrayList<>(t.queue);
|
||||
t.queue.clear();
|
||||
hadDeliveredTurn = t.awaitingCompletion;
|
||||
t.awaitingCompletion = false;
|
||||
t.awaitingPickup = false;
|
||||
}
|
||||
forget.accept(target); // the worker is gone — clear its readiness/presence too (CB-114)
|
||||
for (Pending p : pending) {
|
||||
p.delivered().completeExceptionally(cause);
|
||||
}
|
||||
if (hadDeliveredTurn) {
|
||||
turnListener.onTurnFailed(target);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -23,13 +23,20 @@ public final class StatusPoller {
|
||||
|
||||
private final AgentControl agents;
|
||||
private final Injector injector;
|
||||
private final StatusRefiner refiner;
|
||||
private final long intervalMillis;
|
||||
private volatile boolean running;
|
||||
private Thread thread;
|
||||
|
||||
public StatusPoller(AgentControl agents, Injector injector, long intervalMillis) {
|
||||
this(agents, injector, new StatusRefiner(agents), intervalMillis);
|
||||
}
|
||||
|
||||
public StatusPoller(AgentControl agents, Injector injector, StatusRefiner refiner,
|
||||
long intervalMillis) {
|
||||
this.agents = agents;
|
||||
this.injector = injector;
|
||||
this.refiner = refiner;
|
||||
this.intervalMillis = intervalMillis;
|
||||
}
|
||||
|
||||
@@ -47,7 +54,9 @@ public final class StatusPoller {
|
||||
for (String target : active) {
|
||||
if (!running) return;
|
||||
try {
|
||||
AgentStatus status = agents.status(target);
|
||||
// herdr's agent_status can misreport a settled worker as `unknown`; refine it
|
||||
// against the pane content before it drives delivery/completion (CB-115).
|
||||
AgentStatus status = refiner.refine(target, agents.status(target));
|
||||
injector.onStatus(target, status);
|
||||
} catch (HerdrException e) {
|
||||
// The worker's agent is gone — stop trying and unblock its waiters.
|
||||
|
||||
@@ -0,0 +1,89 @@
|
||||
package dev.ltms.bridged.inject;
|
||||
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.AgentStatus;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
/**
|
||||
* Refines an unreliable {@link AgentStatus#UNKNOWN} into a real state by reading the worker's
|
||||
* terminal content (CB-115).
|
||||
*
|
||||
* <p>Some workers' panes are misclassified by herdr as {@code unknown} even when the worker is
|
||||
* plainly settled at an idle prompt (empty {@code ❯}, "auto mode on" footer, a completed
|
||||
* {@code ⏺} answer above). Left as {@code UNKNOWN} that both <em>wedges delivery</em> — the
|
||||
* status-gated {@link Injector} only injects into an {@link AgentStatus#injectable} worker — and
|
||||
* <em>mis-fires the CB-109 stall failure</em> on a worker that has actually answered. herdr's
|
||||
* {@code agent_status} is a heuristic; the pane content is the ground truth.
|
||||
*
|
||||
* <p>The refinement only ever runs on a raw {@code UNKNOWN} sample (every other status is trusted
|
||||
* as-is), so a healthy worker adds zero extra herdr traffic; a persistently-{@code unknown} worker
|
||||
* costs one extra {@code agent.read} per poll while it has work outstanding. Classification is
|
||||
* deliberately conservative — it upgrades {@code UNKNOWN} to {@link AgentStatus#WORKING} or
|
||||
* {@link AgentStatus#IDLE} only on a clear signal, and leaves a genuinely unclassifiable screen
|
||||
* (e.g. a wedged error state) as {@code UNKNOWN} so the CB-109 stall path can still fail it.
|
||||
*/
|
||||
public final class StatusRefiner {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(StatusRefiner.class);
|
||||
|
||||
/**
|
||||
* herdr {@code agent.read} source used to inspect the pane. {@code detection} is the region
|
||||
* herdr itself uses for status detection (the prompt/footer tail), which is exactly what we
|
||||
* need to tell "idle at prompt" from "mid-turn".
|
||||
*/
|
||||
static final String PROBE_SOURCE = "detection";
|
||||
|
||||
private final AgentControl agents;
|
||||
|
||||
public StatusRefiner(AgentControl agents) {
|
||||
this.agents = agents;
|
||||
}
|
||||
|
||||
/**
|
||||
* Return a trustworthy status for {@code target}. Any non-{@code UNKNOWN} {@code raw} is returned
|
||||
* unchanged; an {@code UNKNOWN} triggers a pane read and content classification. A read failure
|
||||
* leaves it {@code UNKNOWN} (the safe default: no delivery, and the stall path still applies).
|
||||
*/
|
||||
public AgentStatus refine(String target, AgentStatus raw) {
|
||||
if (raw != AgentStatus.UNKNOWN) return raw;
|
||||
String pane;
|
||||
try {
|
||||
pane = agents.read(target, PROBE_SOURCE);
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("status refine read for {} failed; leaving UNKNOWN: {}", target, e.getMessage());
|
||||
return AgentStatus.UNKNOWN;
|
||||
}
|
||||
AgentStatus refined = classify(pane);
|
||||
if (refined != AgentStatus.UNKNOWN) {
|
||||
log.debug("refined {} from UNKNOWN to {} via pane content", target, refined);
|
||||
}
|
||||
return refined;
|
||||
}
|
||||
|
||||
/**
|
||||
* Classify a Claude Code TUI pane tail. Package-private and pure so it is unit-testable without
|
||||
* herdr.
|
||||
*
|
||||
* <ul>
|
||||
* <li>An active-generation marker ({@code esc to interrupt}) ⇒ {@link AgentStatus#WORKING} —
|
||||
* never inject here.</li>
|
||||
* <li>Otherwise, an interactive input prompt with no active-turn marker ({@code ❯}, the
|
||||
* {@code │ >} input box, or the idle {@code auto mode} / shortcuts footer) ⇒
|
||||
* {@link AgentStatus#IDLE} — settled and safe to inject / a completed turn.</li>
|
||||
* <li>Anything else (blank, or an unrecognizable screen) ⇒ {@link AgentStatus#UNKNOWN}.</li>
|
||||
* </ul>
|
||||
*/
|
||||
static AgentStatus classify(String pane) {
|
||||
if (pane == null || pane.isBlank()) return AgentStatus.UNKNOWN;
|
||||
String lower = pane.toLowerCase();
|
||||
// Claude Code shows "(esc to interrupt)" only while a turn is actively generating.
|
||||
if (lower.contains("esc to interrupt")) return AgentStatus.WORKING;
|
||||
// A settled, ready input prompt with no active-turn marker = idle-at-prompt.
|
||||
boolean readyPrompt = pane.contains("❯")
|
||||
|| pane.contains("│ >")
|
||||
|| lower.contains("auto mode on")
|
||||
|| lower.contains("? for shortcuts");
|
||||
return readyPrompt ? AgentStatus.IDLE : AgentStatus.UNKNOWN;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,39 @@
|
||||
package dev.ltms.bridged.inject;
|
||||
|
||||
/**
|
||||
* Notified when a worker's delegated turn is observed to complete — a confirmed
|
||||
* {@code WORKING → IDLE} transition after a delivery. This is the CB-106 completion signal the
|
||||
* {@code CompletionResolver} uses to resolve a blocked send whose worker never called
|
||||
* {@code bridge_reply}. Kept as a seam so the {@link Injector} needs no dependency on the message
|
||||
* layer and stays unit-testable with a capturing fake.
|
||||
*/
|
||||
@FunctionalInterface
|
||||
public interface TurnListener {
|
||||
|
||||
/** A worker's delegated turn finished (worker returned to idle after visibly working). */
|
||||
void onTurnComplete(String target);
|
||||
|
||||
/**
|
||||
* A worker that visibly ran a delegated turn then wedged in a non-idle, non-working state
|
||||
* (CB-109) — e.g. an error screen herdr classifies as {@code unknown} — so no
|
||||
* {@code working → idle} completion boundary will ever arrive. A default no-op keeps this a
|
||||
* functional interface; the completion resolver overrides it to fail the awaiting send.
|
||||
*/
|
||||
default void onTurnFailed(String target) {
|
||||
}
|
||||
|
||||
/**
|
||||
* A message was just delivered into {@code target}'s pane (CB-115). Fired so the completion
|
||||
* resolver can snapshot the pane's pre-turn content: a later {@link #onTurnComplete} whose
|
||||
* scrape is unchanged from this baseline is a <em>misattributed</em> boundary (e.g. the prior
|
||||
* turn's wind-down sampled as this turn's completion on rapid back-to-back sends) and must not
|
||||
* resolve the send with the previous turn's stale answer. A default no-op keeps the interface
|
||||
* functional for callers that don't scrape.
|
||||
*/
|
||||
default void onDelivered(String target) {
|
||||
}
|
||||
|
||||
/** No-op default for callers that only need delivery, not completion signalling. */
|
||||
TurnListener NOOP = _ -> {
|
||||
};
|
||||
}
|
||||
@@ -0,0 +1,38 @@
|
||||
package dev.ltms.bridged.inject;
|
||||
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.Set;
|
||||
|
||||
/**
|
||||
* Tracks which workers are <em>available</em> — their Claude has booted and connected its MCP client
|
||||
* to the bridge (CB-113). This is the reliable readiness signal, unlike herdr's {@code agent_status},
|
||||
* which reports {@code idle} for a worker whose Claude is still booting. Delivering into that boot
|
||||
* window pastes into a not-yet-ready TUI (the text is lost) and wedges the worker's delivery state,
|
||||
* so the {@link Injector} holds the first delivery until the worker is present here.
|
||||
*
|
||||
* <p>Populated from the MCP transport: any MCP request whose connection resolves to a worker terminal
|
||||
* marks that worker present (its {@code initialize} is the first such contact). A worker that never
|
||||
* mounts the bridge MCP is never marked present — its sends stay queued until they time out, which is
|
||||
* correct (it could not have replied anyway).
|
||||
*/
|
||||
public class WorkerPresence {
|
||||
|
||||
private final Set<String> present = ConcurrentHashMap.newKeySet();
|
||||
|
||||
/** Record that {@code terminal}'s worker has connected its MCP client (is available). */
|
||||
public void markPresent(String terminal) {
|
||||
if (terminal != null && !terminal.isBlank()) {
|
||||
present.add(terminal);
|
||||
}
|
||||
}
|
||||
|
||||
/** Whether {@code terminal}'s worker is available (has been seen on the bridge MCP). */
|
||||
public boolean isPresent(String terminal) {
|
||||
return present.contains(terminal);
|
||||
}
|
||||
|
||||
/** Forget a torn-down worker so its terminal id does not linger as "present". */
|
||||
public void forget(String terminal) {
|
||||
present.remove(terminal);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,539 @@
|
||||
package dev.ltms.bridged.mcp;
|
||||
|
||||
import dev.ltms.bridged.guard.GuardException;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.session.WorkerSession;
|
||||
import dev.ltms.bridged.session.WorktreeRequest;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import io.modelcontextprotocol.common.McpTransportContext;
|
||||
import io.modelcontextprotocol.json.McpJsonMapper;
|
||||
import io.modelcontextprotocol.json.jackson3.JacksonMcpJsonMapperSupplier;
|
||||
import io.modelcontextprotocol.server.McpServer;
|
||||
import io.modelcontextprotocol.server.McpSyncServer;
|
||||
import io.modelcontextprotocol.server.McpSyncServerExchange;
|
||||
import io.modelcontextprotocol.server.transport.HttpServletStreamableServerTransportProvider;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import jakarta.servlet.http.HttpServlet;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.function.Function;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
* The MCP SERVER face (CB-105): a Streamable-HTTP MCP server whose tools are <em>thin adapters</em>
|
||||
* over the same {@link MessageService}/{@link Rendezvous} the REST routes use — so the two are
|
||||
* validated by parity, not by re-implementing behaviour. The primary Opus calls {@code bridge_send}
|
||||
* / {@code bridge_status}; the worker calls {@code bridge_reply}.
|
||||
*
|
||||
* <p>Beyond delegation the primary also manages the fleet here (CB-108): {@code bridge_spawn} /
|
||||
* {@code bridge_list} / {@code bridge_stop} adapt {@link ClaudeCodeLauncher} so a worker's whole lifecycle
|
||||
* is driven through MCP, with the subscription boundary still enforced inside {@code ClaudeCodeLauncher}.
|
||||
*
|
||||
* <p>The tool <em>logic</em> lives in package-private static methods returning a
|
||||
* {@link McpSchema.CallToolResult}, so it is unit-testable without standing up the HTTP transport;
|
||||
* the SDK owns the wire protocol. Mount {@link #servlet()} at {@code /mcp} on the daemon's Jetty.
|
||||
*/
|
||||
public final class BridgeMcp {
|
||||
|
||||
private static final long DEFAULT_TIMEOUT_MS = 25_000;
|
||||
private static final long MAX_TIMEOUT_MS = 120_000;
|
||||
// bridge_ask blocks the WORKER's own MCP call, which its client caps near 60s — default under
|
||||
// that so the bridge returns a clean timeout before the client severs the call (CB-205).
|
||||
private static final long ASK_DEFAULT_TIMEOUT_MS = 55_000;
|
||||
private static final long ASK_MAX_TIMEOUT_MS = 115_000;
|
||||
private static final ObjectMapper MAPPER = new ObjectMapper(); // worker-view JSON projections
|
||||
|
||||
/** Transport-context key under which the extractor stashes the resolved caller identity. */
|
||||
static final String CALLER_TERMINAL = "callerTerminal";
|
||||
/** Transport-context key under which the extractor stashes the caller's PID (for cwd inherit). */
|
||||
static final String CALLER_PID = "callerPid";
|
||||
|
||||
private final HttpServletStreamableServerTransportProvider transport;
|
||||
private final McpSyncServer server;
|
||||
|
||||
public BridgeMcp(MessageService messages, Rendezvous rendezvous, ClaudeCodeLauncher workers,
|
||||
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence) {
|
||||
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
|
||||
this.transport = HttpServletStreamableServerTransportProvider.builder()
|
||||
.jsonMapper(json)
|
||||
.mcpEndpoint("/mcp")
|
||||
// Resolve the caller from the connection (peer PID → herdr pane) in one lookup: the
|
||||
// worker terminal for bridge_reply (no spoofable arg), and the PID so bridge_spawn can
|
||||
// inherit the primary's cwd (CB-112). Any contact from a worker marks it available
|
||||
// (CB-113) — its MCP initialize is the reliable "the agent is up" signal.
|
||||
.contextExtractor(req -> {
|
||||
ConnectionIdentity.Caller c = identity.resolve(req.getRemoteAddr(), req.getRemotePort());
|
||||
presence.markPresent(c.terminal()); // no-op for the primary (null terminal)
|
||||
return McpTransportContext.create(Map.of(
|
||||
CALLER_TERMINAL, orEmpty(c.terminal()),
|
||||
CALLER_PID, Long.toString(c.pid())));
|
||||
})
|
||||
.build();
|
||||
this.server = McpServer.sync(transport)
|
||||
.serverInfo("bridge", "0.1.0")
|
||||
.capabilities(McpSchema.ServerCapabilities.builder().tools(true).build())
|
||||
.toolCall(sendTool(), (_, req) -> {
|
||||
Map<String, Object> a = req.arguments();
|
||||
String turnId = str(a, "turnId");
|
||||
if (turnId != null && !turnId.isBlank()) {
|
||||
// Answering a worker's bridge_ask (CB-205): resolve its blocked question and
|
||||
// block for the worker's reply as it resumes the same turn.
|
||||
return answer(messages, turnId, str(a, "content"), timeoutMs(a));
|
||||
}
|
||||
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
|
||||
return Boolean.FALSE.equals(a.get("wait"))
|
||||
? sendAsync(messages, str(a, "sessionId"), str(a, "content"))
|
||||
: send(messages, str(a, "sessionId"), str(a, "content"), timeoutMs(a));
|
||||
})
|
||||
// bridge_reply's identity is the CONNECTION, never an argument.
|
||||
.toolCall(replyTool(), (exchange, req) ->
|
||||
reply(rendezvous, callerTerminal(exchange), str(req.arguments(), "content")))
|
||||
// bridge_ask (CB-205): a worker's mid-turn question — identity from the CONNECTION.
|
||||
.toolCall(askTool(), (exchange, req) ->
|
||||
ask(messages, callerTerminal(exchange), str(req.arguments(), "question"), timeoutMs(req.arguments())))
|
||||
.toolCall(statusTool(), (_, req) ->
|
||||
status(messages, str(req.arguments(), "sessionId")))
|
||||
.toolCall(pollTool(), (_, req) ->
|
||||
poll(messages, str(req.arguments(), "ticket")))
|
||||
// Fleet management (CB-108): spawn/list/stop over ClaudeCodeLauncher.
|
||||
.toolCall(spawnTool(), (exchange, req) -> {
|
||||
Map<String, Object> a = req.arguments();
|
||||
// CB-112: worker inherits the primary's cwd unless the call pins one.
|
||||
// CB-301: carry the caller's identity as the session owner (null for the primary).
|
||||
// CB-301-ext: optional isolated worktree for parallel implementers.
|
||||
String callerCwd = identity.cwdForPid(callerPid(exchange));
|
||||
return spawn(sessions, str(a, "profile"), str(a, "cwd"), callerCwd,
|
||||
callerTerminal(exchange), worktreeRequest(a));
|
||||
})
|
||||
.toolCall(listTool(), (_, _) -> listWorkers(workers, sessions))
|
||||
.toolCall(stopTool(), (_, req) -> stop(sessions, str(req.arguments(), "paneId")))
|
||||
.toolCall(profilesTool(), (_, _) -> profiles(workers))
|
||||
.build();
|
||||
}
|
||||
|
||||
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
|
||||
private static String callerTerminal(McpSyncServerExchange exchange) {
|
||||
Object v = exchange.transportContext().get(CALLER_TERMINAL);
|
||||
String s = v == null ? null : v.toString();
|
||||
return (s == null || s.isBlank()) ? null : s;
|
||||
}
|
||||
|
||||
/** The caller's PID resolved from this call's connection, or {@code -1} if unknown. */
|
||||
private static long callerPid(McpSyncServerExchange exchange) {
|
||||
Object v = exchange.transportContext().get(CALLER_PID);
|
||||
try {
|
||||
return v == null ? -1 : Long.parseLong(v.toString());
|
||||
} catch (NumberFormatException e) {
|
||||
return -1;
|
||||
}
|
||||
}
|
||||
|
||||
private static String orEmpty(String s) {
|
||||
return s == null ? "" : s;
|
||||
}
|
||||
|
||||
/** The Streamable-HTTP servlet to mount at {@code /mcp} on the daemon's Jetty. */
|
||||
public HttpServlet servlet() {
|
||||
return transport;
|
||||
}
|
||||
|
||||
/** Graceful shutdown of the MCP server. */
|
||||
public void close() {
|
||||
server.closeGracefully();
|
||||
}
|
||||
|
||||
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
|
||||
|
||||
/** {@code bridge_send}: delegate {@code content} to a worker session and block for its reply. */
|
||||
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content, Long timeoutMs) {
|
||||
if (isBlank(sessionId) || isBlank(content)) {
|
||||
return error("sessionId and content are required");
|
||||
}
|
||||
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
|
||||
try {
|
||||
return formatReply(messages.send(sessionId, content, timeout), timeout);
|
||||
} catch (HerdrException e) {
|
||||
return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code bridge_send} carrying a {@code turnId}: the primary's answer to a worker's
|
||||
* {@code bridge_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as
|
||||
* it resumes the same turn — surfaced to the primary identically to a normal send.
|
||||
*/
|
||||
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs) {
|
||||
if (isBlank(turnId) || isBlank(content)) {
|
||||
return error("turnId and content are required to answer a worker's question");
|
||||
}
|
||||
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
|
||||
return formatReply(messages.answer(turnId, content, timeout), timeout);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code bridge_ask} (CB-205): a worker pauses its delegated turn to ask the primary, blocking
|
||||
* until the primary answers. The worker is identified by its connection ({@code callerTerminal}),
|
||||
* never an argument — a {@code null} means the caller is not a known worker.
|
||||
*/
|
||||
static McpSchema.CallToolResult ask(MessageService messages, String callerTerminal, String question, Long timeoutMs) {
|
||||
if (callerTerminal == null) {
|
||||
return error("bridge_ask is for workers only — could not identify the calling worker "
|
||||
+ "from the connection");
|
||||
}
|
||||
if (isBlank(question)) {
|
||||
return error("question is required");
|
||||
}
|
||||
long timeout = Math.clamp(timeoutMs == null ? ASK_DEFAULT_TIMEOUT_MS : timeoutMs, 1, ASK_MAX_TIMEOUT_MS);
|
||||
MessageService.AskResult r = messages.ask(callerTerminal, question, timeout);
|
||||
return switch (r.outcome()) {
|
||||
case ANSWERED -> text(r.answer());
|
||||
case NO_WAITER -> error("no primary is awaiting this turn — bridge_ask only works while a "
|
||||
+ "bridge_send delegation is open to answer it");
|
||||
case TIMED_OUT -> text("[no answer within " + timeout + "ms — the primary did not respond; "
|
||||
+ "proceed on your best judgement, then call bridge_reply to end the turn]");
|
||||
};
|
||||
}
|
||||
|
||||
/** Render a {@link MessageService.Reply} as a tool result — shared by {@link #send} and {@link #answer}. */
|
||||
private static McpSchema.CallToolResult formatReply(MessageService.Reply r, long timeout) {
|
||||
return switch (r.outcome()) {
|
||||
case REPLIED -> text(r.text());
|
||||
// The worker's turn finished but it never called bridge_reply — hand back the scraped
|
||||
// transcript tail, flagged so the primary knows it isn't a structured reply.
|
||||
case COMPLETED_UNREPLIED -> text(
|
||||
"[worker finished without a structured bridge_reply — transcript tail follows]\n" + r.text());
|
||||
// The worker ran the turn then wedged (CB-109) — surface the error context.
|
||||
case WORKER_FAILED -> text("[worker failed — turn ended in an unrecoverable state]\n" + r.text());
|
||||
// The worker paused mid-turn to ask (CB-205) — tell the primary how to answer in-turn.
|
||||
case QUESTION -> text("[question] the worker paused to ask before it can finish:\n" + r.text()
|
||||
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + r.turnId()
|
||||
+ "\" and content set to your answer; the worker resumes the same turn.");
|
||||
case STALE_TURN -> error("that question is no longer open — it timed out or was already "
|
||||
+ "answered (turnId stale)");
|
||||
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> text("[no reply within " + timeout + "ms — worker "
|
||||
+ r.outcome().name().toLowerCase().replace("timed_out_", "") + "; retry or poll status]");
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code bridge_send} with {@code wait:false}: delegate {@code content} and return a ticket
|
||||
* immediately (fire-and-poll), so a long task isn't cut off by the caller's MCP call timeout.
|
||||
*/
|
||||
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content) {
|
||||
if (isBlank(sessionId) || isBlank(content)) {
|
||||
return error("sessionId and content are required");
|
||||
}
|
||||
String ticket = messages.sendAsync(sessionId, content);
|
||||
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
|
||||
}
|
||||
|
||||
/** {@code bridge_poll}: check an async delegation by ticket (pending / done+reply / failed). */
|
||||
static McpSchema.CallToolResult poll(MessageService messages, String ticket) {
|
||||
if (isBlank(ticket)) {
|
||||
return error("ticket is required");
|
||||
}
|
||||
MessageService.TaskView v = messages.poll(ticket);
|
||||
if (v == null) {
|
||||
return error("unknown ticket: " + ticket + " (never issued, or expired)");
|
||||
}
|
||||
return switch (v.phase()) {
|
||||
case DONE -> text(v.replySource() != null && v.replySource().equals("transcript")
|
||||
? "[done — worker finished without a structured bridge_reply; transcript tail follows]\n" + v.reply()
|
||||
: v.reply());
|
||||
case PENDING -> text("[pending — " + v.detail() + "]");
|
||||
case FAILED -> text("[failed — " + v.detail() + "]");
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code bridge_reply}: the worker returns its structured answer, resolving the awaiting send.
|
||||
* {@code callerTerminal} is resolved from the connection (never an argument); a {@code null}
|
||||
* means the caller is not a known worker (e.g. the primary called it by mistake).
|
||||
*/
|
||||
static McpSchema.CallToolResult reply(Rendezvous rendezvous, String callerTerminal, String content) {
|
||||
if (callerTerminal == null) {
|
||||
return error("bridge_reply is for workers only — could not identify the calling worker "
|
||||
+ "from the connection");
|
||||
}
|
||||
if (content == null) {
|
||||
return error("content is required");
|
||||
}
|
||||
return rendezvous.resolve(callerTerminal, content)
|
||||
? text("delivered")
|
||||
: error("no send is awaiting a reply for this worker");
|
||||
}
|
||||
|
||||
/** {@code bridge_status}: the live lifecycle status of a worker session. */
|
||||
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
|
||||
if (isBlank(sessionId)) {
|
||||
return error("sessionId is required");
|
||||
}
|
||||
try {
|
||||
return text(messages.status(sessionId).name().toLowerCase());
|
||||
} catch (HerdrException e) {
|
||||
return error("herdr error for session " + sessionId + ": " + e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
// --- fleet management logic (CB-108 / CB-301) --------------------------------------------
|
||||
|
||||
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
|
||||
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
|
||||
return spawn(sessions, profile, null, null, null, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code bridge_spawn}: launch a guard-checked worker for {@code profile} (blank → the default
|
||||
* profile) and return its session id + pane id. The worker's cwd is {@code requestedCwd} if given,
|
||||
* else the profile's config, else {@code callerCwd} (the primary's directory), else the daemon's.
|
||||
* CB-301: the session is registered with {@code ownerTerminal} as its owner.
|
||||
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
|
||||
*/
|
||||
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile,
|
||||
String requestedCwd, String callerCwd,
|
||||
String ownerTerminal, WorktreeRequest worktreeRequest) {
|
||||
try {
|
||||
WorkerSession worker = sessions.acquire(isBlank(profile) ? null : profile,
|
||||
requestedCwd, callerCwd, ownerTerminal, worktreeRequest);
|
||||
return text(json(workerView(worker)));
|
||||
} catch (GuardException e) {
|
||||
return error("subscription boundary: " + e.getMessage());
|
||||
} catch (IllegalArgumentException e) {
|
||||
return error(e.getMessage()); // unknown / no-default profile
|
||||
} catch (HerdrException e) {
|
||||
return error("herdr error spawning worker: " + e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
/** Build a {@link WorktreeRequest} from {@code bridge_spawn}'s optional {@code worktree}/{@code ticket} args. */
|
||||
private static WorktreeRequest worktreeRequest(Map<String, Object> a) {
|
||||
Object w = a.get("worktree");
|
||||
if (w == null || Boolean.FALSE.equals(w)) {
|
||||
return null;
|
||||
}
|
||||
String ticket = str(a, "ticket");
|
||||
if (w instanceof String s) {
|
||||
if (s.isBlank() || "false".equalsIgnoreCase(s)) {
|
||||
return null;
|
||||
}
|
||||
if ("true".equalsIgnoreCase(s)) {
|
||||
if (isBlank(ticket)) {
|
||||
throw new IllegalArgumentException("worktree=true requires a ticket slug");
|
||||
}
|
||||
return new WorktreeRequest(ticket, null);
|
||||
}
|
||||
return new WorktreeRequest(s, null);
|
||||
}
|
||||
if (w instanceof Boolean b && b) {
|
||||
if (isBlank(ticket)) {
|
||||
throw new IllegalArgumentException("worktree=true requires a ticket slug");
|
||||
}
|
||||
return new WorktreeRequest(ticket, null);
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/** {@code bridge_profiles}: the configured worker profiles and the default. */
|
||||
static McpSchema.CallToolResult profiles(ClaudeCodeLauncher workers) {
|
||||
return text(json(Map.of(
|
||||
"profiles", workers.profiles(),
|
||||
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
|
||||
}
|
||||
|
||||
/** {@code bridge_list}: bridge-owned roster merged with live herdr status by paneId. */
|
||||
static McpSchema.CallToolResult listWorkers(ClaudeCodeLauncher workers, SessionManager sessions) {
|
||||
try {
|
||||
Map<String, Agent> live = workers.list().stream()
|
||||
.filter(a -> a.paneId() != null)
|
||||
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
|
||||
List<Map<String, Object>> out = sessions.roster().stream()
|
||||
.map(s -> SessionManager.rosterView(s, live.get(s.paneId())))
|
||||
.toList();
|
||||
return text(json(Map.of("workers", out)));
|
||||
} catch (HerdrException e) {
|
||||
return error("herdr error listing workers: " + e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
/** {@code bridge_stop}: tear a worker down by its pane id. */
|
||||
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
|
||||
if (isBlank(paneId)) {
|
||||
return error("paneId is required");
|
||||
}
|
||||
try {
|
||||
sessions.release(paneId);
|
||||
return text("stopped " + paneId);
|
||||
} catch (HerdrException e) {
|
||||
return error("herdr error stopping " + paneId + ": " + e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
/** CB-301 projection from the authoritative session registry. */
|
||||
private static Map<String, Object> workerView(WorkerSession s) {
|
||||
Map<String, Object> m = new LinkedHashMap<>();
|
||||
m.put("sessionId", s.terminalId());
|
||||
m.put("paneId", s.paneId());
|
||||
m.put("status", s.state().name().toLowerCase());
|
||||
if (s.worktree() != null) {
|
||||
m.put("worktree", s.worktree());
|
||||
}
|
||||
if (s.branch() != null) {
|
||||
m.put("branch", s.branch());
|
||||
}
|
||||
return m;
|
||||
}
|
||||
|
||||
private static String json(Object o) {
|
||||
try {
|
||||
return MAPPER.writeValueAsString(o);
|
||||
} catch (Exception e) {
|
||||
return String.valueOf(o);
|
||||
}
|
||||
}
|
||||
|
||||
// --- tool schemas --------------------------------------------------------------------------
|
||||
|
||||
private static McpSchema.Tool sendTool() {
|
||||
return tool("bridge_send",
|
||||
"Delegate a task to a worker session. By default blocks until the worker replies and "
|
||||
+ "returns its reply (or a 'still working / queued' note on timeout). Pass wait:false "
|
||||
+ "for a long task to return a ticket immediately, then poll it with bridge_poll. To "
|
||||
+ "answer a worker's bridge_ask, pass its turnId (with content) instead of sessionId.",
|
||||
objectSchema(Map.of(
|
||||
"sessionId", stringProp("The worker session id (herdr terminal_id) to delegate to"),
|
||||
"content", stringProp("The task/message to send to the worker (or your answer, with turnId)"),
|
||||
"timeoutMs", Map.of("type", "integer", "description", "Max ms to wait for a reply (blocking mode)"),
|
||||
"wait", Map.of("type", "boolean",
|
||||
"description", "Block for the reply (default true); false returns a ticket to poll"),
|
||||
"turnId", stringProp("When answering a worker's bridge_ask, its question turnId — "
|
||||
+ "routes your answer back into the same turn (omit for a normal delegation)")),
|
||||
List.of("content")));
|
||||
}
|
||||
|
||||
private static McpSchema.Tool askTool() {
|
||||
// No target/session arg — the worker's identity is resolved from the connection.
|
||||
return tool("bridge_ask",
|
||||
"Pause your current delegated turn to ask the primary a question, blocking until it "
|
||||
+ "answers — then resume the same turn with the answer. Use this when only the "
|
||||
+ "primary has a decision or detail you need to continue. You do not address the "
|
||||
+ "primary; identity is resolved from your connection.",
|
||||
objectSchema(Map.of(
|
||||
"question", stringProp("The question to put to the primary"),
|
||||
"timeoutMs", Map.of("type", "integer",
|
||||
"description", "Max ms to wait for the primary's answer")),
|
||||
List.of("question")));
|
||||
}
|
||||
|
||||
private static McpSchema.Tool pollTool() {
|
||||
return tool("bridge_poll",
|
||||
"Check an async delegation (a bridge_send with wait:false) by its ticket: "
|
||||
+ "pending, done (with the worker's reply), or failed.",
|
||||
objectSchema(Map.of(
|
||||
"ticket", stringProp("The ticket returned by bridge_send wait:false")),
|
||||
List.of("ticket")));
|
||||
}
|
||||
|
||||
private static McpSchema.Tool spawnTool() {
|
||||
return tool("bridge_spawn",
|
||||
"Spawn a new off-subscription worker session. Pass a profile (from bridge_profiles) to "
|
||||
+ "pick the backend, or omit it for the default. The worker opens your current "
|
||||
+ "directory by default; pass cwd to pin a different one. Pass worktree:true (with "
|
||||
+ "ticket) or worktree:<ticket-slug> to provision an isolated git worktree. "
|
||||
+ "Returns the worker's sessionId (use with bridge_send) and paneId (use with bridge_stop).",
|
||||
objectSchema(Map.of(
|
||||
"profile", stringProp("Worker profile to spawn (omit for the default profile)"),
|
||||
"cwd", stringProp("Working directory for the worker (omit to inherit yours)"),
|
||||
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
|
||||
"ticket", stringProp("Ticket slug when worktree:true")),
|
||||
List.of()));
|
||||
}
|
||||
|
||||
private static McpSchema.Tool profilesTool() {
|
||||
return tool("bridge_profiles",
|
||||
"List the configured worker profiles (backends) and which one bridge_spawn uses by default.",
|
||||
objectSchema(Map.of(), List.of()));
|
||||
}
|
||||
|
||||
private static McpSchema.Tool listTool() {
|
||||
return tool("bridge_list",
|
||||
"List the worker sessions the bridge tracks — each with its sessionId, paneId, profile, "
|
||||
+ "state, optional worktree/branch/owner, and live herdr status.",
|
||||
objectSchema(Map.of(), List.of()));
|
||||
}
|
||||
|
||||
private static McpSchema.Tool stopTool() {
|
||||
return tool("bridge_stop",
|
||||
"Tear down a worker session by its paneId (from bridge_spawn or bridge_list).",
|
||||
objectSchema(Map.of(
|
||||
"paneId", stringProp("The worker's paneId to stop")),
|
||||
List.of("paneId")));
|
||||
}
|
||||
|
||||
private static McpSchema.Tool replyTool() {
|
||||
// No session/target arg — the worker's identity is resolved from the connection.
|
||||
return tool("bridge_reply",
|
||||
"Return your structured answer for the task you were delegated, "
|
||||
+ "resolving the caller's blocked bridge_send.",
|
||||
objectSchema(Map.of(
|
||||
"content", stringProp("Your reply/answer")),
|
||||
List.of("content")));
|
||||
}
|
||||
|
||||
private static McpSchema.Tool statusTool() {
|
||||
return tool("bridge_status",
|
||||
"Get the live lifecycle status (idle/working/blocked/unknown) of a worker session.",
|
||||
objectSchema(Map.of(
|
||||
"sessionId", stringProp("The worker session id to query")),
|
||||
List.of("sessionId")));
|
||||
}
|
||||
|
||||
// --- small helpers -------------------------------------------------------------------------
|
||||
|
||||
// The SDK 2.0.0 deprecates its own Tool builders without a stable replacement — isolate it here.
|
||||
@SuppressWarnings("deprecation")
|
||||
private static McpSchema.Tool tool(String name, String description, Map<String, Object> inputSchema) {
|
||||
return McpSchema.Tool.builder(name).description(description).inputSchema(inputSchema).build();
|
||||
}
|
||||
|
||||
private static Map<String, Object> objectSchema(Map<String, Object> properties, List<String> required) {
|
||||
return Map.of("type", "object", "properties", properties, "required", required);
|
||||
}
|
||||
|
||||
private static Map<String, Object> stringProp(String description) {
|
||||
return Map.of("type", "string", "description", description);
|
||||
}
|
||||
|
||||
private static McpSchema.CallToolResult text(String s) {
|
||||
return McpSchema.CallToolResult.builder().addTextContent(s == null ? "" : s).build();
|
||||
}
|
||||
|
||||
private static McpSchema.CallToolResult error(String s) {
|
||||
return McpSchema.CallToolResult.builder().addTextContent(s).isError(true).build();
|
||||
}
|
||||
|
||||
private static String str(Map<String, Object> args, String key) {
|
||||
Object v = args.get(key);
|
||||
return v == null ? null : v.toString();
|
||||
}
|
||||
|
||||
private static Long timeoutMs(Map<String, Object> args) {
|
||||
Object v = args.get("timeoutMs");
|
||||
return v instanceof Number n ? n.longValue() : null;
|
||||
}
|
||||
|
||||
private static long clamp(long ms) {
|
||||
return Math.clamp(ms, 1, MAX_TIMEOUT_MS);
|
||||
}
|
||||
|
||||
private static boolean isBlank(String s) {
|
||||
return s == null || s.isBlank();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,66 @@
|
||||
package dev.ltms.bridged.mcp;
|
||||
|
||||
import dev.ltms.bridged.herdr.PaneLocator;
|
||||
|
||||
/**
|
||||
* Resolves <em>who is calling</em> an MCP tool from the connection alone — the anti-spoofing
|
||||
* identity model of the MCP contract. It ties the connection's loopback peer PID (from the OS)
|
||||
* to a herdr agent pane (from herdr), yielding the caller's worker {@code terminal_id}. A caller
|
||||
* that maps to no worker pane — the primary, or an off-host client — resolves to {@code null}.
|
||||
*
|
||||
* <p>Both sources are authoritative and unforgeable: the OS reports the real connecting PID, and
|
||||
* herdr owns the PID→pane mapping. A worker cannot claim to be another worker, nor the primary.
|
||||
* Single-host only (the herd shares the {@code bridged} host); the token path is the split-host
|
||||
* fallback.
|
||||
*/
|
||||
public final class ConnectionIdentity {
|
||||
|
||||
private final PaneLocator panes;
|
||||
private final PeerPidLookup pids;
|
||||
private final ProcessCwdLookup cwds;
|
||||
|
||||
/** Identity only (no cwd resolution — {@link #cwdForPid} returns {@code null}). */
|
||||
public ConnectionIdentity(PaneLocator panes, PeerPidLookup pids) {
|
||||
this(panes, pids, _ -> null);
|
||||
}
|
||||
|
||||
/** Identity plus cwd resolution (CB-112 — inherit the primary's directory on spawn). */
|
||||
public ConnectionIdentity(PaneLocator panes, PeerPidLookup pids, ProcessCwdLookup cwds) {
|
||||
this.panes = panes;
|
||||
this.pids = pids;
|
||||
this.cwds = cwds;
|
||||
}
|
||||
|
||||
/**
|
||||
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
|
||||
* primary / an off-host client) and its {@code pid} (or {@code -1} if not resolvable).
|
||||
*/
|
||||
public record Caller(String terminal, long pid) {
|
||||
}
|
||||
|
||||
/** Resolve the caller's terminal and PID from one peer-PID lookup. */
|
||||
public Caller resolve(String remoteAddr, int remotePort) {
|
||||
if (!isLoopback(remoteAddr)) {
|
||||
return new Caller(null, -1); // only same-host callers can be workers
|
||||
}
|
||||
long pid = pids.pidForLocalPort(remotePort);
|
||||
return new Caller(panes.terminalForPid(pid), pid);
|
||||
}
|
||||
|
||||
/**
|
||||
* The calling worker's {@code terminal_id}, or {@code null} if the caller is not a known
|
||||
* on-host worker (treat as the primary).
|
||||
*/
|
||||
public String callerTerminal(String remoteAddr, int remotePort) {
|
||||
return resolve(remoteAddr, remotePort).terminal();
|
||||
}
|
||||
|
||||
/** The working directory of {@code pid} (the primary's cwd on an MCP spawn), or {@code null}. */
|
||||
public String cwdForPid(long pid) {
|
||||
return pid > 0 ? cwds.cwdForPid(pid) : null;
|
||||
}
|
||||
|
||||
private static boolean isLoopback(String addr) {
|
||||
return "127.0.0.1".equals(addr) || "::1".equals(addr) || "0:0:0:0:0:0:0:1".equals(addr);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,58 @@
|
||||
package dev.ltms.bridged.mcp;
|
||||
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.io.BufferedReader;
|
||||
import java.io.InputStreamReader;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
/**
|
||||
* {@link PeerPidLookup} via {@code lsof} (present on macOS and Linux). For a loopback TCP source
|
||||
* {@code port}, both the client and this daemon appear on that port — so we exclude our own PID
|
||||
* and take the other end, which is the calling process.
|
||||
*/
|
||||
public final class LsofPeerPidLookup implements PeerPidLookup {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(LsofPeerPidLookup.class);
|
||||
|
||||
private final long selfPid = ProcessHandle.current().pid();
|
||||
|
||||
@Override
|
||||
public long pidForLocalPort(int port) {
|
||||
try {
|
||||
Process p = new ProcessBuilder("lsof", "-nP", "-FpP", "-iTCP:" + port)
|
||||
.redirectErrorStream(true).start();
|
||||
long found = -1;
|
||||
try (BufferedReader r = new BufferedReader(
|
||||
new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
|
||||
long current = -1;
|
||||
String line;
|
||||
// -Fp emits records: a 'p<pid>' line, then the ports/files under that pid.
|
||||
while ((line = r.readLine()) != null) {
|
||||
if (line.startsWith("p")) {
|
||||
current = parse(line.substring(1));
|
||||
} else if (current > 0 && current != selfPid) {
|
||||
found = current; // first process on this port that isn't us = the client
|
||||
}
|
||||
}
|
||||
}
|
||||
if (!p.waitFor(2, TimeUnit.SECONDS)) {
|
||||
p.destroyForcibly();
|
||||
}
|
||||
return found;
|
||||
} catch (Exception e) {
|
||||
log.debug("lsof peer-pid lookup for port {} failed: {}", port, e.getMessage());
|
||||
return -1;
|
||||
}
|
||||
}
|
||||
|
||||
private static long parse(String s) {
|
||||
try {
|
||||
return Long.parseLong(s.trim());
|
||||
} catch (NumberFormatException e) {
|
||||
return -1;
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,48 @@
|
||||
package dev.ltms.bridged.mcp;
|
||||
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.io.BufferedReader;
|
||||
import java.io.InputStreamReader;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
/**
|
||||
* {@link ProcessCwdLookup} via {@code lsof} (present on macOS and Linux): {@code lsof -a -p <pid>
|
||||
* -d cwd -Fn} prints the process's cwd on the {@code n…} line. Used to inherit the primary's
|
||||
* working directory for a spawned worker (CB-112).
|
||||
*/
|
||||
public final class LsofProcessCwdLookup implements ProcessCwdLookup {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(LsofProcessCwdLookup.class);
|
||||
|
||||
@Override
|
||||
public String cwdForPid(long pid) {
|
||||
if (pid <= 0) {
|
||||
return null;
|
||||
}
|
||||
try {
|
||||
Process p = new ProcessBuilder("lsof", "-a", "-p", Long.toString(pid), "-d", "cwd", "-Fn")
|
||||
.redirectErrorStream(true).start();
|
||||
String cwd = null;
|
||||
try (BufferedReader r = new BufferedReader(
|
||||
new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
|
||||
String line;
|
||||
while ((line = r.readLine()) != null) {
|
||||
if (line.startsWith("n")) { // 'n<path>' is the file-name field for the cwd fd
|
||||
cwd = line.substring(1);
|
||||
break;
|
||||
}
|
||||
}
|
||||
}
|
||||
if (!p.waitFor(2, TimeUnit.SECONDS)) {
|
||||
p.destroyForcibly();
|
||||
}
|
||||
return (cwd == null || cwd.isBlank()) ? null : cwd;
|
||||
} catch (Exception e) {
|
||||
log.debug("lsof cwd lookup for pid {} failed: {}", pid, e.getMessage());
|
||||
return null;
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,13 @@
|
||||
package dev.ltms.bridged.mcp;
|
||||
|
||||
/**
|
||||
* Resolves the OS PID that owns a loopback TCP source port — the OS half of connection-based MCP
|
||||
* identity. Java exposes no peer PID for a TCP socket, so this shells out. Injectable so
|
||||
* {@link ConnectionIdentity} is testable without a real connection.
|
||||
*/
|
||||
@FunctionalInterface
|
||||
public interface PeerPidLookup {
|
||||
|
||||
/** The PID whose socket has local (source) {@code port} on loopback, or {@code -1} if unknown. */
|
||||
long pidForLocalPort(int port);
|
||||
}
|
||||
@@ -0,0 +1,13 @@
|
||||
package dev.ltms.bridged.mcp;
|
||||
|
||||
/**
|
||||
* Resolves a process's current working directory from its PID — the OS half of CB-112's
|
||||
* "a worker inherits the primary's directory." Injectable so {@link ConnectionIdentity} stays
|
||||
* testable without shelling out.
|
||||
*/
|
||||
@FunctionalInterface
|
||||
public interface ProcessCwdLookup {
|
||||
|
||||
/** The working directory of {@code pid}, or {@code null} if unknown. */
|
||||
String cwdForPid(long pid);
|
||||
}
|
||||
@@ -0,0 +1,385 @@
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.AgentStatus;
|
||||
import dev.ltms.bridged.inject.Injector;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.CompletionException;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.ExecutionException;
|
||||
import java.util.concurrent.ExecutorService;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.TimeoutException;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.concurrent.locks.ReentrantLock;
|
||||
|
||||
/**
|
||||
* The blocking delegation feature (CB-104): deliver {@code content} into a worker and block until
|
||||
* the worker returns a <em>structured reply</em> via {@code bridge_reply} (the {@link Rendezvous}),
|
||||
* then hand that reply back. Delivery is the {@link Injector}'s job (the background poller sends it
|
||||
* when the worker is injectable); this service never drives the injector or scrapes the terminal —
|
||||
* completion is the worker's explicit reply, not a guess about {@code agent_status}.
|
||||
*
|
||||
* <p>Sends are serialized per session so exactly one reply can be outstanding per worker, which is
|
||||
* what lets a reply map unambiguously to its send (no cross-talk between concurrent callers).
|
||||
*
|
||||
* <p>If the worker never replies within the timeout, the caller gets a typed "still working" /
|
||||
* "queued" outcome — the message may still be mid-flight. A finished-but-unreplied turn is caught
|
||||
* by the CB-106 completion fallback (see {@link Rendezvous#resolveCompletion}).
|
||||
*
|
||||
* <p><strong>Async fire-and-poll (CB-107).</strong> A caller's MCP client caps a blocking call at
|
||||
* ~60s, but a real delegated task runs for minutes. {@link #sendAsync} therefore runs the same
|
||||
* blocking {@link #send} on a background virtual thread and hands back a <em>ticket</em> the caller
|
||||
* polls with {@link #poll}. The blocking and async paths share one code path (and the same per-target
|
||||
* serialization), so async inherits the reply + completion resolution behaviour for free.
|
||||
*/
|
||||
public final class MessageService {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(MessageService.class);
|
||||
|
||||
/**
|
||||
* The window a fire-and-poll send waits for resolution — generous, since no caller is blocked on
|
||||
* it; a real delegated task resolves (reply or completion) well within this, and only a genuinely
|
||||
* hung worker rides it out.
|
||||
*/
|
||||
private static final long ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L;
|
||||
|
||||
/** How long a finished (terminal) ticket is retained for polling before it is pruned. */
|
||||
private static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
|
||||
|
||||
/** Outcome of a blocking send. */
|
||||
public enum Outcome {
|
||||
/** The worker called {@code bridge_reply}; {@code text} holds the structured answer. */
|
||||
REPLIED,
|
||||
/**
|
||||
* The worker's delegated turn finished without a {@code bridge_reply} (CB-106 fallback);
|
||||
* {@code text} is the scraped transcript tail rather than a structured answer.
|
||||
*/
|
||||
COMPLETED_UNREPLIED,
|
||||
/**
|
||||
* The worker ran the turn then wedged in an unrecoverable state (CB-109); {@code text} is the
|
||||
* failure context (e.g. the error screen). Terminal, but not a successful completion.
|
||||
*/
|
||||
WORKER_FAILED,
|
||||
/**
|
||||
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
|
||||
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with
|
||||
* {@link #answer(String, String, long)} and the turn resumes.
|
||||
*/
|
||||
QUESTION,
|
||||
/** Timed out after the message was delivered — the worker is still working. */
|
||||
TIMED_OUT_WORKING,
|
||||
/** Timed out before delivery — the message is still queued for the worker. */
|
||||
TIMED_OUT_QUEUED,
|
||||
/** Another send to this session was in flight for the whole window. */
|
||||
BUSY,
|
||||
/**
|
||||
* An answer ({@link #answer(String, String, long)}) referenced a {@code turnId} that is no
|
||||
* longer open — the worker's {@code bridge_ask} already timed out or was answered.
|
||||
*/
|
||||
STALE_TURN
|
||||
}
|
||||
|
||||
/**
|
||||
* @param outcome how the send ended (or paused)
|
||||
* @param text the worker's answer when {@link #completed()} (a structured {@code bridge_reply}
|
||||
* for {@link Outcome#REPLIED}, a scraped transcript tail for
|
||||
* {@link Outcome#COMPLETED_UNREPLIED}), or the question for {@link Outcome#QUESTION},
|
||||
* else {@code null}
|
||||
* @param turnId correlation id for a {@link Outcome#QUESTION} (answered via
|
||||
* {@link #answer(String, String, long)}), else {@code null}
|
||||
*/
|
||||
public record Reply(Outcome outcome, String text, String turnId) {
|
||||
/** A reply with no correlation id (the common terminal outcomes). */
|
||||
public Reply(Outcome outcome, String text) {
|
||||
this(outcome, text, null);
|
||||
}
|
||||
|
||||
/** Whether the worker's turn actually finished with an answer (replied or scraped). */
|
||||
public boolean completed() {
|
||||
return outcome == Outcome.REPLIED || outcome == Outcome.COMPLETED_UNREPLIED;
|
||||
}
|
||||
}
|
||||
|
||||
/** How a worker's {@code bridge_ask} (CB-205) resolved. */
|
||||
public enum AskOutcome {
|
||||
/** The primary answered; {@link AskResult#answer} carries it. */
|
||||
ANSWERED,
|
||||
/** No delegation was open to surface the question to — the worker has no one to ask. */
|
||||
NO_WAITER,
|
||||
/** The primary did not answer within the window. */
|
||||
TIMED_OUT
|
||||
}
|
||||
|
||||
/** The outcome of a worker's {@code bridge_ask}: how it resolved and (if answered) the answer. */
|
||||
public record AskResult(AskOutcome outcome, String answer) {
|
||||
}
|
||||
|
||||
/** Lifecycle phase of an async delegation ticket. */
|
||||
public enum Phase {
|
||||
/** Delegated and in flight — queued for the worker or being worked. */
|
||||
PENDING,
|
||||
/** The worker's turn finished; {@link TaskView#reply} holds the answer. */
|
||||
DONE,
|
||||
/** The delegation could not complete (timed out, worker gone, or busy). */
|
||||
FAILED
|
||||
}
|
||||
|
||||
/**
|
||||
* A poll snapshot of an async delegation.
|
||||
*
|
||||
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, else {@code null}
|
||||
* @param replySource {@code "reply"} (structured {@code bridge_reply}) or {@code "transcript"}
|
||||
* (completion scrape) when {@link Phase#DONE}, else {@code null}
|
||||
* @param detail a human note (live worker status while pending, or the failure reason)
|
||||
*/
|
||||
public record TaskView(String ticket, Phase phase, String reply, String replySource, String detail) {
|
||||
}
|
||||
|
||||
/** An in-flight or finished async delegation, keyed by its ticket. */
|
||||
private record Task(String target, CompletableFuture<Reply> future, long createdNanos) {
|
||||
}
|
||||
|
||||
private final AgentControl agents;
|
||||
private final Injector injector;
|
||||
private final Rendezvous rendezvous;
|
||||
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
|
||||
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
|
||||
private final AtomicLong ticketSeq = new AtomicLong();
|
||||
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
|
||||
Thread.ofVirtual().name("bridge-async-", 0).factory());
|
||||
|
||||
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous) {
|
||||
this.agents = agents;
|
||||
this.injector = injector;
|
||||
this.rendezvous = rendezvous;
|
||||
}
|
||||
|
||||
/** Current lifecycle status of a worker (the {@code GET /sessions/{id}/status} surface). */
|
||||
public AgentStatus status(String target) {
|
||||
return agents.status(target);
|
||||
}
|
||||
|
||||
/**
|
||||
* Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the
|
||||
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
|
||||
*/
|
||||
public Reply send(String target, String content, long timeoutMillis) {
|
||||
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
|
||||
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
|
||||
|
||||
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
|
||||
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
|
||||
}
|
||||
try {
|
||||
CompletableFuture<Void> delivered = injector.enqueue(target, content);
|
||||
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
|
||||
try {
|
||||
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
|
||||
return new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
|
||||
} catch (TimeoutException e) {
|
||||
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
|
||||
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
|
||||
return new Reply(wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null);
|
||||
} catch (ExecutionException e) {
|
||||
Throwable cause = e.getCause();
|
||||
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
|
||||
} catch (InterruptedException e) {
|
||||
Thread.currentThread().interrupt();
|
||||
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
|
||||
} finally {
|
||||
rendezvous.close(target, reply);
|
||||
}
|
||||
} finally {
|
||||
lock.unlock();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* A worker's mid-turn question (CB-205 reverse rendezvous): surface {@code question} to the
|
||||
* primary by resolving its open blocking {@code bridge_send}, then block this (worker) call until
|
||||
* the primary answers via {@link #answer} or {@code timeoutMillis} elapses. Identity is the
|
||||
* worker's own session — it does not address the primary.
|
||||
*
|
||||
* <p>Returns {@link AskOutcome#NO_WAITER} when no delegation is open to surface the question to
|
||||
* (nothing to answer it), {@link AskOutcome#ANSWERED} with the primary's answer, or
|
||||
* {@link AskOutcome#TIMED_OUT} if the primary stayed silent. The worker resumes its turn either
|
||||
* way — an answered ask hands back the answer; an unanswered one leaves it to proceed alone.
|
||||
*/
|
||||
public AskResult ask(String workerSession, String question, long timeoutMillis) {
|
||||
Rendezvous.AskTicket ticket = rendezvous.openAsk(workerSession);
|
||||
// Only the freshly-opening caller surfaces the question; a coalesced duplicate simply blocks on
|
||||
// the shared answer future that the fresh owner is already responsible for.
|
||||
if (ticket.fresh()) {
|
||||
// Register the reverse waiter first, then surface the question — so the answer, which can
|
||||
// arrive the instant the primary reacts, always finds an open waiter to resolve.
|
||||
if (!rendezvous.resolveQuestion(workerSession, question, ticket.turnId())) {
|
||||
rendezvous.closeAsk(ticket.turnId());
|
||||
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
|
||||
}
|
||||
}
|
||||
try {
|
||||
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
|
||||
return new AskResult(AskOutcome.ANSWERED, answer);
|
||||
} catch (TimeoutException e) {
|
||||
log.debug("bridge_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
|
||||
return new AskResult(AskOutcome.TIMED_OUT, null);
|
||||
} catch (ExecutionException e) {
|
||||
Throwable cause = e.getCause();
|
||||
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
|
||||
} catch (InterruptedException e) {
|
||||
Thread.currentThread().interrupt();
|
||||
throw new IllegalStateException("interrupted awaiting the primary's answer for " + workerSession, e);
|
||||
} finally {
|
||||
// Only the fresh owner tears down the shared turn; a duplicate must leave it open.
|
||||
if (ticket.fresh()) {
|
||||
rendezvous.closeAsk(ticket.turnId());
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The primary's answer to a worker's {@code bridge_ask} (CB-205): resolve the worker's blocked
|
||||
* question identified by {@code turnId}, then — like a fresh {@link #send} — block for the worker's
|
||||
* eventual {@code bridge_reply} as it finishes the resumed turn. The worker session is derived from
|
||||
* {@code turnId}, never a caller argument.
|
||||
*
|
||||
* <p>Unlike {@link #send} this does not re-inject through the {@link Injector}: the worker is
|
||||
* mid-turn (already picked up), so the answer flows back through its own open {@code bridge_ask}
|
||||
* call, not a new status-gated delivery. The forward waiter is opened <em>before</em> the worker
|
||||
* is unblocked so a reply that lands the instant it resumes is not lost.
|
||||
*/
|
||||
public Reply answer(String turnId, String content, long timeoutMillis) {
|
||||
String workerSession = rendezvous.askSession(turnId);
|
||||
if (workerSession == null) {
|
||||
return new Reply(Outcome.STALE_TURN, null); // the ask lapsed (timed out or already answered)
|
||||
}
|
||||
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
|
||||
ReentrantLock lock = sessionLocks.computeIfAbsent(workerSession, _ -> new ReentrantLock());
|
||||
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
|
||||
return new Reply(Outcome.BUSY, null);
|
||||
}
|
||||
try {
|
||||
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession);
|
||||
if (!rendezvous.answerAsk(turnId, content)) {
|
||||
rendezvous.close(workerSession, reply);
|
||||
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
|
||||
}
|
||||
try {
|
||||
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
|
||||
return new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
|
||||
} catch (TimeoutException e) {
|
||||
// The worker resumed but hasn't replied yet — no completion fallback arms an answered
|
||||
// turn (it never re-entered the injector), so a silent worker rides out the window.
|
||||
return new Reply(Outcome.TIMED_OUT_WORKING, null);
|
||||
} catch (ExecutionException e) {
|
||||
Throwable cause = e.getCause();
|
||||
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
|
||||
} catch (InterruptedException e) {
|
||||
Thread.currentThread().interrupt();
|
||||
throw new IllegalStateException("interrupted awaiting reply from " + workerSession, e);
|
||||
} finally {
|
||||
rendezvous.close(workerSession, reply);
|
||||
}
|
||||
} finally {
|
||||
lock.unlock();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Fire-and-poll variant of {@link #send}: deliver {@code content} to {@code target} on a
|
||||
* background virtual thread and return immediately with a ticket to {@link #poll}. This is how a
|
||||
* long task is delegated without tripping the caller's MCP client call timeout.
|
||||
*
|
||||
* @return the ticket to poll for the eventual result
|
||||
*/
|
||||
public String sendAsync(String target, String content) {
|
||||
String ticket = "task-" + ticketSeq.incrementAndGet();
|
||||
CompletableFuture<Reply> future =
|
||||
CompletableFuture.supplyAsync(() -> send(target, content, ASYNC_TIMEOUT_MS), asyncExecutor);
|
||||
tasks.put(ticket, new Task(target, future, System.nanoTime()));
|
||||
pruneTerminalTickets();
|
||||
log.debug("async send {} -> {}", ticket, target);
|
||||
return ticket;
|
||||
}
|
||||
|
||||
/**
|
||||
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket;
|
||||
* otherwise a {@link Phase#PENDING} view (with the live worker status as detail), a
|
||||
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason.
|
||||
*/
|
||||
public TaskView poll(String ticket) {
|
||||
Task task = tasks.get(ticket);
|
||||
if (task == null) {
|
||||
return null;
|
||||
}
|
||||
CompletableFuture<Reply> f = task.future();
|
||||
if (!f.isDone()) {
|
||||
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target()));
|
||||
}
|
||||
Reply r;
|
||||
try {
|
||||
r = f.getNow(null);
|
||||
} catch (CompletionException | java.util.concurrent.CancellationException e) {
|
||||
Throwable cause = (e instanceof CompletionException ce && ce.getCause() != null) ? ce.getCause() : e;
|
||||
return new TaskView(ticket, Phase.FAILED, null, null, cause.getMessage());
|
||||
}
|
||||
if (r.completed()) {
|
||||
String source = r.outcome() == Outcome.REPLIED ? "reply" : "transcript";
|
||||
return new TaskView(ticket, Phase.DONE, r.text(), source, null);
|
||||
}
|
||||
// A wedged worker (CB-109) carries the error context as its reason; the timeout/busy
|
||||
// outcomes carry none, so fall back to the outcome name.
|
||||
String detail = r.outcome() == Outcome.WORKER_FAILED && r.text() != null
|
||||
? r.text()
|
||||
: "no reply — " + r.outcome().name().toLowerCase();
|
||||
return new TaskView(ticket, Phase.FAILED, null, null, detail);
|
||||
}
|
||||
|
||||
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
|
||||
private String liveStatus(String target) {
|
||||
try {
|
||||
return agents.status(target).name().toLowerCase();
|
||||
} catch (RuntimeException e) {
|
||||
return "unknown";
|
||||
}
|
||||
}
|
||||
|
||||
/** Drop finished tickets older than the TTL so the registry cannot grow without bound. */
|
||||
private void pruneTerminalTickets() {
|
||||
long cutoff = System.nanoTime() - TICKET_TTL_NANOS;
|
||||
tasks.values().removeIf(t -> t.future().isDone() && t.createdNanos() < cutoff);
|
||||
}
|
||||
|
||||
/** Release the async executor. */
|
||||
public void close() {
|
||||
asyncExecutor.shutdown();
|
||||
}
|
||||
|
||||
/** Map a rendezvous {@link Rendezvous.Kind} onto its send {@link Outcome} (shared by send/answer). */
|
||||
private static Outcome outcomeOf(Rendezvous.Kind kind) {
|
||||
return switch (kind) {
|
||||
case REPLY -> Outcome.REPLIED;
|
||||
case COMPLETION -> Outcome.COMPLETED_UNREPLIED;
|
||||
case FAILED -> Outcome.WORKER_FAILED;
|
||||
case QUESTION -> Outcome.QUESTION;
|
||||
};
|
||||
}
|
||||
|
||||
private static boolean tryLock(ReentrantLock lock, long millis) {
|
||||
try {
|
||||
return lock.tryLock(Math.max(0, millis), TimeUnit.MILLISECONDS);
|
||||
} catch (InterruptedException e) {
|
||||
Thread.currentThread().interrupt();
|
||||
throw new IllegalStateException("interrupted awaiting the session send lock", e);
|
||||
}
|
||||
}
|
||||
|
||||
private static long remainingMillis(long deadlineNanos) {
|
||||
return (deadlineNanos - System.nanoTime()) / 1_000_000L;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,216 @@
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
|
||||
/**
|
||||
* The reply rendezvous: where a blocking {@code bridge_send} awaits how the worker's delegated turn
|
||||
* ends. The sending (primary) request thread {@link #open}s a waiter; it is resolved either by the
|
||||
* worker's explicit {@code bridge_reply} ({@link #resolve}, arriving on a different thread via
|
||||
* {@code POST /sessions/{id}/reply}) or — the CB-106 fallback — by the injector observing the
|
||||
* worker's delegated turn return to idle without a reply ({@link #resolveCompletion}).
|
||||
*
|
||||
* <p>At most one waiter per session — {@link MessageService} serializes sends per session, so a
|
||||
* resolution maps unambiguously to the one outstanding send and cannot be captured by another.
|
||||
*
|
||||
* <p><strong>Waiter identity (CB-116).</strong> The completion/failure fallbacks run asynchronously
|
||||
* and can fire <em>after</em> the turn they belong to has already been resolved by an explicit reply
|
||||
* and a <em>next</em> send has opened its own waiter on the same session. Resolving "whatever waiter
|
||||
* is registered now" would then land turn N's stale scrape on turn N+1's send. So those fallbacks
|
||||
* resolve a <em>specific</em> {@link CompletableFuture} captured when their turn was delivered
|
||||
* ({@link #resolveCompletion(CompletableFuture, String)} /
|
||||
* {@link #resolveFailure(CompletableFuture, String)}): a no-op if that waiter was already resolved,
|
||||
* and it can never touch a later send's waiter.
|
||||
*/
|
||||
public final class Rendezvous {
|
||||
|
||||
/** How a delegated turn ended (or paused). */
|
||||
public enum Kind {
|
||||
/** The worker called {@code bridge_reply} with a structured answer. */
|
||||
REPLY,
|
||||
/** The worker's turn finished without a {@code bridge_reply}; {@code text} is a scrape. */
|
||||
COMPLETION,
|
||||
/** The worker ran the turn then wedged (CB-109); {@code text} is the failure context. */
|
||||
FAILED,
|
||||
/**
|
||||
* The worker paused mid-turn to ask the primary a question (CB-205 reverse rendezvous);
|
||||
* {@code text} is the question and {@code turnId} correlates the primary's answer back to
|
||||
* the worker's blocked {@code bridge_ask}. Not terminal — the turn resumes after the answer.
|
||||
*/
|
||||
QUESTION
|
||||
}
|
||||
|
||||
/**
|
||||
* The resolved outcome of a send: its {@link Kind}, the associated text, and — only for
|
||||
* {@link Kind#QUESTION} — the {@code turnId} the primary answers with (else {@code null}).
|
||||
*/
|
||||
public record Resolution(Kind kind, String text, String turnId) {
|
||||
/** A terminal resolution (reply / completion / failure) with no correlation id. */
|
||||
public Resolution(Kind kind, String text) {
|
||||
this(kind, text, null);
|
||||
}
|
||||
}
|
||||
|
||||
/** A worker's open mid-turn question: the worker session it belongs to and the answer future. */
|
||||
private record AskWaiter(String session, CompletableFuture<String> answer) {
|
||||
}
|
||||
|
||||
/**
|
||||
* Handle to a reverse-rendezvous turn: the {@code turnId}, its answer future, and whether this
|
||||
* call freshly opened it (versus coalescing onto an already-open ask).
|
||||
*/
|
||||
public record AskTicket(String turnId, CompletableFuture<String> answer, boolean fresh) {
|
||||
}
|
||||
|
||||
private final ConcurrentHashMap<String, CompletableFuture<Resolution>> waiters = new ConcurrentHashMap<>();
|
||||
|
||||
/** Reverse rendezvous (CB-205): worker questions awaiting the primary's answer, keyed by {@code turnId}. */
|
||||
private final ConcurrentHashMap<String, AskWaiter> asks = new ConcurrentHashMap<>();
|
||||
private final AtomicLong askSeq = new AtomicLong();
|
||||
/** Per-session index of the currently-open ask, so duplicate bridge_ask calls coalesce onto one turn. */
|
||||
private final ConcurrentHashMap<String, String> openAsksBySession = new ConcurrentHashMap<>();
|
||||
|
||||
/**
|
||||
* Register a waiter for {@code session} — the await side of the public {@code resolve*} methods.
|
||||
* The caller must hold that session's send lock.
|
||||
*/
|
||||
public CompletableFuture<Resolution> open(String session) {
|
||||
CompletableFuture<Resolution> waiter = new CompletableFuture<>();
|
||||
waiters.put(session, waiter);
|
||||
return waiter;
|
||||
}
|
||||
|
||||
/** Remove {@code waiter} for {@code session} (only if it is still the registered one). */
|
||||
void close(String session, CompletableFuture<Resolution> waiter) {
|
||||
waiters.remove(session, waiter);
|
||||
}
|
||||
|
||||
/** Whether a send is currently awaiting a resolution for {@code session}. */
|
||||
public boolean isWaiting(String session) {
|
||||
return waiters.containsKey(session);
|
||||
}
|
||||
|
||||
/**
|
||||
* The waiter currently registered for {@code session}, or {@code null} if none is waiting. The
|
||||
* completion/failure fallbacks capture this at delivery time so they can later resolve that exact
|
||||
* send (see the CB-116 note above) rather than whichever send happens to be waiting when they fire.
|
||||
*/
|
||||
public CompletableFuture<Resolution> currentWaiter(String session) {
|
||||
return waiters.get(session);
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve the send awaiting on {@code session} with the worker's explicit reply {@code content}.
|
||||
*
|
||||
* @return {@code true} if a waiter was resolved; {@code false} if none was waiting (a late or
|
||||
* spurious reply — e.g. the send already timed out)
|
||||
*/
|
||||
public boolean resolve(String session, String content) {
|
||||
return complete(session, new Resolution(Kind.REPLY, content));
|
||||
}
|
||||
|
||||
// --- reverse rendezvous (CB-205 bridge_ask) ------------------------------------------------
|
||||
|
||||
/**
|
||||
* Open a reverse-rendezvous waiter for a worker's mid-turn question. If {@code session} already has
|
||||
* an open ask, coalesce onto it (same {@code turnId}, same answer future). Otherwise atomically mint
|
||||
* a fresh {@code turnId}, register it in both the per-turn and per-session indexes, and hand it back
|
||||
* marked fresh. The caller then {@link #resolveQuestion surfaces the question} to the primary and
|
||||
* blocks on the returned future until the primary {@link #answerAsk answers}.
|
||||
*/
|
||||
public AskTicket openAsk(String session) {
|
||||
while (true) {
|
||||
AskWaiter[] minted = { null };
|
||||
String turnId = openAsksBySession.computeIfAbsent(session, _ -> {
|
||||
String newTurnId = session + "#" + askSeq.incrementAndGet();
|
||||
CompletableFuture<String> answer = new CompletableFuture<>();
|
||||
AskWaiter waiter = new AskWaiter(session, answer);
|
||||
asks.put(newTurnId, waiter);
|
||||
minted[0] = waiter;
|
||||
return newTurnId;
|
||||
});
|
||||
if (minted[0] != null) {
|
||||
return new AskTicket(turnId, minted[0].answer(), true);
|
||||
}
|
||||
AskWaiter existing = asks.get(turnId);
|
||||
if (existing != null) {
|
||||
return new AskTicket(turnId, existing.answer(), false);
|
||||
}
|
||||
// A close raced and removed the waiter after we read the turnId; clear the stale index entry
|
||||
// and retry so a fresh ask is always backed by a registered waiter.
|
||||
openAsksBySession.remove(session, turnId);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Surface a worker's mid-turn {@code question} by resolving the primary's open {@code bridge_send}
|
||||
* with a {@link Kind#QUESTION} carrying {@code turnId}. Same session-keyed semantics as
|
||||
* {@link #resolve}: the one outstanding send for {@code session} unblocks with the question.
|
||||
*
|
||||
* @return {@code true} if a send was awaiting (the question reached the primary); {@code false}
|
||||
* if none was (no delegation is open to answer it)
|
||||
*/
|
||||
public boolean resolveQuestion(String session, String question, String turnId) {
|
||||
return complete(session, new Resolution(Kind.QUESTION, question, turnId));
|
||||
}
|
||||
|
||||
/** The worker session an outstanding ask {@code turnId} belongs to, or {@code null} if unknown/lapsed. */
|
||||
public String askSession(String turnId) {
|
||||
AskWaiter w = asks.get(turnId);
|
||||
return w == null ? null : w.session();
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve a worker's blocked {@code bridge_ask} with the primary's {@code answer}, unblocking it
|
||||
* to resume its turn.
|
||||
*
|
||||
* @return {@code true} if the ask was still open and got the answer; {@code false} if the
|
||||
* {@code turnId} is unknown or the ask already lapsed (timed out / was answered)
|
||||
*/
|
||||
public boolean answerAsk(String turnId, String answer) {
|
||||
AskWaiter w = asks.get(turnId);
|
||||
return w != null && w.answer().complete(answer);
|
||||
}
|
||||
|
||||
/** Drop a reverse-rendezvous turn once its {@code bridge_ask} has resolved (answered or lapsed). */
|
||||
public void closeAsk(String turnId) {
|
||||
AskWaiter w = asks.get(turnId);
|
||||
if (w == null) {
|
||||
return;
|
||||
}
|
||||
// Remove the session index first and only if it still points to this turn, so a concurrent
|
||||
// fresh ask cannot inherit a waiter we are about to drop.
|
||||
openAsksBySession.remove(w.session(), turnId);
|
||||
asks.remove(turnId);
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve a specific captured {@code waiter} as a completion (the delegated turn finished with no
|
||||
* {@code bridge_reply}); {@code text} is the scraped transcript tail. The waiter is the one
|
||||
* captured when this turn was delivered, so a late completion for turn N cannot land on turn N+1's
|
||||
* send (CB-116). A no-op if that waiter was already resolved — a raced {@code bridge_reply} wins.
|
||||
*
|
||||
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
|
||||
*/
|
||||
public boolean resolveCompletion(CompletableFuture<Resolution> waiter, String text) {
|
||||
return waiter != null && waiter.complete(new Resolution(Kind.COMPLETION, text));
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve a specific captured {@code waiter} as a failure — the worker ran the turn but wedged in
|
||||
* an unrecoverable state (CB-109); {@code reason} is the failure context (e.g. the error screen).
|
||||
* Like {@link #resolveCompletion(CompletableFuture, String)} it targets the exact captured send
|
||||
* (CB-116). A no-op if that waiter was already resolved — first resolution wins.
|
||||
*
|
||||
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
|
||||
*/
|
||||
public boolean resolveFailure(CompletableFuture<Resolution> waiter, String reason) {
|
||||
return waiter != null && waiter.complete(new Resolution(Kind.FAILED, reason));
|
||||
}
|
||||
|
||||
private boolean complete(String session, Resolution resolution) {
|
||||
CompletableFuture<Resolution> waiter = waiters.get(session);
|
||||
return waiter != null && waiter.complete(resolution);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,35 @@
|
||||
package dev.ltms.bridged.peer;
|
||||
|
||||
/**
|
||||
* Declared capabilities of a {@link PeerLauncher}. The protocol is the union across all
|
||||
* configured launchers; a verb invoked against a peer that lacks the capability returns a clean
|
||||
* "unsupported for this peer" rather than a crash. Capabilities keep the protocol honest as peers
|
||||
* diversify and prevent the core from assuming "every peer is a Claude in a worktree."
|
||||
*/
|
||||
public enum Capability {
|
||||
|
||||
/**
|
||||
* The peer supports {@code bridge_ask} rendezvous — pausing its delegated turn to ask
|
||||
* the primary a question, then resuming once answered. All Claude Code peers support this.
|
||||
*/
|
||||
MID_TURN_ASK,
|
||||
|
||||
/**
|
||||
* The peer can open its own PR at the end of an implementation turn (CB-302). Opt-in per
|
||||
* profile: granted only when the profile carries a git-forge token ({@code gitTokenEnv}).
|
||||
*/
|
||||
SELF_PR,
|
||||
|
||||
/**
|
||||
* The peer can run inside a provisioned isolated git worktree. All CLI-based peers support
|
||||
* this since their cwd is set at spawn time.
|
||||
*/
|
||||
WORKTREE,
|
||||
|
||||
/**
|
||||
* The spawner can reconcile orphaned peers on boot — workers that outlived a prior daemon
|
||||
* process and whose pane ids died with it (CB-117). Claude Code over herdr supports this
|
||||
* via name-based matching against the herdr agent list.
|
||||
*/
|
||||
ORPHAN_REAP
|
||||
}
|
||||
@@ -0,0 +1,29 @@
|
||||
package dev.ltms.bridged.peer;
|
||||
|
||||
/**
|
||||
* An opaque handle returned by {@link PeerLauncher#spawn(SpawnRequest)}. The core routes on
|
||||
* {@link #id()} (the registry/routing key) and uses {@link #terminalId()} for session tracking;
|
||||
* launcher-private coordinates beyond these are reachable through the concrete implementation.
|
||||
*
|
||||
* <p>A {@link PeerHandle} is returned <em>after</em> the peer process is live — the launcher
|
||||
* has already completed subscription-guarded env/vfs setup, process start, and placement. The
|
||||
* handle is a ticket the core exchanges for the running peer, not a lazy/delayed reference.
|
||||
*/
|
||||
public interface PeerHandle {
|
||||
|
||||
/**
|
||||
* The registry/routing key — an opaque, launcher-assigned identifier. For the herdr-backed
|
||||
* launcher this is the herdr pane id; for other launchers it is whatever their transport
|
||||
* uses. Guaranteed to be non-null and unique among live peers within a single daemon process.
|
||||
*/
|
||||
String id();
|
||||
|
||||
/**
|
||||
* The transport-level session identifier used for message routing and presence tracking.
|
||||
* For the herdr launcher this is the herdr terminal UUID. A non-herdr launcher may return
|
||||
* its own analogous identifier, or {@code null} if the concept does not apply.
|
||||
*/
|
||||
default String terminalId() {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,85 @@
|
||||
package dev.ltms.bridged.peer;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Set;
|
||||
|
||||
/**
|
||||
* SPI for materializing a connected peer — the only way the bridge core creates or tears down
|
||||
* a peer process. Every launcher is a first-party, in-tree adapter selected by (future) profile
|
||||
* config; today's single adapter is the {@code ClaudeCodeLauncher} / Claude Code over herdr.
|
||||
*
|
||||
* <p>The core delegates spawn and teardown to this interface without knowing how the peer is set
|
||||
* up. Environment variables, CLI flags, subscription guards, transport (herdr tab/pane) layout,
|
||||
* and naming conventions are all adapter-private — the core sees only the returned
|
||||
* {@link PeerHandle} whose {@code id()} is the registry/routing key.
|
||||
*
|
||||
* <p>The interface is a superset of what {@code SessionManager} and {@code Bridged.main} call
|
||||
* on the concrete launcher today.
|
||||
*/
|
||||
public interface PeerLauncher {
|
||||
|
||||
/**
|
||||
* The set of {@link Capability capabilities} this launcher declares. A peer whose profile
|
||||
* opts into a git-forge token should include {@link Capability#SELF_PR}; the base set for
|
||||
* the Claude Code herdr adapter is always {@code MID_TURN_ASK, WORKTREE, ORPHAN_REAP}.
|
||||
*/
|
||||
Set<Capability> capabilities();
|
||||
|
||||
/**
|
||||
* {@code profileName}/requestedCwd null/blank → default resolution. Returns after the peer
|
||||
* process is live (env + argv + placement complete). Never returns {@code null}.
|
||||
*
|
||||
* @param req the spawn parameters (profile, requested cwd, caller cwd)
|
||||
* @return a handle whose {@link PeerHandle#id()} is the registry/routing key
|
||||
* @throws IllegalArgumentException if the profile is unknown and no default is configured
|
||||
*/
|
||||
PeerHandle spawn(SpawnRequest req);
|
||||
|
||||
/**
|
||||
* The configured worker profile names — the set of names {@code spawn(profileName)} accepts.
|
||||
*/
|
||||
Set<String> profiles();
|
||||
|
||||
/**
|
||||
* The profile a no-argument {@link #spawn(SpawnRequest)} uses, or {@code null} if none is configured.
|
||||
*/
|
||||
String defaultProfile();
|
||||
|
||||
/**
|
||||
* Resolve the effective working directory for a spawn {@code req} without actually spawning.
|
||||
* Resolution order: requestedCwd → profile cwd → callerCwd → daemon cwd.
|
||||
*
|
||||
* @return the resolved absolute path, never null/blank
|
||||
*/
|
||||
String effectiveCwd(SpawnRequest req);
|
||||
|
||||
/**
|
||||
* The parity-overlay file list for {@code profileName} (default list when unset). Used by
|
||||
* worktree provisioning to copy config files into the isolated checkout before spawning.
|
||||
*/
|
||||
List<String> parityOverlay(String profileName);
|
||||
|
||||
/**
|
||||
* The set of all agents this launcher currently tracks, transport-specific. Each element
|
||||
* exposes at minimum a pane-like {@code id()} matching this launcher's {@link PeerHandle}
|
||||
* scheme, plus transport-level status. Callers merge this set with the session registry to
|
||||
* build a live roster view.
|
||||
*/
|
||||
List<?> list();
|
||||
|
||||
/**
|
||||
* Reap orphaned peers left behind by a prior daemon process. Only peers whose naming scheme
|
||||
* matches this launcher's and whose nonce differs from the current process are eligible.
|
||||
* Best-effort: a failure to list or to stop any one peer is logged and never aborts startup.
|
||||
*
|
||||
* @return the number of orphaned peers reaped
|
||||
*/
|
||||
int reapOrphanWorkers();
|
||||
|
||||
/**
|
||||
* Tear a peer down by its registry/routing key ({@link PeerHandle#id()}). Tolerates an
|
||||
* already-gone peer. Also cleans up launcher-private resources (e.g. empty dedicated tabs)
|
||||
* when safe to do so.
|
||||
*/
|
||||
void stop(String id);
|
||||
}
|
||||
@@ -0,0 +1,13 @@
|
||||
package dev.ltms.bridged.peer;
|
||||
|
||||
/**
|
||||
* Parameters for a {@link PeerLauncher#spawn(SpawnRequest)} call — the peer-neutral
|
||||
* aggregation of what the core knows at delegation time: which profile to use, the caller's
|
||||
* requested working directory, and the caller's own cwd (to inherit when no other cwd is set).
|
||||
*
|
||||
* <p>A null or blank {@code profileName} means "use the launcher's default profile."
|
||||
* A null or blank {@code requestedCwd} means "inherit from config or caller."
|
||||
* A null {@code callerCwd} means "the request came from the daemon itself (not a primary)."
|
||||
*/
|
||||
public record SpawnRequest(String profileName, String requestedCwd, String callerCwd) {
|
||||
}
|
||||
@@ -1,18 +1,29 @@
|
||||
package dev.ltms.bridged.rest;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.bridged.guard.GuardException;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.herdr.HerdrClient;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.worker.WorkerService;
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.session.WorkerSession;
|
||||
import dev.ltms.bridged.session.WorktreeRequest;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import io.javalin.Javalin;
|
||||
import io.javalin.http.Context;
|
||||
import jakarta.servlet.http.HttpServlet;
|
||||
import org.eclipse.jetty.servlet.ServletHolder;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.function.Function;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
* The REST surface — {@code bridged}'s contract and its testability seam. Every
|
||||
@@ -25,22 +36,56 @@ import java.util.Map;
|
||||
*/
|
||||
public final class BridgedApp {
|
||||
|
||||
private final HerdrClient herdr;
|
||||
private final WorkerService workers;
|
||||
/** Default blocking window for a message; kept under typical HTTP idle timeouts. */
|
||||
private static final long DEFAULT_MESSAGE_TIMEOUT_MS = 25_000;
|
||||
private static final long MAX_MESSAGE_TIMEOUT_MS = 120_000;
|
||||
/** Blocking window for a worker's bridge_ask (CB-205); the worker's MCP client caps its own call. */
|
||||
private static final long DEFAULT_ASK_TIMEOUT_MS = 55_000;
|
||||
private static final long MAX_ASK_TIMEOUT_MS = 115_000;
|
||||
|
||||
public BridgedApp(HerdrClient herdr, WorkerService workers) {
|
||||
private final HerdrClient herdr;
|
||||
private final ClaudeCodeLauncher workers;
|
||||
private final SessionManager sessions; // CB-301: authoritative session registry
|
||||
private final MessageService messages;
|
||||
private final Rendezvous rendezvous;
|
||||
private final WorkerPresence presence; // CB-113: which workers are MCP-connected (available)
|
||||
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
|
||||
private final ObjectMapper mapper = new ObjectMapper();
|
||||
|
||||
public BridgedApp(HerdrClient herdr, ClaudeCodeLauncher workers, SessionManager sessions,
|
||||
MessageService messages, Rendezvous rendezvous, WorkerPresence presence,
|
||||
HttpServlet mcpServlet) {
|
||||
this.herdr = herdr;
|
||||
this.workers = workers;
|
||||
this.sessions = sessions;
|
||||
this.messages = messages;
|
||||
this.rendezvous = rendezvous;
|
||||
this.presence = presence;
|
||||
this.mcpServlet = mcpServlet;
|
||||
}
|
||||
|
||||
/** Wire routes onto a fresh, unstarted Javalin instance. Caller starts it. */
|
||||
public Javalin build() {
|
||||
Javalin app = Javalin.create(cfg -> cfg.showJavalinBanner = false);
|
||||
Javalin app = Javalin.create(cfg -> {
|
||||
cfg.showJavalinBanner = false;
|
||||
if (mcpServlet != null) {
|
||||
// The MCP server shares the daemon's port; Jetty routes /mcp to its servlet.
|
||||
cfg.jetty.modifyServletContextHandler(h ->
|
||||
h.addServlet(new ServletHolder(mcpServlet), "/mcp"));
|
||||
}
|
||||
});
|
||||
app.get("/healthz", this::healthz);
|
||||
app.get("/sessions", this::sessions);
|
||||
app.get("/agents", this::agents);
|
||||
app.post("/workers", this::spawnWorker);
|
||||
app.get("/workers", this::listWorkers); // CB-304: registry roster + live herdr status
|
||||
app.get("/profiles", this::profiles); // configured worker profiles
|
||||
app.post("/workers", this::spawnWorker); // optional ?profile= or {"profile":…}
|
||||
app.delete("/workers/{paneId}", this::stopWorker);
|
||||
app.post("/sessions/{id}/message", this::sendMessage); // bridge_send (primary; blocking, wait:false, or answer via turnId)
|
||||
app.post("/sessions/{id}/reply", this::replyMessage); // bridge_reply (worker)
|
||||
app.post("/sessions/{id}/ask", this::askMessage); // bridge_ask (worker → primary, CB-205)
|
||||
app.get("/sessions/{id}/status", this::sessionStatus); // bridge_status
|
||||
app.get("/tasks/{ticket}", this::taskStatus); // poll an async (wait:false) send
|
||||
return app;
|
||||
}
|
||||
|
||||
@@ -81,22 +126,269 @@ public final class BridgedApp {
|
||||
ctx.status(200).json(Map.of("agents", workers.list().stream().map(BridgedApp::view).toList()));
|
||||
}
|
||||
|
||||
/** Spawn a guard-checked worker. 403 if the base_url would breach the subscription boundary. */
|
||||
/** CB-304: bridge-owned roster merged with live herdr status by paneId. */
|
||||
private void listWorkers(Context ctx) {
|
||||
Map<String, Agent> live = workers.list().stream()
|
||||
.filter(a -> a.paneId() != null)
|
||||
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
|
||||
List<Map<String, Object>> out = sessions.roster().stream()
|
||||
.map(s -> SessionManager.rosterView(s, live.get(s.paneId())))
|
||||
.toList();
|
||||
ctx.status(200).json(Map.of("workers", out));
|
||||
}
|
||||
|
||||
/** The configured worker profiles and which one a no-argument spawn uses. */
|
||||
private void profiles(Context ctx) {
|
||||
ctx.status(200).json(Map.of(
|
||||
"profiles", workers.profiles(),
|
||||
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile()));
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn a guard-checked worker. An optional {@code profile} (query param or {@code {"profile":…}}
|
||||
* body) picks which configured profile; omitted → the default. 403 if the base_url would breach
|
||||
* the subscription boundary, 400 for an unknown profile.
|
||||
*/
|
||||
private void spawnWorker(Context ctx) {
|
||||
String profile = ctx.queryParam("profile");
|
||||
String cwd = ctx.queryParam("cwd");
|
||||
String worktree = ctx.queryParam("worktree");
|
||||
String ticket = ctx.queryParam("ticket");
|
||||
if (profile == null || profile.isBlank() || cwd == null || cwd.isBlank()
|
||||
|| worktree == null || worktree.isBlank()) {
|
||||
try {
|
||||
String body = ctx.body();
|
||||
if (!body.isBlank()) {
|
||||
JsonNode b = mapper.readTree(body);
|
||||
if (profile == null || profile.isBlank()) profile = b.path("profile").asText(null);
|
||||
if (cwd == null || cwd.isBlank()) cwd = b.path("cwd").asText(null);
|
||||
if (worktree == null || worktree.isBlank()) worktree = b.path("worktree").asText(null);
|
||||
if (ticket == null || ticket.isBlank()) ticket = b.path("ticket").asText(null);
|
||||
}
|
||||
} catch (Exception ignored) {
|
||||
// A malformed/empty body just means "no overrides" → fall through to defaults.
|
||||
}
|
||||
}
|
||||
WorktreeRequest wt = worktreeRequest(worktree, ticket);
|
||||
try {
|
||||
Agent worker = workers.spawn();
|
||||
// No MCP caller over REST, so callerCwd and ownerTerminal are null.
|
||||
WorkerSession worker = sessions.acquire(blankToNull(profile), blankToNull(cwd), null, null, wt);
|
||||
ctx.status(201).json(view(worker));
|
||||
} catch (GuardException e) {
|
||||
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
|
||||
} catch (IllegalArgumentException e) {
|
||||
ctx.status(400).json(Map.of("error", "unknown_profile", "detail", e.getMessage()));
|
||||
}
|
||||
}
|
||||
|
||||
private static WorktreeRequest worktreeRequest(String worktree, String ticket) {
|
||||
if (worktree == null || worktree.isBlank() || "false".equalsIgnoreCase(worktree)) {
|
||||
return null;
|
||||
}
|
||||
if ("true".equalsIgnoreCase(worktree)) {
|
||||
if (ticket == null || ticket.isBlank()) {
|
||||
throw new IllegalArgumentException("worktree=true requires a ticket slug");
|
||||
}
|
||||
return new WorktreeRequest(ticket, null);
|
||||
}
|
||||
return new WorktreeRequest(worktree, null);
|
||||
}
|
||||
|
||||
private static String blankToNull(String s) {
|
||||
return (s == null || s.isBlank()) ? null : s;
|
||||
}
|
||||
|
||||
/** Tear a worker down by pane id. */
|
||||
private void stopWorker(Context ctx) {
|
||||
workers.stop(ctx.pathParam("paneId"));
|
||||
sessions.release(ctx.pathParam("paneId"));
|
||||
ctx.status(204);
|
||||
}
|
||||
|
||||
/**
|
||||
* The blocking delegation call (CB-104): inject {@code content} into the worker via the
|
||||
* status-gated injector and block until the worker returns a structured {@code bridge_reply}.
|
||||
* Times out with a typed 202 (working / queued / busy) rather than an error — the message may
|
||||
* still land.
|
||||
*/
|
||||
private void sendMessage(Context ctx) {
|
||||
String id = ctx.pathParam("id");
|
||||
String content;
|
||||
String turnId;
|
||||
long timeout;
|
||||
boolean wait;
|
||||
try {
|
||||
JsonNode body = mapper.readTree(ctx.body());
|
||||
content = body.path("content").asText("");
|
||||
turnId = body.path("turnId").asText(null);
|
||||
timeout = body.path("timeoutMs").asLong(DEFAULT_MESSAGE_TIMEOUT_MS);
|
||||
wait = body.path("wait").asBoolean(true); // default: block for the reply (CB-104)
|
||||
} catch (Exception e) {
|
||||
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
|
||||
return;
|
||||
}
|
||||
if (content.isBlank()) {
|
||||
ctx.status(400).json(Map.of("error", "bad_request", "detail", "content is required"));
|
||||
return;
|
||||
}
|
||||
timeout = Math.clamp(timeout, 1, MAX_MESSAGE_TIMEOUT_MS);
|
||||
|
||||
// Answering a worker's bridge_ask (CB-205): always blocks, and derives the worker from turnId.
|
||||
if (turnId != null && !turnId.isBlank()) {
|
||||
writeReply(ctx, id, messages.answer(turnId, content, timeout), timeout);
|
||||
return;
|
||||
}
|
||||
|
||||
if (!wait) {
|
||||
// Fire-and-poll (CB-107): return a ticket immediately; the caller polls GET /tasks/{ticket}.
|
||||
String ticket = messages.sendAsync(id, content);
|
||||
ctx.status(202).json(Map.of("sessionId", id, "ticket", ticket, "status", "accepted"));
|
||||
return;
|
||||
}
|
||||
|
||||
try {
|
||||
writeReply(ctx, id, messages.send(id, content, timeout), timeout);
|
||||
} catch (HerdrException e) {
|
||||
herdrError(ctx, e);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Render a {@link MessageService.Reply} onto the response — shared by a normal send and a
|
||||
* bridge_ask answer. A structured/scraped completion is 200; a worker's mid-turn question a 202
|
||||
* (with its {@code turnId}); a stale answer a 409; every other non-terminal outcome a typed 202.
|
||||
*/
|
||||
private void writeReply(Context ctx, String id, MessageService.Reply reply, long timeout) {
|
||||
switch (reply.outcome()) {
|
||||
case QUESTION -> ctx.status(202).json(Map.of(
|
||||
"sessionId", id, "status", "question",
|
||||
"question", reply.text(), "turnId", reply.turnId()));
|
||||
case STALE_TURN -> ctx.status(409).json(Map.of(
|
||||
"sessionId", id, "error", "stale_turn",
|
||||
"detail", "that question is no longer open (timed out or already answered)"));
|
||||
case REPLIED, COMPLETED_UNREPLIED -> {
|
||||
// replySource distinguishes a structured bridge_reply from the CB-106 completion
|
||||
// fallback (a scrape of the worker's transcript when it finished without replying).
|
||||
String source = reply.outcome() == MessageService.Outcome.REPLIED ? "reply" : "transcript";
|
||||
ctx.status(200).json(Map.of("sessionId", id, "reply", reply.text(), "replySource", source));
|
||||
}
|
||||
default -> ctx.status(202).json(Map.of(
|
||||
"sessionId", id,
|
||||
"status", switch (reply.outcome()) {
|
||||
case TIMED_OUT_WORKING -> "working";
|
||||
case TIMED_OUT_QUEUED -> "queued";
|
||||
case BUSY -> "busy";
|
||||
case WORKER_FAILED -> "failed";
|
||||
default -> "done"; // unreachable (terminal outcomes handled above)
|
||||
},
|
||||
"detail", reply.outcome() == MessageService.Outcome.WORKER_FAILED && reply.text() != null
|
||||
? reply.text()
|
||||
: "no reply within " + timeout + "ms; poll status or retry"));
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* A worker's mid-turn question ({@code bridge_ask}, CB-205) — surfaces to the primary's open
|
||||
* blocking send and blocks until it answers. 200 with the answer, 409 if no delegation is open,
|
||||
* 202 if the primary stayed silent.
|
||||
*/
|
||||
private void askMessage(Context ctx) {
|
||||
String id = ctx.pathParam("id");
|
||||
String question;
|
||||
long timeout;
|
||||
try {
|
||||
JsonNode body = mapper.readTree(ctx.body());
|
||||
question = body.path("question").asText("");
|
||||
timeout = body.path("timeoutMs").asLong(DEFAULT_ASK_TIMEOUT_MS);
|
||||
} catch (Exception e) {
|
||||
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
|
||||
return;
|
||||
}
|
||||
if (question.isBlank()) {
|
||||
ctx.status(400).json(Map.of("error", "bad_request", "detail", "question is required"));
|
||||
return;
|
||||
}
|
||||
timeout = Math.clamp(timeout, 1, MAX_ASK_TIMEOUT_MS);
|
||||
MessageService.AskResult r = messages.ask(id, question, timeout);
|
||||
switch (r.outcome()) {
|
||||
case ANSWERED -> ctx.status(200).json(Map.of("sessionId", id, "answered", true, "answer", r.answer()));
|
||||
case NO_WAITER -> ctx.status(409).json(Map.of(
|
||||
"sessionId", id, "error", "no_pending_send",
|
||||
"detail", "no primary is awaiting this turn to answer a question"));
|
||||
case TIMED_OUT -> ctx.status(202).json(Map.of(
|
||||
"sessionId", id, "status", "no_answer",
|
||||
"detail", "the primary did not answer within " + timeout + "ms"));
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The worker's structured reply ({@code bridge_reply}) — resolves the blocking send awaiting
|
||||
* on this session. 200 if a send was waiting, 409 if none was (late or spurious reply).
|
||||
*/
|
||||
private void replyMessage(Context ctx) {
|
||||
String id = ctx.pathParam("id");
|
||||
String content;
|
||||
try {
|
||||
content = mapper.readTree(ctx.body()).path("content").asText("");
|
||||
} catch (Exception e) {
|
||||
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
|
||||
return;
|
||||
}
|
||||
if (rendezvous.resolve(id, content)) {
|
||||
ctx.status(200).json(Map.of("sessionId", id, "delivered", true));
|
||||
} else {
|
||||
ctx.status(409).json(Map.of(
|
||||
"sessionId", id, "error", "no_pending_send",
|
||||
"detail", "no send is awaiting a reply for this session"));
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Live lifecycle status of a worker (MCP `bridge_status` wraps this in CB-105), plus its
|
||||
* <em>readiness</em> (CB-113): {@code ready} is true once the worker's Claude has connected the
|
||||
* bridge MCP — the reliable "available to receive a task" signal, unlike bare {@code idle}, which
|
||||
* is also true during boot.
|
||||
*/
|
||||
private void sessionStatus(Context ctx) {
|
||||
String id = ctx.pathParam("id");
|
||||
try {
|
||||
ctx.status(200).json(Map.of(
|
||||
"sessionId", id,
|
||||
"status", messages.status(id).name().toLowerCase(),
|
||||
"ready", presence.isPresent(id)));
|
||||
} catch (HerdrException e) {
|
||||
herdrError(ctx, e);
|
||||
}
|
||||
}
|
||||
|
||||
/** Poll an async (wait:false) delegation by ticket. 404 for an unknown/expired ticket. */
|
||||
private void taskStatus(Context ctx) {
|
||||
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"));
|
||||
if (v == null) {
|
||||
ctx.status(404).json(Map.of("error", "unknown_ticket", "detail", "no such task (or it has expired)"));
|
||||
return;
|
||||
}
|
||||
Map<String, Object> body = new LinkedHashMap<>();
|
||||
body.put("ticket", v.ticket());
|
||||
body.put("phase", v.phase().name().toLowerCase());
|
||||
if (v.reply() != null) {
|
||||
body.put("reply", v.reply());
|
||||
body.put("replySource", v.replySource());
|
||||
}
|
||||
if (v.detail() != null) {
|
||||
body.put("detail", v.detail());
|
||||
}
|
||||
ctx.status(200).json(body);
|
||||
}
|
||||
|
||||
/** Map a herdr failure: unknown target → 404, anything else → 502 (herdr is upstream). */
|
||||
private static void herdrError(Context ctx, HerdrException e) {
|
||||
if (e.code() != null && e.code().endsWith("_not_found")) {
|
||||
ctx.status(404).json(Map.of("error", "session_not_found", "detail", e.getMessage()));
|
||||
} else {
|
||||
ctx.status(502).json(Map.of("error", "herdr_error", "detail", e.getMessage()));
|
||||
}
|
||||
}
|
||||
|
||||
/** Stable JSON projection of an agent (null-safe for the start-time shape). */
|
||||
private static Map<String, Object> view(Agent a) {
|
||||
Map<String, Object> m = new LinkedHashMap<>();
|
||||
@@ -109,4 +401,22 @@ public final class BridgedApp {
|
||||
m.put("status", a.status().name().toLowerCase());
|
||||
return m;
|
||||
}
|
||||
|
||||
/** CB-301 projection of an authoritative bridge-owned session. */
|
||||
private static Map<String, Object> view(WorkerSession s) {
|
||||
Map<String, Object> m = new LinkedHashMap<>();
|
||||
m.put("terminalId", s.terminalId());
|
||||
m.put("paneId", s.paneId());
|
||||
m.put("profile", s.profile());
|
||||
m.put("cwd", s.cwd());
|
||||
m.put("ownerTerminal", s.ownerTerminal());
|
||||
m.put("state", s.state().name().toLowerCase());
|
||||
if (s.worktree() != null) {
|
||||
m.put("worktree", s.worktree());
|
||||
}
|
||||
if (s.branch() != null) {
|
||||
m.put("branch", s.branch());
|
||||
}
|
||||
return m;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,179 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.io.BufferedReader;
|
||||
import java.io.IOException;
|
||||
import java.io.InputStreamReader;
|
||||
import java.io.UncheckedIOException;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.nio.file.StandardCopyOption;
|
||||
import java.security.SecureRandom;
|
||||
import java.util.List;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
* Production {@link Worktrees} implementation that shells {@code git} via {@link ProcessBuilder}.
|
||||
* Non-zero exits become {@link WorktreeException}. Worktree directories live under a configurable
|
||||
* root (default: a sibling {@code .bridged-worktrees} of the repo root) so they are never nested
|
||||
* inside the primary working tree.
|
||||
*/
|
||||
public final class GitWorktrees implements Worktrees {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(GitWorktrees.class);
|
||||
|
||||
private final String configuredRoot;
|
||||
private final SecureRandom random = new SecureRandom();
|
||||
private final AtomicLong seq = new AtomicLong();
|
||||
|
||||
/** Default constructor: worktree root is derived per-repo as {@code <repoRoot>/../.bridged-worktrees}. */
|
||||
public GitWorktrees() {
|
||||
this(null);
|
||||
}
|
||||
|
||||
/** @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of the repo root. */
|
||||
public GitWorktrees(String configuredRoot) {
|
||||
this.configuredRoot = configuredRoot;
|
||||
}
|
||||
|
||||
@Override
|
||||
public String add(String repoRoot, String branch, String baseRef) {
|
||||
String base = (baseRef == null || baseRef.isBlank()) ? "HEAD" : baseRef;
|
||||
String nonce = nonce();
|
||||
Path root = resolveRoot(repoRoot);
|
||||
Path path = root.resolve(nonce);
|
||||
try {
|
||||
Files.createDirectories(root);
|
||||
} catch (IOException e) {
|
||||
throw new WorktreeException("cannot create worktree root " + root + ": " + e.getMessage(), e);
|
||||
}
|
||||
String wt = path.toAbsolutePath().toString();
|
||||
log.info("adding worktree branch={} path={} base={}", branch, wt, base);
|
||||
exec("git", "-C", repoRoot, "worktree", "add", wt, "-b", branch, base);
|
||||
return wt;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void remove(String repoRoot, String worktreePath) {
|
||||
Path p = Path.of(worktreePath);
|
||||
if (!Files.exists(p)) {
|
||||
log.debug("worktree {} already gone — nothing to remove", worktreePath);
|
||||
return;
|
||||
}
|
||||
log.info("removing worktree {}", worktreePath);
|
||||
exec("git", "-C", repoRoot, "worktree", "remove", "--force", worktreePath);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
|
||||
if (overlay == null || overlay.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
Path srcRoot = Path.of(repoRoot).toAbsolutePath().normalize();
|
||||
Path dstRoot = Path.of(worktreePath).toAbsolutePath().normalize();
|
||||
for (String rel : overlay) {
|
||||
Path src = srcRoot.resolve(rel).normalize();
|
||||
if (!Files.exists(src)) {
|
||||
log.debug("parity overlay source missing — skipping {}", rel);
|
||||
continue;
|
||||
}
|
||||
Path dst = dstRoot.resolve(rel).normalize();
|
||||
try {
|
||||
Files.createDirectories(dst.getParent());
|
||||
Files.copy(src, dst, StandardCopyOption.REPLACE_EXISTING, StandardCopyOption.COPY_ATTRIBUTES);
|
||||
log.debug("copied parity overlay {}", rel);
|
||||
} catch (IOException e) {
|
||||
throw new WorktreeException("cannot copy overlay " + rel + ": " + e.getMessage(), e);
|
||||
}
|
||||
if (isTracked(dstRoot, rel)) {
|
||||
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", rel);
|
||||
log.debug("marked overlay --skip-worktree {}", rel);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public String repoRoot(String cwd) {
|
||||
String out = exec("git", "-C", cwd, "rev-parse", "--show-toplevel");
|
||||
return Path.of(out.trim()).toAbsolutePath().normalize().toString();
|
||||
}
|
||||
|
||||
/** Resolve the directory that will hold per-session worktree checkouts. */
|
||||
private Path resolveRoot(String repoRoot) {
|
||||
if (configuredRoot != null && !configuredRoot.isBlank()) {
|
||||
return Path.of(configuredRoot).toAbsolutePath().normalize();
|
||||
}
|
||||
Path repo = Path.of(repoRoot).toAbsolutePath().normalize();
|
||||
return repo.resolveSibling(".bridged-worktrees");
|
||||
}
|
||||
|
||||
private String nonce() {
|
||||
return String.format("%06x", random.nextInt(1 << 24)) + "-" + seq.incrementAndGet();
|
||||
}
|
||||
|
||||
private boolean isTracked(Path worktreeRoot, String rel) {
|
||||
return exitCode("git", "-C", worktreeRoot.toString(), "ls-files", "--error-unmatch", rel) == 0;
|
||||
}
|
||||
|
||||
/**
|
||||
* Run a command and return its stdout. Non-zero exit → {@link WorktreeException} with both
|
||||
* stdout and stderr (merged by redirectErrorStream).
|
||||
*/
|
||||
private String exec(String... command) {
|
||||
String out;
|
||||
int code;
|
||||
Process p;
|
||||
try {
|
||||
p = new ProcessBuilder(command).redirectErrorStream(true).start();
|
||||
} catch (IOException e) {
|
||||
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
|
||||
}
|
||||
try (BufferedReader r = new BufferedReader(new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
|
||||
out = r.lines().collect(Collectors.joining("\n"));
|
||||
} catch (IOException e) {
|
||||
p.destroyForcibly();
|
||||
throw new UncheckedIOException(e);
|
||||
}
|
||||
try {
|
||||
if (!p.waitFor(30, TimeUnit.SECONDS)) {
|
||||
p.destroyForcibly();
|
||||
throw new WorktreeException("command timed out: " + String.join(" ", command) + "\n" + out);
|
||||
}
|
||||
code = p.exitValue();
|
||||
} catch (InterruptedException e) {
|
||||
p.destroyForcibly();
|
||||
Thread.currentThread().interrupt();
|
||||
throw new WorktreeException("interrupted waiting for command: " + String.join(" ", command), e);
|
||||
}
|
||||
if (code != 0) {
|
||||
throw new WorktreeException("exit " + code + " for: " + String.join(" ", command)
|
||||
+ (out.isBlank() ? "" : "\n" + out));
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
private int exitCode(String... command) {
|
||||
Process p;
|
||||
try {
|
||||
p = new ProcessBuilder(command).redirectErrorStream(true).start();
|
||||
} catch (IOException e) {
|
||||
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
|
||||
}
|
||||
try {
|
||||
if (!p.waitFor(30, TimeUnit.SECONDS)) {
|
||||
p.destroyForcibly();
|
||||
throw new WorktreeException("command timed out: " + String.join(" ", command));
|
||||
}
|
||||
return p.exitValue();
|
||||
} catch (InterruptedException e) {
|
||||
p.destroyForcibly();
|
||||
Thread.currentThread().interrupt();
|
||||
throw new WorktreeException("interrupted waiting for command: " + String.join(" ", command), e);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,411 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.inject.TurnListener;
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import dev.ltms.bridged.peer.SpawnRequest;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.security.SecureRandom;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Optional;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
/**
|
||||
* Authoritative in-daemon registry of the worker sessions this {@code bridged} process spawned.
|
||||
* Delegates spawn/teardown to a {@link PeerLauncher} (which performs subscription-guarded env
|
||||
* setup and process/materialization) and adds lifecycle tracking, ownership, and deterministic
|
||||
* teardown on top.
|
||||
*
|
||||
* <p>The state machine is intentionally one-shot / no-reuse: every acquired worker is fresh,
|
||||
* and a finished or released worker is torn down, never pooled. {@link #recycle} is a convenience
|
||||
* for {@code release + acquire} with a new distinct pane id.
|
||||
*
|
||||
* <p>The manager implements {@link TurnListener} so the injector's turn boundaries drive
|
||||
* {@code READY → BUSY → DONE} (or {@code FAILED}). It exposes a {@link WorkerPresence} view via
|
||||
* {@link #asPresence()}: any MCP contact from a worker marks it present and simultaneously
|
||||
* transitions the session {@code SPAWNING → READY}.
|
||||
*/
|
||||
public final class SessionManager implements TurnListener {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(SessionManager.class);
|
||||
|
||||
private final PeerLauncher launcher;
|
||||
private final Worktrees worktrees;
|
||||
private final ConcurrentHashMap<String /*paneId*/, WorkerSession> registry = new ConcurrentHashMap<>();
|
||||
private final WorkerPresence presence;
|
||||
private final SecureRandom nonceRandom = new SecureRandom();
|
||||
private final AtomicLong nonceSeq = new AtomicLong();
|
||||
private final LongSupplier nowNanos;
|
||||
private final int contextCap;
|
||||
|
||||
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
|
||||
public SessionManager(PeerLauncher launcher) {
|
||||
this(launcher, new GitWorktrees(), System::nanoTime, 0);
|
||||
}
|
||||
|
||||
/** Backward-compatible constructor with an injectable worktree seam. */
|
||||
public SessionManager(PeerLauncher launcher, Worktrees worktrees) {
|
||||
this(launcher, worktrees, System::nanoTime, 0);
|
||||
}
|
||||
|
||||
/** Test constructor with an injectable clock. */
|
||||
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos) {
|
||||
this(launcher, worktrees, nowNanos, 0);
|
||||
}
|
||||
|
||||
/** Production constructor with a configured context turn cap. */
|
||||
public SessionManager(PeerLauncher launcher, Worktrees worktrees, int contextCap) {
|
||||
this(launcher, worktrees, System::nanoTime, contextCap);
|
||||
}
|
||||
|
||||
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos,
|
||||
int contextCap) {
|
||||
this.launcher = launcher;
|
||||
this.worktrees = worktrees;
|
||||
this.presence = new PresenceBridge(this);
|
||||
this.nowNanos = nowNanos;
|
||||
this.contextCap = contextCap;
|
||||
}
|
||||
|
||||
/**
|
||||
* The single {@link WorkerPresence} view of this manager: it records availability and forwards
|
||||
* the signal to the {@code SPAWNING → READY} transition. Pass this to the {@code Injector} and
|
||||
* {@code BridgeMcp} where they previously accepted a plain {@link WorkerPresence}. The same
|
||||
* instance is returned every call — presence is shared state, so a fresh bridge per call would
|
||||
* fragment the {@code present} set and lose signals across callers.
|
||||
*/
|
||||
public WorkerPresence asPresence() {
|
||||
return presence;
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn a worker and register it as {@link WorkerSession.State#SPAWNING}. The caller's
|
||||
* identity is recorded as {@code ownerTerminal} ({@code null} for daemon/anon callers).
|
||||
*/
|
||||
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
|
||||
String ownerTerminal) {
|
||||
return acquire(profile, requestedCwd, callerCwd, ownerTerminal, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn a worker, optionally inside a fresh git worktree. When {@code wt} is non-null the
|
||||
* worktree is provisioned, parity-overlaid, and its path becomes the worker's cwd. On any
|
||||
* failure before registration the worktree is removed so no dangling checkout is left.
|
||||
*/
|
||||
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
|
||||
String ownerTerminal, WorktreeRequest wt) {
|
||||
if (wt == null) {
|
||||
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd);
|
||||
PeerHandle handle = launcher.spawn(req);
|
||||
String resolvedProfile = (profile == null || profile.isBlank())
|
||||
? launcher.defaultProfile() : profile;
|
||||
String cwd = launcher.effectiveCwd(req);
|
||||
long now = nowNanos.getAsLong();
|
||||
WorkerSession session = new WorkerSession(
|
||||
handle.id(),
|
||||
handle.terminalId(),
|
||||
resolvedProfile,
|
||||
cwd,
|
||||
ownerTerminal,
|
||||
now,
|
||||
now,
|
||||
0,
|
||||
WorkerSession.State.SPAWNING,
|
||||
null,
|
||||
null);
|
||||
registry.put(handle.id(), session);
|
||||
log.debug("acquired session id={} terminal={} profile={} owner={}",
|
||||
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
|
||||
return session;
|
||||
}
|
||||
return acquireWithWorktree(profile, requestedCwd, callerCwd, ownerTerminal, wt);
|
||||
}
|
||||
|
||||
/** Tear a worker down by pane id and remove it from the registry. Idempotent. */
|
||||
public void release(String paneId) {
|
||||
WorkerSession removed = registry.remove(paneId);
|
||||
if (removed != null) {
|
||||
log.debug("releasing session pane={} terminal={} state={}",
|
||||
removed.paneId(), removed.terminalId(), removed.state());
|
||||
}
|
||||
launcher.stop(paneId);
|
||||
if (removed != null && removed.worktree() != null) {
|
||||
worktrees.remove(worktrees.repoRoot(removed.cwd()), removed.worktree());
|
||||
}
|
||||
}
|
||||
|
||||
private WorkerSession acquireWithWorktree(String profile, String requestedCwd, String callerCwd,
|
||||
String ownerTerminal, WorktreeRequest wt) {
|
||||
String resolvedProfile = (profile == null || profile.isBlank())
|
||||
? launcher.defaultProfile() : profile;
|
||||
String repoRoot = worktrees.repoRoot(firstNonBlank(requestedCwd, callerCwd));
|
||||
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
|
||||
String path = null;
|
||||
PeerHandle handle;
|
||||
try {
|
||||
path = worktrees.add(repoRoot, branch, wt.baseRef());
|
||||
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(resolvedProfile));
|
||||
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd));
|
||||
} catch (RuntimeException e) {
|
||||
if (path != null) {
|
||||
try {
|
||||
worktrees.remove(repoRoot, path);
|
||||
} catch (RuntimeException cleanup) {
|
||||
log.warn("failed to clean up worktree {} after spawn error: {}", path, cleanup.getMessage());
|
||||
}
|
||||
}
|
||||
throw e;
|
||||
}
|
||||
long now = nowNanos.getAsLong();
|
||||
WorkerSession session = new WorkerSession(
|
||||
handle.id(),
|
||||
handle.terminalId(),
|
||||
resolvedProfile,
|
||||
resolveCwd(path, profile, callerCwd),
|
||||
ownerTerminal,
|
||||
now,
|
||||
now,
|
||||
0,
|
||||
WorkerSession.State.SPAWNING,
|
||||
path,
|
||||
branch);
|
||||
registry.put(handle.id(), session);
|
||||
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
|
||||
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
|
||||
return session;
|
||||
}
|
||||
|
||||
private String slug(String raw) {
|
||||
return raw == null ? "ticket" : raw.toLowerCase().replaceAll("[^a-z0-9]+", "-").replaceAll("^-+|-+$", "");
|
||||
}
|
||||
|
||||
private String nonce() {
|
||||
return String.format("%06x", nonceRandom.nextInt(1 << 24)) + "-" + nonceSeq.incrementAndGet();
|
||||
}
|
||||
|
||||
private static String firstNonBlank(String... values) {
|
||||
for (String v : values) {
|
||||
if (v != null && !v.isBlank()) return v;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Release the old session and acquire a fresh one with the same profile and working directory.
|
||||
* The new session is guaranteed to have a pane id distinct from the old one (no-reuse invariant).
|
||||
*/
|
||||
public WorkerSession recycle(String paneId) {
|
||||
WorkerSession old = registry.get(paneId);
|
||||
if (old == null) {
|
||||
throw new IllegalArgumentException("no session for paneId " + paneId);
|
||||
}
|
||||
release(paneId);
|
||||
return acquire(old.profile(), old.cwd(), old.cwd(), old.ownerTerminal());
|
||||
}
|
||||
|
||||
/** The session for {@code paneId}, if it is still registered and not released. */
|
||||
public Optional<WorkerSession> get(String paneId) {
|
||||
return Optional.ofNullable(registry.get(paneId));
|
||||
}
|
||||
|
||||
/** Bridge-owned roster: all registered sessions (acquired minus released). */
|
||||
public List<WorkerSession> roster() {
|
||||
return List.copyOf(registry.values());
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-304 merged roster+live view. The registry is authoritative for worktree, branch,
|
||||
* profile, owner, and state; the optional live agent supplies the herdr-reported status.
|
||||
*/
|
||||
public static Map<String, Object> rosterView(WorkerSession session, Agent live) {
|
||||
Map<String, Object> m = new LinkedHashMap<>();
|
||||
m.put("sessionId", session.terminalId());
|
||||
m.put("paneId", session.paneId());
|
||||
m.put("profile", session.profile());
|
||||
m.put("state", session.state().name().toLowerCase());
|
||||
if (session.worktree() != null) {
|
||||
m.put("worktree", session.worktree());
|
||||
}
|
||||
if (session.branch() != null) {
|
||||
m.put("branch", session.branch());
|
||||
}
|
||||
if (session.ownerTerminal() != null) {
|
||||
m.put("owner", session.ownerTerminal());
|
||||
}
|
||||
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
|
||||
return m;
|
||||
}
|
||||
|
||||
/** Lifecycle hook: worker became available on the bridge MCP. */
|
||||
void onReady(String terminalId) {
|
||||
transitionByTerminal(terminalId, WorkerSession.State.SPAWNING, WorkerSession.State.READY);
|
||||
}
|
||||
|
||||
/**
|
||||
* Lifecycle hook: a message was delivered into the worker — it is now busy on a turn.
|
||||
* The turn count is bumped and the activity timestamp is refreshed. A {@code DONE} session
|
||||
* can be re-delivered for multi-turn reuse until it is released.
|
||||
*/
|
||||
@Override
|
||||
public void onDelivered(String target) {
|
||||
WorkerSession current = findByTerminal(target);
|
||||
if (current == null) return;
|
||||
if (current.state() != WorkerSession.State.READY && current.state() != WorkerSession.State.DONE) {
|
||||
return;
|
||||
}
|
||||
long now = nowNanos.getAsLong();
|
||||
WorkerSession updated = current.withState(WorkerSession.State.BUSY).bumpTurn(now);
|
||||
if (replace(current, updated)) {
|
||||
log.debug("session transitioned terminal={} pane={} {} -> BUSY turn={}",
|
||||
target, current.paneId(), current.state(), updated.turnCount());
|
||||
}
|
||||
}
|
||||
|
||||
/** Lifecycle hook: the worker's delegated turn completed successfully. */
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
WorkerSession current = findByTerminal(target);
|
||||
if (current == null || current.state() != WorkerSession.State.BUSY) return;
|
||||
long now = nowNanos.getAsLong();
|
||||
WorkerSession updated = current.withState(WorkerSession.State.DONE).withActivity(now);
|
||||
if (replace(current, updated)) {
|
||||
log.debug("session transitioned terminal={} pane={} BUSY -> DONE turn={}",
|
||||
target, current.paneId(), updated.turnCount());
|
||||
}
|
||||
if (contextCap > 0 && updated.turnCount() >= contextCap) {
|
||||
release(current.paneId());
|
||||
}
|
||||
}
|
||||
|
||||
/** Lifecycle hook: the worker's delegated turn failed. */
|
||||
@Override
|
||||
public void onTurnFailed(String target) {
|
||||
onFailed(target);
|
||||
}
|
||||
|
||||
/** Lifecycle hook: the worker vanished or was dropped mid-life. */
|
||||
void onFailed(String target) {
|
||||
WorkerSession current = findByTerminal(target);
|
||||
if (current == null) return;
|
||||
if (current.state() == WorkerSession.State.RELEASED) return;
|
||||
if (replace(current, current.withState(WorkerSession.State.FAILED))) {
|
||||
log.debug("session marked failed terminal={} pane={}", target, current.paneId());
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Best-effort reap of sessions that have been idle longer than {@code idleTtlNanos}. Only
|
||||
* {@code READY} and {@code DONE} sessions are eligible — never a {@code SPAWNING} or
|
||||
* {@code BUSY} worker. Returns the number of sessions released.
|
||||
*/
|
||||
int reapIdle(long idleTtlNanos) {
|
||||
long now = nowNanos.getAsLong();
|
||||
int reaped = 0;
|
||||
for (WorkerSession s : roster()) {
|
||||
if (s.state() != WorkerSession.State.READY && s.state() != WorkerSession.State.DONE) {
|
||||
continue;
|
||||
}
|
||||
if (now - s.lastActivityAtNanos() > idleTtlNanos) {
|
||||
release(s.paneId());
|
||||
reaped++;
|
||||
}
|
||||
}
|
||||
return reaped;
|
||||
}
|
||||
|
||||
/**
|
||||
* Gracefully drain all registered sessions. For each session that is {@code BUSY}, poll up to
|
||||
* {@code timeoutNanos} for it to leave {@code BUSY}, then release it regardless. Non-busy
|
||||
* sessions are released immediately. A failure releasing one session is logged and does not
|
||||
* abort the rest.
|
||||
*/
|
||||
void drainAll(long timeoutNanos) {
|
||||
long deadline = System.nanoTime() + timeoutNanos;
|
||||
for (WorkerSession s : roster()) {
|
||||
try {
|
||||
if (s.state() == WorkerSession.State.BUSY) {
|
||||
while (System.nanoTime() < deadline) {
|
||||
WorkerSession current = registry.get(s.paneId());
|
||||
if (current == null || current.state() != WorkerSession.State.BUSY) {
|
||||
break;
|
||||
}
|
||||
try {
|
||||
long remaining = deadline - System.nanoTime();
|
||||
Thread.sleep(Math.min(TimeUnit.NANOSECONDS.toMillis(remaining), 50));
|
||||
} catch (InterruptedException e) {
|
||||
Thread.currentThread().interrupt();
|
||||
break;
|
||||
}
|
||||
}
|
||||
}
|
||||
release(s.paneId());
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("drain failed for pane={}; continuing with remaining sessions", s.paneId(), e);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Close this manager by draining all sessions. The timeout comes from configuration when set,
|
||||
* otherwise a sensible default.
|
||||
*/
|
||||
public void close(Integer drainTimeoutSeconds) {
|
||||
int seconds = (drainTimeoutSeconds != null && drainTimeoutSeconds > 0) ? drainTimeoutSeconds : 5;
|
||||
drainAll(TimeUnit.SECONDS.toNanos(seconds));
|
||||
}
|
||||
|
||||
/** Number of sessions currently registered. */
|
||||
public int size() {
|
||||
return registry.size();
|
||||
}
|
||||
|
||||
private WorkerSession findByTerminal(String terminalId) {
|
||||
for (WorkerSession s : registry.values()) {
|
||||
if (terminalId.equals(s.terminalId())) return s;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private void transitionByTerminal(String terminalId, WorkerSession.State from,
|
||||
WorkerSession.State to) {
|
||||
WorkerSession current = findByTerminal(terminalId);
|
||||
if (current == null || current.state() != from) return;
|
||||
long now = nowNanos.getAsLong();
|
||||
if (replace(current, current.withState(to).withActivity(now))) {
|
||||
log.debug("session transitioned terminal={} pane={} {} -> {}",
|
||||
terminalId, current.paneId(), from, to);
|
||||
}
|
||||
}
|
||||
|
||||
private boolean replace(WorkerSession expected, WorkerSession updated) {
|
||||
return registry.replace(expected.paneId(), expected, updated);
|
||||
}
|
||||
|
||||
private String resolveCwd(String requestedCwd, String profileName, String callerCwd) {
|
||||
return launcher.effectiveCwd(new SpawnRequest(profileName, requestedCwd, callerCwd));
|
||||
}
|
||||
|
||||
/** WorkerPresence bridge that also drives the manager's READY transition. */
|
||||
private static final class PresenceBridge extends WorkerPresence {
|
||||
private final SessionManager sessions;
|
||||
|
||||
PresenceBridge(SessionManager sessions) {
|
||||
this.sessions = sessions;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void markPresent(String terminal) {
|
||||
super.markPresent(terminal);
|
||||
sessions.onReady(terminal);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,70 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
/**
|
||||
* Periodic virtual-thread reaper that tears down {@code READY}/{@code DONE} sessions which have
|
||||
* exceeded their idle TTL. Modeled on {@link dev.ltms.bridged.inject.StatusPoller}: a single
|
||||
* virtual-thread loop, idempotent start/stop, and no {@code ScheduledExecutorService}.
|
||||
*/
|
||||
public final class SessionReaper {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
|
||||
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
|
||||
|
||||
private final SessionManager sessions;
|
||||
private final long idleTtlNanos;
|
||||
private final long intervalMillis;
|
||||
private volatile boolean running;
|
||||
private Thread thread;
|
||||
|
||||
/** Construct a reaper with the default 5-second polling interval. */
|
||||
public SessionReaper(SessionManager sessions, long idleTtlSeconds) {
|
||||
this(sessions, idleTtlSeconds, DEFAULT_INTERVAL_MILLIS);
|
||||
}
|
||||
|
||||
/** Construct a reaper with an explicit polling interval (useful for tests). */
|
||||
public SessionReaper(SessionManager sessions, long idleTtlSeconds, long intervalMillis) {
|
||||
this.sessions = sessions;
|
||||
this.idleTtlNanos = TimeUnit.SECONDS.toNanos(idleTtlSeconds);
|
||||
this.intervalMillis = intervalMillis;
|
||||
}
|
||||
|
||||
/** Start the reaper loop on a virtual thread. Idempotent. */
|
||||
public synchronized void start() {
|
||||
if (running) return;
|
||||
running = true;
|
||||
thread = Thread.ofVirtual().name("session-reaper").start(this::loop);
|
||||
log.info("session reaper started (idle ttl {}s, interval {}ms)",
|
||||
TimeUnit.NANOSECONDS.toSeconds(idleTtlNanos), intervalMillis);
|
||||
}
|
||||
|
||||
private void loop() {
|
||||
while (running) {
|
||||
try {
|
||||
sessions.reapIdle(idleTtlNanos);
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("session reaper iteration failed; continuing", e);
|
||||
}
|
||||
sleep();
|
||||
}
|
||||
}
|
||||
|
||||
private void sleep() {
|
||||
try {
|
||||
Thread.sleep(intervalMillis);
|
||||
} catch (InterruptedException e) {
|
||||
Thread.currentThread().interrupt();
|
||||
running = false;
|
||||
}
|
||||
}
|
||||
|
||||
/** Stop the reaper loop. Idempotent. */
|
||||
public synchronized void stop() {
|
||||
running = false;
|
||||
if (thread != null) thread.interrupt();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,58 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
/**
|
||||
* A bridge-owned worker session — the authoritative in-daemon record of a worker this
|
||||
* process spawned. Immutable; state transitions are performed by replacing the record in
|
||||
* {@link SessionManager}'s registry.
|
||||
*
|
||||
* @param paneId herdr pane handle — the registry key and the argument to teardown
|
||||
* @param terminalId herdr terminal handle — the {@code target} for send/read/status
|
||||
* @param profile the worker profile name that spawned this session
|
||||
* @param cwd the resolved working directory the worker started in
|
||||
* @param ownerTerminal the caller that requested this worker ({@code null} = daemon/anon)
|
||||
* @param spawnedAtNanos {@link System#nanoTime()} when the session was registered
|
||||
* @param lastActivityAtNanos {@link System#nanoTime()} of the most recent lifecycle event
|
||||
* @param turnCount number of delegated turns that have been delivered to this session
|
||||
* @param state current lifecycle state in the one-shot FSM
|
||||
*/
|
||||
public record WorkerSession(
|
||||
String paneId,
|
||||
String terminalId,
|
||||
String profile,
|
||||
String cwd,
|
||||
String ownerTerminal,
|
||||
long spawnedAtNanos,
|
||||
long lastActivityAtNanos,
|
||||
int turnCount,
|
||||
State state,
|
||||
String worktree,
|
||||
String branch) {
|
||||
|
||||
/** One-shot worker lifecycle states. */
|
||||
public enum State {
|
||||
SPAWNING,
|
||||
READY,
|
||||
BUSY,
|
||||
DONE,
|
||||
FAILED,
|
||||
RELEASED
|
||||
}
|
||||
|
||||
/** Return a copy of this session in {@code state}. */
|
||||
public WorkerSession withState(State state) {
|
||||
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
|
||||
lastActivityAtNanos, turnCount, state, worktree, branch);
|
||||
}
|
||||
|
||||
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
|
||||
public WorkerSession withActivity(long nowNanos) {
|
||||
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
|
||||
nowNanos, turnCount, state, worktree, branch);
|
||||
}
|
||||
|
||||
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
|
||||
public WorkerSession bumpTurn(long nowNanos) {
|
||||
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
|
||||
nowNanos, turnCount + 1, state, worktree, branch);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,12 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
/** Non-zero exit or I/O failure from a git worktree operation. */
|
||||
public final class WorktreeException extends RuntimeException {
|
||||
public WorktreeException(String message) {
|
||||
super(message);
|
||||
}
|
||||
|
||||
public WorktreeException(String message, Throwable cause) {
|
||||
super(message, cause);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,6 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
/** Ask {@link SessionManager#acquire} to provision an isolated worktree. null ⇒ run in the shared primary tree. */
|
||||
public record WorktreeRequest(String ticketSlug, String baseRef) {
|
||||
// ticketSlug seeds the branch name; baseRef null/blank ⇒ current HEAD of the repo.
|
||||
}
|
||||
@@ -0,0 +1,18 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/** Seam between {@link SessionManager} and git worktree operations. Tests use a recording fake. */
|
||||
public interface Worktrees {
|
||||
/** git -C <repoRoot> worktree add <path> -b <branch> <baseRef|HEAD>. Returns the worktree path. */
|
||||
String add(String repoRoot, String branch, String baseRef);
|
||||
|
||||
/** git -C <repoRoot> worktree remove --force <path>. Idempotent (already-gone tolerated). */
|
||||
void remove(String repoRoot, String worktreePath);
|
||||
|
||||
/** Copy each existing overlay path repoRoot→worktree; mark tracked ones --skip-worktree. */
|
||||
void overlayParity(String repoRoot, String worktreePath, List<String> overlay);
|
||||
|
||||
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
|
||||
String repoRoot(String cwd);
|
||||
}
|
||||
@@ -0,0 +1,488 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.herdr.Tab;
|
||||
import dev.ltms.bridged.herdr.Workspace;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.peer.Capability;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import dev.ltms.bridged.peer.SpawnRequest;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.security.SecureRandom;
|
||||
import java.util.ArrayList;
|
||||
import java.util.EnumSet;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.Function;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
/**
|
||||
* Spawns and lists worker sessions — the safe path from a delegation request to a
|
||||
* running off-subscription Claude.
|
||||
*
|
||||
* <p>The spawn sequence encodes the subscription boundary: build the worker env with
|
||||
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em>
|
||||
* touching herdr, and only then {@code agent.start}. A worker's base_url lives in the
|
||||
* env map handed to herdr and nowhere else; {@code bridged}'s own environment is never
|
||||
* mutated.
|
||||
*
|
||||
* <p>Placement: in the default {@code tab} policy a worker lands in its own tab inside a
|
||||
* dedicated worker space (found-or-created once, then shared), so workers never split or
|
||||
* clutter the user's real work spaces. Teardown removes the worker's pane <em>and</em> its
|
||||
* now-empty tab, tolerating an already-gone worker so a repeated DELETE is harmless.
|
||||
*/
|
||||
public final class ClaudeCodeLauncher implements PeerLauncher {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(ClaudeCodeLauncher.class);
|
||||
|
||||
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
|
||||
private static final int NAME_RETRIES = 8;
|
||||
|
||||
/**
|
||||
* A bridge-spawned worker label {@code claude-<profile>-<nonce>-<seq>} (see
|
||||
* {@link #startUniquelyNamed}); group 1 captures the 6-hex per-process {@code nonce}. The
|
||||
* profile segment may itself contain {@code -}, so the nonce/seq are anchored at the tail.
|
||||
* Names not matching this shape are not workers we started and are never reaped (CB-117).
|
||||
*/
|
||||
private static final Pattern WORKER_NAME = Pattern.compile("claude-.*-([0-9a-f]{6})-\\d+");
|
||||
|
||||
private final AgentControl agents;
|
||||
private final WorkspaceControl spaces;
|
||||
private final SubscriptionGuard guard;
|
||||
private final Map<String, BridgedConfig.Worker> profiles; // profile name → spawn settings
|
||||
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
|
||||
private final Function<String, String> env; // host env lookup (injectable for tests)
|
||||
private final AtomicLong nameSeq = new AtomicLong(); // per-worker counter (also the tab #)
|
||||
|
||||
/**
|
||||
* Standing instruction appended to the worker's system prompt so it returns its result via
|
||||
* {@code bridge_reply}. Injected as a launch flag, so nothing is written to the worker's
|
||||
* profile — it is guidance, and a worker that never replies is caught by the send's timeout.
|
||||
*/
|
||||
static final String REPLY_CHARTER =
|
||||
"You are an off-subscription worker in the claude-bridge fleet. Every message you "
|
||||
+ "receive arrives through the bridge, and the ONLY channel back to the sender is the "
|
||||
+ "bridge_reply MCP tool. Text you write in your terminal is NOT sent anywhere — the "
|
||||
+ "sender cannot see your screen, so an in-terminal answer is silently discarded. "
|
||||
+ "Therefore you MUST end EVERY turn by calling bridge_reply with `content` set to your "
|
||||
+ "complete response. This holds for every message without exception — tasks, questions, "
|
||||
+ "clarifications, acknowledgements, and ordinary back-and-forth conversation. Call "
|
||||
+ "bridge_reply exactly once, as the final action of your turn, with your full answer in "
|
||||
+ "`content`; never wait for confirmation first. If you end a turn without calling "
|
||||
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
|
||||
|
||||
// Per-process token mixed into each worker name so a fresh process (nameSeq back at 0)
|
||||
// cannot collide with same-profile workers that outlived a restart. See startUniquelyNamed.
|
||||
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
|
||||
|
||||
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Function<String, String> env) {
|
||||
this.agents = agents;
|
||||
this.spaces = spaces;
|
||||
this.guard = guard;
|
||||
this.profiles = Map.copyOf(profiles);
|
||||
this.defaultProfile = defaultProfile;
|
||||
this.env = env;
|
||||
}
|
||||
|
||||
/** The configured worker profile names (what {@code spawn(profile)} accepts). */
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return profiles.keySet();
|
||||
}
|
||||
|
||||
/** The parity-overlay file list for {@code profileName} (default list when unset). */
|
||||
@Override
|
||||
public List<String> parityOverlay(String profileName) {
|
||||
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
|
||||
if (name == null || name.isBlank()) {
|
||||
return List.of();
|
||||
}
|
||||
BridgedConfig.Worker cfg = profiles.get(name);
|
||||
return cfg == null ? List.of() : cfg.parityOverlay();
|
||||
}
|
||||
|
||||
/** The profile a no-argument {@link #spawn()} uses, or {@code null} if none is configured. */
|
||||
@Override
|
||||
public String defaultProfile() {
|
||||
return defaultProfile;
|
||||
}
|
||||
|
||||
/** Spawn a worker for the default profile in the resolved default cwd. */
|
||||
public Agent spawn() {
|
||||
return spawn(null, null, null);
|
||||
}
|
||||
|
||||
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
|
||||
public Agent spawn(String profileName) {
|
||||
return spawn(profileName, null, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn a worker. {@code profileName} null/blank → the default profile. The worker's working
|
||||
* directory (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd} (a
|
||||
* spawn argument), else the profile's configured {@code cwd}, else {@code callerCwd} (the
|
||||
* primary's cwd, when the spawn came from the primary over MCP), else the daemon's cwd — never
|
||||
* assumed to be {@code $HOME}. Guard runs before any herdr call.
|
||||
*/
|
||||
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
|
||||
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
|
||||
if (name == null || name.isBlank()) {
|
||||
throw new IllegalArgumentException("no default worker profile is configured — "
|
||||
+ "pass a profile; configured: " + profiles.keySet());
|
||||
}
|
||||
BridgedConfig.Worker cfg = profiles.get(name);
|
||||
if (cfg == null) {
|
||||
throw new IllegalArgumentException("unknown worker profile '" + name
|
||||
+ "' — configured: " + profiles.keySet());
|
||||
}
|
||||
String baseUrl = cfg.baseUrl();
|
||||
guard.assertWorker(baseUrl); // hard stop before we spawn anything
|
||||
|
||||
Map<String, String> workerEnv = new LinkedHashMap<>();
|
||||
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
|
||||
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
|
||||
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
|
||||
String token = env.apply(cfg.tokenEnv());
|
||||
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", token);
|
||||
|
||||
// CB-302: the worker checkpoint (commit → push → open its own PR). Push is free over SSH;
|
||||
// the only incremental grant is PR-create, a repo-scoped forge token injected here — opt-in
|
||||
// per profile via gitTokenEnv, and never mutating bridged's own env. The paired forge host
|
||||
// rides along only when a token is actually granted, so non-implementer profiles get neither.
|
||||
if (cfg.hasGitToken()) {
|
||||
String gitToken = resolveEnv(cfg.gitTokenEnv());
|
||||
if (gitToken != null) {
|
||||
workerEnv.put("GITEA_TOKEN", gitToken);
|
||||
putIfPresent(workerEnv, "GITEA_HOST", resolveEnv(cfg.gitHostEnv()));
|
||||
}
|
||||
}
|
||||
|
||||
// Mount the bridge MCP + reply charter as launch FLAGS (non-invasive: nothing written to
|
||||
// the worker's profile/config dir). Identity is connection-based, so the mount is shared.
|
||||
List<String> argv = argvWithBridge(cfg);
|
||||
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
|
||||
|
||||
return cfg.tabPlacement()
|
||||
? spawnInTab(cfg, workerEnv, argv, cwd)
|
||||
: spawnAsPane(cfg, workerEnv, argv, cwd);
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-112 cwd resolution: spawn arg → profile config → the primary's cwd → the daemon's cwd.
|
||||
* Never returns {@code null}/blank: {@code "."} (the daemon's own working directory) is the
|
||||
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
|
||||
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
|
||||
*/
|
||||
private static String resolveCwd(String requestedCwd, BridgedConfig.Worker cfg, String callerCwd) {
|
||||
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-301: the effective working directory a spawn for {@code profileName} would use, without
|
||||
* actually spawning. Used by {@link dev.ltms.bridged.session.SessionManager} to record the
|
||||
* resolved cwd in the session registry.
|
||||
*/
|
||||
public String effectiveCwd(String profileName, String requestedCwd, String callerCwd) {
|
||||
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
|
||||
if (name == null || name.isBlank()) {
|
||||
throw new IllegalArgumentException("no default worker profile is configured — "
|
||||
+ "pass a profile; configured: " + profiles.keySet());
|
||||
}
|
||||
BridgedConfig.Worker cfg = profiles.get(name);
|
||||
if (cfg == null) {
|
||||
throw new IllegalArgumentException("unknown worker profile '" + name
|
||||
+ "' — configured: " + profiles.keySet());
|
||||
}
|
||||
return resolveCwd(requestedCwd, cfg, callerCwd);
|
||||
}
|
||||
|
||||
private static String firstNonBlank(String... values) {
|
||||
for (String v : values) {
|
||||
if (v != null && !v.isBlank()) return v;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* The launch argv, plus — when {@code worker.mcpUrl} is set — inline {@code --mcp-config} for
|
||||
* the bridge server and {@code --append-system-prompt} for the {@link #REPLY_CHARTER}. Neither
|
||||
* touches the profile's config; both are pure command-line flags.
|
||||
*/
|
||||
private List<String> argvWithBridge(BridgedConfig.Worker cfg) {
|
||||
if (!cfg.hasMcp()) {
|
||||
return cfg.argv();
|
||||
}
|
||||
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
|
||||
+ cfg.mcpUrl() + "\"}}}";
|
||||
List<String> argv = new ArrayList<>(cfg.argv());
|
||||
argv.add("--mcp-config");
|
||||
argv.add(mcpJson);
|
||||
argv.add("--append-system-prompt");
|
||||
argv.add(REPLY_CHARTER);
|
||||
return argv;
|
||||
}
|
||||
|
||||
/** Dedicated worker space → own tab → start the worker (rooted at {@code cwd}) → drop the shell. */
|
||||
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
|
||||
List<String> argv, String cwd) {
|
||||
Workspace space = spaces.ensureWorkspace(cfg.workspace());
|
||||
Tab.Created tab = spaces.createTab(space.workspaceId());
|
||||
log.info("spawning worker profile={} base_url={} space={} tab={} cwd={}",
|
||||
cfg.profile(), cfg.baseUrl(), space.workspaceId(), tab.tab().tabId(), cwd);
|
||||
|
||||
Started started;
|
||||
try {
|
||||
started = startUniquelyNamed(cfg, workerEnv, argv, tab.tab().tabId(), cwd);
|
||||
} catch (RuntimeException e) {
|
||||
// The worker never started — don't leave the tab we just created orphaned.
|
||||
// Best-effort cleanup; never let it mask the real spawn failure.
|
||||
try {
|
||||
spaces.closeTab(tab.tab().tabId());
|
||||
} catch (RuntimeException cleanup) {
|
||||
log.warn("failed to close orphaned tab {} after spawn error: {}",
|
||||
tab.tab().tabId(), cleanup.getMessage());
|
||||
}
|
||||
throw e;
|
||||
}
|
||||
|
||||
// The worker is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell
|
||||
// so the tab holds only the worker; label the tab). They must not fail the spawn or
|
||||
// orphan the running worker — on error we log and still return it so the caller gets
|
||||
// its paneId and can tear it down.
|
||||
if (tab.rootPaneId() != null) {
|
||||
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
|
||||
} else {
|
||||
log.warn("tab {} had no seed pane in the create response; worker tab may hold an extra pane",
|
||||
tab.tab().tabId());
|
||||
}
|
||||
tidy("label tab " + tab.tab().tabId(),
|
||||
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
|
||||
log.info("worker started pane={} tab={} terminal={}",
|
||||
started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
|
||||
return started.agent();
|
||||
}
|
||||
|
||||
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
|
||||
private void tidy(String what, Runnable step) {
|
||||
try {
|
||||
step.run();
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("post-start step failed ({}) — worker is running regardless: {}", what, e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
/** Legacy placement: herdr splits the currently-focused tab; the worker still starts in {@code cwd}. */
|
||||
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
|
||||
List<String> argv, String cwd) {
|
||||
log.info("spawning worker (pane placement) profile={} base_url={} cwd={} argv={}",
|
||||
cfg.profile(), cfg.baseUrl(), cwd, argv);
|
||||
Agent worker = startUniquelyNamed(cfg, workerEnv, argv, null, cwd).agent();
|
||||
log.info("worker started pane={} terminal={}", worker.paneId(), worker.terminalId());
|
||||
return worker;
|
||||
}
|
||||
|
||||
/** A started worker together with the sequence its unique name/label used. */
|
||||
private record Started(Agent agent, long seq) {
|
||||
}
|
||||
|
||||
/**
|
||||
* Start the worker under a unique herdr agent name. herdr requires each running
|
||||
* agent's {@code name} to be distinct (a 2nd {@code name:"claude"} fails
|
||||
* {@code agent_name_taken}) — the exact case that makes multiple workers useful. The name
|
||||
* is {@code claude-<profile>-<nonce>-<seq>}: {@code seq} distinguishes workers within this
|
||||
* process, and the per-process {@code nonce} keeps a fresh process (whose {@code seq}
|
||||
* restarts at 0) from colliding with same-profile workers that outlived a restart. The
|
||||
* retry is a belt-and-braces backstop for the astronomically unlikely nonce+seq clash;
|
||||
* the name is a label only — herdr detects kind and status from terminal output, not it.
|
||||
*/
|
||||
private Started startUniquelyNamed(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
|
||||
List<String> argv, String tabId, String cwd) {
|
||||
HerdrException last = null;
|
||||
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
|
||||
long seq = nameSeq.incrementAndGet();
|
||||
String name = "claude-" + cfg.profile() + "-" + nameNonce + "-" + seq;
|
||||
try {
|
||||
return new Started(agents.start(name, argv, workerEnv, tabId, cwd), seq);
|
||||
} catch (HerdrException e) {
|
||||
if (!"agent_name_taken".equals(e.code())) throw e;
|
||||
log.debug("worker name '{}' taken, retrying", name);
|
||||
last = e;
|
||||
}
|
||||
}
|
||||
throw last;
|
||||
}
|
||||
|
||||
/** All herdr-tracked agents — discovery for "what workers exist". */
|
||||
@Override
|
||||
public List<Agent> list() {
|
||||
return agents.list();
|
||||
}
|
||||
|
||||
/**
|
||||
* Reap worker panes left behind by an earlier daemon process (CB-117). herdr keeps a worker's
|
||||
* pane alive across a daemon restart <em>by design</em>, and that pane's id is held only by its
|
||||
* spawner — so a worker whose owning process exited before issuing the matching teardown leaks
|
||||
* with nothing tracking it (there is no registry; {@link #list()} only asks herdr). On boot we
|
||||
* scan herdr for agents whose name matches our {@code claude-<profile>-<nonce>-<seq>} scheme with
|
||||
* a nonce <em>other</em> than this process's {@link #nameNonce}, and tear each one down (its pane
|
||||
* and, via {@link #stop}, its now-empty dedicated tab). A current-nonce worker is ours and live,
|
||||
* so it is left running; a user's own {@code claude} session carries no such name and is never
|
||||
* touched. Best-effort: a failed listing, or a failure to stop any one worker, is logged and
|
||||
* never aborts startup.
|
||||
*
|
||||
* @return the number of orphaned workers reaped
|
||||
*/
|
||||
@Override
|
||||
public int reapOrphanWorkers() {
|
||||
List<Agent> all;
|
||||
try {
|
||||
all = agents.list();
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("orphan-worker reap skipped — agent.list failed: {}", e.getMessage());
|
||||
return 0;
|
||||
}
|
||||
int reaped = 0;
|
||||
for (Agent a : all) {
|
||||
if (!isForeignWorker(a.name(), nameNonce)) continue;
|
||||
try {
|
||||
stop(a.paneId());
|
||||
reaped++;
|
||||
log.info("reaped orphan worker {} (pane={} tab={}) left by a prior daemon",
|
||||
a.name(), a.paneId(), a.tabId());
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("could not reap orphan worker {} (pane={}): {}",
|
||||
a.name(), a.paneId(), e.getMessage());
|
||||
}
|
||||
}
|
||||
if (reaped > 0) {
|
||||
log.info("orphan-worker reap complete — {} stale worker(s) removed at startup", reaped);
|
||||
}
|
||||
return reaped;
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether {@code name} is a bridge worker started by a <em>different</em> process than
|
||||
* {@code currentNonce} — the reap predicate (CB-117). True only for our naming scheme with a
|
||||
* foreign nonce: a non-worker name (no match, e.g. a user session) or our own live nonce is
|
||||
* excluded. Pure and package-private so the decision is unit-testable without herdr.
|
||||
*/
|
||||
static boolean isForeignWorker(String name, String currentNonce) {
|
||||
String nonce = workerNonce(name);
|
||||
return nonce != null && !nonce.equals(currentNonce);
|
||||
}
|
||||
|
||||
/** The 6-hex nonce embedded in a bridge worker name, or {@code null} if {@code name} isn't one. */
|
||||
static String workerNonce(String name) {
|
||||
if (name == null) return null;
|
||||
Matcher m = WORKER_NAME.matcher(name);
|
||||
return m.matches() ? m.group(1) : null;
|
||||
}
|
||||
|
||||
/** This process's worker-name nonce (a label component only; exposed for reaper tests). */
|
||||
String nameNonce() {
|
||||
return nameNonce;
|
||||
}
|
||||
|
||||
/**
|
||||
* Tear a worker down by pane id: close the pane, and close its tab <em>only</em> when the
|
||||
* worker is that tab's sole occupant. The single-pane check is what makes this safe
|
||||
* regardless of how the worker was placed (or a placement-config change across a restart):
|
||||
* a pane-placement worker sitting in one of the user's shared tabs has siblings, so its
|
||||
* tab is never closed — we only ever remove a tab we created to hold one worker.
|
||||
*
|
||||
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
|
||||
* (repeated DELETE, crashed worker) is treated as success; any other failure propagates so
|
||||
* a genuinely failed teardown is not reported as done.
|
||||
*/
|
||||
@Override
|
||||
public void stop(String paneId) {
|
||||
// Teardown knows only the paneId, not which profile spawned it. Attempt tab cleanup when any
|
||||
// profile uses tab placement (so the bridge may have created a dedicated worker tab); the
|
||||
// single-occupant check below is what actually protects the user's shared tabs.
|
||||
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
|
||||
try {
|
||||
agents.close(paneId);
|
||||
} catch (HerdrException e) {
|
||||
if (!isAlreadyGone(e)) throw e;
|
||||
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
|
||||
}
|
||||
if (loc != null && loc.tabPaneCount() == 1) {
|
||||
spaces.closeTab(loc.tabId());
|
||||
} else if (loc != null) {
|
||||
log.debug("not closing tab {} — it holds {} panes (not a dedicated worker tab)",
|
||||
loc.tabId(), loc.tabPaneCount());
|
||||
}
|
||||
}
|
||||
|
||||
/** Whether any configured profile places workers in their own tab (so tabs may need cleanup). */
|
||||
private boolean usesTabPlacement() {
|
||||
return profiles.values().stream().anyMatch(BridgedConfig.Worker::tabPlacement);
|
||||
}
|
||||
|
||||
/** True when a herdr error means the target is already gone (safe to treat as done). */
|
||||
private static boolean isAlreadyGone(HerdrException e) {
|
||||
return e.code() != null && e.code().endsWith("_not_found");
|
||||
}
|
||||
|
||||
// --- PeerLauncher SPI -------------------------------------------------------------------
|
||||
|
||||
@Override
|
||||
public Set<Capability> capabilities() {
|
||||
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
|
||||
if (hasGitTokenProfile()) {
|
||||
caps.add(Capability.SELF_PR);
|
||||
}
|
||||
return Set.copyOf(caps);
|
||||
}
|
||||
|
||||
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
|
||||
private boolean hasGitTokenProfile() {
|
||||
return profiles.values().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>Delegates to the three-arg {@link #spawn(String, String, String)} and wraps the
|
||||
* resulting herdr {@link Agent} in a {@link WorkerHandle} whose {@link PeerHandle#id()}
|
||||
* equals the agent's paneId.
|
||||
*/
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req) {
|
||||
Agent agent = spawn(req.profileName(), req.requestedCwd(), req.callerCwd());
|
||||
return new WorkerHandle(agent.paneId(), agent.terminalId());
|
||||
}
|
||||
|
||||
/** A concrete {@link PeerHandle} wrapping herdr agent coordinates. */
|
||||
private record WorkerHandle(String id, String terminalId) implements PeerHandle {
|
||||
}
|
||||
|
||||
@Override
|
||||
public String effectiveCwd(SpawnRequest req) {
|
||||
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
|
||||
}
|
||||
|
||||
private static void putIfPresent(Map<String, String> m, String k, String v) {
|
||||
if (v != null && !v.isBlank()) {
|
||||
m.put(k, v);
|
||||
}
|
||||
}
|
||||
|
||||
/** Host env lookup that tolerates an unconfigured (null/blank) var name — returns null then. */
|
||||
private String resolveEnv(String name) {
|
||||
return (name == null || name.isBlank()) ? null : env.apply(name);
|
||||
}
|
||||
}
|
||||
@@ -1,206 +0,0 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.herdr.Tab;
|
||||
import dev.ltms.bridged.herdr.Workspace;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.security.SecureRandom;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.Function;
|
||||
|
||||
/**
|
||||
* Spawns and lists worker sessions — the safe path from a delegation request to a
|
||||
* running off-subscription Claude.
|
||||
*
|
||||
* <p>The spawn sequence encodes the subscription boundary: build the worker env with
|
||||
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em>
|
||||
* touching herdr, and only then {@code agent.start}. A worker's base_url lives in the
|
||||
* env map handed to herdr and nowhere else; {@code bridged}'s own environment is never
|
||||
* mutated.
|
||||
*
|
||||
* <p>Placement: in the default {@code tab} policy a worker lands in its own tab inside a
|
||||
* dedicated worker space (found-or-created once, then shared), so workers never split or
|
||||
* clutter the user's real work spaces. Teardown removes the worker's pane <em>and</em> its
|
||||
* now-empty tab, tolerating an already-gone worker so a repeated DELETE is harmless.
|
||||
*/
|
||||
public final class WorkerService {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(WorkerService.class);
|
||||
|
||||
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
|
||||
private static final int NAME_RETRIES = 8;
|
||||
|
||||
private final AgentControl agents;
|
||||
private final WorkspaceControl spaces;
|
||||
private final SubscriptionGuard guard;
|
||||
private final BridgedConfig.Worker cfg;
|
||||
private final Function<String, String> env; // host env lookup (injectable for tests)
|
||||
private final AtomicLong nameSeq = new AtomicLong(); // per-worker counter (also the tab #)
|
||||
// Per-process token mixed into each worker name so a fresh process (nameSeq back at 0)
|
||||
// cannot collide with same-profile workers that outlived a restart. See startUniquelyNamed.
|
||||
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
|
||||
|
||||
public WorkerService(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||
BridgedConfig.Worker cfg, Function<String, String> env) {
|
||||
this.agents = agents;
|
||||
this.spaces = spaces;
|
||||
this.guard = guard;
|
||||
this.cfg = cfg;
|
||||
this.env = env;
|
||||
}
|
||||
|
||||
/** Spawn a worker for the configured profile. Guard runs before any herdr call. */
|
||||
public Agent spawn() {
|
||||
String baseUrl = cfg.baseUrl();
|
||||
guard.assertWorker(baseUrl); // hard stop before we spawn anything
|
||||
|
||||
Map<String, String> workerEnv = new LinkedHashMap<>();
|
||||
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
|
||||
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
|
||||
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
|
||||
String token = env.apply(cfg.tokenEnv());
|
||||
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", token);
|
||||
|
||||
return cfg.tabPlacement() ? spawnInTab(workerEnv) : spawnAsPane(workerEnv);
|
||||
}
|
||||
|
||||
/** Dedicated worker space → own tab → drop the placeholder shell so only the worker remains. */
|
||||
private Agent spawnInTab(Map<String, String> workerEnv) {
|
||||
Workspace space = spaces.ensureWorkspace(cfg.workspace());
|
||||
Tab.Created tab = spaces.createTab(space.workspaceId());
|
||||
log.info("spawning worker profile={} base_url={} space={} tab={}",
|
||||
cfg.profile(), cfg.baseUrl(), space.workspaceId(), tab.tab().tabId());
|
||||
|
||||
Started started;
|
||||
try {
|
||||
started = startUniquelyNamed(workerEnv, tab.tab().tabId());
|
||||
} catch (RuntimeException e) {
|
||||
// The worker never started — don't leave the tab we just created orphaned.
|
||||
// Best-effort cleanup; never let it mask the real spawn failure.
|
||||
try {
|
||||
spaces.closeTab(tab.tab().tabId());
|
||||
} catch (RuntimeException cleanup) {
|
||||
log.warn("failed to close orphaned tab {} after spawn error: {}",
|
||||
tab.tab().tabId(), cleanup.getMessage());
|
||||
}
|
||||
throw e;
|
||||
}
|
||||
|
||||
// The worker is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell
|
||||
// so the tab holds only the worker; label the tab). They must not fail the spawn or
|
||||
// orphan the running worker — on error we log and still return it so the caller gets
|
||||
// its paneId and can tear it down.
|
||||
if (tab.rootPaneId() != null) {
|
||||
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
|
||||
} else {
|
||||
log.warn("tab {} had no seed pane in the create response; worker tab may hold an extra pane",
|
||||
tab.tab().tabId());
|
||||
}
|
||||
tidy("label tab " + tab.tab().tabId(),
|
||||
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
|
||||
log.info("worker started pane={} tab={} terminal={}",
|
||||
started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
|
||||
return started.agent();
|
||||
}
|
||||
|
||||
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
|
||||
private void tidy(String what, Runnable step) {
|
||||
try {
|
||||
step.run();
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("post-start step failed ({}) — worker is running regardless: {}", what, e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
/** Legacy placement: herdr splits the currently-focused tab. */
|
||||
private Agent spawnAsPane(Map<String, String> workerEnv) {
|
||||
log.info("spawning worker (pane placement) profile={} base_url={} argv={}",
|
||||
cfg.profile(), cfg.baseUrl(), cfg.argv());
|
||||
Agent worker = startUniquelyNamed(workerEnv, null).agent();
|
||||
log.info("worker started pane={} terminal={}", worker.paneId(), worker.terminalId());
|
||||
return worker;
|
||||
}
|
||||
|
||||
/** A started worker together with the sequence its unique name/label used. */
|
||||
private record Started(Agent agent, long seq) {
|
||||
}
|
||||
|
||||
/**
|
||||
* Start the worker under a unique herdr agent name. herdr requires each running
|
||||
* agent's {@code name} to be distinct (a 2nd {@code name:"claude"} fails
|
||||
* {@code agent_name_taken}) — the exact case that makes multiple workers useful. The name
|
||||
* is {@code claude-<profile>-<nonce>-<seq>}: {@code seq} distinguishes workers within this
|
||||
* process, and the per-process {@code nonce} keeps a fresh process (whose {@code seq}
|
||||
* restarts at 0) from colliding with same-profile workers that outlived a restart. The
|
||||
* retry is a belt-and-braces backstop for the astronomically unlikely nonce+seq clash;
|
||||
* the name is a label only — herdr detects kind and status from terminal output, not it.
|
||||
*/
|
||||
private Started startUniquelyNamed(Map<String, String> workerEnv, String tabId) {
|
||||
HerdrException last = null;
|
||||
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
|
||||
long seq = nameSeq.incrementAndGet();
|
||||
String name = "claude-" + cfg.profile() + "-" + nameNonce + "-" + seq;
|
||||
try {
|
||||
return new Started(agents.start(name, cfg.argv(), workerEnv, tabId), seq);
|
||||
} catch (HerdrException e) {
|
||||
if (!"agent_name_taken".equals(e.code())) throw e;
|
||||
log.debug("worker name '{}' taken, retrying", name);
|
||||
last = e;
|
||||
}
|
||||
}
|
||||
throw last;
|
||||
}
|
||||
|
||||
/** All herdr-tracked agents — discovery for "what workers exist". */
|
||||
public List<Agent> list() {
|
||||
return agents.list();
|
||||
}
|
||||
|
||||
/**
|
||||
* Tear a worker down by pane id: close the pane, and close its tab <em>only</em> when the
|
||||
* worker is that tab's sole occupant. The single-pane check is what makes this safe
|
||||
* regardless of how the worker was placed (or a placement-config change across a restart):
|
||||
* a pane-placement worker sitting in one of the user's shared tabs has siblings, so its
|
||||
* tab is never closed — we only ever remove a tab we created to hold one worker.
|
||||
*
|
||||
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
|
||||
* (repeated DELETE, crashed worker) is treated as success; any other failure propagates so
|
||||
* a genuinely failed teardown is not reported as done.
|
||||
*/
|
||||
public void stop(String paneId) {
|
||||
WorkspaceControl.PaneLocation loc = cfg.tabPlacement() ? spaces.locatePane(paneId) : null;
|
||||
try {
|
||||
agents.close(paneId);
|
||||
} catch (HerdrException e) {
|
||||
if (!isAlreadyGone(e)) throw e;
|
||||
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
|
||||
}
|
||||
if (loc != null && loc.tabPaneCount() == 1) {
|
||||
spaces.closeTab(loc.tabId());
|
||||
} else if (loc != null) {
|
||||
log.debug("not closing tab {} — it holds {} panes (not a dedicated worker tab)",
|
||||
loc.tabId(), loc.tabPaneCount());
|
||||
}
|
||||
}
|
||||
|
||||
/** True when a herdr error means the target is already gone (safe to treat as done). */
|
||||
private static boolean isAlreadyGone(HerdrException e) {
|
||||
return e.code() != null && e.code().endsWith("_not_found");
|
||||
}
|
||||
|
||||
private static void putIfPresent(Map<String, String> m, String k, String v) {
|
||||
if (v != null && !v.isBlank()) {
|
||||
m.put(k, v);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -5,6 +5,7 @@ import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
@@ -46,6 +47,43 @@ class BridgedConfigTest {
|
||||
assertTrue(cfg.guard().offSubscriptionHosts().isEmpty());
|
||||
}
|
||||
|
||||
@Test
|
||||
void singleWorkerBecomesAOneEntryProfileMapWithItselfAsDefault(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("single.yaml");
|
||||
Files.writeString(f, """
|
||||
worker:
|
||||
profile: ltms-local
|
||||
baseUrl: http://gx00.gw:8000
|
||||
""");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertEquals(Set.of("ltms-local"), cfg.workerProfiles().keySet(), "legacy worker → one profile");
|
||||
assertEquals("ltms-local", cfg.defaultProfile());
|
||||
}
|
||||
|
||||
@Test
|
||||
void loadsMultipleWorkerProfilesWithADefault(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("multi.yaml");
|
||||
Files.writeString(f, """
|
||||
workers:
|
||||
gx10:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
argv: ["ccs", "gx10"]
|
||||
ollama:
|
||||
baseUrl: http://ollama.ltms.dev
|
||||
argv: ["ccs", "ollama"]
|
||||
defaultWorker: gx10
|
||||
guard:
|
||||
offSubscriptionHosts: [gx10.gw, ollama.ltms.dev]
|
||||
""");
|
||||
|
||||
BridgedConfig cfg = BridgedConfig.load(f);
|
||||
assertEquals(Set.of("gx10", "ollama"), cfg.workerProfiles().keySet());
|
||||
assertEquals("gx10", cfg.defaultProfile());
|
||||
assertEquals("ollama", cfg.workerProfiles().get("ollama").profile(), "profile defaults to its map key");
|
||||
assertEquals("http://gx10.gw:8000", cfg.workerProfiles().get("gx10").baseUrl());
|
||||
}
|
||||
|
||||
@Test
|
||||
void ignoresUnknownKeys(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("future.yaml");
|
||||
|
||||
@@ -0,0 +1,40 @@
|
||||
package dev.ltms.bridged.herdr;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/** Unit-level behaviour of {@link AgentControl} over a fake herdr. */
|
||||
class AgentControlTest {
|
||||
|
||||
/** The {@code text} of every agent.send, in call order. */
|
||||
@SuppressWarnings("unchecked")
|
||||
private static List<String> sendTexts(FakeHerdr herdr) {
|
||||
return herdr.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.send"))
|
||||
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
|
||||
.toList();
|
||||
}
|
||||
|
||||
@Test
|
||||
void sendDeliversThePayloadThenAStandaloneSubmitKey() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
new AgentControl(herdr).send("term_x", "do the thing");
|
||||
|
||||
// The Enter must be its own event — appended to the paste it would be swallowed as text.
|
||||
assertEquals(List.of("do the thing", "\r"), sendTexts(herdr),
|
||||
"payload paste first, then a separate carriage-return keystroke to submit it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void sendPreservesEmbeddedNewlinesAndSubmitsOnlyOnce() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
new AgentControl(herdr).send("term_x", "line1\nline2");
|
||||
|
||||
assertEquals(List.of("line1\nline2", "\r"), sendTexts(herdr),
|
||||
"multiline content is delivered verbatim; a single trailing Enter submits it");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,43 @@
|
||||
package dev.ltms.bridged.herdr;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/** Wire mapping and injectability of {@link AgentStatus}, including the CB-115 {@code done} state. */
|
||||
class AgentStatusTest {
|
||||
|
||||
@Test
|
||||
void mapsTheKnownWireStrings() {
|
||||
assertEquals(AgentStatus.IDLE, AgentStatus.fromWire("idle"));
|
||||
assertEquals(AgentStatus.WORKING, AgentStatus.fromWire("working"));
|
||||
assertEquals(AgentStatus.BLOCKED, AgentStatus.fromWire("blocked"));
|
||||
assertEquals(AgentStatus.DONE, AgentStatus.fromWire("done"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void mapsDoneCaseInsensitively() {
|
||||
assertEquals(AgentStatus.DONE, AgentStatus.fromWire("DONE"));
|
||||
assertEquals(AgentStatus.DONE, AgentStatus.fromWire("Done"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void unknownAndNullFallToUnknown() {
|
||||
assertEquals(AgentStatus.UNKNOWN, AgentStatus.fromWire("unknown"));
|
||||
assertEquals(AgentStatus.UNKNOWN, AgentStatus.fromWire("something-else"));
|
||||
assertEquals(AgentStatus.UNKNOWN, AgentStatus.fromWire(null));
|
||||
}
|
||||
|
||||
@Test
|
||||
void doneIsInjectableLikeIdle() {
|
||||
// The whole point of CB-115: a finished worker herdr reports as `done` must be deliverable,
|
||||
// not treated as UNKNOWN (which wedged delivery and mis-fired the stall failure).
|
||||
assertTrue(AgentStatus.DONE.injectable());
|
||||
assertTrue(AgentStatus.IDLE.injectable());
|
||||
assertTrue(AgentStatus.BLOCKED.injectable());
|
||||
assertFalse(AgentStatus.WORKING.injectable());
|
||||
assertFalse(AgentStatus.UNKNOWN.injectable());
|
||||
}
|
||||
}
|
||||
@@ -16,15 +16,20 @@ public final class FakeHerdr implements HerdrClient {
|
||||
public record Call(String method, Object params) {
|
||||
}
|
||||
|
||||
/** The foreground PID of the one agent pane (term_a) in the canned {@code pane.process_info}. */
|
||||
public static final long WORKER_PID = 4242;
|
||||
|
||||
private final ObjectMapper mapper = new ObjectMapper();
|
||||
public final List<Call> calls = new ArrayList<>();
|
||||
private boolean healthy = true;
|
||||
private final List<String> extraWorkspaces = new ArrayList<>();
|
||||
private final List<String> extraAgents = new ArrayList<>();
|
||||
private int agentNameTakenFor = 0;
|
||||
private int workerTabPaneCount = 1;
|
||||
private String paneCloseErrorCode = null;
|
||||
private String agentSendErrorCode = null;
|
||||
private volatile String agentStatus = "idle"; // what agent.get reports
|
||||
private volatile String agentStatus = "idle"; // steady-state agent.get status
|
||||
private volatile String readText = "worker transcript tail"; // canned agent.read output
|
||||
|
||||
public FakeHerdr healthy(boolean h) {
|
||||
this.healthy = h;
|
||||
@@ -55,12 +60,32 @@ public final class FakeHerdr implements HerdrClient {
|
||||
return this;
|
||||
}
|
||||
|
||||
/** The text {@code agent.read} returns (the CB-106 completion scrape). */
|
||||
public FakeHerdr readText(String text) {
|
||||
this.readText = text;
|
||||
return this;
|
||||
}
|
||||
|
||||
/** Make {@code agent.send} fail with this herdr error code. */
|
||||
public FakeHerdr agentSendFailsWith(String code) {
|
||||
this.agentSendErrorCode = code;
|
||||
return this;
|
||||
}
|
||||
|
||||
|
||||
/**
|
||||
* Seed a named agent into {@code agent.list} (e.g. an orphaned worker for CB-117 reaper tests).
|
||||
* The {@code name} carries the worker label the reaper keys on; {@code paneId}/{@code tabId}
|
||||
* locate its pane for teardown.
|
||||
*/
|
||||
public FakeHerdr withAgent(String name, String terminalId, String paneId, String tabId) {
|
||||
extraAgents.add(("{\"terminal_id\":\"%s\",\"agent\":\"claude\",\"agent_status\":\"idle\","
|
||||
+ "\"name\":\"%s\",\"agent_session\":{\"kind\":\"id\",\"value\":\"sess-%s\"},"
|
||||
+ "\"workspace_id\":\"wQ\",\"tab_id\":\"%s\",\"pane_id\":\"%s\"}")
|
||||
.formatted(terminalId, name, terminalId, tabId, paneId));
|
||||
return this;
|
||||
}
|
||||
|
||||
/** Seed an additional workspace into {@code workspace.list} (e.g. a pre-existing worker space). */
|
||||
public FakeHerdr withWorkspace(String id, String label) {
|
||||
extraWorkspaces.add(("{\"workspace_id\":\"%s\",\"label\":\"%s\",\"focused\":false,"
|
||||
@@ -91,11 +116,12 @@ public final class FakeHerdr implements HerdrClient {
|
||||
{"workspace_id":"w1","label":"dev-mgnl","focused":true,"pane_count":7,"agent_status":"unknown"},
|
||||
{"workspace_id":"w2","label":"ltms","focused":false,"pane_count":5,"agent_status":"done"}%s]}""")
|
||||
.formatted(extraWorkspaces.isEmpty() ? "" : "," + String.join(",", extraWorkspaces)));
|
||||
case "agent.list" -> mapper.readTree("""
|
||||
case "agent.list" -> mapper.readTree(("""
|
||||
{"type":"agent_list","agents":[
|
||||
{"terminal_id":"term_a","agent":"claude","agent_status":"idle",
|
||||
"agent_session":{"kind":"id","value":"sess-1111"},
|
||||
"workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}]}""");
|
||||
"workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}%s]}""")
|
||||
.formatted(extraAgents.isEmpty() ? "" : "," + String.join(",", extraAgents)));
|
||||
case "agent.send" -> {
|
||||
if (agentSendErrorCode != null) {
|
||||
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.send failed",
|
||||
@@ -107,6 +133,8 @@ public final class FakeHerdr implements HerdrClient {
|
||||
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":"claude",
|
||||
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
|
||||
.formatted(agentStatus));
|
||||
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
|
||||
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
|
||||
case "agent.start" -> {
|
||||
long starts = calls.stream().filter(c -> c.method().equals("agent.start")).count();
|
||||
if (starts <= agentNameTakenFor) {
|
||||
@@ -114,10 +142,12 @@ public final class FakeHerdr implements HerdrClient {
|
||||
"herdr error [agent_name_taken]: agent name already used",
|
||||
"agent_name_taken", null);
|
||||
}
|
||||
yield mapper.readTree("""
|
||||
long n = starts - agentNameTakenFor;
|
||||
yield mapper.readTree(("""
|
||||
{"type":"agent_started","agent":{
|
||||
"terminal_id":"term_new","name":"claude","agent_status":"unknown",
|
||||
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"w9:pW"}}""");
|
||||
"terminal_id":"term_new_%d","name":"claude","agent_status":"unknown",
|
||||
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"w9:pW_%d"}}""")
|
||||
.formatted(n, n));
|
||||
}
|
||||
case "workspace.create" -> mapper.readTree("""
|
||||
{"type":"workspace_created",
|
||||
@@ -141,6 +171,21 @@ public final class FakeHerdr implements HerdrClient {
|
||||
case "pane.get" -> mapper.readTree("""
|
||||
{"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9",
|
||||
"tab_id":"w9:t2","agent_status":"idle"}}""");
|
||||
case "pane.list" -> mapper.readTree("""
|
||||
{"type":"pane_list","panes":[
|
||||
{"pane_id":"w2:p7","terminal_id":"term_a","workspace_id":"w2","tab_id":"w2:t7","agent":"claude"},
|
||||
{"pane_id":"w2:p9","terminal_id":"term_shell","workspace_id":"w2","tab_id":"w2:t8"}]}""");
|
||||
case "pane.process_info" -> {
|
||||
Object paneId = params instanceof java.util.Map<?, ?> m ? m.get("pane_id") : null;
|
||||
yield "w2:p7".equals(paneId)
|
||||
? mapper.readTree(("""
|
||||
{"type":"pane_process_info","process_info":{"pane_id":"w2:p7","shell_pid":%d,
|
||||
"foreground_processes":[{"pid":%d,"name":"node","argv0":"claude"}]}}""")
|
||||
.formatted(WORKER_PID, WORKER_PID))
|
||||
: mapper.readTree("""
|
||||
{"type":"pane_process_info","process_info":{"pane_id":"w2:p9","shell_pid":9001,
|
||||
"foreground_processes":[]}}""");
|
||||
}
|
||||
case "pane.close" -> {
|
||||
if (paneCloseErrorCode != null) {
|
||||
throw new HerdrException("herdr error [" + paneCloseErrorCode + "]: pane.close failed",
|
||||
|
||||
@@ -0,0 +1,43 @@
|
||||
package dev.ltms.bridged.herdr;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import org.junit.jupiter.api.Tag;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
import static org.junit.jupiter.api.Assumptions.assumeTrue;
|
||||
|
||||
/**
|
||||
* Contract test for the herdr half of connection-based identity against a REAL herdr: spawn a
|
||||
* harmless probe, read its actual {@code shell_pid} from {@code pane.process_info}, and confirm
|
||||
* {@link PaneLocator} resolves that PID back to the probe's own {@code terminal_id}.
|
||||
*
|
||||
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
|
||||
*/
|
||||
@Tag("contract")
|
||||
class PaneLocatorContractTest {
|
||||
|
||||
@Test
|
||||
void resolvesTheTerminalOwningARealProcessPid() throws Exception {
|
||||
assumeTrue(Files.exists(UnixSocketHerdrClient.defaultSocketPath()), "no herdr socket — skipping");
|
||||
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
Agent probe = agents.start("__pidprobe__", List.of("bash", "-c", "sleep 20"), Map.of());
|
||||
try {
|
||||
JsonNode info = herdr.call("pane.process_info", Map.of("pane_id", probe.paneId()))
|
||||
.path("process_info");
|
||||
long shellPid = info.path("shell_pid").asLong(-1);
|
||||
assertTrue(shellPid > 0, "probe pane should report a shell pid");
|
||||
|
||||
assertEquals(probe.terminalId(), new PaneLocator(herdr).terminalForPid(shellPid),
|
||||
"a real PID must resolve back to its own pane's terminal_id");
|
||||
} finally {
|
||||
agents.close(probe.paneId());
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,27 @@
|
||||
package dev.ltms.bridged.herdr;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/** Unit tests for PID → pane resolution (the herdr half of connection-based MCP identity). */
|
||||
class PaneLocatorTest {
|
||||
|
||||
private final PaneLocator loc = new PaneLocator(new FakeHerdr());
|
||||
|
||||
@Test
|
||||
void resolvesTerminalForAForegroundPid() {
|
||||
assertEquals("term_a", loc.terminalForPid(FakeHerdr.WORKER_PID));
|
||||
}
|
||||
|
||||
@Test
|
||||
void nullForAPidInNoPane() {
|
||||
assertNull(loc.terminalForPid(999_999));
|
||||
}
|
||||
|
||||
@Test
|
||||
void nullForNonPositivePid() {
|
||||
assertNull(loc.terminalForPid(0));
|
||||
assertNull(loc.terminalForPid(-1));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,227 @@
|
||||
package dev.ltms.bridged.inject;
|
||||
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/** Unit behaviour of the CB-106 completion resolver in isolation from the injector. */
|
||||
class CompletionResolverTest {
|
||||
|
||||
@Test
|
||||
void skipsTheScrapeWhenNoSendIsWaiting() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
|
||||
resolver.resolve("term_a", null); // no in-flight turn captured for this target
|
||||
|
||||
assertFalse(herdr.called("agent.read"),
|
||||
"a turn nobody is blocked on must not cost a transcript scrape");
|
||||
}
|
||||
|
||||
@Test
|
||||
void failSkipsTheScrapeWhenNoSendIsWaiting() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
|
||||
resolver.fail("term_a", null); // no in-flight turn, and no registered waiter to fall back to
|
||||
|
||||
assertFalse(herdr.called("agent.read"),
|
||||
"a wedge nobody is blocked on must not cost a transcript scrape");
|
||||
}
|
||||
|
||||
@Test
|
||||
void captureBaselineSkipsTheReadWhenNoSendIsWaiting() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ X\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
|
||||
resolver.captureBaseline("term_a"); // no send to attribute a later completion to
|
||||
|
||||
assertFalse(herdr.called("agent.read"),
|
||||
"with no waiting send there is no turn to baseline — skip the scrape");
|
||||
}
|
||||
|
||||
// --- CB-115 clean scrape: extract the last assistant block ----------------
|
||||
|
||||
@Test
|
||||
void extractsTheLastAssistantBlockStrippingChrome() {
|
||||
String raw = """
|
||||
⏺ Reading the file…
|
||||
|
||||
⏺ Done. The bug was an off-by-one in the loop bound.
|
||||
|
||||
╭──────────────────────────────────────╮
|
||||
│ > │
|
||||
╰──────────────────────────────────────╯
|
||||
⏵⏵ auto mode on · ? for shortcuts
|
||||
""";
|
||||
assertEquals("Done. The bug was an off-by-one in the loop bound.",
|
||||
CompletionResolver.lastAssistantBlock(raw));
|
||||
}
|
||||
|
||||
@Test
|
||||
void keepsMultiLineAssistantContent() {
|
||||
String raw = "⏺ Line one.\nLine two.\n❯ ";
|
||||
assertEquals("Line one.\nLine two.", CompletionResolver.lastAssistantBlock(raw));
|
||||
}
|
||||
|
||||
@Test
|
||||
void fallsBackToRawTextWhenThereIsNoMarker() {
|
||||
String raw = "plain worker output with no glyph";
|
||||
assertEquals("plain worker output with no glyph", CompletionResolver.lastAssistantBlock(raw));
|
||||
}
|
||||
|
||||
@Test
|
||||
void blankScrapeYieldsEmpty() {
|
||||
assertTrue(CompletionResolver.lastAssistantBlock("").isEmpty());
|
||||
assertTrue(CompletionResolver.lastAssistantBlock(null).isEmpty());
|
||||
}
|
||||
|
||||
@Test
|
||||
void stripsSpinnerAndRuleChrome() {
|
||||
String raw = """
|
||||
⏺ Channel check confirmed — your message got through.
|
||||
|
||||
✻ Brewed for 11s
|
||||
|
||||
─────────────────────────────────────
|
||||
""";
|
||||
assertEquals("Channel check confirmed — your message got through.",
|
||||
CompletionResolver.lastAssistantBlock(raw));
|
||||
}
|
||||
|
||||
@Test
|
||||
void cutsANextTurnPromptEchoAndTrailingTipsFromTheBlock() {
|
||||
// The exact turn-2 leak: the scrape captured the settled answer, then a "✻ Cooked" spinner,
|
||||
// then the NEXT turn's echoed prompt, then a "✶ Forming…" spinner and trailing tips/warnings
|
||||
// whose lines (⎿, ⚠) are not themselves chrome-terminated. Stopping at the first boundary
|
||||
// (the ✻ spinner) is what keeps every one of those interface lines out of the reply.
|
||||
String raw = """
|
||||
⏺ Channel confirmed — the bridge reply delivered successfully.
|
||||
|
||||
✻ Cooked for 9s
|
||||
|
||||
❯ Thanks. Now a small task: what is 17 * 23? Show just the number.
|
||||
|
||||
|
||||
|
||||
✶ Forming…
|
||||
⎿ Tip: Name your conversations with /rename
|
||||
⚠ claude.ai connectors are disabled because ANTHROPIC_API_KEY is set
|
||||
""";
|
||||
assertEquals("Channel confirmed — the bridge reply delivered successfully.",
|
||||
CompletionResolver.lastAssistantBlock(raw));
|
||||
}
|
||||
|
||||
// --- CB-115 misattribution guard: suppress a stale (unchanged) completion -------
|
||||
|
||||
@Test
|
||||
void suppressesACompletionWhoseScrapeIsUnchangedFromDelivery() {
|
||||
// Rapid back-to-back turn: the pane still shows the PREVIOUS turn's answer when this turn's
|
||||
// (misattributed) completion boundary fires. The scrape == the delivery baseline, so the
|
||||
// send must NOT be resolved with the stale answer.
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ 391\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
|
||||
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
|
||||
// The turn as captured at delivery: its waiter, and the previous turn's answer still on screen.
|
||||
var turn = new CompletionResolver.InFlight(waiter, "391");
|
||||
resolver.resolve("term_a", turn); // scrape still "391" == baseline → suppress
|
||||
|
||||
assertFalse(waiter.isDone(), "a completion with no output change must not resolve the send");
|
||||
assertTrue(rendezvous.isWaiting("term_a"), "the send stays waiting for a real reply");
|
||||
}
|
||||
|
||||
@Test
|
||||
void resolvesACompletionWhoseScrapeChangedSinceDelivery() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ No, 391 = 17 × 23.\n❯ "); // the worker's real answer
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
// Delivery baseline was the previous turn's "391"; the scrape now differs → resolve.
|
||||
var turn = new CompletionResolver.InFlight(waiter, "391");
|
||||
resolver.resolve("term_a", turn);
|
||||
|
||||
assertTrue(waiter.isDone(), "a completion with new output must resolve the send");
|
||||
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
|
||||
assertEquals("No, 391 = 17 × 23.", waiter.getNow(null).text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void suppressesAnUnchangedCompletionEvenWhenTheBlockExceedsTheScrapeCap() {
|
||||
// The fan-out issue-hunt finding: captureBaseline once stored the RAW (unclipped) assistant
|
||||
// block while resolve compares against a clip()'d tail. For a block longer than MAX_SCRAPE_CHARS
|
||||
// the two capped representations differ even when the pane never changed, so the CB-115
|
||||
// byte-identical guard failed to fire and a stale completion could resolve the send. Both sides
|
||||
// must clip identically; here an unchanged >cap block on rapid back-to-back turns stays suppressed.
|
||||
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
|
||||
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
|
||||
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
|
||||
resolver.captureBaseline("term_a"); // baseline is the clipped >cap block
|
||||
var turn = resolver.inFlight("term_a");
|
||||
assertEquals(CompletionResolver.MAX_SCRAPE_CHARS, turn.baseline().length(),
|
||||
"the delivery baseline is clipped to the same cap resolve() applies to the tail");
|
||||
|
||||
resolver.resolve("term_a", turn); // scrape unchanged → clipped tail == baseline → suppress
|
||||
|
||||
assertFalse(waiter.isDone(),
|
||||
"an unchanged >cap block must still be recognised as stale and suppressed");
|
||||
assertTrue(rendezvous.isWaiting("term_a"), "the send stays waiting for a real reply");
|
||||
}
|
||||
|
||||
@Test
|
||||
void resolvesWhenThereIsNoBaseline() {
|
||||
// No delivery baseline (e.g. the pre-turn read failed) ⇒ never suppress; the completion resolves.
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ hello\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||
|
||||
assertTrue(waiter.isDone(), "with no baseline a completion resolves as before");
|
||||
assertEquals("hello", waiter.getNow(null).text());
|
||||
}
|
||||
|
||||
// --- CB-116 waiter identity: a late completion never crosses into the next turn ---------
|
||||
|
||||
@Test
|
||||
void aLateCompletionForOneTurnNeverResolvesTheNextTurnsWaiter() {
|
||||
// The cross-turn stale reply the conversation test surfaced: turn N's completion fallback
|
||||
// fires AFTER turn N was resolved by an explicit bridge_reply and turn N+1 has opened its own
|
||||
// waiter on the same session. Resolving "whatever is waiting now" would hand turn N's stale
|
||||
// scrape to turn N+1; targeting turn N's captured waiter makes the late completion a no-op.
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ turn N answer\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
|
||||
var waiterN = rendezvous.open("term_a"); // turn N's send
|
||||
// The turn as the injector captured it at delivery (waiter + pre-turn baseline).
|
||||
var turnN = new CompletionResolver.InFlight(waiterN, "an earlier answer");
|
||||
|
||||
// Turn N is resolved by the worker's explicit reply.
|
||||
assertTrue(rendezvous.resolve("term_a", "N replied"));
|
||||
|
||||
// Turn N+1's send opens its own waiter on the same session (replacing the registered one).
|
||||
var waiterN1 = rendezvous.open("term_a");
|
||||
|
||||
resolver.resolve("term_a", turnN); // turn N's completion fallback finally fires
|
||||
|
||||
assertFalse(waiterN1.isDone(), "turn N's late completion must not resolve turn N+1's waiter");
|
||||
assertEquals(Rendezvous.Kind.REPLY, waiterN.getNow(null).kind(),
|
||||
"turn N stays resolved by its own reply");
|
||||
assertTrue(rendezvous.isWaiting("term_a"), "turn N+1 is still awaiting its own resolution");
|
||||
}
|
||||
}
|
||||
@@ -6,8 +6,10 @@ import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.ExecutionException;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
@@ -25,12 +27,18 @@ class InjectorTest {
|
||||
private final FakeHerdr herdr = new FakeHerdr();
|
||||
private final Injector injector = new Injector(new AgentControl(herdr));
|
||||
|
||||
/** Text of every agent.send, in order. */
|
||||
/**
|
||||
* The logical messages delivered, in order. AgentControl.send emits each delivery as two
|
||||
* agent.send calls — the payload, then a standalone Enter keystroke ({@code "\r"}) to submit
|
||||
* it; these tests assert delivery ordering/gating, not the submit event, so drop the bare
|
||||
* carriage returns.
|
||||
*/
|
||||
@SuppressWarnings("unchecked")
|
||||
private List<String> sent() {
|
||||
return herdr.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.send"))
|
||||
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
|
||||
.filter(t -> !t.equals("\r"))
|
||||
.toList();
|
||||
}
|
||||
|
||||
@@ -43,6 +51,47 @@ class InjectorTest {
|
||||
assertEquals(List.of("hello"), sent());
|
||||
}
|
||||
|
||||
@Test
|
||||
void holdsDeliveryUntilTheWorkerIsAvailable() {
|
||||
// CB-113: idle alone is not enough — hold until the worker's MCP is connected (ready).
|
||||
java.util.Set<String> ready = new java.util.HashSet<>();
|
||||
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains);
|
||||
inj.enqueue(T, "task");
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE); // idle but not yet available → held out of the boot window
|
||||
assertEquals(List.of(), sent(), "must not deliver into a not-yet-available worker");
|
||||
|
||||
ready.add(T); // the worker's Claude connects the bridge MCP
|
||||
inj.onStatus(T, AgentStatus.IDLE);
|
||||
assertEquals(List.of("task"), sent(), "delivers once the worker is available");
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private long enterKeystrokes() {
|
||||
return herdr.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.send"))
|
||||
.filter(c -> "\r".equals(((Map<String, Object>) c.params()).get("text")))
|
||||
.count();
|
||||
}
|
||||
|
||||
@Test
|
||||
void resubmitsEnterWhenADeliveredMessageIsNotPickedUp() {
|
||||
// CB-113: the Enter at delivery can race the paste; while the worker stays idle (not picked
|
||||
// up), the injector re-nudges Enter so the pending paste submits.
|
||||
injector.enqueue(T, "task");
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver: paste + one Enter
|
||||
long afterDeliver = enterKeystrokes();
|
||||
|
||||
injector.onStatus(T, AgentStatus.IDLE); // still idle → re-nudge Enter
|
||||
injector.onStatus(T, AgentStatus.IDLE); // and again
|
||||
assertTrue(enterKeystrokes() > afterDeliver, "an unpicked-up delivery re-nudges Enter");
|
||||
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker finally starts
|
||||
long atPickup = enterKeystrokes();
|
||||
injector.onStatus(T, AgentStatus.WORKING);
|
||||
assertEquals(atPickup, enterKeystrokes(), "no more nudges once the worker has picked up");
|
||||
}
|
||||
|
||||
@Test
|
||||
void holdsWhileWorkingThenDeliversOnIdle() {
|
||||
injector.enqueue(T, "later");
|
||||
@@ -120,17 +169,97 @@ class InjectorTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void activeWhileQueuedOrInFlightThenQuietAfterPickup() {
|
||||
void activeWhileQueuedOrInFlightThenQuietAfterTurnCompletes() {
|
||||
assertTrue(injector.activeTargets().isEmpty());
|
||||
injector.enqueue(T, "x");
|
||||
assertEquals(java.util.Set.of(T), injector.activeTargets(), "active while a message is queued");
|
||||
assertEquals(Set.of(T), injector.activeTargets(), "active while a message is queued");
|
||||
|
||||
injector.onStatus(T, AgentStatus.IDLE); // delivers; still in-flight (awaiting pickup)
|
||||
assertEquals(java.util.Set.of(T), injector.activeTargets(),
|
||||
injector.onStatus(T, AgentStatus.IDLE); // delivers; awaiting pickup
|
||||
assertEquals(Set.of(T), injector.activeTargets(),
|
||||
"stays active so the poller can observe the worker pick the message up");
|
||||
|
||||
injector.onStatus(T, AgentStatus.WORKING); // pickup observed → in-flight cleared
|
||||
assertTrue(injector.activeTargets().isEmpty(), "quiet once queue is empty and pickup is seen");
|
||||
injector.onStatus(T, AgentStatus.WORKING); // pickup observed; now awaiting turn completion
|
||||
assertEquals(Set.of(T), injector.activeTargets(),
|
||||
"stays active after pickup so the working→idle completion boundary is observed");
|
||||
|
||||
injector.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete
|
||||
assertTrue(injector.activeTargets().isEmpty(), "quiet once the delegated turn has completed");
|
||||
}
|
||||
|
||||
@Test
|
||||
void firesTurnCompleteOnAConfirmedWorkingThenIdle() {
|
||||
List<String> completed = new ArrayList<>();
|
||||
Injector inj = new Injector(new AgentControl(herdr), completed::add);
|
||||
inj.enqueue(T, "task");
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
inj.onStatus(T, AgentStatus.WORKING); // pickup + turn running
|
||||
assertEquals(List.of(), completed, "no completion until the turn returns to idle");
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete
|
||||
assertEquals(List.of(T), completed, "a confirmed working→idle fires exactly one completion");
|
||||
}
|
||||
|
||||
@Test
|
||||
void doesNotSynthesizeCompletionFromAnUnconfirmedTurn() {
|
||||
List<String> completed = new ArrayList<>();
|
||||
Injector inj = new Injector(new AgentControl(herdr), completed::add);
|
||||
inj.enqueue(T, "task");
|
||||
|
||||
// Deliver, then only ever idle — a `working` sample is never seen. The pickup grace unwedges
|
||||
// the queue but must NOT invent a completion: without a sampled turn there is no trustworthy
|
||||
// "the worker finished the task" signal, so the send should fall through to its timeout.
|
||||
for (int i = 0; i < 15; i++) inj.onStatus(T, AgentStatus.IDLE);
|
||||
assertEquals(List.of(), completed, "no completion is synthesized from an unconfirmed turn");
|
||||
}
|
||||
|
||||
/** Captures both turn-lifecycle callbacks so the CB-109 stall path can be asserted. */
|
||||
private static final class Captor implements TurnListener {
|
||||
final List<String> completed = new ArrayList<>();
|
||||
final List<String> failed = new ArrayList<>();
|
||||
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
completed.add(target);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnFailed(String target) {
|
||||
failed.add(target);
|
||||
}
|
||||
}
|
||||
|
||||
// ~30s of unknown at the 250ms prod poll interval; enough onStatus samples to trip the stall.
|
||||
private static final int STALL_SAMPLES = 130;
|
||||
|
||||
@Test
|
||||
void failsAnOutstandingDelegationWhoseWorkerWedgesInUnknown() {
|
||||
Captor cap = new Captor();
|
||||
Injector inj = new Injector(new AgentControl(herdr), cap);
|
||||
inj.enqueue(T, "task");
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
inj.onStatus(T, AgentStatus.WORKING); // worker starts the turn
|
||||
for (int i = 0; i < STALL_SAMPLES; i++) inj.onStatus(T, AgentStatus.UNKNOWN); // then wedges
|
||||
|
||||
assertEquals(List.of(T), cap.failed, "a sustained unknown streak fails the outstanding send");
|
||||
assertEquals(List.of(), cap.completed, "a wedge is a failure, not a completion");
|
||||
assertTrue(inj.activeTargets().isEmpty(), "the wedged target is reclaimed, not polled forever");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aTransientUnknownGlitchNeitherFailsNorBlocksCompletion() {
|
||||
Captor cap = new Captor();
|
||||
Injector inj = new Injector(new AgentControl(herdr), cap);
|
||||
inj.enqueue(T, "task");
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
inj.onStatus(T, AgentStatus.WORKING); // confirmed turn
|
||||
for (int i = 0; i < 10; i++) inj.onStatus(T, AgentStatus.UNKNOWN); // brief glitch, well under grace
|
||||
inj.onStatus(T, AgentStatus.IDLE); // working → idle: the real completion
|
||||
|
||||
assertEquals(List.of(), cap.failed, "a short unknown blip must not fail the turn");
|
||||
assertEquals(List.of(T), cap.completed, "the streak reset, so the turn still completes");
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -151,6 +280,86 @@ class InjectorTest {
|
||||
assertTrue(f.isCompletedExceptionally(), "queued waiters unblock when the worker vanishes");
|
||||
}
|
||||
|
||||
@Test
|
||||
void dropFailsTheTurnOfADeliveredMessageWhenTheWorkerVanishes() {
|
||||
// CB-110: the message was delivered (no longer queued), so failing queued waiters alone would
|
||||
// leave its send hanging. A vanished worker must fail that in-flight turn too.
|
||||
Captor cap = new Captor();
|
||||
Injector inj = new Injector(new AgentControl(herdr), cap);
|
||||
inj.enqueue(T, "task");
|
||||
inj.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
inj.onStatus(T, AgentStatus.WORKING); // turn running
|
||||
|
||||
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
|
||||
assertEquals(List.of(T), cap.failed, "a worker that vanishes mid-turn fails its in-flight send");
|
||||
}
|
||||
|
||||
@Test
|
||||
void dropFailsADeliveredTurnThatVanishesBeforePickupIsConfirmed() {
|
||||
// Delivered but no WORKING sampled yet (awaitingCompletion=true, awaitingPickup still true,
|
||||
// turnObserved=false) — a distinct state the other two drop tests don't cover. (Gap surfaced
|
||||
// by an off-sub worker's review of CB-110, delegated through the bridge.)
|
||||
Captor cap = new Captor();
|
||||
Injector inj = new Injector(new AgentControl(herdr), cap);
|
||||
inj.enqueue(T, "task");
|
||||
inj.onStatus(T, AgentStatus.IDLE); // deliver; pickup never confirmed
|
||||
|
||||
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
|
||||
assertEquals(List.of(T), cap.failed, "a delivery that vanishes before pickup still fails its send");
|
||||
}
|
||||
|
||||
// ~60s of idle-but-not-ready at the 250ms prod poll interval; enough to trip the readiness grace.
|
||||
private static final int READINESS_SAMPLES = 245;
|
||||
|
||||
@Test
|
||||
void failsAQueuedMessageWhoseWorkerNeverBecomesReady() {
|
||||
// CB-114: herdr keeps reporting the worker idle, but its Claude never connects the bridge MCP,
|
||||
// so the readiness gate never opens. The message must not be held (and the target polled)
|
||||
// forever — after the grace it fails, the caller unblocks via the worker-failure path, the
|
||||
// target is reclaimed, and the never-set presence is cleared.
|
||||
Captor cap = new Captor();
|
||||
List<String> forgotten = new ArrayList<>();
|
||||
Injector inj = new Injector(new AgentControl(herdr), cap, _ -> false, forgotten::add);
|
||||
CompletableFuture<Void> f = inj.enqueue(T, "task");
|
||||
|
||||
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
assertEquals(List.of(), sent(), "a never-ready worker is never delivered to");
|
||||
assertTrue(f.isCompletedExceptionally(), "the caller's future fails instead of hanging forever");
|
||||
assertEquals(List.of(T), cap.failed, "the awaiting send resolves through the worker-failure path");
|
||||
assertEquals(List.of(), cap.completed, "a never-ready worker is a failure, not a completion");
|
||||
assertEquals(List.of(T), forgotten, "the never-ready worker's presence is cleared");
|
||||
assertTrue(inj.activeTargets().isEmpty(), "the target is reclaimed, not polled forever");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWorkerThatBecomesReadyWithinTheGraceIsDeliveredNormally() {
|
||||
// The readiness grace must not fail a worker that is merely slow to boot: once it becomes
|
||||
// available before the grace elapses, delivery proceeds as usual (the counter resets).
|
||||
Set<String> ready = new java.util.HashSet<>();
|
||||
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains, _ -> {
|
||||
});
|
||||
inj.enqueue(T, "task");
|
||||
|
||||
for (int i = 0; i < 100; i++) inj.onStatus(T, AgentStatus.IDLE); // still booting, well under grace
|
||||
assertEquals(List.of(), sent());
|
||||
|
||||
ready.add(T); // MCP connects before the grace elapses
|
||||
inj.onStatus(T, AgentStatus.IDLE);
|
||||
assertEquals(List.of("task"), sent(), "a worker that connects within the grace is delivered to");
|
||||
}
|
||||
|
||||
@Test
|
||||
void dropClearsWorkerPresence() {
|
||||
// CB-114 (finding #1): a vanished worker's readiness must be forgotten so a stale entry cannot
|
||||
// linger past the worker's life (WorkerPresence.forget had no caller before this).
|
||||
List<String> forgotten = new ArrayList<>();
|
||||
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, forgotten::add);
|
||||
inj.enqueue(T, "orphan");
|
||||
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
|
||||
assertEquals(List.of(T), forgotten, "drop clears the gone worker's presence");
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollerDeliversToAnIdleWorker() throws Exception {
|
||||
// End-to-end through the poller: idle worker → message delivered without manual onStatus.
|
||||
@@ -170,7 +379,9 @@ class InjectorTest {
|
||||
@SuppressWarnings("unchecked")
|
||||
Map<String, Object> p = (Map<String, Object>) c.params();
|
||||
return p.get("text").toString();
|
||||
}).toList());
|
||||
})
|
||||
.filter(t -> !t.equals("\r")) // drop the standalone submit keystroke
|
||||
.toList());
|
||||
}
|
||||
|
||||
@Test
|
||||
|
||||
@@ -0,0 +1,88 @@
|
||||
package dev.ltms.bridged.inject;
|
||||
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.AgentStatus;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/** Content-based refinement of an unreliable {@code UNKNOWN} status (CB-115). */
|
||||
class StatusRefinerTest {
|
||||
|
||||
// --- pure classification -------------------------------------------------
|
||||
|
||||
@Test
|
||||
void classifiesAnIdlePromptAsIdle() {
|
||||
String pane = """
|
||||
⏺ All done — the file compiles cleanly.
|
||||
|
||||
╭──────────────────────────────────────╮
|
||||
│ > │
|
||||
╰──────────────────────────────────────╯
|
||||
⏵⏵ auto mode on (shift+tab to cycle)
|
||||
""";
|
||||
assertEquals(AgentStatus.IDLE, StatusRefiner.classify(pane));
|
||||
}
|
||||
|
||||
@Test
|
||||
void classifiesABarePromptGlyphAsIdle() {
|
||||
assertEquals(AgentStatus.IDLE, StatusRefiner.classify("some output\n❯ "));
|
||||
}
|
||||
|
||||
@Test
|
||||
void classifiesActiveGenerationAsWorking() {
|
||||
String pane = """
|
||||
⏺ Working on it…
|
||||
✳ Thinking… (12s · esc to interrupt)
|
||||
""";
|
||||
assertEquals(AgentStatus.WORKING, StatusRefiner.classify(pane));
|
||||
}
|
||||
|
||||
@Test
|
||||
void anEscToInterruptScreenIsWorkingEvenWithAPromptBox() {
|
||||
// "esc to interrupt" wins over a prompt box: the turn is still generating.
|
||||
String pane = "│ > │\n esc to interrupt";
|
||||
assertEquals(AgentStatus.WORKING, StatusRefiner.classify(pane));
|
||||
}
|
||||
|
||||
@Test
|
||||
void anUnrecognizableScreenStaysUnknown() {
|
||||
assertEquals(AgentStatus.UNKNOWN, StatusRefiner.classify("garbled ansi noise with no prompt"));
|
||||
assertEquals(AgentStatus.UNKNOWN, StatusRefiner.classify(""));
|
||||
assertEquals(AgentStatus.UNKNOWN, StatusRefiner.classify(null));
|
||||
}
|
||||
|
||||
// --- refine() wiring -----------------------------------------------------
|
||||
|
||||
@Test
|
||||
void refinePassesNonUnknownStatusesThroughWithoutReading() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
StatusRefiner refiner = new StatusRefiner(new AgentControl(herdr));
|
||||
|
||||
assertEquals(AgentStatus.WORKING, refiner.refine("term_a", AgentStatus.WORKING));
|
||||
assertEquals(AgentStatus.IDLE, refiner.refine("term_a", AgentStatus.IDLE));
|
||||
|
||||
assertFalse(herdr.called("agent.read"),
|
||||
"a trusted status must not cost a pane read");
|
||||
}
|
||||
|
||||
@Test
|
||||
void refineUpgradesUnknownToIdleFromPaneContent() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ answer\n❯ ");
|
||||
StatusRefiner refiner = new StatusRefiner(new AgentControl(herdr));
|
||||
|
||||
assertEquals(AgentStatus.IDLE, refiner.refine("term_a", AgentStatus.UNKNOWN));
|
||||
assertTrue(herdr.called("agent.read"), "an UNKNOWN must trigger a pane read");
|
||||
}
|
||||
|
||||
@Test
|
||||
void refineLeavesUnknownWhenContentIsUnclassifiable() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("nothing recognizable here");
|
||||
StatusRefiner refiner = new StatusRefiner(new AgentControl(herdr));
|
||||
|
||||
assertEquals(AgentStatus.UNKNOWN, refiner.refine("term_a", AgentStatus.UNKNOWN));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,33 @@
|
||||
package dev.ltms.bridged.inject;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/** The CB-113 worker-availability registry. */
|
||||
class WorkerPresenceTest {
|
||||
|
||||
@Test
|
||||
void tracksPresenceAndForgets() {
|
||||
WorkerPresence p = new WorkerPresence();
|
||||
assertFalse(p.isPresent("term_a"), "unseen worker is not available");
|
||||
|
||||
p.markPresent("term_a");
|
||||
assertTrue(p.isPresent("term_a"), "a worker seen on the MCP is available");
|
||||
|
||||
p.forget("term_a");
|
||||
assertFalse(p.isPresent("term_a"), "a torn-down worker is no longer available");
|
||||
}
|
||||
|
||||
@Test
|
||||
void nullOrBlankMarkIsANoOp() {
|
||||
WorkerPresence p = new WorkerPresence();
|
||||
assertDoesNotThrow(() -> {
|
||||
p.markPresent(null);
|
||||
p.markPresent(" ");
|
||||
});
|
||||
assertFalse(p.isPresent(""), "blank/null contacts (the primary) are never present");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,288 @@
|
||||
package dev.ltms.bridged.mcp;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.inject.Injector;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.session.FakeWorktrees;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.session.WorkerSession;
|
||||
import dev.ltms.bridged.session.WorktreeRequest;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* Parity tests for the MCP tool adapters — they must produce the same outcomes as the REST routes,
|
||||
* since both drive the same {@link MessageService}/{@link Rendezvous}. The MCP wire protocol itself
|
||||
* is the SDK's concern; here we test the thin adapter logic directly.
|
||||
*/
|
||||
class BridgeMcpTest {
|
||||
|
||||
private final FakeHerdr herdr = new FakeHerdr();
|
||||
private final AgentControl agents = new AgentControl(herdr);
|
||||
private final Rendezvous rendezvous = new Rendezvous();
|
||||
private final MessageService messages = new MessageService(agents, new Injector(agents), rendezvous);
|
||||
|
||||
private static String textOf(McpSchema.CallToolResult r) {
|
||||
return ((McpSchema.TextContent) r.content().getFirst()).text();
|
||||
}
|
||||
|
||||
private static ClaudeCodeLauncher workerService(FakeHerdr h, String baseUrl, Set<String> allow) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"ltms-local", baseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||
new SubscriptionGuard(allow), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> "tok");
|
||||
}
|
||||
|
||||
private static SessionManager sessionManager(FakeHerdr h, String baseUrl, Set<String> allow) {
|
||||
return new SessionManager(workerService(h, baseUrl, allow));
|
||||
}
|
||||
|
||||
@Test
|
||||
void sendThenReplyRoundTrips() throws Exception {
|
||||
// bridge_send blocks; bridge_reply resolves it with the worker's structured answer.
|
||||
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
|
||||
() -> BridgeMcp.send(messages, "term_a", "review this", 4000L));
|
||||
|
||||
McpSchema.CallToolResult reply = BridgeMcp.reply(rendezvous, "term_a", "LGTM");
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (Boolean.TRUE.equals(reply.isError()) && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
reply = BridgeMcp.reply(rendezvous, "term_a", "LGTM");
|
||||
}
|
||||
assertEquals("delivered", textOf(reply));
|
||||
|
||||
McpSchema.CallToolResult res = send.get(6, TimeUnit.SECONDS);
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
assertEquals("LGTM", textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void asyncSendReturnsATicketThenPollReportsTheReply() throws Exception {
|
||||
// wait:false parity — a ticket is issued, resolved by a reply, and surfaced by bridge_poll.
|
||||
McpSchema.CallToolResult accepted = BridgeMcp.sendAsync(messages, "term_a", "do it");
|
||||
assertNotEquals(Boolean.TRUE, accepted.isError());
|
||||
String out = textOf(accepted);
|
||||
assertTrue(out.contains("ticket="), out);
|
||||
String ticket = out.substring(out.indexOf("ticket=") + "ticket=".length()).trim();
|
||||
|
||||
// Resolve the awaiting send once it has opened (retry past the async-open race).
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
McpSchema.CallToolResult reply = BridgeMcp.reply(rendezvous, "term_a", "async LGTM");
|
||||
while (Boolean.TRUE.equals(reply.isError()) && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
reply = BridgeMcp.reply(rendezvous, "term_a", "async LGTM");
|
||||
}
|
||||
assertEquals("delivered", textOf(reply));
|
||||
|
||||
// Poll until the async send completes and reports the reply.
|
||||
McpSchema.CallToolResult polled = BridgeMcp.poll(messages, ticket);
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (!textOf(polled).contains("async LGTM") && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
polled = BridgeMcp.poll(messages, ticket);
|
||||
}
|
||||
assertEquals("async LGTM", textOf(polled));
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollUnknownTicketIsAnError() {
|
||||
McpSchema.CallToolResult res = BridgeMcp.poll(messages, "task-999");
|
||||
assertTrue(res.isError());
|
||||
assertTrue(textOf(res).contains("unknown ticket"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void sendTimesOutWithAWorkingNote() {
|
||||
McpSchema.CallToolResult res = BridgeMcp.send(messages, "term_a", "hi", 120L);
|
||||
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
|
||||
assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void sendRejectsMissingArgs() {
|
||||
assertTrue(BridgeMcp.send(messages, null, "hi", null).isError());
|
||||
assertTrue(BridgeMcp.send(messages, "term_a", " ", null).isError());
|
||||
}
|
||||
|
||||
@Test
|
||||
void replyWithNoPendingSendIsAnError() {
|
||||
McpSchema.CallToolResult res = BridgeMcp.reply(rendezvous, "term_a", "orphan");
|
||||
assertTrue(res.isError());
|
||||
assertTrue(textOf(res).contains("no send is awaiting"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void askThenAnswerRoundTrips() throws Exception {
|
||||
// The primary delegates and blocks; wait until its waiter is open before the worker asks.
|
||||
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
|
||||
() -> BridgeMcp.send(messages, "term_a", "do X", 5000L));
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"), "the send must be waiting for the ask to surface to");
|
||||
|
||||
// The worker asks mid-turn; the call blocks for the primary's answer.
|
||||
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
|
||||
() -> BridgeMcp.ask(messages, "term_a", "which config?", 5000L));
|
||||
|
||||
// The primary's send unblocks with the question and a turnId to answer on.
|
||||
McpSchema.CallToolResult q = send.get(6, TimeUnit.SECONDS);
|
||||
assertNotEquals(Boolean.TRUE, q.isError());
|
||||
String qt = textOf(q);
|
||||
assertTrue(qt.contains("[question]"), qt);
|
||||
String afterMarker = qt.substring(qt.indexOf("turnId=\"") + "turnId=\"".length());
|
||||
String turnId = afterMarker.substring(0, afterMarker.indexOf('"'));
|
||||
|
||||
// The primary answers via bridge_send(turnId); this blocks again for the worker's reply.
|
||||
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
|
||||
() -> BridgeMcp.answer(messages, turnId, "config.yaml", 5000L));
|
||||
|
||||
// The worker's ask returns the answer — it resumes the same turn.
|
||||
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
|
||||
|
||||
// The resumed worker replies, resolving the answering send (retry past the reopen race).
|
||||
McpSchema.CallToolResult reply = BridgeMcp.reply(rendezvous, "term_a", "done");
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (Boolean.TRUE.equals(reply.isError()) && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
reply = BridgeMcp.reply(rendezvous, "term_a", "done");
|
||||
}
|
||||
assertEquals("delivered", textOf(reply));
|
||||
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
|
||||
}
|
||||
|
||||
@Test
|
||||
void askFromANonWorkerConnectionIsAnError() {
|
||||
McpSchema.CallToolResult res = BridgeMcp.ask(messages, null, "which config?", 500L);
|
||||
assertTrue(res.isError());
|
||||
assertTrue(textOf(res).contains("workers only"), textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void answerToAStaleTurnIsAnError() {
|
||||
McpSchema.CallToolResult res = BridgeMcp.answer(messages, "term_a#999", "too late", 500L);
|
||||
assertTrue(res.isError());
|
||||
assertTrue(textOf(res).contains("no longer open"), textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnReturnsTheNewWorkersSessionAndPane() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null);
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"sessionId\":\"term_new_1\""), out);
|
||||
assertTrue(out.contains("\"paneId\":\"w9:pW_1\""), out);
|
||||
assertTrue(out.contains("\"status\":\"spawning\""), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnRejectsAnOffAllowlistProfileWithoutTouchingHerdr() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res =
|
||||
BridgeMcp.spawn(sessionManager(h, "https://api.anthropic.com", Set.of("gx00.gw")), null);
|
||||
assertTrue(res.isError());
|
||||
assertTrue(textOf(res).contains("subscription boundary"));
|
||||
assertFalse(h.called("agent.start"), "the guard must block before any spawn");
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnRejectsAnUnknownProfileAsAnError() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res =
|
||||
BridgeMcp.spawn(sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), "nope");
|
||||
assertTrue(res.isError());
|
||||
assertTrue(textOf(res).contains("unknown worker profile"), textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnPassesTheRequestedCwdToTheWorker() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "/req/dir", null, null, null);
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
@SuppressWarnings("unchecked")
|
||||
Map<String, Object> start = (Map<String, Object>) h.lastCall("agent.start").params();
|
||||
assertEquals("/req/dir", start.get("cwd"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void profilesListsConfiguredProfilesAndDefault() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res = BridgeMcp.profiles(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("ltms-local"), out);
|
||||
assertTrue(out.contains("\"default\":\"ltms-local\""), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void listReportsTrackedWorkers() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
|
||||
new WorktreeRequest("cb-304", null));
|
||||
|
||||
McpSchema.CallToolResult res = BridgeMcp.listWorkers(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions);
|
||||
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"sessionId\":\"" + s.terminalId() + "\""), out);
|
||||
assertTrue(out.contains("\"paneId\":\"" + s.paneId() + "\""), out);
|
||||
assertTrue(out.contains("\"profile\":\"ltms-local\""), out);
|
||||
assertTrue(out.contains("\"state\":\"spawning\""), out);
|
||||
assertTrue(out.contains("\"worktree\":\"" + s.worktree() + "\""), out);
|
||||
assertTrue(out.contains("\"branch\":\"" + s.branch() + "\""), out);
|
||||
assertTrue(out.contains("\"owner\":\"term_primary\""), out);
|
||||
assertTrue(out.contains("\"liveStatus\":\"unknown\""), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void stopTearsDownAWorkerByPane() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res = BridgeMcp.stop(
|
||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), "w9:pW");
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
assertEquals("stopped w9:pW", textOf(res));
|
||||
assertTrue(h.called("pane.close"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void stopRequiresAPaneId() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
assertTrue(BridgeMcp.stop(
|
||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), " ").isError());
|
||||
}
|
||||
|
||||
@Test
|
||||
void statusReportsLiveAgentStatus() {
|
||||
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
|
||||
AgentControl blockedAgents = new AgentControl(blocked);
|
||||
McpSchema.CallToolResult res = BridgeMcp.status(
|
||||
new MessageService(blockedAgents, new Injector(blockedAgents), rendezvous), "term_a");
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
assertEquals("blocked", textOf(res));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,46 @@
|
||||
package dev.ltms.bridged.mcp;
|
||||
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.PaneLocator;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/** Connection → caller-identity resolution, with the OS peer-PID lookup faked. */
|
||||
class ConnectionIdentityTest {
|
||||
|
||||
private final FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
private ConnectionIdentity with(PeerPidLookup pids) {
|
||||
return new ConnectionIdentity(new PaneLocator(herdr), pids);
|
||||
}
|
||||
|
||||
@Test
|
||||
void resolvesWorkerFromLoopbackPeerPid() {
|
||||
assertEquals("term_a", with(_ -> FakeHerdr.WORKER_PID).callerTerminal("127.0.0.1", 55555));
|
||||
}
|
||||
|
||||
@Test
|
||||
void nullForOffHostCaller() {
|
||||
// A non-loopback peer can't be an on-host worker → treat as primary/unknown.
|
||||
assertNull(with(_ -> FakeHerdr.WORKER_PID).callerTerminal("10.0.0.9", 55555));
|
||||
}
|
||||
|
||||
@Test
|
||||
void nullWhenPidOwnsNoPane() {
|
||||
// e.g. the primary — its PID maps to no worker pane.
|
||||
assertNull(with(_ -> 999_999).callerTerminal("127.0.0.1", 55555));
|
||||
}
|
||||
|
||||
@Test
|
||||
void resolvesTheCallersPidAndCwd() {
|
||||
// CB-112: the primary maps to no pane, but its PID and cwd are still readable.
|
||||
ConnectionIdentity id = new ConnectionIdentity(
|
||||
new PaneLocator(herdr), _ -> 999_999, pid -> pid == 999_999 ? "/main/project" : null);
|
||||
ConnectionIdentity.Caller c = id.resolve("127.0.0.1", 55555);
|
||||
assertNull(c.terminal(), "the primary owns no worker pane");
|
||||
assertEquals(999_999, c.pid());
|
||||
assertEquals("/main/project", id.cwdForPid(c.pid()), "the primary's cwd is resolvable from its PID");
|
||||
assertNull(id.cwdForPid(-1), "no cwd for an unresolved PID");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,214 @@
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.AgentStatus;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.inject.CompletionResolver;
|
||||
import dev.ltms.bridged.inject.Injector;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* The message layer's resolution paths (CB-104 reply + CB-106 completion fallback). The turn is
|
||||
* driven deterministically by feeding {@code onStatus} rather than running a real poller.
|
||||
*/
|
||||
class MessageServiceTest {
|
||||
|
||||
private static final String T = "term_a";
|
||||
|
||||
private final FakeHerdr herdr = new FakeHerdr().readText("BUILD GREEN: 391 files");
|
||||
private final AgentControl agents = new AgentControl(herdr);
|
||||
private final Rendezvous rendezvous = new Rendezvous();
|
||||
private final CompletionResolver completion = new CompletionResolver(agents, rendezvous);
|
||||
private final Injector injector = new Injector(agents, completion);
|
||||
private final MessageService messages = new MessageService(agents, injector, rendezvous);
|
||||
|
||||
/** Run {@code send} on a background thread; the current thread drives the worker's turn. */
|
||||
private CompletableFuture<MessageService.Reply> sendAsync() {
|
||||
return CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 5000));
|
||||
}
|
||||
|
||||
private void awaitWaiting() throws InterruptedException {
|
||||
long deadline = System.currentTimeMillis() + 2000;
|
||||
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting(T), "send should have opened its rendezvous waiter");
|
||||
}
|
||||
|
||||
@Test
|
||||
void completionFallbackResolvesATurnThatNeverCalledBridgeReply() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitWaiting();
|
||||
|
||||
herdr.readText("$ prompt"); // pre-turn pane: no answer yet (baseline reference)
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver the task (baselines the pre-turn content)
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up and works
|
||||
herdr.readText("BUILD GREEN: 391 files"); // the worker's turn produced new output
|
||||
injector.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete, no bridge_reply
|
||||
|
||||
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, reply.outcome(),
|
||||
"an unreplied but finished turn resolves via the completion fallback");
|
||||
assertEquals("BUILD GREEN: 391 files", reply.text(), "the scraped transcript tail is returned");
|
||||
assertTrue(reply.completed(), "a scraped completion still counts as completed");
|
||||
}
|
||||
|
||||
@Test
|
||||
void explicitBridgeReplyResolvesAsReplied() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitWaiting();
|
||||
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker working
|
||||
assertTrue(rendezvous.resolve(T, "LGTM ship it"), "an explicit reply resolves the send");
|
||||
|
||||
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.REPLIED, reply.outcome());
|
||||
assertEquals("LGTM ship it", reply.text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWedgedWorkerResolvesTheSendAsFailedWithTheErrorContext() throws Exception {
|
||||
herdr.readText("API Error: Unable to connect to API (ENOTFOUND)");
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitWaiting();
|
||||
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker starts the turn
|
||||
for (int i = 0; i < 130; i++) injector.onStatus(T, AgentStatus.UNKNOWN); // then wedges (CB-109)
|
||||
|
||||
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome());
|
||||
assertFalse(reply.completed(), "a wedge is terminal but not a successful completion");
|
||||
assertTrue(reply.text().contains("ENOTFOUND"), "the error screen is carried as the failure reason");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWorkerThatVanishesMidTurnResolvesTheSendAsFailed() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitWaiting();
|
||||
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker starts the turn
|
||||
// The worker's pane crashes — the poller sees a *_not_found and drops it (CB-110).
|
||||
injector.drop(T, new HerdrException("worker gone", "pane_not_found", null));
|
||||
|
||||
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome(),
|
||||
"a delivered send whose worker vanishes fails instead of hanging to the timeout");
|
||||
assertFalse(reply.completed());
|
||||
}
|
||||
|
||||
// --- bridge_ask reverse rendezvous (CB-205) ------------------------------------------------
|
||||
|
||||
@Test
|
||||
void askSurfacesAsAQuestionAndTheAnswerResumesTheSameTurn() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up, then pauses to ask
|
||||
|
||||
// The worker asks mid-turn on its own thread; the call blocks for the primary's answer.
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||
|
||||
// The primary's blocking send unblocks with the question and a turnId to answer on.
|
||||
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
|
||||
assertEquals("which config file?", q.text());
|
||||
assertNotNull(q.turnId(), "a question carries a turnId to answer on");
|
||||
|
||||
// The primary answers via bridge_send(turnId); this blocks again for the worker's reply.
|
||||
CompletableFuture<MessageService.Reply> answer =
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
|
||||
|
||||
// The worker's ask returns the answer — it resumes the same turn.
|
||||
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome());
|
||||
assertEquals("config.yaml", a.answer());
|
||||
|
||||
// The resumed worker finishes with a structured reply, resolving the answering send.
|
||||
awaitWaiting(); // the answering send has (re)opened its forward waiter
|
||||
assertTrue(rendezvous.resolve(T, "done"), "the worker's final reply resolves the answering send");
|
||||
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
|
||||
assertEquals("done", done.text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void duplicateAsksFromTheSameSessionCoalesceToOneTurn() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up, then pauses to ask
|
||||
|
||||
// A transport retry: two concurrent bridge_ask calls from the same worker session.
|
||||
CompletableFuture<MessageService.AskResult> ask1 =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||
CompletableFuture<MessageService.AskResult> ask2 =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||
|
||||
// The primary's single blocked send surfaces exactly ONE question (one turnId).
|
||||
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
|
||||
assertEquals("which config file?", q.text());
|
||||
assertNotNull(q.turnId(), "only one turnId should be minted");
|
||||
|
||||
// The primary answers that one turnId; both asks unblock with the same answer.
|
||||
CompletableFuture<MessageService.Reply> answer =
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
|
||||
|
||||
MessageService.AskResult a1 = ask1.get(5, TimeUnit.SECONDS);
|
||||
MessageService.AskResult a2 = ask2.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.AskOutcome.ANSWERED, a1.outcome());
|
||||
assertEquals("config.yaml", a1.answer());
|
||||
assertEquals(MessageService.AskOutcome.ANSWERED, a2.outcome());
|
||||
assertEquals("config.yaml", a2.answer());
|
||||
|
||||
// The resumed worker finishes with a structured reply, resolving the answering send.
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "done"), "the worker's final reply resolves the answering send");
|
||||
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
|
||||
assertEquals("done", done.text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void askWithNoOpenDelegationReturnsNoWaiter() {
|
||||
MessageService.AskResult r = messages.ask(T, "anyone listening?", 500);
|
||||
assertEquals(MessageService.AskOutcome.NO_WAITER, r.outcome(),
|
||||
"a question with no blocked send has no primary to answer it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void askTimesOutWhenThePrimaryNeverAnswers() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
injector.onStatus(T, AgentStatus.WORKING);
|
||||
|
||||
MessageService.AskResult r = messages.ask(T, "still there?", 200); // primary never answers
|
||||
assertEquals(MessageService.AskOutcome.TIMED_OUT, r.outcome());
|
||||
|
||||
// The send itself already unblocked with the question the instant the ask surfaced.
|
||||
MessageService.Reply q = send.get(2, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
|
||||
}
|
||||
|
||||
@Test
|
||||
void answeringAnUnknownTurnIsStale() {
|
||||
MessageService.Reply r = messages.answer(T + "#999", "too late", 500);
|
||||
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
|
||||
"an answer to a turn that never existed (or already lapsed) is stale, not a hang");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,80 @@
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* The reverse rendezvous (CB-205): the {@code bridge_ask} registry that lets a worker pause mid-turn
|
||||
* to ask the primary. Unit-level — the message-layer round-trip is covered in {@link MessageServiceTest}.
|
||||
*/
|
||||
class RendezvousTest {
|
||||
|
||||
private static final String W = "term_a";
|
||||
|
||||
private final Rendezvous rendezvous = new Rendezvous();
|
||||
|
||||
@Test
|
||||
void openAskMintsAUniqueTurnScopedToItsSessionAndCoalescesDuplicates() {
|
||||
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
|
||||
Rendezvous.AskTicket t2 = rendezvous.openAsk(W);
|
||||
assertEquals(t1.turnId(), t2.turnId(), "duplicate asks from the same session coalesce onto one turn");
|
||||
assertTrue(t1.fresh(), "the first ask freshly opens the turn");
|
||||
assertFalse(t2.fresh(), "the coalesced ask rides the existing turn");
|
||||
assertTrue(t1.turnId().startsWith(W + "#"), "the turnId is scoped to the worker session");
|
||||
assertEquals(W, rendezvous.askSession(t1.turnId()));
|
||||
}
|
||||
|
||||
@Test
|
||||
void openAskAfterCloseMintsANewTurn() {
|
||||
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
|
||||
rendezvous.closeAsk(t1.turnId());
|
||||
Rendezvous.AskTicket t2 = rendezvous.openAsk(W);
|
||||
assertNotEquals(t1.turnId(), t2.turnId(), "after closing, a new ask gets a fresh turnId");
|
||||
assertTrue(t2.fresh(), "the reopened ask is fresh");
|
||||
assertEquals(W, rendezvous.askSession(t2.turnId()));
|
||||
}
|
||||
|
||||
@Test
|
||||
void answerAskCompletesTheWaitersFuture() {
|
||||
Rendezvous.AskTicket t = rendezvous.openAsk(W);
|
||||
assertTrue(rendezvous.answerAsk(t.turnId(), "config.yaml"), "answering an open ask succeeds");
|
||||
assertEquals("config.yaml", t.answer().getNow(null), "the answer reaches the blocked worker");
|
||||
}
|
||||
|
||||
@Test
|
||||
void answerAskOnAnUnknownTurnIsFalse() {
|
||||
assertFalse(rendezvous.answerAsk("no-such#1", "x"), "an answer to an unknown turn is a no-op");
|
||||
}
|
||||
|
||||
@Test
|
||||
void resolveQuestionResolvesAnOpenSendWithTheQuestionKindAndTurnId() {
|
||||
CompletableFuture<Rendezvous.Resolution> send = rendezvous.open(W);
|
||||
assertTrue(rendezvous.resolveQuestion(W, "which config?", W + "#7"),
|
||||
"the question resolves the primary's open send");
|
||||
Rendezvous.Resolution r = send.getNow(null);
|
||||
assertEquals(Rendezvous.Kind.QUESTION, r.kind());
|
||||
assertEquals("which config?", r.text());
|
||||
assertEquals(W + "#7", r.turnId(), "the turnId rides along so the primary can answer");
|
||||
}
|
||||
|
||||
@Test
|
||||
void resolveQuestionWithNoOpenSendIsFalse() {
|
||||
assertFalse(rendezvous.resolveQuestion(W, "anyone?", W + "#1"),
|
||||
"no blocked send means no primary to surface the question to");
|
||||
}
|
||||
|
||||
@Test
|
||||
void closeAskRemovesTheTurn() {
|
||||
Rendezvous.AskTicket t = rendezvous.openAsk(W);
|
||||
rendezvous.closeAsk(t.turnId());
|
||||
assertNull(rendezvous.askSession(t.turnId()), "a closed ask is forgotten");
|
||||
assertFalse(rendezvous.answerAsk(t.turnId(), "late"), "a closed ask can no longer be answered");
|
||||
}
|
||||
}
|
||||
@@ -7,7 +7,16 @@ import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.worker.WorkerService;
|
||||
import dev.ltms.bridged.inject.Injector;
|
||||
import dev.ltms.bridged.inject.StatusPoller;
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.session.FakeWorktrees;
|
||||
import dev.ltms.bridged.session.GitWorktrees;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.session.Worktrees;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
@@ -32,10 +41,13 @@ class BridgedAppTest {
|
||||
|
||||
private final ObjectMapper mapper = new ObjectMapper();
|
||||
private final HttpClient http = HttpClient.newHttpClient();
|
||||
private WorkerPresence presence;
|
||||
private Javalin app;
|
||||
private StatusPoller poller;
|
||||
|
||||
@AfterEach
|
||||
void stop() {
|
||||
if (poller != null) poller.stop();
|
||||
if (app != null) app.stop();
|
||||
}
|
||||
|
||||
@@ -44,13 +56,27 @@ class BridgedAppTest {
|
||||
}
|
||||
|
||||
private int start(FakeHerdr herdr, String workerBaseUrl, Set<String> allow, String placement) {
|
||||
return start(herdr, workerBaseUrl, allow, placement, new GitWorktrees());
|
||||
}
|
||||
|
||||
private int start(FakeHerdr herdr, String workerBaseUrl, Set<String> allow, String placement, Worktrees worktrees) {
|
||||
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
|
||||
"ltms-local", workerBaseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||
placement, "bridged-workers", "worker: {profile} #{n}");
|
||||
WorkerService workers = new WorkerService(
|
||||
new AgentControl(herdr), new WorkspaceControl(herdr), new SubscriptionGuard(allow), wcfg,
|
||||
placement, "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
|
||||
agents, new WorkspaceControl(herdr), new SubscriptionGuard(allow),
|
||||
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
|
||||
k -> "BRIDGED_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
|
||||
app = new BridgedApp(herdr, workers).build().start("127.0.0.1", 0);
|
||||
SessionManager sessions = new SessionManager(workers, worktrees);
|
||||
this.presence = sessions.asPresence();
|
||||
Injector injector = new Injector(agents);
|
||||
poller = new StatusPoller(agents, injector, 5); // delivers when the fake reports idle
|
||||
poller.start();
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
MessageService messages = new MessageService(agents, injector, rendezvous);
|
||||
app = new BridgedApp(herdr, workers, sessions, messages, rendezvous, this.presence, null)
|
||||
.build().start("127.0.0.1", 0);
|
||||
return app.port();
|
||||
}
|
||||
|
||||
@@ -58,6 +84,18 @@ class BridgedAppTest {
|
||||
return start(new FakeHerdr(), "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
}
|
||||
|
||||
private HttpResponse<String> postMessage(int port, String json) throws Exception {
|
||||
return postJson(port, "/sessions/term_a/message", json);
|
||||
}
|
||||
|
||||
private HttpResponse<String> postJson(int port, String path, String json) throws Exception {
|
||||
HttpRequest r = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path))
|
||||
.header("Content-Type", "application/json")
|
||||
.POST(HttpRequest.BodyPublishers.ofString(json)).build();
|
||||
return http.send(r, HttpResponse.BodyHandlers.ofString());
|
||||
}
|
||||
|
||||
|
||||
private HttpResponse<String> req(int port, String method, String path) throws Exception {
|
||||
HttpRequest.Builder b = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path));
|
||||
b = switch (method) {
|
||||
@@ -85,7 +123,8 @@ class BridgedAppTest {
|
||||
|
||||
@Test
|
||||
void healthzDegradedWhenHerdrDown() throws Exception {
|
||||
int port = start(new FakeHerdr().healthy(false), "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
FakeHerdr down = new FakeHerdr().healthy(false);
|
||||
int port = start(down, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
HttpResponse<String> res = req(port, "GET", "/healthz");
|
||||
assertEquals(503, res.statusCode());
|
||||
assertEquals("degraded", mapper.readTree(res.body()).get("status").asText());
|
||||
@@ -117,8 +156,8 @@ class BridgedAppTest {
|
||||
HttpResponse<String> res = req(port, "POST", "/workers");
|
||||
assertEquals(201, res.statusCode());
|
||||
JsonNode body = mapper.readTree(res.body());
|
||||
assertEquals("w9:pW", body.get("paneId").asText());
|
||||
assertEquals("w9:t2", body.get("tabId").asText());
|
||||
assertEquals("w9:pW_1", body.get("paneId").asText());
|
||||
assertEquals("spawning", body.get("state").asText());
|
||||
|
||||
// Subscription boundary: agent.start carried base_url + token in its env map.
|
||||
Map<String, Object> start = params(herdr, "agent.start");
|
||||
@@ -137,6 +176,58 @@ class BridgedAppTest {
|
||||
"tab label carries the worker number so siblings stay distinct");
|
||||
}
|
||||
|
||||
@Test
|
||||
void profilesEndpointListsConfiguredProfilesAndDefault() throws Exception {
|
||||
int port = startHealthy();
|
||||
JsonNode body = mapper.readTree(req(port, "GET", "/profiles").body());
|
||||
assertEquals("ltms-local", body.get("default").asText());
|
||||
assertEquals("ltms-local", body.get("profiles").get(0).asText());
|
||||
}
|
||||
|
||||
@Test
|
||||
void workersEndpointReturnsRegistryRosterWithLiveStatus() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"), "tab",
|
||||
new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt"));
|
||||
|
||||
HttpResponse<String> spawn = req(port, "POST", "/workers?worktree=true&ticket=cb-304");
|
||||
assertEquals(201, spawn.statusCode());
|
||||
JsonNode spawned = mapper.readTree(spawn.body());
|
||||
String paneId = spawned.get("paneId").asText();
|
||||
|
||||
HttpResponse<String> res = req(port, "GET", "/workers");
|
||||
assertEquals(200, res.statusCode());
|
||||
JsonNode workers = mapper.readTree(res.body()).get("workers");
|
||||
assertEquals(1, workers.size());
|
||||
JsonNode w = workers.get(0);
|
||||
assertEquals(spawned.get("terminalId").asText(), w.get("sessionId").asText());
|
||||
assertEquals(paneId, w.get("paneId").asText());
|
||||
assertEquals("ltms-local", w.get("profile").asText());
|
||||
assertEquals("spawning", w.get("state").asText());
|
||||
assertTrue(w.has("worktree"), "worktree-backed session exposes worktree");
|
||||
assertTrue(w.has("branch"), "worktree-backed session exposes branch");
|
||||
assertEquals("unknown", w.get("liveStatus").asText(),
|
||||
"liveStatus is unknown when herdr has no matching pane");
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnWithACwdParamRootsTheWorkerThere() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
assertEquals(201, req(port, "POST", "/workers?cwd=/tmp/proj").statusCode());
|
||||
assertEquals("/tmp/proj", params(herdr, "agent.start").get("cwd"), "the worker starts in cwd");
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnWithAnUnknownProfileIs400() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
HttpResponse<String> res = req(port, "POST", "/workers?profile=nope");
|
||||
assertEquals(400, res.statusCode());
|
||||
assertEquals("unknown_profile", mapper.readTree(res.body()).get("error").asText());
|
||||
assertFalse(herdr.called("agent.start"), "an unknown profile must not spawn anything");
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnWorkerReusesExistingWorkerSpace() throws Exception {
|
||||
// A space labelled "bridged-workers" already exists → no second workspace.create.
|
||||
@@ -211,6 +302,133 @@ class BridgedAppTest {
|
||||
assertEquals("w9:t2", params(herdr, "tab.close").get("tab_id"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void messageReturnsTheWorkersStructuredReply() throws Exception {
|
||||
// CB-104 (option C): the blocking send resolves on the worker's bridge_reply, not a scrape.
|
||||
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // poller delivers the injection
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
var send = java.util.concurrent.CompletableFuture.supplyAsync(() -> {
|
||||
try { return postMessage(port, "{\"content\":\"review this\",\"timeoutMs\":4000}"); }
|
||||
catch (Exception e) { throw new RuntimeException(e); }
|
||||
});
|
||||
|
||||
// The worker replies once a send is actually awaiting (retry past the startup race).
|
||||
HttpResponse<String> reply;
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
do {
|
||||
reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"LGTM ship it\"}");
|
||||
if (reply.statusCode() != 409) break;
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
} while (System.currentTimeMillis() < deadline);
|
||||
assertEquals(200, reply.statusCode());
|
||||
|
||||
HttpResponse<String> res = send.get(6, java.util.concurrent.TimeUnit.SECONDS);
|
||||
assertEquals(200, res.statusCode());
|
||||
assertEquals("LGTM ship it", mapper.readTree(res.body()).get("reply").asText());
|
||||
// (injection via agent.send is covered deterministically by the timeout-working test)
|
||||
}
|
||||
|
||||
@Test
|
||||
void asyncSendReturnsATicketThenPollReportsTheReply() throws Exception {
|
||||
// CB-107 fire-and-poll: wait:false returns a ticket immediately; the result is polled.
|
||||
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // poller delivers the injection
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
HttpResponse<String> accepted = postMessage(port, "{\"content\":\"do it\",\"wait\":false}");
|
||||
assertEquals(202, accepted.statusCode());
|
||||
String ticket = mapper.readTree(accepted.body()).get("ticket").asText();
|
||||
assertFalse(ticket.isBlank(), "an async send must return a ticket");
|
||||
|
||||
// The worker replies once the async send is actually awaiting (retry past the startup race).
|
||||
HttpResponse<String> reply;
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
do {
|
||||
reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"async LGTM\"}");
|
||||
if (reply.statusCode() != 409) break;
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
} while (System.currentTimeMillis() < deadline);
|
||||
assertEquals(200, reply.statusCode());
|
||||
|
||||
// Polling the ticket now reports the finished delegation and its reply.
|
||||
JsonNode task;
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
do {
|
||||
task = mapper.readTree(req(port, "GET", "/tasks/" + ticket).body());
|
||||
if ("done".equals(task.path("phase").asText())) break;
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
} while (System.currentTimeMillis() < deadline);
|
||||
assertEquals("done", task.get("phase").asText());
|
||||
assertEquals("async LGTM", task.get("reply").asText());
|
||||
assertEquals("reply", task.get("replySource").asText());
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollUnknownTicketIs404() throws Exception {
|
||||
int port = startHealthy();
|
||||
HttpResponse<String> res = req(port, "GET", "/tasks/task-999");
|
||||
assertEquals(404, res.statusCode());
|
||||
assertEquals("unknown_ticket", mapper.readTree(res.body()).get("error").asText());
|
||||
}
|
||||
|
||||
@Test
|
||||
void replyWithNoPendingSendIsConflict() throws Exception {
|
||||
int port = startHealthy();
|
||||
HttpResponse<String> res = postJson(port, "/sessions/term_a/reply", "{\"content\":\"orphan\"}");
|
||||
assertEquals(409, res.statusCode());
|
||||
assertEquals("no_pending_send", mapper.readTree(res.body()).get("error").asText());
|
||||
}
|
||||
|
||||
@Test
|
||||
void messageTimesOutQueuedWhenWorkerNeverInjectable() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr().agentStatus("working"); // never injectable → never delivered
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":150}");
|
||||
assertEquals(202, res.statusCode());
|
||||
assertEquals("queued", mapper.readTree(res.body()).get("status").asText());
|
||||
assertFalse(herdr.called("agent.send"), "no injection while the worker is mid-turn");
|
||||
}
|
||||
|
||||
@Test
|
||||
void messageTimesOutWorkingWhenDeliveredButNoReply() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // delivered, but nobody replies
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":250}");
|
||||
assertEquals(202, res.statusCode());
|
||||
assertEquals("working", mapper.readTree(res.body()).get("status").asText());
|
||||
assertTrue(herdr.called("agent.send"), "message was injected");
|
||||
}
|
||||
|
||||
@Test
|
||||
void messageRejectsBlankContent() throws Exception {
|
||||
int port = startHealthy();
|
||||
assertEquals(400, postMessage(port, "{}").statusCode());
|
||||
}
|
||||
|
||||
@Test
|
||||
void sessionStatusReportsLiveAgentStatus() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr().agentStatus("blocked");
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
HttpResponse<String> res = req(port, "GET", "/sessions/term_a/status");
|
||||
assertEquals(200, res.statusCode());
|
||||
assertEquals("blocked", mapper.readTree(res.body()).get("status").asText());
|
||||
}
|
||||
|
||||
@Test
|
||||
void sessionStatusReportsReadinessFromMcpPresence() throws Exception {
|
||||
int port = start(new FakeHerdr(), "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
// Not yet seen on the bridge MCP → not ready.
|
||||
assertFalse(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
|
||||
// Worker connects its MCP client → available.
|
||||
presence.markPresent("term_a");
|
||||
assertTrue(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
|
||||
}
|
||||
|
||||
@Test
|
||||
void stopWorkerInPanePlacementClosesOnlyThePane() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
@@ -0,0 +1,130 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
import java.util.Collections;
|
||||
import java.util.List;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.CopyOnWriteArrayList;
|
||||
|
||||
/** Recording fake {@link Worktrees} for CB-301-ext acceptance tests (no live git). */
|
||||
public final class FakeWorktrees implements Worktrees {
|
||||
|
||||
public record AddCall(String repoRoot, String branch, String baseRef) {
|
||||
}
|
||||
|
||||
public record RemoveCall(String repoRoot, String worktreePath) {
|
||||
}
|
||||
|
||||
public record OverlayCall(String repoRoot, String worktreePath,
|
||||
List<String> requested, List<String> copied, List<String> skipWorktree) {
|
||||
}
|
||||
|
||||
public record RepoRootCall(String cwd) {
|
||||
}
|
||||
|
||||
private final List<AddCall> addCalls = new CopyOnWriteArrayList<>();
|
||||
private final List<RemoveCall> removeCalls = new CopyOnWriteArrayList<>();
|
||||
private final List<OverlayCall> overlayCalls = new CopyOnWriteArrayList<>();
|
||||
private final List<RepoRootCall> repoRootCalls = new CopyOnWriteArrayList<>();
|
||||
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
|
||||
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
|
||||
private volatile RuntimeException addFailure;
|
||||
private volatile String repoRoot = "/repo";
|
||||
private volatile String prefix = "/worktrees";
|
||||
|
||||
public FakeWorktrees withRepoRoot(String root) {
|
||||
this.repoRoot = root;
|
||||
return this;
|
||||
}
|
||||
|
||||
public FakeWorktrees withPrefix(String prefix) {
|
||||
this.prefix = prefix;
|
||||
return this;
|
||||
}
|
||||
|
||||
/** Paths that exist in the primary repo and will be copied to the worktree. */
|
||||
public FakeWorktrees exists(String... paths) {
|
||||
Collections.addAll(existingPaths, paths);
|
||||
return this;
|
||||
}
|
||||
|
||||
/** Paths that exist AND are tracked, so overlayParity should --skip-worktree them. */
|
||||
public FakeWorktrees track(String... paths) {
|
||||
exists(paths);
|
||||
Collections.addAll(trackedPaths, paths);
|
||||
return this;
|
||||
}
|
||||
|
||||
/** Make subsequent {@link #add} calls throw (simulates git worktree add failure). */
|
||||
public FakeWorktrees failAdd(String message) {
|
||||
this.addFailure = new WorktreeException(message);
|
||||
return this;
|
||||
}
|
||||
|
||||
@Override
|
||||
public String add(String repoRoot, String branch, String baseRef) {
|
||||
addCalls.add(new AddCall(repoRoot, branch, baseRef));
|
||||
if (addFailure != null) {
|
||||
throw addFailure;
|
||||
}
|
||||
// The branch already carries a unique nonce, so the derived path is distinct per acquire
|
||||
// without an extra counter — keep it a pure function of the branch the test can predict.
|
||||
return prefix + "/" + branch.replace('/', '_');
|
||||
}
|
||||
|
||||
@Override
|
||||
public void remove(String repoRoot, String worktreePath) {
|
||||
removeCalls.add(new RemoveCall(repoRoot, worktreePath));
|
||||
}
|
||||
|
||||
@Override
|
||||
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
|
||||
List<String> copied = new java.util.ArrayList<>();
|
||||
List<String> skipped = new java.util.ArrayList<>();
|
||||
for (String rel : overlay) {
|
||||
if (!existingPaths.contains(rel)) {
|
||||
continue; // missing source is silently skipped
|
||||
}
|
||||
copied.add(rel);
|
||||
if (trackedPaths.contains(rel)) {
|
||||
skipped.add(rel);
|
||||
}
|
||||
}
|
||||
overlayCalls.add(new OverlayCall(repoRoot, worktreePath, List.copyOf(overlay),
|
||||
List.copyOf(copied), List.copyOf(skipped)));
|
||||
}
|
||||
|
||||
@Override
|
||||
public String repoRoot(String cwd) {
|
||||
repoRootCalls.add(new RepoRootCall(cwd));
|
||||
return repoRoot;
|
||||
}
|
||||
|
||||
public List<AddCall> addCalls() {
|
||||
return List.copyOf(addCalls);
|
||||
}
|
||||
|
||||
public List<RemoveCall> removeCalls() {
|
||||
return List.copyOf(removeCalls);
|
||||
}
|
||||
|
||||
public List<OverlayCall> overlayCalls() {
|
||||
return List.copyOf(overlayCalls);
|
||||
}
|
||||
|
||||
public List<RepoRootCall> repoRootCalls() {
|
||||
return List.copyOf(repoRootCalls);
|
||||
}
|
||||
|
||||
public AddCall lastAdd() {
|
||||
return addCalls.isEmpty() ? null : addCalls.getLast();
|
||||
}
|
||||
|
||||
public RemoveCall lastRemove() {
|
||||
return removeCalls.isEmpty() ? null : removeCalls.getLast();
|
||||
}
|
||||
|
||||
public OverlayCall lastOverlay() {
|
||||
return overlayCalls.isEmpty() ? null : overlayCalls.getLast();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,335 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* CB-301 / CB-303 acceptance tests for the authoritative session registry, one-shot lifecycle FSM,
|
||||
* and configurable lifecycle limits (idle TTL, context cap, drain).
|
||||
* No live herdr — everything runs against the same {@link FakeHerdr} the rest of the project uses.
|
||||
*/
|
||||
class SessionManagerTest {
|
||||
|
||||
private SessionManager sessionManager(FakeHerdr herdr) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
return new SessionManager(workers);
|
||||
}
|
||||
|
||||
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock) {
|
||||
return sessionManager(herdr, clock, 0);
|
||||
}
|
||||
|
||||
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
return new SessionManager(workers, new GitWorktrees(), clock, contextCap);
|
||||
}
|
||||
|
||||
@Test
|
||||
void acquireRegistersSpawningSessionWithDistinctPaneId() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
|
||||
WorkerSession a = sessions.acquire("ltms-local", "/work/a", "/caller/a", "term_primary");
|
||||
WorkerSession b = sessions.acquire("ltms-local", "/work/b", "/caller/b", "term_primary");
|
||||
|
||||
assertEquals(WorkerSession.State.SPAWNING, a.state(), "fresh session starts spawning");
|
||||
assertEquals("ltms-local", a.profile());
|
||||
assertEquals("/work/a", a.cwd(), "explicit requested cwd is recorded");
|
||||
assertEquals("term_primary", a.ownerTerminal());
|
||||
assertTrue(a.spawnedAtNanos() > 0);
|
||||
assertNotNull(a.paneId());
|
||||
assertNotNull(a.terminalId());
|
||||
|
||||
assertNotEquals(a.paneId(), b.paneId(), "no pane reuse");
|
||||
assertNotEquals(a.terminalId(), b.terminalId(), "no terminal reuse");
|
||||
assertEquals(2, sessions.roster().size(), "both sessions are registered");
|
||||
}
|
||||
|
||||
@Test
|
||||
void presenceMovesSpawningToReadyAndDeliveredTurnMovesToDone() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
assertEquals(WorkerSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"MCP presence moves SPAWNING → READY");
|
||||
assertTrue(sessions.asPresence().isPresent(terminal), "presence is also recorded");
|
||||
|
||||
sessions.onDelivered(terminal);
|
||||
assertEquals(WorkerSession.State.BUSY, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"delivery moves READY → BUSY");
|
||||
|
||||
sessions.onTurnComplete(terminal);
|
||||
assertEquals(WorkerSession.State.DONE, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"turn completion moves BUSY → DONE");
|
||||
}
|
||||
|
||||
@Test
|
||||
void releaseTearsDownWorkerAndRemovesFromRosterAndIsIdempotent() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", null);
|
||||
String paneId = session.paneId();
|
||||
|
||||
sessions.release(paneId);
|
||||
|
||||
assertTrue(herdr.called("pane.close"), "release tears the worker pane down");
|
||||
assertTrue(sessions.get(paneId).isEmpty(), "released session is no longer retrievable");
|
||||
assertTrue(sessions.roster().isEmpty(), "released session is no longer in the roster");
|
||||
|
||||
assertDoesNotThrow(() -> sessions.release(paneId), "a second release is harmless");
|
||||
}
|
||||
|
||||
@Test
|
||||
void onTurnFailedMovesSessionToFailed() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
sessions.onDelivered(terminal);
|
||||
|
||||
sessions.onTurnFailed(terminal);
|
||||
|
||||
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
|
||||
assertEquals(WorkerSession.State.FAILED, updated.state(), "turn failure moves to FAILED");
|
||||
assertTrue(sessions.roster().contains(updated), "FAILED is still in acquired-minus-released roster");
|
||||
}
|
||||
|
||||
@Test
|
||||
void recycleProducesNewPaneIdAndOldOneIsGone() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
WorkerSession oldSession = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String oldPane = oldSession.paneId();
|
||||
String oldTerminal = oldSession.terminalId();
|
||||
|
||||
WorkerSession fresh = sessions.recycle(oldPane);
|
||||
|
||||
assertNotEquals(oldPane, fresh.paneId(), "recycle yields a new pane id");
|
||||
assertNotEquals(oldTerminal, fresh.terminalId(), "recycle yields a new terminal id");
|
||||
assertEquals(oldSession.profile(), fresh.profile(), "profile is preserved");
|
||||
assertEquals(oldSession.cwd(), fresh.cwd(), "cwd is preserved");
|
||||
assertEquals(oldSession.ownerTerminal(), fresh.ownerTerminal(), "owner is preserved");
|
||||
|
||||
assertTrue(sessions.get(oldPane).isEmpty(), "old pane is deregistered");
|
||||
assertEquals(1, sessions.roster().size(), "only the fresh session remains");
|
||||
assertEquals(fresh.paneId(), sessions.roster().getFirst().paneId());
|
||||
|
||||
long paneCloseCount = herdr.calls.stream()
|
||||
.filter(c -> "pane.close".equals(c.method()))
|
||||
.filter(c -> oldPane.equals(((Map<?, ?>) c.params()).get("pane_id")))
|
||||
.count();
|
||||
assertEquals(1, paneCloseCount, "the old worker was torn down");
|
||||
}
|
||||
|
||||
@Test
|
||||
void rosterReflectsAcquiredMinusReleased() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
WorkerSession a = sessions.acquire("ltms-local", "/a", "/caller", "ownerA");
|
||||
WorkerSession b = sessions.acquire("ltms-local", "/b", "/caller", "ownerB");
|
||||
|
||||
assertEquals(2, sessions.roster().size());
|
||||
assertTrue(sessions.roster().stream().anyMatch(s -> s.paneId().equals(a.paneId())));
|
||||
assertTrue(sessions.roster().stream().anyMatch(s -> s.paneId().equals(b.paneId())));
|
||||
|
||||
sessions.release(a.paneId());
|
||||
|
||||
assertEquals(1, sessions.roster().size());
|
||||
assertEquals(b.paneId(), sessions.roster().getFirst().paneId());
|
||||
}
|
||||
|
||||
// --- CB-303 lifecycle limits ----------------------------------------------------
|
||||
|
||||
@Test
|
||||
void reapIdleDoesNothingWhenNoSessions() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> 0L);
|
||||
|
||||
assertEquals(0, sessions.reapIdle(10));
|
||||
assertTrue(sessions.roster().isEmpty());
|
||||
}
|
||||
|
||||
@Test
|
||||
void readySessionPastIdleTtlIsReaped() {
|
||||
long[] clock = {0};
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
|
||||
clock[0] = 11;
|
||||
assertEquals(1, sessions.reapIdle(10), "READY session past TTL is reaped");
|
||||
assertTrue(sessions.get(session.paneId()).isEmpty(), "reaped session is removed from registry");
|
||||
assertTrue(herdr.called("pane.close"), "reaped session tears the pane down");
|
||||
}
|
||||
|
||||
@Test
|
||||
void readySessionWithinIdleTtlSurvives() {
|
||||
long[] clock = {0};
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
|
||||
clock[0] = 5;
|
||||
assertEquals(0, sessions.reapIdle(10), "READY session within TTL is not reaped");
|
||||
assertEquals(WorkerSession.State.READY,
|
||||
sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"READY session survives");
|
||||
}
|
||||
|
||||
@Test
|
||||
void busySessionPastIdleTtlIsNotReaped() {
|
||||
long[] clock = {0};
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
sessions.onDelivered(terminal);
|
||||
|
||||
clock[0] = 100;
|
||||
assertEquals(0, sessions.reapIdle(10), "BUSY session past TTL is never reaped");
|
||||
assertEquals(WorkerSession.State.BUSY,
|
||||
sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"BUSY session remains");
|
||||
}
|
||||
|
||||
@Test
|
||||
void doneSessionPastIdleTtlIsReaped() {
|
||||
long[] clock = {0};
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
sessions.onDelivered(terminal);
|
||||
sessions.onTurnComplete(terminal);
|
||||
|
||||
clock[0] = 21;
|
||||
assertEquals(1, sessions.reapIdle(20), "DONE session past TTL is reaped");
|
||||
assertTrue(sessions.get(session.paneId()).isEmpty(), "DONE session is removed");
|
||||
}
|
||||
|
||||
@Test
|
||||
void reapIdleReturnsCorrectCountAndSkipsBusy() {
|
||||
long[] clock = {0};
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
|
||||
|
||||
WorkerSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "owner1");
|
||||
WorkerSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "owner2");
|
||||
sessions.asPresence().markPresent(ready.terminalId());
|
||||
sessions.asPresence().markPresent(busy.terminalId());
|
||||
sessions.onDelivered(busy.terminalId());
|
||||
|
||||
clock[0] = 50;
|
||||
assertEquals(1, sessions.reapIdle(30), "only READY past TTL is reaped");
|
||||
assertTrue(sessions.get(ready.paneId()).isEmpty(), "READY session is gone");
|
||||
assertEquals(WorkerSession.State.BUSY,
|
||||
sessions.get(busy.paneId()).orElseThrow().state(),
|
||||
"BUSY session is still registered");
|
||||
}
|
||||
|
||||
@Test
|
||||
void contextCapDisabledSessionSurvivesMultipleTurns() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> 0L, 0);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
|
||||
sessions.onDelivered(terminal);
|
||||
sessions.onTurnComplete(terminal);
|
||||
sessions.onDelivered(terminal);
|
||||
sessions.onTurnComplete(terminal);
|
||||
|
||||
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
|
||||
assertEquals(WorkerSession.State.DONE, updated.state(), "session finishes second turn");
|
||||
assertEquals(2, updated.turnCount(), "turn count tracks both deliveries");
|
||||
long releaseCloseCount = paneCloseCallsFor(herdr, session.paneId());
|
||||
assertEquals(0, releaseCloseCount, "cap disabled — no forced release of the worker pane");
|
||||
}
|
||||
|
||||
@Test
|
||||
void contextCapTwoReleasesAfterSecondComplete() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> 0L, 2);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
|
||||
sessions.onDelivered(terminal);
|
||||
sessions.onTurnComplete(terminal);
|
||||
assertEquals(WorkerSession.State.DONE,
|
||||
sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"first turn completes without release");
|
||||
|
||||
sessions.onDelivered(terminal);
|
||||
sessions.onTurnComplete(terminal);
|
||||
|
||||
assertTrue(sessions.get(session.paneId()).isEmpty(), "session released after cap reached");
|
||||
assertTrue(sessions.roster().isEmpty(), "released session leaves roster");
|
||||
assertEquals(1, paneCloseCallsFor(herdr, session.paneId()),
|
||||
"forced release tears the worker pane down exactly once");
|
||||
}
|
||||
|
||||
@Test
|
||||
void drainAllReleasesBusyAndReadySessionsAndWaitsForBusy() {
|
||||
long[] clock = {0};
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
|
||||
|
||||
WorkerSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "ownerR");
|
||||
WorkerSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "ownerB");
|
||||
sessions.asPresence().markPresent(ready.terminalId());
|
||||
sessions.asPresence().markPresent(busy.terminalId());
|
||||
sessions.onDelivered(busy.terminalId());
|
||||
|
||||
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
|
||||
|
||||
assertTrue(sessions.roster().isEmpty(), "drain clears the roster");
|
||||
assertTrue(sessions.get(ready.paneId()).isEmpty(), "ready session is released");
|
||||
assertTrue(sessions.get(busy.paneId()).isEmpty(), "busy session is released after timeout");
|
||||
assertEquals(1, paneCloseCallsFor(herdr, ready.paneId()),
|
||||
"ready worker pane is torn down");
|
||||
assertEquals(1, paneCloseCallsFor(herdr, busy.paneId()),
|
||||
"busy worker pane is torn down");
|
||||
}
|
||||
|
||||
private static long paneCloseCallsFor(FakeHerdr herdr, String paneId) {
|
||||
return herdr.calls.stream()
|
||||
.filter(c -> "pane.close".equals(c.method()))
|
||||
.filter(c -> paneId.equals(((Map<?, ?>) c.params()).get("pane_id")))
|
||||
.count();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,169 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* CB-301-ext acceptance tests for worktree provisioning and config-parity overlay.
|
||||
* No live git — every Worktrees call is handled by {@link FakeWorktrees} and every herdr
|
||||
* call by {@link FakeHerdr}, matching the project's fake-based test style.
|
||||
*/
|
||||
class WorktreeSessionManagerTest {
|
||||
|
||||
private static ClaudeCodeLauncher workerService(FakeHerdr herdr) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
}
|
||||
|
||||
private static String startCwd(FakeHerdr herdr) {
|
||||
@SuppressWarnings("unchecked")
|
||||
Map<String, Object> start = (Map<String, Object>) herdr.lastCall("agent.start").params();
|
||||
Object cwd = start.get("cwd");
|
||||
return cwd == null ? null : cwd.toString();
|
||||
}
|
||||
|
||||
@Test
|
||||
void sharedTreeAcquireMakesNoWorktreesCallsAndRecordsNullWorktree() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees();
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary");
|
||||
|
||||
assertTrue(worktrees.addCalls().isEmpty(), "shared-tree acquire never adds a worktree");
|
||||
assertTrue(worktrees.repoRootCalls().isEmpty(), "shared-tree acquire never resolves a repo root");
|
||||
assertTrue(worktrees.overlayCalls().isEmpty(), "shared-tree acquire never overlays parity");
|
||||
assertNull(s.worktree(), "shared-tree session has no worktree");
|
||||
assertNull(s.branch(), "shared-tree session has no branch");
|
||||
assertEquals("/caller/proj", s.cwd(), "shared-tree cwd is the caller's cwd");
|
||||
assertEquals("/caller/proj", startCwd(herdr), "spawn receives the caller's cwd");
|
||||
}
|
||||
|
||||
@Test
|
||||
void worktreeAcquireProvisionsAndRecordsPathAndBranch() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
|
||||
new WorktreeRequest("cb-999", null));
|
||||
|
||||
assertEquals(1, worktrees.addCalls().size(), "one worktree was added");
|
||||
FakeWorktrees.AddCall add = worktrees.lastAdd();
|
||||
assertNotNull(add);
|
||||
assertEquals("/repo", add.repoRoot());
|
||||
assertTrue(add.branch().startsWith("worker/cb-999-"), "branch is worker/<slug>-<nonce>: " + add.branch());
|
||||
assertNull(add.baseRef(), "null baseRef is passed through (HEAD default)");
|
||||
|
||||
String expectedPath = "/wt/" + add.branch().replace('/', '_');
|
||||
assertEquals(expectedPath, s.worktree(), "session records the returned worktree path");
|
||||
assertEquals(add.branch(), s.branch(), "session records the branch");
|
||||
assertEquals(expectedPath, startCwd(herdr), "spawn receives the worktree path as cwd");
|
||||
assertEquals(expectedPath, s.cwd(), "session cwd is the worktree path");
|
||||
}
|
||||
|
||||
@Test
|
||||
void worktreeAcquireRunsParityOverlayWithProfileDefaults() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
|
||||
.track(".mcp.json")
|
||||
.exists(".claude/settings.local.json");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
|
||||
sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("cb-888", null));
|
||||
|
||||
assertEquals(1, worktrees.overlayCalls().size());
|
||||
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
|
||||
assertNotNull(overlay);
|
||||
assertEquals("/repo", overlay.repoRoot());
|
||||
assertEquals(List.of(".mcp.json", ".claude/settings.local.json", ".env", ".envrc"),
|
||||
overlay.requested(), "default parity overlay is used when unset");
|
||||
assertEquals(List.of(".mcp.json", ".claude/settings.local.json"), overlay.copied(),
|
||||
"existing paths are copied; missing paths are skipped");
|
||||
assertEquals(List.of(".mcp.json"), overlay.skipWorktree(),
|
||||
"tracked copied paths are --skip-worktree'd");
|
||||
}
|
||||
|
||||
@Test
|
||||
void releaseRemovesWorktreeButDoesNotDeleteBranch() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("cb-666", null));
|
||||
String paneId = s.paneId();
|
||||
|
||||
sessions.release(paneId);
|
||||
|
||||
assertTrue(herdr.called("pane.close"), "release still tears the worker pane down");
|
||||
assertEquals(1, worktrees.removeCalls().size(), "worktree session triggers one remove");
|
||||
FakeWorktrees.RemoveCall remove = worktrees.lastRemove();
|
||||
assertNotNull(remove);
|
||||
assertEquals("/repo", remove.repoRoot());
|
||||
assertEquals(s.worktree(), remove.worktreePath());
|
||||
// The fake records no branch-delete calls because Worktrees.remove only removes the checkout.
|
||||
assertTrue(sessions.get(paneId).isEmpty(), "released session is no longer retrievable");
|
||||
}
|
||||
|
||||
@Test
|
||||
void sharedTreeReleaseMakesNoWorktreesCalls() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees();
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", null);
|
||||
|
||||
sessions.release(s.paneId());
|
||||
|
||||
assertTrue(herdr.called("pane.close"), "release tears the worker pane down");
|
||||
assertTrue(worktrees.removeCalls().isEmpty(), "shared-tree release never removes a worktree");
|
||||
}
|
||||
|
||||
@Test
|
||||
void failedWorktreeAddUnwindsWithoutRegisteringSessionOrSpawning() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().failAdd("worktree add failed");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
|
||||
assertThrows(WorktreeException.class, () ->
|
||||
sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("cb-555", null)));
|
||||
|
||||
assertEquals(0, sessions.size(), "failed acquire leaves no registry entry");
|
||||
assertFalse(herdr.called("agent.start"), "spawn is never reached when add fails");
|
||||
assertTrue(worktrees.removeCalls().isEmpty(), "no worktree was added, so none is removed");
|
||||
}
|
||||
|
||||
@Test
|
||||
void twoWorktreeAcquiresYieldDistinctBranchesAndPaths() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
|
||||
WorkerSession a = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("cb-444", null));
|
||||
WorkerSession b = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("cb-444", null));
|
||||
|
||||
assertNotEquals(a.branch(), b.branch(), "branches are distinct");
|
||||
assertNotEquals(a.worktree(), b.worktree(), "paths are distinct");
|
||||
assertEquals(2, worktrees.addCalls().size());
|
||||
assertEquals(2, sessions.roster().size());
|
||||
}
|
||||
|
||||
}
|
||||
@@ -0,0 +1,321 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.peer.Capability;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.SpawnRequest;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.function.Function;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/** The step-4 launch-flag injection: the bridge MCP + reply charter are appended to the argv. */
|
||||
class ClaudeCodeLauncherTest {
|
||||
|
||||
private ClaudeCodeLauncher service(FakeHerdr herdr, List<String> argv, String mcpUrl) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
argv, "tab", "bridged-workers", "worker: {profile} #{n}", mcpUrl, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private List<String> spawnedArgv(FakeHerdr herdr) {
|
||||
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("argv");
|
||||
}
|
||||
|
||||
@Test
|
||||
void appendsBridgeMcpAndReplyCharterWhenMcpUrlSet() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, List.of("ccs", "ltms-local"), "http://127.0.0.1:8765/mcp").spawn();
|
||||
|
||||
List<String> argv = spawnedArgv(herdr);
|
||||
assertEquals(List.of("ccs", "ltms-local"), argv.subList(0, 2), "base command preserved first");
|
||||
assertTrue(argv.contains("--mcp-config"));
|
||||
assertTrue(argv.stream().anyMatch(a -> a.contains("\"bridge\"") && a.contains("http://127.0.0.1:8765/mcp")),
|
||||
"inline bridge MCP config present");
|
||||
assertTrue(argv.contains("--append-system-prompt"));
|
||||
assertTrue(argv.stream().anyMatch(a -> a.contains("bridge_reply")), "reply charter present");
|
||||
}
|
||||
|
||||
@Test
|
||||
void noBridgeFlagsWhenMcpUrlAbsent() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, List.of("bash", "-c", "sleep 1"), null).spawn();
|
||||
assertEquals(List.of("bash", "-c", "sleep 1"), spawnedArgv(herdr), "argv untouched without mcpUrl");
|
||||
}
|
||||
|
||||
private ClaudeCodeLauncher multiProfile(FakeHerdr herdr) {
|
||||
BridgedConfig.Worker gx10 = new BridgedConfig.Worker("gx10", "http://gx10.gw:8000", "coder",
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "gx10"), "tab", "bridged-workers", "w #{n}", null, null, null);
|
||||
BridgedConfig.Worker ollama = new BridgedConfig.Worker("ollama", "http://ollama.ltms.dev", null,
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "ollama"), "tab", "bridged-workers", "w #{n}", null, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx10.gw", "ollama.ltms.dev")),
|
||||
Map.of("gx10", gx10, "ollama", ollama), "gx10", _ -> "tok");
|
||||
}
|
||||
|
||||
@Test
|
||||
@SuppressWarnings("unchecked")
|
||||
void spawnPicksTheNamedProfilesBaseUrlAndArgv() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
multiProfile(herdr).spawn("ollama");
|
||||
|
||||
Map<String, Object> start = (Map<String, Object>) herdr.lastCall("agent.start").params();
|
||||
Map<String, String> env = (Map<String, String>) start.get("env");
|
||||
assertEquals("http://ollama.ltms.dev", env.get("ANTHROPIC_BASE_URL"), "the named profile's base_url");
|
||||
assertEquals(List.of("ccs", "ollama"), start.get("argv"), "the named profile's launch command");
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnRejectsAnUnknownProfile() {
|
||||
try (FakeHerdr herdr = new FakeHerdr()) {
|
||||
assertThrows(IllegalArgumentException.class, () -> multiProfile(herdr).spawn("nope"));
|
||||
}
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static String startCwd(FakeHerdr herdr) {
|
||||
// The worker's cwd is set on agent.start (an agent pane does not inherit the tab's cwd).
|
||||
Object v = ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("cwd");
|
||||
return v == null ? null : v.toString();
|
||||
}
|
||||
|
||||
@Test
|
||||
void requestedCwdRootsTheWorker() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, List.of("ccs", "ltms-local"), null).spawn("ltms-local", "/work/proj", "/caller/home");
|
||||
assertEquals("/work/proj", startCwd(herdr), "an explicit spawn cwd wins over everything");
|
||||
}
|
||||
|
||||
@Test
|
||||
void profileConfigCwdBeatsTheCallerCwd() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker("ltms-local", "http://gx00.gw:8000", "coder",
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"w #{n}", null, "/pinned/dir", null);
|
||||
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("ltms-local", cfg), "ltms-local", _ -> null);
|
||||
svc.spawn("ltms-local", null, "/caller/home");
|
||||
assertEquals("/pinned/dir", startCwd(herdr), "a profile-pinned cwd overrides the caller's");
|
||||
}
|
||||
|
||||
@Test
|
||||
void inheritsTheCallerCwdWhenNothingElseIsSet() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, List.of("ccs", "ltms-local"), null).spawn("ltms-local", null, "/primary/project");
|
||||
assertEquals("/primary/project", startCwd(herdr), "no explicit/config cwd → inherit the primary's");
|
||||
}
|
||||
|
||||
// --- CB-302 git-forge token injection (worker checkpoint grant) ------------
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static Map<String, String> startEnv(FakeHerdr herdr) {
|
||||
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("env");
|
||||
}
|
||||
|
||||
@Test
|
||||
void injectsForgeTokenAndHostWhenProfileGrantsIt() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"impl", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "impl"), "tab", "bridged-workers", "w #{n}", null, null, null,
|
||||
"GITEA_ACCESS_TOKEN", null); // parityOverlay null; gitHostEnv null → defaults to GITEA_HOST
|
||||
Function<String, String> host = name -> switch (name) {
|
||||
case "GITEA_ACCESS_TOKEN" -> "gt-secret";
|
||||
case "GITEA_HOST" -> "git.ltms.dev";
|
||||
default -> null;
|
||||
};
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("impl", cfg), "impl", host).spawn();
|
||||
|
||||
Map<String, String> env = startEnv(herdr);
|
||||
assertEquals("gt-secret", env.get("GITEA_TOKEN"), "the forge token is injected for a granting profile");
|
||||
assertEquals("git.ltms.dev", env.get("GITEA_HOST"), "the paired forge host rides along with the token");
|
||||
}
|
||||
|
||||
@Test
|
||||
void noForgeTokenWhenProfileDoesNotGrantIt() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
// gitTokenEnv unset (12-arg ctor); the env would resolve a token if asked, proving the gate
|
||||
// is the profile config, not a missing env var.
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker("ltms-local", "http://gx00.gw:8000", "coder",
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("ccs"), "tab", "bridged-workers", "w #{n}",
|
||||
null, null, null);
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("ltms-local", cfg), "ltms-local",
|
||||
_ -> "would-be-secret").spawn();
|
||||
|
||||
Map<String, String> env = startEnv(herdr);
|
||||
assertNull(env.get("GITEA_TOKEN"), "no forge token when the profile does not opt in");
|
||||
assertNull(env.get("GITEA_HOST"), "no forge host without a granted token");
|
||||
}
|
||||
|
||||
// --- CB-117 orphan reap: the pure predicate --------------------------------
|
||||
|
||||
@Test
|
||||
void isForeignWorkerMatchesOurSchemeWithANonSelfNonce() {
|
||||
assertTrue(ClaudeCodeLauncher.isForeignWorker("claude-ollama-be09c2-2", "aaaaaa"),
|
||||
"a bridge worker name with a different nonce is a prior daemon's orphan");
|
||||
assertTrue(ClaudeCodeLauncher.isForeignWorker("claude-gx10-4127af-11", "aaaaaa"),
|
||||
"profile and multi-digit seq are still parsed; foreign nonce ⇒ reap");
|
||||
}
|
||||
|
||||
@Test
|
||||
void isForeignWorkerSparesOurOwnLiveWorkersAndNonWorkers() {
|
||||
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude-ollama-abcdef-3", "abcdef"),
|
||||
"a worker with THIS process's nonce is ours and live — never reap it");
|
||||
assertFalse(ClaudeCodeLauncher.isForeignWorker(null, "abcdef"), "an unnamed agent is not a worker");
|
||||
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude", "abcdef"), "a bare kind name is not a worker");
|
||||
assertFalse(ClaudeCodeLauncher.isForeignWorker("my-repl", "abcdef"), "a user's own label is not a worker");
|
||||
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude-ollama-XYZ123-2", "abcdef"),
|
||||
"a non-hex nonce does not match our scheme");
|
||||
}
|
||||
|
||||
// --- CB-117 orphan reap: the wiring through stop() -------------------------
|
||||
|
||||
private static long paneCloseCount(FakeHerdr herdr, String paneId) {
|
||||
return herdr.calls.stream()
|
||||
.filter(c -> c.method().equals("pane.close"))
|
||||
.filter(c -> paneId.equals(((Map<?, ?>) c.params()).get("pane_id")))
|
||||
.count();
|
||||
}
|
||||
|
||||
@Test
|
||||
void reapsAForeignOrphanButSparesOurOwnWorkerAndUserSessions() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher svc = multiProfile(herdr);
|
||||
herdr.withAgent("claude-ollama-be09c2-2", "term_orphan", "wQ:pF", "wQ:t8") // prior daemon's leak
|
||||
.withAgent("claude-gx10-" + svc.nameNonce() + "-1", "term_mine", "wQ:pMine", "wQ:tMine"); // ours, live
|
||||
// (the fake's default unnamed term_a stands in for a user's own Claude session)
|
||||
|
||||
int reaped = svc.reapOrphanWorkers();
|
||||
|
||||
assertEquals(1, reaped, "exactly the one foreign-nonce orphan is reaped");
|
||||
assertEquals(1, paneCloseCount(herdr, "wQ:pF"), "the orphan's pane is closed");
|
||||
assertEquals(0, paneCloseCount(herdr, "wQ:pMine"), "our own live worker's pane is left running");
|
||||
assertEquals(0, paneCloseCount(herdr, "w2:p7"), "a user's own session is never touched");
|
||||
assertTrue(herdr.called("tab.close"), "the orphan's now-empty dedicated tab is closed too");
|
||||
}
|
||||
|
||||
@Test
|
||||
void reapCountsAnAlreadyGoneOrphanAsReaped() {
|
||||
FakeHerdr herdr = new FakeHerdr().paneCloseFailsWith("pane_not_found");
|
||||
ClaudeCodeLauncher svc = multiProfile(herdr);
|
||||
herdr.withAgent("claude-ollama-0d856d-3", "term_gone", "wQ:pS", "wQ:tD");
|
||||
|
||||
assertEquals(1, svc.reapOrphanWorkers(),
|
||||
"a pane that vanished between list and close is a successful reap, not a failure");
|
||||
}
|
||||
|
||||
@Test
|
||||
void reapIsSkippedWhenHerdrCannotBeListed() {
|
||||
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.list throws
|
||||
assertEquals(0, multiProfile(herdr).reapOrphanWorkers(), "a listing failure reaps nothing and does not throw");
|
||||
}
|
||||
|
||||
// --- PeerHandle indirection ----------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void spawnReturnsPeerHandleWithIdEqualToPaneId() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
|
||||
|
||||
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
|
||||
|
||||
assertNotNull(handle, "spawn must return a non-null handle");
|
||||
assertEquals("w9:pW_1", handle.id(), "handle.id() must equal the agent's paneId");
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnReturnsPeerHandleWithCorrectTerminalId() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
|
||||
|
||||
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, "/caller"));
|
||||
|
||||
assertEquals("term_new_1", handle.terminalId(), "handle.terminalId() must equal the agent's terminalId");
|
||||
}
|
||||
|
||||
@Test
|
||||
void capabilitiesIncludeMidTurnAskWorktreeOrphanReap() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
|
||||
|
||||
Set<Capability> caps = svc.capabilities();
|
||||
|
||||
assertTrue(caps.contains(Capability.MID_TURN_ASK), "every Claude Code peer supports mid-turn ask");
|
||||
assertTrue(caps.contains(Capability.WORKTREE), "every CLI peer supports worktree cwd");
|
||||
assertTrue(caps.contains(Capability.ORPHAN_REAP), "every herdr launcher supports orphan reap");
|
||||
}
|
||||
|
||||
@Test
|
||||
void capabilitiesIncludeSelfPrWhenProfileHasGitToken() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
"impl", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "impl"), "tab", "bridged-workers", "w #{n}", null, null, null,
|
||||
"GITEA_ACCESS_TOKEN", null);
|
||||
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("impl", cfg), "impl",
|
||||
_ -> "tok");
|
||||
|
||||
assertTrue(svc.capabilities().contains(Capability.SELF_PR),
|
||||
"a profile with a git token grants SELF_PR");
|
||||
}
|
||||
|
||||
@Test
|
||||
void capabilitiesExcludeSelfPrWhenNoGitToken() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
|
||||
|
||||
assertFalse(svc.capabilities().contains(Capability.SELF_PR),
|
||||
"no git token profile → no SELF_PR capability");
|
||||
}
|
||||
|
||||
@Test
|
||||
void effectiveCwdViaSpawnRequestMatchesExistingResolution() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
|
||||
|
||||
String cwd = svc.effectiveCwd(new SpawnRequest("ltms-local", "/work/proj", "/caller/home"));
|
||||
|
||||
assertEquals("/work/proj", cwd, "effectiveCwd via SpawnRequest must match the three-arg resolution");
|
||||
}
|
||||
|
||||
@Test
|
||||
void profilesViaPeerLauncherMatchesExistingApi() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher svc = multiProfile(herdr);
|
||||
|
||||
assertEquals(Set.of("gx10", "ollama"), svc.profiles(), "profiles() via PeerLauncher must match");
|
||||
}
|
||||
|
||||
@Test
|
||||
void defaultProfileViaPeerLauncherMatches() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher svc = multiProfile(herdr);
|
||||
|
||||
assertEquals("gx10", svc.defaultProfile(), "defaultProfile() via PeerLauncher must match");
|
||||
}
|
||||
|
||||
@Test
|
||||
void stopViaPeerLauncherTearsDownByHandleId() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
|
||||
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
|
||||
|
||||
svc.stop(handle.id());
|
||||
|
||||
assertTrue(herdr.called("pane.close"), "stop via handle.id() must close the pane");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,131 @@
|
||||
# CB-301 — Session Manager (one-shot, no reuse)
|
||||
|
||||
**Status:** design spec for review → delegate implementation.
|
||||
**Grounded in:** `WorkerService`, `Injector`/`StatusPoller`/`TurnListener`, `MessageService`,
|
||||
`BridgeMcp`, `BridgedApp` (see [wiki 9. Implementation](../wiki/9-Implementation.md)).
|
||||
|
||||
## Problem
|
||||
|
||||
`WorkerService` is **stateless about what it spawned**. Its own Javadoc says it plainly:
|
||||
|
||||
> "there is no registry; `list()` only asks herdr." — `WorkerService.reapOrphanWorkers` (line 288)
|
||||
|
||||
Consequences today:
|
||||
- The daemon cannot answer "which workers did *I* spawn, in what lifecycle state, owned by whom,
|
||||
since when?" without shelling to herdr for a raw agent list (no state, no ownership, no age).
|
||||
- Cleanup of a worker that outlived its owning process depends entirely on the boot-time
|
||||
name-nonce **reaper** (CB-117) — there is no live, authoritative roster during a run.
|
||||
- `bridge_list` (CB-304) can only surface herdr's view, not a bridge-owned roster.
|
||||
- There is no seam for per-session policy (checkpoint on teardown → CB-302; idle_ttl /
|
||||
context_cap / drain → CB-303).
|
||||
|
||||
## Goal & non-goals
|
||||
|
||||
**Goal.** Introduce a `SessionManager` that owns an authoritative in-daemon registry of the worker
|
||||
sessions this daemon process spawned, tracks each one's lifecycle state, and tears each down
|
||||
deterministically. It becomes the single source of truth for the roster and the seam CB-302/303/304
|
||||
build on.
|
||||
|
||||
**Non-goals (explicit — reuse policy chosen: one-shot, no reuse).**
|
||||
- **No pooling / no reuse.** Every delegated task gets a fresh worker; a finished worker is torn
|
||||
down, never handed to a later task. No "warm idle" pool, no `role@profile` keying.
|
||||
- **No auto-teardown *timing*.** *When* a one-shot worker is released (immediately on turn
|
||||
completion vs after an idle grace) is CB-303. CB-301 provides the **mechanism** (`release`) and
|
||||
the registry; CB-303 sets the policy.
|
||||
- **No checkpoint content.** Writing `STATE.md` + commit on teardown is CB-302; CB-301 only exposes
|
||||
the release hook it will attach to.
|
||||
|
||||
"Recycle" under no-reuse is simply **release + fresh acquire** — a helper, not a pool operation.
|
||||
|
||||
## Design
|
||||
|
||||
`SessionManager` **wraps** `WorkerService` (does not replace it). `WorkerService` keeps doing the
|
||||
subscription-guarded spawn/teardown mechanics; `SessionManager` adds the registry, lifecycle, and
|
||||
ownership on top.
|
||||
|
||||
**Package:** new `dev.ltms.bridged.session` — keeps the registry/lifecycle concern separate from
|
||||
the `worker` spawn mechanics. Holds `SessionManager` + `WorkerSession`.
|
||||
|
||||
**`recycle` is IN SCOPE for CB-301** (decided): implement `recycle(paneId, …)` = `release` the old
|
||||
session then `acquire` a fresh one, asserting a new distinct paneId (the no-reuse invariant). It is
|
||||
a thin convenience over the two primitives, shipped now so the no-reuse teardown+respawn path is
|
||||
covered by a test from day one.
|
||||
|
||||
### `WorkerSession` (record or small mutable holder)
|
||||
|
||||
| Field | Source | Notes |
|
||||
|---|---|---|
|
||||
| `paneId` | `Agent.paneId()` | registry key |
|
||||
| `terminalId` | `Agent.terminalId()` | for status/identity joins |
|
||||
| `profile` | spawn arg | which profile spawned it |
|
||||
| `cwd` | resolved cwd | the worker's working dir |
|
||||
| `ownerTerminal` | caller identity (nullable) | the primary/turn that requested it; `null` = daemon/anon |
|
||||
| `spawnedAtNanos` | `System.nanoTime()` | age basis for CB-303 (monotonic; no wall clock in tests) |
|
||||
| `state` | lifecycle FSM | see below |
|
||||
|
||||
State is held in a `ConcurrentHashMap<String /*paneId*/, WorkerSession>`.
|
||||
|
||||
### Lifecycle state machine (one-shot)
|
||||
|
||||
```
|
||||
SPAWNING --ready(MCP present)--> READY
|
||||
READY --onDelivered--> BUSY
|
||||
BUSY --onTurnComplete--> DONE
|
||||
BUSY --onTurnFailed--> FAILED
|
||||
READY|DONE|FAILED --release()--> RELEASED (deregistered)
|
||||
SPAWNING|READY|BUSY|DONE --vanished/drop--> FAILED
|
||||
```
|
||||
|
||||
- Transitions are driven by hooks the manager already has access to:
|
||||
`WorkerPresence.markPresent` → `READY`; `TurnListener.onDelivered/onTurnComplete/onTurnFailed`
|
||||
(the manager implements or decorates `TurnListener`) → `BUSY`/`DONE`/`FAILED`.
|
||||
- `RELEASED` sessions are removed from the registry (teardown is terminal).
|
||||
- Any state → `FAILED` on drop (worker vanished / injector `drop`), mirroring `Injector`.
|
||||
|
||||
### API
|
||||
|
||||
```java
|
||||
final class SessionManager {
|
||||
WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal);
|
||||
void release(String paneId); // deterministic teardown + deregister
|
||||
WorkerSession recycle(String paneId, ...); // release + acquire (no-reuse convenience)
|
||||
Optional<WorkerSession> get(String paneId);
|
||||
List<WorkerSession> roster(); // bridge-owned view (CB-304 consumes this)
|
||||
// lifecycle hooks (package-private): onReady/onDelivered/onComplete/onFailed(target)
|
||||
}
|
||||
```
|
||||
|
||||
- `acquire` = `workerService.spawn(profile, requestedCwd, callerCwd)` → register `SPAWNING`.
|
||||
- `release` = `workerService.stop(paneId)` → deregister. Idempotent (already-gone tolerated, matching
|
||||
`WorkerService.stop`).
|
||||
- `roster` joins the registry with live herdr status for a truthful "roster + live" (CB-304).
|
||||
|
||||
### Integration points
|
||||
|
||||
- **`Bridged.main`** — construct `SessionManager(workerService, ...)`; wire it as/decorating the
|
||||
`TurnListener` alongside `CompletionResolver` so it sees turn boundaries, and give it the
|
||||
`WorkerPresence` signal for `READY`.
|
||||
- **`BridgeMcp.spawn` / `BridgedApp.spawnWorker`** — route spawn through `SessionManager.acquire`
|
||||
(carry `callerTerminal` as `ownerTerminal`). **`bridge_stop` / `DELETE /workers/{paneId}`** →
|
||||
`SessionManager.release`.
|
||||
- **`bridge_list` / `GET /sessions` (CB-304 later)** — read `SessionManager.roster()`.
|
||||
- **`MessageService`** — no change required for one-shot; a later CB-303 auto-release hook can call
|
||||
`release` from `onTurnComplete` under policy.
|
||||
|
||||
## Acceptance (tests, no live herdr — fakes as elsewhere)
|
||||
|
||||
1. `acquire` registers a `SPAWNING` session with the right owner/profile/cwd; a second `acquire`
|
||||
yields a **distinct** paneId and a **distinct** session (no reuse).
|
||||
2. Presence signal moves `SPAWNING → READY`; a delivered turn moves `READY → BUSY → DONE`.
|
||||
3. `release` tears the worker down via `WorkerService.stop` and removes it from `roster()`;
|
||||
a second `release` on the same paneId is a harmless no-op.
|
||||
4. `onTurnFailed` / drop moves the session to `FAILED` and it is absent from the live roster.
|
||||
5. `recycle` produces a new paneId and the old one is gone (no-reuse invariant).
|
||||
6. `roster()` reflects exactly the sessions acquired-minus-released, joined with live status.
|
||||
|
||||
## Seams left open (deliberately)
|
||||
|
||||
- **CB-302** — attach a checkpoint step (`STATE.md` + commit) to the `release` path.
|
||||
- **CB-303** — a policy loop over `roster()` using `spawnedAtNanos`/state to auto-`release` on
|
||||
`idle_ttl`, or drain on `context_cap`.
|
||||
- **CB-304** — `bridge_list` reads `roster()` for a bridge-owned roster + live join.
|
||||
@@ -0,0 +1,179 @@
|
||||
# CB-301-ext — Worktree provisioning + config-parity overlay
|
||||
|
||||
**Status:** design spec for review → delegate implementation.
|
||||
**Extends:** [CB-301 Session Manager](CB-301-Session-Manager.md) (shipped, commit `54d907c`).
|
||||
**Realizes:** the config-parity requirement in [Worker Git Workflow](Worker-Git-Workflow.md).
|
||||
**Grounded in:** `SessionManager`, `WorkerService.spawn/effectiveCwd`, `BridgedConfig.Worker`,
|
||||
`inject/…LsofPeerPidLookup` (the `ProcessBuilder` exec pattern).
|
||||
|
||||
## Problem
|
||||
|
||||
CB-301 gives each worker a session record but every worker still runs in the **primary's own
|
||||
working tree** (`cwd = callerCwd`). One worker at a time is safe; two **parallel implementers** would
|
||||
stomp each other. We need each implementer session to get an **isolated git worktree on its own
|
||||
branch** — *without* degrading the worker: a bare worktree checks out **tracked files only**, so it
|
||||
silently drops the untracked/local config (`.claude/settings.local.json`, the locally-modified
|
||||
`.mcp.json`, `.env`) that makes a session a full peer of the primary. **CB-301-ext provisions the
|
||||
worktree AND hydrates it to config parity**, so a worker differs from the primary only in the LLM
|
||||
provider.
|
||||
|
||||
## Decisions (locked)
|
||||
|
||||
1. **Opt-in, not default.** A worktree is provisioned **only** when the caller requests one. Absent a
|
||||
request, `acquire` behaves exactly as it does today (shared primary tree) — auditors, smoke tests,
|
||||
and conversational workers are unaffected. **Backward compatibility is a hard requirement.**
|
||||
2. **Copy-overlay + `--skip-worktree`, not symlink.** Each parity file is **copied** primary→worktree
|
||||
(isolation-friendly, no symlink type-change noise on tracked files). For a *tracked* overlay file
|
||||
(`.mcp.json`) the worktree copy is then marked `git update-index --skip-worktree`, so the worker's
|
||||
commits can **never** include the parity overlay. Ignored files (`settings.local.json`) stay
|
||||
ignored in the worktree (shared `info/exclude`), so no marking is needed.
|
||||
3. **Branch persists; worktree is disposable.** `release` runs `git worktree remove --force` (the
|
||||
working dir is throwaway) but **never deletes the branch** — the branch holds the worker's commits
|
||||
and its PR (CB-302). Teardown of the checkout ≠ teardown of the work.
|
||||
4. **Git behind a seam.** SessionManager depends on a `Worktrees` interface (production impl shells
|
||||
`git` via `ProcessBuilder`; tests use a fake). No live `git` in unit tests — mirrors the
|
||||
`WorkerService`/`FakeHerdr` seam.
|
||||
|
||||
## Design
|
||||
|
||||
### `WorktreeRequest` (new, nullable = "no worktree")
|
||||
|
||||
```java
|
||||
package dev.ltms.bridged.session;
|
||||
/** Ask acquire() to provision an isolated worktree. null ⇒ run in the shared primary tree. */
|
||||
public record WorktreeRequest(String ticketSlug, String baseRef) {
|
||||
// ticketSlug seeds the branch name; baseRef null/blank ⇒ current HEAD of the repo.
|
||||
}
|
||||
```
|
||||
|
||||
### `WorkerSession` — two nullable fields added
|
||||
|
||||
| Field | Notes |
|
||||
|---|---|
|
||||
| `worktree` | absolute path of the provisioned worktree; `null` ⇒ shared tree |
|
||||
| `branch` | the worker's branch (`worker/<slug>-<nonce>`); `null` ⇒ shared tree |
|
||||
|
||||
Add to the record + `withState`. A `null` worktree keeps every existing test and the shared-tree path
|
||||
untouched.
|
||||
|
||||
### `Worktrees` seam (new)
|
||||
|
||||
```java
|
||||
package dev.ltms.bridged.session;
|
||||
public interface Worktrees {
|
||||
/** git -C <repoRoot> worktree add <path> -b <branch> <baseRef|HEAD>. Returns the worktree path. */
|
||||
String add(String repoRoot, String branch, String baseRef);
|
||||
/** git -C <repoRoot> worktree remove --force <path>. Idempotent (already-gone tolerated). */
|
||||
void remove(String repoRoot, String worktreePath);
|
||||
/** Copy each existing overlay path repoRoot→worktree; mark tracked ones --skip-worktree. */
|
||||
void overlayParity(String repoRoot, String worktreePath, List<String> overlay);
|
||||
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
|
||||
String repoRoot(String cwd);
|
||||
}
|
||||
```
|
||||
|
||||
- **Production impl** `GitWorktrees implements Worktrees` — `ProcessBuilder` per path, `redirectErrorStream(true)`, non-zero exit → a `WorktreeException`. Worktree location = `<worktreeRoot>/<nonce>` where `worktreeRoot` is a daemon setting (default: sibling `../.bridged-worktrees` of the repo root — **outside** the repo, never nested).
|
||||
- `overlayParity` per file: skip if absent in `repoRoot`; else copy into the worktree; if `git -C <wt> ls-files --error-unmatch <path>` succeeds (tracked), run `git -C <wt> update-index --skip-worktree <path>`.
|
||||
|
||||
### `acquire` — extended, old signature preserved
|
||||
|
||||
```java
|
||||
// existing (unchanged): shared tree
|
||||
WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal);
|
||||
// new overload: provision a worktree when wt != null
|
||||
WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal,
|
||||
WorktreeRequest wt);
|
||||
```
|
||||
|
||||
When `wt != null`:
|
||||
1. `repoRoot = worktrees.repoRoot(firstNonBlank(requestedCwd, callerCwd))`.
|
||||
2. `branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce`.
|
||||
3. `path = worktrees.add(repoRoot, branch, wt.baseRef())`.
|
||||
4. `worktrees.overlayParity(repoRoot, path, cfg.parityOverlay())`.
|
||||
5. `spawn(profile, path, callerCwd)` — **the worktree path becomes the worker's cwd** (highest
|
||||
precedence in `WorkerService.resolveCwd`).
|
||||
6. register the session with `worktree=path, branch=branch`.
|
||||
7. **On any failure in 1–5, unwind**: if the worktree was added, `remove` it; do not leave a dangling
|
||||
registry entry. (Guard/spawn already throw before herdr on a bad base_url — unchanged.)
|
||||
|
||||
### `release` — remove the worktree, keep the branch
|
||||
|
||||
```java
|
||||
public void release(String paneId) {
|
||||
WorkerSession s = registry.remove(paneId);
|
||||
workerService.stop(paneId); // existing
|
||||
if (s != null && s.worktree() != null) {
|
||||
worktrees.remove(worktrees.repoRoot(s.cwd()), s.worktree()); // branch is NOT deleted
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Config — `BridgedConfig.Worker.parityOverlay` + a `worktreeRoot`
|
||||
|
||||
- Add `List<String> parityOverlay` to the `Worker` record (12th field). Compact-constructor default
|
||||
when null/empty: `[".mcp.json", ".claude/settings.local.json", ".env", ".envrc"]` (missing paths are
|
||||
silently skipped, so the default is safe across repos). Update `withProfile`.
|
||||
- Add a top-level daemon setting `worktreeRoot` (String, nullable → `<repoParent>/.bridged-worktrees`).
|
||||
- `@JsonIgnoreProperties(ignoreUnknown = true)` already set → additive, no parser breakage.
|
||||
|
||||
### Surface: MCP + REST
|
||||
|
||||
- `bridge_spawn` gains an optional `worktree` arg: `true`, or a ticket slug string. Truthy ⇒ build a
|
||||
`WorktreeRequest(slug, null)` and call the 5-arg `acquire`.
|
||||
- `POST /workers` gains `worktree` (+ optional `ticket`) in the body/query, same mapping.
|
||||
- `workerView`/`view(WorkerSession)` include `worktree` and `branch` **when non-null** (omit for
|
||||
shared-tree sessions, so existing response assertions for shared-tree spawns are unchanged).
|
||||
|
||||
## Flow
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
A["acquire(profile, ..., WorktreeRequest?)"] --> B{"worktree<br/>requested?"}
|
||||
B -->|"no (default)"| C["spawn(cwd = callerCwd)<br/>— shared tree, unchanged"]
|
||||
B -->|yes| D["repoRoot = rev-parse --show-toplevel"]
|
||||
D --> E["git worktree add path -b branch base"]
|
||||
E --> F["overlayParity: copy local config in;<br/>--skip-worktree the tracked ones"]
|
||||
F --> G["spawn(cwd = worktree path)"]
|
||||
G --> H["register worktree + branch on the session"]
|
||||
E -.->|"add/overlay/spawn fails"| X["unwind: remove worktree,<br/>no dangling registry entry"]:::warn
|
||||
C --> R["worker is a full peer of the primary"]:::goal
|
||||
H --> R
|
||||
classDef warn fill:#b7791f,stroke:#7b341e,color:#ffffff;
|
||||
classDef goal fill:#2b6cb0,stroke:#2a4365,color:#ffffff;
|
||||
```
|
||||
|
||||
*Green = the invariant: whether shared-tree (parity for free) or worktree (parity via overlay), the
|
||||
worker matches the primary. Amber = the failure-unwind path.*
|
||||
|
||||
## Acceptance (fake `Worktrees`, no live git)
|
||||
|
||||
1. **Backward-compat:** `acquire` with **no** `WorktreeRequest` makes **zero** `Worktrees` calls,
|
||||
spawns with `cwd = callerCwd`, and records `worktree == null` / `branch == null`. (Every CB-301
|
||||
test still passes.)
|
||||
2. **Provision:** `acquire(..., new WorktreeRequest("cb-999", null))` calls `add(repoRoot,
|
||||
"worker/cb-999-<nonce>", null)`, then `spawn` receives the returned worktree path as
|
||||
`requestedCwd`; the session records that path + branch.
|
||||
3. **Overlay:** `overlayParity` is invoked with the profile's `parityOverlay` (default list when
|
||||
unset); the fake asserts tracked paths were `--skip-worktree`'d and missing paths skipped.
|
||||
4. **Release removes worktree, keeps branch:** releasing a worktree session calls
|
||||
`Worktrees.remove(repoRoot, path)` and performs **no** branch-delete; a shared-tree session's
|
||||
release makes no `Worktrees` calls.
|
||||
5. **Failure unwind:** a fake `add` that throws ⇒ `acquire` throws, the session is **not** registered,
|
||||
and no worker is left running (spawn not reached / torn down).
|
||||
6. **Distinct worktrees:** two worktree acquires yield **distinct** branches and paths (no collision).
|
||||
|
||||
## Constraints & exclusions (standing, non-negotiable)
|
||||
|
||||
- **Only a worker sets `ANTHROPIC_BASE_URL`.** Worktrees touch cwd + files only; env path is unchanged
|
||||
— the guard still runs before any herdr call.
|
||||
- **Worker cwd stays inside the primary repo** (the worktree is a checkout of it) — never `$HOME`.
|
||||
- **`.mcp.json` and `wiki/` never enter a worker commit.** `.mcp.json` is overlaid for *reference* but
|
||||
`--skip-worktree`'d so it can't be staged; `wiki/` is a submodule the worker must not touch. The
|
||||
CB-302 commit step (and the implementer skill) exclude both.
|
||||
- **The overlay list stays explicit + minimal** (trust: local secrets flow to an off-subscription
|
||||
worker). No blanket tree copy.
|
||||
|
||||
## Seams left for later
|
||||
|
||||
- **CB-302** — the worker commit → push → PR checkpoint runs *inside* the worktree on its branch.
|
||||
- **CB-304** — `roster()` rows surface `worktree`/`branch` for the fleet view.
|
||||
@@ -0,0 +1,200 @@
|
||||
# CB-401 — Peer Launcher SPI (Stage 4: pluggable peers)
|
||||
|
||||
**Status:** design note (feature branch `feature/peer-launcher-spi`)
|
||||
**Stage:** 4 — makes the bridge scalable to heterogeneous peers (Claude Code, Codex, …) without
|
||||
the core learning any one peer's environment.
|
||||
|
||||
## 1. Why
|
||||
|
||||
`claude-bridge` is a **communication bus between heterogeneous AI agents** — its stable surface is
|
||||
the protocol (`bridge_spawn / send / poll / reply / ask / list / stop / status`), and that surface
|
||||
should stay provider-neutral. Today the daemon can only materialize one kind of peer: an
|
||||
off-subscription Claude Code CLI over herdr. Everything specific to *how that peer is set up*
|
||||
(`ANTHROPIC_BASE_URL`, the subscription guard, `--mcp-config`/system-prompt flags, `claude-*`
|
||||
naming, git-token injection) is baked directly into the core spawn path.
|
||||
|
||||
The goal of CB-401 is a seam — a **`PeerLauncher` SPI** — so that "how to bring a peer of kind X to
|
||||
life" lives in a swappable adapter the bus *delegates to*, while the bus itself owns only transport,
|
||||
session/turn lifecycle, and routing. This both unlocks a second peer kind (Codex, a human, another
|
||||
Claude) and retroactively gives the CB-301-ext / CB-302 environment features a principled home (an
|
||||
adapter) instead of sitting in core.
|
||||
|
||||
> **Non-goal for CB-401:** dynamic/external plugin loading (arbitrary jars). That is Stage C and
|
||||
> carries a security model of its own (§7). CB-401 delivers a *first-party, in-tree, config-selected*
|
||||
> SPI with exactly one implementation, proving the seam.
|
||||
|
||||
## 2. The boundary
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph core["BRIDGE CORE — provider-neutral"]
|
||||
proto["Protocol verbs<br/>spawn / send / poll / reply / ask / list / stop"]
|
||||
sm["SessionManager<br/>FSM · registry · roster · lifecycle"]
|
||||
msg["Message store / routing"]
|
||||
life["Lifecycle limits<br/>idle_ttl · context_cap · drain"]
|
||||
end
|
||||
subgraph adapters["PEER ADAPTERS — env-specific"]
|
||||
cc["ClaudeCodeLauncher<br/>(today's WorkerService)"]
|
||||
cx["CodexLauncher<br/>(future, CB-402)"]
|
||||
hu["HumanLauncher<br/>(future)"]
|
||||
end
|
||||
sm -->|"delegates spawn/release"| spi{{"PeerLauncher SPI"}}
|
||||
spi --> cc
|
||||
spi --> cx
|
||||
spi --> hu
|
||||
cc -.->|"herdr transport"| herdr["herdr (terminal multiplexer)"]
|
||||
```
|
||||
|
||||
*Figure 1 — the core delegates peer materialization to a launcher chosen by profile; the core never
|
||||
learns a peer's env.*
|
||||
|
||||
## 3. Coupling audit (as-built, main @ `0efb65c`)
|
||||
|
||||
Where Claude/herdr specifics actually live today:
|
||||
|
||||
| Concern | Location | Verdict |
|
||||
|---|---|---|
|
||||
| `ANTHROPIC_BASE_URL` / `ANTHROPIC_MODEL` / `CLAUDE_CONFIG_DIR` / `ANTHROPIC_AUTH_TOKEN` env | `WorkerService.spawn` | **→ adapter** |
|
||||
| `SubscriptionGuard.assertWorker(baseUrl)` (billing boundary) | `WorkerService.spawn` → `guard` | **→ adapter** (it guards an `ANTHROPIC_*` concept) |
|
||||
| `--mcp-config` + `--append-system-prompt REPLY_CHARTER` (Claude Code CLI flags) | `WorkerService.argvWithBridge` | **→ adapter** |
|
||||
| `claude-<profile>-<nonce>-<seq>` naming, `WORKER_NAME` regex, orphan reap (CB-117) | `WorkerService` | **→ adapter** (naming is a herdr-label detail) |
|
||||
| `GITEA_TOKEN` / `GITEA_HOST` injection (CB-302 checkpoint) | `WorkerService.spawn` | **→ adapter** + a **capability** (§6) |
|
||||
| tab/pane placement, worker space, tab labels | `WorkerService.spawnInTab/spawnAsPane` via herdr `WorkspaceControl` | **→ adapter** (herdr transport detail) |
|
||||
| `BridgedConfig.Worker` profile shape (`baseUrl`, `model`, `configDir`, …) | `config` | **mostly adapter-shaped** — see §5 |
|
||||
| FSM, registry, roster, `reapIdle`/`drainAll`/`contextCap`, `rosterView` | `SessionManager` | **stays core** |
|
||||
| turn/completion detection (`TurnListener`, `CompletionResolver`, `StatusPoller`, `WorkerPresence`) | `inject/` | **stays core**, but reads herdr terminal output → transport-coupled (§4b) |
|
||||
| message store & routing | `msg/` | **stays core** |
|
||||
| `Agent` / `paneId` handle | `herdr/` | **generalize** — see §4a |
|
||||
|
||||
**Finding:** the extraction is tractable because ~90% of the coupling is already funnelled through
|
||||
one class (`WorkerService`). Renaming/adapting it to `ClaudeCodeLauncher implements PeerLauncher` and
|
||||
having `SessionManager` depend on the interface is the bulk of Stage A.
|
||||
|
||||
## 4. Two friction points
|
||||
|
||||
### 4a. `paneId` is a herdr handle, not a peer-neutral id
|
||||
|
||||
`SessionManager` keys its registry by `paneId`, MCP/REST route by `paneId`, and `WorkerSession`
|
||||
stores it. `paneId` is a herdr pane handle — meaningless for a peer that isn't a herdr pane.
|
||||
|
||||
**Decision:** introduce an opaque `PeerHandle` the launcher returns. It carries a launcher-assigned
|
||||
**`id`** (the registry/routing key) plus launcher-private coordinates (for herdr: paneId, tabId,
|
||||
terminalId). Stage A keeps `id == paneId` for the Claude adapter so nothing downstream changes value,
|
||||
but the *type* stops being "a herdr pane" — the core routes on `PeerHandle.id()`.
|
||||
|
||||
```mermaid
|
||||
classDiagram
|
||||
class PeerLauncher {
|
||||
<<interface>>
|
||||
+Set~Capability~ capabilities()
|
||||
+PeerHandle spawn(SpawnRequest req)
|
||||
+void release(PeerHandle h)
|
||||
+String effectiveCwd(SpawnRequest req)
|
||||
+List~String~ parityOverlay(String profile)
|
||||
+int reapOrphans()
|
||||
}
|
||||
class PeerHandle {
|
||||
<<interface>>
|
||||
+String id()
|
||||
}
|
||||
class ClaudeCodeLauncher {
|
||||
herdr AgentControl/WorkspaceControl
|
||||
SubscriptionGuard
|
||||
}
|
||||
PeerLauncher <|.. ClaudeCodeLauncher
|
||||
ClaudeCodeLauncher ..> PeerHandle : returns
|
||||
```
|
||||
|
||||
*Figure 2 — the SPI the core sees. `ClaudeCodeLauncher` is today's `WorkerService`, adapted.*
|
||||
|
||||
### 4b. Turn/completion detection reads herdr output
|
||||
|
||||
`inject/` (turn listener, completion resolver, status poller, presence) infers turn boundaries from
|
||||
herdr terminal scraping. That is genuinely peer-transport-specific — a Codex peer would signal turns
|
||||
differently. For CB-401 this stays in core (it's the *Claude/herdr* transport's detector), but §6's
|
||||
capability model is what lets a future non-herdr peer bring its own turn-signalling without the core
|
||||
assuming terminal scraping. **Out of scope for Stage A**; noted so the SPI doesn't accidentally
|
||||
hard-wire "turns come from herdr".
|
||||
|
||||
## 5. Config shape
|
||||
|
||||
`BridgedConfig.Worker` is Claude-shaped (`baseUrl`, `model`, `configDir`, `tokenEnv`). Rather than
|
||||
break existing YAML, CB-401 keeps `workers:` exactly as-is and treats those fields as the
|
||||
**ClaudeCodeLauncher's** profile schema. A future peer kind adds a `kind:` discriminator
|
||||
(default `"claude-code"`) selecting the launcher; unknown-kind → clear config error. No migration of
|
||||
existing configs. (Jackson already ignores unknown keys, so adding `kind` is backward-safe.)
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant MCP as bridge_spawn (MCP/REST)
|
||||
participant SM as SessionManager
|
||||
participant L as PeerLauncher (by profile.kind)
|
||||
participant T as transport (herdr)
|
||||
MCP->>SM: acquire(profile, cwd, owner)
|
||||
SM->>L: spawn(SpawnRequest)
|
||||
L->>L: build env + guard + argv (adapter-private)
|
||||
L->>T: start(name, argv, env, cwd)
|
||||
T-->>L: handle (paneId…)
|
||||
L-->>SM: PeerHandle(id)
|
||||
SM->>SM: register session keyed by handle.id()
|
||||
SM-->>MCP: session view
|
||||
```
|
||||
|
||||
*Figure 3 — spawn delegation. The core's `acquire` is unchanged in shape; only the thing it calls
|
||||
becomes an interface.*
|
||||
|
||||
## 6. Capabilities
|
||||
|
||||
Peers are not uniform. Let each launcher declare a capability set; the protocol is the union and
|
||||
degrades gracefully when a launcher lacks one:
|
||||
|
||||
| Capability | Meaning | Claude Code | Codex (likely) | Human |
|
||||
|---|---|---|---|---|
|
||||
| `MID_TURN_ASK` | supports `bridge_ask` rendezvous | ✓ | ? | ✗ |
|
||||
| `SELF_PR` | can open its own PR at checkpoint (CB-302) | ✓ (opt-in token) | ? | ✗ |
|
||||
| `WORKTREE` | can run in a provisioned git worktree | ✓ | ✓ | ✗ |
|
||||
| `ORPHAN_REAP` | spawner can reconcile orphaned peers on boot | ✓ | ? | ✗ |
|
||||
|
||||
A verb invoked against a peer that lacks the capability returns a clean "unsupported for this peer"
|
||||
rather than a crash. This keeps the protocol honest as peers diversify and prevents the core from
|
||||
assuming "every peer is a Claude in a worktree" (the drift signal from the identity note).
|
||||
|
||||
## 7. Staging & the Stage-C security gate
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
A["Stage A — CB-401<br/>extract PeerLauncher SPI in-tree<br/>ClaudeCodeLauncher = adapted WorkerService<br/>one impl, config-selected"] --> B["Stage B — CB-402+<br/>2nd in-tree adapter (Codex/human)<br/>proves the SPI held"]
|
||||
B --> C["Stage C<br/>dynamic external plugin loading<br/>ServiceLoader / jar discovery"]
|
||||
C -.requires.-> G["Trust & capability model<br/>what env/tokens a plugin may inject"]
|
||||
classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff;
|
||||
class G gate
|
||||
```
|
||||
|
||||
*Figure 4 — deliver A now; B when a real second peer exists; C only if third parties must ship
|
||||
adapters, and only behind a trust model.*
|
||||
|
||||
**Security note (Stage C, not now):** a launcher runs at daemon privilege and touches process
|
||||
spawning **and env/token injection into peers** — the most sensitive path in the system. "Extra
|
||||
plugins" must mean *first-party, in-tree, config-selected* for the foreseeable future. A third-party
|
||||
plugin that can inject env into a peer needs a genuine trust/capability model before it may exist.
|
||||
Regardless of stage, the **primary remains the merge/verify gate** — self-reports over the bus are
|
||||
messages, not verified facts.
|
||||
|
||||
## 8. Stage A scope (this branch, delegation-ready)
|
||||
|
||||
Deliverable for CB-401 Stage A — mechanical, behaviour-preserving:
|
||||
|
||||
1. `PeerLauncher` interface + `PeerHandle` (opaque id) + `SpawnRequest` (profile, requestedCwd,
|
||||
callerCwd) + `Capability` enum, new package `dev.ltms.bridged.peer`.
|
||||
2. `ClaudeCodeLauncher implements PeerLauncher` = today's `WorkerService`, adapted: `spawn(...)`
|
||||
returns a `PeerHandle` (id = paneId), `capabilities()` declares
|
||||
`MID_TURN_ASK, SELF_PR(when token), WORKTREE, ORPHAN_REAP`.
|
||||
3. `SessionManager` depends on `PeerLauncher`, not `WorkerService` concretely; routing keys on
|
||||
`PeerHandle.id()` (== paneId today, so zero value change).
|
||||
4. `Bridged.main` wires the concrete `ClaudeCodeLauncher` behind the interface.
|
||||
5. **No behaviour change, no config change.** Full green gate: `ide_sync` → `ide_diagnostics`
|
||||
(0 errors/0 warnings) → `mvn clean install` with `MVN_EXIT` captured (no masking pipe). All
|
||||
existing tests pass unchanged; add tests only for the new `PeerHandle` indirection.
|
||||
|
||||
Explicitly **out of scope** for Stage A: `kind:` config discriminator, any second adapter, capability
|
||||
*enforcement* at the verb layer (declare only), touching `inject/` turn detection, dynamic loading.
|
||||
@@ -0,0 +1,374 @@
|
||||
# MCP Contract — `bridged`'s unified gateway
|
||||
|
||||
> **Status:** 🟡 Design (2026-07-14). Greenfield — no MCP code exists yet; the pom carries
|
||||
> only Javalin/Jackson. This page defines the tool surface that CB-104 and its followers
|
||||
> implement. It supersedes nothing; it fills the "MCP server face" left open by the
|
||||
> [Architecture](1-Architecture) page.
|
||||
|
||||
`bridged` is the **sole communication gateway** for every Claude session in the bridge. Both
|
||||
the **primary** (Opus, on subscription) and every **worker** (off-subscription Claude Code)
|
||||
mount the *same* MCP server with a single `claude mcp add` line, and talk only through its
|
||||
tools. No Claude session ever addresses a broker, a peer, or the network directly.
|
||||
|
||||
This document defines every MCP tool that face must expose, who may call it, its blocking
|
||||
semantics, and how it maps onto the code already in the tree.
|
||||
|
||||
---
|
||||
|
||||
## 1. Design constraints (non-negotiable)
|
||||
|
||||
These come from the project's core invariants and bound every decision below.
|
||||
|
||||
1. **One server, both roles.** The primary and all workers mount an identical server. The
|
||||
catalog must serve both, and `bridged` must decide *who is calling* from the connection —
|
||||
never from a caller-supplied argument that could be spoofed.
|
||||
2. **Subscription-safe by construction.** No MCP tool ever reads, sets, or forwards
|
||||
`ANTHROPIC_BASE_URL`. Mounting the bridge cannot move a session off subscription.
|
||||
Enforced today by [`SubscriptionGuard`](1-Architecture).
|
||||
3. **Blocking rendezvous, no busy-poll.** The primary consumes a worker's reply through a
|
||||
*single* MCP call that `bridged` holds open — never a cross-turn poll loop that would burn
|
||||
subscription quota.
|
||||
4. **Status-gated delivery.** Anything that puts text into a worker flows through the existing
|
||||
[`Injector`](1-Architecture): delivered only when the worker is `idle`/`blocked`, at most
|
||||
one message per turn.
|
||||
5. **`bridged` owns policy; herdr owns PTYs.** MCP tools express *intent*; `bridged`
|
||||
translates it into guard checks, rendezvous bookkeeping, and herdr `agent.*` calls.
|
||||
|
||||
---
|
||||
|
||||
## 2. Topology
|
||||
|
||||
Both faces live in the one daemon. The **north face** is MCP (this document); the **south
|
||||
face** is the herdr Unix socket. REST/SSE remains only for non-Claude clients and dashboards.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
|
||||
subgraph BD["bridged — standalone daemon"]
|
||||
MCP["MCP server (north face)<br/>bridge_send · bridge_reply<br/>bridge_ask · bridge_status · lifecycle"]
|
||||
RDV["rendezvous registry<br/>(blocking-call waiters)"]
|
||||
INJ["Injector + StatusPoller<br/>(status-gated writer)"]
|
||||
SOCK["herdr socket client (south face)"]
|
||||
MCP --> RDV
|
||||
RDV --> INJ
|
||||
INJ --> SOCK
|
||||
MCP --> SOCK
|
||||
end
|
||||
HERDR["herdr<br/>panes · agent-status"]
|
||||
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
|
||||
|
||||
OPUS -->|"bridge_send (blocks)"| MCP
|
||||
W -.->|"bridge_reply / bridge_ask"| MCP
|
||||
SOCK -->|"agent.start · agent.send<br/>agent.get · pane.close"| HERDR
|
||||
HERDR -->|"drives PTY"| W
|
||||
|
||||
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
|
||||
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
|
||||
class OPUS,W ext
|
||||
class MCP,RDV,INJ,SOCK core
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Identity & addressing
|
||||
|
||||
Because the same server is mounted by everyone, `bridged` resolves the caller's role on every
|
||||
request — this is the linchpin of the whole contract and has no code yet.
|
||||
|
||||
- **Workers are known.** `bridged` spawns every worker
|
||||
([`WorkerService`](1-Architecture)) and records its herdr session UUID / `terminal_id` on
|
||||
the returned [`Agent`]. When a call arrives on a connection that maps to a known worker,
|
||||
the caller is *that* worker — so **workers never pass a target**; routing is implicit.
|
||||
- **The primary is "not a worker".** Any connection that does not map to a known worker is
|
||||
treated as a primary. It addresses workers **explicitly** by `target` — a session UUID,
|
||||
a `terminal_id`, or a friendly `profile` name.
|
||||
- **Turn correlation.** A blocking `bridge_send` registers a *waiter* keyed by worker
|
||||
identity. A worker's later `bridge_reply` / `bridge_ask` on the same identity resolves that
|
||||
waiter. A `turn_id` is minted per exchange so a clarification round-trip
|
||||
(§6.2) rejoins the right turn.
|
||||
|
||||
---
|
||||
|
||||
## 4. Transport
|
||||
|
||||
`bridged` is a long-lived daemon serving **multiple** concurrent clients (one primary + N
|
||||
workers), so a per-client stdio child is the wrong shape. The recommended transport is
|
||||
**streamable-HTTP / SSE** on the same bind as the REST face:
|
||||
|
||||
```bash
|
||||
# identical on primary and every worker
|
||||
claude mcp add --transport http bridged http://127.0.0.1:8080/mcp
|
||||
```
|
||||
|
||||
This adds an MCP-server dependency the pom does not yet carry. See [Open decisions](#10-open-decisions).
|
||||
|
||||
---
|
||||
|
||||
## 5. Tool catalog
|
||||
|
||||
| Tool | Caller | Blocks? | Backing (exists today?) |
|
||||
|---|---|---|---|
|
||||
| [`bridge_send`](#bridge_send) | primary | yes (default) | `Injector.enqueue` ✅ · rendezvous registry ❌ (CB-104) |
|
||||
| [`bridge_reply`](#bridge_reply) | worker | no | rendezvous ❌ · pane injection via `Injector` ✅ |
|
||||
| [`bridge_ask`](#bridge_ask) | worker | yes | reverse rendezvous ❌ |
|
||||
| [`bridge_status`](#bridge_status) | either | no | `AgentControl.status` ✅ · `Injector.activeTargets` ✅ |
|
||||
| [`bridge_spawn`](#lifecycle) | primary | no | `WorkerService.spawn` ✅ (`POST /workers`) |
|
||||
| [`bridge_list`](#lifecycle) | either | no | `WorkerService.list` ✅ (`/agents`) |
|
||||
| [`bridge_stop`](#lifecycle) | primary | no | `WorkerService.stop` ✅ (`DELETE /workers/{paneId}`) |
|
||||
| [`bridge_read`](#bridge_read) | primary | no | `AgentControl.read` ✅ |
|
||||
| [`bridge_cancel`](#bridge_cancel) | primary | no | — ❌ (future) |
|
||||
|
||||
### Core: delegation & rendezvous
|
||||
|
||||
#### `bridge_send`
|
||||
*(primary → worker — the headline tool, CB-104)*
|
||||
|
||||
- **Params:** `message` (required); `target` (optional — defaults to the sole worker / default
|
||||
profile); `timeout_seconds` (default 600); `block` (default `true`); `auto_spawn`
|
||||
(default `true`); `turn_id` (optional — supplied when answering a worker's `bridge_ask`).
|
||||
- **Blocking (`block:true`):** enqueue `message` via the `Injector`, then hold the call open
|
||||
until exactly one of:
|
||||
- worker calls `bridge_reply` → `{ outcome:"reply", text }`
|
||||
- worker calls `bridge_ask` → `{ outcome:"question", text, turn_id }`
|
||||
- worker's `agent_status` reaches done/idle with no reply → `{ outcome:"turn_done", text:<terminal tail> }`
|
||||
- deadline elapses → `{ outcome:"timeout" }`
|
||||
- worker gone → error `worker_gone`
|
||||
- **Detached (`block:false`):** enqueue and return `{ outcome:"dispatched", dispatch_id }`
|
||||
immediately. The eventual reply is injected into the primary's idle pane (§6.3), or drained
|
||||
via `bridge_status` on a split-host primary.
|
||||
|
||||
#### `bridge_reply`
|
||||
*(worker → primary)*
|
||||
|
||||
- **Params:** `text` (required); `final` (default `true`).
|
||||
- **Behavior:** resolve the primary waiter registered against this worker with `text`. If no
|
||||
waiter exists (detached delegation), `bridged` **injects the primary's idle pane** instead.
|
||||
Returns `{ delivered:true, mode:"resolved"|"injected" }`. No `target` — identity is implicit.
|
||||
|
||||
#### `bridge_ask`
|
||||
*(worker → primary — the reverse rendezvous)*
|
||||
|
||||
- **Params:** `question` (required); `timeout_seconds`.
|
||||
- **Behavior:** blocks the *worker's* call. Surfaces the question to the primary (resolving its
|
||||
open `bridge_send` with `outcome:"question"`, or injecting its pane). When the primary
|
||||
answers — a `bridge_send` carrying the matching `turn_id` — that unblocks this call and
|
||||
returns `{ answer }` to the worker, which continues **in the same turn**.
|
||||
|
||||
### Worker lifecycle
|
||||
<a id="lifecycle"></a>
|
||||
Thin adapters over [`WorkerService`](1-Architecture) — parity with the existing REST routes.
|
||||
|
||||
- **`bridge_spawn`** — `{ profile? }` → worker view (`sessionId`, `terminalId`, `paneId`,
|
||||
`status`). Guard-checked; a boundary breach returns error `subscription_boundary` (the
|
||||
REST `403`).
|
||||
- **`bridge_list`** — no params → all workers + `agent_status`. Read-only, either role.
|
||||
- **`bridge_stop`** — `{ target }` → tears down the pane and its dedicated tab. Idempotent.
|
||||
|
||||
### Observability
|
||||
|
||||
#### `bridge_status`
|
||||
*(either role — the README's 4th named tool)*
|
||||
|
||||
- **Params:** `target?`.
|
||||
- **Behavior:** per-worker `agent_status`, queue depth (`Injector.activeTargets`), whether a
|
||||
rendezvous is open, and ids. For the *calling* session it also reports/drains **pending
|
||||
messages addressed to me** — the path a split-host primary's `Stop`-hook uses to wake and
|
||||
collect replies without being injectable. Read-only, non-blocking.
|
||||
|
||||
#### `bridge_read`
|
||||
*(primary)*
|
||||
|
||||
- **Params:** `target`; `source` ∈ `visible | recent | recent_unwrapped | detection`.
|
||||
- **Behavior:** returns the worker's terminal text so the primary can peek at a *detached*
|
||||
worker's progress. Adapter over `AgentControl.read`.
|
||||
|
||||
### Control (future)
|
||||
|
||||
#### `bridge_cancel`
|
||||
*(primary)*
|
||||
|
||||
- **Params:** `target`. Interrupt the worker's current turn / abandon the rendezvous. No
|
||||
backing code yet.
|
||||
|
||||
---
|
||||
|
||||
## 6. Rendezvous flows
|
||||
|
||||
### 6.1 Delegation — happy path
|
||||
|
||||
One blocking call, zero polls.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant P as Primary (Opus)
|
||||
participant B as bridged (MCP + Injector)
|
||||
participant H as herdr
|
||||
participant W as Worker (Claude)
|
||||
|
||||
P->>B: bridge_send("do X", target=w) — blocks
|
||||
B->>B: register waiter(w)
|
||||
B->>H: agent.send(w, "do X") (idle window)
|
||||
H-->>W: prompt injected
|
||||
W->>W: works the turn
|
||||
W->>B: bridge_reply("result")
|
||||
B->>B: resolve waiter(w)
|
||||
B-->>P: { outcome:"reply", text:"result" }
|
||||
```
|
||||
|
||||
### 6.2 Clarification — reverse rendezvous (`bridge_ask`)
|
||||
|
||||
The worker pauses mid-turn to ask; the primary answers; the worker resumes in the same turn.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant P as Primary
|
||||
participant B as bridged
|
||||
participant W as Worker
|
||||
|
||||
P->>B: bridge_send("do X", target=w) — blocks
|
||||
B-->>W: "do X" (injected)
|
||||
W->>B: bridge_ask("which config?") — worker blocks
|
||||
B-->>P: { outcome:"question", text:"which config?", turn_id }
|
||||
P->>B: bridge_send("config.yaml", target=w, turn_id) — blocks again
|
||||
B-->>W: resolve bridge_ask → { answer:"config.yaml" }
|
||||
W->>W: resumes same turn
|
||||
W->>B: bridge_reply("done")
|
||||
B-->>P: { outcome:"reply", text:"done" }
|
||||
```
|
||||
|
||||
### 6.3 Detached delegation — pane injection
|
||||
|
||||
The primary does not block; the reply arrives later in its idle pane.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant P as Primary
|
||||
participant B as bridged
|
||||
participant W as Worker
|
||||
|
||||
P->>B: bridge_send("do X", target=w, block=false)
|
||||
B-->>P: { outcome:"dispatched", dispatch_id }
|
||||
P->>P: continues its own work
|
||||
W->>B: bridge_reply("result")
|
||||
Note over B: no waiter → detached path
|
||||
B->>B: Injector.enqueue(primary_pane, "result")
|
||||
B-->>P: injected into idle pane (status-gated)
|
||||
```
|
||||
|
||||
### 6.4 Uncooperative worker — turn-done fallback
|
||||
|
||||
A worker that never calls `bridge_reply` still returns a result: `bridged` reads its terminal
|
||||
tail when the turn completes.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant P as Primary
|
||||
participant B as bridged
|
||||
participant W as Worker
|
||||
|
||||
P->>B: bridge_send("do X", target=w) — blocks
|
||||
B-->>W: "do X" (injected)
|
||||
W->>W: works, never calls bridge_reply
|
||||
B->>B: StatusPoller sees agent_status → idle/done
|
||||
B->>B: AgentControl.read(w, "recent")
|
||||
B-->>P: { outcome:"turn_done", text:<terminal tail> }
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. Status gating
|
||||
|
||||
Delivery only happens in a safe window. This is the state machine the `Injector` already
|
||||
enforces via `AgentStatus.injectable()`; MCP `bridge_send` is simply its producer.
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> IDLE
|
||||
IDLE --> WORKING: message delivered / picks up
|
||||
WORKING --> IDLE: turn done
|
||||
WORKING --> BLOCKED: awaits input
|
||||
BLOCKED --> WORKING: input delivered
|
||||
IDLE --> UNKNOWN: detection glitch
|
||||
BLOCKED --> UNKNOWN: detection glitch
|
||||
UNKNOWN --> IDLE: re-detected
|
||||
|
||||
note right of IDLE
|
||||
injectable — deliver head of FIFO
|
||||
end note
|
||||
note right of BLOCKED
|
||||
injectable — deliver head of FIFO
|
||||
end note
|
||||
note right of WORKING
|
||||
NOT injectable — counts as pickup
|
||||
end note
|
||||
note right of UNKNOWN
|
||||
NOT injectable, NOT a pickup — wait
|
||||
end note
|
||||
```
|
||||
|
||||
At most one message is delivered per turn: after a send the `Injector` waits for a `WORKING`
|
||||
pickup before delivering the next, with a `PICKUP_GRACE_POLLS` fallback for turns faster than
|
||||
the poll interval. A herdr `events.subscribe` stream can later replace the sampling without
|
||||
touching this state machine.
|
||||
|
||||
---
|
||||
|
||||
## 8. Error model
|
||||
|
||||
| Condition | `bridge_send` result | Notes |
|
||||
|---|---|---|
|
||||
| Worker replies | `{ outcome:"reply" }` | normal |
|
||||
| Worker asks | `{ outcome:"question", turn_id }` | answer with `bridge_send(turn_id)` |
|
||||
| Turn ends, no reply | `{ outcome:"turn_done" }` | terminal tail as text |
|
||||
| Deadline elapsed | `{ outcome:"timeout" }` | message may still be queued/delivered |
|
||||
| Worker vanished | error `worker_gone` | `Injector.drop` fails the queued future |
|
||||
| Guard breach on spawn | error `subscription_boundary` | REST `403` parity |
|
||||
| Delivery failed at herdr | error, message dropped | poisoned message not left blocking the FIFO |
|
||||
|
||||
`bridge_reply` from a worker with no open waiter is **not** an error — it falls through to
|
||||
detached pane injection (§6.3).
|
||||
|
||||
---
|
||||
|
||||
## 9. Mapping to existing code
|
||||
|
||||
The MCP face is a thin adapter layer; nearly every capability already exists behind the REST
|
||||
seam. Only the **rendezvous registry** and the **caller-identity resolver** are new.
|
||||
|
||||
| MCP tool | Existing collaborator | New work |
|
||||
|---|---|---|
|
||||
| `bridge_send` | `Injector.enqueue`, `AgentControl.send` | waiter registry, timeout, outcome mux (CB-104) |
|
||||
| `bridge_reply` / `bridge_ask` | `Injector` (pane injection) | reverse rendezvous, identity resolver |
|
||||
| `bridge_status` | `AgentControl.status`, `Injector.activeTargets` | pending-drain projection |
|
||||
| `bridge_spawn` / `list` / `stop` | `WorkerService.{spawn,list,stop}` | MCP adapter only |
|
||||
| `bridge_read` | `AgentControl.read` | MCP adapter only |
|
||||
|
||||
Because the REST routes in `BridgedApp` already exercise the collaborators, MCP tools are
|
||||
validated by **parity** against those routes, not by re-testing behavior.
|
||||
|
||||
---
|
||||
|
||||
## 10. Open decisions
|
||||
|
||||
1. **`bridge_ask` direction.** This page defines it as *worker-asks-primary* (a genuine reverse
|
||||
channel, matching the "inject the primary's pane" language). The alternative — a synonym for
|
||||
a blocking primary→worker send — is weaker and produces different plumbing. **Recommend
|
||||
worker-asks-primary.**
|
||||
2. **Detached delivery shape.** A `block:false` param on `bridge_send` (keeps the catalog
|
||||
small) vs. a separate `bridge_dispatch` tool. **Recommend the param.**
|
||||
3. **Auto-spawn on send.** `bridge_send` provisions a worker per profile when none exists
|
||||
(simplest primary UX) vs. requiring an explicit `bridge_spawn` first. **Recommend
|
||||
auto-spawn, defaulting on.**
|
||||
4. **Transport & SDK.** Streamable-HTTP/SSE co-located with the REST bind (recommended) vs.
|
||||
stdio. Requires choosing a Java MCP server SDK and adding it to the pom.
|
||||
|
||||
---
|
||||
|
||||
## 11. Implementation staging
|
||||
|
||||
- **CB-104** — blocking `bridge_send` + rendezvous registry + caller-identity resolver
|
||||
(the producer that finally drives the inert `StatusPoller`).
|
||||
- **CB-1xx** — `bridge_reply` / `bridge_ask` reverse rendezvous + detached pane injection.
|
||||
- **CB-1xx** — lifecycle + observability adapters (`bridge_spawn/list/stop/status/read`).
|
||||
- **CB-1xx** — transport wiring + `claude mcp add` docs; parity tests vs. REST.
|
||||
- **Later** — `bridge_cancel`; swap `StatusPoller` for herdr `events.subscribe`.
|
||||
+154
@@ -0,0 +1,154 @@
|
||||
# Team — lead orchestrating a mixed Claude + local-LLM fleet
|
||||
|
||||
The message server (`bridged`) delivers **one turn into one worker**. A **team** is the
|
||||
layer above it: a **Claude team-lead** that fans a job out across a **mixed fleet** of
|
||||
workers — some on Claude, some on the remote local LLM — and reduces their replies. Same
|
||||
`bridged` delivery, same subscription boundary; this doc is only about **orchestration** —
|
||||
who the workers are, how the lead picks one, and how it runs many at once.
|
||||
|
||||
> Delivery mechanics (blocking `POST /message`, status-gated reply envelope) live in the
|
||||
> Message-Server design. Transport rationale is in Approaches. This doc assumes both.
|
||||
|
||||
## The team
|
||||
|
||||
- **Team-lead** — the primary **Opus** (Claude Code, env **CLEAN**, on Pro/Max). Not a
|
||||
worker; a **thin client of `bridged`**. It plans, routes, dispatches, and integrates, and
|
||||
never sets `ANTHROPIC_BASE_URL`.
|
||||
- **Workers** — a herd of `claude` panes in herdr, each an addressable `bridged` session
|
||||
with its **own model/env**:
|
||||
- **Claude workers** (clean env, e.g. Sonnet) — reasoning-heavy or high-accuracy subtasks.
|
||||
- **Local workers** (`ANTHROPIC_BASE_URL=https://ollama.ltms.dev`) — bulk, cheap, or
|
||||
embarrassingly parallel subtasks.
|
||||
|
||||
Every worker is still a *real Claude Code process* (inherits `CLAUDE.md`, hooks, skills,
|
||||
MCP) — only its model differs. Scale each kind horizontally by adding panes.
|
||||
|
||||
### Topology
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
LEAD["lead — Opus<br/>(Claude Code, env CLEAN)"]
|
||||
BD["bridged<br/>message server + router"]
|
||||
HERDR["herdr<br/>panes · agent-status"]
|
||||
WC1["w-claude-1<br/>Sonnet · CLEAN"]
|
||||
WC2["w-claude-2<br/>Sonnet · CLEAN"]
|
||||
WL1["w-local-1<br/>ANTHROPIC_BASE_URL set"]
|
||||
WL2["w-local-2<br/>ANTHROPIC_BASE_URL set"]
|
||||
ANT["api.anthropic.com<br/>(Pro/Max)"]
|
||||
OLL["ollama.ltms.dev<br/>(local model)"]
|
||||
|
||||
LEAD -->|"blocking POST /message (target role)"| BD
|
||||
BD -->|"Unix socket · send_text · events.subscribe"| HERDR
|
||||
HERDR --> WC1 & WC2 & WL1 & WL2
|
||||
WC1 --> ANT
|
||||
WC2 --> ANT
|
||||
WL1 --> OLL
|
||||
WL2 --> OLL
|
||||
|
||||
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
|
||||
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
|
||||
classDef local fill:#6b46c1,stroke:#44337a,color:#ffffff;
|
||||
class LEAD,WC1,WC2 ext
|
||||
class BD,HERDR core
|
||||
class WL1,WL2 local
|
||||
```
|
||||
|
||||
### Roles & routing
|
||||
|
||||
| Role | Env | Model | Route here when… |
|
||||
|---|---|---|---|
|
||||
| `lead` | clean | Opus (sub) | always — it does the routing |
|
||||
| `w-claude-*` | clean | Sonnet (sub) | task needs Claude-grade reasoning / careful edits |
|
||||
| `w-local-*` | `ANTHROPIC_BASE_URL` set | local LLM | task is bulk / cheap / embarrassingly parallel |
|
||||
|
||||
The lead applies this rubric itself, guided by its `CLAUDE.md` team charter (below). Worker
|
||||
selection is **policy in the lead**, not a `bridged` concern — `bridged` just delivers to
|
||||
the session the lead names.
|
||||
|
||||
## Subscription boundary in a team
|
||||
|
||||
Unchanged from the base architecture, and it scales with the fleet: **only local-worker
|
||||
panes** launch with `ANTHROPIC_BASE_URL`. The lead and every Claude worker stay env-clean on
|
||||
the subscription. `bridged` enforces which panes may carry the off-subscription env, so
|
||||
adding workers never widens the boundary.
|
||||
|
||||
## Parallel fan-out (map / reduce)
|
||||
|
||||
The lead's advantage over a single bridge is **concurrency**: independent subtasks go to
|
||||
different workers at once, then results are gathered.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant L as lead (Opus)
|
||||
participant B as bridged
|
||||
participant WC as w-claude-1
|
||||
participant WL as w-local-1
|
||||
|
||||
Note over L: split job → subtask A (reasoning), subtask B (bulk)
|
||||
par A → Claude worker
|
||||
L->>B: POST /message {role: w-claude, prompt: A}
|
||||
B->>WC: send_text into running pane
|
||||
WC-->>B: status working → idle + Stop-hook envelope
|
||||
B-->>L: 200 reply A
|
||||
and B → local worker
|
||||
L->>B: POST /message {role: w-local, prompt: B}
|
||||
B->>WL: send_text into running pane
|
||||
WL-->>B: status working → idle + Stop-hook envelope
|
||||
B-->>L: 200 reply B
|
||||
end
|
||||
Note over L: reduce → integrate A + B into final answer
|
||||
```
|
||||
|
||||
- **Map:** the lead issues N concurrent blocking `POST /message` calls (one per subtask → its
|
||||
chosen worker). Each call blocks only *that* request; `bridged` holds it open until the
|
||||
worker's turn completes (status-gated) and returns the reply envelope.
|
||||
- **Reduce:** the lead collects the N envelopes and integrates. A slow local worker never
|
||||
blocks a fast Claude worker — wall-clock ≈ the slowest single subtask, not the sum.
|
||||
- **Detached / long jobs** use the async broker path instead of a held request (Channel 2 in
|
||||
the base architecture), so the lead never busy-polls across turns.
|
||||
|
||||
Fan-out is bounded by the herd size (pane count) and `bridged`'s concurrency policy, not by
|
||||
the lead.
|
||||
|
||||
## Knowing the roster
|
||||
|
||||
The lead discovers its team from `bridged` (session list / roles) rather than hard-coding
|
||||
pane ids, so workers can be added or restarted without editing the lead. A minimal charter
|
||||
in the lead's `CLAUDE.md` turns Opus into the orchestrator:
|
||||
|
||||
```markdown
|
||||
## Your team (via bridged)
|
||||
You are the team-lead. Delegate through the bridged client — never launch workers yourself.
|
||||
Roster: ask bridged for current sessions/roles.
|
||||
- w-claude-* — Claude Sonnet. Reasoning-heavy / high-accuracy subtasks.
|
||||
- w-local-* — remote local LLM. Bulk, cheap, or parallelizable subtasks.
|
||||
|
||||
Route each subtask by the rubric in the Team design. For independent subtasks, DISPATCH ALL
|
||||
of them (concurrent blocking sends), THEN gather — never serialize independent work.
|
||||
Integrate the reply envelopes; you own the final answer.
|
||||
```
|
||||
|
||||
Wrap the send as a Claude Code skill (`/delegate <role> "<task>"`) so the lead calls one
|
||||
tool instead of hand-rolling the HTTP request.
|
||||
|
||||
## What this layer does NOT change
|
||||
|
||||
- **Delivery** is still `bridged` → herdr `pane.send_text` + status events (Message-Server).
|
||||
- **Completion timing** is still the worker status event; **reply content** still rides the
|
||||
worker `Stop`-hook envelope.
|
||||
- **Single-host** still applies: herdr's socket is local, so the whole herd lives on the
|
||||
`bridged` host. The lead may be remote — it only needs HTTP to `bridged`.
|
||||
|
||||
## Open questions
|
||||
|
||||
- **Routing intelligence:** rubric-in-`CLAUDE.md` (lead decides) vs. a `bridged` role-router
|
||||
(label-based). Start with the former; promote to the latter if routing logic grows.
|
||||
- **Backpressure:** per-role concurrency caps in `bridged` so a fan-out can't exhaust the
|
||||
local gateway.
|
||||
- **Result schema:** whether reply envelopes should carry structured metadata (worker, model,
|
||||
tokens) to help the lead's reduce step.
|
||||
|
||||
## Status
|
||||
|
||||
🟡 Design (2026-07-11). Orchestration layer over the selected `bridged` server; inherits
|
||||
herdr (chosen) + AgentAPI (fallback). Delivery unchanged — see the Message-Server design.
|
||||
@@ -0,0 +1,182 @@
|
||||
# Worker Git Workflow — worktree · branch · PR
|
||||
|
||||
**Status:** design (defining the fleet's working model). Builds on the
|
||||
[CB-301 session manager](CB-301-Session-Manager.md) and the one-shot/no-reuse decision.
|
||||
|
||||
## Guiding principle — a worker is a full peer of the primary
|
||||
|
||||
The whole point of the bridge is **the same Claude Code agent running against a different LLM
|
||||
provider**. A worker must be **functionally identical to the primary in context and knowledge** —
|
||||
same project + user `CLAUDE.md`, same skills, same memory, same MCP tools, **same local/private
|
||||
settings** — and differ **only** in `ANTHROPIC_BASE_URL`/`ANTHROPIC_MODEL`. The worktree exists
|
||||
*solely* for git isolation (a branchable checkout for commits + code reference). It must **never**
|
||||
strip the worker of the configuration a main-tree session has. **Config parity is a hard
|
||||
requirement, not a nice-to-have.**
|
||||
|
||||
## Decisions
|
||||
|
||||
- **One-shot, no reuse** (CB-301) — each task gets a fresh worker, torn down after. The **PR is the
|
||||
durable artifact**; no context is carried across workers.
|
||||
- **Worktree provisioned by the daemon** — `SessionManager` creates a dedicated git worktree +
|
||||
branch per session, **hydrates it to full config parity** (below), and tears it down on release.
|
||||
- **Worker opens its own PR** — the worker commits, pushes its branch, and opens the PR/MR itself,
|
||||
returning the PR URL in its `bridge_reply`.
|
||||
|
||||
## Why worktrees (the hazard being fixed)
|
||||
|
||||
Today every bridge worker inherits the **primary's own working tree** as its cwd
|
||||
(`WorkerService` cwd resolution → caller cwd). A single worker editing at a time is safe, but two
|
||||
**parallel implementers** would stomp each other's files. A worktree per session gives each worker
|
||||
an isolated checkout on its own branch — the precondition for fanning out implementation work.
|
||||
|
||||
## Config parity — the worktree is code-only; it must NOT strip the worker
|
||||
|
||||
**The trap:** `git worktree add` creates a checkout of **tracked files only**. Untracked / gitignored
|
||||
files do **not** come along. So a worker moved from the primary's main tree into a bare worktree
|
||||
**silently loses** exactly the local/private configuration that makes it a full peer — while the
|
||||
shared-tree it runs in *today* gives it all of this for free. Moving to worktrees must **preserve**
|
||||
that, not regress it.
|
||||
|
||||
Classify every source of "what a main session knows", by whether a worktree keeps it:
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
subgraph keeps["Inherited automatically — no action"]
|
||||
A["User-global config<br/>~/.claude/CLAUDE.md, ~/.ccs memory"]:::ok
|
||||
B["Tracked project config<br/>CLAUDE.md, committed .claude/skills, committed .mcp.json"]:::ok
|
||||
C["CLAUDE_CONFIG_DIR<br/>(daemon already injects per profile)"]:::ok
|
||||
D["Bridge MCP<br/>(injected via --mcp-config launch flag)"]:::ok
|
||||
end
|
||||
subgraph gap["LOST by a bare worktree — must be hydrated"]
|
||||
E["settings.local.json<br/>local .claude/* overrides"]:::warn
|
||||
F["local .mcp.json mods<br/>(the M .mcp.json in git status)"]:::warn
|
||||
G[".env / .envrc / direnv<br/>local secrets + tokens"]:::warn
|
||||
H["any other gitignored local config"]:::warn
|
||||
end
|
||||
keeps --> R["Worker = full peer of primary"]:::goal
|
||||
gap -->|"overlay step at provision"| R
|
||||
classDef ok fill:#2f855a,stroke:#22543d,color:#ffffff;
|
||||
classDef warn fill:#b7791f,stroke:#7b341e,color:#ffffff;
|
||||
classDef goal fill:#2b6cb0,stroke:#2a4365,color:#ffffff;
|
||||
```
|
||||
|
||||
*Green is inherited by path (home dir / `CLAUDE_CONFIG_DIR`) or lives in tracked files that the
|
||||
worktree checks out anyway. Amber is the real gap — untracked local config the worktree drops.*
|
||||
|
||||
**Mechanism — worktree hydration (part of `SessionManager.acquire`, after `git worktree add`):**
|
||||
|
||||
1. **Inherit by path, don't copy** — keep the worker's `$HOME`, `CLAUDE_CONFIG_DIR`, and memory dir
|
||||
identical to the primary's. Everything home-scoped (user `CLAUDE.md`, memory, auth, global
|
||||
skills) is already parity for free; only *cwd-relative* project-local files are the gap.
|
||||
2. **Overlay the untracked project-local set** from the primary tree into the worktree — a defined,
|
||||
configurable list: `settings.local.json` (+ any local `.claude/*`), the locally-modified
|
||||
`.mcp.json`, `.env`/`.envrc`, and any other gitignored config the primary depends on. **Symlink**
|
||||
(read-only parity, stays live, nothing to go stale) rather than copy where possible; copy only
|
||||
what a worker may write.
|
||||
3. **Never overlay the git plumbing** — the worktree's own `.git` file/branch is what gives
|
||||
isolation; that's the *one* thing that must differ from the main tree.
|
||||
|
||||
The overlay set lives in config (`BridgedConfig.Worker.parityOverlay` — a list of repo-relative
|
||||
paths, with sane defaults) so it's auditable and per-repo tunable.
|
||||
|
||||
> **Trust note (deliberate).** Hydrating local config means the primary's local secrets/tokens
|
||||
> (`.env`, `.mcp.json` auth, gitea token) become visible to an **off-subscription** worker running
|
||||
> against a third-party model endpoint. That is the accepted consequence of "workers must be full
|
||||
> peers" — but it is a real trust expansion over a worker that only sees tracked code. Keep the
|
||||
> overlay list **explicit and minimal**; don't blanket-symlink the whole tree. `.mcp.json` and
|
||||
> `wiki/` remain **excluded from all worker commits** regardless of being present for reference.
|
||||
|
||||
## Lifecycle
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant P as Primary
|
||||
participant SM as SessionManager (daemon)
|
||||
participant G as git / gitea
|
||||
participant W as Worker
|
||||
|
||||
P->>SM: acquire(ticket, profile)
|
||||
SM->>G: git worktree add wt -b worker/ticket-nonce main
|
||||
SM->>SM: overlay parity config into wt
|
||||
SM->>W: spawn (cwd = wt, on its branch)
|
||||
Note over W: implement in the isolated worktree
|
||||
W->>G: git commit + git push (SSH, same user)
|
||||
W->>G: open PR (branch to main)
|
||||
W-->>P: bridge_reply (prUrl, branch, summary, tests)
|
||||
P->>SM: release(paneId)
|
||||
SM->>G: git worktree remove wt
|
||||
Note over G: branch + PR persist for review/merge
|
||||
P->>G: review PR, merge on green
|
||||
```
|
||||
|
||||
*The worker's "checkpoint" (CB-302) is exactly steps 6–8: commit → push → open PR. This replaces the
|
||||
earlier `STATE.md` idea — a PR is reviewable, mergeable, and self-describing.*
|
||||
|
||||
## Infra facts (verified this session)
|
||||
|
||||
- **Remote:** `ssh://git@git.ltms.dev:2224/lms/claude-bridge.git` (gitea). Push is over **SSH** —
|
||||
a worker running as the same user with the same keys can `git push` **with no extra credential**.
|
||||
- **gitea is NOT in the project `.mcp.json`** (only `jetbrains`, `intellij-index`, `bridged`). The
|
||||
primary's gitea MCP comes from a global/user config, so **workers do not inherit it**. A worker
|
||||
gets only the `bridge` MCP mounted (via `--mcp-config` launch flag).
|
||||
- **No gitea CLI** (`tea`) installed; `glab` is present but is the GitLab CLI (wrong backend).
|
||||
|
||||
**Implication:** `git push` is free for workers; only **PR creation** needs a new mechanism.
|
||||
|
||||
## Open decision — how the worker creates the PR
|
||||
|
||||
| Option | Mechanism | Trade-off |
|
||||
|---|---|---|
|
||||
| **A. gitea REST + token** | Worker `curl`s `POST /api/v1/repos/lms/claude-bridge/pulls` with a scoped token injected by the daemon into the worker env | Minimal, no new server; token lives in the off-subscription worker's env (scope it tightly) |
|
||||
| **B. mount gitea MCP into workers** | Add the gitea MCP to the worker's `--mcp-config` alongside `bridge` | Clean tool call, but the gitea MCP's own auth/token must be provisioned per worker; more moving parts |
|
||||
| **C. install `tea` CLI** | Worker runs `tea pr create` with a token | Another dependency to install + configure; same token question as A |
|
||||
|
||||
**Recommendation: A (gitea REST + a repo-scoped token).** Smallest surface, reuses SSH for push,
|
||||
and the token is a single scoped secret the daemon injects like it already injects
|
||||
`ANTHROPIC_AUTH_TOKEN`. The implementer skill wraps the `curl` in one documented step.
|
||||
|
||||
### Trust / token scope (the real cost of "worker opens its own PR")
|
||||
|
||||
- Off-subscription workers already *could* push (SSH, same user). The **incremental grant is
|
||||
PR-create**, i.e. a gitea API token.
|
||||
- Scope the token **minimally**: the `lms/claude-bridge` repo, `write:repository` (create branch +
|
||||
PR), **not** merge/admin/org. A leaked token can open PRs, not merge them — the primary/human is
|
||||
still the merge gate.
|
||||
- Inject via the daemon (env var, e.g. `GITEA_TOKEN`), never written to the worker's config dir —
|
||||
same non-invasive pattern as the ANTHROPIC token. Guard is unaffected (it concerns
|
||||
`ANTHROPIC_BASE_URL`, not git).
|
||||
|
||||
## Implementation plan
|
||||
|
||||
| Piece | Where | Notes |
|
||||
|---|---|---|
|
||||
| Worktree provision/teardown | **CB-301 ext** — `SessionManager.acquire`/`release`; `WorkerSession` gains `worktree`, `branch` | daemon shells out to `git worktree add/remove` |
|
||||
| **Config-parity overlay** | **CB-301 ext** — `SessionManager.acquire`, after `git worktree add` | symlink/copy the `parityOverlay` set into the worktree so the worker is a full peer; **this is what makes worktrees viable, not a dead-end** |
|
||||
| Overlay config | `BridgedConfig.Worker.parityOverlay` — repo-relative paths, sane defaults | auditable, per-repo tunable; keep explicit + minimal (trust) |
|
||||
| Branch naming | `worker/<ticket-slug>-<nonce>` off `main` (or a configured base) | one branch per session |
|
||||
| Commit + push + PR handoff | **CB-302** — worker-driven, guided by the skill | push = SSH; PR = option A |
|
||||
| Implementer skill | `.claude/skills/implementer/SKILL.md` | worktree-aware playbook (see below); mounts automatically since workers inherit repo cwd |
|
||||
| gitea token injection | `WorkerService` env + `BridgedConfig` | repo-scoped, minimal perms |
|
||||
| PR review + merge | Primary (has gitea MCP + judgment) | merge on green; the human/primary gate stays |
|
||||
|
||||
## Implementer skill (outline)
|
||||
|
||||
A worker-facing playbook (sibling to the existing `reviewer` skill):
|
||||
|
||||
1. **You are in a git worktree on a dedicated branch** — check `git status`/`git branch`; do all
|
||||
work here, never on `main`.
|
||||
2. **Implement the task**; keep commits focused and message them clearly.
|
||||
3. **Push** your branch (`git push -u origin HEAD`).
|
||||
4. **Open a PR** to `main` (option A `curl`, or the decided mechanism) with a title/body describing
|
||||
the change and referencing the ticket.
|
||||
5. **Reply** via `bridge_reply` with the **PR URL**, branch name, files changed, and test names —
|
||||
that reply is the whole handoff.
|
||||
6. Do **not** merge; do **not** touch `.mcp.json` or `wiki/`.
|
||||
|
||||
## Sequencing
|
||||
|
||||
1. Finish + verify **CB-301 core** (in flight) — registry/FSM.
|
||||
2. Extend CB-301 with **worktree provisioning** (this doc) once the PR mechanism is chosen.
|
||||
3. Add the **implementer skill** + **token injection**.
|
||||
4. **CB-302** = wire the worker checkpoint (commit/push/PR) as the release-time handoff.
|
||||
@@ -0,0 +1,137 @@
|
||||
# Worker startup: working directory & the folder-trust prompt
|
||||
|
||||
When `bridged` spawns a worker, the worker CLI may show an **interactive startup prompt** before it
|
||||
is ready to accept a task — most importantly a *"Do you trust the files in this folder?"* dialog. An
|
||||
unattended worker parked on that prompt never becomes injectable: the status-gated injector waits for
|
||||
`idle`/`blocked`, the task is never delivered, and (worst case) a stray Enter answers the dialog
|
||||
wrong. How this is handled depends on two things:
|
||||
|
||||
1. **The worker's working directory** — which folder the CLI is asked to trust.
|
||||
2. **Which CLI launches the worker** — each has its own first-run/trust behaviour.
|
||||
|
||||
This doc pins the current assumption (**`ccs`** is the launcher), how its trust prompt works, the
|
||||
rule that **a worker inherits the primary's directory** (never `$HOME`), and how other CLIs differ.
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
A["bridge_spawn / POST /workers"] --> B{"explicit cwd?<br/>(profile cwd or spawn arg)"}
|
||||
B -->|"yes — told otherwise"| C["use that cwd"]
|
||||
B -->|"no"| D{"caller PID resolvable?<br/>(MCP peer PID)"}
|
||||
D -->|"yes"| E["cwd = the primary's cwd<br/>lsof -a -p PID -d cwd"]
|
||||
D -->|"no (REST / off-host)"| F["cwd = bridged daemon cwd<br/>(never $HOME by assumption)"]
|
||||
C --> G["ensureWorkspace → tab.create → agent.start {cwd}"]
|
||||
E --> G
|
||||
F --> G
|
||||
G --> H{"does the CLI trust this folder?"}
|
||||
H -->|"yes"| I["worker reaches its prompt → injectable"]
|
||||
H -->|"no"| J["worker BLOCKS on the trust dialog<br/>never injectable"]
|
||||
classDef good fill:#2f855a,stroke:#22543d,color:#ffffff;
|
||||
classDef bad fill:#9b2c2c,stroke:#63171b,color:#ffffff;
|
||||
class I good
|
||||
class J bad
|
||||
```
|
||||
|
||||
*Figure 1 — spawn resolves a working directory, then the CLI's trust check gates readiness.*
|
||||
|
||||
## Working directory: inherit the primary's path
|
||||
|
||||
**Rule: a worker opens the same directory the primary (main) session is working in, unless told
|
||||
otherwise. Never assume `$HOME`.** If the primary is in `/Users/you/LTMS/claude-bridge`, its workers
|
||||
open there too — so delegated tasks share the same relative paths and the same (already-trusted)
|
||||
project folder.
|
||||
|
||||
**The herdr seam.** An `agent.start` pane does **not** inherit its tab's or workspace's cwd — it
|
||||
starts in `$HOME` unless told otherwise. herdr's `agent.start` honours an (undocumented) **`cwd`**
|
||||
param, verified live: setting it roots the worker process at that directory. So the worker's cwd is
|
||||
threaded onto `agent.start {…, cwd}`, not the placement step (`workspace.create`/`tab.create` cwd
|
||||
only affect the seed shell, which the bridge closes).
|
||||
|
||||
**Resolution order** (first match wins):
|
||||
|
||||
| # | Source | When |
|
||||
|---|--------|------|
|
||||
| 1 | Explicit `cwd` — a per-profile `cwd:` in config, or a spawn argument | "told otherwise" — pin a fixed workdir |
|
||||
| 2 | The **primary's cwd**, auto-detected from the `bridge_spawn` caller | normal MCP spawn from the primary |
|
||||
| 3 | The `bridged` daemon's own cwd | REST spawn / off-host caller — **never `$HOME`** |
|
||||
|
||||
The primary's cwd (source 2) is discoverable with no new plumbing: `bridged` already resolves the MCP
|
||||
caller's loopback **peer PID** for connection identity (`ConnectionIdentity` → `LsofPeerPidLookup`);
|
||||
the same PID yields its cwd via `lsof -a -p <pid> -d cwd -Fn` (the `n…` line). The primary maps to no
|
||||
worker pane (it is not a worker), but its PID and cwd are still readable.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant P as "Primary (main)"
|
||||
participant B as "bridged"
|
||||
participant O as "OS (lsof)"
|
||||
participant H as "herdr"
|
||||
P->>B: "bridge_spawn {profile} (no cwd)"
|
||||
B->>O: "peer PID for this connection's port"
|
||||
O-->>B: "pid"
|
||||
B->>O: "cwd of pid (lsof -d cwd)"
|
||||
O-->>B: "/Users/you/LTMS/claude-bridge"
|
||||
B->>H: "tab.create (placement)"
|
||||
B->>H: "agent.start {argv, env, tab_id, cwd}"
|
||||
H-->>B: "worker in the primary's directory"
|
||||
```
|
||||
|
||||
*Figure 2 — a no-cwd spawn inherits the primary's directory from the caller's PID.*
|
||||
|
||||
> **Status:** implemented (CB-112). `bridged` threads the resolved `cwd` onto **`agent.start {cwd}`**
|
||||
> (verified: the worker process is rooted there), keeping the single shared worker space. On an MCP
|
||||
> `bridge_spawn` the primary's cwd is auto-detected from the caller's PID; over REST (no MCP caller)
|
||||
> it is the explicit `cwd` param else the daemon's cwd. Both placements (`tab` and legacy `pane`)
|
||||
> carry it, since it rides `agent.start`.
|
||||
|
||||
## Assumed launcher: `ccs` (Claude Code)
|
||||
|
||||
For now the fleet assumes **`ccs`** (Claude Code under the hood) as the worker CLI — `argv: ["ccs",
|
||||
"<profile>"]`. Its startup gate is the **folder-trust dialog**.
|
||||
|
||||
### How `ccs`/Claude Code decides whether to prompt
|
||||
|
||||
Trust is recorded **per-directory, per config dir**. Each `ccs` profile is an isolated instance with
|
||||
its own config dir (`~/.ccs/instances/<profile>/`) and its own `.claude.json`:
|
||||
|
||||
```jsonc
|
||||
// ~/.ccs/instances/<profile>/.claude.json
|
||||
"projects": {
|
||||
"/Users/you/LTMS/claude-bridge": { "hasTrustDialogAccepted": true }, // trusted → no prompt
|
||||
"/Users/you": { "hasTrustDialogAccepted": false } // untrusted → prompts
|
||||
}
|
||||
```
|
||||
|
||||
The worker prompts **iff** its cwd is not marked `hasTrustDialogAccepted: true` for *that instance*.
|
||||
This is why the directory rule above matters: land workers in the primary's project folder and you
|
||||
grant trust **once per profile**, instead of scattering trust across `$HOME` and ad-hoc dirs.
|
||||
|
||||
### Clearing the prompt (ranked)
|
||||
|
||||
1. **Inherit the primary's project dir** (the rule above) and trust that folder once per profile.
|
||||
2. **Pre-trust interactively:** run `ccs <profile>` in the target folder and accept — persists
|
||||
`hasTrustDialogAccepted: true` for that path in the instance's `.claude.json`.
|
||||
3. **Set the flag directly** (scriptable, no interaction): set
|
||||
`projects["<cwd>"].hasTrustDialogAccepted = true` in `~/.ccs/instances/<profile>/.claude.json`.
|
||||
4. **Do not** reach for `--dangerously-skip-permissions` — it disables *all* permission gating, not
|
||||
just this dialog, which defeats running off-subscription workers autonomously.
|
||||
|
||||
## Other CLIs: different launchers, different prompts
|
||||
|
||||
`ccs` is the current assumption, not a hard dependency — a worker profile's `argv` can be any CLI.
|
||||
Each launcher has its **own** first-run/trust gate, so the "clear the prompt" step is CLI-specific
|
||||
and belongs with the profile, not hard-coded:
|
||||
|
||||
| Launcher (`argv`) | Startup gate | How to clear it |
|
||||
|---|---|---|
|
||||
| `ccs <profile>` (Claude Code) | Folder-trust dialog | `hasTrustDialogAccepted` per project in the instance's `.claude.json` (above) |
|
||||
| Other Claude-compatible runtimes via `ccs` (codex, gemini, cursor, …) | Each has its own first-run / trust / login prompt | Per-runtime; document per launcher as it is adopted |
|
||||
| A bare command (`bash -c …`, mechanics probe) | None | n/a — used for non-interactive smoke tests |
|
||||
|
||||
When adding a new launcher, capture its startup-prompt behaviour here (what blocks, and the
|
||||
non-interactive way to satisfy it) so a spawned worker of that kind reaches an injectable prompt
|
||||
unattended.
|
||||
|
||||
## See also
|
||||
|
||||
- `docs/MCP-Contract.md` — the tool surface (`bridge_spawn`, `bridge_profiles`, …).
|
||||
- `wiki/2-Message-Server.md` — the herdr `agent.*` / `workspace.*` schema (`workspace.create {cwd}`).
|
||||
@@ -0,0 +1,2 @@
|
||||
__pycache__/
|
||||
*.pyc
|
||||
@@ -0,0 +1,77 @@
|
||||
# Bridge conversation test (`e2e/`)
|
||||
|
||||
A standard, repeatable **live** end-to-end test of the two-way channel: it drives a real
|
||||
multi-turn conversation between a primary and an off-subscription worker **through the
|
||||
running `bridged` daemon**, captures the full transcript, and grades the channel.
|
||||
|
||||
This is the committed form of the ad-hoc channel test that discovered the CB-115 gaps
|
||||
(herdr `unknown` misclassification wedging delivery, dirty completion scrapes, and workers
|
||||
never calling `bridge_reply` in conversation). Run it after any change to the injector,
|
||||
status handling, completion/failure paths, or the worker reply charter.
|
||||
|
||||
## What it exercises
|
||||
|
||||
Each turn goes through the whole gateway exactly as a primary Opus session would — async
|
||||
fire-and-poll (`POST /sessions/{id}/message {"wait":false}` → `GET /tasks/{ticket}`), so it
|
||||
also validates the path that beats the caller's MCP timeout. It never sets
|
||||
`ANTHROPIC_BASE_URL` and never talks to herdr directly, so it is **subscription-safe by
|
||||
construction** — it only calls the bridge's loopback REST face.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant T as conversation_test.py
|
||||
participant B as bridged (REST)
|
||||
participant W as worker (off-sub)
|
||||
T->>B: POST /workers (spawn)
|
||||
T->>B: GET /sessions/{id}/status (await ready)
|
||||
loop each turn
|
||||
T->>B: POST /sessions/{id}/message {wait:false}
|
||||
B-->>T: ticket
|
||||
B->>W: inject prompt (status-gated)
|
||||
W-->>B: bridge_reply
|
||||
T->>B: GET /tasks/{ticket} (poll)
|
||||
B-->>T: done + reply
|
||||
end
|
||||
T->>B: DELETE /workers/{pane} (stop)
|
||||
```
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- `bridged` is running (default REST on `http://127.0.0.1:8765`) with at least one worker
|
||||
profile configured and its backend reachable.
|
||||
- herdr is up (the daemon needs it).
|
||||
- Python 3 (standard library only — no pip installs).
|
||||
|
||||
## Run
|
||||
|
||||
```bash
|
||||
# spawn the default-profile worker, run the built-in 5-turn conversation, grade, clean up
|
||||
python3 e2e/conversation_test.py
|
||||
|
||||
# pick a profile / reuse a live worker / use your own prompts
|
||||
python3 e2e/conversation_test.py --profile ollama
|
||||
python3 e2e/conversation_test.py --tid term_abc123 --keep-worker
|
||||
python3 e2e/conversation_test.py --prompts my_prompts.txt --out /tmp/run1
|
||||
```
|
||||
|
||||
A prompts file is one prompt per line; blank lines and `#` comments are ignored.
|
||||
|
||||
## Output & grading
|
||||
|
||||
- Writes `transcript.md` (in `--out`, default `e2e/`) — every turn's prompt, worker reply,
|
||||
latency, resolution source, and observed status transitions, with an inline `> **GAP**`
|
||||
note on any non-clean turn.
|
||||
- Prints a per-turn line and an overall summary, and **exits non-zero** if any turn wedged,
|
||||
failed, or returned empty — so it is CI-usable.
|
||||
|
||||
Per-turn grade:
|
||||
|
||||
| Grade | Meaning |
|
||||
|------------|---------------------------------------------------------------------|
|
||||
| `OK` | delivered and resolved by an explicit `bridge_reply` (`source=reply`) |
|
||||
| `DEGRADED` | delivered and answered, but resolved via completion-scrape fallback |
|
||||
| `EMPTY` | turn completed but the reply was empty |
|
||||
| `FAILED` | the worker's turn ended in failure (`phase=failed`) |
|
||||
| `WEDGE` | never resolved within the poll window (delivery wedge / lost turn) |
|
||||
|
||||
`PASS` requires every turn to be `OK` or `DEGRADED`; a clean run is every turn `OK`.
|
||||
@@ -0,0 +1,277 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Live bridge_ask test — the REVERSE rendezvous (CB-205), watched end to end.
|
||||
|
||||
Every other harness drives the forward path: primary `bridge_send` → worker `bridge_reply`.
|
||||
This drives the one that runs the other way. A worker is told to pause its delegated turn,
|
||||
ask the primary a question via `bridge_ask`, and only finish once it has the answer — so the
|
||||
turn round-trips primary→worker→primary→worker inside a SINGLE delegation.
|
||||
|
||||
The mechanics that only this path exercises:
|
||||
|
||||
• a worker's mid-turn question surfacing on the primary's *own* blocked send (Outcome.QUESTION),
|
||||
• the `turnId` correlation that lets the primary answer the exact paused turn,
|
||||
• the answer resuming that same turn and the worker's final `bridge_reply` landing on the
|
||||
re-opened forward waiter (never a stale or cross-wired one).
|
||||
|
||||
It is two blocking REST calls, no polling:
|
||||
|
||||
1. POST /sessions/{id}/message {content: TASK} → blocks, returns 202 {status:"question",
|
||||
question, turnId} (the worker asked)
|
||||
2. POST /sessions/{id}/message {content: ANSWER, turnId} → blocks, returns 200 {reply, replySource}
|
||||
(the worker resumed and replied)
|
||||
|
||||
Like the rest of the suite it talks ONLY to the bridge's REST face on loopback — it never sets
|
||||
ANTHROPIC_BASE_URL and never touches herdr, so it is subscription-safe by construction.
|
||||
|
||||
Usage:
|
||||
python3 bridge_ask_test.py [--base URL] [--profile NAME] [--repo DIR] [--out DIR]
|
||||
[--send-timeout SECS] [--answer-timeout SECS] [--keep-worker]
|
||||
|
||||
--base bridge REST base URL (default http://127.0.0.1:8765)
|
||||
--profile profile for the worker (default: the daemon's default)
|
||||
--repo cwd handed to the worker (default: the bridge repo root)
|
||||
--out output dir for the transcript (default: alongside this file)
|
||||
--send-timeout max wait for the worker to ASK (default 110)
|
||||
--answer-timeout max wait for the worker to REPLY (default 90)
|
||||
--keep-worker do not stop the spawned worker at the end
|
||||
|
||||
Exit code: 0 if the worker asked, the answer resumed the turn, and the final reply reflected
|
||||
the answer; 1 otherwise (CI-usable).
|
||||
"""
|
||||
import argparse
|
||||
import json
|
||||
import pathlib
|
||||
import sys
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from datetime import datetime
|
||||
|
||||
HERE = pathlib.Path(__file__).parent
|
||||
REPO_ROOT = HERE.parent
|
||||
|
||||
# The color the primary will hand back when the worker asks. The final reply must reflect it,
|
||||
# uppercased — proof the answer actually reached the resumed turn (not a value the worker could
|
||||
# have guessed: it is told to ask, and only the primary knows which of red/blue is chosen).
|
||||
ANSWER_COLOR = "blue"
|
||||
|
||||
# A task that CANNOT be completed without asking: the worker is not told which color to choose,
|
||||
# only that the primary will name one when asked. So a correct final reply is only reachable by
|
||||
# actually calling bridge_ask and using the answer.
|
||||
TASK_PROMPT = (
|
||||
"You are a bridge worker in a quick coordination game. You do NOT know which color to pick — "
|
||||
"only the primary does. Do exactly this, in order:\n"
|
||||
"1. Call the `bridge_ask` tool with EXACTLY this question: \"PICK A COLOR: red or blue?\"\n"
|
||||
"2. The primary will answer with one color word. Take that color and uppercase it.\n"
|
||||
"3. Call `bridge_reply` with EXACTLY one line: CHOSEN=<COLOR> (e.g. CHOSEN=GREEN if told green).\n"
|
||||
"Do not guess a color. Do not call bridge_reply before bridge_ask has returned an answer. "
|
||||
"Do nothing else — no file reads, no other tools."
|
||||
)
|
||||
|
||||
|
||||
def http(base, method, path, body=None, timeout=20):
|
||||
"""JSON request. Tolerates an empty body (e.g. 204) → {}. Raises on non-2xx via urllib."""
|
||||
data = json.dumps(body).encode() if body is not None else None
|
||||
req = urllib.request.Request(base + path, data=data, method=method,
|
||||
headers={"Content-Type": "application/json"})
|
||||
with urllib.request.urlopen(req, timeout=timeout) as r:
|
||||
raw = r.read().decode().strip()
|
||||
return json.loads(raw) if raw else {}
|
||||
|
||||
|
||||
def now():
|
||||
return datetime.now().strftime("%H:%M:%S")
|
||||
|
||||
|
||||
def spawn_worker(base, profile, repo):
|
||||
res = http(base, "POST", "/workers", {"profile": profile, "cwd": repo} if profile
|
||||
else {"cwd": repo})
|
||||
tid = res.get("terminalId") or res.get("sessionId")
|
||||
if not tid:
|
||||
raise RuntimeError(f"spawn failed: {res}")
|
||||
pane = res.get("paneId")
|
||||
print(f"[{now()}] spawned {tid} (profile={profile or 'default'}, pane={pane})")
|
||||
return tid, pane
|
||||
|
||||
|
||||
def await_ready(base, tid, timeout=150):
|
||||
deadline = time.time() + timeout
|
||||
while time.time() < deadline:
|
||||
try:
|
||||
st = http(base, "GET", f"/sessions/{tid}/status")
|
||||
except urllib.error.URLError:
|
||||
st = {}
|
||||
if st.get("ready"):
|
||||
print(f"[{now()}] worker ready (status={st.get('status')})")
|
||||
return True
|
||||
time.sleep(3)
|
||||
print(f"[{now()}] WARNING: worker never reported ready within {timeout}s — sending anyway")
|
||||
return False
|
||||
|
||||
|
||||
def post_message(base, tid, body, timeout):
|
||||
"""One BLOCKING send. Returns (http_status, parsed_json). urllib raises on 4xx/5xx, so a
|
||||
stale-turn 409 is surfaced here rather than swallowed."""
|
||||
data = json.dumps(body).encode()
|
||||
req = urllib.request.Request(f"{base}/sessions/{tid}/message", data=data, method="POST",
|
||||
headers={"Content-Type": "application/json"})
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=timeout) as r:
|
||||
raw = r.read().decode().strip()
|
||||
return r.status, (json.loads(raw) if raw else {})
|
||||
except urllib.error.HTTPError as e:
|
||||
raw = e.read().decode().strip()
|
||||
return e.code, (json.loads(raw) if raw else {})
|
||||
|
||||
|
||||
def run(base, profile, repo, send_timeout, answer_timeout):
|
||||
"""Drive the full reverse rendezvous. Returns a result record for grading + transcript."""
|
||||
rec = {"spawned": False, "tid": None, "pane": None, "phase": "spawn",
|
||||
"question": None, "turnId": None, "reply": None, "replySource": None,
|
||||
"detail": None, "ask_latency": None, "answer_latency": None}
|
||||
|
||||
tid, pane = spawn_worker(base, profile, repo)
|
||||
rec.update(tid=tid, pane=pane, spawned=True)
|
||||
await_ready(base, tid)
|
||||
|
||||
# 1) Delegate the ask-forcing task. This blocks until the worker calls bridge_ask, at which
|
||||
# point our own send unblocks carrying the question and the turnId to answer on.
|
||||
print(f"[{now()}] delegating task (blocks until the worker asks; up to {send_timeout}s)…")
|
||||
t0 = time.time()
|
||||
rec["phase"] = "awaiting_question"
|
||||
try:
|
||||
code, resp = post_message(base, tid, {"content": TASK_PROMPT, "timeoutMs": send_timeout * 1000},
|
||||
timeout=send_timeout + 15)
|
||||
except Exception as e: # noqa: BLE001
|
||||
rec.update(phase="send_error", detail=f"delegating send failed: {e}")
|
||||
return rec
|
||||
rec["ask_latency"] = round(time.time() - t0, 1)
|
||||
|
||||
if resp.get("status") != "question":
|
||||
# The worker finished (or stalled) without asking — the whole point didn't happen.
|
||||
rec.update(phase="no_question", detail=f"HTTP {code}: {json.dumps(resp)[:300]}",
|
||||
reply=resp.get("reply"), replySource=resp.get("replySource"))
|
||||
return rec
|
||||
rec.update(phase="question", question=resp.get("question"), turnId=resp.get("turnId"))
|
||||
print(f"[{now()}] worker ASKED ({rec['ask_latency']}s): {rec['question']!r} turnId={rec['turnId']}")
|
||||
|
||||
if not rec["turnId"]:
|
||||
rec.update(phase="no_turnid", detail="question surfaced without a turnId to answer on")
|
||||
return rec
|
||||
|
||||
# 2) Answer on that exact turn. This blocks again until the resumed worker calls bridge_reply.
|
||||
print(f"[{now()}] answering '{ANSWER_COLOR}' on turn {rec['turnId']} (blocks until reply; up to {answer_timeout}s)…")
|
||||
t1 = time.time()
|
||||
rec["phase"] = "awaiting_reply"
|
||||
try:
|
||||
code, resp = post_message(base, tid,
|
||||
{"content": ANSWER_COLOR, "turnId": rec["turnId"],
|
||||
"timeoutMs": answer_timeout * 1000},
|
||||
timeout=answer_timeout + 15)
|
||||
except Exception as e: # noqa: BLE001
|
||||
rec.update(phase="answer_error", detail=f"answer send failed: {e}")
|
||||
return rec
|
||||
rec["answer_latency"] = round(time.time() - t1, 1)
|
||||
|
||||
if code == 409 or resp.get("error") == "stale_turn":
|
||||
rec.update(phase="stale_turn", detail=f"HTTP {code}: {json.dumps(resp)[:300]}")
|
||||
return rec
|
||||
if resp.get("reply") is None:
|
||||
rec.update(phase="no_reply", detail=f"HTTP {code}: {json.dumps(resp)[:300]}")
|
||||
return rec
|
||||
rec.update(phase="replied", reply=resp.get("reply"), replySource=resp.get("replySource"))
|
||||
print(f"[{now()}] worker RESUMED and replied ({rec['answer_latency']}s): {rec['reply']!r} "
|
||||
f"(source={rec['replySource']})")
|
||||
return rec
|
||||
|
||||
|
||||
def grade(rec):
|
||||
"""PASS only if the worker asked, the turn resumed, and the reply reflects the answer."""
|
||||
if rec["phase"] == "no_question":
|
||||
return "NO_ASK", "the worker finished/stalled without ever calling bridge_ask"
|
||||
if rec["phase"] in ("send_error", "answer_error", "spawn"):
|
||||
return "ERROR", rec.get("detail") or "transport error before the round-trip completed"
|
||||
if rec["phase"] == "no_turnid":
|
||||
return "NO_TURNID", "the question surfaced without a turnId — the primary could not answer"
|
||||
if rec["phase"] == "stale_turn":
|
||||
return "STALE", "answering the turn was rejected as stale (it lapsed or was already answered)"
|
||||
if rec["phase"] in ("awaiting_reply", "no_reply"):
|
||||
return "NO_RESUME", "the worker asked but never resumed to a final reply within the window"
|
||||
if rec["phase"] == "replied":
|
||||
reflected = ANSWER_COLOR.upper() in (rec["reply"] or "").upper()
|
||||
if reflected and rec["replySource"] == "reply":
|
||||
return "OK", "asked, resumed the same turn, and the reply reflected the primary's answer"
|
||||
if reflected:
|
||||
return "DEGRADED", f"reply reflected the answer but resolved via {rec['replySource']} " \
|
||||
"(worker did not call bridge_reply cleanly)"
|
||||
return "WRONG_ANSWER", f"the worker replied but did not reflect '{ANSWER_COLOR}' — " \
|
||||
f"the answer may not have reached the resumed turn: {rec['reply']!r}"
|
||||
return "WEDGE", f"unexpected terminal phase {rec['phase']}: {rec.get('detail')}"
|
||||
|
||||
|
||||
def write_transcript(out_dir, rec, meta):
|
||||
path = out_dir / "bridge_ask_transcript.md"
|
||||
g, note = grade(rec)
|
||||
with path.open("w") as f:
|
||||
f.write(f"# Live bridge_ask — reverse rendezvous — {datetime.now():%Y-%m-%d %H:%M}\n\n")
|
||||
f.write(f"One worker paused its delegated turn to ask the primary, then resumed with the "
|
||||
f"answer (profile `{meta['profile']}`). Result: **`{g}`**.\n\n")
|
||||
f.write("## Round-trip\n\n")
|
||||
f.write(f"1. **primary → worker** (delegation): the ask-forcing task.\n")
|
||||
f.write(f"2. **worker → primary** (`bridge_ask`, {rec.get('ask_latency')}s): "
|
||||
f"{rec.get('question')!r} — surfaced on the primary's blocked send as a "
|
||||
f"`question` with `turnId={rec.get('turnId')}`.\n")
|
||||
f.write(f"3. **primary → worker** (answer on that turn): `{ANSWER_COLOR}`.\n")
|
||||
f.write(f"4. **worker → primary** (`bridge_reply`, {rec.get('answer_latency')}s, "
|
||||
f"source={rec.get('replySource')}): {rec.get('reply')!r}\n\n")
|
||||
f.write(f"> **{g}:** {note}\n")
|
||||
if rec.get("detail"):
|
||||
f.write(f">\n> detail: {rec['detail']}\n")
|
||||
return path
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser(description="Live bridge_ask reverse-rendezvous test (CB-205)")
|
||||
ap.add_argument("--base", default="http://127.0.0.1:8765")
|
||||
ap.add_argument("--profile", default=None)
|
||||
ap.add_argument("--repo", default=str(REPO_ROOT))
|
||||
ap.add_argument("--out", default=str(HERE))
|
||||
ap.add_argument("--send-timeout", type=int, default=110, help="max wait for the worker to ASK")
|
||||
ap.add_argument("--answer-timeout", type=int, default=90, help="max wait for the worker to REPLY")
|
||||
ap.add_argument("--keep-worker", action="store_true")
|
||||
args = ap.parse_args()
|
||||
|
||||
print(f"[{now()}] live bridge_ask: 1 primary, 1 worker "
|
||||
f"(profile={args.profile or 'default'}, repo={args.repo})\n")
|
||||
|
||||
rec = {"spawned": False, "pane": None}
|
||||
try:
|
||||
rec = run(args.base, args.profile, args.repo, args.send_timeout, args.answer_timeout)
|
||||
finally:
|
||||
if not args.keep_worker and rec.get("spawned") and rec.get("pane"):
|
||||
try:
|
||||
http(args.base, "DELETE", f"/workers/{rec['pane']}")
|
||||
print(f"[{now()}] stopped worker (pane {rec['pane']})")
|
||||
except Exception as e: # noqa: BLE001
|
||||
print(f"[{now()}] stop failed (ignore): {e}")
|
||||
|
||||
g, note = grade(rec)
|
||||
out_dir = pathlib.Path(args.out)
|
||||
out_dir.mkdir(parents=True, exist_ok=True)
|
||||
path = write_transcript(out_dir, rec, {"profile": args.profile or "default"})
|
||||
|
||||
print()
|
||||
print("=" * 72)
|
||||
print("LIVE bridge_ask SUMMARY — reverse rendezvous (CB-205)")
|
||||
print(f" asked: {rec.get('question')!r} (turnId={rec.get('turnId')}, {rec.get('ask_latency')}s)")
|
||||
print(f" answered: {ANSWER_COLOR!r}")
|
||||
print(f" replied: {rec.get('reply')!r} (source={rec.get('replySource')}, {rec.get('answer_latency')}s)")
|
||||
print(f" transcript: {path}")
|
||||
print(f" RESULT: {g} — {note}")
|
||||
print("=" * 72)
|
||||
|
||||
sys.exit(0 if g == "OK" else 1)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,12 @@
|
||||
# Live bridge_ask — reverse rendezvous — 2026-07-16 16:30
|
||||
|
||||
One worker paused its delegated turn to ask the primary, then resumed with the answer (profile `default`). Result: **`OK`**.
|
||||
|
||||
## Round-trip
|
||||
|
||||
1. **primary → worker** (delegation): the ask-forcing task.
|
||||
2. **worker → primary** (`bridge_ask`, 6.6s): 'PICK A COLOR: red or blue?' — surfaced on the primary's blocked send as a `question` with `turnId=term_656bb47d2c42a9e#1`.
|
||||
3. **primary → worker** (answer on that turn): `blue`.
|
||||
4. **worker → primary** (`bridge_reply`, 7.9s, source=reply): 'CHOSEN=BLUE'
|
||||
|
||||
> **OK:** asked, resumed the same turn, and the reply reflected the primary's answer
|
||||
@@ -0,0 +1,170 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Sustained back-and-forth bridge test — ONE primary, ONE worker, many dependent turns
|
||||
over a fixed wall-clock window (default 5 minutes), through the running `bridged` daemon.
|
||||
|
||||
Where conversation_test.py proves a handful of turns work and issue_hunt_test.py proves
|
||||
fan-out isolation, this proves the channel stays healthy under a *sustained, stateful*
|
||||
conversation: a running-total game the worker must keep in its head across turns. Turn N's
|
||||
prompt does NOT restate the total — the worker has to remember it from turn N-1 — so a
|
||||
correct answer is evidence of genuine multi-turn continuity, not just per-turn liveness.
|
||||
Each reply is machine-checked against the primary's own expected total; on a drift the
|
||||
primary re-anchors (states the correct total once) and keeps going, and drift is reported.
|
||||
|
||||
It talks ONLY to the bridge's REST face on loopback — it never sets ANTHROPIC_BASE_URL and
|
||||
never touches herdr directly, so it is subscription-safe by construction.
|
||||
|
||||
Usage:
|
||||
python3 conversation_sustained_test.py [--base URL] [--profile NAME] [--repo DIR]
|
||||
[--duration SECS] [--turn-timeout SECS] [--keep-worker]
|
||||
|
||||
--base bridge REST base URL (default http://127.0.0.1:8765)
|
||||
--profile worker profile (default: the daemon's default)
|
||||
--repo cwd handed to the worker (default: the bridge repo root)
|
||||
--duration wall-clock window, seconds (default 300 = 5 minutes)
|
||||
--turn-timeout per-turn max wait, seconds (default 150)
|
||||
--keep-worker do not stop the worker at the end
|
||||
|
||||
Exit code: 0 if every turn in the window resolved via a clean bridge_reply with no channel
|
||||
break; 1 otherwise. A live per-turn log streams to stdout so the run can be watched.
|
||||
"""
|
||||
import argparse
|
||||
import pathlib
|
||||
import re
|
||||
import sys
|
||||
import time
|
||||
|
||||
# Reuse the exact REST primitives the other harnesses use (same daemon contract).
|
||||
sys.path.insert(0, str(pathlib.Path(__file__).parent))
|
||||
from issue_hunt_test import http, now, spawn_worker, await_ready, fire, poll # noqa: E402
|
||||
|
||||
HERE = pathlib.Path(__file__).parent
|
||||
REPO_ROOT = HERE.parent
|
||||
|
||||
# The per-turn increments, cycled. Non-trivial and varied so the running total isn't a
|
||||
# predictable multiple the worker could pattern-match without actually tracking it.
|
||||
STEPS = [7, 3, 11, 5, 9, 4, 13, 6, 8, 2]
|
||||
|
||||
RULES = (
|
||||
"Let's play a running-total game across several messages. The total starts at 0. "
|
||||
"In each message I'll tell you to add a number; keep the running total yourself and "
|
||||
"reply via bridge_reply with ONLY the current total as a plain integer — no words, no "
|
||||
"punctuation, just the number. Do not restate the arithmetic. First move: add {step}."
|
||||
)
|
||||
NEXT = ("Add {step}. Reply via bridge_reply with only the new running total.")
|
||||
REANCHOR = ("Let's re-sync — the running total is {total}. Now add {step}. Reply via "
|
||||
"bridge_reply with only the new running total.")
|
||||
|
||||
|
||||
def parse_int(reply):
|
||||
"""Pull the worker's answer integer from its reply (last integer token wins)."""
|
||||
if not reply:
|
||||
return None
|
||||
nums = re.findall(r"-?\d+", reply.replace(",", ""))
|
||||
return int(nums[-1]) if nums else None
|
||||
|
||||
|
||||
def one_turn(base, tid, prompt, turn_timeout):
|
||||
"""Fire one prompt and block on its reply. Returns the poll record."""
|
||||
t0 = time.time()
|
||||
ticket = fire(base, tid, prompt)
|
||||
rec = poll(base, ticket, "worker", t0, turn_timeout)
|
||||
rec["ticket"] = ticket
|
||||
return rec
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser(description="Sustained back-and-forth bridge test (1 primary, 1 worker)")
|
||||
ap.add_argument("--base", default="http://127.0.0.1:8765")
|
||||
ap.add_argument("--profile", default=None)
|
||||
ap.add_argument("--repo", default=str(REPO_ROOT))
|
||||
ap.add_argument("--duration", type=int, default=300)
|
||||
ap.add_argument("--turn-timeout", type=int, default=150)
|
||||
ap.add_argument("--keep-worker", action="store_true")
|
||||
args = ap.parse_args()
|
||||
|
||||
mins = args.duration / 60
|
||||
print(f"[{now()}] sustained back-and-forth: 1 primary <-> 1 worker for {args.duration}s "
|
||||
f"(~{mins:.1f} min) profile={args.profile or 'default'} repo={args.repo}", flush=True)
|
||||
|
||||
tid, pane = spawn_worker(args.base, args.profile, args.repo, "worker")
|
||||
await_ready(args.base, tid, "worker")
|
||||
|
||||
start = time.time()
|
||||
expected = 0 # the primary's authoritative running total
|
||||
reanchor = False # re-state the total next turn after a drift
|
||||
turns, oks, drifts, breaks = 0, 0, 0, 0
|
||||
latencies = []
|
||||
print(f"[{now()}] --- conversation start (worker must keep the total in its head) ---\n", flush=True)
|
||||
|
||||
while time.time() - start < args.duration:
|
||||
turns += 1
|
||||
step = STEPS[(turns - 1) % len(STEPS)]
|
||||
if turns == 1:
|
||||
prompt = RULES.format(step=step)
|
||||
elif reanchor:
|
||||
prompt = REANCHOR.format(total=expected, step=step)
|
||||
reanchor = False
|
||||
else:
|
||||
prompt = NEXT.format(step=step)
|
||||
expected += step
|
||||
|
||||
el = round(time.time() - start)
|
||||
print(f"[{now()}] turn {turns:>2} (t+{el}s) PRIMARY → add {step} (expect total {expected})", flush=True)
|
||||
|
||||
rec = one_turn(args.base, tid, prompt, args.turn_timeout)
|
||||
got = parse_int(rec.get("reply"))
|
||||
lat = rec.get("latency")
|
||||
latencies.append(lat)
|
||||
|
||||
if rec.get("phase") != "done" or not (rec.get("reply") or "").strip():
|
||||
breaks += 1
|
||||
print(f"[{now()}] WORKER ✗ CHANNEL BREAK — phase={rec.get('phase')} "
|
||||
f"source={rec.get('source')} detail={str(rec.get('detail'))[:100]} ({lat}s)\n", flush=True)
|
||||
reanchor = True
|
||||
continue
|
||||
|
||||
src = rec.get("source")
|
||||
badge = "OK " if got == expected else "DRIFT"
|
||||
if got == expected:
|
||||
oks += 1
|
||||
else:
|
||||
drifts += 1
|
||||
reanchor = True # re-sync the worker next turn
|
||||
print(f"[{now()}] WORKER → {str(rec.get('reply')).strip()[:60]!r} = {got} "
|
||||
f"[{badge}] via {src} ({lat}s)", flush=True)
|
||||
if got != expected:
|
||||
print(f"[{now()}] (expected {expected}; will re-anchor next turn)", flush=True)
|
||||
print(flush=True)
|
||||
|
||||
dur = round(time.time() - start)
|
||||
clean = sum(1 for lat in latencies if lat)
|
||||
avg = round(sum(latencies) / len(latencies), 1) if latencies else 0
|
||||
print("=" * 72)
|
||||
print(f"SUSTAINED CONVERSATION SUMMARY — 1 primary <-> 1 worker over {dur}s (~{dur/60:.1f} min)")
|
||||
print(f" turns: {turns}")
|
||||
print(f" clean bridge_reply exchanges: {oks + drifts}/{turns} (channel breaks: {breaks})")
|
||||
print(f" arithmetic correct (continuity held): {oks}/{turns} (drifts: {drifts})")
|
||||
print(f" latency: avg {avg}s over {turns} turns")
|
||||
ok = breaks == 0 and turns >= 2
|
||||
if ok and drifts == 0:
|
||||
print(" RESULT: PASS — every turn resolved via bridge_reply and the worker held the "
|
||||
"running total across the whole window.")
|
||||
elif ok:
|
||||
print(f" RESULT: PASS (channel) — every turn resolved via bridge_reply for the full "
|
||||
f"window; {drifts} arithmetic drift(s) (worker recovered after re-anchor).")
|
||||
else:
|
||||
print(" RESULT: FAIL — the channel broke on at least one turn (see CHANNEL BREAK above).")
|
||||
print("=" * 72, flush=True)
|
||||
|
||||
if not args.keep_worker and pane:
|
||||
try:
|
||||
http(args.base, "DELETE", f"/workers/{pane}")
|
||||
print(f"[{now()}] worker stopped (pane {pane})", flush=True)
|
||||
except Exception as e: # noqa: BLE001
|
||||
print(f"[{now()}] worker stop failed (ignore): {e}", flush=True)
|
||||
|
||||
sys.exit(0 if ok else 1)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,227 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Standard bridge conversation test — a multi-turn primary↔worker exchange through
|
||||
the running `bridged` daemon, fully captured, with automatic gap analysis.
|
||||
|
||||
This is the repeatable form of the ad-hoc channel test that surfaced the CB-115 gaps
|
||||
(herdr `unknown` misclassification, dirty completion scrape, workers not calling
|
||||
bridge_reply). It drives a real off-subscription worker over the live gateway exactly
|
||||
as a primary Opus session would (async fire-and-poll), records every turn, and grades
|
||||
the channel.
|
||||
|
||||
It talks ONLY to the bridge's REST face on loopback — it never sets ANTHROPIC_BASE_URL
|
||||
and never touches herdr directly, so it is subscription-safe by construction.
|
||||
|
||||
Usage:
|
||||
python3 conversation_test.py [--base URL] [--profile NAME] [--tid TERMINAL_ID]
|
||||
[--prompts FILE] [--out DIR] [--poll-timeout SECS]
|
||||
[--keep-worker]
|
||||
|
||||
--base bridge REST base URL (default http://127.0.0.1:8765)
|
||||
--profile worker profile to spawn (default: the daemon's default)
|
||||
--tid reuse an existing worker (skips spawn + readiness wait)
|
||||
--prompts newline-separated prompt file (default: the built-in script)
|
||||
--out output dir for transcript.md (default: alongside this file)
|
||||
--poll-timeout per-turn max wait, seconds (default 300)
|
||||
--keep-worker do not stop a spawned worker at the end
|
||||
|
||||
Exit code: 0 if every turn delivered AND produced a usable reply; 1 otherwise (so it
|
||||
is CI-usable). A per-turn and overall gap report is printed to stdout.
|
||||
"""
|
||||
import argparse
|
||||
import json
|
||||
import pathlib
|
||||
import sys
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from datetime import datetime
|
||||
|
||||
HERE = pathlib.Path(__file__).parent
|
||||
|
||||
# A default conversation: short, varied turns (a question, a follow-up that needs the
|
||||
# prior context, a tiny reasoning task, a meta-question, a close) — enough to exercise
|
||||
# multi-turn delivery + reply on the channel without being a real coding workload.
|
||||
DEFAULT_PROMPTS = [
|
||||
"Hi! Quick check that our channel works. In one sentence, what are you and what model are you running?",
|
||||
"Thanks. Now a small task: what is 17 * 23? Show just the number.",
|
||||
"Good. Remembering that result, is it a prime number? Answer yes or no with a one-line reason.",
|
||||
"Switching topic: name one thing that would make this bridge conversation feel more reliable to you as the worker.",
|
||||
"That's all — please acknowledge and we'll wrap up.",
|
||||
]
|
||||
|
||||
|
||||
def http(base, method, path, body=None, timeout=20):
|
||||
data = json.dumps(body).encode() if body is not None else None
|
||||
req = urllib.request.Request(base + path, data=data, method=method,
|
||||
headers={"Content-Type": "application/json"})
|
||||
with urllib.request.urlopen(req, timeout=timeout) as r:
|
||||
return json.loads(r.read().decode())
|
||||
|
||||
|
||||
def now():
|
||||
return datetime.now().strftime("%H:%M:%S")
|
||||
|
||||
|
||||
def spawn_worker(base, profile):
|
||||
q = f"?profile={profile}" if profile else ""
|
||||
res = http(base, "POST", f"/workers{q}")
|
||||
tid = res.get("terminalId") or res.get("sessionId")
|
||||
if not tid:
|
||||
sys.exit(f"spawn failed: {res}")
|
||||
pane = res.get("paneId")
|
||||
print(f"[{now()}] spawned worker {tid} (profile={profile or 'default'}, pane={pane})")
|
||||
return tid, pane
|
||||
|
||||
|
||||
def await_ready(base, tid, timeout=120):
|
||||
print(f"[{now()}] waiting for worker readiness (bridge MCP connect)…")
|
||||
deadline = time.time() + timeout
|
||||
while time.time() < deadline:
|
||||
try:
|
||||
st = http(base, "GET", f"/sessions/{tid}/status")
|
||||
except urllib.error.URLError:
|
||||
st = {}
|
||||
if st.get("ready"):
|
||||
print(f"[{now()}] worker ready (status={st.get('status')})")
|
||||
return True
|
||||
time.sleep(3)
|
||||
print(f"[{now()}] WARNING: worker never reported ready within {timeout}s — running anyway")
|
||||
return False
|
||||
|
||||
|
||||
def run_turn(base, tid, turn, prompt, poll_timeout):
|
||||
"""Fire one prompt async, poll the ticket to resolution, return a structured record."""
|
||||
t0 = time.time()
|
||||
sent = http(base, "POST", f"/sessions/{tid}/message", {"wait": False, "content": prompt})
|
||||
ticket = sent.get("ticket")
|
||||
samples = [] # (elapsed, phase, live_status)
|
||||
reply = detail = source = phase = None
|
||||
deadline = time.time() + poll_timeout
|
||||
while time.time() < deadline:
|
||||
time.sleep(3)
|
||||
task = http(base, "GET", f"/tasks/{ticket}")
|
||||
phase = task.get("phase")
|
||||
live = (task.get("detail") or "").replace("worker ", "") if phase == "pending" else ""
|
||||
samples.append((round(time.time() - t0, 1), phase, live))
|
||||
if phase in ("done", "failed"):
|
||||
reply = task.get("reply")
|
||||
detail = task.get("detail")
|
||||
source = task.get("replySource")
|
||||
break
|
||||
latency = round(time.time() - t0, 1)
|
||||
|
||||
# compress status samples into a transition string
|
||||
trans, last = [], None
|
||||
for el, ph, live in samples:
|
||||
tag = live if ph == "pending" else ph
|
||||
if tag != last:
|
||||
trans.append(f"{tag}@{el}s")
|
||||
last = tag
|
||||
|
||||
return {
|
||||
"turn": turn, "prompt": prompt, "ticket": ticket, "phase": phase,
|
||||
"source": source, "reply": reply, "detail": detail, "latency": latency,
|
||||
"transitions": " → ".join(trans), "time": now(),
|
||||
}
|
||||
|
||||
|
||||
def grade(rec):
|
||||
"""Classify a turn's outcome. Returns (grade, note)."""
|
||||
phase, source, reply = rec["phase"], rec["source"], rec["reply"]
|
||||
has_reply = bool(reply and reply.strip())
|
||||
if phase == "done" and source == "reply" and has_reply:
|
||||
return "OK", "clean explicit bridge_reply"
|
||||
if phase == "done" and has_reply:
|
||||
return "DEGRADED", f"resolved via {source} (worker did not call bridge_reply)"
|
||||
if phase == "done" and not has_reply:
|
||||
return "EMPTY", "turn completed but reply was empty"
|
||||
if phase == "failed":
|
||||
return "FAILED", f"worker turn failed: {(rec['detail'] or '').strip()[:120]}"
|
||||
return "WEDGE", "never resolved within the poll window (delivery wedge or lost turn)"
|
||||
|
||||
|
||||
def write_transcript(out_dir, records):
|
||||
path = out_dir / "transcript.md"
|
||||
with path.open("w") as f:
|
||||
f.write(f"# Bridge conversation test — {datetime.now():%Y-%m-%d %H:%M}\n\n")
|
||||
for rec in records:
|
||||
g, note = grade(rec)
|
||||
f.write(f"### Turn {rec['turn']} — {rec['time']} "
|
||||
f"(`{g}`, {rec['latency']}s, phase={rec['phase']}, source={rec['source']})\n\n")
|
||||
f.write(f"**PRIMARY:** {rec['prompt']}\n\n")
|
||||
f.write(f"**WORKER:** {rec['reply'] if rec['reply'] else '_(no reply)_ ' + str(rec['detail'])}\n\n")
|
||||
f.write(f"_status: {rec['transitions']}_\n")
|
||||
if g != "OK":
|
||||
f.write(f"\n> **GAP — {g}:** {note}\n")
|
||||
f.write("\n")
|
||||
return path
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser(description="Standard bridge conversation test")
|
||||
ap.add_argument("--base", default="http://127.0.0.1:8765")
|
||||
ap.add_argument("--profile", default=None)
|
||||
ap.add_argument("--tid", default=None)
|
||||
ap.add_argument("--prompts", default=None)
|
||||
ap.add_argument("--out", default=str(HERE))
|
||||
ap.add_argument("--poll-timeout", type=int, default=300)
|
||||
ap.add_argument("--keep-worker", action="store_true")
|
||||
args = ap.parse_args()
|
||||
|
||||
prompts = DEFAULT_PROMPTS
|
||||
if args.prompts:
|
||||
prompts = [ln.strip() for ln in pathlib.Path(args.prompts).read_text().splitlines()
|
||||
if ln.strip() and not ln.startswith("#")]
|
||||
|
||||
spawned = False
|
||||
tid = args.tid
|
||||
pane = None
|
||||
if not tid:
|
||||
tid, pane = spawn_worker(args.base, args.profile)
|
||||
spawned = True
|
||||
await_ready(args.base, tid)
|
||||
|
||||
print(f"[{now()}] running {len(prompts)}-turn conversation on {tid}\n")
|
||||
records = []
|
||||
for i, prompt in enumerate(prompts, 1):
|
||||
rec = run_turn(args.base, tid, i, prompt, args.poll_timeout)
|
||||
g, note = grade(rec)
|
||||
records.append(rec)
|
||||
print(f"[turn {i}] {g:8} {rec['latency']:6}s phase={rec['phase']} source={rec['source']}")
|
||||
print(f" status: {rec['transitions']}")
|
||||
print(f" reply: {(rec['reply'] or '(none) ' + str(rec['detail'])).strip()[:200]}")
|
||||
print(f" note: {note}\n")
|
||||
|
||||
out_dir = pathlib.Path(args.out)
|
||||
out_dir.mkdir(parents=True, exist_ok=True)
|
||||
path = write_transcript(out_dir, records)
|
||||
|
||||
# ---- gap report -------------------------------------------------------
|
||||
grades = [grade(r)[0] for r in records]
|
||||
counts = {g: grades.count(g) for g in ("OK", "DEGRADED", "EMPTY", "FAILED", "WEDGE") if grades.count(g)}
|
||||
print("=" * 68)
|
||||
print(f"CONVERSATION TEST SUMMARY — {len(records)} turns")
|
||||
print(" " + " ".join(f"{g}:{n}" for g, n in counts.items()))
|
||||
print(f" transcript: {path}")
|
||||
ok = all(g in ("OK", "DEGRADED") for g in grades)
|
||||
reply_clean = all(g == "OK" for g in grades)
|
||||
if reply_clean:
|
||||
print(" RESULT: PASS — every turn delivered and got a clean bridge_reply.")
|
||||
elif ok:
|
||||
print(" RESULT: PASS (with notes) — every turn delivered & replied, but some via fallback.")
|
||||
else:
|
||||
print(" RESULT: FAIL — one or more turns wedged, failed, or returned empty (see GAP notes).")
|
||||
print("=" * 68)
|
||||
|
||||
if spawned and pane and not args.keep_worker:
|
||||
try:
|
||||
http(args.base, "DELETE", f"/workers/{pane}")
|
||||
print(f"[{now()}] stopped worker {tid} (pane {pane})")
|
||||
except Exception as e: # noqa: BLE001 - best-effort cleanup
|
||||
print(f"[{now()}] worker stop failed (ignore): {e}")
|
||||
|
||||
sys.exit(0 if ok else 1)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,330 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Standard bridge fan-out test — ONE primary vs MANY workers, concurrently, for
|
||||
issue hunting through the running `bridged` daemon, fully captured, with gap analysis.
|
||||
|
||||
Where conversation_test.py exercises a single worker over multiple turns, this drives
|
||||
the path that only appears under fan-out: the primary spawns N workers, sends each a
|
||||
distinct issue-hunting assignment on a slice of the codebase, fires them all at once,
|
||||
and collects every reply concurrently. That stresses what a single worker never can —
|
||||
|
||||
• simultaneous delivery to many panes (the injector's per-worker, not global, writer),
|
||||
• per-session rendezvous isolation (N blocked sends resolving independently),
|
||||
• reply routing under concurrency (worker A's answer must never resolve worker B's send),
|
||||
|
||||
and, as the payload, whether a fleet of off-subscription workers can actually surface
|
||||
real issues in the repo and report them back structurally via bridge_reply.
|
||||
|
||||
It talks ONLY to the bridge's REST face on loopback — it never sets ANTHROPIC_BASE_URL
|
||||
and never touches herdr directly, so it is subscription-safe by construction.
|
||||
|
||||
Usage:
|
||||
python3 issue_hunt_test.py [--base URL] [--profile NAME] [--repo DIR]
|
||||
[--out DIR] [--poll-timeout SECS] [--keep-workers]
|
||||
[--assignments FILE]
|
||||
|
||||
--base bridge REST base URL (default http://127.0.0.1:8765)
|
||||
--profile profile for every worker (default: the daemon's default)
|
||||
--repo cwd handed to each worker (default: the bridge repo root)
|
||||
--out output dir for the transcript (default: alongside this file)
|
||||
--poll-timeout per-worker max wait, seconds (default 300)
|
||||
--keep-workers do not stop spawned workers at the end
|
||||
--assignments JSON file overriding the built-in assignment list
|
||||
|
||||
Exit code: 0 if every worker delivered AND produced a usable reply with no cross-talk;
|
||||
1 otherwise (CI-usable). A per-worker and fleet-level gap report is printed to stdout.
|
||||
"""
|
||||
import argparse
|
||||
import json
|
||||
import pathlib
|
||||
import sys
|
||||
import threading
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
from datetime import datetime
|
||||
|
||||
HERE = pathlib.Path(__file__).parent
|
||||
REPO_ROOT = HERE.parent
|
||||
|
||||
# Each worker gets a distinct source file to hunt in, plus a `probe` — a token its reply
|
||||
# should mention if it actually addressed ITS assignment (a soft cross-talk detector: a
|
||||
# reply that references only another worker's file is a routing red flag). Targets are the
|
||||
# hot files this project has been iterating on, so a real issue is plausible to find.
|
||||
DEFAULT_ASSIGNMENTS = [
|
||||
{"id": "completion", "probe": "CompletionResolver",
|
||||
"target": "bridged/src/main/java/dev/ltms/bridged/inject/CompletionResolver.java"},
|
||||
{"id": "worker", "probe": "WorkerService",
|
||||
"target": "bridged/src/main/java/dev/ltms/bridged/worker/WorkerService.java"},
|
||||
{"id": "rendezvous", "probe": "Rendezvous",
|
||||
"target": "bridged/src/main/java/dev/ltms/bridged/msg/Rendezvous.java"},
|
||||
]
|
||||
|
||||
PROMPT_TMPL = (
|
||||
"You are one of several issue-hunting workers in the claude-bridge repo (it is your "
|
||||
"current working directory). Your assignment: inspect the file `{target}` and find the "
|
||||
"SINGLE most important real bug, correctness gap, or risk in it. Read the file before "
|
||||
"answering. Reply via bridge_reply with EXACTLY these four lines:\n"
|
||||
"1. {target}:<line>\n"
|
||||
"2. issue: <one sentence>\n"
|
||||
"3. fix: <one line>\n"
|
||||
"4. severity: high|medium|low\n"
|
||||
"Keep it under 90 words. If after reading you find nothing real, reply 'NO ISSUE' and one "
|
||||
"line why. Do NOT hunt in any other file — only `{target}`."
|
||||
)
|
||||
|
||||
|
||||
def http(base, method, path, body=None, timeout=20):
|
||||
"""JSON request. Tolerates an empty body (e.g. 204 No Content on DELETE) → returns {}."""
|
||||
data = json.dumps(body).encode() if body is not None else None
|
||||
req = urllib.request.Request(base + path, data=data, method=method,
|
||||
headers={"Content-Type": "application/json"})
|
||||
with urllib.request.urlopen(req, timeout=timeout) as r:
|
||||
raw = r.read().decode().strip()
|
||||
return json.loads(raw) if raw else {}
|
||||
|
||||
|
||||
def now():
|
||||
return datetime.now().strftime("%H:%M:%S")
|
||||
|
||||
|
||||
def spawn_worker(base, profile, repo, wid):
|
||||
res = http(base, "POST", "/workers", {"profile": profile, "cwd": repo} if profile
|
||||
else {"cwd": repo})
|
||||
tid = res.get("terminalId") or res.get("sessionId")
|
||||
if not tid:
|
||||
raise RuntimeError(f"spawn failed for {wid}: {res}")
|
||||
pane = res.get("paneId")
|
||||
print(f"[{now()}] [{wid}] spawned {tid} (profile={profile or 'default'}, pane={pane})")
|
||||
return tid, pane
|
||||
|
||||
|
||||
def await_ready(base, tid, wid, timeout=150):
|
||||
deadline = time.time() + timeout
|
||||
while time.time() < deadline:
|
||||
try:
|
||||
st = http(base, "GET", f"/sessions/{tid}/status")
|
||||
except urllib.error.URLError:
|
||||
st = {}
|
||||
if st.get("ready"):
|
||||
print(f"[{now()}] [{wid}] ready (status={st.get('status')})")
|
||||
return True
|
||||
time.sleep(3)
|
||||
print(f"[{now()}] [{wid}] WARNING: never reported ready within {timeout}s — sending anyway")
|
||||
return False
|
||||
|
||||
|
||||
def fire(base, tid, prompt):
|
||||
"""Fire one async send; return its ticket (delivery is confirmed by a ticket coming back)."""
|
||||
sent = http(base, "POST", f"/sessions/{tid}/message", {"wait": False, "content": prompt})
|
||||
return sent.get("ticket")
|
||||
|
||||
|
||||
def poll(base, ticket, wid, t0, poll_timeout):
|
||||
"""Poll a ticket to resolution; return (record fields) mirroring conversation_test."""
|
||||
samples, reply, detail, source, phase = [], None, None, None, None
|
||||
deadline = time.time() + poll_timeout
|
||||
while time.time() < deadline:
|
||||
time.sleep(3)
|
||||
task = http(base, "GET", f"/tasks/{ticket}")
|
||||
phase = task.get("phase")
|
||||
live = (task.get("detail") or "").replace("worker ", "") if phase == "pending" else ""
|
||||
samples.append((round(time.time() - t0, 1), phase, live))
|
||||
if phase in ("done", "failed"):
|
||||
reply, detail, source = task.get("reply"), task.get("detail"), task.get("replySource")
|
||||
break
|
||||
trans, last = [], None
|
||||
for el, ph, live in samples:
|
||||
tag = live if ph == "pending" else ph
|
||||
if tag != last:
|
||||
trans.append(f"{tag}@{el}s")
|
||||
last = tag
|
||||
return {"phase": phase, "reply": reply, "detail": detail, "source": source,
|
||||
"latency": round(time.time() - t0, 1), "transitions": " → ".join(trans)}
|
||||
|
||||
|
||||
def grade(rec):
|
||||
phase, source, reply = rec["phase"], rec["source"], rec["reply"]
|
||||
has_reply = bool(reply and reply.strip())
|
||||
if phase == "done" and source == "reply" and has_reply:
|
||||
return "OK", "clean explicit bridge_reply"
|
||||
if phase == "done" and has_reply:
|
||||
return "DEGRADED", f"resolved via {source} (worker did not call bridge_reply)"
|
||||
if phase == "done" and not has_reply:
|
||||
return "EMPTY", "turn completed but reply was empty"
|
||||
if phase == "failed":
|
||||
return "FAILED", f"worker turn failed: {(rec['detail'] or '').strip()[:120]}"
|
||||
return "WEDGE", "never resolved within the poll window (delivery wedge or lost turn)"
|
||||
|
||||
|
||||
def worker_lifecycle(base, profile, repo, poll_timeout, a, barrier):
|
||||
"""Full per-worker path: spawn → ready → (barrier) → fire → poll. Spawn+ready run
|
||||
concurrently across workers; the send waits on the shared `barrier` so every ready
|
||||
worker fires within the same instant — the real simultaneous-delivery stress. A worker
|
||||
that fails to spawn aborts the barrier so the rest don't block forever."""
|
||||
wid = a["id"]
|
||||
rec = {"id": wid, "target": a["target"], "probe": a["probe"], "spawned": False,
|
||||
"tid": None, "pane": None, "ticket": None}
|
||||
try:
|
||||
tid, pane = spawn_worker(base, profile, repo, wid)
|
||||
rec.update(tid=tid, pane=pane, spawned=True)
|
||||
await_ready(base, tid, wid)
|
||||
except Exception as e: # noqa: BLE001
|
||||
barrier.abort() # release peers waiting on the barrier
|
||||
rec.update(phase="failed", detail=f"spawn/ready error: {e}", reply=None,
|
||||
source=None, latency=0.0, transitions="")
|
||||
return rec
|
||||
|
||||
try:
|
||||
barrier.wait(timeout=210) # all ready workers proceed together
|
||||
except (threading.BrokenBarrierError, Exception): # noqa: BLE001
|
||||
pass # a peer died or timed out — fire anyway rather than hang
|
||||
t0 = time.time()
|
||||
prompt = PROMPT_TMPL.format(target=a["target"])
|
||||
try:
|
||||
ticket = fire(base, tid, prompt)
|
||||
rec["ticket"] = ticket
|
||||
print(f"[{now()}] [{wid}] fired (ticket={ticket})")
|
||||
rec.update(poll(base, ticket, wid, t0, poll_timeout))
|
||||
except Exception as e: # noqa: BLE001
|
||||
rec.update(phase="failed", detail=f"send error: {e}", reply=None,
|
||||
source=None, latency=round(time.time() - t0, 1), transitions="")
|
||||
return rec
|
||||
|
||||
|
||||
def crosstalk_report(records):
|
||||
"""Fleet-level isolation checks: distinct tickets, distinct non-empty replies, and each
|
||||
reply addressing its OWN assigned file (probe token present). Returns (list_of_gaps)."""
|
||||
gaps = []
|
||||
tickets = [r.get("ticket") for r in records if r.get("ticket")]
|
||||
if len(tickets) != len(set(tickets)):
|
||||
gaps.append("ticket collision: two workers were handed the same ticket id")
|
||||
replies = {r["id"]: (r.get("reply") or "").strip() for r in records}
|
||||
# identical non-empty replies from distinct assignments ⇒ suspected reply misrouting
|
||||
seen = {}
|
||||
for wid, text in replies.items():
|
||||
if text and text in seen:
|
||||
gaps.append(f"identical reply from '{seen[text]}' and '{wid}' "
|
||||
f"(distinct assignments should not yield byte-identical answers)")
|
||||
elif text:
|
||||
seen[text] = wid
|
||||
# a reply that names ANOTHER worker's file but not its own ⇒ likely cross-routing
|
||||
for r in records:
|
||||
text = (r.get("reply") or "")
|
||||
if not text.strip():
|
||||
continue
|
||||
own = r["probe"] in text or pathlib.Path(r["target"]).name in text
|
||||
others = [o["probe"] for o in records if o["id"] != r["id"] and o["probe"] in text]
|
||||
if not own and others:
|
||||
gaps.append(f"worker '{r['id']}' (assigned {r['probe']}) replied about "
|
||||
f"{others} but not its own file — possible cross-routing")
|
||||
return gaps
|
||||
|
||||
|
||||
def write_transcript(out_dir, records, meta):
|
||||
path = out_dir / "issue_hunt_transcript.md"
|
||||
with path.open("w") as f:
|
||||
f.write(f"# Bridge fan-out issue-hunt — {datetime.now():%Y-%m-%d %H:%M}\n\n")
|
||||
f.write(f"One primary vs **{len(records)} concurrent workers** "
|
||||
f"(profile `{meta['profile']}`), each hunting a distinct file.\n\n")
|
||||
for r in records:
|
||||
g, note = grade(r)
|
||||
f.write(f"### `{r['id']}` — {r['target']} (`{g}`, {r.get('latency')}s, "
|
||||
f"phase={r.get('phase')}, source={r.get('source')})\n\n")
|
||||
f.write(f"**ASSIGNMENT:** find the top issue in `{r['target']}`\n\n")
|
||||
reply = r.get("reply")
|
||||
f.write(f"**WORKER {r['id']}:** {reply if reply else '_(no reply)_ ' + str(r.get('detail'))}\n\n")
|
||||
if r.get("transitions"):
|
||||
f.write(f"_status: {r['transitions']}_\n")
|
||||
if g != "OK":
|
||||
f.write(f"\n> **GAP — {g}:** {note}\n")
|
||||
f.write("\n")
|
||||
if meta["gaps"]:
|
||||
f.write("## Fleet-level gaps\n\n")
|
||||
for gp in meta["gaps"]:
|
||||
f.write(f"- {gp}\n")
|
||||
return path
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser(description="Bridge fan-out issue-hunt test (1 primary, N workers)")
|
||||
ap.add_argument("--base", default="http://127.0.0.1:8765")
|
||||
ap.add_argument("--profile", default=None)
|
||||
ap.add_argument("--repo", default=str(REPO_ROOT))
|
||||
ap.add_argument("--out", default=str(HERE))
|
||||
ap.add_argument("--poll-timeout", type=int, default=300)
|
||||
ap.add_argument("--keep-workers", action="store_true")
|
||||
ap.add_argument("--assignments", default=None)
|
||||
args = ap.parse_args()
|
||||
|
||||
assignments = DEFAULT_ASSIGNMENTS
|
||||
if args.assignments:
|
||||
assignments = json.loads(pathlib.Path(args.assignments).read_text())
|
||||
|
||||
n = len(assignments)
|
||||
barrier = threading.Barrier(n) # releases exactly when all n ready workers reach it
|
||||
print(f"[{now()}] fan-out issue-hunt: 1 primary vs {n} workers "
|
||||
f"(profile={args.profile or 'default'}, repo={args.repo})\n")
|
||||
|
||||
# Spawn + ready + fire + poll all workers concurrently; the barrier makes every send fire
|
||||
# together once all are ready, so delivery pressure hits the daemon simultaneously.
|
||||
records = []
|
||||
with ThreadPoolExecutor(max_workers=n) as ex:
|
||||
futures = [ex.submit(worker_lifecycle, args.base, args.profile, args.repo,
|
||||
args.poll_timeout, a, barrier) for a in assignments]
|
||||
for fut in futures:
|
||||
records.append(fut.result())
|
||||
|
||||
records.sort(key=lambda r: [a["id"] for a in assignments].index(r["id"]))
|
||||
|
||||
print()
|
||||
for r in records:
|
||||
g, note = grade(r)
|
||||
print(f"[{r['id']:11}] {g:8} {str(r.get('latency','?')):6}s "
|
||||
f"phase={r.get('phase')} source={r.get('source')}")
|
||||
print(f" status: {r.get('transitions') or '(none)'}")
|
||||
print(f" reply: {((r.get('reply') or '(none) ' + str(r.get('detail'))).strip()[:200])}")
|
||||
print(f" note: {note}\n")
|
||||
|
||||
gaps = crosstalk_report(records)
|
||||
meta = {"profile": args.profile or "default", "gaps": gaps}
|
||||
out_dir = pathlib.Path(args.out)
|
||||
out_dir.mkdir(parents=True, exist_ok=True)
|
||||
path = write_transcript(out_dir, records, meta)
|
||||
|
||||
grades = [grade(r)[0] for r in records]
|
||||
counts = {g: grades.count(g) for g in ("OK", "DEGRADED", "EMPTY", "FAILED", "WEDGE") if grades.count(g)}
|
||||
print("=" * 72)
|
||||
print(f"FAN-OUT ISSUE-HUNT SUMMARY — 1 primary vs {n} workers")
|
||||
print(" channel: " + " ".join(f"{g}:{v}" for g, v in counts.items()))
|
||||
print(f" transcript: {path}")
|
||||
if gaps:
|
||||
print(" FLEET GAPS:")
|
||||
for gp in gaps:
|
||||
print(f" ⚠ {gp}")
|
||||
else:
|
||||
print(" isolation: clean — distinct tickets, distinct replies, each on its own file")
|
||||
delivered = all(r.get("ticket") for r in records)
|
||||
replied = all(g in ("OK", "DEGRADED") for g in grades)
|
||||
if delivered and replied and not gaps:
|
||||
print(" RESULT: PASS — all workers delivered concurrently, replied, and stayed isolated.")
|
||||
elif delivered and replied:
|
||||
print(" RESULT: PASS (with notes) — all delivered & replied, but see FLEET GAPS.")
|
||||
else:
|
||||
print(" RESULT: FAIL — a worker did not deliver or did not reply (see GAP notes).")
|
||||
print("=" * 72)
|
||||
|
||||
if not args.keep_workers:
|
||||
for r in records:
|
||||
if r.get("spawned") and r.get("pane"):
|
||||
try:
|
||||
http(args.base, "DELETE", f"/workers/{r['pane']}")
|
||||
print(f"[{now()}] [{r['id']}] stopped (pane {r['pane']})")
|
||||
except Exception as e: # noqa: BLE001
|
||||
print(f"[{now()}] [{r['id']}] stop failed (ignore): {e}")
|
||||
|
||||
sys.exit(0 if (delivered and replied and not gaps) else 1)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
+1
-1
Submodule wiki updated: 8c37bb71c1...8e5fd01ac9
Reference in New Issue
Block a user