Compare commits

..

50 Commits

Author SHA1 Message Date
Dai Ha 3aa69a9e32 CB-401 Stage A follow-up: rename WorkerService -> ClaudeCodeLauncher
Name the first-class Claude Code adapter explicitly, per the Peer Launcher SPI:
WorkerService was the de-facto Claude-Code launcher; as an in-tree PeerLauncher impl
it should say so. Pure IDE rename (class + file + WorkerServiceTest) plus stale
Javadoc/comment mentions swept to the new name. No behaviour change.

Gate: IDE diagnostics 0/0 on touched files; mvn clean install BUILD SUCCESS,
MVN_EXIT=0, 183 tests pass. Deferral #1 from issue #3 cleared.
2026-07-18 14:39:56 +02:00
Dai Ha e056c7e1fa CB-401: Stage A - extract PeerLauncher SPI in-tree 2026-07-18 07:48:55 +02:00
Dai Ha d63273d082 CB-401: Peer Launcher SPI design note (Stage 4)
Design-only. Audits the as-built Claude/herdr coupling (concentrated in
WorkerService), defines a PeerLauncher SPI + opaque PeerHandle + capability
model so the bus delegates peer materialization to a config-selected adapter.
ClaudeCodeLauncher = adapted WorkerService. Stages A/B/C with the Stage-C
plugin-loading security gate called out. No production code touched.
2026-07-17 17:49:21 +02:00
Dai Ha 0efb65ca0a CB-303: session lifecycle limits — idle_ttl reaper, context_cap, graceful drain
Verified on primary: ide_diagnostics clean (incl. weak warnings), mvn clean install
BUILD SUCCESS, 173 tests. Delegated impl (worker/cb-303-80ec1a-3, 3 parts), primary-gated.
2026-07-17 10:08:55 +02:00
Dai Ha 9fe04bfb08 CB-304: bridge_list roster + live herdr join (surface worktree/branch); add GET /workers
Verified on primary: ide_diagnostics clean (incl. weak warnings), mvn clean install
BUILD SUCCESS, 165 tests. Delegated impl (worker/cb-304-bd1e4f-2), primary-gated.
2026-07-17 10:08:46 +02:00
Dai Ha 09d3948acf CB-303 part 3: graceful drain on shutdown 2026-07-17 09:59:57 +02:00
Dai Ha 954351a80b CB-303 part 2: context_cap turn budget 2026-07-17 09:56:46 +02:00
Dai Ha 8d51066ddd CB-303 part 1: idle_ttl session reaper (injectable clock + SessionReaper) 2026-07-17 09:52:40 +02:00
Dai Ha 84102baab4 CB-304: bridge_list roster + live herdr join (worktree/branch); add GET /workers 2026-07-17 09:50:12 +02:00
Dai Ha 64e70efdf1 CB-302: worker checkpoint — repo-scoped forge token injection + implementer skill
The worker "checkpoint" is commit → push → open its own PR. Push is free over SSH
(same user, same keys); the only incremental grant is PR-create, so the daemon injects
a repo-scoped gitea token into the worker env — opt-in per profile, never mutating
bridged's own environment.

- BridgedConfig.Worker: gitTokenEnv/gitHostEnv fields (opt-in; gitHostEnv defaults to
  GITEA_HOST). Backward-compat 12-arg constructor keeps pre-CB-302 call sites + YAML
  working. hasGitToken() gates injection.
- WorkerService.spawn: inject GITEA_TOKEN (and paired GITEA_HOST) only when the profile
  grants a token AND the host env resolves one. resolveEnv() tolerates unset var names.
- WorkerServiceTest: injection present for a granting profile; absent when not (proving
  the gate is config, not a missing env var).
- .claude/skills/implementer/SKILL.md: worktree-aware playbook — confirm the worktree/
  branch, implement, commit (never .mcp.json/wiki), push, open PR via gitea REST with
  GITEA_TOKEN, hand off the PR URL via bridge_reply. Never merge; workers can't run IDE
  diagnostics so never claim IDE-clean.

Whole-project gate: mvn clean install green, 164 tests, 0 failures.
2026-07-17 08:51:27 +02:00
Dai Ha 97ecc7136e CB-301-ext: per-worker git worktree + config-parity overlay
Opt-in isolated worktree so parallel implementers don't stomp the shared
tree, hydrated to config parity so a worker differs from the primary only
in LLM provider.

- Worktrees seam (interface) behind SessionManager; GitWorktrees shells git
  via ProcessBuilder (non-zero exit -> WorktreeException), FakeWorktrees for
  tests. No live git in unit tests.
- acquire() 5-arg overload provisions add -> overlayParity -> spawn(cwd=wt)
  -> register, unwinding the worktree on any failure before registration.
  4-arg overload and shared-tree behavior unchanged (backward compatible).
- release() removes the checkout but never deletes the branch (it holds the
  worker's commits + PR, CB-302).
- overlayParity copies local config (.mcp.json, settings.local.json, .env/
  .envrc) into the worktree; tracked ones get --skip-worktree so a worker
  can never stage the parity overlay.
- WorkerSession gains nullable worktree/branch; BridgedConfig.Worker gains
  parityOverlay (default list) + top-level worktreeRoot.
- bridge_spawn / POST /workers gain an optional worktree(+ticket) arg; the
  worker view includes worktree/branch only when non-null.

Verify fixes on the delegated impl: strip trailing dashes in slug()
(^-+|-+$, was ^-+|^-+$); make FakeWorktrees.add a pure fn of the branch
(nonce already unique); MCP worktreeRequest treats blank/"false" string as
no-worktree, matching the REST builder.

162 tests, 0 failures.
2026-07-17 06:46:20 +02:00
Dai Ha f9073e2320 docs: CB-301-ext spec — worktree provisioning + config-parity overlay
Opt-in per-acquire worktree (shared-tree default preserved). git behind a
Worktrees seam (ProcessBuilder impl, fake in tests). acquire provisions
worktree+branch, overlays local config (copy + --skip-worktree on tracked
files so the worker can't commit .mcp.json), spawns with cwd=worktree.
release removes the worktree but keeps the branch (holds commits/PR).
WorkerSession gains nullable worktree/branch; BridgedConfig.Worker gains
parityOverlay. 6 fake-based acceptance tests incl. backward-compat + unwind.
2026-07-17 06:32:09 +02:00
Dai Ha 54d907c314 CB-301: SessionManager — authoritative worker session registry + one-shot FSM
Adds dev.ltms.bridged.session with WorkerSession (immutable record) and
SessionManager wrapping WorkerService: a ConcurrentHashMap registry keyed by
paneId, the one-shot lifecycle FSM (SPAWNING->READY->BUSY->DONE, ->FAILED on
drop/turn-failure, ->RELEASED on teardown), ownership (ownerTerminal), and
recycle = release + fresh acquire (no-reuse invariant). Driven by TurnListener
(BUSY/DONE/FAILED) and a WorkerPresence bridge (READY).

Wiring: Bridged.main constructs it and composes it into the TurnListener
alongside CompletionResolver; bridge_spawn / POST /workers route through
acquire (carrying caller identity as owner); bridge_stop / DELETE /workers
route through release. WorkerService gains effectiveCwd(); WorkerPresence
de-finalized so the manager can present a READY-driving view.

asPresence() returns a single cached bridge (a fresh one per call would
fragment the shared present set). roster() is the registry snapshot; the live
herdr join is left for CB-304. 6 fake-based acceptance tests; full suite green
(155/155).

Delegated to an off-subscription worker against docs/CB-301-Session-Manager.md;
primary verified (ide diagnostics clean, mvn clean install green) + fixed the
asPresence caching bug.
2026-07-16 19:30:56 +02:00
Dai Ha 82c7d6553a docs: worker git workflow — daemon worktree + config parity + worker-opened PR
Worktree is code-only isolation; SessionManager hydrates it to full config
parity (overlay untracked local settings/.mcp.json/.env) so a worker is a
full peer of the primary, differing only in the LLM provider. Worker commits,
pushes over SSH, and opens its own PR (gitea REST + repo-scoped token).
2026-07-16 19:25:41 +02:00
Dai Ha 19b10e3216 docs: CB-301 session-manager design spec (one-shot, no reuse; recycle in scope)
The as-built audit surfaced that WorkerService keeps no registry of what it
spawned (its own Javadoc: 'there is no registry; list() only asks herdr').
CB-301 adds a SessionManager wrapping WorkerService: an authoritative in-daemon
roster with a per-session lifecycle FSM (SPAWNING/READY/BUSY/DONE/RELEASED/
FAILED), deterministic release, and recycle (= release + fresh acquire, no
reuse). Leaves clean seams for CB-302 (checkpoint on release), CB-303 (idle_ttl/
context_cap/drain policy over roster), CB-304 (bridge_list reads roster).
2026-07-16 19:10:13 +02:00
Dai Ha 0f79e6bed5 wiki: bump submodule to page-9 as-built implementation architecture (8e5fd01) 2026-07-16 16:45:31 +02:00
Dai Ha aa0cf814ff e2e: capture bridge_ask transcript confirming single turnId after coalescing
Post-4aa9d03 live run: the duplicate-ask retry now coalesces onto one turn
(turnId=term_...#1, was #2 before the fix). Clean round-trip, RESULT OK.
2026-07-16 16:32:30 +02:00
Dai Ha 4aa9d03da2 CB-205: coalesce duplicate bridge_ask calls onto one turn
A worker's tool execution is single-threaded, so two bridge_ask calls from the
same session can only be a transport retry — yet openAsk minted a fresh turnId
each time and both raced to resolve the primary's single forward waiter, the
loser returning NO_WAITER and leaving a dangling turn (observed as turnId #2 in
the live e2e). Now openAsk is idempotent per session: a second open ask coalesces
onto the existing turnId + answer future (fresh=false), and only the fresh owner
surfaces the question and tears the turn down. closeAsk clears the per-session
index (conditional by value) so a later ask reopens fresh.

Implemented by an off-subscription worker delegated over the bridge; reviewed and
validated on the primary (IDE-clean, mvn clean install 149 green, +2 tests:
concurrent double-ask coalescing + post-close reopen).
2026-07-16 16:27:10 +02:00
Dai Ha 37abfdfd6e wiki: bump submodule to Stage-2-complete roadmap update (0c81b51) 2026-07-16 16:15:15 +02:00
Dai Ha d5c3ede215 CB-202: reviewer-role skill for bridged workers
The playbook a reviewer worker loads when the lead delegates a scoped review:
read the whole scope before judging, stay in the assigned lane, ask the lead
via bridge_ask when the call is genuinely theirs (resuming the same turn with
the answer), and report exactly one structured finding via bridge_reply. Pairs
the already-shipped bridge_reply/bridge_ask tools with the role guidance that
tells a worker how to use them. Mermaid validated with mmdc.
2026-07-16 16:10:36 +02:00
Dai Ha 426855e378 e2e: live bridge_ask reverse-rendezvous harness (CB-205)
Drives the reverse path end to end over REST loopback: a worker is delegated
a task it cannot finish without asking, calls bridge_ask mid-turn, and the
primary answers on the surfaced turnId so the worker resumes the SAME turn.
Two blocking sends, no polling. Verified live: worker asked in ~9s, resumed
and replied CHOSEN=BLUE via clean bridge_reply after the primary answered.

Subscription-safe by construction (REST face only; never sets ANTHROPIC_BASE_URL).
2026-07-16 16:08:21 +02:00
Dai Ha 358c6970b5 CB-205/CB-201: bridge_ask reverse rendezvous + lightweight question/turn kind
A worker can now pause its delegated turn to ask the primary a question and
resume the same turn with the answer — the reverse of bridge_send.

- Rendezvous: QUESTION kind carrying a turnId, plus a reverse-ask registry
  (openAsk/resolveQuestion/askSession/answerAsk/closeAsk).
- MessageService.ask(): surface a worker's question to the primary's open send,
  block for the answer. answer(): resolve the worker's ask by turnId, then block
  for its eventual bridge_reply (session derived from turnId, not an argument).
- BridgeMcp: bridge_ask tool (worker-only, identity from the connection);
  bridge_send routes a turnId to the answer path. Shared formatReply().
- BridgedApp REST parity: POST /sessions/{id}/ask, turnId on /message.
- Lightweight CB-201: the kind vocabulary is the QUESTION outcome + turn_id
  correlation, not a rigid from/to/corr envelope (the connection-identity
  mechanism already covers addressing more robustly).

Tests: MessageServiceTest ask/answer round-trip + NO_WAITER/timeout/stale;
new RendezvousTest for the reverse registry; BridgeMcp ask/answer parity.
mvn: 147 green.
2026-07-16 15:32:56 +02:00
Dai Ha 55ebd5b949 e2e: sustained back-and-forth conversation harness (5-min stateful continuity) 2026-07-16 15:19:21 +02:00
Dai Ha 2a61fe69f1 CB-118: clip the completion baseline so the CB-115 guard survives >cap blocks
captureBaseline stored the raw, unclipped last-assistant block while resolve()
compares against clip(...) capped at MAX_SCRAPE_CHARS. For a block longer than
4000 chars the two capped representations never match even when the pane is
unchanged, defeating the CB-115 misattribution guard and letting a stale
completion resolve a rapid back-to-back send. Clip the baseline identically.

Regression test: an unchanged >cap block stays suppressed.

Surfaced by the fan-out issue-hunt E2E (1 primary -> 3 concurrent workers,
e2e/issue_hunt_test.py, added here). The same hunt's WorkerService.stop() and
Rendezvous.complete() findings were verified as false positives (locatePane is
already guarded; the sender's finally-close already removes the waiter).

Closes #2
2026-07-16 09:11:20 +02:00
Dai Ha ff6aacdc78 CB-117: reap orphaned worker panes on startup
herdr keeps worker panes alive across a daemon restart by design, and a
worker's paneId is held only by its spawner — so a worker whose owning
process exited before its DELETE leaks with nothing tracking it (there is
no registry; list() only asks herdr). Observed as three idle claude-ollama
panes left in the worker space from earlier runs.

On boot, WorkerService.reapOrphanWorkers() scans herdr for agents whose
name matches our claude-<profile>-<nonce>-<seq> scheme with a nonce other
than this process's nameNonce, and tears each down (pane + its now-empty
dedicated tab). A current-nonce worker is ours and live (spared); a user's
own claude session carries no such name (untouched). Keyed on the nonce so
it survives kill -9 and reaps a *previous* daemon's leaks — the actual case
shutdown-hook reaping and an in-memory registry both miss.

- Agent now projects herdr's 'name' (was dropped) so the reaper can key on it.
- isForeignWorker/workerNonce are pure + package-private for unit testing.
- FakeHerdr.withAgent seeds named agents into agent.list.
- Wired best-effort into Bridged startup before serving.

Closes lms/claude-bridge#1
2026-07-16 08:39:37 +02:00
Dai Ha 5f0ec034d9 e2e: standard bridge conversation test harness
A repeatable multi-turn primary↔worker conversation driven entirely through the
bridge's loopback REST face (async fire-and-poll) — never sets ANTHROPIC_BASE_URL
and never touches herdr, so it is subscription-safe by construction. Records every
turn to a transcript, grades each (OK / DEGRADED / EMPTY / FAILED / WEDGE), and
exits non-zero if any turn fails to deliver-and-reply, so it is CI-usable.

This is the repeatable form of the ad-hoc channel test that surfaced the CB-115 and
CB-116 gaps.
2026-07-16 08:27:50 +02:00
Dai Ha 31e34d177b CB-115/CB-116: reliable turn completion — status refinement, clean scrape, waiter identity
A 5-turn primary↔worker conversation test (see the e2e harness) surfaced three
delegation-channel gaps; this closes them.

CB-115 — status + scrape correctness:
- AgentStatus gains DONE (herdr's explicit turn-complete marker) so a finished
  turn is no longer misread as UNKNOWN and left to wedge or false-fail.
- StatusRefiner reclassifies content-bearing UNKNOWN samples (StatusPoller wired
  to it), and the Injector baselines pane content on delivery (TurnListener gains
  onDelivered) to guard completion against previous-turn misattribution.
- CompletionResolver.lastAssistantBlock stops at the first hard TUI boundary, so a
  scrape returns only the assistant answer — no input box, prompt echo, spinner,
  tips or warnings.

CB-116 — waiter identity (the cross-turn stale reply):
- The completion/failure fallback ran on a virtual thread and resolved whichever
  waiter was currently registered for the session. Since the rendezvous holds one
  waiter per session and sends serialize, turn N's late completion could land on
  turn N+1's waiter and deliver turn N's stale scrape as turn N+1's answer. The
  baseline guard missed it because turn N was resolved by bridge_reply, which never
  updates the completion baseline.
- Fix: capture the exact waiter (and pre-turn baseline) when a turn is delivered,
  on the poller thread before any next-turn delivery can overwrite it, and resolve
  THAT waiter — a no-op if it was already resolved. Rendezvous.resolveCompletion/
  resolveFailure now take the captured CompletableFuture; currentWaiter exposes the
  registered one for capture. A late completion for turn N can no longer touch turn
  N+1's send.

Verified: 128 unit tests green (incl. a CB-116 regression asserting a late
completion never resolves the next turn's waiter); a re-run of the conversation
test passes with turn 5 resolving to its own reply rather than turn 4's text.
2026-07-16 08:27:38 +02:00
Dai Ha b2d85af78b wiki: bump submodule to roadmap Stage-1-complete update (4304dc4) 2026-07-16 06:49:27 +02:00
Dai Ha 968a5c68b6 docs(README): update Status to reflect the shipped implementation
The Status section still read 'Design selected' — but bridged is built and
dogfooded (CB-101..114, 105 tests, live MCP tools, multi-profile, cwd inheritance,
readiness gate). Replace it with an accurate shipped/next breakdown, and drop the
bridge_ask overclaim from the gateway bullet (bridge_ask is roadmap, not built).
2026-07-16 06:46:40 +02:00
Dai Ha 3c05823f19 CB-114: resolveCwd never returns null (honor the 'daemon cwd, never $HOME' contract)
Final review-sweep finding: firstNonBlank(requestedCwd, cfg.cwd(), callerCwd,
user.dir) returns null if all are blank (pathological env with user.dir unset),
after which AgentControl drops the cwd and herdr defaults the pane to $HOME —
violating CB-112's documented contract. Append "." (the daemon's own cwd) as a
guaranteed non-blank last resort. Near-impossible trigger; makes the code honor its
own javadoc. 105 tests green.
2026-07-16 06:43:05 +02:00
Dai Ha 2fb46f670f CB-114: readiness-gate timeout + presence cleanup (delegated review findings)
An off-sub worker's review of CB-113 (delegated through the bridge) surfaced two
real gaps in the readiness gate:

- A worker that herdr reports idle but whose Claude never connects the bridge MCP
  (crashed during boot, or wedged on a startup prompt) left its message queued
  forever: ready.test() never passed, the target was polled indefinitely, and the
  caller's future never completed (async waiter hung for the full 30-min window).
  Injector now counts injectable-but-not-ready samples and, after a ~60s grace
  (READINESS_GRACE_POLLS, deliberately longer than the UNKNOWN stall grace since a
  first boot is slower than an in-turn blip), fails the queued messages, fires
  onTurnFailed so blocking/async waiters resolve WORKER_FAILED, and reclaims the
  target. Mirrors the CB-109 UNKNOWN-stall path.

- WorkerPresence.forget had no caller, so a worker's readiness lingered past its
  life. Injector now clears presence via a forget callback on drop() (pane crash)
  and on the readiness timeout.

+3 InjectorTest cases (never-ready failure, ready-within-grace delivery, drop clears
presence). 105 tests green.
2026-07-16 05:01:58 +02:00
Dai Ha 37f7ad4185 CB-113: reliable worker readiness gate + submission nudge
Delegating right after spawn failed: herdr reports 'idle' during the worker's
boot, so the injector delivered into a not-ready TUI (paste lost) and wedged the
worker. Two fixes:

- Readiness gate: a worker is 'available' only once its Claude connects the bridge
  MCP (the daemon observes it via peer-PID->terminal). WorkerPresence tracks it; the
  injector holds the first delivery until present, so it never pastes into the boot
  window. Exposed as 'ready' on GET /sessions/{id}/status.
- Submission nudge: the Enter accompanying a delivery can race the paste (esp. right
  as the TUI becomes ready), leaving text unsubmitted. While a delivered message
  stays idle (not picked up), the injector re-sends Enter each poll until the worker
  starts (WORKING) or the grace expires.

Validated live: spawn + immediate delegate now holds during boot (ready=false),
delivers on MCP-connect, re-nudges Enter, worker replies. 102 tests green.
2026-07-15 19:14:57 +02:00
Dai Ha 926724a279 CB-112: workers inherit the primary's working directory (not $HOME)
A worker now opens the same directory the primary is in, unless told otherwise.
Resolution: explicit spawn cwd → per-profile config cwd → the primary's cwd
(auto-detected from the bridge_spawn caller's PID via lsof -d cwd) → the daemon's
cwd. Never $HOME.

Mechanism (found by live probe, corrects the earlier assumption): an agent.start
pane does NOT inherit its tab's or workspace's cwd — it starts in $HOME. herdr's
agent.start honours an (undocumented) cwd param, so the resolved cwd is threaded
onto agent.start {cwd} (both tab and pane placement), not tab.create.

Surfaces: bridge_spawn {cwd?} + auto-detect via ConnectionIdentity.resolve (peer
PID) + ProcessCwdLookup (lsof); REST POST /workers ?cwd= / body cwd; per-profile
'cwd:' config. Validated live: explicit cwd → worker rooted there; no cwd over
REST → daemon cwd, not $HOME. Also clears the ccs folder-trust prompt when the
project dir is already trusted (see docs/Worker-Startup-and-Trust.md).
2026-07-15 16:33:46 +02:00
Dai Ha 959f04bc96 docs: worker startup — working directory & the folder-trust prompt
Documents (a) the rule that a worker inherits the PRIMARY's working directory,
never $HOME — resolved from an explicit cwd, else the bridge_spawn caller's PID
(lsof -d cwd), else the daemon cwd; herdr's seam is workspace.create {cwd}, since
agent.start has no cwd; (b) ccs (Claude Code) as the assumed launcher and its
folder-trust model (hasTrustDialogAccepted per project in each instance's
.claude.json), so trusting the project dir once per profile clears the prompt;
(c) that other CLIs have their own startup gates, documented per launcher.
Marks the cwd-inheritance as target design (follow-up), not yet wired.
2026-07-15 16:11:18 +02:00
Dai Ha 7f1b6b3a0d CB-111: multi-profile workers — named backends selectable at spawn
The bridge was hard-wired to one worker profile. Config now takes a 'workers'
map keyed by profile name plus 'defaultWorker'; WorkerService holds the map and
gains spawn(profile) (spawn() uses the default). Selection threads through the
surfaces: REST POST /workers ?profile= / {"profile":…} + GET /profiles; MCP
bridge_spawn {profile?} + new bridge_profiles. Each profile's base_url is
guard-checked independently, so gx10 and ollama can run side by side and you
address each worker by its returned sessionId. Backward-compatible: the legacy
singular 'worker:' block still loads as a one-entry profile map.
2026-07-15 15:54:31 +02:00
Dai Ha ab771ea24e CB-110: cover drop of a delivered-but-unpicked-up turn
Adds the drop test for the delivered/awaitingPickup/!turnObserved state (worker
vanished after delivery but before a WORKING sample) — a gap the other two drop
tests missed. Surfaced by an off-sub worker's code review of CB-110, delegated
through the bridge itself.
2026-07-15 15:30:20 +02:00
Dai Ha a629a7ee73 CB-110: fail an in-flight delegation when its worker vanishes
Companion to CB-109. When a worker disappears mid-turn (pane crash → herdr
*_not_found), the status poller drops the target, which failed only *queued*
messages — a message already DELIVERED is out of the queue, so its send's
rendezvous was left hanging until the 30-min async timeout. Injector.drop now
fires onTurnFailed for a delivered-but-unresolved turn (awaitingCompletion), so
the send resolves as WORKER_FAILED. Reuses the CB-109 resolver path; the failure
reason is neutral to cover both wedge (stuck) and vanish (gone).
2026-07-15 15:23:17 +02:00
Dai Ha 2052929768 CB-109: fail a delegation whose worker wedges in an unknown state
Dogfood found the gap: a turn that dies into an error screen herdr reports as
'unknown' (e.g. the worker hitting API ENOTFOUND) never produces a working->idle
boundary, so CB-106 never fires and the async send rides its full 30-min timeout.

The injector now counts consecutive 'unknown' samples while a delegation is
outstanding; any working/idle sample resets the streak, so only a genuine wedge
(~30s continuous unknown) trips it. It then fires TurnListener.onTurnFailed;
CompletionResolver scrapes the error screen and resolves the send via
Rendezvous.resolveFailure (Kind.FAILED -> Outcome.WORKER_FAILED), surfaced as
async phase=failed / REST status=failed / a [worker failed] MCP note, with the
error context as the reason. This also frees a delivery that wedged before pickup,
which the injectable-only pickup grace could never release.
2026-07-15 15:16:55 +02:00
Dai Ha a97c287aee CB-108: fleet-management MCP tools (bridge_spawn / bridge_list / bridge_stop)
The primary could delegate to a worker but not create or reap one over MCP —
spawning was a raw REST POST /workers. BridgeMcp now adapts WorkerService so a
worker's whole lifecycle runs through MCP: bridge_spawn returns the new worker's
sessionId (for bridge_send) and paneId (for bridge_stop); bridge_list projects the
tracked workers; bridge_stop tears one down. The subscription boundary stays
enforced inside WorkerService (bridge_spawn surfaces a guard breach as a tool error
without touching herdr). Tools are thin static adapters, unit-tested by parity.
2026-07-15 14:22:50 +02:00
Dai Ha 8ed2370fbe CB-107: async fire-and-poll delegation (wait:false + ticket poll)
A caller's MCP client caps a blocking bridge_send at ~60s, but a real delegated
task runs for minutes. sendAsync runs the same blocking send on a background
virtual thread and returns a ticket; poll(ticket) reports pending/done/failed.
Async reuses the blocking path (and its per-target serialization), so it inherits
reply + completion resolution for free. Surfaces: REST POST message wait:false ->
202 {ticket} + GET /tasks/{ticket}; MCP bridge_send wait flag + new bridge_poll.
Terminal tickets are pruned after a TTL so the registry stays bounded.
2026-07-15 14:19:14 +02:00
Dai Ha 5b26caca0c CB-106: completion fallback — resolve a send when the worker's turn ends without bridge_reply
The blocking send previously resolved only on an explicit bridge_reply; a real
delegated task (edit files, run a build) finishes and goes idle without ever
calling it, so the send always timed out. The injector now reports a confirmed
working -> idle turn boundary via a TurnListener; CompletionResolver scrapes the
worker's transcript tail and resolves the awaiting send (Rendezvous.resolveCompletion,
Kind.COMPLETION -> Outcome.COMPLETED_UNREPLIED), surfaced as replySource=transcript
at the REST/MCP edges. Completion is synthesized only from a confirmed turn (a
sampled 'working'), never from the pickup-grace path, so it can't race the explicit
reply or fire on a turn that never ran.
2026-07-15 14:14:04 +02:00
Dai Ha 6988bfe88f CB-103: injector submits the prompt (Enter as a separate keystroke)
A delegated message was delivered into the worker's input box but never
submitted, so no task was ever processed: bridge_send blocked until timeout
while the worker sat idle with the prompt typed but not entered.

herdr's agent.send delivers text as a bracketed paste; a trailing carriage
return in that same call is swallowed as literal newline content, not Enter.
AgentControl.send now emits two keystroke events — the payload, then a
standalone "\r" — so the worker actually submits and runs the task.

Verified live (Claude Code v2.1.210, off-sub worker): full hands-off
bridge_send -> auto-submit -> worker computes -> bridge_reply round trip
returns the reply in ~41s. 62 tests green.
2026-07-15 09:40:18 +02:00
Dai Ha 07722a6007 CB-1xx: worker launch via ccs + inline bridge MCP mount (step 4) + port 8765
Spawn a real worker with 'ccs ltms-local' (profile sets CLAUDE_CONFIG_DIR + off-sub base_url; auto mode preconfigured as defaultMode:auto). bridged appends the bridge MCP mount (--mcp-config, inline JSON) and the reply charter (--append-system-prompt) as launch FLAGS — non-invasive, nothing written to the worker's profile (safer than provisioning its config dir, which would clobber it). Identity is connection-based so the mount is shared. config: worker.mcpUrl. Default bind port 8080 -> 8765. The reply charter is guidance; the send timeout catches a non-cooperative worker. 60 tests green, IDE-clean.
2026-07-15 09:19:31 +02:00
Dai Ha fa570ab32a CB-105: connection-based MCP caller identity (peer PID → herdr pane)
Resolve who is calling an MCP tool from the connection, not a spoofable argument (per docs/MCP-Contract.md). ConnectionIdentity ties the loopback peer PID (LsofPeerPidLookup) to a herdr pane (PaneLocator via pane.list/pane.process_info) → the caller's terminal_id; a caller owning no pane is the primary. bridge_reply now takes only content and resolves the worker from the connection (no sessionId). Live-verified: real PID→pane (contract test), and a non-pane MCP caller correctly gets a workers-only error. Single-host; the token path stays the split-host fallback. 58 tests green, IDE-clean.
2026-07-15 09:19:31 +02:00
kevin 46f87b9cd3 wiki: sync docs with MCP contract (idle-edge turn-done, bridge_list/status naming)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 20:01:23 +07:00
Dai Ha 8ffbfd1f06 CB-105: MCP server (bridge_send/bridge_reply/bridge_status) over the REST core
Streamable-HTTP MCP server (io.modelcontextprotocol.sdk:mcp 2.0.0) mounted on the daemon's Jetty at /mcp, exposing three tools as thin adapters over MessageService/Rendezvous — the primary calls bridge_send/bridge_status, the worker calls bridge_reply. Tool logic in unit-testable static methods (parity tests); SDK owns the wire protocol. Resolves the Jackson 2/3 split by pinning jackson-annotations 3.0-rc5 (works for both our Jackson 2.19 and the SDK's Jackson 3). Live-verified: initialize handshake + tools/list return all three tools. Documents the residual Jackson-3 CVE (loopback, trusted clients). 52 tests green, IDE-clean.
2026-07-14 14:21:11 +02:00
Dai Ha 597ac2e562 CB-104: blocking bridge_send with rendezvous reply (POST /sessions/{id}/message + /reply)
Per-session-serialized blocking send that enqueues via the CB-103 injector (poller delivers) and blocks on a rendezvous resolved by the worker's structured bridge_reply, or a typed 200/202 outcome. No status-polling completion, no terminal scrape. GET /sessions/{id}/status. Reworked from an initial poll+scrape draft after a high-effort review found the polling completion unreliable; all findings fixed.
2026-07-14 14:21:11 +02:00
Dai Ha bc04637694 deps: bump to latest patched (jackson 2.19.0, javalin 6.7.0, jetty 11.0.25, logback 1.5.18)
Clears jetty CVE-2024-8184/CVE-2024-6763. Pins all Jetty modules via jetty-bom (no skew). Documents residual advisories with no upstream fix (jetty-http 11.x EOL, logback config-file CVEs, jackson WS-2026-0003) as accepted for this loopback daemon. CLAUDE.md: validate CVEs with the jetbrains analyzer (intellij-index is stale after pom edits).
2026-07-14 14:21:11 +02:00
kevin c7b58f2195 docs: add Team.md under docs/
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 19:14:25 +07:00
kevin 01cfda7965 Add MCP design doc 2026-07-14 12:30:20 +07:00
75 changed files with 9452 additions and 324 deletions
+126
View File
@@ -0,0 +1,126 @@
---
name: implementer
description: Implementer-role playbook for a bridged worker — you are in an isolated git worktree on a dedicated branch; implement the assigned task, commit, push, open your own PR to main, and hand off the PR URL via bridge_reply. You never merge. Load this when you have been delegated an implementation task over bridged.
---
# Implementer worker
You are an **implementer** in the claude-bridge fleet. The lead delegated you one scoped task,
and you are running in an **isolated git worktree on your own branch** — a full peer of the
primary (same `CLAUDE.md`, skills, memory, MCP), differing only in the model behind you and the
branch you sit on. Your job for this turn: **implement the task, then hand off a PR the lead can
review and merge.** You do the work; the lead (or human) is the merge gate — you never merge.
Delivery mechanics (how the task reached you, how your reply resolves the lead's blocked send)
are in [`docs/MCP-Contract.md`](../../../docs/MCP-Contract.md); the worktree/PR model is in
[`docs/Worker-Git-Workflow.md`](../../../docs/Worker-Git-Workflow.md). You only need the steps
below.
## 1. Confirm where you are — a worktree on a dedicated branch
Before touching anything, verify your ground truth:
```bash
git rev-parse --show-toplevel # your worktree root — NOT the primary's main tree
git branch --show-current # your dedicated branch: worker/<ticket>-<nonce>
git status # should be clean at the start
```
Do **all** work here, on this branch. **Never** switch to `main`, never `git checkout main`,
never rebase onto or push to `main` directly. The branch is your isolation — respect it.
## 2. Implement the task
- Implement exactly the scope the lead named. Keep changes focused; if you notice something out
of scope, note it in your reply rather than expanding the diff.
- Match the surrounding code's style, naming, and idioms. Follow project `CLAUDE.md`.
- **You cannot run the IDE MCP tools** (intellij-index / jetbrains are the primary's, not yours).
So **never claim a file is "IDE-clean" or "diagnostics-clean"** — you cannot verify that. State
only what you actually ran (e.g. `mvn`, a test) and its real output. A fabricated clean claim is
worse than an honest "I could not verify inspections here."
- Run whatever build/test you can and **report the true result** — including failures.
## 3. Commit — focused, and never the excluded files
```bash
git add <the files you changed>
git commit -m "<ticket>: <clear one-line summary>"
```
**Excluded from every commit, always:** `.mcp.json` (the primary's local, session-modified copy —
present only for parity) and `wiki/` (a separate submodule). Stage files explicitly; do **not**
`git add -A` / `git add .` blindly, or you risk staging them. If `.mcp.json` shows as modified,
leave it — it is flagged `--skip-worktree` and is not yours to commit.
## 4. Push your branch
```bash
git push -u origin HEAD
```
Push is over SSH as the same user — no extra credential needed. Push the branch as-is; do not
force-push over anything you did not create.
## 5. Open your own PR to `main`
Open the PR via the gitea REST API. The daemon injected a **repo-scoped token** (`GITEA_TOKEN`)
and the forge host (`GITEA_HOST`) into your env for exactly this — the token can create a PR but
**cannot merge** (that stays the lead/human gate).
```bash
API="${GITEA_HOST%/}/api/v1/repos/lms/claude-bridge/pulls"
BRANCH="$(git branch --show-current)"
curl -sS -X POST "$API" \
-H "Authorization: token ${GITEA_TOKEN}" \
-H "Content-Type: application/json" \
-d "$(cat <<JSON
{"head": "${BRANCH}", "base": "main",
"title": "<ticket>: <concise change summary>",
"body": "<what changed and why; reference the ticket; note tests run and their result>"}
JSON
)"
```
The response JSON includes `"html_url"` — that is your PR URL. If the call fails (non-2xx), read
the error body, fix the cause if it is yours (e.g. branch not pushed yet), and report the failure
honestly in your reply rather than inventing a URL. If `GITEA_TOKEN` is unset, your profile was
not granted PR-create — push the branch (step 4) and report the branch name so the lead opens the
PR.
## 6. Reply via `bridge_reply` — the PR is the handoff
End your turn with **exactly one** `bridge_reply`. That reply is the entire handoff — the lead
cannot see your terminal. Include:
```
PR: <html_url from step 5, or "not created: <reason>" + branch name>
branch: <your branch>
files: <the files you changed>
tests: <what you ran and its REAL result — or "not run: <why>">
summary: <2-3 lines: what you implemented and any caveat the reviewer needs>
```
Then stop. **Do not merge. Do not touch `.mcp.json` or `wiki/`.** One reply closes the turn.
```mermaid
sequenceDiagram
autonumber
participant L as Lead
participant B as bridged
participant I as Implementer (you)
participant G as git / gitea
L->>B: bridge_send(task) — blocks
B-->>I: your assignment (in a worktree on your branch)
I->>I: implement + build/test here
I->>G: git commit (never .mcp.json / wiki)
I->>G: git push -u origin HEAD
I->>G: POST /pulls (GITEA_TOKEN) — open PR to main
G-->>I: html_url
I->>B: bridge_reply(PR url, branch, files, tests)
B-->>L: { outcome:"reply", text }
Note over L,G: lead reviews the PR, merges on green — you never merge
```
*The implement turn: work in the worktree, commit → push → open the PR, hand off the URL. The
lead is the merge gate.*
+84
View File
@@ -0,0 +1,84 @@
---
name: reviewer
description: Reviewer-role playbook for a bridged worker — read the assigned scope, find the real issues, ask the lead via bridge_ask when a decision is genuinely theirs, and report the finding via bridge_reply. Load this when you have been delegated a code review over bridged.
---
# Reviewer worker
You are a **reviewer** in the claude-bridge fleet. The lead delegated you one scoped review
over `bridged`, and your whole job is **this single turn**: examine the scope it named, and
report back. You are not the owner of the code and you do not merge anything — you surface
what the owner needs to know, then hand the turn back.
Delivery mechanics (how the task reached you, how your reply resolves the lead's blocked
send) are in [`docs/MCP-Contract.md`](../../../docs/MCP-Contract.md); you only need the three
rules below.
## 1. Read the whole scope before you judge
The delegation names your scope — a file, a diff, a PR, a function. **Read all of it first.**
A review that fires on a snippet misses the caller that makes it safe (or the one that makes
it a bug). Reviewing only part of the scope and guessing the rest is the most common way a
reviewer worker is wrong.
## 2. Stay in your lane
- Review **only** the assigned scope. If you notice something elsewhere, mention it in one
line — do **not** go hunt it. Wandering is how two workers end up reporting the same thing
and neither covers what it was given.
- Do **not** edit files, run the build, or spawn other workers. You review; the owner acts.
- You never set `ANTHROPIC_BASE_URL` and never touch herdr — you are a Claude Code process,
not part of the transport.
## 3. When the decision is the lead's — ask, don't guess
Some things you cannot resolve from the code: an ambiguous requirement, a missing acceptance
criterion, "is this behavior intended or a bug?", or a choice between two defensible fixes.
Guessing there produces a confident-but-wrong finding. Instead **pause and ask the lead** with
`bridge_ask` — a single crisp question. The call blocks; when the lead answers you **resume
the same turn** with the answer and finish. Ask only when the answer changes your finding;
don't narrate options you could decide yourself.
```mermaid
sequenceDiagram
participant L as Lead
participant B as bridged
participant R as Reviewer (you)
L->>B: bridge_send(review scope) — blocks
B-->>R: your assignment
R->>R: read the full scope
opt a decision only the lead can make
R->>B: bridge_ask("intended, or a bug?") — you block
B-->>L: { outcome:"question", turn_id }
L->>B: bridge_send(answer, turn_id)
B-->>R: { answer } — you resume the SAME turn
end
R->>B: bridge_reply(structured finding) — ends your turn
B-->>L: { outcome:"reply", text }
```
*The review turn, with the optional `bridge_ask` detour when the call is the lead's to make.*
## 4. Report with `bridge_reply` — one structured finding
End your turn with **exactly one** `bridge_reply`. Report the **single most important** real
issue in the scope, in these four lines, under ~90 words:
```
1. <path>:<line>
2. issue: <one sentence — what is wrong and why it matters>
3. fix: <one line — the concrete change>
4. severity: high | medium | low
```
- Found nothing real after reading? Reply `NO ISSUE` and one line saying why — a clean review
is a valid result, and a fabricated issue is worse than none.
- **Severity:** `high` = wrong result, data loss, security, or a hang/crash on a real path ·
`medium` = a real bug on an edge path, or a correctness risk under load/concurrency ·
`low` = clarity, a latent foot-gun, or a smell with no current failure.
- Be specific and verifiable: a line number and a one-line repro beat an adjective. If you
can't point to where it goes wrong, you haven't found it yet.
One reply closes the turn. If you asked mid-turn, the answer you got is already folded into
this finding — you do not ask again after replying.
+9
View File
@@ -20,6 +20,15 @@ module is **`bridged`**. Always pass these to IDE MCP tools:
test run). A per-file-clean file can still break the build or another module. This is the
whole-project gate before declaring work done or committing.
**Whenever dependencies change (or a `pom.xml` edit), validate CVEs with
`jetbrains get_file_problems{filePath: "bridged/pom.xml"}`** — its Mend.io check reflects the
dependencies on disk. (Note: `ide_diagnostics` / intellij-index does NOT re-resolve dependencies
after a pom edit without a full Maven reimport, so it reports stale CVE results — don't trust it
for this.) Treat a CVE warning like any other: bump to a patched version and confirm
`mvn clean install` still passes. If the latest available version is still flagged (EOL line,
"insufficient information", or config-file-only advisories), document it as accepted in the pom
rather than chasing a fix that doesn't exist.
### Use IDE MCP tools for navigation, refactoring, and diagnostics only
- **Navigate (prefer over Grep/Read for symbols):** `ide_find_definition`, `ide_find_class`,
+27 -5
View File
@@ -50,8 +50,9 @@ flowchart LR
plain daemon (no Anthropic quota), so it may poll/subscribe freely.
- **One gateway (unified MCP setup):** `bridged` is the **sole communication path** for every
Claude session. Primary and workers each mount it as an MCP server (one `claude mcp add`
line, same on both) and talk over MCP tools — `bridge_send` / `bridge_reply` / `bridge_ask` /
`bridge_status`. **No Claude session ever addresses a broker, a peer, or the network
line, same on both) and talk over MCP tools — `bridge_send` / `bridge_reply` /
`bridge_status` (with `bridge_ask` planned for the blocked-worker path). **No Claude session
ever addresses a broker, a peer, or the network
directly**; any queue is `bridged`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so
mounting the bridge is subscription-safe by construction.
- **How the primary consumes a reply:** a single **blocking MCP call** (`bridge_send`);
@@ -85,6 +86,27 @@ Gitea wiki.
## Status
🟢 Design — herdr-centric **`bridged`** message server selected as the primary approach
(2026-07-11), superseding the AgentAPI plan (2026-07-08). AgentAPI retained as fallback
injector.
🟢 **Implemented & dogfooded** — the herdr-centric **`bridged`** message server is built and in
real use: an Opus primary delegates tasks to off-subscription workers that reply through the
bridge (code reviews delegated this way have produced committed bug fixes). Selected as the
primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a
fallback injector.
**Shipped** (Java 25 · Maven · 105 tests green — unit/acceptance + live-herdr contract tests):
- **Core gateway** — herdr socket client (contract-tested vs live 0.7.0); guard-checked worker
spawn with `ANTHROPIC_BASE_URL` injected only into the worker's env; status-gated injector;
blocking `bridge_send` with reply rendezvous; MCP server as a thin adapter over the REST core.
- **MCP tools** — `bridge_send` / `bridge_reply` / `bridge_status` (messaging) and `bridge_spawn`
/ `bridge_list` / `bridge_stop` / `bridge_profiles` / `bridge_poll` (fleet). Caller identity is
connection-based (loopback peer PID → herdr pane), so the same mount serves primary and workers.
- **Delivery reliability** — completion fallback (a confirmed `working→idle` turn resolves a
send); async fire-and-poll (beats the caller's MCP call timeout for long tasks); and failure
detection for wedged (`unknown`), vanished, and never-ready workers so a send never hangs.
- **Fleet** — multiple worker profiles, each with an independent base_url guard check; workers
inherit the primary's working directory (never `$HOME`); a readiness gate holds delivery until
a worker's Claude has connected the bridge MCP (no paste lost into its boot window).
**Next** (see the [roadmap](wiki/8-Roadmap.md)) — structured envelope schema, `bridge_ask`
(blocked-worker path), session lifecycle / recycle / `idle_ttl`, split-host, and hardening
(auth/TLS, `/metrics`, CI, systemd).
+46 -14
View File
@@ -6,28 +6,60 @@
# REST + MCP listen address. Keep it on loopback — bridged is same-host in Stage-1.
bind:
host: 127.0.0.1
port: 8080
port: 8765
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
# How a worker session is spawned. Stage-1 uses the existing ccs `ltms-local`
# profile, whose .claude.json routes to the gx00 vLLM below.
worker:
profile: ltms-local
baseUrl: http://gx00.gw:8000 # the gx00 vLLM (models: coder / deepseek-v4-flash)
model: coder
# Placement: each worker lands in its OWN tab inside a dedicated worker space, so it
# never splits or clutters your real work spaces. Use `pane` for the legacy behaviour
# (split the currently-focused tab).
placement: tab # tab | pane
workspace: bridged-workers # the dedicated worker space (found-or-created, shared)
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
# How worker sessions are spawned. Define one or more named profiles (backends) under
# `workers`; each key is the profile name (also the ccs profile). `defaultWorker` picks
# which one a no-argument spawn uses (bridge_spawn with no profile / POST /workers).
#
# Shared knobs (placement/workspace/tabLabel) can be repeated per profile; they usually match.
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
# Use `pane` for the legacy behaviour (split the focused tab).
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
# (--append-system-prompt) as launch flags; nothing is written to the profile.
# tokenEnv → host env var holding the worker's auth token (value never stored in config);
# omit for a backend that needs no token (e.g. a local ollama).
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
# cwd on an MCP spawn, else the daemon's cwd — never $HOME. See
# docs/Worker-Startup-and-Trust.md.
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
workers:
gx10: # ccs profile name (NOT a hostname)
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
model: coder
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
ollama:
baseUrl: http://ollama.ltms.dev # local/self-hosted; usually no token
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}"
mcpUrl: http://127.0.0.1:8765/mcp
argv: ["ccs", "ollama"]
defaultWorker: gx10
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
# must carry none. Grounded in ltms-local's real endpoints.
# must carry none. Every profile above must have its host listed here.
guard:
offSubscriptionHosts:
- gx00.gw
- gx01.gw
- ollama.ltms.dev
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
# contextCap → force-release a session after this many delegated turns
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
# lifecycle:
# idleTtlSeconds: 300
# contextCap: 10
# drainTimeoutSeconds: 5
+53 -4
View File
@@ -17,13 +17,54 @@
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<mainClass>dev.ltms.bridged.Bridged</mainClass>
<jackson.version>2.18.2</jackson.version>
<javalin.version>6.3.0</javalin.version>
<jackson.version>2.19.0</jackson.version>
<javalin.version>6.7.0</javalin.version>
<jetty.version>11.0.25</jetty.version>
<mcp.version>2.0.0</mcp.version>
<slf4j.version>2.0.16</slf4j.version>
<logback.version>1.5.15</logback.version>
<logback.version>1.5.18</logback.version>
<junit.version>5.11.4</junit.version>
</properties>
<!--
Dependency security (validate with the JetBrains analyzer's Mend.io check on this pom).
Deps are pinned to the latest available versions. Residual advisories with NO upstream fix,
accepted for this loopback-bound daemon that processes no untrusted config:
- jetty-http 11.0.25 (via Javalin): CVE-2026-2332, CVE-2025-11143 — Jetty 11 is EOL;
fixed only in Jetty 12, which needs a Javalin major (6.x rides Jetty 11).
- logback-core 1.5.18: CVE-2025-11226, CVE-2026-1225 — both require a MALICIOUS
logback.xml (attacker with config write already has code execution); ours is trusted.
- jackson-core 2.19.0: WS-2026-0003 — "insufficient information", no fixed version published.
- tools.jackson.core (Jackson 3) 3.0.3 via the MCP SDK: CVE-2026-29062 (nesting-depth
resource exhaustion). The SDK 2.0.0 is pinned to Jackson 3.0.3 + jackson-annotations
3.0-rc5; bumping Jackson 3 to the patched 3.2.x breaks the SDK (annotation mismatch).
Only the loopback /mcp endpoint parses this JSON, from trusted local Claude clients.
The 11.0.23 -> 11.0.25 bump did clear jetty CVE-2024-8184 (5.9) and CVE-2024-6763.
-->
<!-- Force the latest patched Jetty 11.x across all Javalin-pulled Jetty modules (no version
skew). Javalin 6.x rides Jetty 11; a move to Jetty 12 needs a Javalin major. -->
<dependencyManagement>
<dependencies>
<dependency>
<groupId>org.eclipse.jetty</groupId>
<artifactId>jetty-bom</artifactId>
<version>${jetty.version}</version>
<type>pom</type>
<scope>import</scope>
</dependency>
<!-- The MCP SDK (Jackson 3) needs jackson-annotations with JsonFormat.Shape.POJO
(the 3.0 line); it shares the com.fasterxml.jackson.annotation package with our
Jackson 2.19 databind, so both must resolve to the same jar. 3.0 is built to work
with Jackson 2.19 databind too — pin it to reconcile the two. -->
<dependency>
<groupId>com.fasterxml.jackson.core</groupId>
<artifactId>jackson-annotations</artifactId>
<version>3.0-rc5</version>
</dependency>
</dependencies>
</dependencyManagement>
<dependencies>
<!-- JSON + YAML (config, herdr wire format, REST bodies) -->
<dependency>
@@ -44,6 +85,14 @@
<version>${javalin.version}</version>
</dependency>
<!-- MCP server: the SERVER face. Streamable-HTTP servlet mounted on Javalin's Jetty at
/mcp, exposing bridge_send/bridge_reply/bridge_status as thin adapters over REST. -->
<dependency>
<groupId>io.modelcontextprotocol.sdk</groupId>
<artifactId>mcp</artifactId>
<version>${mcp.version}</version>
</dependency>
<!-- Logging -->
<dependency>
<groupId>org.slf4j</groupId>
@@ -116,7 +165,7 @@
</profile>
<profile>
<id>contract</id>
<properties><excludedGroups></excludedGroups></properties>
<properties><excludedGroups/></properties>
</profile>
</profiles>
</project>
@@ -3,12 +3,26 @@ package dev.ltms.bridged;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.rest.BridgedApp;
import dev.ltms.bridged.worker.WorkerService;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.session.SessionReaper;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -40,20 +54,89 @@ public final class Bridged {
: UnixSocketHerdrClient.defaultSocketPath();
UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect(socket, new com.fasterxml.jackson.databind.ObjectMapper());
Runtime.getRuntime().addShutdownHook(new Thread(herdr::close));
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
WorkerService workers = new WorkerService(agents, spaces, guard, cfg.worker(), System::getenv);
PeerLauncher workers = new ClaudeCodeLauncher(agents, spaces, guard,
cfg.workerProfiles(), cfg.defaultProfile(), System::getenv);
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died with
// the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
int contextCap = 0;
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
&& cfg.lifecycle().contextCap() > 0) {
contextCap = cfg.lifecycle().contextCap();
}
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()), contextCap);
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
final SessionReaper reaper;
if (cfg.lifecycle() != null
&& cfg.lifecycle().idleTtlSeconds() != null
&& cfg.lifecycle().idleTtlSeconds() > 0) {
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
reaper.start();
} else {
reaper = null;
}
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
Injector injector = new Injector(agents);
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
Rendezvous rendezvous = new Rendezvous();
CompletionResolver completion = new CompletionResolver(agents, rendezvous);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
WorkerPresence presence = sessions.asPresence();
TurnListener turnListener = new TurnListener() {
@Override
public void onTurnComplete(String target) {
completion.onTurnComplete(target);
sessions.onTurnComplete(target);
}
@Override
public void onDelivered(String target) {
completion.onDelivered(target);
sessions.onDelivered(target);
}
@Override
public void onTurnFailed(String target) {
completion.onTurnFailed(target);
sessions.onTurnFailed(target);
}
};
Injector injector = new Injector(agents, turnListener, presence::isPresent, presence::forget);
StatusPoller poller = new StatusPoller(agents, injector, INJECT_POLL_MILLIS);
poller.start();
Runtime.getRuntime().addShutdownHook(new Thread(poller::stop));
Javalin app = new BridgedApp(herdr, workers).build();
MessageService messages = new MessageService(agents, injector, rendezvous);
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// Cast to ClaudeCodeLauncher: BridgeMcp is not yet migrated to PeerLauncher (Stage A scope).
BridgeMcp mcp = new BridgeMcp(messages, rendezvous, (ClaudeCodeLauncher) workers, sessions, identity, presence);
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
// last. This replaces the earlier independent hooks that could race and close herdr early.
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
poller.stop();
messages.close();
mcp.close();
if (reaper != null) reaper.stop();
herdr.close();
}));
Javalin app = new BridgedApp(herdr, (ClaudeCodeLauncher) workers, sessions, messages, rendezvous, presence, mcp.servlet()).build();
app.start(cfg.bind().host(), cfg.bind().port());
log.info("bridged listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
@@ -8,7 +8,9 @@ import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
/**
@@ -16,23 +18,33 @@ import java.util.Set;
* {@code bridged.example.yaml}). Unknown keys are ignored so config can grow ahead
* of the code.
*
* @param bind REST/MCP listen host:port
* @param herdrSocket path to herdr's Unix socket ({@code null} → client default)
* @param worker worker-spawn settings
* @param guard subscription-boundary allowlist
* @param bind REST/MCP listen host:port
* @param herdrSocket path to herdr's Unix socket ({@code null} → client default)
* @param worker single worker profile (legacy; superseded by {@code workers})
* @param workers named worker profiles, keyed by profile name (multi-backend fleet)
* @param defaultWorker which {@code workers} key a no-argument spawn uses ({@code null} → the
* single {@code worker}, or the sole/first profile)
* @param guard subscription-boundary allowlist
* @param worktreeRoot nullable root directory for provisioned worktrees; defaults to a sibling
* of the repo root
* @param lifecycle session lifecycle limits ({@code null} = all disabled)
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record BridgedConfig(
Bind bind,
String herdrSocket,
Worker worker,
Guard guard) {
Map<String, Worker> workers,
String defaultWorker,
Guard guard,
String worktreeRoot,
Lifecycle lifecycle) {
@JsonIgnoreProperties(ignoreUnknown = true)
public record Bind(String host, int port) {
public Bind {
if (host == null || host.isBlank()) host = "127.0.0.1";
if (port <= 0) port = 8080;
if (port <= 0) port = 8765;
}
}
@@ -53,17 +65,63 @@ public record BridgedConfig(
* @param tabLabel template for a worker tab's label; {@code {profile}}/{@code {model}}
* and {@code {n}} (per-worker number, to keep sibling tabs distinct)
* are substituted (default {@code "worker: {profile} #{n}"})
* @param mcpUrl bridge MCP URL to provision into the worker's {@code configDir} so it
* can call {@code bridge_reply} ({@code null}/blank → no provisioning; the
* worker won't reply, only the fallback/timeout resolves the send)
* @param cwd fixed working directory for this profile's workers (CB-112 "told otherwise");
* {@code null}/blank → inherit the primary's cwd, else the daemon's
* @param parityOverlay repo-relative paths copied primary→worktree for config parity; null/empty
* defaults to a sensible set of local config files
* @param gitTokenEnv name of the host env var holding the git-forge API token; when set, its
* value is injected as {@code GITEA_TOKEN} so the worker can open its own PR
* at checkpoint (CB-302). {@code null}/blank ⇒ no token is injected
* (minimal-grant default — push over SSH stays free, PR-create is opt-in)
* @param gitHostEnv name of the host env var holding the forge host (default {@code GITEA_HOST});
* injected as {@code GITEA_HOST} <em>only</em> when {@code gitTokenEnv} is set
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel) {
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd,
List<String> parityOverlay,
String gitTokenEnv, String gitHostEnv) {
public Worker {
argv = (argv == null || argv.isEmpty()) ? List.of("claude") : List.copyOf(argv);
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_WORKER_TOKEN" : tokenEnv;
placement = (placement == null || placement.isBlank()) ? "tab" : placement.toLowerCase();
workspace = (workspace == null || workspace.isBlank()) ? "bridged-workers" : workspace;
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? "worker: {profile} #{n}" : tabLabel;
parityOverlay = (parityOverlay == null || parityOverlay.isEmpty())
? List.of(".mcp.json", ".claude/settings.local.json", ".env", ".envrc")
: List.copyOf(parityOverlay);
// gitTokenEnv stays null when unset (opt-in). gitHostEnv defaults so operators enabling
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
gitHostEnv = (gitHostEnv == null || gitHostEnv.isBlank()) ? "GITEA_HOST" : gitHostEnv;
}
/**
* Backward-compatible constructor without the CB-302 git-forge fields — the worker is
* granted no PR-create token (push over SSH is unaffected). Keeps pre-CB-302 call sites
* (and any {@code workers:} YAML that omits the git keys) working unchanged.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, null, null);
}
/** A copy with {@code profile} set — used to default a profile to its {@code workers} key. */
public Worker withProfile(String p) {
return new Worker(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv);
}
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
public boolean hasGitToken() {
return gitTokenEnv != null && !gitTokenEnv.isBlank();
}
/** True when workers should land in their own tab in the worker space. */
@@ -71,6 +129,11 @@ public record BridgedConfig(
return "tab".equals(placement);
}
/** True when the bridge MCP should be mounted into a spawned worker (via launch flags). */
public boolean hasMcp() {
return mcpUrl != null && !mcpUrl.isBlank();
}
/**
* Render {@link #tabLabel} for the {@code n}-th worker (substitutes
* {@code {profile}}/{@code {model}}/{@code {n}}), so sibling worker tabs are distinct.
@@ -83,6 +146,21 @@ public record BridgedConfig(
}
}
/**
* Session lifecycle limits. All knobs are opt-in: {@code null} or {@code 0} disables the
* feature so existing configs keep the previous behaviour.
*
* @param idleTtlSeconds max seconds a {@code READY}/{@code DONE} session may sit idle
* before it is reaped ({@code null} → disabled)
* @param contextCap max delegated turns a session serves before force-release
* ({@code null} → disabled)
* @param drainTimeoutSeconds seconds to wait for {@code BUSY} sessions to finish before
* forced teardown on shutdown (default 5 when unset)
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Lifecycle(Integer idleTtlSeconds, Integer contextCap, Integer drainTimeoutSeconds) {
}
/**
* Subscription boundary. Only these hosts may back a worker's
* {@code ANTHROPIC_BASE_URL}; the primary must carry none.
@@ -100,6 +178,40 @@ public record BridgedConfig(
}
}
/**
* The effective worker profiles, keyed by profile name. Prefers the {@code workers} map (each
* value's {@code profile} defaulted to its key); falls back to the legacy singular {@code worker}
* (keyed by its own profile). Empty if neither is configured.
*/
public Map<String, Worker> workerProfiles() {
if (workers != null && !workers.isEmpty()) {
Map<String, Worker> out = new LinkedHashMap<>();
workers.forEach((name, w) -> out.put(name,
(w.profile() == null || w.profile().isBlank()) ? w.withProfile(name) : w));
return Map.copyOf(out);
}
if (worker != null) {
String name = (worker.profile() == null || worker.profile().isBlank()) ? "default" : worker.profile();
return Map.of(name, worker);
}
return Map.of();
}
/**
* The profile a no-argument spawn uses: {@code defaultWorker} if set, else the legacy single
* {@code worker}'s profile, else the sole/first configured profile, else {@code null}.
*/
public String defaultProfile() {
if (defaultWorker != null && !defaultWorker.isBlank()) {
return defaultWorker;
}
if (worker != null && worker.profile() != null && !worker.profile().isBlank()) {
return worker.profile();
}
Map<String, Worker> p = workerProfiles();
return p.isEmpty() ? null : p.keySet().iterator().next();
}
private static final ObjectMapper YAML = new ObjectMapper(new YAMLFactory());
/** Load and validate config from {@code path}. */
@@ -116,6 +228,7 @@ public record BridgedConfig(
public BridgedConfig withDefaults() {
Bind b = bind != null ? bind : new Bind(null, 0);
Guard g = guard != null ? guard : new Guard(List.of());
return new BridgedConfig(b, herdrSocket, worker, g);
Lifecycle l = lifecycle != null ? lifecycle : new Lifecycle(null, null, null);
return new BridgedConfig(b, herdrSocket, worker, workers, defaultWorker, g, worktreeRoot, l);
}
}
@@ -15,6 +15,10 @@ import com.fasterxml.jackson.databind.JsonNode;
* @param agentType detected agent kind, e.g. {@code "claude"}, or {@code null} before herdr
* has detected it (the start-time shape)
* @param status current lifecycle state
* @param name the unique label the agent was started with — for a bridge worker this is
* {@code claude-<profile>-<nonce>-<seq>} (CB-117 keys orphan reaping on the
* nonce); {@code null} for agents the bridge did not start, e.g. a user's own
* Claude session
*/
public record Agent(
String terminalId,
@@ -23,7 +27,8 @@ public record Agent(
String tabId,
String sessionId,
String agentType,
AgentStatus status) {
AgentStatus status,
String name) {
/** Project a herdr {@code agent} node. Tolerates the start-time shape (no session yet). */
public static Agent from(JsonNode a) {
@@ -42,6 +47,7 @@ public record Agent(
a.path("tab_id").asText(null),
sessionId,
type,
AgentStatus.fromWire(a.path("agent_status").asText(null)));
AgentStatus.fromWire(a.path("agent_status").asText(null)),
a.path("name").asText(null));
}
}
@@ -18,6 +18,17 @@ import java.util.Map;
*/
public final class AgentControl {
/**
* The keystroke that submits a prompt in the Claude Code TUI: a carriage return (Enter).
* It must be delivered as its <em>own</em> {@code agent.send} call — herdr delivers a message's
* text as a bracketed paste, and a {@code "\r"} appended to that same text is swallowed as
* literal newline content, not a submit. Sent as a separate keystroke event it lands outside
* the paste and submits. (A bare {@code "\n"} inserts a newline either way.) Verified live
* against Claude Code v2.1.210: an injected task stayed unsubmitted with {@code "text\r"} in
* one call, and submitted the instant a standalone {@code "\r"} was sent.
*/
static final String SUBMIT_KEY = "\r";
private final HerdrClient herdr;
public AgentControl(HerdrClient herdr) {
@@ -36,12 +47,20 @@ public final class AgentControl {
return start(name, argv, env, null);
}
/**
* Spawn an agent into a specific tab. With a non-null {@code tabId} the worker lands
* in that tab (the placement policy's dedicated worker tab); with {@code null} herdr
* splits the currently-focused tab (legacy pane placement).
*/
/** Spawn an agent into {@code tabId} at herdr's default cwd. */
public Agent start(String name, List<String> argv, Map<String, String> env, String tabId) {
return start(name, argv, env, tabId, null);
}
/**
* Spawn an agent. With a non-null {@code tabId} the worker lands in that tab (the placement
* policy's dedicated worker tab); with {@code null} herdr splits the currently-focused tab
* (legacy pane placement). A non-blank {@code cwd} sets the worker process's working directory —
* {@code agent.start} honours {@code cwd} directly (an agent pane does <em>not</em> inherit the
* tab's or workspace's cwd, so this is the only way to root a worker in the primary's directory;
* CB-112).
*/
public Agent start(String name, List<String> argv, Map<String, String> env, String tabId, String cwd) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("name", name);
params.put("argv", argv);
@@ -49,13 +68,31 @@ public final class AgentControl {
if (tabId != null) {
params.put("tab_id", tabId);
}
if (cwd != null && !cwd.isBlank()) {
params.put("cwd", cwd);
}
JsonNode result = herdr.call("agent.start", params);
return Agent.from(result.get("agent"));
}
/** Deliver {@code text} to an agent (its next prompt input). */
/**
* Deliver {@code text} to an agent as its next prompt <em>and submit it</em> — two keystroke
* events: the message (a bracketed paste, so any embedded newlines are preserved verbatim),
* then a standalone {@link #SUBMIT_KEY} (Enter) that actually submits it. Without the second
* event the text just sits in the worker's input box, never processed (see {@link #SUBMIT_KEY}).
*/
public void send(String target, String text) {
herdr.call("agent.send", Map.of("target", target, "text", text));
herdr.call("agent.send", Map.of("target", target, "text", SUBMIT_KEY));
}
/**
* Re-send the submit keystroke (Enter) to {@code target}. The Enter that accompanies a delivery
* can race the paste — especially right as the worker's TUI becomes interactive — leaving the
* text unsubmitted; the injector nudges it with this until the worker actually picks up (CB-113).
*/
public void submit(String target) {
herdr.call("agent.send", Map.of("target", target, "text", SUBMIT_KEY));
}
/**
@@ -2,28 +2,38 @@ package dev.ltms.bridged.herdr;
/**
* A herdr agent's lifecycle state, as reported by {@code agent_status}. Drives the
* status-gated injector: a worker is safe to inject into only when {@link #IDLE} or
* {@link #BLOCKED}, never mid-turn ({@link #WORKING}).
* status-gated injector: a worker is safe to inject into only when {@link #IDLE},
* {@link #BLOCKED}, or {@link #DONE}, never mid-turn ({@link #WORKING}).
*/
public enum AgentStatus {
IDLE,
WORKING,
BLOCKED,
/**
* The worker has finished its turn and is settled at an idle prompt. herdr emits this
* (observed live alongside {@code idle}) as a turn-complete marker; earlier code mapped the
* unrecognized string to {@link #UNKNOWN}, which both wedged delivery (not {@link #injectable})
* and mis-fired the CB-109 stall-failure on a worker that had actually answered. It is a
* turn-boundary equivalent to {@link #IDLE}: injectable, and a {@code working → done} edge is a
* real completion.
*/
DONE,
UNKNOWN;
/** Map herdr's wire string ({@code idle|working|blocked|unknown}) to the enum. */
/** Map herdr's wire string ({@code idle|working|blocked|done|unknown}) to the enum. */
public static AgentStatus fromWire(String s) {
if (s == null) return UNKNOWN;
return switch (s.toLowerCase()) {
case "idle" -> IDLE;
case "working" -> WORKING;
case "blocked" -> BLOCKED;
case "done" -> DONE;
default -> UNKNOWN;
};
}
/** Whether {@code bridged} may inject a message now without stepping on a live turn. */
public boolean injectable() {
return this == IDLE || this == BLOCKED;
return this == IDLE || this == BLOCKED || this == DONE;
}
}
@@ -0,0 +1,59 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import java.util.Map;
/**
* Resolves which herdr pane a process belongs to — the herdr half of connection-based MCP
* identity (CB-105). Given the PID that opened an MCP connection, {@link #terminalForPid} finds
* the agent pane whose process tree contains it, so {@code bridged} can tell <em>which worker</em>
* is calling without the worker sending anything spoofable.
*
* <p>herdr owns the PID→pane truth: {@code pane.process_info} reports each pane's {@code shell_pid}
* and foreground process PIDs. This scans agent panes; a spawn-time {@code pid→terminal} cache is
* the obvious optimization once wired into {@code ClaudeCodeLauncher}.
*/
public final class PaneLocator {
private final HerdrClient herdr;
public PaneLocator(HerdrClient herdr) {
this.herdr = herdr;
}
/**
* The {@code terminal_id} of the agent pane whose process tree contains {@code pid}, or
* {@code null} if no agent pane owns it (e.g. the caller is the primary, or off-host).
*/
public String terminalForPid(long pid) {
if (pid <= 0) {
return null;
}
for (JsonNode pane : herdr.call("pane.list", Map.of()).path("panes")) {
String paneId = pane.path("pane_id").asText(null);
if (paneId != null && paneOwnsPid(paneId, pid)) {
return pane.path("terminal_id").asText(null);
}
}
return null;
}
private boolean paneOwnsPid(String paneId, long pid) {
JsonNode info;
try {
info = herdr.call("pane.process_info", Map.of("pane_id", paneId)).path("process_info");
} catch (HerdrException e) {
return false; // pane vanished mid-scan — just skip it
}
if (info.path("shell_pid").asLong(-1) == pid) {
return true;
}
for (JsonNode p : info.path("foreground_processes")) {
if (p.path("pid").asLong(-1) == pid) {
return true;
}
}
return false;
}
}
@@ -65,13 +65,13 @@ public final class WorkspaceControl {
}
/**
* A brand-new tab in {@code workspaceId} plus the placeholder shell pane herdr seeds
* it with. Start the worker into the tab, then {@code pane.close} the root pane so the
* tab holds only the worker.
* A brand-new tab in {@code workspaceId} plus the placeholder shell pane herdr seeds it with.
* Start the worker into the tab, then {@code pane.close} the root pane so the tab holds only the
* worker. (The worker's own cwd is set on {@code agent.start}, not here — an {@code agent.start}
* pane does not inherit the tab's cwd; see {@code AgentControl.start}.)
*/
public Tab.Created createTab(String workspaceId) {
JsonNode result = herdr.call("tab.create", Map.of("workspace_id", workspaceId));
return Tab.Created.from(result);
return Tab.Created.from(herdr.call("tab.create", Map.of("workspace_id", workspaceId)));
}
/** Give a worker's tab a human label in the tab bar. */
@@ -0,0 +1,256 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.msg.Rendezvous;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
/**
* The CB-106 completion fallback: bridges the {@link Injector}'s turn-completion signal to the
* {@link Rendezvous} so a blocking {@code bridge_send} resolves even when the worker finishes its
* task without ever calling {@code bridge_reply} — the common case for a real delegated coding task.
*
* <p>On a confirmed {@code working → idle} boundary it scrapes the worker's recent transcript and
* resolves the awaiting send with that tail (a {@link Rendezvous.Kind#COMPLETION} resolution, so the
* caller can tell a scrape from a structured reply). It scrapes only when a send is actually waiting
* — a fleet worker's own turns, or a send that already timed out, cost no herdr traffic. An explicit
* {@code bridge_reply} that raced in first wins; {@link Rendezvous#resolveCompletion} is then a no-op.
*
* <p>It also handles the CB-109 stall signal ({@link #onTurnFailed}): a worker that ran a turn then
* wedged in an {@code unknown} state resolves the send as a failure (with the error screen as
* context) rather than leaving it to time out.
*
* <p>The scrape is cleaned to the last {@code ⏺} assistant block (stripping TUI chrome) and guarded
* against misattribution (CB-115): the pane content is baselined on delivery ({@link #onDelivered}),
* and a completion whose scrape is unchanged from that baseline — the previous turn's wind-down
* sampled as this turn's boundary on a rapid back-to-back send — is suppressed rather than resolving
* the send with a stale answer.
*
* <p><strong>Waiter-specific resolution (CB-116).</strong> On delivery we also capture the exact
* {@link Rendezvous} waiter this turn belongs to, and the completion/failure fallbacks resolve
* <em>that</em> waiter — never "whatever send is waiting now". A completion fallback runs on a virtual
* thread and can land after the worker's {@code bridge_reply} already resolved the turn and the
* <em>next</em> send opened its own waiter on the same session; resolving the current waiter would
* then deliver turn N's stale scrape as turn N+1's answer. Targeting the captured waiter makes a late
* completion a harmless no-op (its waiter is already done) instead of a cross-turn stale reply.
*
* <p>Wired as the {@link Injector}'s {@link TurnListener}; the handlers hand off to a virtual thread
* so the scrape's herdr round-trip never stalls the status poller. The captured waiter is read on the
* poller thread (before any next-turn delivery can overwrite it) and passed into the virtual thread.
*/
public final class CompletionResolver implements TurnListener {
private static final Logger log = LoggerFactory.getLogger(CompletionResolver.class);
/**
* herdr {@code agent.read} source for the completion scrape. {@code recent} returns the tail of
* the transcript (the worker's last output), which is what a delegator wants when the worker
* didn't structure a reply.
*/
static final String SCRAPE_SOURCE = "recent";
/** Cap the scraped tail so a long transcript can't return an unbounded blob. */
static final int MAX_SCRAPE_CHARS = 4000;
private final AgentControl agents;
private final Rendezvous rendezvous;
/**
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
* delivering send opened, plus the assistant block present when it was delivered.
*
* <p>The {@code waiter} is what makes a late fallback safe (CB-116): we resolve it, not "whoever
* is waiting now", so a completion that fires after the next send has opened its own waiter is a
* no-op rather than a cross-turn stale reply. The {@code baseline} is the CB-115 staleness
* reference: a completion scrape equal to it means the worker produced no new output (the previous
* turn's wind-down sampled as this boundary), so it is suppressed. Overwritten on each delivery;
* cleared when the turn resolves. Package-private so tests can capture and replay a specific turn.
*/
record InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline) {
}
private final ConcurrentHashMap<String, InFlight> inFlight = new ConcurrentHashMap<>();
public CompletionResolver(AgentControl agents, Rendezvous rendezvous) {
this.agents = agents;
this.rendezvous = rendezvous;
}
@Override
public void onDelivered(String target) {
// Capture the exact waiter this turn belongs to (CB-116) and snapshot the pane's pre-turn
// content — what it shows *before* the just-delivered turn produces output — as the staleness
// reference (CB-115). Done synchronously (like the delivering send itself) so both are in
// place before this turn's completion can fire.
captureBaseline(target);
}
/** Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link #onDelivered}). */
void captureBaseline(String target) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null) {
inFlight.remove(target); // no send is waiting on this delivery — nothing to resolve later
return;
}
String baseline;
try {
// Clip to the same cap resolve() applies to the tail (line ~134): the CB-115 misattribution
// guard compares baseline.equals(tail), so both sides must be the same capped representation.
// An unclipped baseline vs a clipped tail would never match for a >MAX_SCRAPE_CHARS block,
// defeating the guard and letting a stale completion resolve the send.
baseline = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
} catch (RuntimeException e) {
baseline = null; // fail open: no baseline ⇒ no suppression
log.debug("delivery baseline for {} failed: {}", target, e.getMessage());
}
inFlight.put(target, new InFlight(waiter, baseline));
}
/** The turn currently baselined for {@code target}, or {@code null} — a test hook for the captureBaseline path. */
InFlight inFlight(String target) {
return inFlight.get(target);
}
@Override
public void onTurnComplete(String target) {
// Read the in-flight turn on the poller thread — before any next-turn delivery can overwrite
// it — then off-load the scrape (a herdr round-trip we must not block polling on) to a vthread.
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("completion-" + target).start(() -> resolve(target, turn));
}
@Override
public void onTurnFailed(String target) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn));
}
/** Synchronous resolve (the unit-testable core of {@link #onTurnComplete}). */
void resolve(String target, InFlight turn) {
CompletableFuture<Rendezvous.Resolution> waiter = turn == null ? null : turn.waiter();
if (waiter == null || waiter.isDone()) {
// Nobody is blocked on THIS turn (it had no send, or its bridge_reply already won). Skip
// the scrape; resolving the current waiter here would be the CB-116 cross-turn stale reply.
inFlight.remove(target, turn);
return;
}
String tail;
boolean scrapeFailed = false;
try {
tail = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
} catch (RuntimeException e) {
// The worker finished but we couldn't read its screen — still resolve the send so the
// caller unblocks; an empty tail beats hanging until the caller's timeout.
log.warn("completion scrape for {} failed; resolving with an empty tail: {}",
target, e.getMessage());
tail = "";
scrapeFailed = true;
}
// Misattribution guard (CB-115): if the scrape is byte-identical to the pane content at
// delivery, this turn produced no new output — the boundary belongs to the previous turn's
// wind-down (common on rapid back-to-back sends). Suppress rather than resolve the send with
// a stale answer; the real bridge_reply (or a later genuine completion) resolves it instead.
// A scrape that failed to read is exempt — an empty tail there is "couldn't see", not "no change".
String baseline = turn.baseline();
if (!scrapeFailed && baseline != null && baseline.equals(tail)) {
log.debug("suppressing misattributed completion for {} (no output change since delivery)",
target);
return; // keep the in-flight record: a later genuine completion still needs it
}
if (rendezvous.resolveCompletion(waiter, tail)) {
inFlight.remove(target, turn);
log.debug("resolved send to {} via turn-completion fallback ({} chars scraped)",
target, tail.length());
}
}
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
void fail(String target, InFlight turn) {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fall back to the currently-registered waiter (unambiguous — that send never completed, so
// no next turn exists to confuse it with).
CompletableFuture<Rendezvous.Resolution> waiter =
turn != null ? turn.waiter() : rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
inFlight.remove(target, turn); // nobody blocked on this worker — nothing to fail
return;
}
String reason;
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
}
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.debug("failed send to {} via turn-stall fallback", target);
}
}
private static String clip(String s) {
if (s == null) return "";
String trimmed = s.strip();
return trimmed.length() <= MAX_SCRAPE_CHARS
? trimmed
: trimmed.substring(trimmed.length() - MAX_SCRAPE_CHARS);
}
/**
* Extract the last assistant message from a raw Claude Code pane scrape (CB-115). Claude Code
* prefixes each assistant turn with {@code ⏺}; the delegator wants that answer, not the TUI
* chrome around it. Take everything from the final {@code ⏺} onward and stop at the <em>first</em>
* hard interface boundary below it — the spinner/status line, input box, {@code ❯} prompt (which
* may echo the <em>next</em> turn's text), footer, or tips/warnings. Stopping at the first
* boundary (rather than trimming only trailing chrome) is what keeps a following turn's echoed
* prompt out of this reply. Blank lines are not boundaries, so a multi-paragraph answer survives;
* trailing blanks are trimmed at the end. With no {@code ⏺} marker (an unusual render) the whole
* text is scanned the same way, so we never lose the reply.
*
* <p>Package-private and pure so it is unit-testable without herdr.
*/
static String lastAssistantBlock(String raw) {
if (raw == null || raw.isBlank()) return "";
int marker = raw.lastIndexOf('⏺');
String block = marker >= 0 ? raw.substring(marker + 1) : raw;
StringBuilder out = new StringBuilder();
int kept = 0;
for (String line : block.split("\n", -1)) {
if (isBoundary(line)) break; // first TUI boundary ends the assistant message
if (kept++ > 0) out.append('\n');
out.append(line);
}
return out.toString().strip();
}
/**
* A hard TUI boundary line that marks the end of an assistant message and the start of interface
* chrome (input box, prompt, spinner, footer, tips/warnings). Blank lines are <em>not</em>
* boundaries — an answer may contain them — so they are kept and trimmed only if trailing.
*/
private static boolean isBoundary(String line) {
String t = line.strip();
if (t.isEmpty()) return false;
// A horizontal rule / all box-drawing separators (e.g. "──────").
if (t.chars().allMatch(c -> c == '─' || c == '—' || c == '━' || c == '═' || c == '-')) {
return true;
}
String lower = t.toLowerCase();
return t.startsWith("╭") || t.startsWith("│") || t.startsWith("╰") || t.startsWith("┌")
|| t.startsWith("└") || t.startsWith("❯") || t.startsWith("⏵")
|| t.startsWith("⎿") || t.startsWith("⚠")
// Status/spinner lines Claude Code renders below a settled or in-flight turn,
// e.g. "✻ Baked for 21s", "✶ Forming…".
|| t.startsWith("✻") || t.startsWith("✳") || t.startsWith("✽") || t.startsWith("·")
|| t.startsWith("●") || t.startsWith("◐") || t.startsWith("✢") || t.startsWith("✶")
|| lower.contains("auto mode") || lower.contains("for shortcuts")
|| lower.contains("esc to interrupt") || lower.contains("bypass permissions");
}
}
@@ -12,6 +12,8 @@ import java.util.List;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Consumer;
import java.util.function.Predicate;
import java.util.stream.Collectors;
/**
@@ -32,6 +34,14 @@ import java.util.stream.Collectors;
* a transient {@code unknown} — counts as a real pickup, so a detection glitch can't prematurely
* release the latch. Perfectly reliable turn boundaries require a herdr {@code events.subscribe}
* stream; that is the intended upgrade and would replace only the sampling, not this queue.
*
* <p><strong>Turn completion (CB-106).</strong> Beyond delivery, the injector reports when a
* delegated turn <em>finishes</em>: after a delivery is picked up (a real {@code working} sample),
* the next injectable sample is a confirmed {@code working → idle} boundary and fires
* {@link TurnListener#onTurnComplete}. Completion is only ever synthesized from a <em>confirmed</em>
* turn — the pickup-grace path (a turn too fast to sample) unwedges the queue but does not fire
* completion, since without a sampled {@code working} there is no trustworthy "the worker just
* finished the task" signal to act on.
*/
public final class Injector {
@@ -45,11 +55,68 @@ public final class Injector {
*/
private static final int PICKUP_GRACE_POLLS = 8;
/**
* How many consecutive {@code unknown} samples while a delegation is outstanding before we
* declare it stalled and fire {@link TurnListener#onTurnFailed} (CB-109). A worker wedged in a
* state herdr can't classify (e.g. an API-error screen) stays {@code unknown} indefinitely and
* would otherwise never resolve; any {@code working}/{@code idle} sample resets the streak, so a
* transient detection glitch cannot trip it. At the 250ms poll interval this is ~30s — far longer
* than any real detection blip, and still vastly better than the async send's timeout.
*/
private static final int TURN_STALL_GRACE_POLLS = 120;
/**
* How many consecutive injectable samples a queued-but-undelivered message may wait on the
* {@link #ready} gate before we give up and fail it (CB-114). The gate holds a message out of a
* worker's boot window (herdr reports {@code idle} while its Claude is still starting), but a
* worker whose Claude crashes during boot — or never connects the bridge MCP — stays "idle and
* not ready" forever: {@link #ready} never accepts it, the message is never delivered, and the
* target would be polled indefinitely with its caller's future never completing. After this
* grace the queued messages are failed and the target released. At the 250ms poll interval this
* is ~60s — deliberately longer than {@link #TURN_STALL_GRACE_POLLS}, since a first boot (spawn
* + model load + MCP connect) legitimately takes longer than an in-turn detection blip.
*/
private static final int READINESS_GRACE_POLLS = 240;
private final AgentControl agents;
private final TurnListener turnListener;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
private final Consumer<String> forget; // CB-114: clear a gone worker's readiness/presence
private final ConcurrentHashMap<String, Target> targets = new ConcurrentHashMap<>();
/** Delivery only; completion signalling is a no-op and every target is treated as available. */
public Injector(AgentControl agents) {
this(agents, TurnListener.NOOP);
}
/** Delivery plus turn-completion signalling (CB-106); every target is treated as available. */
public Injector(AgentControl agents, TurnListener turnListener) {
this(agents, turnListener, _ -> true);
}
/**
* Delivery, completion signalling (CB-106), and a readiness gate (CB-113): a message is delivered
* only when {@code ready} accepts the target — i.e. the worker's Claude has connected the bridge
* MCP. This holds the first delivery out of the worker's boot window, where herdr already reports
* {@code idle} but the TUI would drop an injected paste.
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready) {
this(agents, turnListener, ready, _ -> {
});
}
/**
* Delivery, completion signalling (CB-106), a readiness gate (CB-113), and readiness cleanup
* (CB-114): {@code forget} is invoked with a target when its worker is gone — dropped
* (pane crash) or timed out on the readiness gate — so its stale presence/readiness is cleared
* and does not linger past the worker's life.
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
this.agents = agents;
this.turnListener = turnListener;
this.ready = ready;
this.forget = forget;
}
/** A pending message and the future that completes when it has been delivered. */
@@ -59,8 +126,12 @@ public final class Injector {
/** Per-worker delivery state, guarded by its own monitor (single writer per worker). */
private static final class Target {
final Deque<Pending> queue = new ArrayDeque<>();
boolean awaitingPickup; // sent a message, waiting for the worker to pick it up
boolean awaitingPickup; // sent a message, waiting for the worker to pick it up
int injectableSincePickup; // consecutive injectable samples while awaitingPickup
boolean awaitingCompletion; // a delivered message's turn is not yet known-complete
boolean turnObserved; // saw a real `working` sample since that delivery (turn ran)
int unknownSinceTurn; // consecutive `unknown` samples while a delegation is outstanding (CB-109)
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
synchronized void add(Pending p) {
queue.add(p);
@@ -99,63 +170,154 @@ public final class Injector {
Pending sent = null;
RuntimeException sendError = null;
boolean turnCompleted = false;
boolean turnFailed = false;
boolean resubmit = false;
List<Pending> notReady = null; // queued messages failed because the worker never became ready
synchronized (t) {
if (status == AgentStatus.WORKING) {
// Definitive pickup: the worker is busy on our last message.
// Definitive pickup: the worker is busy on our last message, and (if a delivery is
// outstanding) a real turn is now confirmed to be running.
t.awaitingPickup = false;
t.injectableSincePickup = 0;
t.unknownSinceTurn = 0;
t.notReadySincePoll = 0;
if (t.awaitingCompletion) t.turnObserved = true;
} else if (status.injectable()) { // IDLE or BLOCKED
if (t.awaitingPickup && ++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
// Pickup edge was never sampled (turn faster than the poll, or status lag).
// The worker has plainly moved on — release the latch rather than wedge.
t.awaitingPickup = false;
t.injectableSincePickup = 0;
t.unknownSinceTurn = 0;
if (t.awaitingPickup) {
if (++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
// Pickup edge was never sampled (turn faster than the poll, or status lag).
// Release the latch rather than wedge — and give up on synthesizing a
// completion for this message, since without a confirmed `working` we cannot
// trust that a task-processing turn actually ran.
t.awaitingPickup = false;
t.injectableSincePickup = 0;
t.awaitingCompletion = false;
t.turnObserved = false;
} else {
// Delivered but still idle → the worker hasn't picked it up; the submit
// keystroke likely raced the paste (esp. right as the TUI became ready).
// Re-nudge Enter (CB-113) until the worker starts (WORKING) or the grace ends.
resubmit = true;
}
}
if (!t.awaitingPickup) {
Pending p = t.queue.peek();
if (p != null) {
try {
agents.send(target, p.text());
t.queue.poll();
t.awaitingPickup = true;
t.injectableSincePickup = 0;
sent = p;
} catch (RuntimeException e) {
// Delivery failed at herdr; drop the poisoned message and surface it
// rather than blocking the queue behind it.
t.queue.poll();
sent = p;
sendError = e;
// A confirmed turn (a `working` sample was seen) that has now returned to idle is
// a trustworthy `working → idle` completion boundary.
if (t.awaitingCompletion && t.turnObserved) {
t.awaitingCompletion = false;
t.turnObserved = false;
turnCompleted = true;
}
// Deliver the next queued message only once the prior turn is fully settled, so a
// completion is never confused with the pickup of the following message — and only
// once the worker is available (CB-113), so we never paste into its boot window.
if (!t.awaitingCompletion) {
Pending p = t.queue.peek();
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
try {
agents.send(target, p.text());
t.queue.poll();
t.awaitingPickup = true;
t.awaitingCompletion = true;
t.turnObserved = false;
t.injectableSincePickup = 0;
sent = p;
} catch (RuntimeException e) {
// Delivery failed at herdr; drop the poisoned message and surface it
// rather than blocking the queue behind it.
t.queue.poll();
sent = p;
sendError = e;
}
} else if (p != null && ++t.notReadySincePoll >= READINESS_GRACE_POLLS) {
// The worker has been idle-but-not-ready for the whole grace: its Claude
// never connected the bridge MCP (crashed during boot, or wedged on a
// startup prompt). The readiness gate would hold this message forever, so
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
t.queue.clear();
t.notReadySincePoll = 0;
}
}
}
} else {
// UNKNOWN (or any other non-injectable, non-working): not a safe window nor a
// reliable pickup signal, so we never deliver or release the pickup latch here. But
// an outstanding delegation whose worker has gone unresponsive — stuck in a state
// herdr can't classify (CB-109) — will never yield a working→idle boundary. After a
// sustained streak, declare it failed so the awaiting send resolves rather than
// riding out the async timeout. (This also frees a delivery that wedged before it
// was ever picked up, which the injectable-only pickup grace could never release.)
if (t.awaitingCompletion && ++t.unknownSinceTurn >= TURN_STALL_GRACE_POLLS) {
t.awaitingPickup = false;
t.awaitingCompletion = false;
t.turnObserved = false;
t.unknownSinceTurn = 0;
turnFailed = true;
}
}
// UNKNOWN (and any other non-injectable, non-working): do nothing — neither a safe
// window nor a reliable pickup signal, so we must not deliver or release the latch.
// Reclaim the entry once the worker is fully quiescent, so the map cannot grow without
// bound across many short-lived workers.
if (t.queue.isEmpty() && !t.awaitingPickup) {
// Reclaim the entry once the worker is fully quiescent (nothing queued, no pickup or
// completion awaited), so the map cannot grow without bound across short-lived workers.
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion) {
targets.remove(target, t);
}
}
// Fire listeners / herdr calls after releasing the monitor so nothing runs on the poller
// thread while it holds the target lock.
if (resubmit) {
try {
agents.submit(target); // nudge a raced Enter so the pending paste submits
} catch (RuntimeException e) {
log.debug("resubmit to {} failed (will retry next poll): {}", target, e.getMessage());
}
}
if (notReady != null) {
// Worker never became available: forget its (never-set) readiness, unblock every queued
// caller, and route the awaiting send through the same failure path as a stalled turn so
// a blocking or async waiter resolves WORKER_FAILED rather than riding out the timeout.
forget.accept(target);
RuntimeException cause = new IllegalStateException(
target + " never became available (no bridge MCP connection within the boot window)");
for (Pending p : notReady) {
p.delivered().completeExceptionally(cause);
}
turnListener.onTurnFailed(target);
}
if (turnCompleted) {
turnListener.onTurnComplete(target);
}
if (turnFailed) {
turnListener.onTurnFailed(target);
}
if (sent != null) {
if (sendError != null) {
log.warn("inject to {} failed, dropped message: {}", target, sendError.getMessage());
sent.delivered().completeExceptionally(sendError);
} else {
// Baseline the pane's pre-turn content so a misattributed completion (no new output)
// can't resolve this send with the previous turn's stale answer (CB-115).
turnListener.onDelivered(target);
sent.delivered().complete(null);
}
}
}
/** Targets the poller must keep sampling: those with a queued message or an awaited pickup. */
/**
* Targets the poller must keep sampling: those with a queued message, an awaited pickup, or an
* awaited turn completion (so the {@code working → idle} boundary is observed).
*/
public Set<String> activeTargets() {
return targets.entrySet().stream()
.filter(e -> {
synchronized (e.getValue()) {
return !e.getValue().queue.isEmpty() || e.getValue().awaitingPickup;
Target t = e.getValue();
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion;
}
})
.map(java.util.Map.Entry::getKey)
@@ -163,20 +325,31 @@ public final class Injector {
}
/**
* Forget a target whose worker is gone, failing every still-queued message so awaiting
* callers unblock instead of hanging forever. Futures are completed after the monitor is
* released.
* Forget a target whose worker is gone, failing every still-queued message so awaiting callers
* unblock instead of hanging forever. If a message had already been <em>delivered</em> but its
* turn was not yet resolved (CB-110 — the worker vanished mid-turn, e.g. its pane crashed), fire
* {@link TurnListener#onTurnFailed} for it: a delivered message is no longer in the queue, so
* failing queued waiters alone would leave that send's rendezvous hanging until the async
* timeout. Futures and listeners are completed after the monitor is released.
*/
public void drop(String target, Throwable cause) {
Target t = targets.remove(target);
if (t == null) return;
List<Pending> pending;
boolean hadDeliveredTurn;
synchronized (t) {
pending = new ArrayList<>(t.queue);
t.queue.clear();
hadDeliveredTurn = t.awaitingCompletion;
t.awaitingCompletion = false;
t.awaitingPickup = false;
}
forget.accept(target); // the worker is gone — clear its readiness/presence too (CB-114)
for (Pending p : pending) {
p.delivered().completeExceptionally(cause);
}
if (hadDeliveredTurn) {
turnListener.onTurnFailed(target);
}
}
}
@@ -23,13 +23,20 @@ public final class StatusPoller {
private final AgentControl agents;
private final Injector injector;
private final StatusRefiner refiner;
private final long intervalMillis;
private volatile boolean running;
private Thread thread;
public StatusPoller(AgentControl agents, Injector injector, long intervalMillis) {
this(agents, injector, new StatusRefiner(agents), intervalMillis);
}
public StatusPoller(AgentControl agents, Injector injector, StatusRefiner refiner,
long intervalMillis) {
this.agents = agents;
this.injector = injector;
this.refiner = refiner;
this.intervalMillis = intervalMillis;
}
@@ -47,7 +54,9 @@ public final class StatusPoller {
for (String target : active) {
if (!running) return;
try {
AgentStatus status = agents.status(target);
// herdr's agent_status can misreport a settled worker as `unknown`; refine it
// against the pane content before it drives delivery/completion (CB-115).
AgentStatus status = refiner.refine(target, agents.status(target));
injector.onStatus(target, status);
} catch (HerdrException e) {
// The worker's agent is gone — stop trying and unblock its waiters.
@@ -0,0 +1,89 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
/**
* Refines an unreliable {@link AgentStatus#UNKNOWN} into a real state by reading the worker's
* terminal content (CB-115).
*
* <p>Some workers' panes are misclassified by herdr as {@code unknown} even when the worker is
* plainly settled at an idle prompt (empty {@code ❯}, "auto mode on" footer, a completed
* {@code ⏺} answer above). Left as {@code UNKNOWN} that both <em>wedges delivery</em> — the
* status-gated {@link Injector} only injects into an {@link AgentStatus#injectable} worker — and
* <em>mis-fires the CB-109 stall failure</em> on a worker that has actually answered. herdr's
* {@code agent_status} is a heuristic; the pane content is the ground truth.
*
* <p>The refinement only ever runs on a raw {@code UNKNOWN} sample (every other status is trusted
* as-is), so a healthy worker adds zero extra herdr traffic; a persistently-{@code unknown} worker
* costs one extra {@code agent.read} per poll while it has work outstanding. Classification is
* deliberately conservative — it upgrades {@code UNKNOWN} to {@link AgentStatus#WORKING} or
* {@link AgentStatus#IDLE} only on a clear signal, and leaves a genuinely unclassifiable screen
* (e.g. a wedged error state) as {@code UNKNOWN} so the CB-109 stall path can still fail it.
*/
public final class StatusRefiner {
private static final Logger log = LoggerFactory.getLogger(StatusRefiner.class);
/**
* herdr {@code agent.read} source used to inspect the pane. {@code detection} is the region
* herdr itself uses for status detection (the prompt/footer tail), which is exactly what we
* need to tell "idle at prompt" from "mid-turn".
*/
static final String PROBE_SOURCE = "detection";
private final AgentControl agents;
public StatusRefiner(AgentControl agents) {
this.agents = agents;
}
/**
* Return a trustworthy status for {@code target}. Any non-{@code UNKNOWN} {@code raw} is returned
* unchanged; an {@code UNKNOWN} triggers a pane read and content classification. A read failure
* leaves it {@code UNKNOWN} (the safe default: no delivery, and the stall path still applies).
*/
public AgentStatus refine(String target, AgentStatus raw) {
if (raw != AgentStatus.UNKNOWN) return raw;
String pane;
try {
pane = agents.read(target, PROBE_SOURCE);
} catch (RuntimeException e) {
log.debug("status refine read for {} failed; leaving UNKNOWN: {}", target, e.getMessage());
return AgentStatus.UNKNOWN;
}
AgentStatus refined = classify(pane);
if (refined != AgentStatus.UNKNOWN) {
log.debug("refined {} from UNKNOWN to {} via pane content", target, refined);
}
return refined;
}
/**
* Classify a Claude Code TUI pane tail. Package-private and pure so it is unit-testable without
* herdr.
*
* <ul>
* <li>An active-generation marker ({@code esc to interrupt}) ⇒ {@link AgentStatus#WORKING} —
* never inject here.</li>
* <li>Otherwise, an interactive input prompt with no active-turn marker ({@code ❯}, the
* {@code │ >} input box, or the idle {@code auto mode} / shortcuts footer) ⇒
* {@link AgentStatus#IDLE} — settled and safe to inject / a completed turn.</li>
* <li>Anything else (blank, or an unrecognizable screen) ⇒ {@link AgentStatus#UNKNOWN}.</li>
* </ul>
*/
static AgentStatus classify(String pane) {
if (pane == null || pane.isBlank()) return AgentStatus.UNKNOWN;
String lower = pane.toLowerCase();
// Claude Code shows "(esc to interrupt)" only while a turn is actively generating.
if (lower.contains("esc to interrupt")) return AgentStatus.WORKING;
// A settled, ready input prompt with no active-turn marker = idle-at-prompt.
boolean readyPrompt = pane.contains("❯")
|| pane.contains("│ >")
|| lower.contains("auto mode on")
|| lower.contains("? for shortcuts");
return readyPrompt ? AgentStatus.IDLE : AgentStatus.UNKNOWN;
}
}
@@ -0,0 +1,39 @@
package dev.ltms.bridged.inject;
/**
* Notified when a worker's delegated turn is observed to complete — a confirmed
* {@code WORKING → IDLE} transition after a delivery. This is the CB-106 completion signal the
* {@code CompletionResolver} uses to resolve a blocked send whose worker never called
* {@code bridge_reply}. Kept as a seam so the {@link Injector} needs no dependency on the message
* layer and stays unit-testable with a capturing fake.
*/
@FunctionalInterface
public interface TurnListener {
/** A worker's delegated turn finished (worker returned to idle after visibly working). */
void onTurnComplete(String target);
/**
* A worker that visibly ran a delegated turn then wedged in a non-idle, non-working state
* (CB-109) — e.g. an error screen herdr classifies as {@code unknown} — so no
* {@code working → idle} completion boundary will ever arrive. A default no-op keeps this a
* functional interface; the completion resolver overrides it to fail the awaiting send.
*/
default void onTurnFailed(String target) {
}
/**
* A message was just delivered into {@code target}'s pane (CB-115). Fired so the completion
* resolver can snapshot the pane's pre-turn content: a later {@link #onTurnComplete} whose
* scrape is unchanged from this baseline is a <em>misattributed</em> boundary (e.g. the prior
* turn's wind-down sampled as this turn's completion on rapid back-to-back sends) and must not
* resolve the send with the previous turn's stale answer. A default no-op keeps the interface
* functional for callers that don't scrape.
*/
default void onDelivered(String target) {
}
/** No-op default for callers that only need delivery, not completion signalling. */
TurnListener NOOP = _ -> {
};
}
@@ -0,0 +1,38 @@
package dev.ltms.bridged.inject;
import java.util.concurrent.ConcurrentHashMap;
import java.util.Set;
/**
* Tracks which workers are <em>available</em> — their Claude has booted and connected its MCP client
* to the bridge (CB-113). This is the reliable readiness signal, unlike herdr's {@code agent_status},
* which reports {@code idle} for a worker whose Claude is still booting. Delivering into that boot
* window pastes into a not-yet-ready TUI (the text is lost) and wedges the worker's delivery state,
* so the {@link Injector} holds the first delivery until the worker is present here.
*
* <p>Populated from the MCP transport: any MCP request whose connection resolves to a worker terminal
* marks that worker present (its {@code initialize} is the first such contact). A worker that never
* mounts the bridge MCP is never marked present — its sends stay queued until they time out, which is
* correct (it could not have replied anyway).
*/
public class WorkerPresence {
private final Set<String> present = ConcurrentHashMap.newKeySet();
/** Record that {@code terminal}'s worker has connected its MCP client (is available). */
public void markPresent(String terminal) {
if (terminal != null && !terminal.isBlank()) {
present.add(terminal);
}
}
/** Whether {@code terminal}'s worker is available (has been seen on the bridge MCP). */
public boolean isPresent(String terminal) {
return present.contains(terminal);
}
/** Forget a torn-down worker so its terminal id does not linger as "present". */
public void forget(String terminal) {
present.remove(terminal);
}
}
@@ -0,0 +1,539 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.modelcontextprotocol.common.McpTransportContext;
import io.modelcontextprotocol.json.McpJsonMapper;
import io.modelcontextprotocol.json.jackson3.JacksonMcpJsonMapperSupplier;
import io.modelcontextprotocol.server.McpServer;
import io.modelcontextprotocol.server.McpSyncServer;
import io.modelcontextprotocol.server.McpSyncServerExchange;
import io.modelcontextprotocol.server.transport.HttpServletStreamableServerTransportProvider;
import io.modelcontextprotocol.spec.McpSchema;
import com.fasterxml.jackson.databind.ObjectMapper;
import jakarta.servlet.http.HttpServlet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.function.Function;
import java.util.stream.Collectors;
/**
* The MCP SERVER face (CB-105): a Streamable-HTTP MCP server whose tools are <em>thin adapters</em>
* over the same {@link MessageService}/{@link Rendezvous} the REST routes use — so the two are
* validated by parity, not by re-implementing behaviour. The primary Opus calls {@code bridge_send}
* / {@code bridge_status}; the worker calls {@code bridge_reply}.
*
* <p>Beyond delegation the primary also manages the fleet here (CB-108): {@code bridge_spawn} /
* {@code bridge_list} / {@code bridge_stop} adapt {@link ClaudeCodeLauncher} so a worker's whole lifecycle
* is driven through MCP, with the subscription boundary still enforced inside {@code ClaudeCodeLauncher}.
*
* <p>The tool <em>logic</em> lives in package-private static methods returning a
* {@link McpSchema.CallToolResult}, so it is unit-testable without standing up the HTTP transport;
* the SDK owns the wire protocol. Mount {@link #servlet()} at {@code /mcp} on the daemon's Jetty.
*/
public final class BridgeMcp {
private static final long DEFAULT_TIMEOUT_MS = 25_000;
private static final long MAX_TIMEOUT_MS = 120_000;
// bridge_ask blocks the WORKER's own MCP call, which its client caps near 60s — default under
// that so the bridge returns a clean timeout before the client severs the call (CB-205).
private static final long ASK_DEFAULT_TIMEOUT_MS = 55_000;
private static final long ASK_MAX_TIMEOUT_MS = 115_000;
private static final ObjectMapper MAPPER = new ObjectMapper(); // worker-view JSON projections
/** Transport-context key under which the extractor stashes the resolved caller identity. */
static final String CALLER_TERMINAL = "callerTerminal";
/** Transport-context key under which the extractor stashes the caller's PID (for cwd inherit). */
static final String CALLER_PID = "callerPid";
private final HttpServletStreamableServerTransportProvider transport;
private final McpSyncServer server;
public BridgeMcp(MessageService messages, Rendezvous rendezvous, ClaudeCodeLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence) {
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
.jsonMapper(json)
.mcpEndpoint("/mcp")
// Resolve the caller from the connection (peer PID → herdr pane) in one lookup: the
// worker terminal for bridge_reply (no spoofable arg), and the PID so bridge_spawn can
// inherit the primary's cwd (CB-112). Any contact from a worker marks it available
// (CB-113) — its MCP initialize is the reliable "the agent is up" signal.
.contextExtractor(req -> {
ConnectionIdentity.Caller c = identity.resolve(req.getRemoteAddr(), req.getRemotePort());
presence.markPresent(c.terminal()); // no-op for the primary (null terminal)
return McpTransportContext.create(Map.of(
CALLER_TERMINAL, orEmpty(c.terminal()),
CALLER_PID, Long.toString(c.pid())));
})
.build();
this.server = McpServer.sync(transport)
.serverInfo("bridge", "0.1.0")
.capabilities(McpSchema.ServerCapabilities.builder().tools(true).build())
.toolCall(sendTool(), (_, req) -> {
Map<String, Object> a = req.arguments();
String turnId = str(a, "turnId");
if (turnId != null && !turnId.isBlank()) {
// Answering a worker's bridge_ask (CB-205): resolve its blocked question and
// block for the worker's reply as it resumes the same turn.
return answer(messages, turnId, str(a, "content"), timeoutMs(a));
}
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, str(a, "sessionId"), str(a, "content"))
: send(messages, str(a, "sessionId"), str(a, "content"), timeoutMs(a));
})
// bridge_reply's identity is the CONNECTION, never an argument.
.toolCall(replyTool(), (exchange, req) ->
reply(rendezvous, callerTerminal(exchange), str(req.arguments(), "content")))
// bridge_ask (CB-205): a worker's mid-turn question — identity from the CONNECTION.
.toolCall(askTool(), (exchange, req) ->
ask(messages, callerTerminal(exchange), str(req.arguments(), "question"), timeoutMs(req.arguments())))
.toolCall(statusTool(), (_, req) ->
status(messages, str(req.arguments(), "sessionId")))
.toolCall(pollTool(), (_, req) ->
poll(messages, str(req.arguments(), "ticket")))
// Fleet management (CB-108): spawn/list/stop over ClaudeCodeLauncher.
.toolCall(spawnTool(), (exchange, req) -> {
Map<String, Object> a = req.arguments();
// CB-112: worker inherits the primary's cwd unless the call pins one.
// CB-301: carry the caller's identity as the session owner (null for the primary).
// CB-301-ext: optional isolated worktree for parallel implementers.
String callerCwd = identity.cwdForPid(callerPid(exchange));
return spawn(sessions, str(a, "profile"), str(a, "cwd"), callerCwd,
callerTerminal(exchange), worktreeRequest(a));
})
.toolCall(listTool(), (_, _) -> listWorkers(workers, sessions))
.toolCall(stopTool(), (_, req) -> stop(sessions, str(req.arguments(), "paneId")))
.toolCall(profilesTool(), (_, _) -> profiles(workers))
.build();
}
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
private static String callerTerminal(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_TERMINAL);
String s = v == null ? null : v.toString();
return (s == null || s.isBlank()) ? null : s;
}
/** The caller's PID resolved from this call's connection, or {@code -1} if unknown. */
private static long callerPid(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_PID);
try {
return v == null ? -1 : Long.parseLong(v.toString());
} catch (NumberFormatException e) {
return -1;
}
}
private static String orEmpty(String s) {
return s == null ? "" : s;
}
/** The Streamable-HTTP servlet to mount at {@code /mcp} on the daemon's Jetty. */
public HttpServlet servlet() {
return transport;
}
/** Graceful shutdown of the MCP server. */
public void close() {
server.closeGracefully();
}
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
/** {@code bridge_send}: delegate {@code content} to a worker session and block for its reply. */
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content, Long timeoutMs) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
try {
return formatReply(messages.send(sessionId, content, timeout), timeout);
} catch (HerdrException e) {
return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
}
}
/**
* {@code bridge_send} carrying a {@code turnId}: the primary's answer to a worker's
* {@code bridge_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as
* it resumes the same turn — surfaced to the primary identically to a normal send.
*/
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs) {
if (isBlank(turnId) || isBlank(content)) {
return error("turnId and content are required to answer a worker's question");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
return formatReply(messages.answer(turnId, content, timeout), timeout);
}
/**
* {@code bridge_ask} (CB-205): a worker pauses its delegated turn to ask the primary, blocking
* until the primary answers. The worker is identified by its connection ({@code callerTerminal}),
* never an argument — a {@code null} means the caller is not a known worker.
*/
static McpSchema.CallToolResult ask(MessageService messages, String callerTerminal, String question, Long timeoutMs) {
if (callerTerminal == null) {
return error("bridge_ask is for workers only — could not identify the calling worker "
+ "from the connection");
}
if (isBlank(question)) {
return error("question is required");
}
long timeout = Math.clamp(timeoutMs == null ? ASK_DEFAULT_TIMEOUT_MS : timeoutMs, 1, ASK_MAX_TIMEOUT_MS);
MessageService.AskResult r = messages.ask(callerTerminal, question, timeout);
return switch (r.outcome()) {
case ANSWERED -> text(r.answer());
case NO_WAITER -> error("no primary is awaiting this turn — bridge_ask only works while a "
+ "bridge_send delegation is open to answer it");
case TIMED_OUT -> text("[no answer within " + timeout + "ms — the primary did not respond; "
+ "proceed on your best judgement, then call bridge_reply to end the turn]");
};
}
/** Render a {@link MessageService.Reply} as a tool result — shared by {@link #send} and {@link #answer}. */
private static McpSchema.CallToolResult formatReply(MessageService.Reply r, long timeout) {
return switch (r.outcome()) {
case REPLIED -> text(r.text());
// The worker's turn finished but it never called bridge_reply — hand back the scraped
// transcript tail, flagged so the primary knows it isn't a structured reply.
case COMPLETED_UNREPLIED -> text(
"[worker finished without a structured bridge_reply — transcript tail follows]\n" + r.text());
// The worker ran the turn then wedged (CB-109) — surface the error context.
case WORKER_FAILED -> text("[worker failed — turn ended in an unrecoverable state]\n" + r.text());
// The worker paused mid-turn to ask (CB-205) — tell the primary how to answer in-turn.
case QUESTION -> text("[question] the worker paused to ask before it can finish:\n" + r.text()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + r.turnId()
+ "\" and content set to your answer; the worker resumes the same turn.");
case STALE_TURN -> error("that question is no longer open — it timed out or was already "
+ "answered (turnId stale)");
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> text("[no reply within " + timeout + "ms — worker "
+ r.outcome().name().toLowerCase().replace("timed_out_", "") + "; retry or poll status]");
};
}
/**
* {@code bridge_send} with {@code wait:false}: delegate {@code content} and return a ticket
* immediately (fire-and-poll), so a long task isn't cut off by the caller's MCP call timeout.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
String ticket = messages.sendAsync(sessionId, content);
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
}
/** {@code bridge_poll}: check an async delegation by ticket (pending / done+reply / failed). */
static McpSchema.CallToolResult poll(MessageService messages, String ticket) {
if (isBlank(ticket)) {
return error("ticket is required");
}
MessageService.TaskView v = messages.poll(ticket);
if (v == null) {
return error("unknown ticket: " + ticket + " (never issued, or expired)");
}
return switch (v.phase()) {
case DONE -> text(v.replySource() != null && v.replySource().equals("transcript")
? "[done — worker finished without a structured bridge_reply; transcript tail follows]\n" + v.reply()
: v.reply());
case PENDING -> text("[pending — " + v.detail() + "]");
case FAILED -> text("[failed — " + v.detail() + "]");
};
}
/**
* {@code bridge_reply}: the worker returns its structured answer, resolving the awaiting send.
* {@code callerTerminal} is resolved from the connection (never an argument); a {@code null}
* means the caller is not a known worker (e.g. the primary called it by mistake).
*/
static McpSchema.CallToolResult reply(Rendezvous rendezvous, String callerTerminal, String content) {
if (callerTerminal == null) {
return error("bridge_reply is for workers only — could not identify the calling worker "
+ "from the connection");
}
if (content == null) {
return error("content is required");
}
return rendezvous.resolve(callerTerminal, content)
? text("delivered")
: error("no send is awaiting a reply for this worker");
}
/** {@code bridge_status}: the live lifecycle status of a worker session. */
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
if (isBlank(sessionId)) {
return error("sessionId is required");
}
try {
return text(messages.status(sessionId).name().toLowerCase());
} catch (HerdrException e) {
return error("herdr error for session " + sessionId + ": " + e.getMessage());
}
}
// --- fleet management logic (CB-108 / CB-301) --------------------------------------------
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
return spawn(sessions, profile, null, null, null, null);
}
/**
* {@code bridge_spawn}: launch a guard-checked worker for {@code profile} (blank → the default
* profile) and return its session id + pane id. The worker's cwd is {@code requestedCwd} if given,
* else the profile's config, else {@code callerCwd} (the primary's directory), else the daemon's.
* CB-301: the session is registered with {@code ownerTerminal} as its owner.
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
*/
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile,
String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest worktreeRequest) {
try {
WorkerSession worker = sessions.acquire(isBlank(profile) ? null : profile,
requestedCwd, callerCwd, ownerTerminal, worktreeRequest);
return text(json(workerView(worker)));
} catch (GuardException e) {
return error("subscription boundary: " + e.getMessage());
} catch (IllegalArgumentException e) {
return error(e.getMessage()); // unknown / no-default profile
} catch (HerdrException e) {
return error("herdr error spawning worker: " + e.getMessage());
}
}
/** Build a {@link WorktreeRequest} from {@code bridge_spawn}'s optional {@code worktree}/{@code ticket} args. */
private static WorktreeRequest worktreeRequest(Map<String, Object> a) {
Object w = a.get("worktree");
if (w == null || Boolean.FALSE.equals(w)) {
return null;
}
String ticket = str(a, "ticket");
if (w instanceof String s) {
if (s.isBlank() || "false".equalsIgnoreCase(s)) {
return null;
}
if ("true".equalsIgnoreCase(s)) {
if (isBlank(ticket)) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return new WorktreeRequest(s, null);
}
if (w instanceof Boolean b && b) {
if (isBlank(ticket)) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return null;
}
/** {@code bridge_profiles}: the configured worker profiles and the default. */
static McpSchema.CallToolResult profiles(ClaudeCodeLauncher workers) {
return text(json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
}
/** {@code bridge_list}: bridge-owned roster merged with live herdr status by paneId. */
static McpSchema.CallToolResult listWorkers(ClaudeCodeLauncher workers, SessionManager sessions) {
try {
Map<String, Agent> live = workers.list().stream()
.filter(a -> a.paneId() != null)
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.paneId())))
.toList();
return text(json(Map.of("workers", out)));
} catch (HerdrException e) {
return error("herdr error listing workers: " + e.getMessage());
}
}
/** {@code bridge_stop}: tear a worker down by its pane id. */
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
if (isBlank(paneId)) {
return error("paneId is required");
}
try {
sessions.release(paneId);
return text("stopped " + paneId);
} catch (HerdrException e) {
return error("herdr error stopping " + paneId + ": " + e.getMessage());
}
}
/** CB-301 projection from the authoritative session registry. */
private static Map<String, Object> workerView(WorkerSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", s.terminalId());
m.put("paneId", s.paneId());
m.put("status", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
return m;
}
private static String json(Object o) {
try {
return MAPPER.writeValueAsString(o);
} catch (Exception e) {
return String.valueOf(o);
}
}
// --- tool schemas --------------------------------------------------------------------------
private static McpSchema.Tool sendTool() {
return tool("bridge_send",
"Delegate a task to a worker session. By default blocks until the worker replies and "
+ "returns its reply (or a 'still working / queued' note on timeout). Pass wait:false "
+ "for a long task to return a ticket immediately, then poll it with bridge_poll. To "
+ "answer a worker's bridge_ask, pass its turnId (with content) instead of sessionId.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id (herdr terminal_id) to delegate to"),
"content", stringProp("The task/message to send to the worker (or your answer, with turnId)"),
"timeoutMs", Map.of("type", "integer", "description", "Max ms to wait for a reply (blocking mode)"),
"wait", Map.of("type", "boolean",
"description", "Block for the reply (default true); false returns a ticket to poll"),
"turnId", stringProp("When answering a worker's bridge_ask, its question turnId — "
+ "routes your answer back into the same turn (omit for a normal delegation)")),
List.of("content")));
}
private static McpSchema.Tool askTool() {
// No target/session arg — the worker's identity is resolved from the connection.
return tool("bridge_ask",
"Pause your current delegated turn to ask the primary a question, blocking until it "
+ "answers — then resume the same turn with the answer. Use this when only the "
+ "primary has a decision or detail you need to continue. You do not address the "
+ "primary; identity is resolved from your connection.",
objectSchema(Map.of(
"question", stringProp("The question to put to the primary"),
"timeoutMs", Map.of("type", "integer",
"description", "Max ms to wait for the primary's answer")),
List.of("question")));
}
private static McpSchema.Tool pollTool() {
return tool("bridge_poll",
"Check an async delegation (a bridge_send with wait:false) by its ticket: "
+ "pending, done (with the worker's reply), or failed.",
objectSchema(Map.of(
"ticket", stringProp("The ticket returned by bridge_send wait:false")),
List.of("ticket")));
}
private static McpSchema.Tool spawnTool() {
return tool("bridge_spawn",
"Spawn a new off-subscription worker session. Pass a profile (from bridge_profiles) to "
+ "pick the backend, or omit it for the default. The worker opens your current "
+ "directory by default; pass cwd to pin a different one. Pass worktree:true (with "
+ "ticket) or worktree:<ticket-slug> to provision an isolated git worktree. "
+ "Returns the worker's sessionId (use with bridge_send) and paneId (use with bridge_stop).",
objectSchema(Map.of(
"profile", stringProp("Worker profile to spawn (omit for the default profile)"),
"cwd", stringProp("Working directory for the worker (omit to inherit yours)"),
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
"ticket", stringProp("Ticket slug when worktree:true")),
List.of()));
}
private static McpSchema.Tool profilesTool() {
return tool("bridge_profiles",
"List the configured worker profiles (backends) and which one bridge_spawn uses by default.",
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool listTool() {
return tool("bridge_list",
"List the worker sessions the bridge tracks — each with its sessionId, paneId, profile, "
+ "state, optional worktree/branch/owner, and live herdr status.",
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool stopTool() {
return tool("bridge_stop",
"Tear down a worker session by its paneId (from bridge_spawn or bridge_list).",
objectSchema(Map.of(
"paneId", stringProp("The worker's paneId to stop")),
List.of("paneId")));
}
private static McpSchema.Tool replyTool() {
// No session/target arg — the worker's identity is resolved from the connection.
return tool("bridge_reply",
"Return your structured answer for the task you were delegated, "
+ "resolving the caller's blocked bridge_send.",
objectSchema(Map.of(
"content", stringProp("Your reply/answer")),
List.of("content")));
}
private static McpSchema.Tool statusTool() {
return tool("bridge_status",
"Get the live lifecycle status (idle/working/blocked/unknown) of a worker session.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id to query")),
List.of("sessionId")));
}
// --- small helpers -------------------------------------------------------------------------
// The SDK 2.0.0 deprecates its own Tool builders without a stable replacement — isolate it here.
@SuppressWarnings("deprecation")
private static McpSchema.Tool tool(String name, String description, Map<String, Object> inputSchema) {
return McpSchema.Tool.builder(name).description(description).inputSchema(inputSchema).build();
}
private static Map<String, Object> objectSchema(Map<String, Object> properties, List<String> required) {
return Map.of("type", "object", "properties", properties, "required", required);
}
private static Map<String, Object> stringProp(String description) {
return Map.of("type", "string", "description", description);
}
private static McpSchema.CallToolResult text(String s) {
return McpSchema.CallToolResult.builder().addTextContent(s == null ? "" : s).build();
}
private static McpSchema.CallToolResult error(String s) {
return McpSchema.CallToolResult.builder().addTextContent(s).isError(true).build();
}
private static String str(Map<String, Object> args, String key) {
Object v = args.get(key);
return v == null ? null : v.toString();
}
private static Long timeoutMs(Map<String, Object> args) {
Object v = args.get("timeoutMs");
return v instanceof Number n ? n.longValue() : null;
}
private static long clamp(long ms) {
return Math.clamp(ms, 1, MAX_TIMEOUT_MS);
}
private static boolean isBlank(String s) {
return s == null || s.isBlank();
}
}
@@ -0,0 +1,66 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.herdr.PaneLocator;
/**
* Resolves <em>who is calling</em> an MCP tool from the connection alone — the anti-spoofing
* identity model of the MCP contract. It ties the connection's loopback peer PID (from the OS)
* to a herdr agent pane (from herdr), yielding the caller's worker {@code terminal_id}. A caller
* that maps to no worker pane — the primary, or an off-host client — resolves to {@code null}.
*
* <p>Both sources are authoritative and unforgeable: the OS reports the real connecting PID, and
* herdr owns the PID→pane mapping. A worker cannot claim to be another worker, nor the primary.
* Single-host only (the herd shares the {@code bridged} host); the token path is the split-host
* fallback.
*/
public final class ConnectionIdentity {
private final PaneLocator panes;
private final PeerPidLookup pids;
private final ProcessCwdLookup cwds;
/** Identity only (no cwd resolution — {@link #cwdForPid} returns {@code null}). */
public ConnectionIdentity(PaneLocator panes, PeerPidLookup pids) {
this(panes, pids, _ -> null);
}
/** Identity plus cwd resolution (CB-112 — inherit the primary's directory on spawn). */
public ConnectionIdentity(PaneLocator panes, PeerPidLookup pids, ProcessCwdLookup cwds) {
this.panes = panes;
this.pids = pids;
this.cwds = cwds;
}
/**
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
* primary / an off-host client) and its {@code pid} (or {@code -1} if not resolvable).
*/
public record Caller(String terminal, long pid) {
}
/** Resolve the caller's terminal and PID from one peer-PID lookup. */
public Caller resolve(String remoteAddr, int remotePort) {
if (!isLoopback(remoteAddr)) {
return new Caller(null, -1); // only same-host callers can be workers
}
long pid = pids.pidForLocalPort(remotePort);
return new Caller(panes.terminalForPid(pid), pid);
}
/**
* The calling worker's {@code terminal_id}, or {@code null} if the caller is not a known
* on-host worker (treat as the primary).
*/
public String callerTerminal(String remoteAddr, int remotePort) {
return resolve(remoteAddr, remotePort).terminal();
}
/** The working directory of {@code pid} (the primary's cwd on an MCP spawn), or {@code null}. */
public String cwdForPid(long pid) {
return pid > 0 ? cwds.cwdForPid(pid) : null;
}
private static boolean isLoopback(String addr) {
return "127.0.0.1".equals(addr) || "::1".equals(addr) || "0:0:0:0:0:0:0:1".equals(addr);
}
}
@@ -0,0 +1,58 @@
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.BufferedReader;
import java.io.InputStreamReader;
import java.nio.charset.StandardCharsets;
import java.util.concurrent.TimeUnit;
/**
* {@link PeerPidLookup} via {@code lsof} (present on macOS and Linux). For a loopback TCP source
* {@code port}, both the client and this daemon appear on that port — so we exclude our own PID
* and take the other end, which is the calling process.
*/
public final class LsofPeerPidLookup implements PeerPidLookup {
private static final Logger log = LoggerFactory.getLogger(LsofPeerPidLookup.class);
private final long selfPid = ProcessHandle.current().pid();
@Override
public long pidForLocalPort(int port) {
try {
Process p = new ProcessBuilder("lsof", "-nP", "-FpP", "-iTCP:" + port)
.redirectErrorStream(true).start();
long found = -1;
try (BufferedReader r = new BufferedReader(
new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
long current = -1;
String line;
// -Fp emits records: a 'p<pid>' line, then the ports/files under that pid.
while ((line = r.readLine()) != null) {
if (line.startsWith("p")) {
current = parse(line.substring(1));
} else if (current > 0 && current != selfPid) {
found = current; // first process on this port that isn't us = the client
}
}
}
if (!p.waitFor(2, TimeUnit.SECONDS)) {
p.destroyForcibly();
}
return found;
} catch (Exception e) {
log.debug("lsof peer-pid lookup for port {} failed: {}", port, e.getMessage());
return -1;
}
}
private static long parse(String s) {
try {
return Long.parseLong(s.trim());
} catch (NumberFormatException e) {
return -1;
}
}
}
@@ -0,0 +1,48 @@
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.BufferedReader;
import java.io.InputStreamReader;
import java.nio.charset.StandardCharsets;
import java.util.concurrent.TimeUnit;
/**
* {@link ProcessCwdLookup} via {@code lsof} (present on macOS and Linux): {@code lsof -a -p <pid>
* -d cwd -Fn} prints the process's cwd on the {@code n…} line. Used to inherit the primary's
* working directory for a spawned worker (CB-112).
*/
public final class LsofProcessCwdLookup implements ProcessCwdLookup {
private static final Logger log = LoggerFactory.getLogger(LsofProcessCwdLookup.class);
@Override
public String cwdForPid(long pid) {
if (pid <= 0) {
return null;
}
try {
Process p = new ProcessBuilder("lsof", "-a", "-p", Long.toString(pid), "-d", "cwd", "-Fn")
.redirectErrorStream(true).start();
String cwd = null;
try (BufferedReader r = new BufferedReader(
new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
String line;
while ((line = r.readLine()) != null) {
if (line.startsWith("n")) { // 'n<path>' is the file-name field for the cwd fd
cwd = line.substring(1);
break;
}
}
}
if (!p.waitFor(2, TimeUnit.SECONDS)) {
p.destroyForcibly();
}
return (cwd == null || cwd.isBlank()) ? null : cwd;
} catch (Exception e) {
log.debug("lsof cwd lookup for pid {} failed: {}", pid, e.getMessage());
return null;
}
}
}
@@ -0,0 +1,13 @@
package dev.ltms.bridged.mcp;
/**
* Resolves the OS PID that owns a loopback TCP source port — the OS half of connection-based MCP
* identity. Java exposes no peer PID for a TCP socket, so this shells out. Injectable so
* {@link ConnectionIdentity} is testable without a real connection.
*/
@FunctionalInterface
public interface PeerPidLookup {
/** The PID whose socket has local (source) {@code port} on loopback, or {@code -1} if unknown. */
long pidForLocalPort(int port);
}
@@ -0,0 +1,13 @@
package dev.ltms.bridged.mcp;
/**
* Resolves a process's current working directory from its PID — the OS half of CB-112's
* "a worker inherits the primary's directory." Injectable so {@link ConnectionIdentity} stays
* testable without shelling out.
*/
@FunctionalInterface
public interface ProcessCwdLookup {
/** The working directory of {@code pid}, or {@code null} if unknown. */
String cwdForPid(long pid);
}
@@ -0,0 +1,385 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.inject.Injector;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.CompletionException;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
import java.util.concurrent.atomic.AtomicLong;
import java.util.concurrent.locks.ReentrantLock;
/**
* The blocking delegation feature (CB-104): deliver {@code content} into a worker and block until
* the worker returns a <em>structured reply</em> via {@code bridge_reply} (the {@link Rendezvous}),
* then hand that reply back. Delivery is the {@link Injector}'s job (the background poller sends it
* when the worker is injectable); this service never drives the injector or scrapes the terminal —
* completion is the worker's explicit reply, not a guess about {@code agent_status}.
*
* <p>Sends are serialized per session so exactly one reply can be outstanding per worker, which is
* what lets a reply map unambiguously to its send (no cross-talk between concurrent callers).
*
* <p>If the worker never replies within the timeout, the caller gets a typed "still working" /
* "queued" outcome — the message may still be mid-flight. A finished-but-unreplied turn is caught
* by the CB-106 completion fallback (see {@link Rendezvous#resolveCompletion}).
*
* <p><strong>Async fire-and-poll (CB-107).</strong> A caller's MCP client caps a blocking call at
* ~60s, but a real delegated task runs for minutes. {@link #sendAsync} therefore runs the same
* blocking {@link #send} on a background virtual thread and hands back a <em>ticket</em> the caller
* polls with {@link #poll}. The blocking and async paths share one code path (and the same per-target
* serialization), so async inherits the reply + completion resolution behaviour for free.
*/
public final class MessageService {
private static final Logger log = LoggerFactory.getLogger(MessageService.class);
/**
* The window a fire-and-poll send waits for resolution — generous, since no caller is blocked on
* it; a real delegated task resolves (reply or completion) well within this, and only a genuinely
* hung worker rides it out.
*/
private static final long ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L;
/** How long a finished (terminal) ticket is retained for polling before it is pruned. */
private static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
/** Outcome of a blocking send. */
public enum Outcome {
/** The worker called {@code bridge_reply}; {@code text} holds the structured answer. */
REPLIED,
/**
* The worker's delegated turn finished without a {@code bridge_reply} (CB-106 fallback);
* {@code text} is the scraped transcript tail rather than a structured answer.
*/
COMPLETED_UNREPLIED,
/**
* The worker ran the turn then wedged in an unrecoverable state (CB-109); {@code text} is the
* failure context (e.g. the error screen). Terminal, but not a successful completion.
*/
WORKER_FAILED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with
* {@link #answer(String, String, long)} and the turn resumes.
*/
QUESTION,
/** Timed out after the message was delivered — the worker is still working. */
TIMED_OUT_WORKING,
/** Timed out before delivery — the message is still queued for the worker. */
TIMED_OUT_QUEUED,
/** Another send to this session was in flight for the whole window. */
BUSY,
/**
* An answer ({@link #answer(String, String, long)}) referenced a {@code turnId} that is no
* longer open — the worker's {@code bridge_ask} already timed out or was answered.
*/
STALE_TURN
}
/**
* @param outcome how the send ended (or paused)
* @param text the worker's answer when {@link #completed()} (a structured {@code bridge_reply}
* for {@link Outcome#REPLIED}, a scraped transcript tail for
* {@link Outcome#COMPLETED_UNREPLIED}), or the question for {@link Outcome#QUESTION},
* else {@code null}
* @param turnId correlation id for a {@link Outcome#QUESTION} (answered via
* {@link #answer(String, String, long)}), else {@code null}
*/
public record Reply(Outcome outcome, String text, String turnId) {
/** A reply with no correlation id (the common terminal outcomes). */
public Reply(Outcome outcome, String text) {
this(outcome, text, null);
}
/** Whether the worker's turn actually finished with an answer (replied or scraped). */
public boolean completed() {
return outcome == Outcome.REPLIED || outcome == Outcome.COMPLETED_UNREPLIED;
}
}
/** How a worker's {@code bridge_ask} (CB-205) resolved. */
public enum AskOutcome {
/** The primary answered; {@link AskResult#answer} carries it. */
ANSWERED,
/** No delegation was open to surface the question to — the worker has no one to ask. */
NO_WAITER,
/** The primary did not answer within the window. */
TIMED_OUT
}
/** The outcome of a worker's {@code bridge_ask}: how it resolved and (if answered) the answer. */
public record AskResult(AskOutcome outcome, String answer) {
}
/** Lifecycle phase of an async delegation ticket. */
public enum Phase {
/** Delegated and in flight — queued for the worker or being worked. */
PENDING,
/** The worker's turn finished; {@link TaskView#reply} holds the answer. */
DONE,
/** The delegation could not complete (timed out, worker gone, or busy). */
FAILED
}
/**
* A poll snapshot of an async delegation.
*
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, else {@code null}
* @param replySource {@code "reply"} (structured {@code bridge_reply}) or {@code "transcript"}
* (completion scrape) when {@link Phase#DONE}, else {@code null}
* @param detail a human note (live worker status while pending, or the failure reason)
*/
public record TaskView(String ticket, Phase phase, String reply, String replySource, String detail) {
}
/** An in-flight or finished async delegation, keyed by its ticket. */
private record Task(String target, CompletableFuture<Reply> future, long createdNanos) {
}
private final AgentControl agents;
private final Injector injector;
private final Rendezvous rendezvous;
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
Thread.ofVirtual().name("bridge-async-", 0).factory());
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous) {
this.agents = agents;
this.injector = injector;
this.rendezvous = rendezvous;
}
/** Current lifecycle status of a worker (the {@code GET /sessions/{id}/status} surface). */
public AgentStatus status(String target) {
return agents.status(target);
}
/**
* Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
*/
public Reply send(String target, String content, long timeoutMillis) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
}
try {
CompletableFuture<Void> delivered = injector.enqueue(target, content);
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
return new Reply(wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
} finally {
rendezvous.close(target, reply);
}
} finally {
lock.unlock();
}
}
/**
* A worker's mid-turn question (CB-205 reverse rendezvous): surface {@code question} to the
* primary by resolving its open blocking {@code bridge_send}, then block this (worker) call until
* the primary answers via {@link #answer} or {@code timeoutMillis} elapses. Identity is the
* worker's own session — it does not address the primary.
*
* <p>Returns {@link AskOutcome#NO_WAITER} when no delegation is open to surface the question to
* (nothing to answer it), {@link AskOutcome#ANSWERED} with the primary's answer, or
* {@link AskOutcome#TIMED_OUT} if the primary stayed silent. The worker resumes its turn either
* way — an answered ask hands back the answer; an unanswered one leaves it to proceed alone.
*/
public AskResult ask(String workerSession, String question, long timeoutMillis) {
Rendezvous.AskTicket ticket = rendezvous.openAsk(workerSession);
// Only the freshly-opening caller surfaces the question; a coalesced duplicate simply blocks on
// the shared answer future that the fresh owner is already responsible for.
if (ticket.fresh()) {
// Register the reverse waiter first, then surface the question — so the answer, which can
// arrive the instant the primary reacts, always finds an open waiter to resolve.
if (!rendezvous.resolveQuestion(workerSession, question, ticket.turnId())) {
rendezvous.closeAsk(ticket.turnId());
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
}
}
try {
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
return new AskResult(AskOutcome.ANSWERED, answer);
} catch (TimeoutException e) {
log.debug("bridge_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
return new AskResult(AskOutcome.TIMED_OUT, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting the primary's answer for " + workerSession, e);
} finally {
// Only the fresh owner tears down the shared turn; a duplicate must leave it open.
if (ticket.fresh()) {
rendezvous.closeAsk(ticket.turnId());
}
}
}
/**
* The primary's answer to a worker's {@code bridge_ask} (CB-205): resolve the worker's blocked
* question identified by {@code turnId}, then — like a fresh {@link #send} — block for the worker's
* eventual {@code bridge_reply} as it finishes the resumed turn. The worker session is derived from
* {@code turnId}, never a caller argument.
*
* <p>Unlike {@link #send} this does not re-inject through the {@link Injector}: the worker is
* mid-turn (already picked up), so the answer flows back through its own open {@code bridge_ask}
* call, not a new status-gated delivery. The forward waiter is opened <em>before</em> the worker
* is unblocked so a reply that lands the instant it resumes is not lost.
*/
public Reply answer(String turnId, String content, long timeoutMillis) {
String workerSession = rendezvous.askSession(turnId);
if (workerSession == null) {
return new Reply(Outcome.STALE_TURN, null); // the ask lapsed (timed out or already answered)
}
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(workerSession, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null);
}
try {
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession);
if (!rendezvous.answerAsk(turnId, content)) {
rendezvous.close(workerSession, reply);
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
}
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
} catch (TimeoutException e) {
// The worker resumed but hasn't replied yet — no completion fallback arms an answered
// turn (it never re-entered the injector), so a silent worker rides out the window.
return new Reply(Outcome.TIMED_OUT_WORKING, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + workerSession, e);
} finally {
rendezvous.close(workerSession, reply);
}
} finally {
lock.unlock();
}
}
/**
* Fire-and-poll variant of {@link #send}: deliver {@code content} to {@code target} on a
* background virtual thread and return immediately with a ticket to {@link #poll}. This is how a
* long task is delegated without tripping the caller's MCP client call timeout.
*
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content) {
String ticket = "task-" + ticketSeq.incrementAndGet();
CompletableFuture<Reply> future =
CompletableFuture.supplyAsync(() -> send(target, content, ASYNC_TIMEOUT_MS), asyncExecutor);
tasks.put(ticket, new Task(target, future, System.nanoTime()));
pruneTerminalTickets();
log.debug("async send {} -> {}", ticket, target);
return ticket;
}
/**
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket;
* otherwise a {@link Phase#PENDING} view (with the live worker status as detail), a
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason.
*/
public TaskView poll(String ticket) {
Task task = tasks.get(ticket);
if (task == null) {
return null;
}
CompletableFuture<Reply> f = task.future();
if (!f.isDone()) {
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target()));
}
Reply r;
try {
r = f.getNow(null);
} catch (CompletionException | java.util.concurrent.CancellationException e) {
Throwable cause = (e instanceof CompletionException ce && ce.getCause() != null) ? ce.getCause() : e;
return new TaskView(ticket, Phase.FAILED, null, null, cause.getMessage());
}
if (r.completed()) {
String source = r.outcome() == Outcome.REPLIED ? "reply" : "transcript";
return new TaskView(ticket, Phase.DONE, r.text(), source, null);
}
// A wedged worker (CB-109) carries the error context as its reason; the timeout/busy
// outcomes carry none, so fall back to the outcome name.
String detail = r.outcome() == Outcome.WORKER_FAILED && r.text() != null
? r.text()
: "no reply — " + r.outcome().name().toLowerCase();
return new TaskView(ticket, Phase.FAILED, null, null, detail);
}
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
private String liveStatus(String target) {
try {
return agents.status(target).name().toLowerCase();
} catch (RuntimeException e) {
return "unknown";
}
}
/** Drop finished tickets older than the TTL so the registry cannot grow without bound. */
private void pruneTerminalTickets() {
long cutoff = System.nanoTime() - TICKET_TTL_NANOS;
tasks.values().removeIf(t -> t.future().isDone() && t.createdNanos() < cutoff);
}
/** Release the async executor. */
public void close() {
asyncExecutor.shutdown();
}
/** Map a rendezvous {@link Rendezvous.Kind} onto its send {@link Outcome} (shared by send/answer). */
private static Outcome outcomeOf(Rendezvous.Kind kind) {
return switch (kind) {
case REPLY -> Outcome.REPLIED;
case COMPLETION -> Outcome.COMPLETED_UNREPLIED;
case FAILED -> Outcome.WORKER_FAILED;
case QUESTION -> Outcome.QUESTION;
};
}
private static boolean tryLock(ReentrantLock lock, long millis) {
try {
return lock.tryLock(Math.max(0, millis), TimeUnit.MILLISECONDS);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting the session send lock", e);
}
}
private static long remainingMillis(long deadlineNanos) {
return (deadlineNanos - System.nanoTime()) / 1_000_000L;
}
}
@@ -0,0 +1,216 @@
package dev.ltms.bridged.msg;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicLong;
/**
* The reply rendezvous: where a blocking {@code bridge_send} awaits how the worker's delegated turn
* ends. The sending (primary) request thread {@link #open}s a waiter; it is resolved either by the
* worker's explicit {@code bridge_reply} ({@link #resolve}, arriving on a different thread via
* {@code POST /sessions/{id}/reply}) or — the CB-106 fallback — by the injector observing the
* worker's delegated turn return to idle without a reply ({@link #resolveCompletion}).
*
* <p>At most one waiter per session — {@link MessageService} serializes sends per session, so a
* resolution maps unambiguously to the one outstanding send and cannot be captured by another.
*
* <p><strong>Waiter identity (CB-116).</strong> The completion/failure fallbacks run asynchronously
* and can fire <em>after</em> the turn they belong to has already been resolved by an explicit reply
* and a <em>next</em> send has opened its own waiter on the same session. Resolving "whatever waiter
* is registered now" would then land turn N's stale scrape on turn N+1's send. So those fallbacks
* resolve a <em>specific</em> {@link CompletableFuture} captured when their turn was delivered
* ({@link #resolveCompletion(CompletableFuture, String)} /
* {@link #resolveFailure(CompletableFuture, String)}): a no-op if that waiter was already resolved,
* and it can never touch a later send's waiter.
*/
public final class Rendezvous {
/** How a delegated turn ended (or paused). */
public enum Kind {
/** The worker called {@code bridge_reply} with a structured answer. */
REPLY,
/** The worker's turn finished without a {@code bridge_reply}; {@code text} is a scrape. */
COMPLETION,
/** The worker ran the turn then wedged (CB-109); {@code text} is the failure context. */
FAILED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205 reverse rendezvous);
* {@code text} is the question and {@code turnId} correlates the primary's answer back to
* the worker's blocked {@code bridge_ask}. Not terminal — the turn resumes after the answer.
*/
QUESTION
}
/**
* The resolved outcome of a send: its {@link Kind}, the associated text, and — only for
* {@link Kind#QUESTION} — the {@code turnId} the primary answers with (else {@code null}).
*/
public record Resolution(Kind kind, String text, String turnId) {
/** A terminal resolution (reply / completion / failure) with no correlation id. */
public Resolution(Kind kind, String text) {
this(kind, text, null);
}
}
/** A worker's open mid-turn question: the worker session it belongs to and the answer future. */
private record AskWaiter(String session, CompletableFuture<String> answer) {
}
/**
* Handle to a reverse-rendezvous turn: the {@code turnId}, its answer future, and whether this
* call freshly opened it (versus coalescing onto an already-open ask).
*/
public record AskTicket(String turnId, CompletableFuture<String> answer, boolean fresh) {
}
private final ConcurrentHashMap<String, CompletableFuture<Resolution>> waiters = new ConcurrentHashMap<>();
/** Reverse rendezvous (CB-205): worker questions awaiting the primary's answer, keyed by {@code turnId}. */
private final ConcurrentHashMap<String, AskWaiter> asks = new ConcurrentHashMap<>();
private final AtomicLong askSeq = new AtomicLong();
/** Per-session index of the currently-open ask, so duplicate bridge_ask calls coalesce onto one turn. */
private final ConcurrentHashMap<String, String> openAsksBySession = new ConcurrentHashMap<>();
/**
* Register a waiter for {@code session} — the await side of the public {@code resolve*} methods.
* The caller must hold that session's send lock.
*/
public CompletableFuture<Resolution> open(String session) {
CompletableFuture<Resolution> waiter = new CompletableFuture<>();
waiters.put(session, waiter);
return waiter;
}
/** Remove {@code waiter} for {@code session} (only if it is still the registered one). */
void close(String session, CompletableFuture<Resolution> waiter) {
waiters.remove(session, waiter);
}
/** Whether a send is currently awaiting a resolution for {@code session}. */
public boolean isWaiting(String session) {
return waiters.containsKey(session);
}
/**
* The waiter currently registered for {@code session}, or {@code null} if none is waiting. The
* completion/failure fallbacks capture this at delivery time so they can later resolve that exact
* send (see the CB-116 note above) rather than whichever send happens to be waiting when they fire.
*/
public CompletableFuture<Resolution> currentWaiter(String session) {
return waiters.get(session);
}
/**
* Resolve the send awaiting on {@code session} with the worker's explicit reply {@code content}.
*
* @return {@code true} if a waiter was resolved; {@code false} if none was waiting (a late or
* spurious reply — e.g. the send already timed out)
*/
public boolean resolve(String session, String content) {
return complete(session, new Resolution(Kind.REPLY, content));
}
// --- reverse rendezvous (CB-205 bridge_ask) ------------------------------------------------
/**
* Open a reverse-rendezvous waiter for a worker's mid-turn question. If {@code session} already has
* an open ask, coalesce onto it (same {@code turnId}, same answer future). Otherwise atomically mint
* a fresh {@code turnId}, register it in both the per-turn and per-session indexes, and hand it back
* marked fresh. The caller then {@link #resolveQuestion surfaces the question} to the primary and
* blocks on the returned future until the primary {@link #answerAsk answers}.
*/
public AskTicket openAsk(String session) {
while (true) {
AskWaiter[] minted = { null };
String turnId = openAsksBySession.computeIfAbsent(session, _ -> {
String newTurnId = session + "#" + askSeq.incrementAndGet();
CompletableFuture<String> answer = new CompletableFuture<>();
AskWaiter waiter = new AskWaiter(session, answer);
asks.put(newTurnId, waiter);
minted[0] = waiter;
return newTurnId;
});
if (minted[0] != null) {
return new AskTicket(turnId, minted[0].answer(), true);
}
AskWaiter existing = asks.get(turnId);
if (existing != null) {
return new AskTicket(turnId, existing.answer(), false);
}
// A close raced and removed the waiter after we read the turnId; clear the stale index entry
// and retry so a fresh ask is always backed by a registered waiter.
openAsksBySession.remove(session, turnId);
}
}
/**
* Surface a worker's mid-turn {@code question} by resolving the primary's open {@code bridge_send}
* with a {@link Kind#QUESTION} carrying {@code turnId}. Same session-keyed semantics as
* {@link #resolve}: the one outstanding send for {@code session} unblocks with the question.
*
* @return {@code true} if a send was awaiting (the question reached the primary); {@code false}
* if none was (no delegation is open to answer it)
*/
public boolean resolveQuestion(String session, String question, String turnId) {
return complete(session, new Resolution(Kind.QUESTION, question, turnId));
}
/** The worker session an outstanding ask {@code turnId} belongs to, or {@code null} if unknown/lapsed. */
public String askSession(String turnId) {
AskWaiter w = asks.get(turnId);
return w == null ? null : w.session();
}
/**
* Resolve a worker's blocked {@code bridge_ask} with the primary's {@code answer}, unblocking it
* to resume its turn.
*
* @return {@code true} if the ask was still open and got the answer; {@code false} if the
* {@code turnId} is unknown or the ask already lapsed (timed out / was answered)
*/
public boolean answerAsk(String turnId, String answer) {
AskWaiter w = asks.get(turnId);
return w != null && w.answer().complete(answer);
}
/** Drop a reverse-rendezvous turn once its {@code bridge_ask} has resolved (answered or lapsed). */
public void closeAsk(String turnId) {
AskWaiter w = asks.get(turnId);
if (w == null) {
return;
}
// Remove the session index first and only if it still points to this turn, so a concurrent
// fresh ask cannot inherit a waiter we are about to drop.
openAsksBySession.remove(w.session(), turnId);
asks.remove(turnId);
}
/**
* Resolve a specific captured {@code waiter} as a completion (the delegated turn finished with no
* {@code bridge_reply}); {@code text} is the scraped transcript tail. The waiter is the one
* captured when this turn was delivered, so a late completion for turn N cannot land on turn N+1's
* send (CB-116). A no-op if that waiter was already resolved — a raced {@code bridge_reply} wins.
*
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
*/
public boolean resolveCompletion(CompletableFuture<Resolution> waiter, String text) {
return waiter != null && waiter.complete(new Resolution(Kind.COMPLETION, text));
}
/**
* Resolve a specific captured {@code waiter} as a failure — the worker ran the turn but wedged in
* an unrecoverable state (CB-109); {@code reason} is the failure context (e.g. the error screen).
* Like {@link #resolveCompletion(CompletableFuture, String)} it targets the exact captured send
* (CB-116). A no-op if that waiter was already resolved — first resolution wins.
*
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
*/
public boolean resolveFailure(CompletableFuture<Resolution> waiter, String reason) {
return waiter != null && waiter.complete(new Resolution(Kind.FAILED, reason));
}
private boolean complete(String session, Resolution resolution) {
CompletableFuture<Resolution> waiter = waiters.get(session);
return waiter != null && waiter.complete(resolution);
}
}
@@ -0,0 +1,35 @@
package dev.ltms.bridged.peer;
/**
* Declared capabilities of a {@link PeerLauncher}. The protocol is the union across all
* configured launchers; a verb invoked against a peer that lacks the capability returns a clean
* "unsupported for this peer" rather than a crash. Capabilities keep the protocol honest as peers
* diversify and prevent the core from assuming "every peer is a Claude in a worktree."
*/
public enum Capability {
/**
* The peer supports {@code bridge_ask} rendezvous — pausing its delegated turn to ask
* the primary a question, then resuming once answered. All Claude Code peers support this.
*/
MID_TURN_ASK,
/**
* The peer can open its own PR at the end of an implementation turn (CB-302). Opt-in per
* profile: granted only when the profile carries a git-forge token ({@code gitTokenEnv}).
*/
SELF_PR,
/**
* The peer can run inside a provisioned isolated git worktree. All CLI-based peers support
* this since their cwd is set at spawn time.
*/
WORKTREE,
/**
* The spawner can reconcile orphaned peers on boot — workers that outlived a prior daemon
* process and whose pane ids died with it (CB-117). Claude Code over herdr supports this
* via name-based matching against the herdr agent list.
*/
ORPHAN_REAP
}
@@ -0,0 +1,29 @@
package dev.ltms.bridged.peer;
/**
* An opaque handle returned by {@link PeerLauncher#spawn(SpawnRequest)}. The core routes on
* {@link #id()} (the registry/routing key) and uses {@link #terminalId()} for session tracking;
* launcher-private coordinates beyond these are reachable through the concrete implementation.
*
* <p>A {@link PeerHandle} is returned <em>after</em> the peer process is live — the launcher
* has already completed subscription-guarded env/vfs setup, process start, and placement. The
* handle is a ticket the core exchanges for the running peer, not a lazy/delayed reference.
*/
public interface PeerHandle {
/**
* The registry/routing key — an opaque, launcher-assigned identifier. For the herdr-backed
* launcher this is the herdr pane id; for other launchers it is whatever their transport
* uses. Guaranteed to be non-null and unique among live peers within a single daemon process.
*/
String id();
/**
* The transport-level session identifier used for message routing and presence tracking.
* For the herdr launcher this is the herdr terminal UUID. A non-herdr launcher may return
* its own analogous identifier, or {@code null} if the concept does not apply.
*/
default String terminalId() {
return null;
}
}
@@ -0,0 +1,85 @@
package dev.ltms.bridged.peer;
import java.util.List;
import java.util.Set;
/**
* SPI for materializing a connected peer — the only way the bridge core creates or tears down
* a peer process. Every launcher is a first-party, in-tree adapter selected by (future) profile
* config; today's single adapter is the {@code ClaudeCodeLauncher} / Claude Code over herdr.
*
* <p>The core delegates spawn and teardown to this interface without knowing how the peer is set
* up. Environment variables, CLI flags, subscription guards, transport (herdr tab/pane) layout,
* and naming conventions are all adapter-private — the core sees only the returned
* {@link PeerHandle} whose {@code id()} is the registry/routing key.
*
* <p>The interface is a superset of what {@code SessionManager} and {@code Bridged.main} call
* on the concrete launcher today.
*/
public interface PeerLauncher {
/**
* The set of {@link Capability capabilities} this launcher declares. A peer whose profile
* opts into a git-forge token should include {@link Capability#SELF_PR}; the base set for
* the Claude Code herdr adapter is always {@code MID_TURN_ASK, WORKTREE, ORPHAN_REAP}.
*/
Set<Capability> capabilities();
/**
* {@code profileName}/requestedCwd null/blank → default resolution. Returns after the peer
* process is live (env + argv + placement complete). Never returns {@code null}.
*
* @param req the spawn parameters (profile, requested cwd, caller cwd)
* @return a handle whose {@link PeerHandle#id()} is the registry/routing key
* @throws IllegalArgumentException if the profile is unknown and no default is configured
*/
PeerHandle spawn(SpawnRequest req);
/**
* The configured worker profile names — the set of names {@code spawn(profileName)} accepts.
*/
Set<String> profiles();
/**
* The profile a no-argument {@link #spawn(SpawnRequest)} uses, or {@code null} if none is configured.
*/
String defaultProfile();
/**
* Resolve the effective working directory for a spawn {@code req} without actually spawning.
* Resolution order: requestedCwd → profile cwd → callerCwd → daemon cwd.
*
* @return the resolved absolute path, never null/blank
*/
String effectiveCwd(SpawnRequest req);
/**
* The parity-overlay file list for {@code profileName} (default list when unset). Used by
* worktree provisioning to copy config files into the isolated checkout before spawning.
*/
List<String> parityOverlay(String profileName);
/**
* The set of all agents this launcher currently tracks, transport-specific. Each element
* exposes at minimum a pane-like {@code id()} matching this launcher's {@link PeerHandle}
* scheme, plus transport-level status. Callers merge this set with the session registry to
* build a live roster view.
*/
List<?> list();
/**
* Reap orphaned peers left behind by a prior daemon process. Only peers whose naming scheme
* matches this launcher's and whose nonce differs from the current process are eligible.
* Best-effort: a failure to list or to stop any one peer is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
int reapOrphanWorkers();
/**
* Tear a peer down by its registry/routing key ({@link PeerHandle#id()}). Tolerates an
* already-gone peer. Also cleans up launcher-private resources (e.g. empty dedicated tabs)
* when safe to do so.
*/
void stop(String id);
}
@@ -0,0 +1,13 @@
package dev.ltms.bridged.peer;
/**
* Parameters for a {@link PeerLauncher#spawn(SpawnRequest)} call — the peer-neutral
* aggregation of what the core knows at delegation time: which profile to use, the caller's
* requested working directory, and the caller's own cwd (to inherit when no other cwd is set).
*
* <p>A null or blank {@code profileName} means "use the launcher's default profile."
* A null or blank {@code requestedCwd} means "inherit from config or caller."
* A null {@code callerCwd} means "the request came from the daemon itself (not a primary)."
*/
public record SpawnRequest(String profileName, String requestedCwd, String callerCwd) {
}
@@ -1,18 +1,29 @@
package dev.ltms.bridged.rest;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.worker.WorkerService;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.javalin.Javalin;
import io.javalin.http.Context;
import jakarta.servlet.http.HttpServlet;
import org.eclipse.jetty.servlet.ServletHolder;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.function.Function;
import java.util.stream.Collectors;
/**
* The REST surface — {@code bridged}'s contract and its testability seam. Every
@@ -25,22 +36,56 @@ import java.util.Map;
*/
public final class BridgedApp {
private final HerdrClient herdr;
private final WorkerService workers;
/** Default blocking window for a message; kept under typical HTTP idle timeouts. */
private static final long DEFAULT_MESSAGE_TIMEOUT_MS = 25_000;
private static final long MAX_MESSAGE_TIMEOUT_MS = 120_000;
/** Blocking window for a worker's bridge_ask (CB-205); the worker's MCP client caps its own call. */
private static final long DEFAULT_ASK_TIMEOUT_MS = 55_000;
private static final long MAX_ASK_TIMEOUT_MS = 115_000;
public BridgedApp(HerdrClient herdr, WorkerService workers) {
private final HerdrClient herdr;
private final ClaudeCodeLauncher workers;
private final SessionManager sessions; // CB-301: authoritative session registry
private final MessageService messages;
private final Rendezvous rendezvous;
private final WorkerPresence presence; // CB-113: which workers are MCP-connected (available)
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
private final ObjectMapper mapper = new ObjectMapper();
public BridgedApp(HerdrClient herdr, ClaudeCodeLauncher workers, SessionManager sessions,
MessageService messages, Rendezvous rendezvous, WorkerPresence presence,
HttpServlet mcpServlet) {
this.herdr = herdr;
this.workers = workers;
this.sessions = sessions;
this.messages = messages;
this.rendezvous = rendezvous;
this.presence = presence;
this.mcpServlet = mcpServlet;
}
/** Wire routes onto a fresh, unstarted Javalin instance. Caller starts it. */
public Javalin build() {
Javalin app = Javalin.create(cfg -> cfg.showJavalinBanner = false);
Javalin app = Javalin.create(cfg -> {
cfg.showJavalinBanner = false;
if (mcpServlet != null) {
// The MCP server shares the daemon's port; Jetty routes /mcp to its servlet.
cfg.jetty.modifyServletContextHandler(h ->
h.addServlet(new ServletHolder(mcpServlet), "/mcp"));
}
});
app.get("/healthz", this::healthz);
app.get("/sessions", this::sessions);
app.get("/agents", this::agents);
app.post("/workers", this::spawnWorker);
app.get("/workers", this::listWorkers); // CB-304: registry roster + live herdr status
app.get("/profiles", this::profiles); // configured worker profiles
app.post("/workers", this::spawnWorker); // optional ?profile= or {"profile":…}
app.delete("/workers/{paneId}", this::stopWorker);
app.post("/sessions/{id}/message", this::sendMessage); // bridge_send (primary; blocking, wait:false, or answer via turnId)
app.post("/sessions/{id}/reply", this::replyMessage); // bridge_reply (worker)
app.post("/sessions/{id}/ask", this::askMessage); // bridge_ask (worker → primary, CB-205)
app.get("/sessions/{id}/status", this::sessionStatus); // bridge_status
app.get("/tasks/{ticket}", this::taskStatus); // poll an async (wait:false) send
return app;
}
@@ -81,22 +126,269 @@ public final class BridgedApp {
ctx.status(200).json(Map.of("agents", workers.list().stream().map(BridgedApp::view).toList()));
}
/** Spawn a guard-checked worker. 403 if the base_url would breach the subscription boundary. */
/** CB-304: bridge-owned roster merged with live herdr status by paneId. */
private void listWorkers(Context ctx) {
Map<String, Agent> live = workers.list().stream()
.filter(a -> a.paneId() != null)
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.paneId())))
.toList();
ctx.status(200).json(Map.of("workers", out));
}
/** The configured worker profiles and which one a no-argument spawn uses. */
private void profiles(Context ctx) {
ctx.status(200).json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile()));
}
/**
* Spawn a guard-checked worker. An optional {@code profile} (query param or {@code {"profile":…}}
* body) picks which configured profile; omitted → the default. 403 if the base_url would breach
* the subscription boundary, 400 for an unknown profile.
*/
private void spawnWorker(Context ctx) {
String profile = ctx.queryParam("profile");
String cwd = ctx.queryParam("cwd");
String worktree = ctx.queryParam("worktree");
String ticket = ctx.queryParam("ticket");
if (profile == null || profile.isBlank() || cwd == null || cwd.isBlank()
|| worktree == null || worktree.isBlank()) {
try {
String body = ctx.body();
if (!body.isBlank()) {
JsonNode b = mapper.readTree(body);
if (profile == null || profile.isBlank()) profile = b.path("profile").asText(null);
if (cwd == null || cwd.isBlank()) cwd = b.path("cwd").asText(null);
if (worktree == null || worktree.isBlank()) worktree = b.path("worktree").asText(null);
if (ticket == null || ticket.isBlank()) ticket = b.path("ticket").asText(null);
}
} catch (Exception ignored) {
// A malformed/empty body just means "no overrides" → fall through to defaults.
}
}
WorktreeRequest wt = worktreeRequest(worktree, ticket);
try {
Agent worker = workers.spawn();
// No MCP caller over REST, so callerCwd and ownerTerminal are null.
WorkerSession worker = sessions.acquire(blankToNull(profile), blankToNull(cwd), null, null, wt);
ctx.status(201).json(view(worker));
} catch (GuardException e) {
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
} catch (IllegalArgumentException e) {
ctx.status(400).json(Map.of("error", "unknown_profile", "detail", e.getMessage()));
}
}
private static WorktreeRequest worktreeRequest(String worktree, String ticket) {
if (worktree == null || worktree.isBlank() || "false".equalsIgnoreCase(worktree)) {
return null;
}
if ("true".equalsIgnoreCase(worktree)) {
if (ticket == null || ticket.isBlank()) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return new WorktreeRequest(worktree, null);
}
private static String blankToNull(String s) {
return (s == null || s.isBlank()) ? null : s;
}
/** Tear a worker down by pane id. */
private void stopWorker(Context ctx) {
workers.stop(ctx.pathParam("paneId"));
sessions.release(ctx.pathParam("paneId"));
ctx.status(204);
}
/**
* The blocking delegation call (CB-104): inject {@code content} into the worker via the
* status-gated injector and block until the worker returns a structured {@code bridge_reply}.
* Times out with a typed 202 (working / queued / busy) rather than an error — the message may
* still land.
*/
private void sendMessage(Context ctx) {
String id = ctx.pathParam("id");
String content;
String turnId;
long timeout;
boolean wait;
try {
JsonNode body = mapper.readTree(ctx.body());
content = body.path("content").asText("");
turnId = body.path("turnId").asText(null);
timeout = body.path("timeoutMs").asLong(DEFAULT_MESSAGE_TIMEOUT_MS);
wait = body.path("wait").asBoolean(true); // default: block for the reply (CB-104)
} catch (Exception e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
if (content.isBlank()) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "content is required"));
return;
}
timeout = Math.clamp(timeout, 1, MAX_MESSAGE_TIMEOUT_MS);
// Answering a worker's bridge_ask (CB-205): always blocks, and derives the worker from turnId.
if (turnId != null && !turnId.isBlank()) {
writeReply(ctx, id, messages.answer(turnId, content, timeout), timeout);
return;
}
if (!wait) {
// Fire-and-poll (CB-107): return a ticket immediately; the caller polls GET /tasks/{ticket}.
String ticket = messages.sendAsync(id, content);
ctx.status(202).json(Map.of("sessionId", id, "ticket", ticket, "status", "accepted"));
return;
}
try {
writeReply(ctx, id, messages.send(id, content, timeout), timeout);
} catch (HerdrException e) {
herdrError(ctx, e);
}
}
/**
* Render a {@link MessageService.Reply} onto the response — shared by a normal send and a
* bridge_ask answer. A structured/scraped completion is 200; a worker's mid-turn question a 202
* (with its {@code turnId}); a stale answer a 409; every other non-terminal outcome a typed 202.
*/
private void writeReply(Context ctx, String id, MessageService.Reply reply, long timeout) {
switch (reply.outcome()) {
case QUESTION -> ctx.status(202).json(Map.of(
"sessionId", id, "status", "question",
"question", reply.text(), "turnId", reply.turnId()));
case STALE_TURN -> ctx.status(409).json(Map.of(
"sessionId", id, "error", "stale_turn",
"detail", "that question is no longer open (timed out or already answered)"));
case REPLIED, COMPLETED_UNREPLIED -> {
// replySource distinguishes a structured bridge_reply from the CB-106 completion
// fallback (a scrape of the worker's transcript when it finished without replying).
String source = reply.outcome() == MessageService.Outcome.REPLIED ? "reply" : "transcript";
ctx.status(200).json(Map.of("sessionId", id, "reply", reply.text(), "replySource", source));
}
default -> ctx.status(202).json(Map.of(
"sessionId", id,
"status", switch (reply.outcome()) {
case TIMED_OUT_WORKING -> "working";
case TIMED_OUT_QUEUED -> "queued";
case BUSY -> "busy";
case WORKER_FAILED -> "failed";
default -> "done"; // unreachable (terminal outcomes handled above)
},
"detail", reply.outcome() == MessageService.Outcome.WORKER_FAILED && reply.text() != null
? reply.text()
: "no reply within " + timeout + "ms; poll status or retry"));
}
}
/**
* A worker's mid-turn question ({@code bridge_ask}, CB-205) — surfaces to the primary's open
* blocking send and blocks until it answers. 200 with the answer, 409 if no delegation is open,
* 202 if the primary stayed silent.
*/
private void askMessage(Context ctx) {
String id = ctx.pathParam("id");
String question;
long timeout;
try {
JsonNode body = mapper.readTree(ctx.body());
question = body.path("question").asText("");
timeout = body.path("timeoutMs").asLong(DEFAULT_ASK_TIMEOUT_MS);
} catch (Exception e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
if (question.isBlank()) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "question is required"));
return;
}
timeout = Math.clamp(timeout, 1, MAX_ASK_TIMEOUT_MS);
MessageService.AskResult r = messages.ask(id, question, timeout);
switch (r.outcome()) {
case ANSWERED -> ctx.status(200).json(Map.of("sessionId", id, "answered", true, "answer", r.answer()));
case NO_WAITER -> ctx.status(409).json(Map.of(
"sessionId", id, "error", "no_pending_send",
"detail", "no primary is awaiting this turn to answer a question"));
case TIMED_OUT -> ctx.status(202).json(Map.of(
"sessionId", id, "status", "no_answer",
"detail", "the primary did not answer within " + timeout + "ms"));
}
}
/**
* The worker's structured reply ({@code bridge_reply}) — resolves the blocking send awaiting
* on this session. 200 if a send was waiting, 409 if none was (late or spurious reply).
*/
private void replyMessage(Context ctx) {
String id = ctx.pathParam("id");
String content;
try {
content = mapper.readTree(ctx.body()).path("content").asText("");
} catch (Exception e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
if (rendezvous.resolve(id, content)) {
ctx.status(200).json(Map.of("sessionId", id, "delivered", true));
} else {
ctx.status(409).json(Map.of(
"sessionId", id, "error", "no_pending_send",
"detail", "no send is awaiting a reply for this session"));
}
}
/**
* Live lifecycle status of a worker (MCP `bridge_status` wraps this in CB-105), plus its
* <em>readiness</em> (CB-113): {@code ready} is true once the worker's Claude has connected the
* bridge MCP — the reliable "available to receive a task" signal, unlike bare {@code idle}, which
* is also true during boot.
*/
private void sessionStatus(Context ctx) {
String id = ctx.pathParam("id");
try {
ctx.status(200).json(Map.of(
"sessionId", id,
"status", messages.status(id).name().toLowerCase(),
"ready", presence.isPresent(id)));
} catch (HerdrException e) {
herdrError(ctx, e);
}
}
/** Poll an async (wait:false) delegation by ticket. 404 for an unknown/expired ticket. */
private void taskStatus(Context ctx) {
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"));
if (v == null) {
ctx.status(404).json(Map.of("error", "unknown_ticket", "detail", "no such task (or it has expired)"));
return;
}
Map<String, Object> body = new LinkedHashMap<>();
body.put("ticket", v.ticket());
body.put("phase", v.phase().name().toLowerCase());
if (v.reply() != null) {
body.put("reply", v.reply());
body.put("replySource", v.replySource());
}
if (v.detail() != null) {
body.put("detail", v.detail());
}
ctx.status(200).json(body);
}
/** Map a herdr failure: unknown target → 404, anything else → 502 (herdr is upstream). */
private static void herdrError(Context ctx, HerdrException e) {
if (e.code() != null && e.code().endsWith("_not_found")) {
ctx.status(404).json(Map.of("error", "session_not_found", "detail", e.getMessage()));
} else {
ctx.status(502).json(Map.of("error", "herdr_error", "detail", e.getMessage()));
}
}
/** Stable JSON projection of an agent (null-safe for the start-time shape). */
private static Map<String, Object> view(Agent a) {
Map<String, Object> m = new LinkedHashMap<>();
@@ -109,4 +401,22 @@ public final class BridgedApp {
m.put("status", a.status().name().toLowerCase());
return m;
}
/** CB-301 projection of an authoritative bridge-owned session. */
private static Map<String, Object> view(WorkerSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("terminalId", s.terminalId());
m.put("paneId", s.paneId());
m.put("profile", s.profile());
m.put("cwd", s.cwd());
m.put("ownerTerminal", s.ownerTerminal());
m.put("state", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
return m;
}
}
@@ -0,0 +1,179 @@
package dev.ltms.bridged.session;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.BufferedReader;
import java.io.IOException;
import java.io.InputStreamReader;
import java.io.UncheckedIOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.security.SecureRandom;
import java.util.List;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.stream.Collectors;
/**
* Production {@link Worktrees} implementation that shells {@code git} via {@link ProcessBuilder}.
* Non-zero exits become {@link WorktreeException}. Worktree directories live under a configurable
* root (default: a sibling {@code .bridged-worktrees} of the repo root) so they are never nested
* inside the primary working tree.
*/
public final class GitWorktrees implements Worktrees {
private static final Logger log = LoggerFactory.getLogger(GitWorktrees.class);
private final String configuredRoot;
private final SecureRandom random = new SecureRandom();
private final AtomicLong seq = new AtomicLong();
/** Default constructor: worktree root is derived per-repo as {@code <repoRoot>/../.bridged-worktrees}. */
public GitWorktrees() {
this(null);
}
/** @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of the repo root. */
public GitWorktrees(String configuredRoot) {
this.configuredRoot = configuredRoot;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
String base = (baseRef == null || baseRef.isBlank()) ? "HEAD" : baseRef;
String nonce = nonce();
Path root = resolveRoot(repoRoot);
Path path = root.resolve(nonce);
try {
Files.createDirectories(root);
} catch (IOException e) {
throw new WorktreeException("cannot create worktree root " + root + ": " + e.getMessage(), e);
}
String wt = path.toAbsolutePath().toString();
log.info("adding worktree branch={} path={} base={}", branch, wt, base);
exec("git", "-C", repoRoot, "worktree", "add", wt, "-b", branch, base);
return wt;
}
@Override
public void remove(String repoRoot, String worktreePath) {
Path p = Path.of(worktreePath);
if (!Files.exists(p)) {
log.debug("worktree {} already gone — nothing to remove", worktreePath);
return;
}
log.info("removing worktree {}", worktreePath);
exec("git", "-C", repoRoot, "worktree", "remove", "--force", worktreePath);
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
if (overlay == null || overlay.isEmpty()) {
return;
}
Path srcRoot = Path.of(repoRoot).toAbsolutePath().normalize();
Path dstRoot = Path.of(worktreePath).toAbsolutePath().normalize();
for (String rel : overlay) {
Path src = srcRoot.resolve(rel).normalize();
if (!Files.exists(src)) {
log.debug("parity overlay source missing — skipping {}", rel);
continue;
}
Path dst = dstRoot.resolve(rel).normalize();
try {
Files.createDirectories(dst.getParent());
Files.copy(src, dst, StandardCopyOption.REPLACE_EXISTING, StandardCopyOption.COPY_ATTRIBUTES);
log.debug("copied parity overlay {}", rel);
} catch (IOException e) {
throw new WorktreeException("cannot copy overlay " + rel + ": " + e.getMessage(), e);
}
if (isTracked(dstRoot, rel)) {
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", rel);
log.debug("marked overlay --skip-worktree {}", rel);
}
}
}
@Override
public String repoRoot(String cwd) {
String out = exec("git", "-C", cwd, "rev-parse", "--show-toplevel");
return Path.of(out.trim()).toAbsolutePath().normalize().toString();
}
/** Resolve the directory that will hold per-session worktree checkouts. */
private Path resolveRoot(String repoRoot) {
if (configuredRoot != null && !configuredRoot.isBlank()) {
return Path.of(configuredRoot).toAbsolutePath().normalize();
}
Path repo = Path.of(repoRoot).toAbsolutePath().normalize();
return repo.resolveSibling(".bridged-worktrees");
}
private String nonce() {
return String.format("%06x", random.nextInt(1 << 24)) + "-" + seq.incrementAndGet();
}
private boolean isTracked(Path worktreeRoot, String rel) {
return exitCode("git", "-C", worktreeRoot.toString(), "ls-files", "--error-unmatch", rel) == 0;
}
/**
* Run a command and return its stdout. Non-zero exit → {@link WorktreeException} with both
* stdout and stderr (merged by redirectErrorStream).
*/
private String exec(String... command) {
String out;
int code;
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
try (BufferedReader r = new BufferedReader(new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
out = r.lines().collect(Collectors.joining("\n"));
} catch (IOException e) {
p.destroyForcibly();
throw new UncheckedIOException(e);
}
try {
if (!p.waitFor(30, TimeUnit.SECONDS)) {
p.destroyForcibly();
throw new WorktreeException("command timed out: " + String.join(" ", command) + "\n" + out);
}
code = p.exitValue();
} catch (InterruptedException e) {
p.destroyForcibly();
Thread.currentThread().interrupt();
throw new WorktreeException("interrupted waiting for command: " + String.join(" ", command), e);
}
if (code != 0) {
throw new WorktreeException("exit " + code + " for: " + String.join(" ", command)
+ (out.isBlank() ? "" : "\n" + out));
}
return out;
}
private int exitCode(String... command) {
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
try {
if (!p.waitFor(30, TimeUnit.SECONDS)) {
p.destroyForcibly();
throw new WorktreeException("command timed out: " + String.join(" ", command));
}
return p.exitValue();
} catch (InterruptedException e) {
p.destroyForcibly();
Thread.currentThread().interrupt();
throw new WorktreeException("interrupted waiting for command: " + String.join(" ", command), e);
}
}
}
@@ -0,0 +1,411 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.LongSupplier;
/**
* Authoritative in-daemon registry of the worker sessions this {@code bridged} process spawned.
* Delegates spawn/teardown to a {@link PeerLauncher} (which performs subscription-guarded env
* setup and process/materialization) and adds lifecycle tracking, ownership, and deterministic
* teardown on top.
*
* <p>The state machine is intentionally one-shot / no-reuse: every acquired worker is fresh,
* and a finished or released worker is torn down, never pooled. {@link #recycle} is a convenience
* for {@code release + acquire} with a new distinct pane id.
*
* <p>The manager implements {@link TurnListener} so the injector's turn boundaries drive
* {@code READY → BUSY → DONE} (or {@code FAILED}). It exposes a {@link WorkerPresence} view via
* {@link #asPresence()}: any MCP contact from a worker marks it present and simultaneously
* transitions the session {@code SPAWNING → READY}.
*/
public final class SessionManager implements TurnListener {
private static final Logger log = LoggerFactory.getLogger(SessionManager.class);
private final PeerLauncher launcher;
private final Worktrees worktrees;
private final ConcurrentHashMap<String /*paneId*/, WorkerSession> registry = new ConcurrentHashMap<>();
private final WorkerPresence presence;
private final SecureRandom nonceRandom = new SecureRandom();
private final AtomicLong nonceSeq = new AtomicLong();
private final LongSupplier nowNanos;
private final int contextCap;
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
public SessionManager(PeerLauncher launcher) {
this(launcher, new GitWorktrees(), System::nanoTime, 0);
}
/** Backward-compatible constructor with an injectable worktree seam. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees) {
this(launcher, worktrees, System::nanoTime, 0);
}
/** Test constructor with an injectable clock. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos) {
this(launcher, worktrees, nowNanos, 0);
}
/** Production constructor with a configured context turn cap. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, int contextCap) {
this(launcher, worktrees, System::nanoTime, contextCap);
}
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos,
int contextCap) {
this.launcher = launcher;
this.worktrees = worktrees;
this.presence = new PresenceBridge(this);
this.nowNanos = nowNanos;
this.contextCap = contextCap;
}
/**
* The single {@link WorkerPresence} view of this manager: it records availability and forwards
* the signal to the {@code SPAWNING → READY} transition. Pass this to the {@code Injector} and
* {@code BridgeMcp} where they previously accepted a plain {@link WorkerPresence}. The same
* instance is returned every call — presence is shared state, so a fresh bridge per call would
* fragment the {@code present} set and lose signals across callers.
*/
public WorkerPresence asPresence() {
return presence;
}
/**
* Spawn a worker and register it as {@link WorkerSession.State#SPAWNING}. The caller's
* identity is recorded as {@code ownerTerminal} ({@code null} for daemon/anon callers).
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal) {
return acquire(profile, requestedCwd, callerCwd, ownerTerminal, null);
}
/**
* Spawn a worker, optionally inside a fresh git worktree. When {@code wt} is non-null the
* worktree is provisioned, parity-overlaid, and its path becomes the worker's cwd. On any
* failure before registration the worktree is removed so no dangling checkout is left.
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
if (wt == null) {
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd);
PeerHandle handle = launcher.spawn(req);
String resolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
String cwd = launcher.effectiveCwd(req);
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
cwd,
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
null,
null);
registry.put(handle.id(), session);
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
return session;
}
return acquireWithWorktree(profile, requestedCwd, callerCwd, ownerTerminal, wt);
}
/** Tear a worker down by pane id and remove it from the registry. Idempotent. */
public void release(String paneId) {
WorkerSession removed = registry.remove(paneId);
if (removed != null) {
log.debug("releasing session pane={} terminal={} state={}",
removed.paneId(), removed.terminalId(), removed.state());
}
launcher.stop(paneId);
if (removed != null && removed.worktree() != null) {
worktrees.remove(worktrees.repoRoot(removed.cwd()), removed.worktree());
}
}
private WorkerSession acquireWithWorktree(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
String resolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
String repoRoot = worktrees.repoRoot(firstNonBlank(requestedCwd, callerCwd));
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
String path = null;
PeerHandle handle;
try {
path = worktrees.add(repoRoot, branch, wt.baseRef());
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(resolvedProfile));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd));
} catch (RuntimeException e) {
if (path != null) {
try {
worktrees.remove(repoRoot, path);
} catch (RuntimeException cleanup) {
log.warn("failed to clean up worktree {} after spawn error: {}", path, cleanup.getMessage());
}
}
throw e;
}
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
resolveCwd(path, profile, callerCwd),
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
path,
branch);
registry.put(handle.id(), session);
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
return session;
}
private String slug(String raw) {
return raw == null ? "ticket" : raw.toLowerCase().replaceAll("[^a-z0-9]+", "-").replaceAll("^-+|-+$", "");
}
private String nonce() {
return String.format("%06x", nonceRandom.nextInt(1 << 24)) + "-" + nonceSeq.incrementAndGet();
}
private static String firstNonBlank(String... values) {
for (String v : values) {
if (v != null && !v.isBlank()) return v;
}
return null;
}
/**
* Release the old session and acquire a fresh one with the same profile and working directory.
* The new session is guaranteed to have a pane id distinct from the old one (no-reuse invariant).
*/
public WorkerSession recycle(String paneId) {
WorkerSession old = registry.get(paneId);
if (old == null) {
throw new IllegalArgumentException("no session for paneId " + paneId);
}
release(paneId);
return acquire(old.profile(), old.cwd(), old.cwd(), old.ownerTerminal());
}
/** The session for {@code paneId}, if it is still registered and not released. */
public Optional<WorkerSession> get(String paneId) {
return Optional.ofNullable(registry.get(paneId));
}
/** Bridge-owned roster: all registered sessions (acquired minus released). */
public List<WorkerSession> roster() {
return List.copyOf(registry.values());
}
/**
* CB-304 merged roster+live view. The registry is authoritative for worktree, branch,
* profile, owner, and state; the optional live agent supplies the herdr-reported status.
*/
public static Map<String, Object> rosterView(WorkerSession session, Agent live) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", session.terminalId());
m.put("paneId", session.paneId());
m.put("profile", session.profile());
m.put("state", session.state().name().toLowerCase());
if (session.worktree() != null) {
m.put("worktree", session.worktree());
}
if (session.branch() != null) {
m.put("branch", session.branch());
}
if (session.ownerTerminal() != null) {
m.put("owner", session.ownerTerminal());
}
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
return m;
}
/** Lifecycle hook: worker became available on the bridge MCP. */
void onReady(String terminalId) {
transitionByTerminal(terminalId, WorkerSession.State.SPAWNING, WorkerSession.State.READY);
}
/**
* Lifecycle hook: a message was delivered into the worker — it is now busy on a turn.
* The turn count is bumped and the activity timestamp is refreshed. A {@code DONE} session
* can be re-delivered for multi-turn reuse until it is released.
*/
@Override
public void onDelivered(String target) {
WorkerSession current = findByTerminal(target);
if (current == null) return;
if (current.state() != WorkerSession.State.READY && current.state() != WorkerSession.State.DONE) {
return;
}
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.BUSY).bumpTurn(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} {} -> BUSY turn={}",
target, current.paneId(), current.state(), updated.turnCount());
}
}
/** Lifecycle hook: the worker's delegated turn completed successfully. */
@Override
public void onTurnComplete(String target) {
WorkerSession current = findByTerminal(target);
if (current == null || current.state() != WorkerSession.State.BUSY) return;
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.DONE).withActivity(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} BUSY -> DONE turn={}",
target, current.paneId(), updated.turnCount());
}
if (contextCap > 0 && updated.turnCount() >= contextCap) {
release(current.paneId());
}
}
/** Lifecycle hook: the worker's delegated turn failed. */
@Override
public void onTurnFailed(String target) {
onFailed(target);
}
/** Lifecycle hook: the worker vanished or was dropped mid-life. */
void onFailed(String target) {
WorkerSession current = findByTerminal(target);
if (current == null) return;
if (current.state() == WorkerSession.State.RELEASED) return;
if (replace(current, current.withState(WorkerSession.State.FAILED))) {
log.debug("session marked failed terminal={} pane={}", target, current.paneId());
}
}
/**
* Best-effort reap of sessions that have been idle longer than {@code idleTtlNanos}. Only
* {@code READY} and {@code DONE} sessions are eligible — never a {@code SPAWNING} or
* {@code BUSY} worker. Returns the number of sessions released.
*/
int reapIdle(long idleTtlNanos) {
long now = nowNanos.getAsLong();
int reaped = 0;
for (WorkerSession s : roster()) {
if (s.state() != WorkerSession.State.READY && s.state() != WorkerSession.State.DONE) {
continue;
}
if (now - s.lastActivityAtNanos() > idleTtlNanos) {
release(s.paneId());
reaped++;
}
}
return reaped;
}
/**
* Gracefully drain all registered sessions. For each session that is {@code BUSY}, poll up to
* {@code timeoutNanos} for it to leave {@code BUSY}, then release it regardless. Non-busy
* sessions are released immediately. A failure releasing one session is logged and does not
* abort the rest.
*/
void drainAll(long timeoutNanos) {
long deadline = System.nanoTime() + timeoutNanos;
for (WorkerSession s : roster()) {
try {
if (s.state() == WorkerSession.State.BUSY) {
while (System.nanoTime() < deadline) {
WorkerSession current = registry.get(s.paneId());
if (current == null || current.state() != WorkerSession.State.BUSY) {
break;
}
try {
long remaining = deadline - System.nanoTime();
Thread.sleep(Math.min(TimeUnit.NANOSECONDS.toMillis(remaining), 50));
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
break;
}
}
}
release(s.paneId());
} catch (RuntimeException e) {
log.warn("drain failed for pane={}; continuing with remaining sessions", s.paneId(), e);
}
}
}
/**
* Close this manager by draining all sessions. The timeout comes from configuration when set,
* otherwise a sensible default.
*/
public void close(Integer drainTimeoutSeconds) {
int seconds = (drainTimeoutSeconds != null && drainTimeoutSeconds > 0) ? drainTimeoutSeconds : 5;
drainAll(TimeUnit.SECONDS.toNanos(seconds));
}
/** Number of sessions currently registered. */
public int size() {
return registry.size();
}
private WorkerSession findByTerminal(String terminalId) {
for (WorkerSession s : registry.values()) {
if (terminalId.equals(s.terminalId())) return s;
}
return null;
}
private void transitionByTerminal(String terminalId, WorkerSession.State from,
WorkerSession.State to) {
WorkerSession current = findByTerminal(terminalId);
if (current == null || current.state() != from) return;
long now = nowNanos.getAsLong();
if (replace(current, current.withState(to).withActivity(now))) {
log.debug("session transitioned terminal={} pane={} {} -> {}",
terminalId, current.paneId(), from, to);
}
}
private boolean replace(WorkerSession expected, WorkerSession updated) {
return registry.replace(expected.paneId(), expected, updated);
}
private String resolveCwd(String requestedCwd, String profileName, String callerCwd) {
return launcher.effectiveCwd(new SpawnRequest(profileName, requestedCwd, callerCwd));
}
/** WorkerPresence bridge that also drives the manager's READY transition. */
private static final class PresenceBridge extends WorkerPresence {
private final SessionManager sessions;
PresenceBridge(SessionManager sessions) {
this.sessions = sessions;
}
@Override
public void markPresent(String terminal) {
super.markPresent(terminal);
sessions.onReady(terminal);
}
}
}
@@ -0,0 +1,70 @@
package dev.ltms.bridged.session;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.TimeUnit;
/**
* Periodic virtual-thread reaper that tears down {@code READY}/{@code DONE} sessions which have
* exceeded their idle TTL. Modeled on {@link dev.ltms.bridged.inject.StatusPoller}: a single
* virtual-thread loop, idempotent start/stop, and no {@code ScheduledExecutorService}.
*/
public final class SessionReaper {
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
private final SessionManager sessions;
private final long idleTtlNanos;
private final long intervalMillis;
private volatile boolean running;
private Thread thread;
/** Construct a reaper with the default 5-second polling interval. */
public SessionReaper(SessionManager sessions, long idleTtlSeconds) {
this(sessions, idleTtlSeconds, DEFAULT_INTERVAL_MILLIS);
}
/** Construct a reaper with an explicit polling interval (useful for tests). */
public SessionReaper(SessionManager sessions, long idleTtlSeconds, long intervalMillis) {
this.sessions = sessions;
this.idleTtlNanos = TimeUnit.SECONDS.toNanos(idleTtlSeconds);
this.intervalMillis = intervalMillis;
}
/** Start the reaper loop on a virtual thread. Idempotent. */
public synchronized void start() {
if (running) return;
running = true;
thread = Thread.ofVirtual().name("session-reaper").start(this::loop);
log.info("session reaper started (idle ttl {}s, interval {}ms)",
TimeUnit.NANOSECONDS.toSeconds(idleTtlNanos), intervalMillis);
}
private void loop() {
while (running) {
try {
sessions.reapIdle(idleTtlNanos);
} catch (RuntimeException e) {
log.warn("session reaper iteration failed; continuing", e);
}
sleep();
}
}
private void sleep() {
try {
Thread.sleep(intervalMillis);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
running = false;
}
}
/** Stop the reaper loop. Idempotent. */
public synchronized void stop() {
running = false;
if (thread != null) thread.interrupt();
}
}
@@ -0,0 +1,58 @@
package dev.ltms.bridged.session;
/**
* A bridge-owned worker session — the authoritative in-daemon record of a worker this
* process spawned. Immutable; state transitions are performed by replacing the record in
* {@link SessionManager}'s registry.
*
* @param paneId herdr pane handle — the registry key and the argument to teardown
* @param terminalId herdr terminal handle — the {@code target} for send/read/status
* @param profile the worker profile name that spawned this session
* @param cwd the resolved working directory the worker started in
* @param ownerTerminal the caller that requested this worker ({@code null} = daemon/anon)
* @param spawnedAtNanos {@link System#nanoTime()} when the session was registered
* @param lastActivityAtNanos {@link System#nanoTime()} of the most recent lifecycle event
* @param turnCount number of delegated turns that have been delivered to this session
* @param state current lifecycle state in the one-shot FSM
*/
public record WorkerSession(
String paneId,
String terminalId,
String profile,
String cwd,
String ownerTerminal,
long spawnedAtNanos,
long lastActivityAtNanos,
int turnCount,
State state,
String worktree,
String branch) {
/** One-shot worker lifecycle states. */
public enum State {
SPAWNING,
READY,
BUSY,
DONE,
FAILED,
RELEASED
}
/** Return a copy of this session in {@code state}. */
public WorkerSession withState(State state) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch);
}
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
public WorkerSession withActivity(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount, state, worktree, branch);
}
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
public WorkerSession bumpTurn(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount + 1, state, worktree, branch);
}
}
@@ -0,0 +1,12 @@
package dev.ltms.bridged.session;
/** Non-zero exit or I/O failure from a git worktree operation. */
public final class WorktreeException extends RuntimeException {
public WorktreeException(String message) {
super(message);
}
public WorktreeException(String message, Throwable cause) {
super(message, cause);
}
}
@@ -0,0 +1,6 @@
package dev.ltms.bridged.session;
/** Ask {@link SessionManager#acquire} to provision an isolated worktree. null ⇒ run in the shared primary tree. */
public record WorktreeRequest(String ticketSlug, String baseRef) {
// ticketSlug seeds the branch name; baseRef null/blank ⇒ current HEAD of the repo.
}
@@ -0,0 +1,18 @@
package dev.ltms.bridged.session;
import java.util.List;
/** Seam between {@link SessionManager} and git worktree operations. Tests use a recording fake. */
public interface Worktrees {
/** git -C <repoRoot> worktree add <path> -b <branch> <baseRef|HEAD>. Returns the worktree path. */
String add(String repoRoot, String branch, String baseRef);
/** git -C <repoRoot> worktree remove --force <path>. Idempotent (already-gone tolerated). */
void remove(String repoRoot, String worktreePath);
/** Copy each existing overlay path repoRoot→worktree; mark tracked ones --skip-worktree. */
void overlayParity(String repoRoot, String worktreePath, List<String> overlay);
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
String repoRoot(String cwd);
}
@@ -0,0 +1,488 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.EnumSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* Spawns and lists worker sessions — the safe path from a delegation request to a
* running off-subscription Claude.
*
* <p>The spawn sequence encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em>
* touching herdr, and only then {@code agent.start}. A worker's base_url lives in the
* env map handed to herdr and nowhere else; {@code bridged}'s own environment is never
* mutated.
*
* <p>Placement: in the default {@code tab} policy a worker lands in its own tab inside a
* dedicated worker space (found-or-created once, then shared), so workers never split or
* clutter the user's real work spaces. Teardown removes the worker's pane <em>and</em> its
* now-empty tab, tolerating an already-gone worker so a repeated DELETE is harmless.
*/
public final class ClaudeCodeLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(ClaudeCodeLauncher.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
/**
* A bridge-spawned worker label {@code claude-<profile>-<nonce>-<seq>} (see
* {@link #startUniquelyNamed}); group 1 captures the 6-hex per-process {@code nonce}. The
* profile segment may itself contain {@code -}, so the nonce/seq are anchored at the tail.
* Names not matching this shape are not workers we started and are never reaped (CB-117).
*/
private static final Pattern WORKER_NAME = Pattern.compile("claude-.*-([0-9a-f]{6})-\\d+");
private final AgentControl agents;
private final WorkspaceControl spaces;
private final SubscriptionGuard guard;
private final Map<String, BridgedConfig.Worker> profiles; // profile name → spawn settings
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
private final Function<String, String> env; // host env lookup (injectable for tests)
private final AtomicLong nameSeq = new AtomicLong(); // per-worker counter (also the tab #)
/**
* Standing instruction appended to the worker's system prompt so it returns its result via
* {@code bridge_reply}. Injected as a launch flag, so nothing is written to the worker's
* profile — it is guidance, and a worker that never replies is caught by the send's timeout.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet. Every message you "
+ "receive arrives through the bridge, and the ONLY channel back to the sender is the "
+ "bridge_reply MCP tool. Text you write in your terminal is NOT sent anywhere — the "
+ "sender cannot see your screen, so an in-terminal answer is silently discarded. "
+ "Therefore you MUST end EVERY turn by calling bridge_reply with `content` set to your "
+ "complete response. This holds for every message without exception — tasks, questions, "
+ "clarifications, acknowledgements, and ordinary back-and-forth conversation. Call "
+ "bridge_reply exactly once, as the final action of your turn, with your full answer in "
+ "`content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
// Per-process token mixed into each worker name so a fresh process (nameSeq back at 0)
// cannot collide with same-profile workers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this.agents = agents;
this.spaces = spaces;
this.guard = guard;
this.profiles = Map.copyOf(profiles);
this.defaultProfile = defaultProfile;
this.env = env;
}
/** The configured worker profile names (what {@code spawn(profile)} accepts). */
@Override
public Set<String> profiles() {
return profiles.keySet();
}
/** The parity-overlay file list for {@code profileName} (default list when unset). */
@Override
public List<String> parityOverlay(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
return List.of();
}
BridgedConfig.Worker cfg = profiles.get(name);
return cfg == null ? List.of() : cfg.parityOverlay();
}
/** The profile a no-argument {@link #spawn()} uses, or {@code null} if none is configured. */
@Override
public String defaultProfile() {
return defaultProfile;
}
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawn(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawn(profileName, null, null);
}
/**
* Spawn a worker. {@code profileName} null/blank → the default profile. The worker's working
* directory (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd} (a
* spawn argument), else the profile's configured {@code cwd}, else {@code callerCwd} (the
* primary's cwd, when the spawn came from the primary over MCP), else the daemon's cwd — never
* assumed to be {@code $HOME}. Guard runs before any herdr call.
*/
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
throw new IllegalArgumentException("no default worker profile is configured — "
+ "pass a profile; configured: " + profiles.keySet());
}
BridgedConfig.Worker cfg = profiles.get(name);
if (cfg == null) {
throw new IllegalArgumentException("unknown worker profile '" + name
+ "' — configured: " + profiles.keySet());
}
String baseUrl = cfg.baseUrl();
guard.assertWorker(baseUrl); // hard stop before we spawn anything
Map<String, String> workerEnv = new LinkedHashMap<>();
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
String token = env.apply(cfg.tokenEnv());
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", token);
// CB-302: the worker checkpoint (commit → push → open its own PR). Push is free over SSH;
// the only incremental grant is PR-create, a repo-scoped forge token injected here — opt-in
// per profile via gitTokenEnv, and never mutating bridged's own env. The paired forge host
// rides along only when a token is actually granted, so non-implementer profiles get neither.
if (cfg.hasGitToken()) {
String gitToken = resolveEnv(cfg.gitTokenEnv());
if (gitToken != null) {
workerEnv.put("GITEA_TOKEN", gitToken);
putIfPresent(workerEnv, "GITEA_HOST", resolveEnv(cfg.gitHostEnv()));
}
}
// Mount the bridge MCP + reply charter as launch FLAGS (non-invasive: nothing written to
// the worker's profile/config dir). Identity is connection-based, so the mount is shared.
List<String> argv = argvWithBridge(cfg);
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
return cfg.tabPlacement()
? spawnInTab(cfg, workerEnv, argv, cwd)
: spawnAsPane(cfg, workerEnv, argv, cwd);
}
/**
* CB-112 cwd resolution: spawn arg → profile config → the primary's cwd → the daemon's cwd.
* Never returns {@code null}/blank: {@code "."} (the daemon's own working directory) is the
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
*/
private static String resolveCwd(String requestedCwd, BridgedConfig.Worker cfg, String callerCwd) {
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
}
/**
* CB-301: the effective working directory a spawn for {@code profileName} would use, without
* actually spawning. Used by {@link dev.ltms.bridged.session.SessionManager} to record the
* resolved cwd in the session registry.
*/
public String effectiveCwd(String profileName, String requestedCwd, String callerCwd) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
throw new IllegalArgumentException("no default worker profile is configured — "
+ "pass a profile; configured: " + profiles.keySet());
}
BridgedConfig.Worker cfg = profiles.get(name);
if (cfg == null) {
throw new IllegalArgumentException("unknown worker profile '" + name
+ "' — configured: " + profiles.keySet());
}
return resolveCwd(requestedCwd, cfg, callerCwd);
}
private static String firstNonBlank(String... values) {
for (String v : values) {
if (v != null && !v.isBlank()) return v;
}
return null;
}
/**
* The launch argv, plus — when {@code worker.mcpUrl} is set — inline {@code --mcp-config} for
* the bridge server and {@code --append-system-prompt} for the {@link #REPLY_CHARTER}. Neither
* touches the profile's config; both are pure command-line flags.
*/
private List<String> argvWithBridge(BridgedConfig.Worker cfg) {
if (!cfg.hasMcp()) {
return cfg.argv();
}
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
List<String> argv = new ArrayList<>(cfg.argv());
argv.add("--mcp-config");
argv.add(mcpJson);
argv.add("--append-system-prompt");
argv.add(REPLY_CHARTER);
return argv;
}
/** Dedicated worker space → own tab → start the worker (rooted at {@code cwd}) → drop the shell. */
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId());
log.info("spawning worker profile={} base_url={} space={} tab={} cwd={}",
cfg.profile(), cfg.baseUrl(), space.workspaceId(), tab.tab().tabId(), cwd);
Started started;
try {
started = startUniquelyNamed(cfg, workerEnv, argv, tab.tab().tabId(), cwd);
} catch (RuntimeException e) {
// The worker never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("failed to close orphaned tab {} after spawn error: {}",
tab.tab().tabId(), cleanup.getMessage());
}
throw e;
}
// The worker is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell
// so the tab holds only the worker; label the tab). They must not fail the spawn or
// orphan the running worker — on error we log and still return it so the caller gets
// its paneId and can tear it down.
if (tab.rootPaneId() != null) {
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
} else {
log.warn("tab {} had no seed pane in the create response; worker tab may hold an extra pane",
tab.tab().tabId());
}
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
log.info("worker started pane={} tab={} terminal={}",
started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
return started.agent();
}
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
private void tidy(String what, Runnable step) {
try {
step.run();
} catch (RuntimeException e) {
log.warn("post-start step failed ({}) — worker is running regardless: {}", what, e.getMessage());
}
}
/** Legacy placement: herdr splits the currently-focused tab; the worker still starts in {@code cwd}. */
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
log.info("spawning worker (pane placement) profile={} base_url={} cwd={} argv={}",
cfg.profile(), cfg.baseUrl(), cwd, argv);
Agent worker = startUniquelyNamed(cfg, workerEnv, argv, null, cwd).agent();
log.info("worker started pane={} terminal={}", worker.paneId(), worker.terminalId());
return worker;
}
/** A started worker together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
/**
* Start the worker under a unique herdr agent name. herdr requires each running
* agent's {@code name} to be distinct (a 2nd {@code name:"claude"} fails
* {@code agent_name_taken}) — the exact case that makes multiple workers useful. The name
* is {@code claude-<profile>-<nonce>-<seq>}: {@code seq} distinguishes workers within this
* process, and the per-process {@code nonce} keeps a fresh process (whose {@code seq}
* restarts at 0) from colliding with same-profile workers that outlived a restart. The
* retry is a belt-and-braces backstop for the astronomically unlikely nonce+seq clash;
* the name is a label only — herdr detects kind and status from terminal output, not it.
*/
private Started startUniquelyNamed(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String tabId, String cwd) {
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = "claude-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(agents.start(name, argv, workerEnv, tabId, cwd), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("worker name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
/** All herdr-tracked agents — discovery for "what workers exist". */
@Override
public List<Agent> list() {
return agents.list();
}
/**
* Reap worker panes left behind by an earlier daemon process (CB-117). herdr keeps a worker's
* pane alive across a daemon restart <em>by design</em>, and that pane's id is held only by its
* spawner — so a worker whose owning process exited before issuing the matching teardown leaks
* with nothing tracking it (there is no registry; {@link #list()} only asks herdr). On boot we
* scan herdr for agents whose name matches our {@code claude-<profile>-<nonce>-<seq>} scheme with
* a nonce <em>other</em> than this process's {@link #nameNonce}, and tear each one down (its pane
* and, via {@link #stop}, its now-empty dedicated tab). A current-nonce worker is ours and live,
* so it is left running; a user's own {@code claude} session carries no such name and is never
* touched. Best-effort: a failed listing, or a failure to stop any one worker, is logged and
* never aborts startup.
*
* @return the number of orphaned workers reaped
*/
@Override
public int reapOrphanWorkers() {
List<Agent> all;
try {
all = agents.list();
} catch (RuntimeException e) {
log.warn("orphan-worker reap skipped — agent.list failed: {}", e.getMessage());
return 0;
}
int reaped = 0;
for (Agent a : all) {
if (!isForeignWorker(a.name(), nameNonce)) continue;
try {
stop(a.paneId());
reaped++;
log.info("reaped orphan worker {} (pane={} tab={}) left by a prior daemon",
a.name(), a.paneId(), a.tabId());
} catch (RuntimeException e) {
log.warn("could not reap orphan worker {} (pane={}): {}",
a.name(), a.paneId(), e.getMessage());
}
}
if (reaped > 0) {
log.info("orphan-worker reap complete — {} stale worker(s) removed at startup", reaped);
}
return reaped;
}
/**
* Whether {@code name} is a bridge worker started by a <em>different</em> process than
* {@code currentNonce} — the reap predicate (CB-117). True only for our naming scheme with a
* foreign nonce: a non-worker name (no match, e.g. a user session) or our own live nonce is
* excluded. Pure and package-private so the decision is unit-testable without herdr.
*/
static boolean isForeignWorker(String name, String currentNonce) {
String nonce = workerNonce(name);
return nonce != null && !nonce.equals(currentNonce);
}
/** The 6-hex nonce embedded in a bridge worker name, or {@code null} if {@code name} isn't one. */
static String workerNonce(String name) {
if (name == null) return null;
Matcher m = WORKER_NAME.matcher(name);
return m.matches() ? m.group(1) : null;
}
/** This process's worker-name nonce (a label component only; exposed for reaper tests). */
String nameNonce() {
return nameNonce;
}
/**
* Tear a worker down by pane id: close the pane, and close its tab <em>only</em> when the
* worker is that tab's sole occupant. The single-pane check is what makes this safe
* regardless of how the worker was placed (or a placement-config change across a restart):
* a pane-placement worker sitting in one of the user's shared tabs has siblings, so its
* tab is never closed — we only ever remove a tab we created to hold one worker.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed worker) is treated as success; any other failure propagates so
* a genuinely failed teardown is not reported as done.
*/
@Override
public void stop(String paneId) {
// Teardown knows only the paneId, not which profile spawned it. Attempt tab cleanup when any
// profile uses tab placement (so the bridge may have created a dedicated worker tab); the
// single-occupant check below is what actually protects the user's shared tabs.
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) throw e;
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
spaces.closeTab(loc.tabId());
} else if (loc != null) {
log.debug("not closing tab {} — it holds {} panes (not a dedicated worker tab)",
loc.tabId(), loc.tabPaneCount());
}
}
/** Whether any configured profile places workers in their own tab (so tabs may need cleanup). */
private boolean usesTabPlacement() {
return profiles.values().stream().anyMatch(BridgedConfig.Worker::tabPlacement);
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
// --- PeerLauncher SPI -------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profiles.values().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
/**
* {@inheritDoc}
*
* <p>Delegates to the three-arg {@link #spawn(String, String, String)} and wraps the
* resulting herdr {@link Agent} in a {@link WorkerHandle} whose {@link PeerHandle#id()}
* equals the agent's paneId.
*/
@Override
public PeerHandle spawn(SpawnRequest req) {
Agent agent = spawn(req.profileName(), req.requestedCwd(), req.callerCwd());
return new WorkerHandle(agent.paneId(), agent.terminalId());
}
/** A concrete {@link PeerHandle} wrapping herdr agent coordinates. */
private record WorkerHandle(String id, String terminalId) implements PeerHandle {
}
@Override
public String effectiveCwd(SpawnRequest req) {
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
}
private static void putIfPresent(Map<String, String> m, String k, String v) {
if (v != null && !v.isBlank()) {
m.put(k, v);
}
}
/** Host env lookup that tolerates an unconfigured (null/blank) var name — returns null then. */
private String resolveEnv(String name) {
return (name == null || name.isBlank()) ? null : env.apply(name);
}
}
@@ -1,206 +0,0 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
/**
* Spawns and lists worker sessions — the safe path from a delegation request to a
* running off-subscription Claude.
*
* <p>The spawn sequence encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em>
* touching herdr, and only then {@code agent.start}. A worker's base_url lives in the
* env map handed to herdr and nowhere else; {@code bridged}'s own environment is never
* mutated.
*
* <p>Placement: in the default {@code tab} policy a worker lands in its own tab inside a
* dedicated worker space (found-or-created once, then shared), so workers never split or
* clutter the user's real work spaces. Teardown removes the worker's pane <em>and</em> its
* now-empty tab, tolerating an already-gone worker so a repeated DELETE is harmless.
*/
public final class WorkerService {
private static final Logger log = LoggerFactory.getLogger(WorkerService.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
private final AgentControl agents;
private final WorkspaceControl spaces;
private final SubscriptionGuard guard;
private final BridgedConfig.Worker cfg;
private final Function<String, String> env; // host env lookup (injectable for tests)
private final AtomicLong nameSeq = new AtomicLong(); // per-worker counter (also the tab #)
// Per-process token mixed into each worker name so a fresh process (nameSeq back at 0)
// cannot collide with same-profile workers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
public WorkerService(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
BridgedConfig.Worker cfg, Function<String, String> env) {
this.agents = agents;
this.spaces = spaces;
this.guard = guard;
this.cfg = cfg;
this.env = env;
}
/** Spawn a worker for the configured profile. Guard runs before any herdr call. */
public Agent spawn() {
String baseUrl = cfg.baseUrl();
guard.assertWorker(baseUrl); // hard stop before we spawn anything
Map<String, String> workerEnv = new LinkedHashMap<>();
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
String token = env.apply(cfg.tokenEnv());
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", token);
return cfg.tabPlacement() ? spawnInTab(workerEnv) : spawnAsPane(workerEnv);
}
/** Dedicated worker space → own tab → drop the placeholder shell so only the worker remains. */
private Agent spawnInTab(Map<String, String> workerEnv) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId());
log.info("spawning worker profile={} base_url={} space={} tab={}",
cfg.profile(), cfg.baseUrl(), space.workspaceId(), tab.tab().tabId());
Started started;
try {
started = startUniquelyNamed(workerEnv, tab.tab().tabId());
} catch (RuntimeException e) {
// The worker never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("failed to close orphaned tab {} after spawn error: {}",
tab.tab().tabId(), cleanup.getMessage());
}
throw e;
}
// The worker is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell
// so the tab holds only the worker; label the tab). They must not fail the spawn or
// orphan the running worker — on error we log and still return it so the caller gets
// its paneId and can tear it down.
if (tab.rootPaneId() != null) {
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
} else {
log.warn("tab {} had no seed pane in the create response; worker tab may hold an extra pane",
tab.tab().tabId());
}
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
log.info("worker started pane={} tab={} terminal={}",
started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
return started.agent();
}
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
private void tidy(String what, Runnable step) {
try {
step.run();
} catch (RuntimeException e) {
log.warn("post-start step failed ({}) — worker is running regardless: {}", what, e.getMessage());
}
}
/** Legacy placement: herdr splits the currently-focused tab. */
private Agent spawnAsPane(Map<String, String> workerEnv) {
log.info("spawning worker (pane placement) profile={} base_url={} argv={}",
cfg.profile(), cfg.baseUrl(), cfg.argv());
Agent worker = startUniquelyNamed(workerEnv, null).agent();
log.info("worker started pane={} terminal={}", worker.paneId(), worker.terminalId());
return worker;
}
/** A started worker together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
/**
* Start the worker under a unique herdr agent name. herdr requires each running
* agent's {@code name} to be distinct (a 2nd {@code name:"claude"} fails
* {@code agent_name_taken}) — the exact case that makes multiple workers useful. The name
* is {@code claude-<profile>-<nonce>-<seq>}: {@code seq} distinguishes workers within this
* process, and the per-process {@code nonce} keeps a fresh process (whose {@code seq}
* restarts at 0) from colliding with same-profile workers that outlived a restart. The
* retry is a belt-and-braces backstop for the astronomically unlikely nonce+seq clash;
* the name is a label only — herdr detects kind and status from terminal output, not it.
*/
private Started startUniquelyNamed(Map<String, String> workerEnv, String tabId) {
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = "claude-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(agents.start(name, cfg.argv(), workerEnv, tabId), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("worker name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
/** All herdr-tracked agents — discovery for "what workers exist". */
public List<Agent> list() {
return agents.list();
}
/**
* Tear a worker down by pane id: close the pane, and close its tab <em>only</em> when the
* worker is that tab's sole occupant. The single-pane check is what makes this safe
* regardless of how the worker was placed (or a placement-config change across a restart):
* a pane-placement worker sitting in one of the user's shared tabs has siblings, so its
* tab is never closed — we only ever remove a tab we created to hold one worker.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed worker) is treated as success; any other failure propagates so
* a genuinely failed teardown is not reported as done.
*/
public void stop(String paneId) {
WorkspaceControl.PaneLocation loc = cfg.tabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) throw e;
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
spaces.closeTab(loc.tabId());
} else if (loc != null) {
log.debug("not closing tab {} — it holds {} panes (not a dedicated worker tab)",
loc.tabId(), loc.tabPaneCount());
}
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
private static void putIfPresent(Map<String, String> m, String k, String v) {
if (v != null && !v.isBlank()) {
m.put(k, v);
}
}
}
@@ -5,6 +5,7 @@ import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
@@ -46,6 +47,43 @@ class BridgedConfigTest {
assertTrue(cfg.guard().offSubscriptionHosts().isEmpty());
}
@Test
void singleWorkerBecomesAOneEntryProfileMapWithItselfAsDefault(@TempDir Path dir) throws Exception {
Path f = dir.resolve("single.yaml");
Files.writeString(f, """
worker:
profile: ltms-local
baseUrl: http://gx00.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(Set.of("ltms-local"), cfg.workerProfiles().keySet(), "legacy worker → one profile");
assertEquals("ltms-local", cfg.defaultProfile());
}
@Test
void loadsMultipleWorkerProfilesWithADefault(@TempDir Path dir) throws Exception {
Path f = dir.resolve("multi.yaml");
Files.writeString(f, """
workers:
gx10:
baseUrl: http://gx10.gw:8000
argv: ["ccs", "gx10"]
ollama:
baseUrl: http://ollama.ltms.dev
argv: ["ccs", "ollama"]
defaultWorker: gx10
guard:
offSubscriptionHosts: [gx10.gw, ollama.ltms.dev]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(Set.of("gx10", "ollama"), cfg.workerProfiles().keySet());
assertEquals("gx10", cfg.defaultProfile());
assertEquals("ollama", cfg.workerProfiles().get("ollama").profile(), "profile defaults to its map key");
assertEquals("http://gx10.gw:8000", cfg.workerProfiles().get("gx10").baseUrl());
}
@Test
void ignoresUnknownKeys(@TempDir Path dir) throws Exception {
Path f = dir.resolve("future.yaml");
@@ -0,0 +1,40 @@
package dev.ltms.bridged.herdr;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
/** Unit-level behaviour of {@link AgentControl} over a fake herdr. */
class AgentControlTest {
/** The {@code text} of every agent.send, in call order. */
@SuppressWarnings("unchecked")
private static List<String> sendTexts(FakeHerdr herdr) {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.toList();
}
@Test
void sendDeliversThePayloadThenAStandaloneSubmitKey() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_x", "do the thing");
// The Enter must be its own event — appended to the paste it would be swallowed as text.
assertEquals(List.of("do the thing", "\r"), sendTexts(herdr),
"payload paste first, then a separate carriage-return keystroke to submit it");
}
@Test
void sendPreservesEmbeddedNewlinesAndSubmitsOnlyOnce() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_x", "line1\nline2");
assertEquals(List.of("line1\nline2", "\r"), sendTexts(herdr),
"multiline content is delivered verbatim; a single trailing Enter submits it");
}
}
@@ -0,0 +1,43 @@
package dev.ltms.bridged.herdr;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Wire mapping and injectability of {@link AgentStatus}, including the CB-115 {@code done} state. */
class AgentStatusTest {
@Test
void mapsTheKnownWireStrings() {
assertEquals(AgentStatus.IDLE, AgentStatus.fromWire("idle"));
assertEquals(AgentStatus.WORKING, AgentStatus.fromWire("working"));
assertEquals(AgentStatus.BLOCKED, AgentStatus.fromWire("blocked"));
assertEquals(AgentStatus.DONE, AgentStatus.fromWire("done"));
}
@Test
void mapsDoneCaseInsensitively() {
assertEquals(AgentStatus.DONE, AgentStatus.fromWire("DONE"));
assertEquals(AgentStatus.DONE, AgentStatus.fromWire("Done"));
}
@Test
void unknownAndNullFallToUnknown() {
assertEquals(AgentStatus.UNKNOWN, AgentStatus.fromWire("unknown"));
assertEquals(AgentStatus.UNKNOWN, AgentStatus.fromWire("something-else"));
assertEquals(AgentStatus.UNKNOWN, AgentStatus.fromWire(null));
}
@Test
void doneIsInjectableLikeIdle() {
// The whole point of CB-115: a finished worker herdr reports as `done` must be deliverable,
// not treated as UNKNOWN (which wedged delivery and mis-fired the stall failure).
assertTrue(AgentStatus.DONE.injectable());
assertTrue(AgentStatus.IDLE.injectable());
assertTrue(AgentStatus.BLOCKED.injectable());
assertFalse(AgentStatus.WORKING.injectable());
assertFalse(AgentStatus.UNKNOWN.injectable());
}
}
@@ -16,15 +16,20 @@ public final class FakeHerdr implements HerdrClient {
public record Call(String method, Object params) {
}
/** The foreground PID of the one agent pane (term_a) in the canned {@code pane.process_info}. */
public static final long WORKER_PID = 4242;
private final ObjectMapper mapper = new ObjectMapper();
public final List<Call> calls = new ArrayList<>();
private boolean healthy = true;
private final List<String> extraWorkspaces = new ArrayList<>();
private final List<String> extraAgents = new ArrayList<>();
private int agentNameTakenFor = 0;
private int workerTabPaneCount = 1;
private String paneCloseErrorCode = null;
private String agentSendErrorCode = null;
private volatile String agentStatus = "idle"; // what agent.get reports
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String readText = "worker transcript tail"; // canned agent.read output
public FakeHerdr healthy(boolean h) {
this.healthy = h;
@@ -55,12 +60,32 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/** The text {@code agent.read} returns (the CB-106 completion scrape). */
public FakeHerdr readText(String text) {
this.readText = text;
return this;
}
/** Make {@code agent.send} fail with this herdr error code. */
public FakeHerdr agentSendFailsWith(String code) {
this.agentSendErrorCode = code;
return this;
}
/**
* Seed a named agent into {@code agent.list} (e.g. an orphaned worker for CB-117 reaper tests).
* The {@code name} carries the worker label the reaper keys on; {@code paneId}/{@code tabId}
* locate its pane for teardown.
*/
public FakeHerdr withAgent(String name, String terminalId, String paneId, String tabId) {
extraAgents.add(("{\"terminal_id\":\"%s\",\"agent\":\"claude\",\"agent_status\":\"idle\","
+ "\"name\":\"%s\",\"agent_session\":{\"kind\":\"id\",\"value\":\"sess-%s\"},"
+ "\"workspace_id\":\"wQ\",\"tab_id\":\"%s\",\"pane_id\":\"%s\"}")
.formatted(terminalId, name, terminalId, tabId, paneId));
return this;
}
/** Seed an additional workspace into {@code workspace.list} (e.g. a pre-existing worker space). */
public FakeHerdr withWorkspace(String id, String label) {
extraWorkspaces.add(("{\"workspace_id\":\"%s\",\"label\":\"%s\",\"focused\":false,"
@@ -91,11 +116,12 @@ public final class FakeHerdr implements HerdrClient {
{"workspace_id":"w1","label":"dev-mgnl","focused":true,"pane_count":7,"agent_status":"unknown"},
{"workspace_id":"w2","label":"ltms","focused":false,"pane_count":5,"agent_status":"done"}%s]}""")
.formatted(extraWorkspaces.isEmpty() ? "" : "," + String.join(",", extraWorkspaces)));
case "agent.list" -> mapper.readTree("""
case "agent.list" -> mapper.readTree(("""
{"type":"agent_list","agents":[
{"terminal_id":"term_a","agent":"claude","agent_status":"idle",
"agent_session":{"kind":"id","value":"sess-1111"},
"workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}]}""");
"workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}%s]}""")
.formatted(extraAgents.isEmpty() ? "" : "," + String.join(",", extraAgents)));
case "agent.send" -> {
if (agentSendErrorCode != null) {
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.send failed",
@@ -107,6 +133,8 @@ public final class FakeHerdr implements HerdrClient {
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":"claude",
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentStatus));
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
case "agent.start" -> {
long starts = calls.stream().filter(c -> c.method().equals("agent.start")).count();
if (starts <= agentNameTakenFor) {
@@ -114,10 +142,12 @@ public final class FakeHerdr implements HerdrClient {
"herdr error [agent_name_taken]: agent name already used",
"agent_name_taken", null);
}
yield mapper.readTree("""
long n = starts - agentNameTakenFor;
yield mapper.readTree(("""
{"type":"agent_started","agent":{
"terminal_id":"term_new","name":"claude","agent_status":"unknown",
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"w9:pW"}}""");
"terminal_id":"term_new_%d","name":"claude","agent_status":"unknown",
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"w9:pW_%d"}}""")
.formatted(n, n));
}
case "workspace.create" -> mapper.readTree("""
{"type":"workspace_created",
@@ -141,6 +171,21 @@ public final class FakeHerdr implements HerdrClient {
case "pane.get" -> mapper.readTree("""
{"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9",
"tab_id":"w9:t2","agent_status":"idle"}}""");
case "pane.list" -> mapper.readTree("""
{"type":"pane_list","panes":[
{"pane_id":"w2:p7","terminal_id":"term_a","workspace_id":"w2","tab_id":"w2:t7","agent":"claude"},
{"pane_id":"w2:p9","terminal_id":"term_shell","workspace_id":"w2","tab_id":"w2:t8"}]}""");
case "pane.process_info" -> {
Object paneId = params instanceof java.util.Map<?, ?> m ? m.get("pane_id") : null;
yield "w2:p7".equals(paneId)
? mapper.readTree(("""
{"type":"pane_process_info","process_info":{"pane_id":"w2:p7","shell_pid":%d,
"foreground_processes":[{"pid":%d,"name":"node","argv0":"claude"}]}}""")
.formatted(WORKER_PID, WORKER_PID))
: mapper.readTree("""
{"type":"pane_process_info","process_info":{"pane_id":"w2:p9","shell_pid":9001,
"foreground_processes":[]}}""");
}
case "pane.close" -> {
if (paneCloseErrorCode != null) {
throw new HerdrException("herdr error [" + paneCloseErrorCode + "]: pane.close failed",
@@ -0,0 +1,43 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import java.nio.file.Files;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the herdr half of connection-based identity against a REAL herdr: spawn a
* harmless probe, read its actual {@code shell_pid} from {@code pane.process_info}, and confirm
* {@link PaneLocator} resolves that PID back to the probe's own {@code terminal_id}.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@Tag("contract")
class PaneLocatorContractTest {
@Test
void resolvesTheTerminalOwningARealProcessPid() throws Exception {
assumeTrue(Files.exists(UnixSocketHerdrClient.defaultSocketPath()), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
AgentControl agents = new AgentControl(herdr);
Agent probe = agents.start("__pidprobe__", List.of("bash", "-c", "sleep 20"), Map.of());
try {
JsonNode info = herdr.call("pane.process_info", Map.of("pane_id", probe.paneId()))
.path("process_info");
long shellPid = info.path("shell_pid").asLong(-1);
assertTrue(shellPid > 0, "probe pane should report a shell pid");
assertEquals(probe.terminalId(), new PaneLocator(herdr).terminalForPid(shellPid),
"a real PID must resolve back to its own pane's terminal_id");
} finally {
agents.close(probe.paneId());
}
}
}
}
@@ -0,0 +1,27 @@
package dev.ltms.bridged.herdr;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/** Unit tests for PID → pane resolution (the herdr half of connection-based MCP identity). */
class PaneLocatorTest {
private final PaneLocator loc = new PaneLocator(new FakeHerdr());
@Test
void resolvesTerminalForAForegroundPid() {
assertEquals("term_a", loc.terminalForPid(FakeHerdr.WORKER_PID));
}
@Test
void nullForAPidInNoPane() {
assertNull(loc.terminalForPid(999_999));
}
@Test
void nullForNonPositivePid() {
assertNull(loc.terminalForPid(0));
assertNull(loc.terminalForPid(-1));
}
}
@@ -0,0 +1,227 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.msg.Rendezvous;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Unit behaviour of the CB-106 completion resolver in isolation from the injector. */
class CompletionResolverTest {
@Test
void skipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
resolver.resolve("term_a", null); // no in-flight turn captured for this target
assertFalse(herdr.called("agent.read"),
"a turn nobody is blocked on must not cost a transcript scrape");
}
@Test
void failSkipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
resolver.fail("term_a", null); // no in-flight turn, and no registered waiter to fall back to
assertFalse(herdr.called("agent.read"),
"a wedge nobody is blocked on must not cost a transcript scrape");
}
@Test
void captureBaselineSkipsTheReadWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ X\n❯ ");
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
resolver.captureBaseline("term_a"); // no send to attribute a later completion to
assertFalse(herdr.called("agent.read"),
"with no waiting send there is no turn to baseline — skip the scrape");
}
// --- CB-115 clean scrape: extract the last assistant block ----------------
@Test
void extractsTheLastAssistantBlockStrippingChrome() {
String raw = """
⏺ Reading the file…
⏺ Done. The bug was an off-by-one in the loop bound.
╭──────────────────────────────────────╮
│ > │
╰──────────────────────────────────────╯
⏵⏵ auto mode on · ? for shortcuts
""";
assertEquals("Done. The bug was an off-by-one in the loop bound.",
CompletionResolver.lastAssistantBlock(raw));
}
@Test
void keepsMultiLineAssistantContent() {
String raw = "⏺ Line one.\nLine two.\n❯ ";
assertEquals("Line one.\nLine two.", CompletionResolver.lastAssistantBlock(raw));
}
@Test
void fallsBackToRawTextWhenThereIsNoMarker() {
String raw = "plain worker output with no glyph";
assertEquals("plain worker output with no glyph", CompletionResolver.lastAssistantBlock(raw));
}
@Test
void blankScrapeYieldsEmpty() {
assertTrue(CompletionResolver.lastAssistantBlock("").isEmpty());
assertTrue(CompletionResolver.lastAssistantBlock(null).isEmpty());
}
@Test
void stripsSpinnerAndRuleChrome() {
String raw = """
⏺ Channel check confirmed — your message got through.
✻ Brewed for 11s
─────────────────────────────────────
""";
assertEquals("Channel check confirmed — your message got through.",
CompletionResolver.lastAssistantBlock(raw));
}
@Test
void cutsANextTurnPromptEchoAndTrailingTipsFromTheBlock() {
// The exact turn-2 leak: the scrape captured the settled answer, then a "✻ Cooked" spinner,
// then the NEXT turn's echoed prompt, then a "✶ Forming…" spinner and trailing tips/warnings
// whose lines (⎿, ⚠) are not themselves chrome-terminated. Stopping at the first boundary
// (the ✻ spinner) is what keeps every one of those interface lines out of the reply.
String raw = """
⏺ Channel confirmed — the bridge reply delivered successfully.
✻ Cooked for 9s
❯ Thanks. Now a small task: what is 17 * 23? Show just the number.
✶ Forming…
⎿ Tip: Name your conversations with /rename
⚠ claude.ai connectors are disabled because ANTHROPIC_API_KEY is set
""";
assertEquals("Channel confirmed — the bridge reply delivered successfully.",
CompletionResolver.lastAssistantBlock(raw));
}
// --- CB-115 misattribution guard: suppress a stale (unchanged) completion -------
@Test
void suppressesACompletionWhoseScrapeIsUnchangedFromDelivery() {
// Rapid back-to-back turn: the pane still shows the PREVIOUS turn's answer when this turn's
// (misattributed) completion boundary fires. The scrape == the delivery baseline, so the
// send must NOT be resolved with the stale answer.
FakeHerdr herdr = new FakeHerdr().readText("⏺ 391\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
// The turn as captured at delivery: its waiter, and the previous turn's answer still on screen.
var turn = new CompletionResolver.InFlight(waiter, "391");
resolver.resolve("term_a", turn); // scrape still "391" == baseline → suppress
assertFalse(waiter.isDone(), "a completion with no output change must not resolve the send");
assertTrue(rendezvous.isWaiting("term_a"), "the send stays waiting for a real reply");
}
@Test
void resolvesACompletionWhoseScrapeChangedSinceDelivery() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ No, 391 = 17 × 23.\n❯ "); // the worker's real answer
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
// Delivery baseline was the previous turn's "391"; the scrape now differs → resolve.
var turn = new CompletionResolver.InFlight(waiter, "391");
resolver.resolve("term_a", turn);
assertTrue(waiter.isDone(), "a completion with new output must resolve the send");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
assertEquals("No, 391 = 17 × 23.", waiter.getNow(null).text());
}
@Test
void suppressesAnUnchangedCompletionEvenWhenTheBlockExceedsTheScrapeCap() {
// The fan-out issue-hunt finding: captureBaseline once stored the RAW (unclipped) assistant
// block while resolve compares against a clip()'d tail. For a block longer than MAX_SCRAPE_CHARS
// the two capped representations differ even when the pane never changed, so the CB-115
// byte-identical guard failed to fire and a stale completion could resolve the send. Both sides
// must clip identically; here an unchanged >cap block on rapid back-to-back turns stays suppressed.
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
resolver.captureBaseline("term_a"); // baseline is the clipped >cap block
var turn = resolver.inFlight("term_a");
assertEquals(CompletionResolver.MAX_SCRAPE_CHARS, turn.baseline().length(),
"the delivery baseline is clipped to the same cap resolve() applies to the tail");
resolver.resolve("term_a", turn); // scrape unchanged → clipped tail == baseline → suppress
assertFalse(waiter.isDone(),
"an unchanged >cap block must still be recognised as stale and suppressed");
assertTrue(rendezvous.isWaiting("term_a"), "the send stays waiting for a real reply");
}
@Test
void resolvesWhenThereIsNoBaseline() {
// No delivery baseline (e.g. the pre-turn read failed) ⇒ never suppress; the completion resolves.
FakeHerdr herdr = new FakeHerdr().readText("⏺ hello\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "with no baseline a completion resolves as before");
assertEquals("hello", waiter.getNow(null).text());
}
// --- CB-116 waiter identity: a late completion never crosses into the next turn ---------
@Test
void aLateCompletionForOneTurnNeverResolvesTheNextTurnsWaiter() {
// The cross-turn stale reply the conversation test surfaced: turn N's completion fallback
// fires AFTER turn N was resolved by an explicit bridge_reply and turn N+1 has opened its own
// waiter on the same session. Resolving "whatever is waiting now" would hand turn N's stale
// scrape to turn N+1; targeting turn N's captured waiter makes the late completion a no-op.
FakeHerdr herdr = new FakeHerdr().readText("⏺ turn N answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiterN = rendezvous.open("term_a"); // turn N's send
// The turn as the injector captured it at delivery (waiter + pre-turn baseline).
var turnN = new CompletionResolver.InFlight(waiterN, "an earlier answer");
// Turn N is resolved by the worker's explicit reply.
assertTrue(rendezvous.resolve("term_a", "N replied"));
// Turn N+1's send opens its own waiter on the same session (replacing the registered one).
var waiterN1 = rendezvous.open("term_a");
resolver.resolve("term_a", turnN); // turn N's completion fallback finally fires
assertFalse(waiterN1.isDone(), "turn N's late completion must not resolve turn N+1's waiter");
assertEquals(Rendezvous.Kind.REPLY, waiterN.getNow(null).kind(),
"turn N stays resolved by its own reply");
assertTrue(rendezvous.isWaiting("term_a"), "turn N+1 is still awaiting its own resolution");
}
}
@@ -6,8 +6,10 @@ import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.TimeUnit;
@@ -25,12 +27,18 @@ class InjectorTest {
private final FakeHerdr herdr = new FakeHerdr();
private final Injector injector = new Injector(new AgentControl(herdr));
/** Text of every agent.send, in order. */
/**
* The logical messages delivered, in order. AgentControl.send emits each delivery as two
* agent.send calls — the payload, then a standalone Enter keystroke ({@code "\r"}) to submit
* it; these tests assert delivery ordering/gating, not the submit event, so drop the bare
* carriage returns.
*/
@SuppressWarnings("unchecked")
private List<String> sent() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.filter(t -> !t.equals("\r"))
.toList();
}
@@ -43,6 +51,47 @@ class InjectorTest {
assertEquals(List.of("hello"), sent());
}
@Test
void holdsDeliveryUntilTheWorkerIsAvailable() {
// CB-113: idle alone is not enough — hold until the worker's MCP is connected (ready).
java.util.Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // idle but not yet available → held out of the boot window
assertEquals(List.of(), sent(), "must not deliver into a not-yet-available worker");
ready.add(T); // the worker's Claude connects the bridge MCP
inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("task"), sent(), "delivers once the worker is available");
}
@SuppressWarnings("unchecked")
private long enterKeystrokes() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> "\r".equals(((Map<String, Object>) c.params()).get("text")))
.count();
}
@Test
void resubmitsEnterWhenADeliveredMessageIsNotPickedUp() {
// CB-113: the Enter at delivery can race the paste; while the worker stays idle (not picked
// up), the injector re-nudges Enter so the pending paste submits.
injector.enqueue(T, "task");
injector.onStatus(T, AgentStatus.IDLE); // deliver: paste + one Enter
long afterDeliver = enterKeystrokes();
injector.onStatus(T, AgentStatus.IDLE); // still idle → re-nudge Enter
injector.onStatus(T, AgentStatus.IDLE); // and again
assertTrue(enterKeystrokes() > afterDeliver, "an unpicked-up delivery re-nudges Enter");
injector.onStatus(T, AgentStatus.WORKING); // worker finally starts
long atPickup = enterKeystrokes();
injector.onStatus(T, AgentStatus.WORKING);
assertEquals(atPickup, enterKeystrokes(), "no more nudges once the worker has picked up");
}
@Test
void holdsWhileWorkingThenDeliversOnIdle() {
injector.enqueue(T, "later");
@@ -120,17 +169,97 @@ class InjectorTest {
}
@Test
void activeWhileQueuedOrInFlightThenQuietAfterPickup() {
void activeWhileQueuedOrInFlightThenQuietAfterTurnCompletes() {
assertTrue(injector.activeTargets().isEmpty());
injector.enqueue(T, "x");
assertEquals(java.util.Set.of(T), injector.activeTargets(), "active while a message is queued");
assertEquals(Set.of(T), injector.activeTargets(), "active while a message is queued");
injector.onStatus(T, AgentStatus.IDLE); // delivers; still in-flight (awaiting pickup)
assertEquals(java.util.Set.of(T), injector.activeTargets(),
injector.onStatus(T, AgentStatus.IDLE); // delivers; awaiting pickup
assertEquals(Set.of(T), injector.activeTargets(),
"stays active so the poller can observe the worker pick the message up");
injector.onStatus(T, AgentStatus.WORKING); // pickup observed → in-flight cleared
assertTrue(injector.activeTargets().isEmpty(), "quiet once queue is empty and pickup is seen");
injector.onStatus(T, AgentStatus.WORKING); // pickup observed; now awaiting turn completion
assertEquals(Set.of(T), injector.activeTargets(),
"stays active after pickup so the working→idle completion boundary is observed");
injector.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete
assertTrue(injector.activeTargets().isEmpty(), "quiet once the delegated turn has completed");
}
@Test
void firesTurnCompleteOnAConfirmedWorkingThenIdle() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // pickup + turn running
assertEquals(List.of(), completed, "no completion until the turn returns to idle");
inj.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete
assertEquals(List.of(T), completed, "a confirmed working→idle fires exactly one completion");
}
@Test
void doesNotSynthesizeCompletionFromAnUnconfirmedTurn() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
// Deliver, then only ever idle — a `working` sample is never seen. The pickup grace unwedges
// the queue but must NOT invent a completion: without a sampled turn there is no trustworthy
// "the worker finished the task" signal, so the send should fall through to its timeout.
for (int i = 0; i < 15; i++) inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of(), completed, "no completion is synthesized from an unconfirmed turn");
}
/** Captures both turn-lifecycle callbacks so the CB-109 stall path can be asserted. */
private static final class Captor implements TurnListener {
final List<String> completed = new ArrayList<>();
final List<String> failed = new ArrayList<>();
@Override
public void onTurnComplete(String target) {
completed.add(target);
}
@Override
public void onTurnFailed(String target) {
failed.add(target);
}
}
// ~30s of unknown at the 250ms prod poll interval; enough onStatus samples to trip the stall.
private static final int STALL_SAMPLES = 130;
@Test
void failsAnOutstandingDelegationWhoseWorkerWedgesInUnknown() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // worker starts the turn
for (int i = 0; i < STALL_SAMPLES; i++) inj.onStatus(T, AgentStatus.UNKNOWN); // then wedges
assertEquals(List.of(T), cap.failed, "a sustained unknown streak fails the outstanding send");
assertEquals(List.of(), cap.completed, "a wedge is a failure, not a completion");
assertTrue(inj.activeTargets().isEmpty(), "the wedged target is reclaimed, not polled forever");
}
@Test
void aTransientUnknownGlitchNeitherFailsNorBlocksCompletion() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // confirmed turn
for (int i = 0; i < 10; i++) inj.onStatus(T, AgentStatus.UNKNOWN); // brief glitch, well under grace
inj.onStatus(T, AgentStatus.IDLE); // working → idle: the real completion
assertEquals(List.of(), cap.failed, "a short unknown blip must not fail the turn");
assertEquals(List.of(T), cap.completed, "the streak reset, so the turn still completes");
}
@Test
@@ -151,6 +280,86 @@ class InjectorTest {
assertTrue(f.isCompletedExceptionally(), "queued waiters unblock when the worker vanishes");
}
@Test
void dropFailsTheTurnOfADeliveredMessageWhenTheWorkerVanishes() {
// CB-110: the message was delivered (no longer queued), so failing queued waiters alone would
// leave its send hanging. A vanished worker must fail that in-flight turn too.
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // turn running
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), cap.failed, "a worker that vanishes mid-turn fails its in-flight send");
}
@Test
void dropFailsADeliveredTurnThatVanishesBeforePickupIsConfirmed() {
// Delivered but no WORKING sampled yet (awaitingCompletion=true, awaitingPickup still true,
// turnObserved=false) — a distinct state the other two drop tests don't cover. (Gap surfaced
// by an off-sub worker's review of CB-110, delegated through the bridge.)
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver; pickup never confirmed
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), cap.failed, "a delivery that vanishes before pickup still fails its send");
}
// ~60s of idle-but-not-ready at the 250ms prod poll interval; enough to trip the readiness grace.
private static final int READINESS_SAMPLES = 245;
@Test
void failsAQueuedMessageWhoseWorkerNeverBecomesReady() {
// CB-114: herdr keeps reporting the worker idle, but its Claude never connects the bridge MCP,
// so the readiness gate never opens. The message must not be held (and the target polled)
// forever — after the grace it fails, the caller unblocks via the worker-failure path, the
// target is reclaimed, and the never-set presence is cleared.
Captor cap = new Captor();
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), cap, _ -> false, forgotten::add);
CompletableFuture<Void> f = inj.enqueue(T, "task");
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of(), sent(), "a never-ready worker is never delivered to");
assertTrue(f.isCompletedExceptionally(), "the caller's future fails instead of hanging forever");
assertEquals(List.of(T), cap.failed, "the awaiting send resolves through the worker-failure path");
assertEquals(List.of(), cap.completed, "a never-ready worker is a failure, not a completion");
assertEquals(List.of(T), forgotten, "the never-ready worker's presence is cleared");
assertTrue(inj.activeTargets().isEmpty(), "the target is reclaimed, not polled forever");
}
@Test
void aWorkerThatBecomesReadyWithinTheGraceIsDeliveredNormally() {
// The readiness grace must not fail a worker that is merely slow to boot: once it becomes
// available before the grace elapses, delivery proceeds as usual (the counter resets).
Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains, _ -> {
});
inj.enqueue(T, "task");
for (int i = 0; i < 100; i++) inj.onStatus(T, AgentStatus.IDLE); // still booting, well under grace
assertEquals(List.of(), sent());
ready.add(T); // MCP connects before the grace elapses
inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("task"), sent(), "a worker that connects within the grace is delivered to");
}
@Test
void dropClearsWorkerPresence() {
// CB-114 (finding #1): a vanished worker's readiness must be forgotten so a stale entry cannot
// linger past the worker's life (WorkerPresence.forget had no caller before this).
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, forgotten::add);
inj.enqueue(T, "orphan");
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), forgotten, "drop clears the gone worker's presence");
}
@Test
void pollerDeliversToAnIdleWorker() throws Exception {
// End-to-end through the poller: idle worker → message delivered without manual onStatus.
@@ -170,7 +379,9 @@ class InjectorTest {
@SuppressWarnings("unchecked")
Map<String, Object> p = (Map<String, Object>) c.params();
return p.get("text").toString();
}).toList());
})
.filter(t -> !t.equals("\r")) // drop the standalone submit keystroke
.toList());
}
@Test
@@ -0,0 +1,88 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Content-based refinement of an unreliable {@code UNKNOWN} status (CB-115). */
class StatusRefinerTest {
// --- pure classification -------------------------------------------------
@Test
void classifiesAnIdlePromptAsIdle() {
String pane = """
⏺ All done — the file compiles cleanly.
╭──────────────────────────────────────╮
│ > │
╰──────────────────────────────────────╯
⏵⏵ auto mode on (shift+tab to cycle)
""";
assertEquals(AgentStatus.IDLE, StatusRefiner.classify(pane));
}
@Test
void classifiesABarePromptGlyphAsIdle() {
assertEquals(AgentStatus.IDLE, StatusRefiner.classify("some output\n❯ "));
}
@Test
void classifiesActiveGenerationAsWorking() {
String pane = """
⏺ Working on it…
✳ Thinking… (12s · esc to interrupt)
""";
assertEquals(AgentStatus.WORKING, StatusRefiner.classify(pane));
}
@Test
void anEscToInterruptScreenIsWorkingEvenWithAPromptBox() {
// "esc to interrupt" wins over a prompt box: the turn is still generating.
String pane = "│ > │\n esc to interrupt";
assertEquals(AgentStatus.WORKING, StatusRefiner.classify(pane));
}
@Test
void anUnrecognizableScreenStaysUnknown() {
assertEquals(AgentStatus.UNKNOWN, StatusRefiner.classify("garbled ansi noise with no prompt"));
assertEquals(AgentStatus.UNKNOWN, StatusRefiner.classify(""));
assertEquals(AgentStatus.UNKNOWN, StatusRefiner.classify(null));
}
// --- refine() wiring -----------------------------------------------------
@Test
void refinePassesNonUnknownStatusesThroughWithoutReading() {
FakeHerdr herdr = new FakeHerdr();
StatusRefiner refiner = new StatusRefiner(new AgentControl(herdr));
assertEquals(AgentStatus.WORKING, refiner.refine("term_a", AgentStatus.WORKING));
assertEquals(AgentStatus.IDLE, refiner.refine("term_a", AgentStatus.IDLE));
assertFalse(herdr.called("agent.read"),
"a trusted status must not cost a pane read");
}
@Test
void refineUpgradesUnknownToIdleFromPaneContent() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ answer\n❯ ");
StatusRefiner refiner = new StatusRefiner(new AgentControl(herdr));
assertEquals(AgentStatus.IDLE, refiner.refine("term_a", AgentStatus.UNKNOWN));
assertTrue(herdr.called("agent.read"), "an UNKNOWN must trigger a pane read");
}
@Test
void refineLeavesUnknownWhenContentIsUnclassifiable() {
FakeHerdr herdr = new FakeHerdr().readText("nothing recognizable here");
StatusRefiner refiner = new StatusRefiner(new AgentControl(herdr));
assertEquals(AgentStatus.UNKNOWN, refiner.refine("term_a", AgentStatus.UNKNOWN));
}
}
@@ -0,0 +1,33 @@
package dev.ltms.bridged.inject;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** The CB-113 worker-availability registry. */
class WorkerPresenceTest {
@Test
void tracksPresenceAndForgets() {
WorkerPresence p = new WorkerPresence();
assertFalse(p.isPresent("term_a"), "unseen worker is not available");
p.markPresent("term_a");
assertTrue(p.isPresent("term_a"), "a worker seen on the MCP is available");
p.forget("term_a");
assertFalse(p.isPresent("term_a"), "a torn-down worker is no longer available");
}
@Test
void nullOrBlankMarkIsANoOp() {
WorkerPresence p = new WorkerPresence();
assertDoesNotThrow(() -> {
p.markPresent(null);
p.markPresent(" ");
});
assertFalse(p.isPresent(""), "blank/null contacts (the primary) are never present");
}
}
@@ -0,0 +1,288 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
/**
* Parity tests for the MCP tool adapters — they must produce the same outcomes as the REST routes,
* since both drive the same {@link MessageService}/{@link Rendezvous}. The MCP wire protocol itself
* is the SDK's concern; here we test the thin adapter logic directly.
*/
class BridgeMcpTest {
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final MessageService messages = new MessageService(agents, new Injector(agents), rendezvous);
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
private static ClaudeCodeLauncher workerService(FakeHerdr h, String baseUrl, Set<String> allow) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", baseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(allow), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> "tok");
}
private static SessionManager sessionManager(FakeHerdr h, String baseUrl, Set<String> allow) {
return new SessionManager(workerService(h, baseUrl, allow));
}
@Test
void sendThenReplyRoundTrips() throws Exception {
// bridge_send blocks; bridge_reply resolves it with the worker's structured answer.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> BridgeMcp.send(messages, "term_a", "review this", 4000L));
McpSchema.CallToolResult reply = BridgeMcp.reply(rendezvous, "term_a", "LGTM");
long deadline = System.currentTimeMillis() + 3000;
while (Boolean.TRUE.equals(reply.isError()) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(10);
reply = BridgeMcp.reply(rendezvous, "term_a", "LGTM");
}
assertEquals("delivered", textOf(reply));
McpSchema.CallToolResult res = send.get(6, TimeUnit.SECONDS);
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("LGTM", textOf(res));
}
@Test
void asyncSendReturnsATicketThenPollReportsTheReply() throws Exception {
// wait:false parity — a ticket is issued, resolved by a reply, and surfaced by bridge_poll.
McpSchema.CallToolResult accepted = BridgeMcp.sendAsync(messages, "term_a", "do it");
assertNotEquals(Boolean.TRUE, accepted.isError());
String out = textOf(accepted);
assertTrue(out.contains("ticket="), out);
String ticket = out.substring(out.indexOf("ticket=") + "ticket=".length()).trim();
// Resolve the awaiting send once it has opened (retry past the async-open race).
long deadline = System.currentTimeMillis() + 3000;
McpSchema.CallToolResult reply = BridgeMcp.reply(rendezvous, "term_a", "async LGTM");
while (Boolean.TRUE.equals(reply.isError()) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(10);
reply = BridgeMcp.reply(rendezvous, "term_a", "async LGTM");
}
assertEquals("delivered", textOf(reply));
// Poll until the async send completes and reports the reply.
McpSchema.CallToolResult polled = BridgeMcp.poll(messages, ticket);
deadline = System.currentTimeMillis() + 3000;
while (!textOf(polled).contains("async LGTM") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(10);
polled = BridgeMcp.poll(messages, ticket);
}
assertEquals("async LGTM", textOf(polled));
}
@Test
void pollUnknownTicketIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.poll(messages, "task-999");
assertTrue(res.isError());
assertTrue(textOf(res).contains("unknown ticket"));
}
@Test
void sendTimesOutWithAWorkingNote() {
McpSchema.CallToolResult res = BridgeMcp.send(messages, "term_a", "hi", 120L);
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res));
}
@Test
void sendRejectsMissingArgs() {
assertTrue(BridgeMcp.send(messages, null, "hi", null).isError());
assertTrue(BridgeMcp.send(messages, "term_a", " ", null).isError());
}
@Test
void replyWithNoPendingSendIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.reply(rendezvous, "term_a", "orphan");
assertTrue(res.isError());
assertTrue(textOf(res).contains("no send is awaiting"));
}
@Test
void askThenAnswerRoundTrips() throws Exception {
// The primary delegates and blocks; wait until its waiter is open before the worker asks.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> BridgeMcp.send(messages, "term_a", "do X", 5000L));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "the send must be waiting for the ask to surface to");
// The worker asks mid-turn; the call blocks for the primary's answer.
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
() -> BridgeMcp.ask(messages, "term_a", "which config?", 5000L));
// The primary's send unblocks with the question and a turnId to answer on.
McpSchema.CallToolResult q = send.get(6, TimeUnit.SECONDS);
assertNotEquals(Boolean.TRUE, q.isError());
String qt = textOf(q);
assertTrue(qt.contains("[question]"), qt);
String afterMarker = qt.substring(qt.indexOf("turnId=\"") + "turnId=\"".length());
String turnId = afterMarker.substring(0, afterMarker.indexOf('"'));
// The primary answers via bridge_send(turnId); this blocks again for the worker's reply.
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> BridgeMcp.answer(messages, turnId, "config.yaml", 5000L));
// The worker's ask returns the answer — it resumes the same turn.
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
// The resumed worker replies, resolving the answering send (retry past the reopen race).
McpSchema.CallToolResult reply = BridgeMcp.reply(rendezvous, "term_a", "done");
deadline = System.currentTimeMillis() + 3000;
while (Boolean.TRUE.equals(reply.isError()) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(10);
reply = BridgeMcp.reply(rendezvous, "term_a", "done");
}
assertEquals("delivered", textOf(reply));
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
}
@Test
void askFromANonWorkerConnectionIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.ask(messages, null, "which config?", 500L);
assertTrue(res.isError());
assertTrue(textOf(res).contains("workers only"), textOf(res));
}
@Test
void answerToAStaleTurnIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.answer(messages, "term_a#999", "too late", 500L);
assertTrue(res.isError());
assertTrue(textOf(res).contains("no longer open"), textOf(res));
}
@Test
void spawnReturnsTheNewWorkersSessionAndPane() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"sessionId\":\"term_new_1\""), out);
assertTrue(out.contains("\"paneId\":\"w9:pW_1\""), out);
assertTrue(out.contains("\"status\":\"spawning\""), out);
}
@Test
void spawnRejectsAnOffAllowlistProfileWithoutTouchingHerdr() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res =
BridgeMcp.spawn(sessionManager(h, "https://api.anthropic.com", Set.of("gx00.gw")), null);
assertTrue(res.isError());
assertTrue(textOf(res).contains("subscription boundary"));
assertFalse(h.called("agent.start"), "the guard must block before any spawn");
}
@Test
void spawnRejectsAnUnknownProfileAsAnError() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res =
BridgeMcp.spawn(sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), "nope");
assertTrue(res.isError());
assertTrue(textOf(res).contains("unknown worker profile"), textOf(res));
}
@Test
void spawnPassesTheRequestedCwdToTheWorker() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "/req/dir", null, null, null);
assertNotEquals(Boolean.TRUE, res.isError());
@SuppressWarnings("unchecked")
Map<String, Object> start = (Map<String, Object>) h.lastCall("agent.start").params();
assertEquals("/req/dir", start.get("cwd"));
}
@Test
void profilesListsConfiguredProfilesAndDefault() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.profiles(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("ltms-local"), out);
assertTrue(out.contains("\"default\":\"ltms-local\""), out);
}
@Test
void listReportsTrackedWorkers() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-304", null));
McpSchema.CallToolResult res = BridgeMcp.listWorkers(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"sessionId\":\"" + s.terminalId() + "\""), out);
assertTrue(out.contains("\"paneId\":\"" + s.paneId() + "\""), out);
assertTrue(out.contains("\"profile\":\"ltms-local\""), out);
assertTrue(out.contains("\"state\":\"spawning\""), out);
assertTrue(out.contains("\"worktree\":\"" + s.worktree() + "\""), out);
assertTrue(out.contains("\"branch\":\"" + s.branch() + "\""), out);
assertTrue(out.contains("\"owner\":\"term_primary\""), out);
assertTrue(out.contains("\"liveStatus\":\"unknown\""), out);
}
@Test
void stopTearsDownAWorkerByPane() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.stop(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), "w9:pW");
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("stopped w9:pW", textOf(res));
assertTrue(h.called("pane.close"));
}
@Test
void stopRequiresAPaneId() {
FakeHerdr h = new FakeHerdr();
assertTrue(BridgeMcp.stop(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), " ").isError());
}
@Test
void statusReportsLiveAgentStatus() {
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
AgentControl blockedAgents = new AgentControl(blocked);
McpSchema.CallToolResult res = BridgeMcp.status(
new MessageService(blockedAgents, new Injector(blockedAgents), rendezvous), "term_a");
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("blocked", textOf(res));
}
}
@@ -0,0 +1,46 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/** Connection → caller-identity resolution, with the OS peer-PID lookup faked. */
class ConnectionIdentityTest {
private final FakeHerdr herdr = new FakeHerdr();
private ConnectionIdentity with(PeerPidLookup pids) {
return new ConnectionIdentity(new PaneLocator(herdr), pids);
}
@Test
void resolvesWorkerFromLoopbackPeerPid() {
assertEquals("term_a", with(_ -> FakeHerdr.WORKER_PID).callerTerminal("127.0.0.1", 55555));
}
@Test
void nullForOffHostCaller() {
// A non-loopback peer can't be an on-host worker → treat as primary/unknown.
assertNull(with(_ -> FakeHerdr.WORKER_PID).callerTerminal("10.0.0.9", 55555));
}
@Test
void nullWhenPidOwnsNoPane() {
// e.g. the primary — its PID maps to no worker pane.
assertNull(with(_ -> 999_999).callerTerminal("127.0.0.1", 55555));
}
@Test
void resolvesTheCallersPidAndCwd() {
// CB-112: the primary maps to no pane, but its PID and cwd are still readable.
ConnectionIdentity id = new ConnectionIdentity(
new PaneLocator(herdr), _ -> 999_999, pid -> pid == 999_999 ? "/main/project" : null);
ConnectionIdentity.Caller c = id.resolve("127.0.0.1", 55555);
assertNull(c.terminal(), "the primary owns no worker pane");
assertEquals(999_999, c.pid());
assertEquals("/main/project", id.cwdForPid(c.pid()), "the primary's cwd is resolvable from its PID");
assertNull(id.cwdForPid(-1), "no cwd for an unresolved PID");
}
}
@@ -0,0 +1,214 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.Injector;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* The message layer's resolution paths (CB-104 reply + CB-106 completion fallback). The turn is
* driven deterministically by feeding {@code onStatus} rather than running a real poller.
*/
class MessageServiceTest {
private static final String T = "term_a";
private final FakeHerdr herdr = new FakeHerdr().readText("BUILD GREEN: 391 files");
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final CompletionResolver completion = new CompletionResolver(agents, rendezvous);
private final Injector injector = new Injector(agents, completion);
private final MessageService messages = new MessageService(agents, injector, rendezvous);
/** Run {@code send} on a background thread; the current thread drives the worker's turn. */
private CompletableFuture<MessageService.Reply> sendAsync() {
return CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 5000));
}
private void awaitWaiting() throws InterruptedException {
long deadline = System.currentTimeMillis() + 2000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T), "send should have opened its rendezvous waiter");
}
@Test
void completionFallbackResolvesATurnThatNeverCalledBridgeReply() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
herdr.readText("$ prompt"); // pre-turn pane: no answer yet (baseline reference)
injector.onStatus(T, AgentStatus.IDLE); // deliver the task (baselines the pre-turn content)
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up and works
herdr.readText("BUILD GREEN: 391 files"); // the worker's turn produced new output
injector.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete, no bridge_reply
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, reply.outcome(),
"an unreplied but finished turn resolves via the completion fallback");
assertEquals("BUILD GREEN: 391 files", reply.text(), "the scraped transcript tail is returned");
assertTrue(reply.completed(), "a scraped completion still counts as completed");
}
@Test
void explicitBridgeReplyResolvesAsReplied() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker working
assertTrue(rendezvous.resolve(T, "LGTM ship it"), "an explicit reply resolves the send");
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, reply.outcome());
assertEquals("LGTM ship it", reply.text());
}
@Test
void aWedgedWorkerResolvesTheSendAsFailedWithTheErrorContext() throws Exception {
herdr.readText("API Error: Unable to connect to API (ENOTFOUND)");
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker starts the turn
for (int i = 0; i < 130; i++) injector.onStatus(T, AgentStatus.UNKNOWN); // then wedges (CB-109)
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome());
assertFalse(reply.completed(), "a wedge is terminal but not a successful completion");
assertTrue(reply.text().contains("ENOTFOUND"), "the error screen is carried as the failure reason");
}
@Test
void aWorkerThatVanishesMidTurnResolvesTheSendAsFailed() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker starts the turn
// The worker's pane crashes — the poller sees a *_not_found and drops it (CB-110).
injector.drop(T, new HerdrException("worker gone", "pane_not_found", null));
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome(),
"a delivered send whose worker vanishes fails instead of hanging to the timeout");
assertFalse(reply.completed());
}
// --- bridge_ask reverse rendezvous (CB-205) ------------------------------------------------
@Test
void askSurfacesAsAQuestionAndTheAnswerResumesTheSameTurn() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up, then pauses to ask
// The worker asks mid-turn on its own thread; the call blocks for the primary's answer.
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
// The primary's blocking send unblocks with the question and a turnId to answer on.
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertEquals("which config file?", q.text());
assertNotNull(q.turnId(), "a question carries a turnId to answer on");
// The primary answers via bridge_send(turnId); this blocks again for the worker's reply.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
// The worker's ask returns the answer — it resumes the same turn.
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome());
assertEquals("config.yaml", a.answer());
// The resumed worker finishes with a structured reply, resolving the answering send.
awaitWaiting(); // the answering send has (re)opened its forward waiter
assertTrue(rendezvous.resolve(T, "done"), "the worker's final reply resolves the answering send");
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
assertEquals("done", done.text());
}
@Test
void duplicateAsksFromTheSameSessionCoalesceToOneTurn() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up, then pauses to ask
// A transport retry: two concurrent bridge_ask calls from the same worker session.
CompletableFuture<MessageService.AskResult> ask1 =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
CompletableFuture<MessageService.AskResult> ask2 =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
// The primary's single blocked send surfaces exactly ONE question (one turnId).
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertEquals("which config file?", q.text());
assertNotNull(q.turnId(), "only one turnId should be minted");
// The primary answers that one turnId; both asks unblock with the same answer.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
MessageService.AskResult a1 = ask1.get(5, TimeUnit.SECONDS);
MessageService.AskResult a2 = ask2.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a1.outcome());
assertEquals("config.yaml", a1.answer());
assertEquals(MessageService.AskOutcome.ANSWERED, a2.outcome());
assertEquals("config.yaml", a2.answer());
// The resumed worker finishes with a structured reply, resolving the answering send.
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"), "the worker's final reply resolves the answering send");
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
assertEquals("done", done.text());
}
@Test
void askWithNoOpenDelegationReturnsNoWaiter() {
MessageService.AskResult r = messages.ask(T, "anyone listening?", 500);
assertEquals(MessageService.AskOutcome.NO_WAITER, r.outcome(),
"a question with no blocked send has no primary to answer it");
}
@Test
void askTimesOutWhenThePrimaryNeverAnswers() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
MessageService.AskResult r = messages.ask(T, "still there?", 200); // primary never answers
assertEquals(MessageService.AskOutcome.TIMED_OUT, r.outcome());
// The send itself already unblocked with the question the instant the ask surfaced.
MessageService.Reply q = send.get(2, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
}
@Test
void answeringAnUnknownTurnIsStale() {
MessageService.Reply r = messages.answer(T + "#999", "too late", 500);
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
"an answer to a turn that never existed (or already lapsed) is stale, not a hang");
}
}
@@ -0,0 +1,80 @@
package dev.ltms.bridged.msg;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CompletableFuture;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* The reverse rendezvous (CB-205): the {@code bridge_ask} registry that lets a worker pause mid-turn
* to ask the primary. Unit-level — the message-layer round-trip is covered in {@link MessageServiceTest}.
*/
class RendezvousTest {
private static final String W = "term_a";
private final Rendezvous rendezvous = new Rendezvous();
@Test
void openAskMintsAUniqueTurnScopedToItsSessionAndCoalescesDuplicates() {
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
Rendezvous.AskTicket t2 = rendezvous.openAsk(W);
assertEquals(t1.turnId(), t2.turnId(), "duplicate asks from the same session coalesce onto one turn");
assertTrue(t1.fresh(), "the first ask freshly opens the turn");
assertFalse(t2.fresh(), "the coalesced ask rides the existing turn");
assertTrue(t1.turnId().startsWith(W + "#"), "the turnId is scoped to the worker session");
assertEquals(W, rendezvous.askSession(t1.turnId()));
}
@Test
void openAskAfterCloseMintsANewTurn() {
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
rendezvous.closeAsk(t1.turnId());
Rendezvous.AskTicket t2 = rendezvous.openAsk(W);
assertNotEquals(t1.turnId(), t2.turnId(), "after closing, a new ask gets a fresh turnId");
assertTrue(t2.fresh(), "the reopened ask is fresh");
assertEquals(W, rendezvous.askSession(t2.turnId()));
}
@Test
void answerAskCompletesTheWaitersFuture() {
Rendezvous.AskTicket t = rendezvous.openAsk(W);
assertTrue(rendezvous.answerAsk(t.turnId(), "config.yaml"), "answering an open ask succeeds");
assertEquals("config.yaml", t.answer().getNow(null), "the answer reaches the blocked worker");
}
@Test
void answerAskOnAnUnknownTurnIsFalse() {
assertFalse(rendezvous.answerAsk("no-such#1", "x"), "an answer to an unknown turn is a no-op");
}
@Test
void resolveQuestionResolvesAnOpenSendWithTheQuestionKindAndTurnId() {
CompletableFuture<Rendezvous.Resolution> send = rendezvous.open(W);
assertTrue(rendezvous.resolveQuestion(W, "which config?", W + "#7"),
"the question resolves the primary's open send");
Rendezvous.Resolution r = send.getNow(null);
assertEquals(Rendezvous.Kind.QUESTION, r.kind());
assertEquals("which config?", r.text());
assertEquals(W + "#7", r.turnId(), "the turnId rides along so the primary can answer");
}
@Test
void resolveQuestionWithNoOpenSendIsFalse() {
assertFalse(rendezvous.resolveQuestion(W, "anyone?", W + "#1"),
"no blocked send means no primary to surface the question to");
}
@Test
void closeAskRemovesTheTurn() {
Rendezvous.AskTicket t = rendezvous.openAsk(W);
rendezvous.closeAsk(t.turnId());
assertNull(rendezvous.askSession(t.turnId()), "a closed ask is forgotten");
assertFalse(rendezvous.answerAsk(t.turnId(), "late"), "a closed ask can no longer be answered");
}
}
@@ -7,7 +7,16 @@ import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.WorkerService;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.Worktrees;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
@@ -32,10 +41,13 @@ class BridgedAppTest {
private final ObjectMapper mapper = new ObjectMapper();
private final HttpClient http = HttpClient.newHttpClient();
private WorkerPresence presence;
private Javalin app;
private StatusPoller poller;
@AfterEach
void stop() {
if (poller != null) poller.stop();
if (app != null) app.stop();
}
@@ -44,13 +56,27 @@ class BridgedAppTest {
}
private int start(FakeHerdr herdr, String workerBaseUrl, Set<String> allow, String placement) {
return start(herdr, workerBaseUrl, allow, placement, new GitWorktrees());
}
private int start(FakeHerdr herdr, String workerBaseUrl, Set<String> allow, String placement, Worktrees worktrees) {
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
"ltms-local", workerBaseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
placement, "bridged-workers", "worker: {profile} #{n}");
WorkerService workers = new WorkerService(
new AgentControl(herdr), new WorkspaceControl(herdr), new SubscriptionGuard(allow), wcfg,
placement, "bridged-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(allow),
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
k -> "BRIDGED_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
app = new BridgedApp(herdr, workers).build().start("127.0.0.1", 0);
SessionManager sessions = new SessionManager(workers, worktrees);
this.presence = sessions.asPresence();
Injector injector = new Injector(agents);
poller = new StatusPoller(agents, injector, 5); // delivers when the fake reports idle
poller.start();
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
app = new BridgedApp(herdr, workers, sessions, messages, rendezvous, this.presence, null)
.build().start("127.0.0.1", 0);
return app.port();
}
@@ -58,6 +84,18 @@ class BridgedAppTest {
return start(new FakeHerdr(), "http://gx00.gw:8000", Set.of("gx00.gw"));
}
private HttpResponse<String> postMessage(int port, String json) throws Exception {
return postJson(port, "/sessions/term_a/message", json);
}
private HttpResponse<String> postJson(int port, String path, String json) throws Exception {
HttpRequest r = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path))
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(json)).build();
return http.send(r, HttpResponse.BodyHandlers.ofString());
}
private HttpResponse<String> req(int port, String method, String path) throws Exception {
HttpRequest.Builder b = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path));
b = switch (method) {
@@ -85,7 +123,8 @@ class BridgedAppTest {
@Test
void healthzDegradedWhenHerdrDown() throws Exception {
int port = start(new FakeHerdr().healthy(false), "http://gx00.gw:8000", Set.of("gx00.gw"));
FakeHerdr down = new FakeHerdr().healthy(false);
int port = start(down, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "GET", "/healthz");
assertEquals(503, res.statusCode());
assertEquals("degraded", mapper.readTree(res.body()).get("status").asText());
@@ -117,8 +156,8 @@ class BridgedAppTest {
HttpResponse<String> res = req(port, "POST", "/workers");
assertEquals(201, res.statusCode());
JsonNode body = mapper.readTree(res.body());
assertEquals("w9:pW", body.get("paneId").asText());
assertEquals("w9:t2", body.get("tabId").asText());
assertEquals("w9:pW_1", body.get("paneId").asText());
assertEquals("spawning", body.get("state").asText());
// Subscription boundary: agent.start carried base_url + token in its env map.
Map<String, Object> start = params(herdr, "agent.start");
@@ -137,6 +176,58 @@ class BridgedAppTest {
"tab label carries the worker number so siblings stay distinct");
}
@Test
void profilesEndpointListsConfiguredProfilesAndDefault() throws Exception {
int port = startHealthy();
JsonNode body = mapper.readTree(req(port, "GET", "/profiles").body());
assertEquals("ltms-local", body.get("default").asText());
assertEquals("ltms-local", body.get("profiles").get(0).asText());
}
@Test
void workersEndpointReturnsRegistryRosterWithLiveStatus() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"), "tab",
new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt"));
HttpResponse<String> spawn = req(port, "POST", "/workers?worktree=true&ticket=cb-304");
assertEquals(201, spawn.statusCode());
JsonNode spawned = mapper.readTree(spawn.body());
String paneId = spawned.get("paneId").asText();
HttpResponse<String> res = req(port, "GET", "/workers");
assertEquals(200, res.statusCode());
JsonNode workers = mapper.readTree(res.body()).get("workers");
assertEquals(1, workers.size());
JsonNode w = workers.get(0);
assertEquals(spawned.get("terminalId").asText(), w.get("sessionId").asText());
assertEquals(paneId, w.get("paneId").asText());
assertEquals("ltms-local", w.get("profile").asText());
assertEquals("spawning", w.get("state").asText());
assertTrue(w.has("worktree"), "worktree-backed session exposes worktree");
assertTrue(w.has("branch"), "worktree-backed session exposes branch");
assertEquals("unknown", w.get("liveStatus").asText(),
"liveStatus is unknown when herdr has no matching pane");
}
@Test
void spawnWithACwdParamRootsTheWorkerThere() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(201, req(port, "POST", "/workers?cwd=/tmp/proj").statusCode());
assertEquals("/tmp/proj", params(herdr, "agent.start").get("cwd"), "the worker starts in cwd");
}
@Test
void spawnWithAnUnknownProfileIs400() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "POST", "/workers?profile=nope");
assertEquals(400, res.statusCode());
assertEquals("unknown_profile", mapper.readTree(res.body()).get("error").asText());
assertFalse(herdr.called("agent.start"), "an unknown profile must not spawn anything");
}
@Test
void spawnWorkerReusesExistingWorkerSpace() throws Exception {
// A space labelled "bridged-workers" already exists → no second workspace.create.
@@ -211,6 +302,133 @@ class BridgedAppTest {
assertEquals("w9:t2", params(herdr, "tab.close").get("tab_id"));
}
@Test
void messageReturnsTheWorkersStructuredReply() throws Exception {
// CB-104 (option C): the blocking send resolves on the worker's bridge_reply, not a scrape.
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // poller delivers the injection
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
var send = java.util.concurrent.CompletableFuture.supplyAsync(() -> {
try { return postMessage(port, "{\"content\":\"review this\",\"timeoutMs\":4000}"); }
catch (Exception e) { throw new RuntimeException(e); }
});
// The worker replies once a send is actually awaiting (retry past the startup race).
HttpResponse<String> reply;
long deadline = System.currentTimeMillis() + 3000;
do {
reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"LGTM ship it\"}");
if (reply.statusCode() != 409) break;
//noinspection BusyWait
Thread.sleep(10);
} while (System.currentTimeMillis() < deadline);
assertEquals(200, reply.statusCode());
HttpResponse<String> res = send.get(6, java.util.concurrent.TimeUnit.SECONDS);
assertEquals(200, res.statusCode());
assertEquals("LGTM ship it", mapper.readTree(res.body()).get("reply").asText());
// (injection via agent.send is covered deterministically by the timeout-working test)
}
@Test
void asyncSendReturnsATicketThenPollReportsTheReply() throws Exception {
// CB-107 fire-and-poll: wait:false returns a ticket immediately; the result is polled.
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // poller delivers the injection
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> accepted = postMessage(port, "{\"content\":\"do it\",\"wait\":false}");
assertEquals(202, accepted.statusCode());
String ticket = mapper.readTree(accepted.body()).get("ticket").asText();
assertFalse(ticket.isBlank(), "an async send must return a ticket");
// The worker replies once the async send is actually awaiting (retry past the startup race).
HttpResponse<String> reply;
long deadline = System.currentTimeMillis() + 3000;
do {
reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"async LGTM\"}");
if (reply.statusCode() != 409) break;
//noinspection BusyWait
Thread.sleep(10);
} while (System.currentTimeMillis() < deadline);
assertEquals(200, reply.statusCode());
// Polling the ticket now reports the finished delegation and its reply.
JsonNode task;
deadline = System.currentTimeMillis() + 3000;
do {
task = mapper.readTree(req(port, "GET", "/tasks/" + ticket).body());
if ("done".equals(task.path("phase").asText())) break;
//noinspection BusyWait
Thread.sleep(10);
} while (System.currentTimeMillis() < deadline);
assertEquals("done", task.get("phase").asText());
assertEquals("async LGTM", task.get("reply").asText());
assertEquals("reply", task.get("replySource").asText());
}
@Test
void pollUnknownTicketIs404() throws Exception {
int port = startHealthy();
HttpResponse<String> res = req(port, "GET", "/tasks/task-999");
assertEquals(404, res.statusCode());
assertEquals("unknown_ticket", mapper.readTree(res.body()).get("error").asText());
}
@Test
void replyWithNoPendingSendIsConflict() throws Exception {
int port = startHealthy();
HttpResponse<String> res = postJson(port, "/sessions/term_a/reply", "{\"content\":\"orphan\"}");
assertEquals(409, res.statusCode());
assertEquals("no_pending_send", mapper.readTree(res.body()).get("error").asText());
}
@Test
void messageTimesOutQueuedWhenWorkerNeverInjectable() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("working"); // never injectable → never delivered
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":150}");
assertEquals(202, res.statusCode());
assertEquals("queued", mapper.readTree(res.body()).get("status").asText());
assertFalse(herdr.called("agent.send"), "no injection while the worker is mid-turn");
}
@Test
void messageTimesOutWorkingWhenDeliveredButNoReply() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // delivered, but nobody replies
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":250}");
assertEquals(202, res.statusCode());
assertEquals("working", mapper.readTree(res.body()).get("status").asText());
assertTrue(herdr.called("agent.send"), "message was injected");
}
@Test
void messageRejectsBlankContent() throws Exception {
int port = startHealthy();
assertEquals(400, postMessage(port, "{}").statusCode());
}
@Test
void sessionStatusReportsLiveAgentStatus() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("blocked");
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "GET", "/sessions/term_a/status");
assertEquals(200, res.statusCode());
assertEquals("blocked", mapper.readTree(res.body()).get("status").asText());
}
@Test
void sessionStatusReportsReadinessFromMcpPresence() throws Exception {
int port = start(new FakeHerdr(), "http://gx00.gw:8000", Set.of("gx00.gw"));
// Not yet seen on the bridge MCP → not ready.
assertFalse(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
// Worker connects its MCP client → available.
presence.markPresent("term_a");
assertTrue(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
}
@Test
void stopWorkerInPanePlacementClosesOnlyThePane() throws Exception {
FakeHerdr herdr = new FakeHerdr();
@@ -0,0 +1,130 @@
package dev.ltms.bridged.session;
import java.util.Collections;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.CopyOnWriteArrayList;
/** Recording fake {@link Worktrees} for CB-301-ext acceptance tests (no live git). */
public final class FakeWorktrees implements Worktrees {
public record AddCall(String repoRoot, String branch, String baseRef) {
}
public record RemoveCall(String repoRoot, String worktreePath) {
}
public record OverlayCall(String repoRoot, String worktreePath,
List<String> requested, List<String> copied, List<String> skipWorktree) {
}
public record RepoRootCall(String cwd) {
}
private final List<AddCall> addCalls = new CopyOnWriteArrayList<>();
private final List<RemoveCall> removeCalls = new CopyOnWriteArrayList<>();
private final List<OverlayCall> overlayCalls = new CopyOnWriteArrayList<>();
private final List<RepoRootCall> repoRootCalls = new CopyOnWriteArrayList<>();
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
private volatile RuntimeException addFailure;
private volatile String repoRoot = "/repo";
private volatile String prefix = "/worktrees";
public FakeWorktrees withRepoRoot(String root) {
this.repoRoot = root;
return this;
}
public FakeWorktrees withPrefix(String prefix) {
this.prefix = prefix;
return this;
}
/** Paths that exist in the primary repo and will be copied to the worktree. */
public FakeWorktrees exists(String... paths) {
Collections.addAll(existingPaths, paths);
return this;
}
/** Paths that exist AND are tracked, so overlayParity should --skip-worktree them. */
public FakeWorktrees track(String... paths) {
exists(paths);
Collections.addAll(trackedPaths, paths);
return this;
}
/** Make subsequent {@link #add} calls throw (simulates git worktree add failure). */
public FakeWorktrees failAdd(String message) {
this.addFailure = new WorktreeException(message);
return this;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
addCalls.add(new AddCall(repoRoot, branch, baseRef));
if (addFailure != null) {
throw addFailure;
}
// The branch already carries a unique nonce, so the derived path is distinct per acquire
// without an extra counter — keep it a pure function of the branch the test can predict.
return prefix + "/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
removeCalls.add(new RemoveCall(repoRoot, worktreePath));
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
List<String> copied = new java.util.ArrayList<>();
List<String> skipped = new java.util.ArrayList<>();
for (String rel : overlay) {
if (!existingPaths.contains(rel)) {
continue; // missing source is silently skipped
}
copied.add(rel);
if (trackedPaths.contains(rel)) {
skipped.add(rel);
}
}
overlayCalls.add(new OverlayCall(repoRoot, worktreePath, List.copyOf(overlay),
List.copyOf(copied), List.copyOf(skipped)));
}
@Override
public String repoRoot(String cwd) {
repoRootCalls.add(new RepoRootCall(cwd));
return repoRoot;
}
public List<AddCall> addCalls() {
return List.copyOf(addCalls);
}
public List<RemoveCall> removeCalls() {
return List.copyOf(removeCalls);
}
public List<OverlayCall> overlayCalls() {
return List.copyOf(overlayCalls);
}
public List<RepoRootCall> repoRootCalls() {
return List.copyOf(repoRootCalls);
}
public AddCall lastAdd() {
return addCalls.isEmpty() ? null : addCalls.getLast();
}
public RemoveCall lastRemove() {
return removeCalls.isEmpty() ? null : removeCalls.getLast();
}
public OverlayCall lastOverlay() {
return overlayCalls.isEmpty() ? null : overlayCalls.getLast();
}
}
@@ -0,0 +1,335 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-301 / CB-303 acceptance tests for the authoritative session registry, one-shot lifecycle FSM,
* and configurable lifecycle limits (idle TTL, context cap, drain).
* No live herdr — everything runs against the same {@link FakeHerdr} the rest of the project uses.
*/
class SessionManagerTest {
private SessionManager sessionManager(FakeHerdr herdr) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
return new SessionManager(workers);
}
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock) {
return sessionManager(herdr, clock, 0);
}
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
return new SessionManager(workers, new GitWorktrees(), clock, contextCap);
}
@Test
void acquireRegistersSpawningSessionWithDistinctPaneId() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession a = sessions.acquire("ltms-local", "/work/a", "/caller/a", "term_primary");
WorkerSession b = sessions.acquire("ltms-local", "/work/b", "/caller/b", "term_primary");
assertEquals(WorkerSession.State.SPAWNING, a.state(), "fresh session starts spawning");
assertEquals("ltms-local", a.profile());
assertEquals("/work/a", a.cwd(), "explicit requested cwd is recorded");
assertEquals("term_primary", a.ownerTerminal());
assertTrue(a.spawnedAtNanos() > 0);
assertNotNull(a.paneId());
assertNotNull(a.terminalId());
assertNotEquals(a.paneId(), b.paneId(), "no pane reuse");
assertNotEquals(a.terminalId(), b.terminalId(), "no terminal reuse");
assertEquals(2, sessions.roster().size(), "both sessions are registered");
}
@Test
void presenceMovesSpawningToReadyAndDeliveredTurnMovesToDone() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
assertEquals(WorkerSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"MCP presence moves SPAWNING → READY");
assertTrue(sessions.asPresence().isPresent(terminal), "presence is also recorded");
sessions.onDelivered(terminal);
assertEquals(WorkerSession.State.BUSY, sessions.get(session.paneId()).orElseThrow().state(),
"delivery moves READY → BUSY");
sessions.onTurnComplete(terminal);
assertEquals(WorkerSession.State.DONE, sessions.get(session.paneId()).orElseThrow().state(),
"turn completion moves BUSY → DONE");
}
@Test
void releaseTearsDownWorkerAndRemovesFromRosterAndIsIdempotent() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", null);
String paneId = session.paneId();
sessions.release(paneId);
assertTrue(herdr.called("pane.close"), "release tears the worker pane down");
assertTrue(sessions.get(paneId).isEmpty(), "released session is no longer retrievable");
assertTrue(sessions.roster().isEmpty(), "released session is no longer in the roster");
assertDoesNotThrow(() -> sessions.release(paneId), "a second release is harmless");
}
@Test
void onTurnFailedMovesSessionToFailed() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnFailed(terminal);
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(WorkerSession.State.FAILED, updated.state(), "turn failure moves to FAILED");
assertTrue(sessions.roster().contains(updated), "FAILED is still in acquired-minus-released roster");
}
@Test
void recycleProducesNewPaneIdAndOldOneIsGone() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession oldSession = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String oldPane = oldSession.paneId();
String oldTerminal = oldSession.terminalId();
WorkerSession fresh = sessions.recycle(oldPane);
assertNotEquals(oldPane, fresh.paneId(), "recycle yields a new pane id");
assertNotEquals(oldTerminal, fresh.terminalId(), "recycle yields a new terminal id");
assertEquals(oldSession.profile(), fresh.profile(), "profile is preserved");
assertEquals(oldSession.cwd(), fresh.cwd(), "cwd is preserved");
assertEquals(oldSession.ownerTerminal(), fresh.ownerTerminal(), "owner is preserved");
assertTrue(sessions.get(oldPane).isEmpty(), "old pane is deregistered");
assertEquals(1, sessions.roster().size(), "only the fresh session remains");
assertEquals(fresh.paneId(), sessions.roster().getFirst().paneId());
long paneCloseCount = herdr.calls.stream()
.filter(c -> "pane.close".equals(c.method()))
.filter(c -> oldPane.equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
assertEquals(1, paneCloseCount, "the old worker was torn down");
}
@Test
void rosterReflectsAcquiredMinusReleased() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession a = sessions.acquire("ltms-local", "/a", "/caller", "ownerA");
WorkerSession b = sessions.acquire("ltms-local", "/b", "/caller", "ownerB");
assertEquals(2, sessions.roster().size());
assertTrue(sessions.roster().stream().anyMatch(s -> s.paneId().equals(a.paneId())));
assertTrue(sessions.roster().stream().anyMatch(s -> s.paneId().equals(b.paneId())));
sessions.release(a.paneId());
assertEquals(1, sessions.roster().size());
assertEquals(b.paneId(), sessions.roster().getFirst().paneId());
}
// --- CB-303 lifecycle limits ----------------------------------------------------
@Test
void reapIdleDoesNothingWhenNoSessions() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L);
assertEquals(0, sessions.reapIdle(10));
assertTrue(sessions.roster().isEmpty());
}
@Test
void readySessionPastIdleTtlIsReaped() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
clock[0] = 11;
assertEquals(1, sessions.reapIdle(10), "READY session past TTL is reaped");
assertTrue(sessions.get(session.paneId()).isEmpty(), "reaped session is removed from registry");
assertTrue(herdr.called("pane.close"), "reaped session tears the pane down");
}
@Test
void readySessionWithinIdleTtlSurvives() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
clock[0] = 5;
assertEquals(0, sessions.reapIdle(10), "READY session within TTL is not reaped");
assertEquals(WorkerSession.State.READY,
sessions.get(session.paneId()).orElseThrow().state(),
"READY session survives");
}
@Test
void busySessionPastIdleTtlIsNotReaped() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
clock[0] = 100;
assertEquals(0, sessions.reapIdle(10), "BUSY session past TTL is never reaped");
assertEquals(WorkerSession.State.BUSY,
sessions.get(session.paneId()).orElseThrow().state(),
"BUSY session remains");
}
@Test
void doneSessionPastIdleTtlIsReaped() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
clock[0] = 21;
assertEquals(1, sessions.reapIdle(20), "DONE session past TTL is reaped");
assertTrue(sessions.get(session.paneId()).isEmpty(), "DONE session is removed");
}
@Test
void reapIdleReturnsCorrectCountAndSkipsBusy() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "owner1");
WorkerSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "owner2");
sessions.asPresence().markPresent(ready.terminalId());
sessions.asPresence().markPresent(busy.terminalId());
sessions.onDelivered(busy.terminalId());
clock[0] = 50;
assertEquals(1, sessions.reapIdle(30), "only READY past TTL is reaped");
assertTrue(sessions.get(ready.paneId()).isEmpty(), "READY session is gone");
assertEquals(WorkerSession.State.BUSY,
sessions.get(busy.paneId()).orElseThrow().state(),
"BUSY session is still registered");
}
@Test
void contextCapDisabledSessionSurvivesMultipleTurns() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 0);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(WorkerSession.State.DONE, updated.state(), "session finishes second turn");
assertEquals(2, updated.turnCount(), "turn count tracks both deliveries");
long releaseCloseCount = paneCloseCallsFor(herdr, session.paneId());
assertEquals(0, releaseCloseCount, "cap disabled — no forced release of the worker pane");
}
@Test
void contextCapTwoReleasesAfterSecondComplete() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 2);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
assertEquals(WorkerSession.State.DONE,
sessions.get(session.paneId()).orElseThrow().state(),
"first turn completes without release");
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
assertTrue(sessions.get(session.paneId()).isEmpty(), "session released after cap reached");
assertTrue(sessions.roster().isEmpty(), "released session leaves roster");
assertEquals(1, paneCloseCallsFor(herdr, session.paneId()),
"forced release tears the worker pane down exactly once");
}
@Test
void drainAllReleasesBusyAndReadySessionsAndWaitsForBusy() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "ownerR");
WorkerSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "ownerB");
sessions.asPresence().markPresent(ready.terminalId());
sessions.asPresence().markPresent(busy.terminalId());
sessions.onDelivered(busy.terminalId());
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
assertTrue(sessions.roster().isEmpty(), "drain clears the roster");
assertTrue(sessions.get(ready.paneId()).isEmpty(), "ready session is released");
assertTrue(sessions.get(busy.paneId()).isEmpty(), "busy session is released after timeout");
assertEquals(1, paneCloseCallsFor(herdr, ready.paneId()),
"ready worker pane is torn down");
assertEquals(1, paneCloseCallsFor(herdr, busy.paneId()),
"busy worker pane is torn down");
}
private static long paneCloseCallsFor(FakeHerdr herdr, String paneId) {
return herdr.calls.stream()
.filter(c -> "pane.close".equals(c.method()))
.filter(c -> paneId.equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
}
}
@@ -0,0 +1,169 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-301-ext acceptance tests for worktree provisioning and config-parity overlay.
* No live git — every Worktrees call is handled by {@link FakeWorktrees} and every herdr
* call by {@link FakeHerdr}, matching the project's fake-based test style.
*/
class WorktreeSessionManagerTest {
private static ClaudeCodeLauncher workerService(FakeHerdr herdr) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
}
private static String startCwd(FakeHerdr herdr) {
@SuppressWarnings("unchecked")
Map<String, Object> start = (Map<String, Object>) herdr.lastCall("agent.start").params();
Object cwd = start.get("cwd");
return cwd == null ? null : cwd.toString();
}
@Test
void sharedTreeAcquireMakesNoWorktreesCallsAndRecordsNullWorktree() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary");
assertTrue(worktrees.addCalls().isEmpty(), "shared-tree acquire never adds a worktree");
assertTrue(worktrees.repoRootCalls().isEmpty(), "shared-tree acquire never resolves a repo root");
assertTrue(worktrees.overlayCalls().isEmpty(), "shared-tree acquire never overlays parity");
assertNull(s.worktree(), "shared-tree session has no worktree");
assertNull(s.branch(), "shared-tree session has no branch");
assertEquals("/caller/proj", s.cwd(), "shared-tree cwd is the caller's cwd");
assertEquals("/caller/proj", startCwd(herdr), "spawn receives the caller's cwd");
}
@Test
void worktreeAcquireProvisionsAndRecordsPathAndBranch() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-999", null));
assertEquals(1, worktrees.addCalls().size(), "one worktree was added");
FakeWorktrees.AddCall add = worktrees.lastAdd();
assertNotNull(add);
assertEquals("/repo", add.repoRoot());
assertTrue(add.branch().startsWith("worker/cb-999-"), "branch is worker/<slug>-<nonce>: " + add.branch());
assertNull(add.baseRef(), "null baseRef is passed through (HEAD default)");
String expectedPath = "/wt/" + add.branch().replace('/', '_');
assertEquals(expectedPath, s.worktree(), "session records the returned worktree path");
assertEquals(add.branch(), s.branch(), "session records the branch");
assertEquals(expectedPath, startCwd(herdr), "spawn receives the worktree path as cwd");
assertEquals(expectedPath, s.cwd(), "session cwd is the worktree path");
}
@Test
void worktreeAcquireRunsParityOverlayWithProfileDefaults() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.track(".mcp.json")
.exists(".claude/settings.local.json");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-888", null));
assertEquals(1, worktrees.overlayCalls().size());
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
assertNotNull(overlay);
assertEquals("/repo", overlay.repoRoot());
assertEquals(List.of(".mcp.json", ".claude/settings.local.json", ".env", ".envrc"),
overlay.requested(), "default parity overlay is used when unset");
assertEquals(List.of(".mcp.json", ".claude/settings.local.json"), overlay.copied(),
"existing paths are copied; missing paths are skipped");
assertEquals(List.of(".mcp.json"), overlay.skipWorktree(),
"tracked copied paths are --skip-worktree'd");
}
@Test
void releaseRemovesWorktreeButDoesNotDeleteBranch() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-666", null));
String paneId = s.paneId();
sessions.release(paneId);
assertTrue(herdr.called("pane.close"), "release still tears the worker pane down");
assertEquals(1, worktrees.removeCalls().size(), "worktree session triggers one remove");
FakeWorktrees.RemoveCall remove = worktrees.lastRemove();
assertNotNull(remove);
assertEquals("/repo", remove.repoRoot());
assertEquals(s.worktree(), remove.worktreePath());
// The fake records no branch-delete calls because Worktrees.remove only removes the checkout.
assertTrue(sessions.get(paneId).isEmpty(), "released session is no longer retrievable");
}
@Test
void sharedTreeReleaseMakesNoWorktreesCalls() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", null);
sessions.release(s.paneId());
assertTrue(herdr.called("pane.close"), "release tears the worker pane down");
assertTrue(worktrees.removeCalls().isEmpty(), "shared-tree release never removes a worktree");
}
@Test
void failedWorktreeAddUnwindsWithoutRegisteringSessionOrSpawning() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().failAdd("worktree add failed");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
assertThrows(WorktreeException.class, () ->
sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-555", null)));
assertEquals(0, sessions.size(), "failed acquire leaves no registry entry");
assertFalse(herdr.called("agent.start"), "spawn is never reached when add fails");
assertTrue(worktrees.removeCalls().isEmpty(), "no worktree was added, so none is removed");
}
@Test
void twoWorktreeAcquiresYieldDistinctBranchesAndPaths() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession a = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-444", null));
WorkerSession b = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-444", null));
assertNotEquals(a.branch(), b.branch(), "branches are distinct");
assertNotEquals(a.worktree(), b.worktree(), "paths are distinct");
assertEquals(2, worktrees.addCalls().size());
assertEquals(2, sessions.roster().size());
}
}
@@ -0,0 +1,321 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.*;
/** The step-4 launch-flag injection: the bridge MCP + reply charter are appended to the argv. */
class ClaudeCodeLauncherTest {
private ClaudeCodeLauncher service(FakeHerdr herdr, List<String> argv, String mcpUrl) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
argv, "tab", "bridged-workers", "worker: {profile} #{n}", mcpUrl, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
}
@SuppressWarnings("unchecked")
private List<String> spawnedArgv(FakeHerdr herdr) {
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("argv");
}
@Test
void appendsBridgeMcpAndReplyCharterWhenMcpUrlSet() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("ccs", "ltms-local"), "http://127.0.0.1:8765/mcp").spawn();
List<String> argv = spawnedArgv(herdr);
assertEquals(List.of("ccs", "ltms-local"), argv.subList(0, 2), "base command preserved first");
assertTrue(argv.contains("--mcp-config"));
assertTrue(argv.stream().anyMatch(a -> a.contains("\"bridge\"") && a.contains("http://127.0.0.1:8765/mcp")),
"inline bridge MCP config present");
assertTrue(argv.contains("--append-system-prompt"));
assertTrue(argv.stream().anyMatch(a -> a.contains("bridge_reply")), "reply charter present");
}
@Test
void noBridgeFlagsWhenMcpUrlAbsent() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("bash", "-c", "sleep 1"), null).spawn();
assertEquals(List.of("bash", "-c", "sleep 1"), spawnedArgv(herdr), "argv untouched without mcpUrl");
}
private ClaudeCodeLauncher multiProfile(FakeHerdr herdr) {
BridgedConfig.Worker gx10 = new BridgedConfig.Worker("gx10", "http://gx10.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "gx10"), "tab", "bridged-workers", "w #{n}", null, null, null);
BridgedConfig.Worker ollama = new BridgedConfig.Worker("ollama", "http://ollama.ltms.dev", null,
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "ollama"), "tab", "bridged-workers", "w #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx10.gw", "ollama.ltms.dev")),
Map.of("gx10", gx10, "ollama", ollama), "gx10", _ -> "tok");
}
@Test
@SuppressWarnings("unchecked")
void spawnPicksTheNamedProfilesBaseUrlAndArgv() {
FakeHerdr herdr = new FakeHerdr();
multiProfile(herdr).spawn("ollama");
Map<String, Object> start = (Map<String, Object>) herdr.lastCall("agent.start").params();
Map<String, String> env = (Map<String, String>) start.get("env");
assertEquals("http://ollama.ltms.dev", env.get("ANTHROPIC_BASE_URL"), "the named profile's base_url");
assertEquals(List.of("ccs", "ollama"), start.get("argv"), "the named profile's launch command");
}
@Test
void spawnRejectsAnUnknownProfile() {
try (FakeHerdr herdr = new FakeHerdr()) {
assertThrows(IllegalArgumentException.class, () -> multiProfile(herdr).spawn("nope"));
}
}
@SuppressWarnings("unchecked")
private static String startCwd(FakeHerdr herdr) {
// The worker's cwd is set on agent.start (an agent pane does not inherit the tab's cwd).
Object v = ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("cwd");
return v == null ? null : v.toString();
}
@Test
void requestedCwdRootsTheWorker() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("ccs", "ltms-local"), null).spawn("ltms-local", "/work/proj", "/caller/home");
assertEquals("/work/proj", startCwd(herdr), "an explicit spawn cwd wins over everything");
}
@Test
void profileConfigCwdBeatsTheCallerCwd() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker("ltms-local", "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"w #{n}", null, "/pinned/dir", null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("ltms-local", cfg), "ltms-local", _ -> null);
svc.spawn("ltms-local", null, "/caller/home");
assertEquals("/pinned/dir", startCwd(herdr), "a profile-pinned cwd overrides the caller's");
}
@Test
void inheritsTheCallerCwdWhenNothingElseIsSet() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("ccs", "ltms-local"), null).spawn("ltms-local", null, "/primary/project");
assertEquals("/primary/project", startCwd(herdr), "no explicit/config cwd → inherit the primary's");
}
// --- CB-302 git-forge token injection (worker checkpoint grant) ------------
@SuppressWarnings("unchecked")
private static Map<String, String> startEnv(FakeHerdr herdr) {
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("env");
}
@Test
void injectsForgeTokenAndHostWhenProfileGrantsIt() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"impl", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "impl"), "tab", "bridged-workers", "w #{n}", null, null, null,
"GITEA_ACCESS_TOKEN", null); // parityOverlay null; gitHostEnv null → defaults to GITEA_HOST
Function<String, String> host = name -> switch (name) {
case "GITEA_ACCESS_TOKEN" -> "gt-secret";
case "GITEA_HOST" -> "git.ltms.dev";
default -> null;
};
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("impl", cfg), "impl", host).spawn();
Map<String, String> env = startEnv(herdr);
assertEquals("gt-secret", env.get("GITEA_TOKEN"), "the forge token is injected for a granting profile");
assertEquals("git.ltms.dev", env.get("GITEA_HOST"), "the paired forge host rides along with the token");
}
@Test
void noForgeTokenWhenProfileDoesNotGrantIt() {
FakeHerdr herdr = new FakeHerdr();
// gitTokenEnv unset (12-arg ctor); the env would resolve a token if asked, proving the gate
// is the profile config, not a missing env var.
BridgedConfig.Worker cfg = new BridgedConfig.Worker("ltms-local", "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("ccs"), "tab", "bridged-workers", "w #{n}",
null, null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("ltms-local", cfg), "ltms-local",
_ -> "would-be-secret").spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("GITEA_TOKEN"), "no forge token when the profile does not opt in");
assertNull(env.get("GITEA_HOST"), "no forge host without a granted token");
}
// --- CB-117 orphan reap: the pure predicate --------------------------------
@Test
void isForeignWorkerMatchesOurSchemeWithANonSelfNonce() {
assertTrue(ClaudeCodeLauncher.isForeignWorker("claude-ollama-be09c2-2", "aaaaaa"),
"a bridge worker name with a different nonce is a prior daemon's orphan");
assertTrue(ClaudeCodeLauncher.isForeignWorker("claude-gx10-4127af-11", "aaaaaa"),
"profile and multi-digit seq are still parsed; foreign nonce ⇒ reap");
}
@Test
void isForeignWorkerSparesOurOwnLiveWorkersAndNonWorkers() {
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude-ollama-abcdef-3", "abcdef"),
"a worker with THIS process's nonce is ours and live — never reap it");
assertFalse(ClaudeCodeLauncher.isForeignWorker(null, "abcdef"), "an unnamed agent is not a worker");
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude", "abcdef"), "a bare kind name is not a worker");
assertFalse(ClaudeCodeLauncher.isForeignWorker("my-repl", "abcdef"), "a user's own label is not a worker");
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude-ollama-XYZ123-2", "abcdef"),
"a non-hex nonce does not match our scheme");
}
// --- CB-117 orphan reap: the wiring through stop() -------------------------
private static long paneCloseCount(FakeHerdr herdr, String paneId) {
return herdr.calls.stream()
.filter(c -> c.method().equals("pane.close"))
.filter(c -> paneId.equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
}
@Test
void reapsAForeignOrphanButSparesOurOwnWorkerAndUserSessions() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = multiProfile(herdr);
herdr.withAgent("claude-ollama-be09c2-2", "term_orphan", "wQ:pF", "wQ:t8") // prior daemon's leak
.withAgent("claude-gx10-" + svc.nameNonce() + "-1", "term_mine", "wQ:pMine", "wQ:tMine"); // ours, live
// (the fake's default unnamed term_a stands in for a user's own Claude session)
int reaped = svc.reapOrphanWorkers();
assertEquals(1, reaped, "exactly the one foreign-nonce orphan is reaped");
assertEquals(1, paneCloseCount(herdr, "wQ:pF"), "the orphan's pane is closed");
assertEquals(0, paneCloseCount(herdr, "wQ:pMine"), "our own live worker's pane is left running");
assertEquals(0, paneCloseCount(herdr, "w2:p7"), "a user's own session is never touched");
assertTrue(herdr.called("tab.close"), "the orphan's now-empty dedicated tab is closed too");
}
@Test
void reapCountsAnAlreadyGoneOrphanAsReaped() {
FakeHerdr herdr = new FakeHerdr().paneCloseFailsWith("pane_not_found");
ClaudeCodeLauncher svc = multiProfile(herdr);
herdr.withAgent("claude-ollama-0d856d-3", "term_gone", "wQ:pS", "wQ:tD");
assertEquals(1, svc.reapOrphanWorkers(),
"a pane that vanished between list and close is a successful reap, not a failure");
}
@Test
void reapIsSkippedWhenHerdrCannotBeListed() {
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.list throws
assertEquals(0, multiProfile(herdr).reapOrphanWorkers(), "a listing failure reaps nothing and does not throw");
}
// --- PeerHandle indirection ----------------------------------------------------------------
@Test
void spawnReturnsPeerHandleWithIdEqualToPaneId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn must return a non-null handle");
assertEquals("w9:pW_1", handle.id(), "handle.id() must equal the agent's paneId");
}
@Test
void spawnReturnsPeerHandleWithCorrectTerminalId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, "/caller"));
assertEquals("term_new_1", handle.terminalId(), "handle.terminalId() must equal the agent's terminalId");
}
@Test
void capabilitiesIncludeMidTurnAskWorktreeOrphanReap() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
Set<Capability> caps = svc.capabilities();
assertTrue(caps.contains(Capability.MID_TURN_ASK), "every Claude Code peer supports mid-turn ask");
assertTrue(caps.contains(Capability.WORKTREE), "every CLI peer supports worktree cwd");
assertTrue(caps.contains(Capability.ORPHAN_REAP), "every herdr launcher supports orphan reap");
}
@Test
void capabilitiesIncludeSelfPrWhenProfileHasGitToken() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"impl", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "impl"), "tab", "bridged-workers", "w #{n}", null, null, null,
"GITEA_ACCESS_TOKEN", null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("impl", cfg), "impl",
_ -> "tok");
assertTrue(svc.capabilities().contains(Capability.SELF_PR),
"a profile with a git token grants SELF_PR");
}
@Test
void capabilitiesExcludeSelfPrWhenNoGitToken() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
assertFalse(svc.capabilities().contains(Capability.SELF_PR),
"no git token profile → no SELF_PR capability");
}
@Test
void effectiveCwdViaSpawnRequestMatchesExistingResolution() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
String cwd = svc.effectiveCwd(new SpawnRequest("ltms-local", "/work/proj", "/caller/home"));
assertEquals("/work/proj", cwd, "effectiveCwd via SpawnRequest must match the three-arg resolution");
}
@Test
void profilesViaPeerLauncherMatchesExistingApi() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = multiProfile(herdr);
assertEquals(Set.of("gx10", "ollama"), svc.profiles(), "profiles() via PeerLauncher must match");
}
@Test
void defaultProfileViaPeerLauncherMatches() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = multiProfile(herdr);
assertEquals("gx10", svc.defaultProfile(), "defaultProfile() via PeerLauncher must match");
}
@Test
void stopViaPeerLauncherTearsDownByHandleId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
svc.stop(handle.id());
assertTrue(herdr.called("pane.close"), "stop via handle.id() must close the pane");
}
}
+131
View File
@@ -0,0 +1,131 @@
# CB-301 — Session Manager (one-shot, no reuse)
**Status:** design spec for review → delegate implementation.
**Grounded in:** `WorkerService`, `Injector`/`StatusPoller`/`TurnListener`, `MessageService`,
`BridgeMcp`, `BridgedApp` (see [wiki 9. Implementation](../wiki/9-Implementation.md)).
## Problem
`WorkerService` is **stateless about what it spawned**. Its own Javadoc says it plainly:
> "there is no registry; `list()` only asks herdr." — `WorkerService.reapOrphanWorkers` (line 288)
Consequences today:
- The daemon cannot answer "which workers did *I* spawn, in what lifecycle state, owned by whom,
since when?" without shelling to herdr for a raw agent list (no state, no ownership, no age).
- Cleanup of a worker that outlived its owning process depends entirely on the boot-time
name-nonce **reaper** (CB-117) — there is no live, authoritative roster during a run.
- `bridge_list` (CB-304) can only surface herdr's view, not a bridge-owned roster.
- There is no seam for per-session policy (checkpoint on teardown → CB-302; idle_ttl /
context_cap / drain → CB-303).
## Goal & non-goals
**Goal.** Introduce a `SessionManager` that owns an authoritative in-daemon registry of the worker
sessions this daemon process spawned, tracks each one's lifecycle state, and tears each down
deterministically. It becomes the single source of truth for the roster and the seam CB-302/303/304
build on.
**Non-goals (explicit — reuse policy chosen: one-shot, no reuse).**
- **No pooling / no reuse.** Every delegated task gets a fresh worker; a finished worker is torn
down, never handed to a later task. No "warm idle" pool, no `role@profile` keying.
- **No auto-teardown *timing*.** *When* a one-shot worker is released (immediately on turn
completion vs after an idle grace) is CB-303. CB-301 provides the **mechanism** (`release`) and
the registry; CB-303 sets the policy.
- **No checkpoint content.** Writing `STATE.md` + commit on teardown is CB-302; CB-301 only exposes
the release hook it will attach to.
"Recycle" under no-reuse is simply **release + fresh acquire** — a helper, not a pool operation.
## Design
`SessionManager` **wraps** `WorkerService` (does not replace it). `WorkerService` keeps doing the
subscription-guarded spawn/teardown mechanics; `SessionManager` adds the registry, lifecycle, and
ownership on top.
**Package:** new `dev.ltms.bridged.session` — keeps the registry/lifecycle concern separate from
the `worker` spawn mechanics. Holds `SessionManager` + `WorkerSession`.
**`recycle` is IN SCOPE for CB-301** (decided): implement `recycle(paneId, …)` = `release` the old
session then `acquire` a fresh one, asserting a new distinct paneId (the no-reuse invariant). It is
a thin convenience over the two primitives, shipped now so the no-reuse teardown+respawn path is
covered by a test from day one.
### `WorkerSession` (record or small mutable holder)
| Field | Source | Notes |
|---|---|---|
| `paneId` | `Agent.paneId()` | registry key |
| `terminalId` | `Agent.terminalId()` | for status/identity joins |
| `profile` | spawn arg | which profile spawned it |
| `cwd` | resolved cwd | the worker's working dir |
| `ownerTerminal` | caller identity (nullable) | the primary/turn that requested it; `null` = daemon/anon |
| `spawnedAtNanos` | `System.nanoTime()` | age basis for CB-303 (monotonic; no wall clock in tests) |
| `state` | lifecycle FSM | see below |
State is held in a `ConcurrentHashMap<String /*paneId*/, WorkerSession>`.
### Lifecycle state machine (one-shot)
```
SPAWNING --ready(MCP present)--> READY
READY --onDelivered--> BUSY
BUSY --onTurnComplete--> DONE
BUSY --onTurnFailed--> FAILED
READY|DONE|FAILED --release()--> RELEASED (deregistered)
SPAWNING|READY|BUSY|DONE --vanished/drop--> FAILED
```
- Transitions are driven by hooks the manager already has access to:
`WorkerPresence.markPresent` → `READY`; `TurnListener.onDelivered/onTurnComplete/onTurnFailed`
(the manager implements or decorates `TurnListener`) → `BUSY`/`DONE`/`FAILED`.
- `RELEASED` sessions are removed from the registry (teardown is terminal).
- Any state → `FAILED` on drop (worker vanished / injector `drop`), mirroring `Injector`.
### API
```java
final class SessionManager {
WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal);
void release(String paneId); // deterministic teardown + deregister
WorkerSession recycle(String paneId, ...); // release + acquire (no-reuse convenience)
Optional<WorkerSession> get(String paneId);
List<WorkerSession> roster(); // bridge-owned view (CB-304 consumes this)
// lifecycle hooks (package-private): onReady/onDelivered/onComplete/onFailed(target)
}
```
- `acquire` = `workerService.spawn(profile, requestedCwd, callerCwd)` → register `SPAWNING`.
- `release` = `workerService.stop(paneId)` → deregister. Idempotent (already-gone tolerated, matching
`WorkerService.stop`).
- `roster` joins the registry with live herdr status for a truthful "roster + live" (CB-304).
### Integration points
- **`Bridged.main`** — construct `SessionManager(workerService, ...)`; wire it as/decorating the
`TurnListener` alongside `CompletionResolver` so it sees turn boundaries, and give it the
`WorkerPresence` signal for `READY`.
- **`BridgeMcp.spawn` / `BridgedApp.spawnWorker`** — route spawn through `SessionManager.acquire`
(carry `callerTerminal` as `ownerTerminal`). **`bridge_stop` / `DELETE /workers/{paneId}`** →
`SessionManager.release`.
- **`bridge_list` / `GET /sessions` (CB-304 later)** — read `SessionManager.roster()`.
- **`MessageService`** — no change required for one-shot; a later CB-303 auto-release hook can call
`release` from `onTurnComplete` under policy.
## Acceptance (tests, no live herdr — fakes as elsewhere)
1. `acquire` registers a `SPAWNING` session with the right owner/profile/cwd; a second `acquire`
yields a **distinct** paneId and a **distinct** session (no reuse).
2. Presence signal moves `SPAWNING → READY`; a delivered turn moves `READY → BUSY → DONE`.
3. `release` tears the worker down via `WorkerService.stop` and removes it from `roster()`;
a second `release` on the same paneId is a harmless no-op.
4. `onTurnFailed` / drop moves the session to `FAILED` and it is absent from the live roster.
5. `recycle` produces a new paneId and the old one is gone (no-reuse invariant).
6. `roster()` reflects exactly the sessions acquired-minus-released, joined with live status.
## Seams left open (deliberately)
- **CB-302** — attach a checkpoint step (`STATE.md` + commit) to the `release` path.
- **CB-303** — a policy loop over `roster()` using `spawnedAtNanos`/state to auto-`release` on
`idle_ttl`, or drain on `context_cap`.
- **CB-304** — `bridge_list` reads `roster()` for a bridge-owned roster + live join.
+179
View File
@@ -0,0 +1,179 @@
# CB-301-ext — Worktree provisioning + config-parity overlay
**Status:** design spec for review → delegate implementation.
**Extends:** [CB-301 Session Manager](CB-301-Session-Manager.md) (shipped, commit `54d907c`).
**Realizes:** the config-parity requirement in [Worker Git Workflow](Worker-Git-Workflow.md).
**Grounded in:** `SessionManager`, `WorkerService.spawn/effectiveCwd`, `BridgedConfig.Worker`,
`inject/…LsofPeerPidLookup` (the `ProcessBuilder` exec pattern).
## Problem
CB-301 gives each worker a session record but every worker still runs in the **primary's own
working tree** (`cwd = callerCwd`). One worker at a time is safe; two **parallel implementers** would
stomp each other. We need each implementer session to get an **isolated git worktree on its own
branch** — *without* degrading the worker: a bare worktree checks out **tracked files only**, so it
silently drops the untracked/local config (`.claude/settings.local.json`, the locally-modified
`.mcp.json`, `.env`) that makes a session a full peer of the primary. **CB-301-ext provisions the
worktree AND hydrates it to config parity**, so a worker differs from the primary only in the LLM
provider.
## Decisions (locked)
1. **Opt-in, not default.** A worktree is provisioned **only** when the caller requests one. Absent a
request, `acquire` behaves exactly as it does today (shared primary tree) — auditors, smoke tests,
and conversational workers are unaffected. **Backward compatibility is a hard requirement.**
2. **Copy-overlay + `--skip-worktree`, not symlink.** Each parity file is **copied** primary→worktree
(isolation-friendly, no symlink type-change noise on tracked files). For a *tracked* overlay file
(`.mcp.json`) the worktree copy is then marked `git update-index --skip-worktree`, so the worker's
commits can **never** include the parity overlay. Ignored files (`settings.local.json`) stay
ignored in the worktree (shared `info/exclude`), so no marking is needed.
3. **Branch persists; worktree is disposable.** `release` runs `git worktree remove --force` (the
working dir is throwaway) but **never deletes the branch** — the branch holds the worker's commits
and its PR (CB-302). Teardown of the checkout ≠ teardown of the work.
4. **Git behind a seam.** SessionManager depends on a `Worktrees` interface (production impl shells
`git` via `ProcessBuilder`; tests use a fake). No live `git` in unit tests — mirrors the
`WorkerService`/`FakeHerdr` seam.
## Design
### `WorktreeRequest` (new, nullable = "no worktree")
```java
package dev.ltms.bridged.session;
/** Ask acquire() to provision an isolated worktree. null ⇒ run in the shared primary tree. */
public record WorktreeRequest(String ticketSlug, String baseRef) {
// ticketSlug seeds the branch name; baseRef null/blank ⇒ current HEAD of the repo.
}
```
### `WorkerSession` — two nullable fields added
| Field | Notes |
|---|---|
| `worktree` | absolute path of the provisioned worktree; `null` ⇒ shared tree |
| `branch` | the worker's branch (`worker/<slug>-<nonce>`); `null` ⇒ shared tree |
Add to the record + `withState`. A `null` worktree keeps every existing test and the shared-tree path
untouched.
### `Worktrees` seam (new)
```java
package dev.ltms.bridged.session;
public interface Worktrees {
/** git -C <repoRoot> worktree add <path> -b <branch> <baseRef|HEAD>. Returns the worktree path. */
String add(String repoRoot, String branch, String baseRef);
/** git -C <repoRoot> worktree remove --force <path>. Idempotent (already-gone tolerated). */
void remove(String repoRoot, String worktreePath);
/** Copy each existing overlay path repoRoot→worktree; mark tracked ones --skip-worktree. */
void overlayParity(String repoRoot, String worktreePath, List<String> overlay);
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
String repoRoot(String cwd);
}
```
- **Production impl** `GitWorktrees implements Worktrees` — `ProcessBuilder` per path, `redirectErrorStream(true)`, non-zero exit → a `WorktreeException`. Worktree location = `<worktreeRoot>/<nonce>` where `worktreeRoot` is a daemon setting (default: sibling `../.bridged-worktrees` of the repo root — **outside** the repo, never nested).
- `overlayParity` per file: skip if absent in `repoRoot`; else copy into the worktree; if `git -C <wt> ls-files --error-unmatch <path>` succeeds (tracked), run `git -C <wt> update-index --skip-worktree <path>`.
### `acquire` — extended, old signature preserved
```java
// existing (unchanged): shared tree
WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal);
// new overload: provision a worktree when wt != null
WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal,
WorktreeRequest wt);
```
When `wt != null`:
1. `repoRoot = worktrees.repoRoot(firstNonBlank(requestedCwd, callerCwd))`.
2. `branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce`.
3. `path = worktrees.add(repoRoot, branch, wt.baseRef())`.
4. `worktrees.overlayParity(repoRoot, path, cfg.parityOverlay())`.
5. `spawn(profile, path, callerCwd)` — **the worktree path becomes the worker's cwd** (highest
precedence in `WorkerService.resolveCwd`).
6. register the session with `worktree=path, branch=branch`.
7. **On any failure in 1–5, unwind**: if the worktree was added, `remove` it; do not leave a dangling
registry entry. (Guard/spawn already throw before herdr on a bad base_url — unchanged.)
### `release` — remove the worktree, keep the branch
```java
public void release(String paneId) {
WorkerSession s = registry.remove(paneId);
workerService.stop(paneId); // existing
if (s != null && s.worktree() != null) {
worktrees.remove(worktrees.repoRoot(s.cwd()), s.worktree()); // branch is NOT deleted
}
}
```
### Config — `BridgedConfig.Worker.parityOverlay` + a `worktreeRoot`
- Add `List<String> parityOverlay` to the `Worker` record (12th field). Compact-constructor default
when null/empty: `[".mcp.json", ".claude/settings.local.json", ".env", ".envrc"]` (missing paths are
silently skipped, so the default is safe across repos). Update `withProfile`.
- Add a top-level daemon setting `worktreeRoot` (String, nullable → `<repoParent>/.bridged-worktrees`).
- `@JsonIgnoreProperties(ignoreUnknown = true)` already set → additive, no parser breakage.
### Surface: MCP + REST
- `bridge_spawn` gains an optional `worktree` arg: `true`, or a ticket slug string. Truthy ⇒ build a
`WorktreeRequest(slug, null)` and call the 5-arg `acquire`.
- `POST /workers` gains `worktree` (+ optional `ticket`) in the body/query, same mapping.
- `workerView`/`view(WorkerSession)` include `worktree` and `branch` **when non-null** (omit for
shared-tree sessions, so existing response assertions for shared-tree spawns are unchanged).
## Flow
```mermaid
flowchart TD
A["acquire(profile, ..., WorktreeRequest?)"] --> B{"worktree<br/>requested?"}
B -->|"no (default)"| C["spawn(cwd = callerCwd)<br/>— shared tree, unchanged"]
B -->|yes| D["repoRoot = rev-parse --show-toplevel"]
D --> E["git worktree add path -b branch base"]
E --> F["overlayParity: copy local config in;<br/>--skip-worktree the tracked ones"]
F --> G["spawn(cwd = worktree path)"]
G --> H["register worktree + branch on the session"]
E -.->|"add/overlay/spawn fails"| X["unwind: remove worktree,<br/>no dangling registry entry"]:::warn
C --> R["worker is a full peer of the primary"]:::goal
H --> R
classDef warn fill:#b7791f,stroke:#7b341e,color:#ffffff;
classDef goal fill:#2b6cb0,stroke:#2a4365,color:#ffffff;
```
*Green = the invariant: whether shared-tree (parity for free) or worktree (parity via overlay), the
worker matches the primary. Amber = the failure-unwind path.*
## Acceptance (fake `Worktrees`, no live git)
1. **Backward-compat:** `acquire` with **no** `WorktreeRequest` makes **zero** `Worktrees` calls,
spawns with `cwd = callerCwd`, and records `worktree == null` / `branch == null`. (Every CB-301
test still passes.)
2. **Provision:** `acquire(..., new WorktreeRequest("cb-999", null))` calls `add(repoRoot,
"worker/cb-999-<nonce>", null)`, then `spawn` receives the returned worktree path as
`requestedCwd`; the session records that path + branch.
3. **Overlay:** `overlayParity` is invoked with the profile's `parityOverlay` (default list when
unset); the fake asserts tracked paths were `--skip-worktree`'d and missing paths skipped.
4. **Release removes worktree, keeps branch:** releasing a worktree session calls
`Worktrees.remove(repoRoot, path)` and performs **no** branch-delete; a shared-tree session's
release makes no `Worktrees` calls.
5. **Failure unwind:** a fake `add` that throws ⇒ `acquire` throws, the session is **not** registered,
and no worker is left running (spawn not reached / torn down).
6. **Distinct worktrees:** two worktree acquires yield **distinct** branches and paths (no collision).
## Constraints & exclusions (standing, non-negotiable)
- **Only a worker sets `ANTHROPIC_BASE_URL`.** Worktrees touch cwd + files only; env path is unchanged
— the guard still runs before any herdr call.
- **Worker cwd stays inside the primary repo** (the worktree is a checkout of it) — never `$HOME`.
- **`.mcp.json` and `wiki/` never enter a worker commit.** `.mcp.json` is overlaid for *reference* but
`--skip-worktree`'d so it can't be staged; `wiki/` is a submodule the worker must not touch. The
CB-302 commit step (and the implementer skill) exclude both.
- **The overlay list stays explicit + minimal** (trust: local secrets flow to an off-subscription
worker). No blanket tree copy.
## Seams left for later
- **CB-302** — the worker commit → push → PR checkpoint runs *inside* the worktree on its branch.
- **CB-304** — `roster()` rows surface `worktree`/`branch` for the fleet view.
+200
View File
@@ -0,0 +1,200 @@
# CB-401 — Peer Launcher SPI (Stage 4: pluggable peers)
**Status:** design note (feature branch `feature/peer-launcher-spi`)
**Stage:** 4 — makes the bridge scalable to heterogeneous peers (Claude Code, Codex, …) without
the core learning any one peer's environment.
## 1. Why
`claude-bridge` is a **communication bus between heterogeneous AI agents** — its stable surface is
the protocol (`bridge_spawn / send / poll / reply / ask / list / stop / status`), and that surface
should stay provider-neutral. Today the daemon can only materialize one kind of peer: an
off-subscription Claude Code CLI over herdr. Everything specific to *how that peer is set up*
(`ANTHROPIC_BASE_URL`, the subscription guard, `--mcp-config`/system-prompt flags, `claude-*`
naming, git-token injection) is baked directly into the core spawn path.
The goal of CB-401 is a seam — a **`PeerLauncher` SPI** — so that "how to bring a peer of kind X to
life" lives in a swappable adapter the bus *delegates to*, while the bus itself owns only transport,
session/turn lifecycle, and routing. This both unlocks a second peer kind (Codex, a human, another
Claude) and retroactively gives the CB-301-ext / CB-302 environment features a principled home (an
adapter) instead of sitting in core.
> **Non-goal for CB-401:** dynamic/external plugin loading (arbitrary jars). That is Stage C and
> carries a security model of its own (§7). CB-401 delivers a *first-party, in-tree, config-selected*
> SPI with exactly one implementation, proving the seam.
## 2. The boundary
```mermaid
flowchart LR
subgraph core["BRIDGE CORE — provider-neutral"]
proto["Protocol verbs<br/>spawn / send / poll / reply / ask / list / stop"]
sm["SessionManager<br/>FSM · registry · roster · lifecycle"]
msg["Message store / routing"]
life["Lifecycle limits<br/>idle_ttl · context_cap · drain"]
end
subgraph adapters["PEER ADAPTERS — env-specific"]
cc["ClaudeCodeLauncher<br/>(today's WorkerService)"]
cx["CodexLauncher<br/>(future, CB-402)"]
hu["HumanLauncher<br/>(future)"]
end
sm -->|"delegates spawn/release"| spi{{"PeerLauncher SPI"}}
spi --> cc
spi --> cx
spi --> hu
cc -.->|"herdr transport"| herdr["herdr (terminal multiplexer)"]
```
*Figure 1 — the core delegates peer materialization to a launcher chosen by profile; the core never
learns a peer's env.*
## 3. Coupling audit (as-built, main @ `0efb65c`)
Where Claude/herdr specifics actually live today:
| Concern | Location | Verdict |
|---|---|---|
| `ANTHROPIC_BASE_URL` / `ANTHROPIC_MODEL` / `CLAUDE_CONFIG_DIR` / `ANTHROPIC_AUTH_TOKEN` env | `WorkerService.spawn` | **→ adapter** |
| `SubscriptionGuard.assertWorker(baseUrl)` (billing boundary) | `WorkerService.spawn` → `guard` | **→ adapter** (it guards an `ANTHROPIC_*` concept) |
| `--mcp-config` + `--append-system-prompt REPLY_CHARTER` (Claude Code CLI flags) | `WorkerService.argvWithBridge` | **→ adapter** |
| `claude-<profile>-<nonce>-<seq>` naming, `WORKER_NAME` regex, orphan reap (CB-117) | `WorkerService` | **→ adapter** (naming is a herdr-label detail) |
| `GITEA_TOKEN` / `GITEA_HOST` injection (CB-302 checkpoint) | `WorkerService.spawn` | **→ adapter** + a **capability** (§6) |
| tab/pane placement, worker space, tab labels | `WorkerService.spawnInTab/spawnAsPane` via herdr `WorkspaceControl` | **→ adapter** (herdr transport detail) |
| `BridgedConfig.Worker` profile shape (`baseUrl`, `model`, `configDir`, …) | `config` | **mostly adapter-shaped** — see §5 |
| FSM, registry, roster, `reapIdle`/`drainAll`/`contextCap`, `rosterView` | `SessionManager` | **stays core** |
| turn/completion detection (`TurnListener`, `CompletionResolver`, `StatusPoller`, `WorkerPresence`) | `inject/` | **stays core**, but reads herdr terminal output → transport-coupled (§4b) |
| message store & routing | `msg/` | **stays core** |
| `Agent` / `paneId` handle | `herdr/` | **generalize** — see §4a |
**Finding:** the extraction is tractable because ~90% of the coupling is already funnelled through
one class (`WorkerService`). Renaming/adapting it to `ClaudeCodeLauncher implements PeerLauncher` and
having `SessionManager` depend on the interface is the bulk of Stage A.
## 4. Two friction points
### 4a. `paneId` is a herdr handle, not a peer-neutral id
`SessionManager` keys its registry by `paneId`, MCP/REST route by `paneId`, and `WorkerSession`
stores it. `paneId` is a herdr pane handle — meaningless for a peer that isn't a herdr pane.
**Decision:** introduce an opaque `PeerHandle` the launcher returns. It carries a launcher-assigned
**`id`** (the registry/routing key) plus launcher-private coordinates (for herdr: paneId, tabId,
terminalId). Stage A keeps `id == paneId` for the Claude adapter so nothing downstream changes value,
but the *type* stops being "a herdr pane" — the core routes on `PeerHandle.id()`.
```mermaid
classDiagram
class PeerLauncher {
<<interface>>
+Set~Capability~ capabilities()
+PeerHandle spawn(SpawnRequest req)
+void release(PeerHandle h)
+String effectiveCwd(SpawnRequest req)
+List~String~ parityOverlay(String profile)
+int reapOrphans()
}
class PeerHandle {
<<interface>>
+String id()
}
class ClaudeCodeLauncher {
herdr AgentControl/WorkspaceControl
SubscriptionGuard
}
PeerLauncher <|.. ClaudeCodeLauncher
ClaudeCodeLauncher ..> PeerHandle : returns
```
*Figure 2 — the SPI the core sees. `ClaudeCodeLauncher` is today's `WorkerService`, adapted.*
### 4b. Turn/completion detection reads herdr output
`inject/` (turn listener, completion resolver, status poller, presence) infers turn boundaries from
herdr terminal scraping. That is genuinely peer-transport-specific — a Codex peer would signal turns
differently. For CB-401 this stays in core (it's the *Claude/herdr* transport's detector), but §6's
capability model is what lets a future non-herdr peer bring its own turn-signalling without the core
assuming terminal scraping. **Out of scope for Stage A**; noted so the SPI doesn't accidentally
hard-wire "turns come from herdr".
## 5. Config shape
`BridgedConfig.Worker` is Claude-shaped (`baseUrl`, `model`, `configDir`, `tokenEnv`). Rather than
break existing YAML, CB-401 keeps `workers:` exactly as-is and treats those fields as the
**ClaudeCodeLauncher's** profile schema. A future peer kind adds a `kind:` discriminator
(default `"claude-code"`) selecting the launcher; unknown-kind → clear config error. No migration of
existing configs. (Jackson already ignores unknown keys, so adding `kind` is backward-safe.)
```mermaid
sequenceDiagram
participant MCP as bridge_spawn (MCP/REST)
participant SM as SessionManager
participant L as PeerLauncher (by profile.kind)
participant T as transport (herdr)
MCP->>SM: acquire(profile, cwd, owner)
SM->>L: spawn(SpawnRequest)
L->>L: build env + guard + argv (adapter-private)
L->>T: start(name, argv, env, cwd)
T-->>L: handle (paneId…)
L-->>SM: PeerHandle(id)
SM->>SM: register session keyed by handle.id()
SM-->>MCP: session view
```
*Figure 3 — spawn delegation. The core's `acquire` is unchanged in shape; only the thing it calls
becomes an interface.*
## 6. Capabilities
Peers are not uniform. Let each launcher declare a capability set; the protocol is the union and
degrades gracefully when a launcher lacks one:
| Capability | Meaning | Claude Code | Codex (likely) | Human |
|---|---|---|---|---|
| `MID_TURN_ASK` | supports `bridge_ask` rendezvous | ✓ | ? | ✗ |
| `SELF_PR` | can open its own PR at checkpoint (CB-302) | ✓ (opt-in token) | ? | ✗ |
| `WORKTREE` | can run in a provisioned git worktree | ✓ | ✓ | ✗ |
| `ORPHAN_REAP` | spawner can reconcile orphaned peers on boot | ✓ | ? | ✗ |
A verb invoked against a peer that lacks the capability returns a clean "unsupported for this peer"
rather than a crash. This keeps the protocol honest as peers diversify and prevents the core from
assuming "every peer is a Claude in a worktree" (the drift signal from the identity note).
## 7. Staging & the Stage-C security gate
```mermaid
flowchart TD
A["Stage A — CB-401<br/>extract PeerLauncher SPI in-tree<br/>ClaudeCodeLauncher = adapted WorkerService<br/>one impl, config-selected"] --> B["Stage B — CB-402+<br/>2nd in-tree adapter (Codex/human)<br/>proves the SPI held"]
B --> C["Stage C<br/>dynamic external plugin loading<br/>ServiceLoader / jar discovery"]
C -.requires.-> G["Trust & capability model<br/>what env/tokens a plugin may inject"]
classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff;
class G gate
```
*Figure 4 — deliver A now; B when a real second peer exists; C only if third parties must ship
adapters, and only behind a trust model.*
**Security note (Stage C, not now):** a launcher runs at daemon privilege and touches process
spawning **and env/token injection into peers** — the most sensitive path in the system. "Extra
plugins" must mean *first-party, in-tree, config-selected* for the foreseeable future. A third-party
plugin that can inject env into a peer needs a genuine trust/capability model before it may exist.
Regardless of stage, the **primary remains the merge/verify gate** — self-reports over the bus are
messages, not verified facts.
## 8. Stage A scope (this branch, delegation-ready)
Deliverable for CB-401 Stage A — mechanical, behaviour-preserving:
1. `PeerLauncher` interface + `PeerHandle` (opaque id) + `SpawnRequest` (profile, requestedCwd,
callerCwd) + `Capability` enum, new package `dev.ltms.bridged.peer`.
2. `ClaudeCodeLauncher implements PeerLauncher` = today's `WorkerService`, adapted: `spawn(...)`
returns a `PeerHandle` (id = paneId), `capabilities()` declares
`MID_TURN_ASK, SELF_PR(when token), WORKTREE, ORPHAN_REAP`.
3. `SessionManager` depends on `PeerLauncher`, not `WorkerService` concretely; routing keys on
`PeerHandle.id()` (== paneId today, so zero value change).
4. `Bridged.main` wires the concrete `ClaudeCodeLauncher` behind the interface.
5. **No behaviour change, no config change.** Full green gate: `ide_sync` → `ide_diagnostics`
(0 errors/0 warnings) → `mvn clean install` with `MVN_EXIT` captured (no masking pipe). All
existing tests pass unchanged; add tests only for the new `PeerHandle` indirection.
Explicitly **out of scope** for Stage A: `kind:` config discriminator, any second adapter, capability
*enforcement* at the verb layer (declare only), touching `inject/` turn detection, dynamic loading.
+374
View File
@@ -0,0 +1,374 @@
# MCP Contract — `bridged`'s unified gateway
> **Status:** 🟡 Design (2026-07-14). Greenfield — no MCP code exists yet; the pom carries
> only Javalin/Jackson. This page defines the tool surface that CB-104 and its followers
> implement. It supersedes nothing; it fills the "MCP server face" left open by the
> [Architecture](1-Architecture) page.
`bridged` is the **sole communication gateway** for every Claude session in the bridge. Both
the **primary** (Opus, on subscription) and every **worker** (off-subscription Claude Code)
mount the *same* MCP server with a single `claude mcp add` line, and talk only through its
tools. No Claude session ever addresses a broker, a peer, or the network directly.
This document defines every MCP tool that face must expose, who may call it, its blocking
semantics, and how it maps onto the code already in the tree.
---
## 1. Design constraints (non-negotiable)
These come from the project's core invariants and bound every decision below.
1. **One server, both roles.** The primary and all workers mount an identical server. The
catalog must serve both, and `bridged` must decide *who is calling* from the connection —
never from a caller-supplied argument that could be spoofed.
2. **Subscription-safe by construction.** No MCP tool ever reads, sets, or forwards
`ANTHROPIC_BASE_URL`. Mounting the bridge cannot move a session off subscription.
Enforced today by [`SubscriptionGuard`](1-Architecture).
3. **Blocking rendezvous, no busy-poll.** The primary consumes a worker's reply through a
*single* MCP call that `bridged` holds open — never a cross-turn poll loop that would burn
subscription quota.
4. **Status-gated delivery.** Anything that puts text into a worker flows through the existing
[`Injector`](1-Architecture): delivered only when the worker is `idle`/`blocked`, at most
one message per turn.
5. **`bridged` owns policy; herdr owns PTYs.** MCP tools express *intent*; `bridged`
translates it into guard checks, rendezvous bookkeeping, and herdr `agent.*` calls.
---
## 2. Topology
Both faces live in the one daemon. The **north face** is MCP (this document); the **south
face** is the herdr Unix socket. REST/SSE remains only for non-Claude clients and dashboards.
```mermaid
flowchart LR
OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
subgraph BD["bridged — standalone daemon"]
MCP["MCP server (north face)<br/>bridge_send · bridge_reply<br/>bridge_ask · bridge_status · lifecycle"]
RDV["rendezvous registry<br/>(blocking-call waiters)"]
INJ["Injector + StatusPoller<br/>(status-gated writer)"]
SOCK["herdr socket client (south face)"]
MCP --> RDV
RDV --> INJ
INJ --> SOCK
MCP --> SOCK
end
HERDR["herdr<br/>panes · agent-status"]
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
OPUS -->|"bridge_send (blocks)"| MCP
W -.->|"bridge_reply / bridge_ask"| MCP
SOCK -->|"agent.start · agent.send<br/>agent.get · pane.close"| HERDR
HERDR -->|"drives PTY"| W
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
class OPUS,W ext
class MCP,RDV,INJ,SOCK core
```
---
## 3. Identity & addressing
Because the same server is mounted by everyone, `bridged` resolves the caller's role on every
request — this is the linchpin of the whole contract and has no code yet.
- **Workers are known.** `bridged` spawns every worker
([`WorkerService`](1-Architecture)) and records its herdr session UUID / `terminal_id` on
the returned [`Agent`]. When a call arrives on a connection that maps to a known worker,
the caller is *that* worker — so **workers never pass a target**; routing is implicit.
- **The primary is "not a worker".** Any connection that does not map to a known worker is
treated as a primary. It addresses workers **explicitly** by `target` — a session UUID,
a `terminal_id`, or a friendly `profile` name.
- **Turn correlation.** A blocking `bridge_send` registers a *waiter* keyed by worker
identity. A worker's later `bridge_reply` / `bridge_ask` on the same identity resolves that
waiter. A `turn_id` is minted per exchange so a clarification round-trip
(§6.2) rejoins the right turn.
---
## 4. Transport
`bridged` is a long-lived daemon serving **multiple** concurrent clients (one primary + N
workers), so a per-client stdio child is the wrong shape. The recommended transport is
**streamable-HTTP / SSE** on the same bind as the REST face:
```bash
# identical on primary and every worker
claude mcp add --transport http bridged http://127.0.0.1:8080/mcp
```
This adds an MCP-server dependency the pom does not yet carry. See [Open decisions](#10-open-decisions).
---
## 5. Tool catalog
| Tool | Caller | Blocks? | Backing (exists today?) |
|---|---|---|---|
| [`bridge_send`](#bridge_send) | primary | yes (default) | `Injector.enqueue` ✅ · rendezvous registry ❌ (CB-104) |
| [`bridge_reply`](#bridge_reply) | worker | no | rendezvous ❌ · pane injection via `Injector` ✅ |
| [`bridge_ask`](#bridge_ask) | worker | yes | reverse rendezvous ❌ |
| [`bridge_status`](#bridge_status) | either | no | `AgentControl.status` ✅ · `Injector.activeTargets` ✅ |
| [`bridge_spawn`](#lifecycle) | primary | no | `WorkerService.spawn` ✅ (`POST /workers`) |
| [`bridge_list`](#lifecycle) | either | no | `WorkerService.list` ✅ (`/agents`) |
| [`bridge_stop`](#lifecycle) | primary | no | `WorkerService.stop` ✅ (`DELETE /workers/{paneId}`) |
| [`bridge_read`](#bridge_read) | primary | no | `AgentControl.read` ✅ |
| [`bridge_cancel`](#bridge_cancel) | primary | no | — ❌ (future) |
### Core: delegation & rendezvous
#### `bridge_send`
*(primary → worker — the headline tool, CB-104)*
- **Params:** `message` (required); `target` (optional — defaults to the sole worker / default
profile); `timeout_seconds` (default 600); `block` (default `true`); `auto_spawn`
(default `true`); `turn_id` (optional — supplied when answering a worker's `bridge_ask`).
- **Blocking (`block:true`):** enqueue `message` via the `Injector`, then hold the call open
until exactly one of:
- worker calls `bridge_reply` → `{ outcome:"reply", text }`
- worker calls `bridge_ask` → `{ outcome:"question", text, turn_id }`
- worker's `agent_status` reaches done/idle with no reply → `{ outcome:"turn_done", text:<terminal tail> }`
- deadline elapses → `{ outcome:"timeout" }`
- worker gone → error `worker_gone`
- **Detached (`block:false`):** enqueue and return `{ outcome:"dispatched", dispatch_id }`
immediately. The eventual reply is injected into the primary's idle pane (§6.3), or drained
via `bridge_status` on a split-host primary.
#### `bridge_reply`
*(worker → primary)*
- **Params:** `text` (required); `final` (default `true`).
- **Behavior:** resolve the primary waiter registered against this worker with `text`. If no
waiter exists (detached delegation), `bridged` **injects the primary's idle pane** instead.
Returns `{ delivered:true, mode:"resolved"|"injected" }`. No `target` — identity is implicit.
#### `bridge_ask`
*(worker → primary — the reverse rendezvous)*
- **Params:** `question` (required); `timeout_seconds`.
- **Behavior:** blocks the *worker's* call. Surfaces the question to the primary (resolving its
open `bridge_send` with `outcome:"question"`, or injecting its pane). When the primary
answers — a `bridge_send` carrying the matching `turn_id` — that unblocks this call and
returns `{ answer }` to the worker, which continues **in the same turn**.
### Worker lifecycle
<a id="lifecycle"></a>
Thin adapters over [`WorkerService`](1-Architecture) — parity with the existing REST routes.
- **`bridge_spawn`** — `{ profile? }` → worker view (`sessionId`, `terminalId`, `paneId`,
`status`). Guard-checked; a boundary breach returns error `subscription_boundary` (the
REST `403`).
- **`bridge_list`** — no params → all workers + `agent_status`. Read-only, either role.
- **`bridge_stop`** — `{ target }` → tears down the pane and its dedicated tab. Idempotent.
### Observability
#### `bridge_status`
*(either role — the README's 4th named tool)*
- **Params:** `target?`.
- **Behavior:** per-worker `agent_status`, queue depth (`Injector.activeTargets`), whether a
rendezvous is open, and ids. For the *calling* session it also reports/drains **pending
messages addressed to me** — the path a split-host primary's `Stop`-hook uses to wake and
collect replies without being injectable. Read-only, non-blocking.
#### `bridge_read`
*(primary)*
- **Params:** `target`; `source` ∈ `visible | recent | recent_unwrapped | detection`.
- **Behavior:** returns the worker's terminal text so the primary can peek at a *detached*
worker's progress. Adapter over `AgentControl.read`.
### Control (future)
#### `bridge_cancel`
*(primary)*
- **Params:** `target`. Interrupt the worker's current turn / abandon the rendezvous. No
backing code yet.
---
## 6. Rendezvous flows
### 6.1 Delegation — happy path
One blocking call, zero polls.
```mermaid
sequenceDiagram
participant P as Primary (Opus)
participant B as bridged (MCP + Injector)
participant H as herdr
participant W as Worker (Claude)
P->>B: bridge_send("do X", target=w) — blocks
B->>B: register waiter(w)
B->>H: agent.send(w, "do X") (idle window)
H-->>W: prompt injected
W->>W: works the turn
W->>B: bridge_reply("result")
B->>B: resolve waiter(w)
B-->>P: { outcome:"reply", text:"result" }
```
### 6.2 Clarification — reverse rendezvous (`bridge_ask`)
The worker pauses mid-turn to ask; the primary answers; the worker resumes in the same turn.
```mermaid
sequenceDiagram
participant P as Primary
participant B as bridged
participant W as Worker
P->>B: bridge_send("do X", target=w) — blocks
B-->>W: "do X" (injected)
W->>B: bridge_ask("which config?") — worker blocks
B-->>P: { outcome:"question", text:"which config?", turn_id }
P->>B: bridge_send("config.yaml", target=w, turn_id) — blocks again
B-->>W: resolve bridge_ask → { answer:"config.yaml" }
W->>W: resumes same turn
W->>B: bridge_reply("done")
B-->>P: { outcome:"reply", text:"done" }
```
### 6.3 Detached delegation — pane injection
The primary does not block; the reply arrives later in its idle pane.
```mermaid
sequenceDiagram
participant P as Primary
participant B as bridged
participant W as Worker
P->>B: bridge_send("do X", target=w, block=false)
B-->>P: { outcome:"dispatched", dispatch_id }
P->>P: continues its own work
W->>B: bridge_reply("result")
Note over B: no waiter → detached path
B->>B: Injector.enqueue(primary_pane, "result")
B-->>P: injected into idle pane (status-gated)
```
### 6.4 Uncooperative worker — turn-done fallback
A worker that never calls `bridge_reply` still returns a result: `bridged` reads its terminal
tail when the turn completes.
```mermaid
sequenceDiagram
participant P as Primary
participant B as bridged
participant W as Worker
P->>B: bridge_send("do X", target=w) — blocks
B-->>W: "do X" (injected)
W->>W: works, never calls bridge_reply
B->>B: StatusPoller sees agent_status → idle/done
B->>B: AgentControl.read(w, "recent")
B-->>P: { outcome:"turn_done", text:<terminal tail> }
```
---
## 7. Status gating
Delivery only happens in a safe window. This is the state machine the `Injector` already
enforces via `AgentStatus.injectable()`; MCP `bridge_send` is simply its producer.
```mermaid
stateDiagram-v2
[*] --> IDLE
IDLE --> WORKING: message delivered / picks up
WORKING --> IDLE: turn done
WORKING --> BLOCKED: awaits input
BLOCKED --> WORKING: input delivered
IDLE --> UNKNOWN: detection glitch
BLOCKED --> UNKNOWN: detection glitch
UNKNOWN --> IDLE: re-detected
note right of IDLE
injectable — deliver head of FIFO
end note
note right of BLOCKED
injectable — deliver head of FIFO
end note
note right of WORKING
NOT injectable — counts as pickup
end note
note right of UNKNOWN
NOT injectable, NOT a pickup — wait
end note
```
At most one message is delivered per turn: after a send the `Injector` waits for a `WORKING`
pickup before delivering the next, with a `PICKUP_GRACE_POLLS` fallback for turns faster than
the poll interval. A herdr `events.subscribe` stream can later replace the sampling without
touching this state machine.
---
## 8. Error model
| Condition | `bridge_send` result | Notes |
|---|---|---|
| Worker replies | `{ outcome:"reply" }` | normal |
| Worker asks | `{ outcome:"question", turn_id }` | answer with `bridge_send(turn_id)` |
| Turn ends, no reply | `{ outcome:"turn_done" }` | terminal tail as text |
| Deadline elapsed | `{ outcome:"timeout" }` | message may still be queued/delivered |
| Worker vanished | error `worker_gone` | `Injector.drop` fails the queued future |
| Guard breach on spawn | error `subscription_boundary` | REST `403` parity |
| Delivery failed at herdr | error, message dropped | poisoned message not left blocking the FIFO |
`bridge_reply` from a worker with no open waiter is **not** an error — it falls through to
detached pane injection (§6.3).
---
## 9. Mapping to existing code
The MCP face is a thin adapter layer; nearly every capability already exists behind the REST
seam. Only the **rendezvous registry** and the **caller-identity resolver** are new.
| MCP tool | Existing collaborator | New work |
|---|---|---|
| `bridge_send` | `Injector.enqueue`, `AgentControl.send` | waiter registry, timeout, outcome mux (CB-104) |
| `bridge_reply` / `bridge_ask` | `Injector` (pane injection) | reverse rendezvous, identity resolver |
| `bridge_status` | `AgentControl.status`, `Injector.activeTargets` | pending-drain projection |
| `bridge_spawn` / `list` / `stop` | `WorkerService.{spawn,list,stop}` | MCP adapter only |
| `bridge_read` | `AgentControl.read` | MCP adapter only |
Because the REST routes in `BridgedApp` already exercise the collaborators, MCP tools are
validated by **parity** against those routes, not by re-testing behavior.
---
## 10. Open decisions
1. **`bridge_ask` direction.** This page defines it as *worker-asks-primary* (a genuine reverse
channel, matching the "inject the primary's pane" language). The alternative — a synonym for
a blocking primary→worker send — is weaker and produces different plumbing. **Recommend
worker-asks-primary.**
2. **Detached delivery shape.** A `block:false` param on `bridge_send` (keeps the catalog
small) vs. a separate `bridge_dispatch` tool. **Recommend the param.**
3. **Auto-spawn on send.** `bridge_send` provisions a worker per profile when none exists
(simplest primary UX) vs. requiring an explicit `bridge_spawn` first. **Recommend
auto-spawn, defaulting on.**
4. **Transport & SDK.** Streamable-HTTP/SSE co-located with the REST bind (recommended) vs.
stdio. Requires choosing a Java MCP server SDK and adding it to the pom.
---
## 11. Implementation staging
- **CB-104** — blocking `bridge_send` + rendezvous registry + caller-identity resolver
(the producer that finally drives the inert `StatusPoller`).
- **CB-1xx** — `bridge_reply` / `bridge_ask` reverse rendezvous + detached pane injection.
- **CB-1xx** — lifecycle + observability adapters (`bridge_spawn/list/stop/status/read`).
- **CB-1xx** — transport wiring + `claude mcp add` docs; parity tests vs. REST.
- **Later** — `bridge_cancel`; swap `StatusPoller` for herdr `events.subscribe`.
+154
View File
@@ -0,0 +1,154 @@
# Team — lead orchestrating a mixed Claude + local-LLM fleet
The message server (`bridged`) delivers **one turn into one worker**. A **team** is the
layer above it: a **Claude team-lead** that fans a job out across a **mixed fleet** of
workers — some on Claude, some on the remote local LLM — and reduces their replies. Same
`bridged` delivery, same subscription boundary; this doc is only about **orchestration** —
who the workers are, how the lead picks one, and how it runs many at once.
> Delivery mechanics (blocking `POST /message`, status-gated reply envelope) live in the
> Message-Server design. Transport rationale is in Approaches. This doc assumes both.
## The team
- **Team-lead** — the primary **Opus** (Claude Code, env **CLEAN**, on Pro/Max). Not a
worker; a **thin client of `bridged`**. It plans, routes, dispatches, and integrates, and
never sets `ANTHROPIC_BASE_URL`.
- **Workers** — a herd of `claude` panes in herdr, each an addressable `bridged` session
with its **own model/env**:
- **Claude workers** (clean env, e.g. Sonnet) — reasoning-heavy or high-accuracy subtasks.
- **Local workers** (`ANTHROPIC_BASE_URL=https://ollama.ltms.dev`) — bulk, cheap, or
embarrassingly parallel subtasks.
Every worker is still a *real Claude Code process* (inherits `CLAUDE.md`, hooks, skills,
MCP) — only its model differs. Scale each kind horizontally by adding panes.
### Topology
```mermaid
flowchart TB
LEAD["lead — Opus<br/>(Claude Code, env CLEAN)"]
BD["bridged<br/>message server + router"]
HERDR["herdr<br/>panes · agent-status"]
WC1["w-claude-1<br/>Sonnet · CLEAN"]
WC2["w-claude-2<br/>Sonnet · CLEAN"]
WL1["w-local-1<br/>ANTHROPIC_BASE_URL set"]
WL2["w-local-2<br/>ANTHROPIC_BASE_URL set"]
ANT["api.anthropic.com<br/>(Pro/Max)"]
OLL["ollama.ltms.dev<br/>(local model)"]
LEAD -->|"blocking POST /message (target role)"| BD
BD -->|"Unix socket · send_text · events.subscribe"| HERDR
HERDR --> WC1 & WC2 & WL1 & WL2
WC1 --> ANT
WC2 --> ANT
WL1 --> OLL
WL2 --> OLL
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
classDef local fill:#6b46c1,stroke:#44337a,color:#ffffff;
class LEAD,WC1,WC2 ext
class BD,HERDR core
class WL1,WL2 local
```
### Roles & routing
| Role | Env | Model | Route here when… |
|---|---|---|---|
| `lead` | clean | Opus (sub) | always — it does the routing |
| `w-claude-*` | clean | Sonnet (sub) | task needs Claude-grade reasoning / careful edits |
| `w-local-*` | `ANTHROPIC_BASE_URL` set | local LLM | task is bulk / cheap / embarrassingly parallel |
The lead applies this rubric itself, guided by its `CLAUDE.md` team charter (below). Worker
selection is **policy in the lead**, not a `bridged` concern — `bridged` just delivers to
the session the lead names.
## Subscription boundary in a team
Unchanged from the base architecture, and it scales with the fleet: **only local-worker
panes** launch with `ANTHROPIC_BASE_URL`. The lead and every Claude worker stay env-clean on
the subscription. `bridged` enforces which panes may carry the off-subscription env, so
adding workers never widens the boundary.
## Parallel fan-out (map / reduce)
The lead's advantage over a single bridge is **concurrency**: independent subtasks go to
different workers at once, then results are gathered.
```mermaid
sequenceDiagram
participant L as lead (Opus)
participant B as bridged
participant WC as w-claude-1
participant WL as w-local-1
Note over L: split job → subtask A (reasoning), subtask B (bulk)
par A → Claude worker
L->>B: POST /message {role: w-claude, prompt: A}
B->>WC: send_text into running pane
WC-->>B: status working → idle + Stop-hook envelope
B-->>L: 200 reply A
and B → local worker
L->>B: POST /message {role: w-local, prompt: B}
B->>WL: send_text into running pane
WL-->>B: status working → idle + Stop-hook envelope
B-->>L: 200 reply B
end
Note over L: reduce → integrate A + B into final answer
```
- **Map:** the lead issues N concurrent blocking `POST /message` calls (one per subtask → its
chosen worker). Each call blocks only *that* request; `bridged` holds it open until the
worker's turn completes (status-gated) and returns the reply envelope.
- **Reduce:** the lead collects the N envelopes and integrates. A slow local worker never
blocks a fast Claude worker — wall-clock ≈ the slowest single subtask, not the sum.
- **Detached / long jobs** use the async broker path instead of a held request (Channel 2 in
the base architecture), so the lead never busy-polls across turns.
Fan-out is bounded by the herd size (pane count) and `bridged`'s concurrency policy, not by
the lead.
## Knowing the roster
The lead discovers its team from `bridged` (session list / roles) rather than hard-coding
pane ids, so workers can be added or restarted without editing the lead. A minimal charter
in the lead's `CLAUDE.md` turns Opus into the orchestrator:
```markdown
## Your team (via bridged)
You are the team-lead. Delegate through the bridged client — never launch workers yourself.
Roster: ask bridged for current sessions/roles.
- w-claude-* — Claude Sonnet. Reasoning-heavy / high-accuracy subtasks.
- w-local-* — remote local LLM. Bulk, cheap, or parallelizable subtasks.
Route each subtask by the rubric in the Team design. For independent subtasks, DISPATCH ALL
of them (concurrent blocking sends), THEN gather — never serialize independent work.
Integrate the reply envelopes; you own the final answer.
```
Wrap the send as a Claude Code skill (`/delegate <role> "<task>"`) so the lead calls one
tool instead of hand-rolling the HTTP request.
## What this layer does NOT change
- **Delivery** is still `bridged` → herdr `pane.send_text` + status events (Message-Server).
- **Completion timing** is still the worker status event; **reply content** still rides the
worker `Stop`-hook envelope.
- **Single-host** still applies: herdr's socket is local, so the whole herd lives on the
`bridged` host. The lead may be remote — it only needs HTTP to `bridged`.
## Open questions
- **Routing intelligence:** rubric-in-`CLAUDE.md` (lead decides) vs. a `bridged` role-router
(label-based). Start with the former; promote to the latter if routing logic grows.
- **Backpressure:** per-role concurrency caps in `bridged` so a fan-out can't exhaust the
local gateway.
- **Result schema:** whether reply envelopes should carry structured metadata (worker, model,
tokens) to help the lead's reduce step.
## Status
🟡 Design (2026-07-11). Orchestration layer over the selected `bridged` server; inherits
herdr (chosen) + AgentAPI (fallback). Delivery unchanged — see the Message-Server design.
+182
View File
@@ -0,0 +1,182 @@
# Worker Git Workflow — worktree · branch · PR
**Status:** design (defining the fleet's working model). Builds on the
[CB-301 session manager](CB-301-Session-Manager.md) and the one-shot/no-reuse decision.
## Guiding principle — a worker is a full peer of the primary
The whole point of the bridge is **the same Claude Code agent running against a different LLM
provider**. A worker must be **functionally identical to the primary in context and knowledge** —
same project + user `CLAUDE.md`, same skills, same memory, same MCP tools, **same local/private
settings** — and differ **only** in `ANTHROPIC_BASE_URL`/`ANTHROPIC_MODEL`. The worktree exists
*solely* for git isolation (a branchable checkout for commits + code reference). It must **never**
strip the worker of the configuration a main-tree session has. **Config parity is a hard
requirement, not a nice-to-have.**
## Decisions
- **One-shot, no reuse** (CB-301) — each task gets a fresh worker, torn down after. The **PR is the
durable artifact**; no context is carried across workers.
- **Worktree provisioned by the daemon** — `SessionManager` creates a dedicated git worktree +
branch per session, **hydrates it to full config parity** (below), and tears it down on release.
- **Worker opens its own PR** — the worker commits, pushes its branch, and opens the PR/MR itself,
returning the PR URL in its `bridge_reply`.
## Why worktrees (the hazard being fixed)
Today every bridge worker inherits the **primary's own working tree** as its cwd
(`WorkerService` cwd resolution → caller cwd). A single worker editing at a time is safe, but two
**parallel implementers** would stomp each other's files. A worktree per session gives each worker
an isolated checkout on its own branch — the precondition for fanning out implementation work.
## Config parity — the worktree is code-only; it must NOT strip the worker
**The trap:** `git worktree add` creates a checkout of **tracked files only**. Untracked / gitignored
files do **not** come along. So a worker moved from the primary's main tree into a bare worktree
**silently loses** exactly the local/private configuration that makes it a full peer — while the
shared-tree it runs in *today* gives it all of this for free. Moving to worktrees must **preserve**
that, not regress it.
Classify every source of "what a main session knows", by whether a worktree keeps it:
```mermaid
flowchart TD
subgraph keeps["Inherited automatically — no action"]
A["User-global config<br/>~/.claude/CLAUDE.md, ~/.ccs memory"]:::ok
B["Tracked project config<br/>CLAUDE.md, committed .claude/skills, committed .mcp.json"]:::ok
C["CLAUDE_CONFIG_DIR<br/>(daemon already injects per profile)"]:::ok
D["Bridge MCP<br/>(injected via --mcp-config launch flag)"]:::ok
end
subgraph gap["LOST by a bare worktree — must be hydrated"]
E["settings.local.json<br/>local .claude/* overrides"]:::warn
F["local .mcp.json mods<br/>(the M .mcp.json in git status)"]:::warn
G[".env / .envrc / direnv<br/>local secrets + tokens"]:::warn
H["any other gitignored local config"]:::warn
end
keeps --> R["Worker = full peer of primary"]:::goal
gap -->|"overlay step at provision"| R
classDef ok fill:#2f855a,stroke:#22543d,color:#ffffff;
classDef warn fill:#b7791f,stroke:#7b341e,color:#ffffff;
classDef goal fill:#2b6cb0,stroke:#2a4365,color:#ffffff;
```
*Green is inherited by path (home dir / `CLAUDE_CONFIG_DIR`) or lives in tracked files that the
worktree checks out anyway. Amber is the real gap — untracked local config the worktree drops.*
**Mechanism — worktree hydration (part of `SessionManager.acquire`, after `git worktree add`):**
1. **Inherit by path, don't copy** — keep the worker's `$HOME`, `CLAUDE_CONFIG_DIR`, and memory dir
identical to the primary's. Everything home-scoped (user `CLAUDE.md`, memory, auth, global
skills) is already parity for free; only *cwd-relative* project-local files are the gap.
2. **Overlay the untracked project-local set** from the primary tree into the worktree — a defined,
configurable list: `settings.local.json` (+ any local `.claude/*`), the locally-modified
`.mcp.json`, `.env`/`.envrc`, and any other gitignored config the primary depends on. **Symlink**
(read-only parity, stays live, nothing to go stale) rather than copy where possible; copy only
what a worker may write.
3. **Never overlay the git plumbing** — the worktree's own `.git` file/branch is what gives
isolation; that's the *one* thing that must differ from the main tree.
The overlay set lives in config (`BridgedConfig.Worker.parityOverlay` — a list of repo-relative
paths, with sane defaults) so it's auditable and per-repo tunable.
> **Trust note (deliberate).** Hydrating local config means the primary's local secrets/tokens
> (`.env`, `.mcp.json` auth, gitea token) become visible to an **off-subscription** worker running
> against a third-party model endpoint. That is the accepted consequence of "workers must be full
> peers" — but it is a real trust expansion over a worker that only sees tracked code. Keep the
> overlay list **explicit and minimal**; don't blanket-symlink the whole tree. `.mcp.json` and
> `wiki/` remain **excluded from all worker commits** regardless of being present for reference.
## Lifecycle
```mermaid
sequenceDiagram
autonumber
participant P as Primary
participant SM as SessionManager (daemon)
participant G as git / gitea
participant W as Worker
P->>SM: acquire(ticket, profile)
SM->>G: git worktree add wt -b worker/ticket-nonce main
SM->>SM: overlay parity config into wt
SM->>W: spawn (cwd = wt, on its branch)
Note over W: implement in the isolated worktree
W->>G: git commit + git push (SSH, same user)
W->>G: open PR (branch to main)
W-->>P: bridge_reply (prUrl, branch, summary, tests)
P->>SM: release(paneId)
SM->>G: git worktree remove wt
Note over G: branch + PR persist for review/merge
P->>G: review PR, merge on green
```
*The worker's "checkpoint" (CB-302) is exactly steps 6–8: commit → push → open PR. This replaces the
earlier `STATE.md` idea — a PR is reviewable, mergeable, and self-describing.*
## Infra facts (verified this session)
- **Remote:** `ssh://git@git.ltms.dev:2224/lms/claude-bridge.git` (gitea). Push is over **SSH** —
a worker running as the same user with the same keys can `git push` **with no extra credential**.
- **gitea is NOT in the project `.mcp.json`** (only `jetbrains`, `intellij-index`, `bridged`). The
primary's gitea MCP comes from a global/user config, so **workers do not inherit it**. A worker
gets only the `bridge` MCP mounted (via `--mcp-config` launch flag).
- **No gitea CLI** (`tea`) installed; `glab` is present but is the GitLab CLI (wrong backend).
**Implication:** `git push` is free for workers; only **PR creation** needs a new mechanism.
## Open decision — how the worker creates the PR
| Option | Mechanism | Trade-off |
|---|---|---|
| **A. gitea REST + token** | Worker `curl`s `POST /api/v1/repos/lms/claude-bridge/pulls` with a scoped token injected by the daemon into the worker env | Minimal, no new server; token lives in the off-subscription worker's env (scope it tightly) |
| **B. mount gitea MCP into workers** | Add the gitea MCP to the worker's `--mcp-config` alongside `bridge` | Clean tool call, but the gitea MCP's own auth/token must be provisioned per worker; more moving parts |
| **C. install `tea` CLI** | Worker runs `tea pr create` with a token | Another dependency to install + configure; same token question as A |
**Recommendation: A (gitea REST + a repo-scoped token).** Smallest surface, reuses SSH for push,
and the token is a single scoped secret the daemon injects like it already injects
`ANTHROPIC_AUTH_TOKEN`. The implementer skill wraps the `curl` in one documented step.
### Trust / token scope (the real cost of "worker opens its own PR")
- Off-subscription workers already *could* push (SSH, same user). The **incremental grant is
PR-create**, i.e. a gitea API token.
- Scope the token **minimally**: the `lms/claude-bridge` repo, `write:repository` (create branch +
PR), **not** merge/admin/org. A leaked token can open PRs, not merge them — the primary/human is
still the merge gate.
- Inject via the daemon (env var, e.g. `GITEA_TOKEN`), never written to the worker's config dir —
same non-invasive pattern as the ANTHROPIC token. Guard is unaffected (it concerns
`ANTHROPIC_BASE_URL`, not git).
## Implementation plan
| Piece | Where | Notes |
|---|---|---|
| Worktree provision/teardown | **CB-301 ext** — `SessionManager.acquire`/`release`; `WorkerSession` gains `worktree`, `branch` | daemon shells out to `git worktree add/remove` |
| **Config-parity overlay** | **CB-301 ext** — `SessionManager.acquire`, after `git worktree add` | symlink/copy the `parityOverlay` set into the worktree so the worker is a full peer; **this is what makes worktrees viable, not a dead-end** |
| Overlay config | `BridgedConfig.Worker.parityOverlay` — repo-relative paths, sane defaults | auditable, per-repo tunable; keep explicit + minimal (trust) |
| Branch naming | `worker/<ticket-slug>-<nonce>` off `main` (or a configured base) | one branch per session |
| Commit + push + PR handoff | **CB-302** — worker-driven, guided by the skill | push = SSH; PR = option A |
| Implementer skill | `.claude/skills/implementer/SKILL.md` | worktree-aware playbook (see below); mounts automatically since workers inherit repo cwd |
| gitea token injection | `WorkerService` env + `BridgedConfig` | repo-scoped, minimal perms |
| PR review + merge | Primary (has gitea MCP + judgment) | merge on green; the human/primary gate stays |
## Implementer skill (outline)
A worker-facing playbook (sibling to the existing `reviewer` skill):
1. **You are in a git worktree on a dedicated branch** — check `git status`/`git branch`; do all
work here, never on `main`.
2. **Implement the task**; keep commits focused and message them clearly.
3. **Push** your branch (`git push -u origin HEAD`).
4. **Open a PR** to `main` (option A `curl`, or the decided mechanism) with a title/body describing
the change and referencing the ticket.
5. **Reply** via `bridge_reply` with the **PR URL**, branch name, files changed, and test names —
that reply is the whole handoff.
6. Do **not** merge; do **not** touch `.mcp.json` or `wiki/`.
## Sequencing
1. Finish + verify **CB-301 core** (in flight) — registry/FSM.
2. Extend CB-301 with **worktree provisioning** (this doc) once the PR mechanism is chosen.
3. Add the **implementer skill** + **token injection**.
4. **CB-302** = wire the worker checkpoint (commit/push/PR) as the release-time handoff.
+137
View File
@@ -0,0 +1,137 @@
# Worker startup: working directory & the folder-trust prompt
When `bridged` spawns a worker, the worker CLI may show an **interactive startup prompt** before it
is ready to accept a task — most importantly a *"Do you trust the files in this folder?"* dialog. An
unattended worker parked on that prompt never becomes injectable: the status-gated injector waits for
`idle`/`blocked`, the task is never delivered, and (worst case) a stray Enter answers the dialog
wrong. How this is handled depends on two things:
1. **The worker's working directory** — which folder the CLI is asked to trust.
2. **Which CLI launches the worker** — each has its own first-run/trust behaviour.
This doc pins the current assumption (**`ccs`** is the launcher), how its trust prompt works, the
rule that **a worker inherits the primary's directory** (never `$HOME`), and how other CLIs differ.
```mermaid
flowchart TD
A["bridge_spawn / POST /workers"] --> B{"explicit cwd?<br/>(profile cwd or spawn arg)"}
B -->|"yes — told otherwise"| C["use that cwd"]
B -->|"no"| D{"caller PID resolvable?<br/>(MCP peer PID)"}
D -->|"yes"| E["cwd = the primary's cwd<br/>lsof -a -p PID -d cwd"]
D -->|"no (REST / off-host)"| F["cwd = bridged daemon cwd<br/>(never $HOME by assumption)"]
C --> G["ensureWorkspace → tab.create → agent.start {cwd}"]
E --> G
F --> G
G --> H{"does the CLI trust this folder?"}
H -->|"yes"| I["worker reaches its prompt → injectable"]
H -->|"no"| J["worker BLOCKS on the trust dialog<br/>never injectable"]
classDef good fill:#2f855a,stroke:#22543d,color:#ffffff;
classDef bad fill:#9b2c2c,stroke:#63171b,color:#ffffff;
class I good
class J bad
```
*Figure 1 — spawn resolves a working directory, then the CLI's trust check gates readiness.*
## Working directory: inherit the primary's path
**Rule: a worker opens the same directory the primary (main) session is working in, unless told
otherwise. Never assume `$HOME`.** If the primary is in `/Users/you/LTMS/claude-bridge`, its workers
open there too — so delegated tasks share the same relative paths and the same (already-trusted)
project folder.
**The herdr seam.** An `agent.start` pane does **not** inherit its tab's or workspace's cwd — it
starts in `$HOME` unless told otherwise. herdr's `agent.start` honours an (undocumented) **`cwd`**
param, verified live: setting it roots the worker process at that directory. So the worker's cwd is
threaded onto `agent.start {…, cwd}`, not the placement step (`workspace.create`/`tab.create` cwd
only affect the seed shell, which the bridge closes).
**Resolution order** (first match wins):
| # | Source | When |
|---|--------|------|
| 1 | Explicit `cwd` — a per-profile `cwd:` in config, or a spawn argument | "told otherwise" — pin a fixed workdir |
| 2 | The **primary's cwd**, auto-detected from the `bridge_spawn` caller | normal MCP spawn from the primary |
| 3 | The `bridged` daemon's own cwd | REST spawn / off-host caller — **never `$HOME`** |
The primary's cwd (source 2) is discoverable with no new plumbing: `bridged` already resolves the MCP
caller's loopback **peer PID** for connection identity (`ConnectionIdentity` → `LsofPeerPidLookup`);
the same PID yields its cwd via `lsof -a -p <pid> -d cwd -Fn` (the `n…` line). The primary maps to no
worker pane (it is not a worker), but its PID and cwd are still readable.
```mermaid
sequenceDiagram
participant P as "Primary (main)"
participant B as "bridged"
participant O as "OS (lsof)"
participant H as "herdr"
P->>B: "bridge_spawn {profile} (no cwd)"
B->>O: "peer PID for this connection's port"
O-->>B: "pid"
B->>O: "cwd of pid (lsof -d cwd)"
O-->>B: "/Users/you/LTMS/claude-bridge"
B->>H: "tab.create (placement)"
B->>H: "agent.start {argv, env, tab_id, cwd}"
H-->>B: "worker in the primary's directory"
```
*Figure 2 — a no-cwd spawn inherits the primary's directory from the caller's PID.*
> **Status:** implemented (CB-112). `bridged` threads the resolved `cwd` onto **`agent.start {cwd}`**
> (verified: the worker process is rooted there), keeping the single shared worker space. On an MCP
> `bridge_spawn` the primary's cwd is auto-detected from the caller's PID; over REST (no MCP caller)
> it is the explicit `cwd` param else the daemon's cwd. Both placements (`tab` and legacy `pane`)
> carry it, since it rides `agent.start`.
## Assumed launcher: `ccs` (Claude Code)
For now the fleet assumes **`ccs`** (Claude Code under the hood) as the worker CLI — `argv: ["ccs",
"<profile>"]`. Its startup gate is the **folder-trust dialog**.
### How `ccs`/Claude Code decides whether to prompt
Trust is recorded **per-directory, per config dir**. Each `ccs` profile is an isolated instance with
its own config dir (`~/.ccs/instances/<profile>/`) and its own `.claude.json`:
```jsonc
// ~/.ccs/instances/<profile>/.claude.json
"projects": {
"/Users/you/LTMS/claude-bridge": { "hasTrustDialogAccepted": true }, // trusted → no prompt
"/Users/you": { "hasTrustDialogAccepted": false } // untrusted → prompts
}
```
The worker prompts **iff** its cwd is not marked `hasTrustDialogAccepted: true` for *that instance*.
This is why the directory rule above matters: land workers in the primary's project folder and you
grant trust **once per profile**, instead of scattering trust across `$HOME` and ad-hoc dirs.
### Clearing the prompt (ranked)
1. **Inherit the primary's project dir** (the rule above) and trust that folder once per profile.
2. **Pre-trust interactively:** run `ccs <profile>` in the target folder and accept — persists
`hasTrustDialogAccepted: true` for that path in the instance's `.claude.json`.
3. **Set the flag directly** (scriptable, no interaction): set
`projects["<cwd>"].hasTrustDialogAccepted = true` in `~/.ccs/instances/<profile>/.claude.json`.
4. **Do not** reach for `--dangerously-skip-permissions` — it disables *all* permission gating, not
just this dialog, which defeats running off-subscription workers autonomously.
## Other CLIs: different launchers, different prompts
`ccs` is the current assumption, not a hard dependency — a worker profile's `argv` can be any CLI.
Each launcher has its **own** first-run/trust gate, so the "clear the prompt" step is CLI-specific
and belongs with the profile, not hard-coded:
| Launcher (`argv`) | Startup gate | How to clear it |
|---|---|---|
| `ccs <profile>` (Claude Code) | Folder-trust dialog | `hasTrustDialogAccepted` per project in the instance's `.claude.json` (above) |
| Other Claude-compatible runtimes via `ccs` (codex, gemini, cursor, …) | Each has its own first-run / trust / login prompt | Per-runtime; document per launcher as it is adopted |
| A bare command (`bash -c …`, mechanics probe) | None | n/a — used for non-interactive smoke tests |
When adding a new launcher, capture its startup-prompt behaviour here (what blocks, and the
non-interactive way to satisfy it) so a spawned worker of that kind reaches an injectable prompt
unattended.
## See also
- `docs/MCP-Contract.md` — the tool surface (`bridge_spawn`, `bridge_profiles`, …).
- `wiki/2-Message-Server.md` — the herdr `agent.*` / `workspace.*` schema (`workspace.create {cwd}`).
+2
View File
@@ -0,0 +1,2 @@
__pycache__/
*.pyc
+77
View File
@@ -0,0 +1,77 @@
# Bridge conversation test (`e2e/`)
A standard, repeatable **live** end-to-end test of the two-way channel: it drives a real
multi-turn conversation between a primary and an off-subscription worker **through the
running `bridged` daemon**, captures the full transcript, and grades the channel.
This is the committed form of the ad-hoc channel test that discovered the CB-115 gaps
(herdr `unknown` misclassification wedging delivery, dirty completion scrapes, and workers
never calling `bridge_reply` in conversation). Run it after any change to the injector,
status handling, completion/failure paths, or the worker reply charter.
## What it exercises
Each turn goes through the whole gateway exactly as a primary Opus session would — async
fire-and-poll (`POST /sessions/{id}/message {"wait":false}` → `GET /tasks/{ticket}`), so it
also validates the path that beats the caller's MCP timeout. It never sets
`ANTHROPIC_BASE_URL` and never talks to herdr directly, so it is **subscription-safe by
construction** — it only calls the bridge's loopback REST face.
```mermaid
sequenceDiagram
participant T as conversation_test.py
participant B as bridged (REST)
participant W as worker (off-sub)
T->>B: POST /workers (spawn)
T->>B: GET /sessions/{id}/status (await ready)
loop each turn
T->>B: POST /sessions/{id}/message {wait:false}
B-->>T: ticket
B->>W: inject prompt (status-gated)
W-->>B: bridge_reply
T->>B: GET /tasks/{ticket} (poll)
B-->>T: done + reply
end
T->>B: DELETE /workers/{pane} (stop)
```
## Prerequisites
- `bridged` is running (default REST on `http://127.0.0.1:8765`) with at least one worker
profile configured and its backend reachable.
- herdr is up (the daemon needs it).
- Python 3 (standard library only — no pip installs).
## Run
```bash
# spawn the default-profile worker, run the built-in 5-turn conversation, grade, clean up
python3 e2e/conversation_test.py
# pick a profile / reuse a live worker / use your own prompts
python3 e2e/conversation_test.py --profile ollama
python3 e2e/conversation_test.py --tid term_abc123 --keep-worker
python3 e2e/conversation_test.py --prompts my_prompts.txt --out /tmp/run1
```
A prompts file is one prompt per line; blank lines and `#` comments are ignored.
## Output & grading
- Writes `transcript.md` (in `--out`, default `e2e/`) — every turn's prompt, worker reply,
latency, resolution source, and observed status transitions, with an inline `> **GAP**`
note on any non-clean turn.
- Prints a per-turn line and an overall summary, and **exits non-zero** if any turn wedged,
failed, or returned empty — so it is CI-usable.
Per-turn grade:
| Grade | Meaning |
|------------|---------------------------------------------------------------------|
| `OK` | delivered and resolved by an explicit `bridge_reply` (`source=reply`) |
| `DEGRADED` | delivered and answered, but resolved via completion-scrape fallback |
| `EMPTY` | turn completed but the reply was empty |
| `FAILED` | the worker's turn ended in failure (`phase=failed`) |
| `WEDGE` | never resolved within the poll window (delivery wedge / lost turn) |
`PASS` requires every turn to be `OK` or `DEGRADED`; a clean run is every turn `OK`.
+277
View File
@@ -0,0 +1,277 @@
#!/usr/bin/env python3
"""Live bridge_ask test — the REVERSE rendezvous (CB-205), watched end to end.
Every other harness drives the forward path: primary `bridge_send` → worker `bridge_reply`.
This drives the one that runs the other way. A worker is told to pause its delegated turn,
ask the primary a question via `bridge_ask`, and only finish once it has the answer — so the
turn round-trips primary→worker→primary→worker inside a SINGLE delegation.
The mechanics that only this path exercises:
• a worker's mid-turn question surfacing on the primary's *own* blocked send (Outcome.QUESTION),
• the `turnId` correlation that lets the primary answer the exact paused turn,
• the answer resuming that same turn and the worker's final `bridge_reply` landing on the
re-opened forward waiter (never a stale or cross-wired one).
It is two blocking REST calls, no polling:
1. POST /sessions/{id}/message {content: TASK} → blocks, returns 202 {status:"question",
question, turnId} (the worker asked)
2. POST /sessions/{id}/message {content: ANSWER, turnId} → blocks, returns 200 {reply, replySource}
(the worker resumed and replied)
Like the rest of the suite it talks ONLY to the bridge's REST face on loopback — it never sets
ANTHROPIC_BASE_URL and never touches herdr, so it is subscription-safe by construction.
Usage:
python3 bridge_ask_test.py [--base URL] [--profile NAME] [--repo DIR] [--out DIR]
[--send-timeout SECS] [--answer-timeout SECS] [--keep-worker]
--base bridge REST base URL (default http://127.0.0.1:8765)
--profile profile for the worker (default: the daemon's default)
--repo cwd handed to the worker (default: the bridge repo root)
--out output dir for the transcript (default: alongside this file)
--send-timeout max wait for the worker to ASK (default 110)
--answer-timeout max wait for the worker to REPLY (default 90)
--keep-worker do not stop the spawned worker at the end
Exit code: 0 if the worker asked, the answer resumed the turn, and the final reply reflected
the answer; 1 otherwise (CI-usable).
"""
import argparse
import json
import pathlib
import sys
import time
import urllib.error
import urllib.request
from datetime import datetime
HERE = pathlib.Path(__file__).parent
REPO_ROOT = HERE.parent
# The color the primary will hand back when the worker asks. The final reply must reflect it,
# uppercased — proof the answer actually reached the resumed turn (not a value the worker could
# have guessed: it is told to ask, and only the primary knows which of red/blue is chosen).
ANSWER_COLOR = "blue"
# A task that CANNOT be completed without asking: the worker is not told which color to choose,
# only that the primary will name one when asked. So a correct final reply is only reachable by
# actually calling bridge_ask and using the answer.
TASK_PROMPT = (
"You are a bridge worker in a quick coordination game. You do NOT know which color to pick — "
"only the primary does. Do exactly this, in order:\n"
"1. Call the `bridge_ask` tool with EXACTLY this question: \"PICK A COLOR: red or blue?\"\n"
"2. The primary will answer with one color word. Take that color and uppercase it.\n"
"3. Call `bridge_reply` with EXACTLY one line: CHOSEN=<COLOR> (e.g. CHOSEN=GREEN if told green).\n"
"Do not guess a color. Do not call bridge_reply before bridge_ask has returned an answer. "
"Do nothing else — no file reads, no other tools."
)
def http(base, method, path, body=None, timeout=20):
"""JSON request. Tolerates an empty body (e.g. 204) → {}. Raises on non-2xx via urllib."""
data = json.dumps(body).encode() if body is not None else None
req = urllib.request.Request(base + path, data=data, method=method,
headers={"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=timeout) as r:
raw = r.read().decode().strip()
return json.loads(raw) if raw else {}
def now():
return datetime.now().strftime("%H:%M:%S")
def spawn_worker(base, profile, repo):
res = http(base, "POST", "/workers", {"profile": profile, "cwd": repo} if profile
else {"cwd": repo})
tid = res.get("terminalId") or res.get("sessionId")
if not tid:
raise RuntimeError(f"spawn failed: {res}")
pane = res.get("paneId")
print(f"[{now()}] spawned {tid} (profile={profile or 'default'}, pane={pane})")
return tid, pane
def await_ready(base, tid, timeout=150):
deadline = time.time() + timeout
while time.time() < deadline:
try:
st = http(base, "GET", f"/sessions/{tid}/status")
except urllib.error.URLError:
st = {}
if st.get("ready"):
print(f"[{now()}] worker ready (status={st.get('status')})")
return True
time.sleep(3)
print(f"[{now()}] WARNING: worker never reported ready within {timeout}s — sending anyway")
return False
def post_message(base, tid, body, timeout):
"""One BLOCKING send. Returns (http_status, parsed_json). urllib raises on 4xx/5xx, so a
stale-turn 409 is surfaced here rather than swallowed."""
data = json.dumps(body).encode()
req = urllib.request.Request(f"{base}/sessions/{tid}/message", data=data, method="POST",
headers={"Content-Type": "application/json"})
try:
with urllib.request.urlopen(req, timeout=timeout) as r:
raw = r.read().decode().strip()
return r.status, (json.loads(raw) if raw else {})
except urllib.error.HTTPError as e:
raw = e.read().decode().strip()
return e.code, (json.loads(raw) if raw else {})
def run(base, profile, repo, send_timeout, answer_timeout):
"""Drive the full reverse rendezvous. Returns a result record for grading + transcript."""
rec = {"spawned": False, "tid": None, "pane": None, "phase": "spawn",
"question": None, "turnId": None, "reply": None, "replySource": None,
"detail": None, "ask_latency": None, "answer_latency": None}
tid, pane = spawn_worker(base, profile, repo)
rec.update(tid=tid, pane=pane, spawned=True)
await_ready(base, tid)
# 1) Delegate the ask-forcing task. This blocks until the worker calls bridge_ask, at which
# point our own send unblocks carrying the question and the turnId to answer on.
print(f"[{now()}] delegating task (blocks until the worker asks; up to {send_timeout}s)…")
t0 = time.time()
rec["phase"] = "awaiting_question"
try:
code, resp = post_message(base, tid, {"content": TASK_PROMPT, "timeoutMs": send_timeout * 1000},
timeout=send_timeout + 15)
except Exception as e: # noqa: BLE001
rec.update(phase="send_error", detail=f"delegating send failed: {e}")
return rec
rec["ask_latency"] = round(time.time() - t0, 1)
if resp.get("status") != "question":
# The worker finished (or stalled) without asking — the whole point didn't happen.
rec.update(phase="no_question", detail=f"HTTP {code}: {json.dumps(resp)[:300]}",
reply=resp.get("reply"), replySource=resp.get("replySource"))
return rec
rec.update(phase="question", question=resp.get("question"), turnId=resp.get("turnId"))
print(f"[{now()}] worker ASKED ({rec['ask_latency']}s): {rec['question']!r} turnId={rec['turnId']}")
if not rec["turnId"]:
rec.update(phase="no_turnid", detail="question surfaced without a turnId to answer on")
return rec
# 2) Answer on that exact turn. This blocks again until the resumed worker calls bridge_reply.
print(f"[{now()}] answering '{ANSWER_COLOR}' on turn {rec['turnId']} (blocks until reply; up to {answer_timeout}s)…")
t1 = time.time()
rec["phase"] = "awaiting_reply"
try:
code, resp = post_message(base, tid,
{"content": ANSWER_COLOR, "turnId": rec["turnId"],
"timeoutMs": answer_timeout * 1000},
timeout=answer_timeout + 15)
except Exception as e: # noqa: BLE001
rec.update(phase="answer_error", detail=f"answer send failed: {e}")
return rec
rec["answer_latency"] = round(time.time() - t1, 1)
if code == 409 or resp.get("error") == "stale_turn":
rec.update(phase="stale_turn", detail=f"HTTP {code}: {json.dumps(resp)[:300]}")
return rec
if resp.get("reply") is None:
rec.update(phase="no_reply", detail=f"HTTP {code}: {json.dumps(resp)[:300]}")
return rec
rec.update(phase="replied", reply=resp.get("reply"), replySource=resp.get("replySource"))
print(f"[{now()}] worker RESUMED and replied ({rec['answer_latency']}s): {rec['reply']!r} "
f"(source={rec['replySource']})")
return rec
def grade(rec):
"""PASS only if the worker asked, the turn resumed, and the reply reflects the answer."""
if rec["phase"] == "no_question":
return "NO_ASK", "the worker finished/stalled without ever calling bridge_ask"
if rec["phase"] in ("send_error", "answer_error", "spawn"):
return "ERROR", rec.get("detail") or "transport error before the round-trip completed"
if rec["phase"] == "no_turnid":
return "NO_TURNID", "the question surfaced without a turnId — the primary could not answer"
if rec["phase"] == "stale_turn":
return "STALE", "answering the turn was rejected as stale (it lapsed or was already answered)"
if rec["phase"] in ("awaiting_reply", "no_reply"):
return "NO_RESUME", "the worker asked but never resumed to a final reply within the window"
if rec["phase"] == "replied":
reflected = ANSWER_COLOR.upper() in (rec["reply"] or "").upper()
if reflected and rec["replySource"] == "reply":
return "OK", "asked, resumed the same turn, and the reply reflected the primary's answer"
if reflected:
return "DEGRADED", f"reply reflected the answer but resolved via {rec['replySource']} " \
"(worker did not call bridge_reply cleanly)"
return "WRONG_ANSWER", f"the worker replied but did not reflect '{ANSWER_COLOR}' — " \
f"the answer may not have reached the resumed turn: {rec['reply']!r}"
return "WEDGE", f"unexpected terminal phase {rec['phase']}: {rec.get('detail')}"
def write_transcript(out_dir, rec, meta):
path = out_dir / "bridge_ask_transcript.md"
g, note = grade(rec)
with path.open("w") as f:
f.write(f"# Live bridge_ask — reverse rendezvous — {datetime.now():%Y-%m-%d %H:%M}\n\n")
f.write(f"One worker paused its delegated turn to ask the primary, then resumed with the "
f"answer (profile `{meta['profile']}`). Result: **`{g}`**.\n\n")
f.write("## Round-trip\n\n")
f.write(f"1. **primary → worker** (delegation): the ask-forcing task.\n")
f.write(f"2. **worker → primary** (`bridge_ask`, {rec.get('ask_latency')}s): "
f"{rec.get('question')!r} — surfaced on the primary's blocked send as a "
f"`question` with `turnId={rec.get('turnId')}`.\n")
f.write(f"3. **primary → worker** (answer on that turn): `{ANSWER_COLOR}`.\n")
f.write(f"4. **worker → primary** (`bridge_reply`, {rec.get('answer_latency')}s, "
f"source={rec.get('replySource')}): {rec.get('reply')!r}\n\n")
f.write(f"> **{g}:** {note}\n")
if rec.get("detail"):
f.write(f">\n> detail: {rec['detail']}\n")
return path
def main():
ap = argparse.ArgumentParser(description="Live bridge_ask reverse-rendezvous test (CB-205)")
ap.add_argument("--base", default="http://127.0.0.1:8765")
ap.add_argument("--profile", default=None)
ap.add_argument("--repo", default=str(REPO_ROOT))
ap.add_argument("--out", default=str(HERE))
ap.add_argument("--send-timeout", type=int, default=110, help="max wait for the worker to ASK")
ap.add_argument("--answer-timeout", type=int, default=90, help="max wait for the worker to REPLY")
ap.add_argument("--keep-worker", action="store_true")
args = ap.parse_args()
print(f"[{now()}] live bridge_ask: 1 primary, 1 worker "
f"(profile={args.profile or 'default'}, repo={args.repo})\n")
rec = {"spawned": False, "pane": None}
try:
rec = run(args.base, args.profile, args.repo, args.send_timeout, args.answer_timeout)
finally:
if not args.keep_worker and rec.get("spawned") and rec.get("pane"):
try:
http(args.base, "DELETE", f"/workers/{rec['pane']}")
print(f"[{now()}] stopped worker (pane {rec['pane']})")
except Exception as e: # noqa: BLE001
print(f"[{now()}] stop failed (ignore): {e}")
g, note = grade(rec)
out_dir = pathlib.Path(args.out)
out_dir.mkdir(parents=True, exist_ok=True)
path = write_transcript(out_dir, rec, {"profile": args.profile or "default"})
print()
print("=" * 72)
print("LIVE bridge_ask SUMMARY — reverse rendezvous (CB-205)")
print(f" asked: {rec.get('question')!r} (turnId={rec.get('turnId')}, {rec.get('ask_latency')}s)")
print(f" answered: {ANSWER_COLOR!r}")
print(f" replied: {rec.get('reply')!r} (source={rec.get('replySource')}, {rec.get('answer_latency')}s)")
print(f" transcript: {path}")
print(f" RESULT: {g} — {note}")
print("=" * 72)
sys.exit(0 if g == "OK" else 1)
if __name__ == "__main__":
main()
+12
View File
@@ -0,0 +1,12 @@
# Live bridge_ask — reverse rendezvous — 2026-07-16 16:30
One worker paused its delegated turn to ask the primary, then resumed with the answer (profile `default`). Result: **`OK`**.
## Round-trip
1. **primary → worker** (delegation): the ask-forcing task.
2. **worker → primary** (`bridge_ask`, 6.6s): 'PICK A COLOR: red or blue?' — surfaced on the primary's blocked send as a `question` with `turnId=term_656bb47d2c42a9e#1`.
3. **primary → worker** (answer on that turn): `blue`.
4. **worker → primary** (`bridge_reply`, 7.9s, source=reply): 'CHOSEN=BLUE'
> **OK:** asked, resumed the same turn, and the reply reflected the primary's answer
+170
View File
@@ -0,0 +1,170 @@
#!/usr/bin/env python3
"""Sustained back-and-forth bridge test — ONE primary, ONE worker, many dependent turns
over a fixed wall-clock window (default 5 minutes), through the running `bridged` daemon.
Where conversation_test.py proves a handful of turns work and issue_hunt_test.py proves
fan-out isolation, this proves the channel stays healthy under a *sustained, stateful*
conversation: a running-total game the worker must keep in its head across turns. Turn N's
prompt does NOT restate the total — the worker has to remember it from turn N-1 — so a
correct answer is evidence of genuine multi-turn continuity, not just per-turn liveness.
Each reply is machine-checked against the primary's own expected total; on a drift the
primary re-anchors (states the correct total once) and keeps going, and drift is reported.
It talks ONLY to the bridge's REST face on loopback — it never sets ANTHROPIC_BASE_URL and
never touches herdr directly, so it is subscription-safe by construction.
Usage:
python3 conversation_sustained_test.py [--base URL] [--profile NAME] [--repo DIR]
[--duration SECS] [--turn-timeout SECS] [--keep-worker]
--base bridge REST base URL (default http://127.0.0.1:8765)
--profile worker profile (default: the daemon's default)
--repo cwd handed to the worker (default: the bridge repo root)
--duration wall-clock window, seconds (default 300 = 5 minutes)
--turn-timeout per-turn max wait, seconds (default 150)
--keep-worker do not stop the worker at the end
Exit code: 0 if every turn in the window resolved via a clean bridge_reply with no channel
break; 1 otherwise. A live per-turn log streams to stdout so the run can be watched.
"""
import argparse
import pathlib
import re
import sys
import time
# Reuse the exact REST primitives the other harnesses use (same daemon contract).
sys.path.insert(0, str(pathlib.Path(__file__).parent))
from issue_hunt_test import http, now, spawn_worker, await_ready, fire, poll # noqa: E402
HERE = pathlib.Path(__file__).parent
REPO_ROOT = HERE.parent
# The per-turn increments, cycled. Non-trivial and varied so the running total isn't a
# predictable multiple the worker could pattern-match without actually tracking it.
STEPS = [7, 3, 11, 5, 9, 4, 13, 6, 8, 2]
RULES = (
"Let's play a running-total game across several messages. The total starts at 0. "
"In each message I'll tell you to add a number; keep the running total yourself and "
"reply via bridge_reply with ONLY the current total as a plain integer — no words, no "
"punctuation, just the number. Do not restate the arithmetic. First move: add {step}."
)
NEXT = ("Add {step}. Reply via bridge_reply with only the new running total.")
REANCHOR = ("Let's re-sync — the running total is {total}. Now add {step}. Reply via "
"bridge_reply with only the new running total.")
def parse_int(reply):
"""Pull the worker's answer integer from its reply (last integer token wins)."""
if not reply:
return None
nums = re.findall(r"-?\d+", reply.replace(",", ""))
return int(nums[-1]) if nums else None
def one_turn(base, tid, prompt, turn_timeout):
"""Fire one prompt and block on its reply. Returns the poll record."""
t0 = time.time()
ticket = fire(base, tid, prompt)
rec = poll(base, ticket, "worker", t0, turn_timeout)
rec["ticket"] = ticket
return rec
def main():
ap = argparse.ArgumentParser(description="Sustained back-and-forth bridge test (1 primary, 1 worker)")
ap.add_argument("--base", default="http://127.0.0.1:8765")
ap.add_argument("--profile", default=None)
ap.add_argument("--repo", default=str(REPO_ROOT))
ap.add_argument("--duration", type=int, default=300)
ap.add_argument("--turn-timeout", type=int, default=150)
ap.add_argument("--keep-worker", action="store_true")
args = ap.parse_args()
mins = args.duration / 60
print(f"[{now()}] sustained back-and-forth: 1 primary <-> 1 worker for {args.duration}s "
f"(~{mins:.1f} min) profile={args.profile or 'default'} repo={args.repo}", flush=True)
tid, pane = spawn_worker(args.base, args.profile, args.repo, "worker")
await_ready(args.base, tid, "worker")
start = time.time()
expected = 0 # the primary's authoritative running total
reanchor = False # re-state the total next turn after a drift
turns, oks, drifts, breaks = 0, 0, 0, 0
latencies = []
print(f"[{now()}] --- conversation start (worker must keep the total in its head) ---\n", flush=True)
while time.time() - start < args.duration:
turns += 1
step = STEPS[(turns - 1) % len(STEPS)]
if turns == 1:
prompt = RULES.format(step=step)
elif reanchor:
prompt = REANCHOR.format(total=expected, step=step)
reanchor = False
else:
prompt = NEXT.format(step=step)
expected += step
el = round(time.time() - start)
print(f"[{now()}] turn {turns:>2} (t+{el}s) PRIMARY → add {step} (expect total {expected})", flush=True)
rec = one_turn(args.base, tid, prompt, args.turn_timeout)
got = parse_int(rec.get("reply"))
lat = rec.get("latency")
latencies.append(lat)
if rec.get("phase") != "done" or not (rec.get("reply") or "").strip():
breaks += 1
print(f"[{now()}] WORKER ✗ CHANNEL BREAK — phase={rec.get('phase')} "
f"source={rec.get('source')} detail={str(rec.get('detail'))[:100]} ({lat}s)\n", flush=True)
reanchor = True
continue
src = rec.get("source")
badge = "OK " if got == expected else "DRIFT"
if got == expected:
oks += 1
else:
drifts += 1
reanchor = True # re-sync the worker next turn
print(f"[{now()}] WORKER → {str(rec.get('reply')).strip()[:60]!r} = {got} "
f"[{badge}] via {src} ({lat}s)", flush=True)
if got != expected:
print(f"[{now()}] (expected {expected}; will re-anchor next turn)", flush=True)
print(flush=True)
dur = round(time.time() - start)
clean = sum(1 for lat in latencies if lat)
avg = round(sum(latencies) / len(latencies), 1) if latencies else 0
print("=" * 72)
print(f"SUSTAINED CONVERSATION SUMMARY — 1 primary <-> 1 worker over {dur}s (~{dur/60:.1f} min)")
print(f" turns: {turns}")
print(f" clean bridge_reply exchanges: {oks + drifts}/{turns} (channel breaks: {breaks})")
print(f" arithmetic correct (continuity held): {oks}/{turns} (drifts: {drifts})")
print(f" latency: avg {avg}s over {turns} turns")
ok = breaks == 0 and turns >= 2
if ok and drifts == 0:
print(" RESULT: PASS — every turn resolved via bridge_reply and the worker held the "
"running total across the whole window.")
elif ok:
print(f" RESULT: PASS (channel) — every turn resolved via bridge_reply for the full "
f"window; {drifts} arithmetic drift(s) (worker recovered after re-anchor).")
else:
print(" RESULT: FAIL — the channel broke on at least one turn (see CHANNEL BREAK above).")
print("=" * 72, flush=True)
if not args.keep_worker and pane:
try:
http(args.base, "DELETE", f"/workers/{pane}")
print(f"[{now()}] worker stopped (pane {pane})", flush=True)
except Exception as e: # noqa: BLE001
print(f"[{now()}] worker stop failed (ignore): {e}", flush=True)
sys.exit(0 if ok else 1)
if __name__ == "__main__":
main()
+227
View File
@@ -0,0 +1,227 @@
#!/usr/bin/env python3
"""Standard bridge conversation test — a multi-turn primary↔worker exchange through
the running `bridged` daemon, fully captured, with automatic gap analysis.
This is the repeatable form of the ad-hoc channel test that surfaced the CB-115 gaps
(herdr `unknown` misclassification, dirty completion scrape, workers not calling
bridge_reply). It drives a real off-subscription worker over the live gateway exactly
as a primary Opus session would (async fire-and-poll), records every turn, and grades
the channel.
It talks ONLY to the bridge's REST face on loopback — it never sets ANTHROPIC_BASE_URL
and never touches herdr directly, so it is subscription-safe by construction.
Usage:
python3 conversation_test.py [--base URL] [--profile NAME] [--tid TERMINAL_ID]
[--prompts FILE] [--out DIR] [--poll-timeout SECS]
[--keep-worker]
--base bridge REST base URL (default http://127.0.0.1:8765)
--profile worker profile to spawn (default: the daemon's default)
--tid reuse an existing worker (skips spawn + readiness wait)
--prompts newline-separated prompt file (default: the built-in script)
--out output dir for transcript.md (default: alongside this file)
--poll-timeout per-turn max wait, seconds (default 300)
--keep-worker do not stop a spawned worker at the end
Exit code: 0 if every turn delivered AND produced a usable reply; 1 otherwise (so it
is CI-usable). A per-turn and overall gap report is printed to stdout.
"""
import argparse
import json
import pathlib
import sys
import time
import urllib.error
import urllib.request
from datetime import datetime
HERE = pathlib.Path(__file__).parent
# A default conversation: short, varied turns (a question, a follow-up that needs the
# prior context, a tiny reasoning task, a meta-question, a close) — enough to exercise
# multi-turn delivery + reply on the channel without being a real coding workload.
DEFAULT_PROMPTS = [
"Hi! Quick check that our channel works. In one sentence, what are you and what model are you running?",
"Thanks. Now a small task: what is 17 * 23? Show just the number.",
"Good. Remembering that result, is it a prime number? Answer yes or no with a one-line reason.",
"Switching topic: name one thing that would make this bridge conversation feel more reliable to you as the worker.",
"That's all — please acknowledge and we'll wrap up.",
]
def http(base, method, path, body=None, timeout=20):
data = json.dumps(body).encode() if body is not None else None
req = urllib.request.Request(base + path, data=data, method=method,
headers={"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=timeout) as r:
return json.loads(r.read().decode())
def now():
return datetime.now().strftime("%H:%M:%S")
def spawn_worker(base, profile):
q = f"?profile={profile}" if profile else ""
res = http(base, "POST", f"/workers{q}")
tid = res.get("terminalId") or res.get("sessionId")
if not tid:
sys.exit(f"spawn failed: {res}")
pane = res.get("paneId")
print(f"[{now()}] spawned worker {tid} (profile={profile or 'default'}, pane={pane})")
return tid, pane
def await_ready(base, tid, timeout=120):
print(f"[{now()}] waiting for worker readiness (bridge MCP connect)…")
deadline = time.time() + timeout
while time.time() < deadline:
try:
st = http(base, "GET", f"/sessions/{tid}/status")
except urllib.error.URLError:
st = {}
if st.get("ready"):
print(f"[{now()}] worker ready (status={st.get('status')})")
return True
time.sleep(3)
print(f"[{now()}] WARNING: worker never reported ready within {timeout}s — running anyway")
return False
def run_turn(base, tid, turn, prompt, poll_timeout):
"""Fire one prompt async, poll the ticket to resolution, return a structured record."""
t0 = time.time()
sent = http(base, "POST", f"/sessions/{tid}/message", {"wait": False, "content": prompt})
ticket = sent.get("ticket")
samples = [] # (elapsed, phase, live_status)
reply = detail = source = phase = None
deadline = time.time() + poll_timeout
while time.time() < deadline:
time.sleep(3)
task = http(base, "GET", f"/tasks/{ticket}")
phase = task.get("phase")
live = (task.get("detail") or "").replace("worker ", "") if phase == "pending" else ""
samples.append((round(time.time() - t0, 1), phase, live))
if phase in ("done", "failed"):
reply = task.get("reply")
detail = task.get("detail")
source = task.get("replySource")
break
latency = round(time.time() - t0, 1)
# compress status samples into a transition string
trans, last = [], None
for el, ph, live in samples:
tag = live if ph == "pending" else ph
if tag != last:
trans.append(f"{tag}@{el}s")
last = tag
return {
"turn": turn, "prompt": prompt, "ticket": ticket, "phase": phase,
"source": source, "reply": reply, "detail": detail, "latency": latency,
"transitions": " → ".join(trans), "time": now(),
}
def grade(rec):
"""Classify a turn's outcome. Returns (grade, note)."""
phase, source, reply = rec["phase"], rec["source"], rec["reply"]
has_reply = bool(reply and reply.strip())
if phase == "done" and source == "reply" and has_reply:
return "OK", "clean explicit bridge_reply"
if phase == "done" and has_reply:
return "DEGRADED", f"resolved via {source} (worker did not call bridge_reply)"
if phase == "done" and not has_reply:
return "EMPTY", "turn completed but reply was empty"
if phase == "failed":
return "FAILED", f"worker turn failed: {(rec['detail'] or '').strip()[:120]}"
return "WEDGE", "never resolved within the poll window (delivery wedge or lost turn)"
def write_transcript(out_dir, records):
path = out_dir / "transcript.md"
with path.open("w") as f:
f.write(f"# Bridge conversation test — {datetime.now():%Y-%m-%d %H:%M}\n\n")
for rec in records:
g, note = grade(rec)
f.write(f"### Turn {rec['turn']} — {rec['time']} "
f"(`{g}`, {rec['latency']}s, phase={rec['phase']}, source={rec['source']})\n\n")
f.write(f"**PRIMARY:** {rec['prompt']}\n\n")
f.write(f"**WORKER:** {rec['reply'] if rec['reply'] else '_(no reply)_ ' + str(rec['detail'])}\n\n")
f.write(f"_status: {rec['transitions']}_\n")
if g != "OK":
f.write(f"\n> **GAP — {g}:** {note}\n")
f.write("\n")
return path
def main():
ap = argparse.ArgumentParser(description="Standard bridge conversation test")
ap.add_argument("--base", default="http://127.0.0.1:8765")
ap.add_argument("--profile", default=None)
ap.add_argument("--tid", default=None)
ap.add_argument("--prompts", default=None)
ap.add_argument("--out", default=str(HERE))
ap.add_argument("--poll-timeout", type=int, default=300)
ap.add_argument("--keep-worker", action="store_true")
args = ap.parse_args()
prompts = DEFAULT_PROMPTS
if args.prompts:
prompts = [ln.strip() for ln in pathlib.Path(args.prompts).read_text().splitlines()
if ln.strip() and not ln.startswith("#")]
spawned = False
tid = args.tid
pane = None
if not tid:
tid, pane = spawn_worker(args.base, args.profile)
spawned = True
await_ready(args.base, tid)
print(f"[{now()}] running {len(prompts)}-turn conversation on {tid}\n")
records = []
for i, prompt in enumerate(prompts, 1):
rec = run_turn(args.base, tid, i, prompt, args.poll_timeout)
g, note = grade(rec)
records.append(rec)
print(f"[turn {i}] {g:8} {rec['latency']:6}s phase={rec['phase']} source={rec['source']}")
print(f" status: {rec['transitions']}")
print(f" reply: {(rec['reply'] or '(none) ' + str(rec['detail'])).strip()[:200]}")
print(f" note: {note}\n")
out_dir = pathlib.Path(args.out)
out_dir.mkdir(parents=True, exist_ok=True)
path = write_transcript(out_dir, records)
# ---- gap report -------------------------------------------------------
grades = [grade(r)[0] for r in records]
counts = {g: grades.count(g) for g in ("OK", "DEGRADED", "EMPTY", "FAILED", "WEDGE") if grades.count(g)}
print("=" * 68)
print(f"CONVERSATION TEST SUMMARY — {len(records)} turns")
print(" " + " ".join(f"{g}:{n}" for g, n in counts.items()))
print(f" transcript: {path}")
ok = all(g in ("OK", "DEGRADED") for g in grades)
reply_clean = all(g == "OK" for g in grades)
if reply_clean:
print(" RESULT: PASS — every turn delivered and got a clean bridge_reply.")
elif ok:
print(" RESULT: PASS (with notes) — every turn delivered & replied, but some via fallback.")
else:
print(" RESULT: FAIL — one or more turns wedged, failed, or returned empty (see GAP notes).")
print("=" * 68)
if spawned and pane and not args.keep_worker:
try:
http(args.base, "DELETE", f"/workers/{pane}")
print(f"[{now()}] stopped worker {tid} (pane {pane})")
except Exception as e: # noqa: BLE001 - best-effort cleanup
print(f"[{now()}] worker stop failed (ignore): {e}")
sys.exit(0 if ok else 1)
if __name__ == "__main__":
main()
+330
View File
@@ -0,0 +1,330 @@
#!/usr/bin/env python3
"""Standard bridge fan-out test — ONE primary vs MANY workers, concurrently, for
issue hunting through the running `bridged` daemon, fully captured, with gap analysis.
Where conversation_test.py exercises a single worker over multiple turns, this drives
the path that only appears under fan-out: the primary spawns N workers, sends each a
distinct issue-hunting assignment on a slice of the codebase, fires them all at once,
and collects every reply concurrently. That stresses what a single worker never can —
• simultaneous delivery to many panes (the injector's per-worker, not global, writer),
• per-session rendezvous isolation (N blocked sends resolving independently),
• reply routing under concurrency (worker A's answer must never resolve worker B's send),
and, as the payload, whether a fleet of off-subscription workers can actually surface
real issues in the repo and report them back structurally via bridge_reply.
It talks ONLY to the bridge's REST face on loopback — it never sets ANTHROPIC_BASE_URL
and never touches herdr directly, so it is subscription-safe by construction.
Usage:
python3 issue_hunt_test.py [--base URL] [--profile NAME] [--repo DIR]
[--out DIR] [--poll-timeout SECS] [--keep-workers]
[--assignments FILE]
--base bridge REST base URL (default http://127.0.0.1:8765)
--profile profile for every worker (default: the daemon's default)
--repo cwd handed to each worker (default: the bridge repo root)
--out output dir for the transcript (default: alongside this file)
--poll-timeout per-worker max wait, seconds (default 300)
--keep-workers do not stop spawned workers at the end
--assignments JSON file overriding the built-in assignment list
Exit code: 0 if every worker delivered AND produced a usable reply with no cross-talk;
1 otherwise (CI-usable). A per-worker and fleet-level gap report is printed to stdout.
"""
import argparse
import json
import pathlib
import sys
import threading
import time
import urllib.error
import urllib.request
from concurrent.futures import ThreadPoolExecutor
from datetime import datetime
HERE = pathlib.Path(__file__).parent
REPO_ROOT = HERE.parent
# Each worker gets a distinct source file to hunt in, plus a `probe` — a token its reply
# should mention if it actually addressed ITS assignment (a soft cross-talk detector: a
# reply that references only another worker's file is a routing red flag). Targets are the
# hot files this project has been iterating on, so a real issue is plausible to find.
DEFAULT_ASSIGNMENTS = [
{"id": "completion", "probe": "CompletionResolver",
"target": "bridged/src/main/java/dev/ltms/bridged/inject/CompletionResolver.java"},
{"id": "worker", "probe": "WorkerService",
"target": "bridged/src/main/java/dev/ltms/bridged/worker/WorkerService.java"},
{"id": "rendezvous", "probe": "Rendezvous",
"target": "bridged/src/main/java/dev/ltms/bridged/msg/Rendezvous.java"},
]
PROMPT_TMPL = (
"You are one of several issue-hunting workers in the claude-bridge repo (it is your "
"current working directory). Your assignment: inspect the file `{target}` and find the "
"SINGLE most important real bug, correctness gap, or risk in it. Read the file before "
"answering. Reply via bridge_reply with EXACTLY these four lines:\n"
"1. {target}:<line>\n"
"2. issue: <one sentence>\n"
"3. fix: <one line>\n"
"4. severity: high|medium|low\n"
"Keep it under 90 words. If after reading you find nothing real, reply 'NO ISSUE' and one "
"line why. Do NOT hunt in any other file — only `{target}`."
)
def http(base, method, path, body=None, timeout=20):
"""JSON request. Tolerates an empty body (e.g. 204 No Content on DELETE) → returns {}."""
data = json.dumps(body).encode() if body is not None else None
req = urllib.request.Request(base + path, data=data, method=method,
headers={"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=timeout) as r:
raw = r.read().decode().strip()
return json.loads(raw) if raw else {}
def now():
return datetime.now().strftime("%H:%M:%S")
def spawn_worker(base, profile, repo, wid):
res = http(base, "POST", "/workers", {"profile": profile, "cwd": repo} if profile
else {"cwd": repo})
tid = res.get("terminalId") or res.get("sessionId")
if not tid:
raise RuntimeError(f"spawn failed for {wid}: {res}")
pane = res.get("paneId")
print(f"[{now()}] [{wid}] spawned {tid} (profile={profile or 'default'}, pane={pane})")
return tid, pane
def await_ready(base, tid, wid, timeout=150):
deadline = time.time() + timeout
while time.time() < deadline:
try:
st = http(base, "GET", f"/sessions/{tid}/status")
except urllib.error.URLError:
st = {}
if st.get("ready"):
print(f"[{now()}] [{wid}] ready (status={st.get('status')})")
return True
time.sleep(3)
print(f"[{now()}] [{wid}] WARNING: never reported ready within {timeout}s — sending anyway")
return False
def fire(base, tid, prompt):
"""Fire one async send; return its ticket (delivery is confirmed by a ticket coming back)."""
sent = http(base, "POST", f"/sessions/{tid}/message", {"wait": False, "content": prompt})
return sent.get("ticket")
def poll(base, ticket, wid, t0, poll_timeout):
"""Poll a ticket to resolution; return (record fields) mirroring conversation_test."""
samples, reply, detail, source, phase = [], None, None, None, None
deadline = time.time() + poll_timeout
while time.time() < deadline:
time.sleep(3)
task = http(base, "GET", f"/tasks/{ticket}")
phase = task.get("phase")
live = (task.get("detail") or "").replace("worker ", "") if phase == "pending" else ""
samples.append((round(time.time() - t0, 1), phase, live))
if phase in ("done", "failed"):
reply, detail, source = task.get("reply"), task.get("detail"), task.get("replySource")
break
trans, last = [], None
for el, ph, live in samples:
tag = live if ph == "pending" else ph
if tag != last:
trans.append(f"{tag}@{el}s")
last = tag
return {"phase": phase, "reply": reply, "detail": detail, "source": source,
"latency": round(time.time() - t0, 1), "transitions": " → ".join(trans)}
def grade(rec):
phase, source, reply = rec["phase"], rec["source"], rec["reply"]
has_reply = bool(reply and reply.strip())
if phase == "done" and source == "reply" and has_reply:
return "OK", "clean explicit bridge_reply"
if phase == "done" and has_reply:
return "DEGRADED", f"resolved via {source} (worker did not call bridge_reply)"
if phase == "done" and not has_reply:
return "EMPTY", "turn completed but reply was empty"
if phase == "failed":
return "FAILED", f"worker turn failed: {(rec['detail'] or '').strip()[:120]}"
return "WEDGE", "never resolved within the poll window (delivery wedge or lost turn)"
def worker_lifecycle(base, profile, repo, poll_timeout, a, barrier):
"""Full per-worker path: spawn → ready → (barrier) → fire → poll. Spawn+ready run
concurrently across workers; the send waits on the shared `barrier` so every ready
worker fires within the same instant — the real simultaneous-delivery stress. A worker
that fails to spawn aborts the barrier so the rest don't block forever."""
wid = a["id"]
rec = {"id": wid, "target": a["target"], "probe": a["probe"], "spawned": False,
"tid": None, "pane": None, "ticket": None}
try:
tid, pane = spawn_worker(base, profile, repo, wid)
rec.update(tid=tid, pane=pane, spawned=True)
await_ready(base, tid, wid)
except Exception as e: # noqa: BLE001
barrier.abort() # release peers waiting on the barrier
rec.update(phase="failed", detail=f"spawn/ready error: {e}", reply=None,
source=None, latency=0.0, transitions="")
return rec
try:
barrier.wait(timeout=210) # all ready workers proceed together
except (threading.BrokenBarrierError, Exception): # noqa: BLE001
pass # a peer died or timed out — fire anyway rather than hang
t0 = time.time()
prompt = PROMPT_TMPL.format(target=a["target"])
try:
ticket = fire(base, tid, prompt)
rec["ticket"] = ticket
print(f"[{now()}] [{wid}] fired (ticket={ticket})")
rec.update(poll(base, ticket, wid, t0, poll_timeout))
except Exception as e: # noqa: BLE001
rec.update(phase="failed", detail=f"send error: {e}", reply=None,
source=None, latency=round(time.time() - t0, 1), transitions="")
return rec
def crosstalk_report(records):
"""Fleet-level isolation checks: distinct tickets, distinct non-empty replies, and each
reply addressing its OWN assigned file (probe token present). Returns (list_of_gaps)."""
gaps = []
tickets = [r.get("ticket") for r in records if r.get("ticket")]
if len(tickets) != len(set(tickets)):
gaps.append("ticket collision: two workers were handed the same ticket id")
replies = {r["id"]: (r.get("reply") or "").strip() for r in records}
# identical non-empty replies from distinct assignments ⇒ suspected reply misrouting
seen = {}
for wid, text in replies.items():
if text and text in seen:
gaps.append(f"identical reply from '{seen[text]}' and '{wid}' "
f"(distinct assignments should not yield byte-identical answers)")
elif text:
seen[text] = wid
# a reply that names ANOTHER worker's file but not its own ⇒ likely cross-routing
for r in records:
text = (r.get("reply") or "")
if not text.strip():
continue
own = r["probe"] in text or pathlib.Path(r["target"]).name in text
others = [o["probe"] for o in records if o["id"] != r["id"] and o["probe"] in text]
if not own and others:
gaps.append(f"worker '{r['id']}' (assigned {r['probe']}) replied about "
f"{others} but not its own file — possible cross-routing")
return gaps
def write_transcript(out_dir, records, meta):
path = out_dir / "issue_hunt_transcript.md"
with path.open("w") as f:
f.write(f"# Bridge fan-out issue-hunt — {datetime.now():%Y-%m-%d %H:%M}\n\n")
f.write(f"One primary vs **{len(records)} concurrent workers** "
f"(profile `{meta['profile']}`), each hunting a distinct file.\n\n")
for r in records:
g, note = grade(r)
f.write(f"### `{r['id']}` — {r['target']} (`{g}`, {r.get('latency')}s, "
f"phase={r.get('phase')}, source={r.get('source')})\n\n")
f.write(f"**ASSIGNMENT:** find the top issue in `{r['target']}`\n\n")
reply = r.get("reply")
f.write(f"**WORKER {r['id']}:** {reply if reply else '_(no reply)_ ' + str(r.get('detail'))}\n\n")
if r.get("transitions"):
f.write(f"_status: {r['transitions']}_\n")
if g != "OK":
f.write(f"\n> **GAP — {g}:** {note}\n")
f.write("\n")
if meta["gaps"]:
f.write("## Fleet-level gaps\n\n")
for gp in meta["gaps"]:
f.write(f"- {gp}\n")
return path
def main():
ap = argparse.ArgumentParser(description="Bridge fan-out issue-hunt test (1 primary, N workers)")
ap.add_argument("--base", default="http://127.0.0.1:8765")
ap.add_argument("--profile", default=None)
ap.add_argument("--repo", default=str(REPO_ROOT))
ap.add_argument("--out", default=str(HERE))
ap.add_argument("--poll-timeout", type=int, default=300)
ap.add_argument("--keep-workers", action="store_true")
ap.add_argument("--assignments", default=None)
args = ap.parse_args()
assignments = DEFAULT_ASSIGNMENTS
if args.assignments:
assignments = json.loads(pathlib.Path(args.assignments).read_text())
n = len(assignments)
barrier = threading.Barrier(n) # releases exactly when all n ready workers reach it
print(f"[{now()}] fan-out issue-hunt: 1 primary vs {n} workers "
f"(profile={args.profile or 'default'}, repo={args.repo})\n")
# Spawn + ready + fire + poll all workers concurrently; the barrier makes every send fire
# together once all are ready, so delivery pressure hits the daemon simultaneously.
records = []
with ThreadPoolExecutor(max_workers=n) as ex:
futures = [ex.submit(worker_lifecycle, args.base, args.profile, args.repo,
args.poll_timeout, a, barrier) for a in assignments]
for fut in futures:
records.append(fut.result())
records.sort(key=lambda r: [a["id"] for a in assignments].index(r["id"]))
print()
for r in records:
g, note = grade(r)
print(f"[{r['id']:11}] {g:8} {str(r.get('latency','?')):6}s "
f"phase={r.get('phase')} source={r.get('source')}")
print(f" status: {r.get('transitions') or '(none)'}")
print(f" reply: {((r.get('reply') or '(none) ' + str(r.get('detail'))).strip()[:200])}")
print(f" note: {note}\n")
gaps = crosstalk_report(records)
meta = {"profile": args.profile or "default", "gaps": gaps}
out_dir = pathlib.Path(args.out)
out_dir.mkdir(parents=True, exist_ok=True)
path = write_transcript(out_dir, records, meta)
grades = [grade(r)[0] for r in records]
counts = {g: grades.count(g) for g in ("OK", "DEGRADED", "EMPTY", "FAILED", "WEDGE") if grades.count(g)}
print("=" * 72)
print(f"FAN-OUT ISSUE-HUNT SUMMARY — 1 primary vs {n} workers")
print(" channel: " + " ".join(f"{g}:{v}" for g, v in counts.items()))
print(f" transcript: {path}")
if gaps:
print(" FLEET GAPS:")
for gp in gaps:
print(f" ⚠ {gp}")
else:
print(" isolation: clean — distinct tickets, distinct replies, each on its own file")
delivered = all(r.get("ticket") for r in records)
replied = all(g in ("OK", "DEGRADED") for g in grades)
if delivered and replied and not gaps:
print(" RESULT: PASS — all workers delivered concurrently, replied, and stayed isolated.")
elif delivered and replied:
print(" RESULT: PASS (with notes) — all delivered & replied, but see FLEET GAPS.")
else:
print(" RESULT: FAIL — a worker did not deliver or did not reply (see GAP notes).")
print("=" * 72)
if not args.keep_workers:
for r in records:
if r.get("spawned") and r.get("pane"):
try:
http(args.base, "DELETE", f"/workers/{r['pane']}")
print(f"[{now()}] [{r['id']}] stopped (pane {r['pane']})")
except Exception as e: # noqa: BLE001
print(f"[{now()}] [{r['id']}] stop failed (ignore): {e}")
sys.exit(0 if (delivered and replied and not gaps) else 1)
if __name__ == "__main__":
main()
+1 -1
Submodule wiki updated: 8c37bb71c1...8e5fd01ac9