Compare commits

...

94 Commits

Author SHA1 Message Date
kevin 871b595954 CB-518: state the primary's orchestration as an explicit, ordered flow
CI / build (pull_request) Successful in 1m23s
CB-517 moved orchestration policy into CLAUDE.md but left it as a bullet
list, so the procedure was implicit: the order of operations had to be
reconstructed from a parallelisation bullet, and nothing said when to
review or when to tear down. A policy you have to reassemble on each
task is one you will reassemble differently each task. Restate the
primary's half as a numbered 0-8 flow — role check, split, gate, spawn
all, send all, collect, verify, review, adjudicate — so that following
it is checkable against the tool calls rather than a matter of recall.

Two steps carry the load. Spawn and send are separate on purpose:
folding them into one loop is what silently serialises work that was
meant to fan out. And review is now its own step ahead of the merge
rather than a clause inside it, because the two have opposite owners —
reviewers fan out over the diff (never the implementer of the scope
they review, and briefed from the diff rather than the author's
rationale, which carries the same blind spot), while adjudication, the
merge and teardown stay with the primary. Merging on a reviewer's word
is delegating the gate by proxy, so the step says so outright.

Nothing is dropped. The six bullets that trailed the tool table are
relocated into the step that owns each — profile explicitness into
spawn, playbook naming and self-containment into send, claim
verification into its own step, the ~60s blocking-send cap into a note
beneath the flow — and the table stays as the intent→tool lookup.

The wiki pointer moves with it. The block is canonical only if its
template matches byte for byte, so the template was produced by
splicing the block out of CLAUDE.md rather than by editing it in
parallel, and the sync check the repo documents passes. Bumping the
pointer in the same commit keeps charter and template versioned
together, as CB-517 did.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw1FosEt3Noix5GG9wJ2r
2026-08-04 22:32:19 +07:00
Dai Ha b67b1585c2 CB-517: deploy LavinMQ as a pinned, durable, self-restarting broker
CI / build (push) Successful in 1m23s
The CB-307 durable ReplyInbox needs an AMQP broker, but the one behind it
was run ad hoc and had simply vanished from the host — which takes the
whole daemon with it, since AmqpReplyInbox.open throws and Bridged.java:187
does not guard it. A missing broker is a hard startup failure, not a
degraded mode, so 'how the broker runs' is part of the system, not a local
detail.

Pinned to 2.9.1 (:latest would move the broker under a running daemon),
data on a named volume (held-but-unacked replies are the entire point of
Stage 2 — a plain 'compose down' would discard exactly what durability
protects), and restart: unless-stopped so it comes back after a reboot
instead of disappearing again.

Ports are bound to 127.0.0.1 deliberately: LavinMQ ships a default
guest/guest account, which is only acceptable while nothing off-host can
reach it.

Verified by driving the production AmqpReplyInbox against this deployment
(publish/peek/dedup/FIFO/ack, then reconnect): 8/8 including redelivery of
the unacked message. That pairing had never been exercised — the
@Tag("contract") test runs against a RabbitMQ container, and is excluded
from the default build, so mvn clean install covers the broker path zero
times.
2026-08-04 16:17:21 +02:00
Dai Ha 979b2b5632 CB-517: add bridge_whoami and make the bridge prompt a portable charter
CI / build (push) Successful in 1m25s
The communication rules lived only in two opt-in skills, so nothing
always-on told the primary how to orchestrate and nothing guaranteed a
worker loaded its playbook. Move protocol and policy into CLAUDE.md,
which a worker inherits for free (its worktree is a checkout of this
repo), and leave the skills as pure per-job procedure.

bridge_whoami closes the load-bearing gap: every tool already consumed
the caller identity ConnectionIdentity resolves from the connection, but
none reported it, so an agent had to infer its own role from side
channels the daemon does not control. Guessing fails asymmetrically — a
primary acting as a worker is refused by the authz gate and learns at
once, while a worker acting as the primary ends its turn without
bridge_reply and the sender silently receives nothing. The tool reuses
the same Principal the gate is built on, so the two cannot disagree; the
primary gets role only (handing it a sessionId it does not own would
invite the forged reply Authz refuses), and a worker missing from the
registry still gets role + sessionId rather than 'unknown'.

The CLAUDE.md block is written to be copied as-is into any project that
mounts the bridge: repo-local details (Authz paths, the .mcp.json/wiki
exclusions, the skill names) moved below it into a project addendum, and
every role-inference fallback is stated one-way — the mount-name signal
only holds for mcp__bridge__* (the launcher fixes it), not for the
primary's mount, which each project names itself. The wiki carries the
block verbatim as the template, with a sync check.

Because this repo IS the bridge, that block is shipped surface, not
documentation: the addendum adds a mandatory checklist mapping each part
of the code to the part of the prompt it can invalidate.

Also: delegate-by-default policy for the primary — the test is not 'could
I do this faster myself' but 'can I write a brief good enough for a
worker'.

mvn clean install: 356 tests green (353 + 3 for whoami); ide_diagnostics
clean on both changed files.
2026-08-04 16:01:16 +02:00
kevin cf4ad186ab CB-516: fail a delegation when its worker session is released
CI / build (push) Successful in 1m30s
Tearing a worker down left the send that was waiting on it stranded. The
rendezvous waiter stayed open, so a blocking bridge_send kept blocking and an
async one kept reporting PENDING until ASYNC_TIMEOUT_MS — thirty minutes —
even though the worker provably no longer existed and the delegation could
never complete.

Observed repeatedly this session: a task sitting at
{"phase":"pending","detail":"worker unknown"} long after I had personally
deleted the pane. The maddening part is that poll() already HAD the evidence —
it calls liveStatus() to build that detail string, gets back "unknown", and
reports PENDING anyway. It also never reached /metrics under any outcome, so a
stalled delegation was invisible to both the task view and the dashboard. The
only way I ever diagnosed one was reading the worker's pane over the herdr
socket by hand.

Fixed at the choke point rather than by guessing from status strings:
SessionManager.release() is the single path every teardown funnels through
(REST stop, MCP stop, idle-TTL reaper, recycle, shutdown drain), so it now
notifies a release listener with the terminalId, and Bridged.main wires that to
MessageService.abandon(). Abandon fails the open waiter with a real reason.

Deliberately NOT done by inferring "gone" from liveStatus(): that method
collapses a vanished pane, a wedged worker and a herdr hiccup into the same
"unknown" string, so acting on it would fail live delegations during a
transient blip. An explicit lifecycle signal cannot be ambiguous.

Notified before launcher.stop() so a blocked caller fails fast, and wrapped so
a listener failure can never prevent the teardown it is reacting to.

Resolving as a failure rather than letting it time out also means the outcome
is counted — a torn-down delegation now shows up as
sends_total{outcome="failed"} instead of nothing at all.

353 tests (was 346): abandon fails a waiting send / is a no-op with no waiter /
never clobbers a send the worker already answered, an abandoned async task
polls FAILED rather than PENDING, release notifies with the right terminal,
releasing an unknown pane notifies nobody, and a throwing listener does not
block teardown.

Verified live on the running daemon, reproducing the original scenario:
  async send        -> {"phase":"pending","detail":"worker working"}
  DELETE the worker -> 204
  poll              -> {"phase":"failed","detail":"the worker session was
                        released before it replied"}
  /metrics          -> bridged_sends_total{outcome="failed"} 1
Previously that poll returned PENDING for thirty minutes and the counter stayed
empty.

NOT addressed here, and worth deciding separately: a worker that simply never
replies (rather than being released) still rides out the full 30-minute
ASYNC_TIMEOUT_MS. That is a policy question about long-running delegations, not
a correctness bug.
2026-08-01 23:57:45 +07:00
kevin 5100f215cf CB-515: regression-protect the turn-attribution guards
CI / build (push) Successful in 1m33s
Five tests pinning the invariants that decide WHICH turn a reply belongs to.
These protect against a silent correctness bug — a reply attributed to the
wrong turn — not against a crash, which is why they were worth picking over
higher-percentage coverage gaps.

Chosen by blast radius, not by uncovered-line count. Both guards are compound
conditions with a side that never executed, i.e. exactly the shape where a
clause can be deleted as "redundant" and every existing test still passes.

CompletionResolver:
- The CB-115 misattribution guard suppresses a completion when the scrape is
  byte-identical to the pane at delivery. Its !scrapeFailed clause was
  unexercised: delete it and a FAILED read is misread as "no output change",
  so the send is suppressed and hangs to the caller's timeout instead of
  resolving. The new test sets the baseline to "" so the empty tail from a
  failed read would byte-match and wrongly suppress — built to die precisely
  when that clause dies.
- The fail() guard leaves an already-resolved waiter alone. The new test also
  asserts agent.read is never called, so the worker is not scraped for a send
  nobody is waiting on.

Rendezvous: a second resolution of an already-completed waiter returns false
and does not overwrite the first value, for both resolveCompletion and
resolveFailure.

Verified by sabotage, one guard at a time: removing !scrapeFailed reds
resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent; removing the
isDone() clause reds failLeavesAnAlreadyResolvedWaiterUntouchedAndSkipsTheScrape.
(The first attempt at the second sabotage reported a false pass — the patch hit
an identically-worded guard earlier in the file. Line-targeted and re-run.)

346 tests, was 341. Worker-implemented on the local-vLLM profile; it noticed
three of the eight cases I asked for already existed and said so with names
rather than duplicating them.

Also of note: the first delegation of this ticket wedged the worker — the pane
showed a zsh parse error and it went idle with an untouched worktree, task stuck
pending. The retry differed only in phrasing the same requirements as prose
instead of quoting Java boolean expressions. Filed as a bridge robustness
concern: injected content shares a channel with control, and a wedged turn is
invisible in both the task view and /metrics.
2026-08-01 23:48:05 +07:00
kevin f129e9b7cd Merge CB-514: MessageService coverage for timeout, answer, poll and lock edges
CI / build (push) Successful in 1m29s
Six tests on the delivery core, the most load-bearing class in the project,
which sat at 74% with its hardest paths unexercised: answer() (the CB-205
ask-answer resolution) had 11 of 23 lines uncovered, send()'s timeout branches
7 of 21, poll() 7 of 18, and tryLock 3 of 4.

Covered: TIMED_OUT_QUEUED vs TIMED_OUT_WORKING (undelivered vs delivered-but-
silent), an answered worker that never sends its follow-up reply, an unknown
ticket, a completed async ticket reporting its reply and replySource, and a
second concurrent send to the same session returning BUSY rather than hanging.

Worker-implemented on the local-vLLM profile, self-verified: it reported
'Tests run: 341, Failures: 0' and an independent run of its branch agrees
exactly. Additive only — 101 lines in one test file, no main/ source touched.

Quality is good on its own terms, not just green: the async-completion test
polls to a deadline instead of sleeping and hoping, the contention test
releases the blocked send so the test thread is not left pinned, and every
assertion is on a specific Outcome rather than 'nothing threw' — the failure
mode an earlier worker produced in CB-510.

Branch pushed by the worker; PR left to the primary since GITEA_TOKEN is not
granted to that profile by design.
2026-08-01 23:26:54 +07:00
kevin 4aed45de19 CB-514: add MessageService coverage for timeout, answer, poll, and lock edges 2026-08-01 23:24:54 +07:00
kevin 1e6daa5c73 CB-513: test the MCP-side authorization gate (BridgeMcp 27.4% -> 57.4%)
CI / build (push) Successful in 1m13s
CB-505 claimed authorization is "enforced on both entry paths". It is — but
only REST was ever tested. Coverage showed BridgeMcp.deny(), principal(),
callerTerminal(), worktreeRequest() and every tool-registration lambda at ZERO
executed lines: no test had ever constructed a BridgeMcp, because the existing
BridgeMcpTest calls only the static handler methods. So the MCP half of the
security control had ten REST tests' worth of nothing behind it.

An unexercised security control is a claim, not a control.

Made testable by separating policy from plumbing rather than by reaching for a
mocking library the project does not use:
- denyFor(Principal, Action, target) is the decision — testable directly.
- deny(exchange, ...) shrinks to pulling the caller out of the SDK exchange.
- principalFrom(role, terminal, pid) extracts identity reconstruction from
  McpSyncServerExchange, an SDK type with no fake available.

Moved the `authz == null` enforcement switch OUT of the exchange-facing wrapper
and INTO denyFor. Found by a failing test: as written, any future tool calling
denyFor directly would have silently skipped the gate. The switch now lives with
the decision it governs.

New BridgeMcpAuthzTest constructs a real BridgeMcp — which is why coverage moved
so far, since that also runs the constructor and all the tool wiring — and pins
the table on this path: primary orchestrates, worker cannot; worker replies only
as itself; the primary cannot forge a worker reply; anonymous gets nothing; and
401-shaped vs 403-shaped refusals are counted apart.

Verified as real controls, not decoration: with the gate forced open, 5 of the 9
fail. 335 tests (was 326).
2026-08-01 23:11:56 +07:00
kevin 4c015d76b7 Merge CB-512: wire bridged_push_nudges_total (worker-implemented, self-verified)
CI / build (push) Successful in 1m30s
Fixes one of the three counters declared in BridgedMetrics but never
incremented, so bridged_push_nudges_total{outcome=delivered|exhausted} now
actually appears on /metrics. An absent series reads as 'no push failures ever'
rather than 'not measured', which is the misleading case.

Implemented end to end by an opencode worker on the local-vLLM profile in an
isolated worktree, and this is the first delegation where the worker verified
its own work: it ran mvn, hit a real test failure, iterated, and reported
'Tests run: 326, Failures: 0' — which matches an independent run of its branch
exactly. Every earlier delegation reported results it had no way to check,
because CB-511 had not yet given workers a PATH with a toolchain.

It also committed and pushed its own branch unprompted, and was straight about
the one thing it could not do: opening the PR, since GITEA_HOST/GITEA_TOKEN are
not granted to that profile ('URL rejected: No host part'). That grant is opt-in
per profile by design, so the merge is primary-side as intended.

Diff reviewed and correct on every constraint, including the subtle ones:
delivered counted inside the try after a successful send (not the catch),
exhausted only on the reminder-cap branch and not the other two STOP paths, and
a single Metrics instance moved above pushLoop and shared with MessageService.
2026-08-01 23:01:07 +07:00
kevin 37a11cd168 CB-512: wire bridged_push_nudges_total metric increments 2026-08-01 22:56:32 +07:00
kevin 22ad24db6c CB-511: give workers a toolchain — propagate the daemon PATH, add profile env:
CI / build (push) Successful in 1m17s
Workers could not run `mvn` or `java`. Every delegated task that asked for a
build came back "mvn is not on PATH", and the worker was right.

Root cause: HerdrPeerLauncher seeded the worker environment with an EMPTY map,
so bridged passed only the vars it explicitly set (OPENCODE_CONFIG, GITEA_TOKEN,
ANTHROPIC_*) and never PATH. herdr merges that map into its own process env, so
a worker inherited whatever PATH the herdr SERVER was started with. On this host
that server (pid 79870, PPID 1) had been up since Jul 4 with a PATH containing
neither the JDK nor Maven. Confirmed on a live worker: its PATH was byte-identical
to herdr's, and the only var bridged had contributed was OPENCODE_CONFIG.

The failure was invisible and non-deterministic: the fleet's capabilities
depended on how a long-lived daemon happened to be launched weeks earlier. There
are three herdr processes on this box with three different PATHs; the one owning
the socket is the one without a toolchain. bridged itself HAD Maven on PATH the
whole time — it just never passed it on.

It also quietly contradicted the project's own principle that "a worker is a
full peer of the primary", and the implementer skill's instruction to build,
commit and open a PR. Every delegation so far has depended on the primary
running the build gate.

Fix: baseEnv(cfg) seeds each worker with the daemon's own PATH, then applies the
profile's new optional env: map. Adapter-specific vars are layered on top and
therefore win — that ordering is load-bearing, not incidental: it stops an env:
entry from overwriting ANTHROPIC_BASE_URL and slipping past SubscriptionGuard,
which is checked against the profile's baseUrl alone. Pinned by a test.

Because the default is now the daemon's PATH, both supervision units set PATH
explicitly — launchd and systemd do not source a login shell, so under CB-504
the daemon (and every worker) would otherwise get a bare /usr/bin:/bin and this
bug would silently return in production.

324 tests (was 321): daemon-PATH propagation, profile env: passthrough including
an explicit PATH override, and the guard-bypass ordering.

Verified live: daemon restarted, worker spawned, and asked to run the tools —
"Apache Maven 3.9.16", "java version 25.0.2". Previously both were absent.
2026-08-01 22:34:30 +07:00
kevin e94c1b8841 CB-510: SessionReaper wrapper tests (0% -> 86.7%)
SessionReaper had no tests at all. Its TTL *policy* was already well covered
(SessionManager.reapIdle, 6 cases in SessionManagerTest); what was untested was
the thread wrapper around it — idempotent start/stop and whether the loop
actually runs and actually stops.

Observed through an injected clock rather than by sleeping and hoping: reapIdle
reads nowNanos exactly once per call, so the tick count IS the iteration count.
Waits are bounded polls, not fixed sleeps, and nothing asserts an exact
timing-derived number — flaky counts would be worse than no test.

321 tests (was 318); line coverage 66.9% -> 67.9%.

Drafted by an opencode worker on the new local-vLLM profile (branch
worker/cb-510-session-reaper-test-cd1793-1). Its structure and setup were good
and it was honest that it could not run mvn. But its third test asserted
NOTHING — it started the reaper, slept, stopped it, and relied on "no throw",
with a comment claiming that proved the loop had run. It did not: verified by
sabotage, all three of its tests passed against a start() replaced with an
immediate return.

Rewritten so the assertions can fail for the right reason. Same sabotage now
fails 2 of 3 (the third only pins stop()-before-start(), where "does not throw"
genuinely is the contract). Uncomfortably on the nose given this task began as
a hunt for tests that do not mean anything.
2026-08-01 21:42:06 +07:00
kevin cc0ec65714 CB-509: add JaCoCo coverage reporting
Build-time tooling only — never a compile or runtime dependency, so it adds
nothing to the shipped jar and no new transitive surface to the artifact.
(Noting per CLAUDE.md that the pom CVE gate could not be run: no JetBrains MCP
server is connected this session.)

Report at target/site/jacoco/index.html, machine-readable at jacoco.csv.

Deliberately NO check rule or threshold. A coverage gate rewards writing tests
that merely execute lines, which is the exact failure mode this codebase has
already been bitten by — CB-507 shipped a null-argument NPE with 311 green
tests because FakeWorktrees.repoRoot records its argument instead of shelling
out, so the broken line was covered and still wrong. Coverage is a map of where
to look, not a target to hit.

Baseline: 66.9% line, 60.6% branch, 76.4% method.
2026-08-01 21:34:59 +07:00
kevin d67d30c58a CB-508: let an opencode profile pin its own OpenAI-compatible endpoint
CI / build (push) Successful in 1m57s
Points an opencode worker at a local vLLM (or llama.cpp / LM Studio / TGI)
instead of opencode's own gateway. opencode has no ANTHROPIC_BASE_URL seam, so
this could not be a config-only change: setting baseUrl on a kind: opencode
profile now makes the launcher emit a custom `provider` block into the
generated opencode.json, using @ai-sdk/openai-compatible.

The provider id comes from the provider half of the model: selector, so one
field drives both the generated declaration and the -m flag and the two cannot
drift apart. A bare model name with a baseUrl set is rejected at spawn with a
message saying how to fix it — silently falling back to the default gateway
would leave a worker talking to the wrong LLM while looking perfectly healthy.

A bare host:port gets /v1 appended (where these servers mount the API); a URL
that already carries a path is used verbatim. tokenEnv, when set, becomes the
provider apiKey; local servers generally ignore it but the AI SDK requires a
non-empty value, so a placeholder is used otherwise.

Two supporting changes:
- writeConfig previously ran only when a bridge MCP url was set. A pinned
  endpoint needs the config file too, so it now runs when either applies, and
  the mcp/instructions half is emitted conditionally.
- The config is now built with Jackson instead of string concatenation. The
  provider block is nested and interpolates operator-supplied values (URL,
  model id, api key), so escaping has to be real rather than a hand-rolled
  two-character replace.

No guard entry is required even with baseUrl set. SubscriptionGuard exists to
stop a worker borrowing the primary's Anthropic subscription, and an opencode
process has no Anthropic credential path at all — the asymmetry with the Claude
adapter reusing the same field is deliberate and documented at the call site.

Also fixes a brittle assertion in the existing MCP-mount test, which matched
the substring "\"type\": \"remote\"" and broke on Jackson's spacing. It now
parses the generated JSON and asserts on structure; whitespace is the
formatter's business, not the contract's.

318 tests (was 311): 5 new covering provider generation, /v1 normalisation,
path-preserving URLs, the missing-prefix rejection, MCP+provider coexistence,
and that no baseUrl still means no provider block.

Verified live end to end: daemon restarted on this build, worker spawned on the
opencode-local profile, generated config carries baseURL
http://127.0.0.1:8000/v1, and a blocking bridge_send returned
{"reply":"LOCAL-OK","replySource":"reply"} — a structured reply, not the
completion fallback. The worker pane reports
"Build · deepseek-v4-flash local-vllm (bridged)", confirming traffic reached
the local server rather than silently falling back.
2026-08-01 20:03:00 +07:00
kevin d75ee1cca5 CB-507: regression tests for worktree cwd resolution
Two cases in WorktreeSessionManagerTest, covering the gap that let the NPE ship
(313 tests, was 311).

1. worktreeAcquireWithNoRequestedOrCallerCwdStillResolvesANonNullRepoRoot —
   the null/null case a plain REST spawn produces.
2. worktreeAcquireHonoursTheProfileConfiguredCwd — the quieter second bug on
   the same line, where a pinned per-profile cwd: was ignored entirely.

Both assert on the cwd RECORDED by FakeWorktrees rather than expecting a throw.
That is deliberate: FakeWorktrees.repoRoot only records its argument and returns
a canned root, so a null passes through the fake harmlessly while the real
GitWorktrees runs `git -C null` and NPEs. The fake being more permissive than
the real seam is exactly why 311 tests stayed green over a broken feature —
asserting "an exception was raised" would be untestable here and would give
false confidence.

Verified as genuine regressions, not tautologies: with the pre-CB-507
expression restored both fail, with the messages they were written to give
(expected: not <null>, and expected </pinned/dir> but was <null>). Restored
after.

Drafted by an opencode-free worker over the bridge in an isolated worktree
(branch worker/cb-507-regression-test-11591f-4). Its test 1 was correct as
written. Test 2 was wrong and went red: it passed "/pinned/dir" as the 4th
constructor argument, which is configDir, not cwd (the 11th, after mcpUrl), so
cwd stayed null and the chain fell through to the daemon cwd. Corrected on
integration, along with removing two unused locals and adding the rationale
comments.
2026-08-01 19:50:31 +07:00
kevin 6804676a96 CB-507: fix NPE on a worktree spawn with no cwd (HTTP 500 over REST)
POST /workers?worktree=true returned HTTP 500 with a NullPointerException out
of ProcessBuilder.start(): acquireWithWorktree resolved the repo root from
firstNonBlank(requestedCwd, callerCwd), and a plain REST spawn supplies
neither (BridgedApp hardcodes callerCwd=null, "no MCP caller over REST"). Both
null yielded null, putting `git -C null rev-parse --show-toplevel` on the
command line.

Now resolved through launcher.effectiveCwd, the CB-112 chain used everywhere
else (requested -> profile cwd -> caller -> daemon cwd -> "."), which is
documented never to return null. The non-worktree path in this same class
already went through it; only the worktree branch was missed.

Also fixes a second latent bug in the same line: firstNonBlank never consulted
the profile's configured cwd:, so a worktree spawn silently ignored a pinned
per-profile working directory. effectiveCwd honours it.

Removes firstNonBlank, now dead (this was its only call site) — javac ignores
an unused private method but IDE inspections flag it, and CLAUDE.md requires a
clean bill.

Why 311 tests missed it: the null/null case only arises over REST, and
WorktreeSessionManagerTest always passes an explicit cwd. Over MCP callerCwd is
populated from the caller PID, so the feature worked there. This is the third
REST-vs-MCP divergence found this month, after CB-505's path-trusted session id.

The one-line change was implemented by an opencode-free worker over the bridge
in an isolated worktree (branch worker/cb-507-worktree-cwd-npe-3e9c3b-3); the
dead-helper cleanup and the explanatory comment were added on integration.
A regression test is still outstanding and is being delegated separately.
2026-08-01 19:43:08 +07:00
kevin 3e5d742ac7 CB-506: keep the test suite out of the production audit log
main/resources/logback.xml routes the `audit` logger to a RollingFileAppender
at logs/audit.log — the CB-505 security trail. AuditLogTest and
BridgedAppAuthTest exercise that same logger, so every `mvn test` appended
fabricated records to the production file.

They are byte-identical to genuine ones: runs of denied/forbidden
SPAWN/STOP/SEND from worker:term_a, which read exactly like an intrusion
attempt. logs/audit.2026-07-29.0.log is 38 fabricated records out of 76 — half
that day's security log is test fixtures, and nothing distinguishes them.

Fix is one new file, src/test/resources/logback-test.xml: logback prefers it on
the test classpath, so tests get a console-only config with no file appender
and main/resources/logback.xml is untouched. The `audit` logger stays ENABLED
(INFO, additivity=false) because AuditLogTest attaches its own ListAppender and
asserts on emitted records — setting it OFF would have silently gutted those
assertions.

Verified: 311 tests green, and logs/audit.log line count is identical before
and after a full `mvn clean install` (zero new records).

Implemented by an opencode-free worker over the bridge in an isolated worktree
(branch worker/cb-506-audit-test-isolation-e4aa9c-2); it correctly reported it
could not run mvn rather than fabricating a result, so the build gate and the
before/after audit-count check were run primary-side. Header comment added on
integration.
2026-08-01 19:20:29 +07:00
kevin 6da2a71050 CI: drop upload-artifact — unsupported on this Gitea instance
CI / build (push) Successful in 1m44s
Run 2 built clean (311 tests, BUILD SUCCESS, Maven 3.6.3 on Java 25.0.4 —
JAVA_HOME from setup-java correctly beat the JRE apt pulled in) but the job
still went red on the artifact step:

  GHESNotSupportedError: @actions/artifact v2.0.0+, upload-artifact@v4+ and
  download-artifact@v4+ are not currently supported on GHES.

Gitea Actions presents as GHES, so v4 artifact upload cannot work here. The
artifact was unretrievable regardless, so replace it with a failure-only step
that cats the failing surefire .txt reports into the job log, where they are
readable. Guarded with 'exit 0' so the dump itself can never mask the real
failure.
2026-07-29 23:26:21 +07:00
kevin e32ac39faf CI: install Maven — setup-java provides the JDK only
CI / build (push) Failing after 2m1s
First CI run failed at 'Build and test' with exit code 127 (command not
found): actions/setup-java@v4 provisions a JDK but not Maven, and the runner
image has no mvn on PATH. The sibling lms/alms workflow apt-installs both;
this workflow switched to setup-java for JDK 25 (the image's default-jdk is
too old for maven.compiler.release=25) and dropped the maven install with it.

Adds an explicit Maven install with Apache Maven 3.9.16 (2bdd9fddda4b155ebf8000e807eb73fd829a51d5)
Maven home: /Users/appbuilder/Tool/apache-maven-3.9.16
Java version: 25.0.2, vendor: Oracle Corporation, runtime: /Users/appbuilder/Tool/jdk-25.0.2.jdk/Contents/Home
Default locale: en_VN, platform encoding: UTF-8
OS name: "mac os x", version: "26.5.1", arch: "aarch64", family: "mac" so the log proves which JDK
it resolved — JAVA_HOME from setup-java must win over the JRE apt drags in.
2026-07-29 23:21:55 +07:00
kevin daa243d37a example config: opencode profile uses the verified zero-credential free tier
CI / build (push) Failing after 2m21s
The commented CB-402 block suggested google/gemini-2.5-pro, which needs
credentials. Replaced with the opencode/*-free gateway models proven during the
dogfood to work with no auth at all, and noted that the free model names change
so `opencode models` is the source of truth.
2026-07-29 22:49:58 +07:00
kevin 2773ab600d CB-402: live dogfood complete — Stage B verified against opencode 1.18.5
Closes the one known-unverified item before cross-host. CB-402 merged in
ded226a with increment 5 (the §5 live checklist) deferred; it has now run.

Provider question (§7 Q1) resolved with no credentials needed: opencode's own
gateway serves free-tier models. `opencode auth list` reports 0 credentials,
yet `opencode run -m opencode/north-mini-code-free` answers. Distinct from the
primary's subscription by construction, and needs no guard entry — opencode
carries no ANTHROPIC_BASE_URL, so SubscriptionGuard never applies to it.

The schema-drift risk was the real one and it did not bite. The adapter was
designed against opencode 1.1.31; installed is 1.18.5. The generated config
still validates unchanged (type:"remote" + instructions:[path]), and
`OPENCODE_CONFIG=… opencode mcp list` reports the bridge connected. Pinned as
a verified fact for 1.18.5.

Full lifecycle through REST: spawn (201, kind-routed to OpenCodeLauncher) ->
CB-306 gate passed ~0.6s -> ready -> send -> {"replySource":"reply"} (a
STRUCTURED bridge_reply, not the CB-115 completion fallback) -> delete (204,
tolerant teardown).

Unplanned cross-validation with CB-501: the audit trail recorded the reply as
role=WORKER actor=worker:term_657c… — connection-based identity classified an
opencode process as a worker with no opencode-specific handling. The identity
model is peer-kind-agnostic, which is what CB-308 needs when the roster
stretches across hosts.

Stage 5 verified live on the same run: /workers (CB-304) answers where the
13-day-old daemon 404'd, /metrics counted the delegation
(sends_total{outcome=replied} 1, replies_total{path=rendezvous} 1,
inbox_depth 0), and the audit log captured SPAWN/SEND/REPLY with correct roles.

Adds the opencode-free dogfood profile to the local bridged.yaml (gitignored;
recorded here for reproducibility) and docs/CB-402 §8 as-built.
2026-07-29 22:49:15 +07:00
kevin 19cdf8dc9f CB-505 fix: audit lines were not valid JSON
The first cut spliced the timestamp on via a logback pattern:

    {"ts":"%d{...}",%replace(%msg){'^\{',''}%n

Logback's variable substitution chokes on the literal braces
("All tokens consumed but was expecting }"), so the encoder failed to
configure. Caught by running the real jar and noticing logback had dumped its
internal status — which it only does when something failed to parse. The build
was green throughout: nothing asserted the audit trail was machine-readable.

AuditLog now emits the complete object including its own ISO-8601 "ts", and the
appender pattern is a bare %msg. Adds AuditLogTest, which parses each emitted
line with Jackson (so a malformed record fails the build) and pins that hostile
ids cannot escape their field to forge a second record.

311 tests green; logback now configures with zero internal errors.
2026-07-29 22:32:18 +07:00
kevin 9daf1ec5ba CB-5xx: Stage 5 hardening — auth, authz+audit, metrics, CI, supervision
Closes out single-host before the cross-host work. Sequenced BEFORE CB-308
deliberately: federation's own gating concern is the trust model, and it
inherits whatever identity shape lands here.

The finding this stage is built around: bridged had exactly ONE security
control, the loopback bind. ConnectionIdentity resolves a worker from its
connection (unforgeable), but every caller that was not a recognised worker
pane fell through to being treated as the PRIMARY -- the most privileged role
on the bus. Latent today; load-bearing the moment a bind widens.

CB-501 auth:
- Role/Principal/CallerResolver: connection identity first, bearer token
  second, ANONYMOUS third. Inverts the old default so absence of identity
  means nothing, not everything.
- Worker identity is never token-gated, so enabling auth cannot lock the
  fleet out of bridge_reply.
- Constant-time token compare (MessageDigest.isEqual).
- validateAuthExposure(): a non-loopback bind under loopback-trust now
  REFUSES TO START. Makes the dangerous config unrepresentable rather than
  merely documented.
- TLS terminates at a reverse proxy by design (D3), not in the JVM.

CB-505 authz + audit, enforced on BOTH entry paths:
- The docs describe MCP as "a thin adapter over the REST core"; at code level
  it is not. BridgeMcp calls MessageService directly, and /mcp is a raw
  servlet on Jetty's context handler that never traverses Javalin's before
  filter. Enforcing only at REST would have left /mcp open.
- Load-bearing rule is own-session-only: a worker may reply/ask only as
  itself. Structurally true over MCP already; over REST the session id in the
  URL path had simply been trusted.
- Audit: JSON lines to a dedicated appender, additivity=false. Never records
  message content -- this bus carries source and prompts.

CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer
instead of the specced Micrometer, because this pom already hand-pins
jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against
skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated
dependency CVE gate could not be run (no JetBrains MCP server connected).
Instrumented at MessageService, the single funnel both surfaces share.

CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner.
Needs no contract-exclusion flag -- the pom's default-excludes profile
already sets excludedGroups=contract, so plain `mvn clean install` IS the
mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older).

CB-504 supervision: launchd agent (the real target -- this host is macOS,
there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds.
Ordering directives are advisory, so the actual fix is that startup now waits
up to 30s for the herdr socket and then serves degraded, instead of crashing
into a restart loop on a boot-order race.

Also fixes drift found while surveying:
- bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config
  binds via plain Jackson with ignoreUnknown, so uncommenting it would have
  been silently dropped and the default kept. Now camelCase, with a test that
  loads the shipped example and one that pins every documented knob's
  spelling -- no test had ever loaded that file.
- Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay,
  gitTokenEnv, gitHostEnv, configDir, primary:).
- README "Next" listed bridge_ask and session lifecycle as upcoming; both
  shipped long ago.
- docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre-
  implementation" for work already merged.

307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS.
Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check
could not be run -- no JetBrains/intellij-index MCP server is connected this
session. mvn clean install is the only gate that ran.
2026-07-29 22:29:26 +07:00
Dai Ha c9f0ca9359 wiki: bump submodule pointer to b646be1 (chapter 10 — Cross-Host Messaging & Broker Topology) 2026-07-28 16:33:51 +02:00
Dai Ha 84c8a2d2f0 CB-500 §11: resolve distributed-sandbox topology (gateway-per-host × local sandboxes)
Clarifies the open fork from §4/§9: Development A (sandbox launcher) and CB-308
(per-host federation) COMPOSE — each host runs a bridged gateway whose launcher
spawns agents into that host's LOCAL sandboxes; the broker moves messages +
presence, never keystrokes.

The forcing fact: delivery is herdr keystroke-injection into a locally-owned PTY,
so "a sandboxed agent on another host" ≡ "a sandbox spawned by that host's
gateway" (a remote container with no local herdr can't be injected into). Rules
out a central daemon reaching remote PTYs.

Adds Figure 11 (composed topology) + Figure 12 (remote-delegation sequence: the
local?inject:publish fork with a sandboxed far side — both injection points stay
local, only the middle hop crosses the broker), the two forced reachability
changes (host-routable mcpUrl; PTY in the local gateway's herdr), and a
maps-to-existing-seams table (CB-308 gateway × Dev-A launcher, CB-117 reap,
CB-303 container lifecycle, CB-308 #5 trust). No new pillars. Both diagrams
mmdc-validated; §4/§9 updated to point at §11.
2026-07-28 16:26:10 +02:00
Dai Ha ded226abfe Merge CB-402: opencode second peer adapter (Stage B of the Peer Launcher SPI)
Proves the CB-401 PeerLauncher SPI is genuinely provider-neutral by landing a
second adapter — opencode — that shares NONE of Claude Code's private launch
seams (no ANTHROPIC_BASE_URL, no SubscriptionGuard). Merged as one unit:

- Incr 2 (6e37722): kind: discriminator on BridgedConfig.Worker (claude-code |
  opencode), argv defaults to the kind binary, isClaudeCode()/isOpenCode().
- Incr 3 (b034f10): OpenCodeLauncher extends HerdrPeerLauncher — diverges only
  in buildLaunch (file-based MCP mount via OPENCODE_CONFIG + reply charter under
  instructions[], model via -m provider/model, opencode- name/reap prefix).
- Incr 4 (11f8709): CompositePeerLauncher routes the fleet by kind (spawn/cwd/
  parity by profile, stop by pane-id owner, list dedup, reap/caps/profiles
  union); Bridged.main partitions profiles by kind → composite. BridgeMcp +
  BridgedApp migrated onto the PeerLauncher SPI (the two (ClaudeCodeLauncher)
  casts removed; dead rendezvous params dropped).

Incr 1 (extract HerdrPeerLauncher base) already on main at ffce30a.
266 tests green, mvn BUILD SUCCESS. Live opencode-on-Gemini dogfood tracked
separately (needs a running-daemon restart + a resolved Gemini provider).
2026-07-28 16:17:34 +02:00
Dai Ha 9b8d55bc18 CB-500: design note — multi-tier coordination (Stage 6)
Scopes the lead's next direction on top of the peer-launcher arc, as a
proposal (ticket split deferred):
- A · sandboxed, role-specific workers — a sandbox is a *placement*, so it
  slots into the CB-401 SPI as a new kind: sandbox adapter exactly the way
  CB-402's opencode slotted in as a new provider (SPI proven placement-neutral,
  not just provider-neutral). Ownership line held: bridge launches INTO a
  peer-owned image, never provisions the IDE/dev-tools inside it.
- B · main-agent pairs (Opus + cloud) — both mains are MCP clients so neither
  can be called into; each needs a pull inbox → depends on CB-308's per-agent
  channels. PrimaryRegistry single-slot → multi-slot.
- C · orchestrator tier — SessionManager recursed one tier up
  (orchestrator:mains :: main:workers) + context scoping; re-roots the human
  from a live primary to the orchestrator.

10 mermaid diagrams (component + sequence per development, ownership guardrail,
staging graph), all mmdc-validated and theme-safe. §7 pins the bus-vs-env-manager
boundary as an acceptance criterion on A; §8 stages A → CB-308 substrate → B → C.
2026-07-28 15:20:16 +02:00
Dai Ha 11f8709286 CB-402 Increment 4: CompositePeerLauncher — route the fleet by kind
Introduce the router the core holds when more than one adapter is
configured: one HerdrPeerLauncher per peer kind, dispatched by profile
(spawn/effectiveCwd/parityOverlay), by pane id (stop, via a spawn-time
owner map), and fanned out + combined for the fleet-wide queries
(list dedup by pane id, reap/caps union, profiles union). The ctor
rejects an empty adapter list and a profile two adapters both claim.

Wire it in Bridged.main: partition workerProfiles() by kind (claude-code
is the always-present default adapter; opencode is added when any profile
opts in) and front both with the composite. This lets BridgeMcp and
BridgedApp finally take the PeerLauncher SPI instead of a concrete
ClaudeCodeLauncher — the two (ClaudeCodeLauncher) casts in Bridged are
gone. list() elements are cast to herdr Agent at the point of the
herdr-specific roster view, where that assumption actually lives.

Add BridgedConfig.Worker.isClaudeCode()/isOpenCode() kind predicates
(the wiring uses isOpenCode; both are unit-tested). Drop the long-dead
'rendezvous' constructor param threaded into BridgeMcp and BridgedApp.

10 CompositePeerLauncherTest cases over two real adapters on one
FakeHerdr: profile routing (observed via the started agent's
claude-/opencode- name prefix), default resolution, unknown-profile
and duplicate-profile rejection, caps union, list dedup, reap sum,
and stop teardown. 266 tests green.
2026-07-22 05:43:28 +02:00
Dai Ha b034f105c0 CB-402 Increment 3: OpenCodeLauncher — the SPI-proving second adapter
A HerdrPeerLauncher subclass for opencode, a provider-agnostic terminal
coding agent. It reuses every line of shared base transport (tab/pane
placement, CB-306 readiness gate, unique naming + CB-117 reap, teardown,
listing, cwd) and diverges only in buildLaunch:

  - No subscription boundary: no ANTHROPIC_BASE_URL, no SubscriptionGuard
    (the guard is a Claude-private concern, not part of the SPI).
  - File-based MCP mount + instructions: writes an ephemeral opencode.json
    declaring the bridge as a remote MCP server + a reply-charter file under
    instructions, pointed at via OPENCODE_CONFIG (opencode has no inline
    --mcp-config / --append-system-prompt).
  - Model selected with -m provider/model, not an env var.
  - 'opencode' name prefix so reap matches opencode-* panes only.

configRoot is injectable so tests inspect the generated config/charter under
a @TempDir. 10 tests cover config content, model flag, git-token grant,
capabilities, reap predicate, both production ctors, and the readiness gate
(throws PeerUnreachable on timeout, reaps only the worker pane).
2026-07-22 05:30:46 +02:00
Dai Ha 6e37722383 CB-402 Increment 2: kind: discriminator on Worker profiles
Add a `kind` field to BridgedConfig.Worker — "claude-code" (default) or
"opencode" — the discriminator the CompositePeerLauncher will route spawn/reap
by so each adapter drives only its own peer kind. Normalised to lower-case;
blank/absent ⇒ claude-code, so every existing config and call site is
unchanged. argv now defaults to the kind's own binary (claude vs opencode)
rather than always `claude`, so an opencode profile never inherits the Claude
command.

kind is appended at the record tail; a new 14-arg back-compat constructor
(git fields, no kind) keeps the CB-302 call sites working, and the existing
12-arg constructor is untouched. Also drop the never-used Primary(String)
legacy constructor to keep the file warning-clean.

example.yaml documents the key and carries a commented opencode-gemini
profile. Tests cover default/normalisation/argv-defaulting. 245 tests green,
config files 0 IDE problems.
2026-07-22 05:20:04 +02:00
Dai Ha ffce30afa2 CB-402 Increment 1: extract HerdrPeerLauncher abstract base
Behaviour-preserving refactor ahead of the second-adapter work. All herdr
transport shared by any peer kind — tab/pane placement, the CB-306
spawn-readiness gate, unique naming + CB-117 orphan reap, teardown, listing,
cwd resolution, and the peer-neutral git-forge grant — moves into a new
abstract HerdrPeerLauncher (Template-Method base). ClaudeCodeLauncher becomes
a final subclass supplying only the two Claude-specific seams: the `claude`
name prefix and buildLaunch(), which encodes the subscription boundary
(ANTHROPIC_BASE_URL + SubscriptionGuard assert, inline --mcp-config and
--append-system-prompt reply charter).

The base owns the injectable clock + sleeper for the readiness gate; the poll
interval is baked into the sleeper, so the vestigial spawnReadyPollMs field is
dropped from the base and from the full testability constructor (the explicit
sleeper already encodes it). The 6-arg and 8-arg production constructors keep
their signatures; three full-ctor test sites drop the now-unused poll argument.

No behaviour change: 242 tests green, both refactored files 0 IDE problems.
2026-07-22 05:15:21 +02:00
Dai Ha e724a59f2d CB-402: design note — opencode second peer adapter (Stage B)
Design-note-first for gitea #7. Extract HerdrPeerLauncher abstract base
(Template Method) + kind: discriminator + OpenCodeLauncher + routing
CompositePeerLauncher; finish the Stage-A caster migration off
ClaudeCodeLauncher. opencode proves the SPI for a non-Claude peer
(no subscription guard, OPENCODE_CONFIG MCP mount, config-based charter).
2026-07-22 04:59:58 +02:00
Dai Ha 4bf855d225 wiki: bump submodule pointer to 4d1548a (CB-307 as-built)
Advance the wiki submodule pointer to include the CB-307 Stage 2/3 as-built
docs plus the intervening CB-401/delegation-directive commits. Standalone
pointer bump — not bundled into a feature commit.
2026-07-19 18:10:01 +02:00
Dai Ha d4c9704007 CB-307: primary-gate cleanup of push-loop (drop dead clock, IDE 0/0)
Primary verification pass over the delegated push-loop delivery:
- Remove the unused LongSupplier clock threaded into ReplyPushLoop
  (timing is the scheduler's; the field was never read) from the
  component, Bridged wiring, and both test call sites.
- Collapse the single-statement WAIT_BUSY switch arm (redundant block).
- Drop now-dead test scaffolding: the always-"idle" recordingClient
  param and unused AgentStatus/AtomicReference imports.

IDE diagnostics 0/0 on all changed files; mvn clean install green
(242 tests, 0 failures).
2026-07-19 11:03:34 +02:00
Dai Ha 7c252b5f5f CB-307 Increment 3: bridge_ack tool — per-msgId ack refinement
- MessageService.ackReply(target, msgId) delegates to inbox.ack
- BridgeMcp registers bridge_ack tool with target/msgId args
- Tests: valid/invalid args, ack surface via BridgeMcp
2026-07-19 10:21:06 +02:00
Dai Ha f756933879 CB-307 Increment 2: ReplyPushLoop — status-gated push loop (mechanism b)
- ReplyPushLoop: dedicated scheduled loop for nudging the primary
- Package-private decide() method for pure decision logic (unit-testable)
- Status-gated injection via AgentControl.status().injectable()
- Bounded reminders (cap + backoff), idempotent per target
- Wire into MessageService.reply after inbox.publish on no-waiter branch
- Wire into Bridged.main (constructor + shutdown hook)
- Primary config record updated with pushReminders/pushBackoffMs knobs
- Tests: decide() matrix, nudge injection, idempotency, cap enforcement
2026-07-19 10:16:24 +02:00
Dai Ha a1aecbf4fc CB-307 Increment 1: PrimaryRegistry + config + wiring
- PrimaryRegistry: thread-safe single-slot registry with pin support
- Primary config record (last positional, like Broker)
- Wire capture in BridgeMcp (bridge_send and bridge_spawn handlers)
- Construct PrimaryRegistry in Bridged.main
- Tests: PrimaryRegistryTest + BridgedConfigTest primary config cases
2026-07-19 09:59:02 +02:00
Dai Ha 131e7b1ccd CB-307: lock push-loop injection to dedicated status-gated loop (mechanism b)
Chose a small dedicated scheduled loop over AgentControl.send guarded by an
injectable status check, instead of reusing the worker Injector (which couples
to WorkerPresence/StatusPoller). Isolated + unit-testable via injected clock.
Records live ground truth: this primary resolves to term_656c8cc03e1f0b1 (w2:pY)
— confirms the primary runs in a herdr pane so the push path is exercisable.
2026-07-19 09:44:21 +02:00
Dai Ha d0ac6c435f CB-307: design note for active push-to-primary + bounded reminder loop
The reliability layer over the durable inbox (Stage 2, 2bc5f3a): push a nudge
into the primary's own herdr pane the moment a no-waiter reply lands, remind on
a bounded backoff until the primary drains (ack=drain), degrade to pull when the
primary pane isn't resolvable. Grounds the seam: the primary terminal_id is
already derivable via ConnectionIdentity/PaneLocator, just discarded today.

Refs gitea #5.
2026-07-19 09:41:22 +02:00
Dai Ha 2bc5f3a057 CB-307 Stage 2: AmqpReplyInbox — durable, cross-restart reply delivery
Behind the existing ReplyInbox port, add an AMQP-backed adapter selected by a
`broker:` block in config (absent → the in-memory soft-state inbox; present →
AMQP). Mapping is consume-and-hold with deferred manual ack: each target owns a
durable queue `agent.<target>.inbox`; a manual-ack consumer pulls persistent
messages into an in-memory held map (dedup by msgId) but does not ack; peek
returns the snapshot; ack acks the broker delivery-tag and drops it. A crash
before caller-ack leaves messages unacked, so the broker redelivers on
reconnect — genuine durability with the port contract preserved. bridged still
owns no persistence; the broker does.

- msg/AmqpReplyInbox: the adapter (single synchronized channel; recovery
  listener clears held on reconnect so fresh delivery-tags repopulate).
- config/BridgedConfig: nullable Broker(uri) record; isConfigured() gates it.
- Bridged.main: select adapter; close the AMQP connection in the ordered
  shutdown hook (no-op for the in-memory inbox).
- deps: com.rabbitmq:amqp-client (main); testcontainers rabbitmq/junit-jupiter
  (test). Pinned commons-compress 1.27.1 + commons-lang3 3.18.0 to clear the
  test-scope CVEs those pull. Production default LavinMQ; RabbitMQ URI-swap.
- tests: BridgedConfigTest broker-selection cases (hermetic); AmqpReplyInbox
  contract test (@Tag("contract"), Testcontainers RabbitMQ) proving
  publish/peek/ack, msgId dedup, and cross-restart redelivery. Excluded from
  the default build so `mvn clean install` stays hermetic (210 green).
2026-07-19 07:30:29 +02:00
Dai Ha ba6b4a5da9 CB-307 Stage 1: reply-inbox port + in-memory adapter — hold stranded worker replies instead of dropping them
Problem: the reverse (worker->primary) path was Rendezvous, a map of LIVE blocking
waiters only. A bridge_reply arriving with no open send hit Rendezvous.complete()
-> no waiter -> returned false -> the reply was silently DISCARDED (worker saw an
error / REST 409). No message-id/dedup/ack existed anywhere.

Stage 1 (no broker, soft-state) behind one port:
- ReplyInbox port + InboxMessage record; InMemoryReplyInbox adapter (per-target
  FIFO via LinkedHashMap, dedup by msgId, thread-safe). Soft-state, not persistence.
- MessageService.reply(session, content): resolve an open send, else publish to the
  inbox with a minted UUID (was a silent drop). drainReplies(target) = peek + ack.
- BridgeMcp.reply / BridgedApp.replyMessage repointed off bare Rendezvous.resolve
  onto messages.reply -> no-waiter is now SUCCESS (queued), not error / 409.
- Drain surface: bridge_poll gains optional target; REST GET /sessions/{id}/replies.
- Rendezvous left untouched. QUESTION path (bridge_ask) NOT queued (interactive,
  keeps NO_WAITER); completion/failure fallbacks NOT queued (captured-waiter).
- Bridged.main wires new InMemoryReplyInbox(); no broker: config yet (Stage 2 = AMQP).

Tests: +19 (188 -> 207), 0 failures/0 errors. New InMemoryReplyInboxTest (12) +
MessageService/BridgeMcp/BridgedApp coverage incl. guards proving a QUESTION and a
completion fallback are never queued.

Implemented via delegation to an off-sub gx10 worker in a pre-trusted worktree;
primary-verified (mvn clean install green, 207 tests) and committed by the primary
because the worker's completion replies were lost to the very bug this fixes.

Refs CB-307 (gitea #5), Stage 1 of 2.
2026-07-18 21:12:51 +02:00
Dai Ha da5a987df0 CB-307/CB-308 design notes: reliable-delivery Stage-1 delegation spec + multi-host federation proposal
- docs/CB-307-Reliable-Delivery.md: ReplyInbox port + in-memory adapter spec
  (Stage 1, no broker); publish at the Rendezvous no-waiter drop seam, drain by target.
- docs/CB-308-Multi-Host-Federation.md: per-host gateway + per-agent broker channels
  + federated roster proposal (gitea #6), wiki-ready with theme-safe mermaid.
2026-07-18 20:24:16 +02:00
Dai Ha 7dd6c46156 CB-306: spawn-readiness gate — ClaudeCodeLauncher blocks until the worker is injectable or throws PeerUnreachableException
The launcher now polls AgentControl.status(paneId) after starting the pane.
It returns the handle only once the worker reports an injectable state
(IDLE/BLOCKED/DONE). If the timeout elapses while still UNKNOWN, the
pane is self-reaped and a PeerUnreachableException is thrown — no orphan
left behind. The gate is disabled when spawnReadyTimeoutMs == 0 (legacy
non-blocking spawn, the default for the 6-arg constructor).

Key changes:
- PeerUnreachableException (new) in dev.ltms.bridged.peer
- BridgedConfig: spawnReadyTimeoutMs (default 20000), spawnReadyPollMs (default 300)
- ClaudeCodeLauncher: 3 constructor overloads:
  (a) 6-arg backward-compat: gate disabled (timeout=0)
  (b) 8-arg production: gate with config knobs + real clock/sleep
  (c) 10-arg testability: full seam (LongSupplier clock + Runnable sleeper)
- waitUntilInjectableOrThrow() loop in spawn(SpawnRequest)
- sleepUninterruptibly() helper for the production sleeper
- BridgeMcp.spawn + BridgedApp.spawnWorker catch PeerUnreachableException
  → clean tool error / 502 response (not an uncaught 500)
- SessionManager.acquire inherently registers nothing on throw (both
  worktree and non-worktree paths) — confirmed by new test

Tests:
- ClaudeCodeLauncherTest: 4 new tests
  - unknown→idle: returns handle, no pane.close
  - always-unknown: throws PeerUnreachableException, pane closed,
    clock advanced past timeout
  - timeout=0 (6-arg ctor): no agent.get calls, returns handle
  - timeout=0 (10-arg ctor): no orphan pane close
- SessionManagerTest: 1 new test
  - acquire → PeerUnreachableException: roster remains empty
Total: 188 tests, all pass (no existing test changed semantics)
2026-07-18 16:17:43 +02:00
Dai Ha 3a5cdc5108 CB-306: spawn-readiness gate design note (delegation spec) 2026-07-18 15:41:43 +02:00
Dai Ha 3aa69a9e32 CB-401 Stage A follow-up: rename WorkerService -> ClaudeCodeLauncher
Name the first-class Claude Code adapter explicitly, per the Peer Launcher SPI:
WorkerService was the de-facto Claude-Code launcher; as an in-tree PeerLauncher impl
it should say so. Pure IDE rename (class + file + WorkerServiceTest) plus stale
Javadoc/comment mentions swept to the new name. No behaviour change.

Gate: IDE diagnostics 0/0 on touched files; mvn clean install BUILD SUCCESS,
MVN_EXIT=0, 183 tests pass. Deferral #1 from issue #3 cleared.
2026-07-18 14:39:56 +02:00
Dai Ha e056c7e1fa CB-401: Stage A - extract PeerLauncher SPI in-tree 2026-07-18 07:48:55 +02:00
Dai Ha d63273d082 CB-401: Peer Launcher SPI design note (Stage 4)
Design-only. Audits the as-built Claude/herdr coupling (concentrated in
WorkerService), defines a PeerLauncher SPI + opaque PeerHandle + capability
model so the bus delegates peer materialization to a config-selected adapter.
ClaudeCodeLauncher = adapted WorkerService. Stages A/B/C with the Stage-C
plugin-loading security gate called out. No production code touched.
2026-07-17 17:49:21 +02:00
Dai Ha 0efb65ca0a CB-303: session lifecycle limits — idle_ttl reaper, context_cap, graceful drain
Verified on primary: ide_diagnostics clean (incl. weak warnings), mvn clean install
BUILD SUCCESS, 173 tests. Delegated impl (worker/cb-303-80ec1a-3, 3 parts), primary-gated.
2026-07-17 10:08:55 +02:00
Dai Ha 9fe04bfb08 CB-304: bridge_list roster + live herdr join (surface worktree/branch); add GET /workers
Verified on primary: ide_diagnostics clean (incl. weak warnings), mvn clean install
BUILD SUCCESS, 165 tests. Delegated impl (worker/cb-304-bd1e4f-2), primary-gated.
2026-07-17 10:08:46 +02:00
Dai Ha 09d3948acf CB-303 part 3: graceful drain on shutdown 2026-07-17 09:59:57 +02:00
Dai Ha 954351a80b CB-303 part 2: context_cap turn budget 2026-07-17 09:56:46 +02:00
Dai Ha 8d51066ddd CB-303 part 1: idle_ttl session reaper (injectable clock + SessionReaper) 2026-07-17 09:52:40 +02:00
Dai Ha 84102baab4 CB-304: bridge_list roster + live herdr join (worktree/branch); add GET /workers 2026-07-17 09:50:12 +02:00
Dai Ha 64e70efdf1 CB-302: worker checkpoint — repo-scoped forge token injection + implementer skill
The worker "checkpoint" is commit → push → open its own PR. Push is free over SSH
(same user, same keys); the only incremental grant is PR-create, so the daemon injects
a repo-scoped gitea token into the worker env — opt-in per profile, never mutating
bridged's own environment.

- BridgedConfig.Worker: gitTokenEnv/gitHostEnv fields (opt-in; gitHostEnv defaults to
  GITEA_HOST). Backward-compat 12-arg constructor keeps pre-CB-302 call sites + YAML
  working. hasGitToken() gates injection.
- WorkerService.spawn: inject GITEA_TOKEN (and paired GITEA_HOST) only when the profile
  grants a token AND the host env resolves one. resolveEnv() tolerates unset var names.
- WorkerServiceTest: injection present for a granting profile; absent when not (proving
  the gate is config, not a missing env var).
- .claude/skills/implementer/SKILL.md: worktree-aware playbook — confirm the worktree/
  branch, implement, commit (never .mcp.json/wiki), push, open PR via gitea REST with
  GITEA_TOKEN, hand off the PR URL via bridge_reply. Never merge; workers can't run IDE
  diagnostics so never claim IDE-clean.

Whole-project gate: mvn clean install green, 164 tests, 0 failures.
2026-07-17 08:51:27 +02:00
Dai Ha 97ecc7136e CB-301-ext: per-worker git worktree + config-parity overlay
Opt-in isolated worktree so parallel implementers don't stomp the shared
tree, hydrated to config parity so a worker differs from the primary only
in LLM provider.

- Worktrees seam (interface) behind SessionManager; GitWorktrees shells git
  via ProcessBuilder (non-zero exit -> WorktreeException), FakeWorktrees for
  tests. No live git in unit tests.
- acquire() 5-arg overload provisions add -> overlayParity -> spawn(cwd=wt)
  -> register, unwinding the worktree on any failure before registration.
  4-arg overload and shared-tree behavior unchanged (backward compatible).
- release() removes the checkout but never deletes the branch (it holds the
  worker's commits + PR, CB-302).
- overlayParity copies local config (.mcp.json, settings.local.json, .env/
  .envrc) into the worktree; tracked ones get --skip-worktree so a worker
  can never stage the parity overlay.
- WorkerSession gains nullable worktree/branch; BridgedConfig.Worker gains
  parityOverlay (default list) + top-level worktreeRoot.
- bridge_spawn / POST /workers gain an optional worktree(+ticket) arg; the
  worker view includes worktree/branch only when non-null.

Verify fixes on the delegated impl: strip trailing dashes in slug()
(^-+|-+$, was ^-+|^-+$); make FakeWorktrees.add a pure fn of the branch
(nonce already unique); MCP worktreeRequest treats blank/"false" string as
no-worktree, matching the REST builder.

162 tests, 0 failures.
2026-07-17 06:46:20 +02:00
Dai Ha f9073e2320 docs: CB-301-ext spec — worktree provisioning + config-parity overlay
Opt-in per-acquire worktree (shared-tree default preserved). git behind a
Worktrees seam (ProcessBuilder impl, fake in tests). acquire provisions
worktree+branch, overlays local config (copy + --skip-worktree on tracked
files so the worker can't commit .mcp.json), spawns with cwd=worktree.
release removes the worktree but keeps the branch (holds commits/PR).
WorkerSession gains nullable worktree/branch; BridgedConfig.Worker gains
parityOverlay. 6 fake-based acceptance tests incl. backward-compat + unwind.
2026-07-17 06:32:09 +02:00
Dai Ha 54d907c314 CB-301: SessionManager — authoritative worker session registry + one-shot FSM
Adds dev.ltms.bridged.session with WorkerSession (immutable record) and
SessionManager wrapping WorkerService: a ConcurrentHashMap registry keyed by
paneId, the one-shot lifecycle FSM (SPAWNING->READY->BUSY->DONE, ->FAILED on
drop/turn-failure, ->RELEASED on teardown), ownership (ownerTerminal), and
recycle = release + fresh acquire (no-reuse invariant). Driven by TurnListener
(BUSY/DONE/FAILED) and a WorkerPresence bridge (READY).

Wiring: Bridged.main constructs it and composes it into the TurnListener
alongside CompletionResolver; bridge_spawn / POST /workers route through
acquire (carrying caller identity as owner); bridge_stop / DELETE /workers
route through release. WorkerService gains effectiveCwd(); WorkerPresence
de-finalized so the manager can present a READY-driving view.

asPresence() returns a single cached bridge (a fresh one per call would
fragment the shared present set). roster() is the registry snapshot; the live
herdr join is left for CB-304. 6 fake-based acceptance tests; full suite green
(155/155).

Delegated to an off-subscription worker against docs/CB-301-Session-Manager.md;
primary verified (ide diagnostics clean, mvn clean install green) + fixed the
asPresence caching bug.
2026-07-16 19:30:56 +02:00
Dai Ha 82c7d6553a docs: worker git workflow — daemon worktree + config parity + worker-opened PR
Worktree is code-only isolation; SessionManager hydrates it to full config
parity (overlay untracked local settings/.mcp.json/.env) so a worker is a
full peer of the primary, differing only in the LLM provider. Worker commits,
pushes over SSH, and opens its own PR (gitea REST + repo-scoped token).
2026-07-16 19:25:41 +02:00
Dai Ha 19b10e3216 docs: CB-301 session-manager design spec (one-shot, no reuse; recycle in scope)
The as-built audit surfaced that WorkerService keeps no registry of what it
spawned (its own Javadoc: 'there is no registry; list() only asks herdr').
CB-301 adds a SessionManager wrapping WorkerService: an authoritative in-daemon
roster with a per-session lifecycle FSM (SPAWNING/READY/BUSY/DONE/RELEASED/
FAILED), deterministic release, and recycle (= release + fresh acquire, no
reuse). Leaves clean seams for CB-302 (checkpoint on release), CB-303 (idle_ttl/
context_cap/drain policy over roster), CB-304 (bridge_list reads roster).
2026-07-16 19:10:13 +02:00
Dai Ha 0f79e6bed5 wiki: bump submodule to page-9 as-built implementation architecture (8e5fd01) 2026-07-16 16:45:31 +02:00
Dai Ha aa0cf814ff e2e: capture bridge_ask transcript confirming single turnId after coalescing
Post-4aa9d03 live run: the duplicate-ask retry now coalesces onto one turn
(turnId=term_...#1, was #2 before the fix). Clean round-trip, RESULT OK.
2026-07-16 16:32:30 +02:00
Dai Ha 4aa9d03da2 CB-205: coalesce duplicate bridge_ask calls onto one turn
A worker's tool execution is single-threaded, so two bridge_ask calls from the
same session can only be a transport retry — yet openAsk minted a fresh turnId
each time and both raced to resolve the primary's single forward waiter, the
loser returning NO_WAITER and leaving a dangling turn (observed as turnId #2 in
the live e2e). Now openAsk is idempotent per session: a second open ask coalesces
onto the existing turnId + answer future (fresh=false), and only the fresh owner
surfaces the question and tears the turn down. closeAsk clears the per-session
index (conditional by value) so a later ask reopens fresh.

Implemented by an off-subscription worker delegated over the bridge; reviewed and
validated on the primary (IDE-clean, mvn clean install 149 green, +2 tests:
concurrent double-ask coalescing + post-close reopen).
2026-07-16 16:27:10 +02:00
Dai Ha 37abfdfd6e wiki: bump submodule to Stage-2-complete roadmap update (0c81b51) 2026-07-16 16:15:15 +02:00
Dai Ha d5c3ede215 CB-202: reviewer-role skill for bridged workers
The playbook a reviewer worker loads when the lead delegates a scoped review:
read the whole scope before judging, stay in the assigned lane, ask the lead
via bridge_ask when the call is genuinely theirs (resuming the same turn with
the answer), and report exactly one structured finding via bridge_reply. Pairs
the already-shipped bridge_reply/bridge_ask tools with the role guidance that
tells a worker how to use them. Mermaid validated with mmdc.
2026-07-16 16:10:36 +02:00
Dai Ha 426855e378 e2e: live bridge_ask reverse-rendezvous harness (CB-205)
Drives the reverse path end to end over REST loopback: a worker is delegated
a task it cannot finish without asking, calls bridge_ask mid-turn, and the
primary answers on the surfaced turnId so the worker resumes the SAME turn.
Two blocking sends, no polling. Verified live: worker asked in ~9s, resumed
and replied CHOSEN=BLUE via clean bridge_reply after the primary answered.

Subscription-safe by construction (REST face only; never sets ANTHROPIC_BASE_URL).
2026-07-16 16:08:21 +02:00
Dai Ha 358c6970b5 CB-205/CB-201: bridge_ask reverse rendezvous + lightweight question/turn kind
A worker can now pause its delegated turn to ask the primary a question and
resume the same turn with the answer — the reverse of bridge_send.

- Rendezvous: QUESTION kind carrying a turnId, plus a reverse-ask registry
  (openAsk/resolveQuestion/askSession/answerAsk/closeAsk).
- MessageService.ask(): surface a worker's question to the primary's open send,
  block for the answer. answer(): resolve the worker's ask by turnId, then block
  for its eventual bridge_reply (session derived from turnId, not an argument).
- BridgeMcp: bridge_ask tool (worker-only, identity from the connection);
  bridge_send routes a turnId to the answer path. Shared formatReply().
- BridgedApp REST parity: POST /sessions/{id}/ask, turnId on /message.
- Lightweight CB-201: the kind vocabulary is the QUESTION outcome + turn_id
  correlation, not a rigid from/to/corr envelope (the connection-identity
  mechanism already covers addressing more robustly).

Tests: MessageServiceTest ask/answer round-trip + NO_WAITER/timeout/stale;
new RendezvousTest for the reverse registry; BridgeMcp ask/answer parity.
mvn: 147 green.
2026-07-16 15:32:56 +02:00
Dai Ha 55ebd5b949 e2e: sustained back-and-forth conversation harness (5-min stateful continuity) 2026-07-16 15:19:21 +02:00
Dai Ha 2a61fe69f1 CB-118: clip the completion baseline so the CB-115 guard survives >cap blocks
captureBaseline stored the raw, unclipped last-assistant block while resolve()
compares against clip(...) capped at MAX_SCRAPE_CHARS. For a block longer than
4000 chars the two capped representations never match even when the pane is
unchanged, defeating the CB-115 misattribution guard and letting a stale
completion resolve a rapid back-to-back send. Clip the baseline identically.

Regression test: an unchanged >cap block stays suppressed.

Surfaced by the fan-out issue-hunt E2E (1 primary -> 3 concurrent workers,
e2e/issue_hunt_test.py, added here). The same hunt's WorkerService.stop() and
Rendezvous.complete() findings were verified as false positives (locatePane is
already guarded; the sender's finally-close already removes the waiter).

Closes #2
2026-07-16 09:11:20 +02:00
Dai Ha ff6aacdc78 CB-117: reap orphaned worker panes on startup
herdr keeps worker panes alive across a daemon restart by design, and a
worker's paneId is held only by its spawner — so a worker whose owning
process exited before its DELETE leaks with nothing tracking it (there is
no registry; list() only asks herdr). Observed as three idle claude-ollama
panes left in the worker space from earlier runs.

On boot, WorkerService.reapOrphanWorkers() scans herdr for agents whose
name matches our claude-<profile>-<nonce>-<seq> scheme with a nonce other
than this process's nameNonce, and tears each down (pane + its now-empty
dedicated tab). A current-nonce worker is ours and live (spared); a user's
own claude session carries no such name (untouched). Keyed on the nonce so
it survives kill -9 and reaps a *previous* daemon's leaks — the actual case
shutdown-hook reaping and an in-memory registry both miss.

- Agent now projects herdr's 'name' (was dropped) so the reaper can key on it.
- isForeignWorker/workerNonce are pure + package-private for unit testing.
- FakeHerdr.withAgent seeds named agents into agent.list.
- Wired best-effort into Bridged startup before serving.

Closes lms/claude-bridge#1
2026-07-16 08:39:37 +02:00
Dai Ha 5f0ec034d9 e2e: standard bridge conversation test harness
A repeatable multi-turn primary↔worker conversation driven entirely through the
bridge's loopback REST face (async fire-and-poll) — never sets ANTHROPIC_BASE_URL
and never touches herdr, so it is subscription-safe by construction. Records every
turn to a transcript, grades each (OK / DEGRADED / EMPTY / FAILED / WEDGE), and
exits non-zero if any turn fails to deliver-and-reply, so it is CI-usable.

This is the repeatable form of the ad-hoc channel test that surfaced the CB-115 and
CB-116 gaps.
2026-07-16 08:27:50 +02:00
Dai Ha 31e34d177b CB-115/CB-116: reliable turn completion — status refinement, clean scrape, waiter identity
A 5-turn primary↔worker conversation test (see the e2e harness) surfaced three
delegation-channel gaps; this closes them.

CB-115 — status + scrape correctness:
- AgentStatus gains DONE (herdr's explicit turn-complete marker) so a finished
  turn is no longer misread as UNKNOWN and left to wedge or false-fail.
- StatusRefiner reclassifies content-bearing UNKNOWN samples (StatusPoller wired
  to it), and the Injector baselines pane content on delivery (TurnListener gains
  onDelivered) to guard completion against previous-turn misattribution.
- CompletionResolver.lastAssistantBlock stops at the first hard TUI boundary, so a
  scrape returns only the assistant answer — no input box, prompt echo, spinner,
  tips or warnings.

CB-116 — waiter identity (the cross-turn stale reply):
- The completion/failure fallback ran on a virtual thread and resolved whichever
  waiter was currently registered for the session. Since the rendezvous holds one
  waiter per session and sends serialize, turn N's late completion could land on
  turn N+1's waiter and deliver turn N's stale scrape as turn N+1's answer. The
  baseline guard missed it because turn N was resolved by bridge_reply, which never
  updates the completion baseline.
- Fix: capture the exact waiter (and pre-turn baseline) when a turn is delivered,
  on the poller thread before any next-turn delivery can overwrite it, and resolve
  THAT waiter — a no-op if it was already resolved. Rendezvous.resolveCompletion/
  resolveFailure now take the captured CompletableFuture; currentWaiter exposes the
  registered one for capture. A late completion for turn N can no longer touch turn
  N+1's send.

Verified: 128 unit tests green (incl. a CB-116 regression asserting a late
completion never resolves the next turn's waiter); a re-run of the conversation
test passes with turn 5 resolving to its own reply rather than turn 4's text.
2026-07-16 08:27:38 +02:00
Dai Ha b2d85af78b wiki: bump submodule to roadmap Stage-1-complete update (4304dc4) 2026-07-16 06:49:27 +02:00
Dai Ha 968a5c68b6 docs(README): update Status to reflect the shipped implementation
The Status section still read 'Design selected' — but bridged is built and
dogfooded (CB-101..114, 105 tests, live MCP tools, multi-profile, cwd inheritance,
readiness gate). Replace it with an accurate shipped/next breakdown, and drop the
bridge_ask overclaim from the gateway bullet (bridge_ask is roadmap, not built).
2026-07-16 06:46:40 +02:00
Dai Ha 3c05823f19 CB-114: resolveCwd never returns null (honor the 'daemon cwd, never $HOME' contract)
Final review-sweep finding: firstNonBlank(requestedCwd, cfg.cwd(), callerCwd,
user.dir) returns null if all are blank (pathological env with user.dir unset),
after which AgentControl drops the cwd and herdr defaults the pane to $HOME —
violating CB-112's documented contract. Append "." (the daemon's own cwd) as a
guaranteed non-blank last resort. Near-impossible trigger; makes the code honor its
own javadoc. 105 tests green.
2026-07-16 06:43:05 +02:00
Dai Ha 2fb46f670f CB-114: readiness-gate timeout + presence cleanup (delegated review findings)
An off-sub worker's review of CB-113 (delegated through the bridge) surfaced two
real gaps in the readiness gate:

- A worker that herdr reports idle but whose Claude never connects the bridge MCP
  (crashed during boot, or wedged on a startup prompt) left its message queued
  forever: ready.test() never passed, the target was polled indefinitely, and the
  caller's future never completed (async waiter hung for the full 30-min window).
  Injector now counts injectable-but-not-ready samples and, after a ~60s grace
  (READINESS_GRACE_POLLS, deliberately longer than the UNKNOWN stall grace since a
  first boot is slower than an in-turn blip), fails the queued messages, fires
  onTurnFailed so blocking/async waiters resolve WORKER_FAILED, and reclaims the
  target. Mirrors the CB-109 UNKNOWN-stall path.

- WorkerPresence.forget had no caller, so a worker's readiness lingered past its
  life. Injector now clears presence via a forget callback on drop() (pane crash)
  and on the readiness timeout.

+3 InjectorTest cases (never-ready failure, ready-within-grace delivery, drop clears
presence). 105 tests green.
2026-07-16 05:01:58 +02:00
Dai Ha 37f7ad4185 CB-113: reliable worker readiness gate + submission nudge
Delegating right after spawn failed: herdr reports 'idle' during the worker's
boot, so the injector delivered into a not-ready TUI (paste lost) and wedged the
worker. Two fixes:

- Readiness gate: a worker is 'available' only once its Claude connects the bridge
  MCP (the daemon observes it via peer-PID->terminal). WorkerPresence tracks it; the
  injector holds the first delivery until present, so it never pastes into the boot
  window. Exposed as 'ready' on GET /sessions/{id}/status.
- Submission nudge: the Enter accompanying a delivery can race the paste (esp. right
  as the TUI becomes ready), leaving text unsubmitted. While a delivered message
  stays idle (not picked up), the injector re-sends Enter each poll until the worker
  starts (WORKING) or the grace expires.

Validated live: spawn + immediate delegate now holds during boot (ready=false),
delivers on MCP-connect, re-nudges Enter, worker replies. 102 tests green.
2026-07-15 19:14:57 +02:00
Dai Ha 926724a279 CB-112: workers inherit the primary's working directory (not $HOME)
A worker now opens the same directory the primary is in, unless told otherwise.
Resolution: explicit spawn cwd → per-profile config cwd → the primary's cwd
(auto-detected from the bridge_spawn caller's PID via lsof -d cwd) → the daemon's
cwd. Never $HOME.

Mechanism (found by live probe, corrects the earlier assumption): an agent.start
pane does NOT inherit its tab's or workspace's cwd — it starts in $HOME. herdr's
agent.start honours an (undocumented) cwd param, so the resolved cwd is threaded
onto agent.start {cwd} (both tab and pane placement), not tab.create.

Surfaces: bridge_spawn {cwd?} + auto-detect via ConnectionIdentity.resolve (peer
PID) + ProcessCwdLookup (lsof); REST POST /workers ?cwd= / body cwd; per-profile
'cwd:' config. Validated live: explicit cwd → worker rooted there; no cwd over
REST → daemon cwd, not $HOME. Also clears the ccs folder-trust prompt when the
project dir is already trusted (see docs/Worker-Startup-and-Trust.md).
2026-07-15 16:33:46 +02:00
Dai Ha 959f04bc96 docs: worker startup — working directory & the folder-trust prompt
Documents (a) the rule that a worker inherits the PRIMARY's working directory,
never $HOME — resolved from an explicit cwd, else the bridge_spawn caller's PID
(lsof -d cwd), else the daemon cwd; herdr's seam is workspace.create {cwd}, since
agent.start has no cwd; (b) ccs (Claude Code) as the assumed launcher and its
folder-trust model (hasTrustDialogAccepted per project in each instance's
.claude.json), so trusting the project dir once per profile clears the prompt;
(c) that other CLIs have their own startup gates, documented per launcher.
Marks the cwd-inheritance as target design (follow-up), not yet wired.
2026-07-15 16:11:18 +02:00
Dai Ha 7f1b6b3a0d CB-111: multi-profile workers — named backends selectable at spawn
The bridge was hard-wired to one worker profile. Config now takes a 'workers'
map keyed by profile name plus 'defaultWorker'; WorkerService holds the map and
gains spawn(profile) (spawn() uses the default). Selection threads through the
surfaces: REST POST /workers ?profile= / {"profile":…} + GET /profiles; MCP
bridge_spawn {profile?} + new bridge_profiles. Each profile's base_url is
guard-checked independently, so gx10 and ollama can run side by side and you
address each worker by its returned sessionId. Backward-compatible: the legacy
singular 'worker:' block still loads as a one-entry profile map.
2026-07-15 15:54:31 +02:00
Dai Ha ab771ea24e CB-110: cover drop of a delivered-but-unpicked-up turn
Adds the drop test for the delivered/awaitingPickup/!turnObserved state (worker
vanished after delivery but before a WORKING sample) — a gap the other two drop
tests missed. Surfaced by an off-sub worker's code review of CB-110, delegated
through the bridge itself.
2026-07-15 15:30:20 +02:00
Dai Ha a629a7ee73 CB-110: fail an in-flight delegation when its worker vanishes
Companion to CB-109. When a worker disappears mid-turn (pane crash → herdr
*_not_found), the status poller drops the target, which failed only *queued*
messages — a message already DELIVERED is out of the queue, so its send's
rendezvous was left hanging until the 30-min async timeout. Injector.drop now
fires onTurnFailed for a delivered-but-unresolved turn (awaitingCompletion), so
the send resolves as WORKER_FAILED. Reuses the CB-109 resolver path; the failure
reason is neutral to cover both wedge (stuck) and vanish (gone).
2026-07-15 15:23:17 +02:00
Dai Ha 2052929768 CB-109: fail a delegation whose worker wedges in an unknown state
Dogfood found the gap: a turn that dies into an error screen herdr reports as
'unknown' (e.g. the worker hitting API ENOTFOUND) never produces a working->idle
boundary, so CB-106 never fires and the async send rides its full 30-min timeout.

The injector now counts consecutive 'unknown' samples while a delegation is
outstanding; any working/idle sample resets the streak, so only a genuine wedge
(~30s continuous unknown) trips it. It then fires TurnListener.onTurnFailed;
CompletionResolver scrapes the error screen and resolves the send via
Rendezvous.resolveFailure (Kind.FAILED -> Outcome.WORKER_FAILED), surfaced as
async phase=failed / REST status=failed / a [worker failed] MCP note, with the
error context as the reason. This also frees a delivery that wedged before pickup,
which the injectable-only pickup grace could never release.
2026-07-15 15:16:55 +02:00
Dai Ha a97c287aee CB-108: fleet-management MCP tools (bridge_spawn / bridge_list / bridge_stop)
The primary could delegate to a worker but not create or reap one over MCP —
spawning was a raw REST POST /workers. BridgeMcp now adapts WorkerService so a
worker's whole lifecycle runs through MCP: bridge_spawn returns the new worker's
sessionId (for bridge_send) and paneId (for bridge_stop); bridge_list projects the
tracked workers; bridge_stop tears one down. The subscription boundary stays
enforced inside WorkerService (bridge_spawn surfaces a guard breach as a tool error
without touching herdr). Tools are thin static adapters, unit-tested by parity.
2026-07-15 14:22:50 +02:00
Dai Ha 8ed2370fbe CB-107: async fire-and-poll delegation (wait:false + ticket poll)
A caller's MCP client caps a blocking bridge_send at ~60s, but a real delegated
task runs for minutes. sendAsync runs the same blocking send on a background
virtual thread and returns a ticket; poll(ticket) reports pending/done/failed.
Async reuses the blocking path (and its per-target serialization), so it inherits
reply + completion resolution for free. Surfaces: REST POST message wait:false ->
202 {ticket} + GET /tasks/{ticket}; MCP bridge_send wait flag + new bridge_poll.
Terminal tickets are pruned after a TTL so the registry stays bounded.
2026-07-15 14:19:14 +02:00
Dai Ha 5b26caca0c CB-106: completion fallback — resolve a send when the worker's turn ends without bridge_reply
The blocking send previously resolved only on an explicit bridge_reply; a real
delegated task (edit files, run a build) finishes and goes idle without ever
calling it, so the send always timed out. The injector now reports a confirmed
working -> idle turn boundary via a TurnListener; CompletionResolver scrapes the
worker's transcript tail and resolves the awaiting send (Rendezvous.resolveCompletion,
Kind.COMPLETION -> Outcome.COMPLETED_UNREPLIED), surfaced as replySource=transcript
at the REST/MCP edges. Completion is synthesized only from a confirmed turn (a
sampled 'working'), never from the pickup-grace path, so it can't race the explicit
reply or fire on a turn that never ran.
2026-07-15 14:14:04 +02:00
Dai Ha 6988bfe88f CB-103: injector submits the prompt (Enter as a separate keystroke)
A delegated message was delivered into the worker's input box but never
submitted, so no task was ever processed: bridge_send blocked until timeout
while the worker sat idle with the prompt typed but not entered.

herdr's agent.send delivers text as a bracketed paste; a trailing carriage
return in that same call is swallowed as literal newline content, not Enter.
AgentControl.send now emits two keystroke events — the payload, then a
standalone "\r" — so the worker actually submits and runs the task.

Verified live (Claude Code v2.1.210, off-sub worker): full hands-off
bridge_send -> auto-submit -> worker computes -> bridge_reply round trip
returns the reply in ~41s. 62 tests green.
2026-07-15 09:40:18 +02:00
Dai Ha 07722a6007 CB-1xx: worker launch via ccs + inline bridge MCP mount (step 4) + port 8765
Spawn a real worker with 'ccs ltms-local' (profile sets CLAUDE_CONFIG_DIR + off-sub base_url; auto mode preconfigured as defaultMode:auto). bridged appends the bridge MCP mount (--mcp-config, inline JSON) and the reply charter (--append-system-prompt) as launch FLAGS — non-invasive, nothing written to the worker's profile (safer than provisioning its config dir, which would clobber it). Identity is connection-based so the mount is shared. config: worker.mcpUrl. Default bind port 8080 -> 8765. The reply charter is guidance; the send timeout catches a non-cooperative worker. 60 tests green, IDE-clean.
2026-07-15 09:19:31 +02:00
Dai Ha fa570ab32a CB-105: connection-based MCP caller identity (peer PID → herdr pane)
Resolve who is calling an MCP tool from the connection, not a spoofable argument (per docs/MCP-Contract.md). ConnectionIdentity ties the loopback peer PID (LsofPeerPidLookup) to a herdr pane (PaneLocator via pane.list/pane.process_info) → the caller's terminal_id; a caller owning no pane is the primary. bridge_reply now takes only content and resolves the worker from the connection (no sessionId). Live-verified: real PID→pane (contract test), and a non-pane MCP caller correctly gets a workers-only error. Single-host; the token path stays the split-host fallback. 58 tests green, IDE-clean.
2026-07-15 09:19:31 +02:00
kevin 46f87b9cd3 wiki: sync docs with MCP contract (idle-edge turn-done, bridge_list/status naming)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 20:01:23 +07:00
Dai Ha 8ffbfd1f06 CB-105: MCP server (bridge_send/bridge_reply/bridge_status) over the REST core
Streamable-HTTP MCP server (io.modelcontextprotocol.sdk:mcp 2.0.0) mounted on the daemon's Jetty at /mcp, exposing three tools as thin adapters over MessageService/Rendezvous — the primary calls bridge_send/bridge_status, the worker calls bridge_reply. Tool logic in unit-testable static methods (parity tests); SDK owns the wire protocol. Resolves the Jackson 2/3 split by pinning jackson-annotations 3.0-rc5 (works for both our Jackson 2.19 and the SDK's Jackson 3). Live-verified: initialize handshake + tools/list return all three tools. Documents the residual Jackson-3 CVE (loopback, trusted clients). 52 tests green, IDE-clean.
2026-07-14 14:21:11 +02:00
Dai Ha 597ac2e562 CB-104: blocking bridge_send with rendezvous reply (POST /sessions/{id}/message + /reply)
Per-session-serialized blocking send that enqueues via the CB-103 injector (poller delivers) and blocks on a rendezvous resolved by the worker's structured bridge_reply, or a typed 200/202 outcome. No status-polling completion, no terminal scrape. GET /sessions/{id}/status. Reworked from an initial poll+scrape draft after a high-effort review found the polling completion unreliable; all findings fixed.
2026-07-14 14:21:11 +02:00
Dai Ha bc04637694 deps: bump to latest patched (jackson 2.19.0, javalin 6.7.0, jetty 11.0.25, logback 1.5.18)
Clears jetty CVE-2024-8184/CVE-2024-6763. Pins all Jetty modules via jetty-bom (no skew). Documents residual advisories with no upstream fix (jetty-http 11.x EOL, logback config-file CVEs, jackson WS-2026-0003) as accepted for this loopback daemon. CLAUDE.md: validate CVEs with the jetbrains analyzer (intellij-index is stale after pom edits).
2026-07-14 14:21:11 +02:00
kevin c7b58f2195 docs: add Team.md under docs/
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 19:14:25 +07:00
kevin 01cfda7965 Add MCP design doc 2026-07-14 12:30:20 +07:00
118 changed files with 17688 additions and 327 deletions
+109
View File
@@ -0,0 +1,109 @@
---
name: implementer
description: Implementer-role procedure for a bridged worker — verify your worktree, implement the scope, commit, push, open your own PR, and hand off the PR URL. Load this when the lead delegates you an implementation task over bridged.
---
# Implementer worker — procedure
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
never merge, never commit `.mcp.json` or `wiki/`) is in **`CLAUDE.md` → Bridge communication →
Worker** and already applies. This skill is only the *implement-and-hand-off procedure*.
You run in an **isolated git worktree on your own branch** — a full peer of the primary (same
repo, `CLAUDE.md`, skills, MCP), differing in the model behind you and the branch you sit on.
The worktree model is documented in [`docs/Worker-Git-Workflow.md`](../../../docs/Worker-Git-Workflow.md).
## 1. Confirm where you are
Before touching anything:
```bash
git rev-parse --show-toplevel # your worktree root — NOT the primary's main tree
git branch --show-current # your dedicated branch: worker/<ticket>-<nonce>
git status # should be clean at the start
```
Do **all** work here, on this branch. Never `git checkout main`, never rebase onto or push to
`main`. The branch is your isolation — respect it.
## 2. Implement
- Implement exactly the scope the lead named. Keep the diff focused; note anything out of scope
in your reply instead of widening it.
- Match the surrounding code's style, naming, and idioms.
- Run whatever build/test you can — `mvn clean install` from the module root. Read its **full**
output; a piped `mvn ... | tail` hides failures.
## 3. Commit
```bash
git add <the files you changed> # explicitly — never `git add -A` / `git add .`
git commit -m "<ticket>: <clear one-line summary>"
```
`.mcp.json` will show as modified. Leave it — it is `--skip-worktree` and not yours to commit.
## 4. Push
```bash
git push -u origin HEAD
```
Push is over SSH as the same user — no extra credential needed. Never force-push over anything
you did not create.
## 5. Open your own PR to `main`
Via the gitea REST API. The daemon injected a **repo-scoped token** (`GITEA_TOKEN`) and the forge
host (`GITEA_HOST`) into your env for exactly this — the token can create a PR but **cannot
merge**.
```bash
API="${GITEA_HOST%/}/api/v1/repos/lms/claude-bridge/pulls"
BRANCH="$(git branch --show-current)"
curl -sS -X POST "$API" \
-H "Authorization: token ${GITEA_TOKEN}" \
-H "Content-Type: application/json" \
-d "$(cat <<JSON
{"head": "${BRANCH}", "base": "main",
"title": "<ticket>: <concise change summary>",
"body": "<what changed and why; reference the ticket; note tests run and their result>"}
JSON
)"
```
The response JSON carries `"html_url"` — that is your PR URL. On a non-2xx, read the error body,
fix it if the cause is yours (e.g. branch not pushed yet), and report the failure rather than
inventing a URL. If `GITEA_TOKEN` is unset your profile was not granted PR-create: push the branch
and report its name so the lead opens the PR.
## 6. Hand off — what goes in `bridge_reply`
The reply is the entire handoff; the lead cannot see your terminal.
```
PR: <html_url from step 5, or "not created: <reason>" + branch name>
branch: <your branch>
files: <the files you changed>
tests: <what you ran and its REAL result — or "not run: <why>">
summary: <2-3 lines: what you implemented and any caveat the reviewer needs>
```
```mermaid
sequenceDiagram
autonumber
participant L as Lead
participant I as Implementer (you)
participant G as git / gitea
L->>I: delegated task (you are in a worktree on your branch)
I->>I: implement + build/test here
I->>G: git commit (never .mcp.json / wiki)
I->>G: git push -u origin HEAD
I->>G: POST /pulls (GITEA_TOKEN) — open PR to main
G-->>I: html_url
I->>L: bridge_reply(PR url, branch, files, tests)
Note over L,G: the lead reviews the PR and merges on green — you never merge
```
*The implement turn: work in the worktree, commit → push → open the PR, hand off the URL.*
+51
View File
@@ -0,0 +1,51 @@
---
name: reviewer
description: Reviewer-role procedure for a bridged worker — how to work a review scope and the exact shape of the finding to report. Load this when the lead delegates you a code review over bridged.
---
# Reviewer worker — procedure
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and already applies. This
skill is only the *review procedure*: how to work the scope, and the exact shape of what you
send back.
## 1. Read the whole scope before you judge
The delegation names your scope — a file, a diff, a PR, a function. **Read all of it first.**
A review that fires on a snippet misses the caller that makes it safe (or the one that makes it
a bug). Reviewing part of the scope and guessing the rest is the most common way a reviewer is
wrong.
## 2. Stay in the scope
- Review **only** what you were assigned. Something elsewhere looks wrong? One line in your
reply — do not go hunt it. Wandering is how two reviewers report the same thing and neither
covers what it was given.
- Do **not** edit files or run the build. You review; the owner acts.
## 3. Reach for `bridge_ask` only for a genuine fork
Ambiguous requirement, a missing acceptance criterion, "intended or a bug?", or two defensible
fixes with different consequences — those are the lead's call, and guessing produces a
confident-but-wrong finding. Anything you could settle by reading more code is yours to settle.
## 4. The finding — what goes in `bridge_reply`
Report the **single most important** real issue in the scope, in these four lines, under
~90 words:
```
1. <path>:<line>
2. issue: <one sentence — what is wrong and why it matters>
3. fix: <one line — the concrete change>
4. severity: high | medium | low
```
- **Nothing real after reading?** Reply `NO ISSUE` and one line saying why. A clean review is a
valid result; a fabricated issue is worse than none.
- **Severity:** `high` = wrong result, data loss, security, or a hang/crash on a real path ·
`medium` = a real bug on an edge path, or a correctness risk under load/concurrency ·
`low` = clarity, a latent foot-gun, or a smell with no current failure.
- Be specific and verifiable: a line number and a one-line repro beat an adjective. If you can't
point at where it goes wrong, you haven't found it yet.
+55
View File
@@ -0,0 +1,55 @@
name: CI
on:
push:
branches: [main]
pull_request:
branches: [main]
jobs:
build:
runs-on: ubuntu-latest
steps:
# The wiki submodule is docs only and is not needed to build — leave it unfetched so CI
# does not depend on the wiki repo being reachable.
- uses: actions/checkout@v4
# The runner image ships an older default-jdk; bridged sets maven.compiler.release=25, so
# provision the JDK explicitly rather than apt-installing whatever "default" means today.
- name: Set up JDK 25
uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '25'
cache: maven
# setup-java provisions the JDK only — it does NOT install Maven, and the runner image has
# no mvn on PATH (a bare `mvn` exits 127). Install it separately. apt pulls a default JRE as
# a dependency; JAVA_HOME from setup-java still wins, which the version check below proves.
- name: Install Maven
run: |
apt-get update && apt-get install -y --no-install-recommends maven
mvn -version
- name: Build and test
working-directory: bridged
# This IS the mock-socket surface CB-503 asks for: the pom's `default-excludes` profile
# already sets excludedGroups=contract, so the @Tag("contract") tests — which need a live
# herdr socket and a RabbitMQ container — are excluded without any flag here. Everything
# that runs does so against the fake UDS herdr and fake ccs/claude stubs.
run: mvn -B clean install
# Deliberately NOT actions/upload-artifact: this Gitea instance presents as GHES, and
# @actions/artifact v2+ (i.e. upload-artifact@v4) refuses to run there —
# "GHESNotSupportedError ... not currently supported on GHES", which red-Xes an otherwise
# green build. Since the artifact could not be retrieved anyway, dump the failing tests into
# the log instead, where they are actually readable.
- name: Failing test output
if: failure()
working-directory: bridged
run: |
for f in target/surefire-reports/*.txt; do
[ -f "$f" ] || continue
grep -qE "Failures: [1-9]|Errors: [1-9]" "$f" && { echo "===== $f ====="; cat "$f"; }
done
exit 0
+198
View File
@@ -1,7 +1,196 @@
# claude-bridge — project instructions
## Bridge communication (enforced — read this first)
> **Canonical block.** Everything down to §Layering is the portable bridge charter, copied verbatim
> into every project that mounts the bridge MCP. Keep it byte-identical with the template in the
> wiki ([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable
> CLAUDE.md block*); improvements go to the template first, then out to each project. Anything
> specific to *this* repo lives under §Project addendum below, never inline above it.
If no `bridge_*` MCP tools are mounted in this session, this section does not apply — skip it.
`bridged` is the **sole communication gateway** between agents here. The orchestrating session (the
**primary**) and every delegated peer (a **worker**) mount the *same* MCP server and talk only
through its `bridge_*` tools. No session addresses a peer, a broker, or the network directly.
### Which role am I? — settle this before acting
**Both roles read this file.** A worker runs in a git worktree of this same repo, so it inherits
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
**Call `bridge_whoami`.** It returns `{"role":"primary"}` or `{"role":"worker","sessionId":…,
"profile":…,"worktree":…,"branch":…}`, resolved by the daemon from your connection — unforgeable,
and the same resolution its authorization gate uses. Don't infer what you can ask.
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
fires: the reply charter in your system prompt (*"You are an off-subscription worker in the
claude-bridge fleet"*) ⇒ **worker**; bridge tools prefixed `mcp__bridge__*` ⇒ **worker** (the
launcher fixes that mount name; a primary's mount is named by whoever wrote its `.mcp.json`, so it
varies); `ANTHROPIC_BASE_URL` set ⇒ **worker** (Claude-model workers run on a clean env, so its
*absence* proves nothing). **Still unsure ⇒ act as a worker.** The two mistakes are not symmetric: a
primary acting as a worker is refused by the authorization gate — loud and self-correcting — while a
worker acting as the primary ends its turn with no `bridge_reply`, and the sender silently receives
nothing. Fail toward the recoverable error.
### Invariants — both roles, no exceptions
1. **Never set, export, or forward `ANTHROPIC_BASE_URL`** (or `ANTHROPIC_AUTH_TOKEN`). The primary
stays on subscription; only the bridge puts a worker off it, at spawn. Mounting the bridge must
never move a session across that boundary.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/send/drain are primary-only; reply/ask are
worker-only-and-only-as-itself. A call outside your role is refused, not queued.
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
re-send because a call looks slow — the bridge delivers when the peer is `idle`/`blocked`.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
### Primary (lead) — run this on every task, in order
**Delegate by default — that is the job.** With the bridge mounted you are an orchestrator on a
metered subscription, and workers are cheap, parallel, and disposable. The default answer to "who
does this?" is **a worker**, not you. Reach for `bridge_send` before you reach for `Edit`. The steps
below are the procedure — run them in order, every task, not only the big ones.
0. **Know your role** — `bridge_whoami`, once per session, before anything else.
1. **Split.** Write the unit list. Every unit carries: scope · the files or PR in question ·
acceptance criteria · exactly what to report back. A unit with no acceptance criteria is not
ready to delegate — refine it or keep it.
2. **Gate each unit** on one question: **"can I write a brief good enough for a worker to
succeed?"** — *not* "could I do this faster myself?" (usually you could; doing it yourself costs
your context and your subscription, while a wasted worker turn costs a worker turn). Yes ⇒
delegate. The keep-list is closed: the conversation with the user, decomposition and planning,
the final judgment call, verification, merges, and anything that depends on context only you
hold. Nothing else is yours by default.
3. **Spawn every delegated unit first** — `bridge_spawn{profile, worktree:true, ticket}`, one per
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model and cost, not in
tier, so the default is rarely what you want.
4. **Then send them all** — `bridge_send{sessionId, content, wait:false}`. Line 1 of every brief is
`Load the <name> skill.` naming the worker's playbook; those skills are opt-in and that line is
what makes them reliable. Where the project ships no such skill, spell the procedure out in the
brief instead. The brief is self-contained — the worker sees your message and the repo, nothing
of your context, your plan, or your screen.
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`bridge_status`, never by reading its terminal.
6. **Verify yourself.** Re-run the build and the checks. A worker mounts only the bridge MCP and
cannot run your other tooling, and a piped command (`… | tail`) hides failures behind a zero
exit — never promote a worker's "clean" to a fact.
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
as it lands; don't wait for the last implementer. Under ~50 changed lines, skip the fan-out and
read it yourself.
8. **Adjudicate, merge, tear down — yours alone.** Read the diff yourself: fully if it is small,
targeted at the reported findings and the risky paths if it is large. Reviewer findings direct
your attention; they never substitute for it. Then merge, then `bridge_stop{paneId}`.
**Steps 3 and 4 are separate on purpose** — spawning and sending in one loop is how parallel work
silently becomes serial, and it is the most common way this layer is wasted. For the same reason,
prefer `wait:false` + `bridge_poll` for anything non-trivial: a blocking `bridge_send` is capped by
*your own* MCP client call timeout (~60s), well below the task's real runtime.
**Delegating does not delegate responsibility.** Workers open PRs; you are the gate. Never delegate
the merge — and merging on a reviewer's word is delegating it by proxy.
| Intent | Tool |
|---|---|
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a worker | `bridge_spawn{profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` · one worker's state: `bridge_status{sessionId}` |
| Delegate (blocking) | `bridge_send{sessionId, content}` |
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
| Answer a worker's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
| Tear down | `bridge_stop{paneId}` |
### Worker — the turn contract
1. **Load the playbook skill the lead named** before doing anything else.
2. **Do the assigned scope only.** Note anything you spot outside it in one line; don't go hunt it.
3. **`bridge_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
defensible fixes, "bug or intended?"). It blocks and you resume the *same* turn with the answer.
Don't ask what you could decide yourself.
4. **End the turn with exactly one `bridge_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `bridge_reply` ⇒ the sender gets nothing and the exchange stalls.
5. **Report honestly.** State only what you actually ran and its real output, including failures.
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
so never claim the result of a check you had no way to run.
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
### Where each rule lives (don't duplicate — extend the right layer)
| Layer | Scope | Reaches |
|---|---|---|
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every worker, at launch, every peer kind |
| **this section** | protocol + orchestration policy | primary **and** every Claude worker — tracked in git, so worktrees inherit it |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a worker told to load one |
| the bridge's own docs | design detail, flows, error model | on demand |
A rule belongs in **exactly one** layer — the outermost one that must obey it. Peers that don't read
`CLAUDE.md` (non-Claude adapters) get the charter only, so any rule *they* must obey belongs in the
charter, not here.
## Project addendum — claude-bridge (not part of the canonical block)
- **This repo is the bridge.** The daemon is `bridged`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/BridgeMcp` (tools), `auth/Authz` (the role table),
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
`reviewer` (scoped review → one structured finding). Name one in every delegation.
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
session's context.
### The prompt is part of the product — update it with the code (mandatory)
This repo *is* the bridge, so the canonical block above is not documentation about someone else's
system: it is the instruction surface this codebase ships. **Every change here must end by asking
whether the block still tells the truth.** A code change that silently invalidates it is an
incomplete change — the agents reading it have no other source.
Before you call any work done, check the row that matches what you touched:
| You changed… | Re-read and update… |
|---|---|
| a `bridge_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
| `Authz` / the role table | invariant 3, and the primary-only vs worker-only claims |
| `ConnectionIdentity` / how a caller is resolved | the `bridge_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__bridge__*`), and the layering table's top row |
| the injector / status gating | invariant 4 |
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
Then **propagate**: the block in this file and the template in the wiki
([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
block*) must stay byte-identical, and other projects carrying the block need the same edit. Verify
rather than trust:
```bash
python3 - <<'PY'
import pathlib
c = pathlib.Path("CLAUDE.md").read_text()
w = pathlib.Path("wiki/7-Use-Cases.md").read_text()
S, E = "## Bridge communication (enforced", "## Project addendum — claude-bridge"
block = c[c.index(S):c.index(E)].rstrip() + "\n"
i = w.index("```markdown\n") + len("```markdown\n")
print("in sync:", w[i:w.index("\n```\n", i) + 1] == block)
PY
```
## IDE MCP tools & validation workflow (enforced)
> **Primary only.** Workers have no IDE MCP mount — if you are a worker, skip this section and
> report the build/test output you actually ran (see §Bridge communication → Worker).
Two IDE MCP servers are connected: **intellij-index** (semantic code intelligence) and
**jetbrains** (file problems, reformat, debugger). IntelliJ has multiple projects open; our
module is **`bridged`**. Always pass these to IDE MCP tools:
@@ -20,6 +209,15 @@ module is **`bridged`**. Always pass these to IDE MCP tools:
test run). A per-file-clean file can still break the build or another module. This is the
whole-project gate before declaring work done or committing.
**Whenever dependencies change (or a `pom.xml` edit), validate CVEs with
`jetbrains get_file_problems{filePath: "bridged/pom.xml"}`** — its Mend.io check reflects the
dependencies on disk. (Note: `ide_diagnostics` / intellij-index does NOT re-resolve dependencies
after a pom edit without a full Maven reimport, so it reports stale CVE results — don't trust it
for this.) Treat a CVE warning like any other: bump to a patched version and confirm
`mvn clean install` still passes. If the latest available version is still flagged (EOL line,
"insufficient information", or config-file-only advisories), document it as accepted in the pom
rather than chasing a fix that doesn't exist.
### Use IDE MCP tools for navigation, refactoring, and diagnostics only
- **Navigate (prefer over Grep/Read for symbols):** `ide_find_definition`, `ide_find_class`,
+38 -5
View File
@@ -50,8 +50,9 @@ flowchart LR
plain daemon (no Anthropic quota), so it may poll/subscribe freely.
- **One gateway (unified MCP setup):** `bridged` is the **sole communication path** for every
Claude session. Primary and workers each mount it as an MCP server (one `claude mcp add`
line, same on both) and talk over MCP tools — `bridge_send` / `bridge_reply` / `bridge_ask` /
`bridge_status`. **No Claude session ever addresses a broker, a peer, or the network
line, same on both) and talk over MCP tools — `bridge_send` / `bridge_reply` /
`bridge_status` (with `bridge_ask` planned for the blocked-worker path). **No Claude session
ever addresses a broker, a peer, or the network
directly**; any queue is `bridged`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so
mounting the bridge is subscription-safe by construction.
- **How the primary consumes a reply:** a single **blocking MCP call** (`bridge_send`);
@@ -85,6 +86,38 @@ Gitea wiki.
## Status
🟢 Design — herdr-centric **`bridged`** message server selected as the primary approach
(2026-07-11), superseding the AgentAPI plan (2026-07-08). AgentAPI retained as fallback
injector.
🟢 **Implemented & dogfooded** — the herdr-centric **`bridged`** message server is built and in
real use: an Opus primary delegates tasks to off-subscription workers that reply through the
bridge (code reviews delegated this way have produced committed bug fixes). Selected as the
primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a
fallback injector.
**Shipped** (Java 25 · Maven · 266 unit/acceptance tests green; the live-herdr and broker contract
tests run separately via `mvn test -Pcontract`):
- **Core gateway** — herdr socket client (contract-tested vs live 0.7.0); guard-checked worker
spawn with `ANTHROPIC_BASE_URL` injected only into the worker's env; status-gated injector;
blocking `bridge_send` with reply rendezvous; MCP server as a thin adapter over the REST core.
- **MCP tools** — `bridge_send` / `bridge_reply` / `bridge_status` (messaging) and `bridge_spawn`
/ `bridge_list` / `bridge_stop` / `bridge_profiles` / `bridge_poll` (fleet). Caller identity is
connection-based (loopback peer PID → herdr pane), so the same mount serves primary and workers.
- **Delivery reliability** — completion fallback (a confirmed `working→idle` turn resolves a
send); async fire-and-poll (beats the caller's MCP call timeout for long tasks); and failure
detection for wedged (`unknown`), vanished, and never-ready workers so a send never hangs.
- **Fleet** — multiple worker profiles, each with an independent base_url guard check; workers
inherit the primary's working directory (never `$HOME`); a readiness gate holds delivery until
a worker's Claude has connected the bridge MCP (no paste lost into its boot window).
- **Blocked-worker path** — `bridge_ask` reverse rendezvous: a worker pauses its delegated turn to
ask the primary and resumes the *same* turn with the answer (CB-205).
- **Session lifecycle** — session manager with spawn/reuse/recycle, `idle_ttl` reaper, `context_cap`,
and graceful drain on shutdown (CB-301/CB-303); per-worker git worktrees on their own branch with
a config-parity overlay, so parallel implementers never stomp each other (CB-301-ext).
- **Reliable worker→primary delivery** — a durable `ReplyInbox` (in-memory by default, AMQP/LavinMQ
for cross-restart durability) holds a reply that arrives with no open send, and an active
status-gated push loop nudges the primary to drain it (CB-307).
- **Pluggable peers** — a `PeerLauncher` SPI with two in-tree adapters, `claude-code` and `opencode`,
routed by a `kind:` discriminator (CB-401/CB-402).
**Next** (see the [roadmap](wiki/8-Roadmap.md)) — Stage 5 hardening (auth/TLS, `/metrics`, CI,
service supervision, per-session authz + audit), then cross-host: CB-308 multi-host federation and
CB-500 multi-tier coordination.
+3
View File
@@ -5,6 +5,9 @@ dependency-reduced-pom.xml
# Local runtime config (copy from bridged.example.yaml)
bridged.yaml
# CB-505 audit trail + daemon stdout/stderr — runtime records, never source
logs/
# Editor / OS
*.iml
.idea/
+185 -15
View File
@@ -3,31 +3,201 @@
# bridged is the sole gateway between primary/worker Claude sessions and herdr.
# It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL.
# REST + MCP listen address. Keep it on loopback — bridged is same-host in Stage-1.
# REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token`
# below — bridged REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
bind:
host: 127.0.0.1
port: 8080
port: 8765
# API authentication (CB-501). Governs how a caller that is NOT an on-host worker pane proves it
# is the primary. Worker identity never depends on this: a loopback peer PID that maps to a herdr
# pane is unforgeable and is always honoured, so turning auth on cannot lock the fleet out.
#
# mode: loopback-trust → DEFAULT, and the historical behaviour: any loopback caller that is not
# a worker is the primary, no credential needed. Sound ONLY because the
# OS refuses remote connections to a loopback socket.
# mode: token → such a caller must send `Authorization: Bearer <token>`; without it it
# is anonymous and authorized for nothing. REQUIRED for a non-loopback
# bind — the daemon fails fast otherwise, because "unauthenticated ⇒
# primary" on a reachable port would hand spawn/stop/send to anyone.
# tokenEnv → host env var holding the token (never the literal value). Default
# BRIDGED_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
#
# TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and
# let it own certificate lifecycle, e.g.
# location / { proxy_pass http://127.0.0.1:8765; proxy_set_header Authorization $http_authorization; }
# The broker link gets TLS from its own URI (amqps://…) — see `broker` below.
# auth:
# mode: token
# tokenEnv: BRIDGED_API_TOKEN
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
# How a worker session is spawned. Stage-1 uses the existing ccs `ltms-local`
# profile, whose .claude.json routes to the gx00 vLLM below.
worker:
profile: ltms-local
baseUrl: http://gx00.gw:8000 # the gx00 vLLM (models: coder / deepseek-v4-flash)
model: coder
# Placement: each worker lands in its OWN tab inside a dedicated worker space, so it
# never splits or clutters your real work spaces. Use `pane` for the legacy behaviour
# (split the currently-focused tab).
placement: tab # tab | pane
workspace: bridged-workers # the dedicated worker space (found-or-created, shared)
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
# How worker sessions are spawned. Define one or more named profiles (backends) under
# `workers`; each key is the profile name (also the ccs profile). `defaultWorker` picks
# which one a no-argument spawn uses (bridge_spawn with no profile / POST /workers).
#
# Shared knobs (placement/workspace/tabLabel) can be repeated per profile; they usually match.
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
# Use `pane` for the legacy behaviour (split the focused tab).
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
# (--append-system-prompt) as launch flags; nothing is written to the profile.
# tokenEnv → host env var holding the worker's auth token (value never stored in config);
# omit for a backend that needs no token (e.g. a local ollama).
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
# cwd on an MCP spawn, else the daemon's cwd — never $HOME. See
# docs/Worker-Startup-and-Trust.md.
# configDir → CLAUDE_CONFIG_DIR for the worker, so it inherits that profile's
# skills/MCP/hooks. Omit to leave the worker on the host default.
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
# worktree sees the same local config (CB-301-ext). Omit for the default set:
# [.mcp.json, .claude/settings.local.json, .env, .envrc].
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
# Opt-in by design — omit and the worker gets no PR-create grant (push over
# SSH is unaffected). The token value itself is never stored in this file.
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
# env → extra environment for this profile's workers, as a literal key/value map
# (CB-511). Use it to give workers a toolchain.
#
# A worker's environment does NOT come from your shell. bridged hands herdr an
# explicit env map and herdr merges it into ITS OWN process env — so before
# CB-511 a worker inherited whatever PATH the herdr server happened to be
# started with, which on a long-lived herdr can predate your toolchain entirely
# and leave workers unable to run `mvn` or `java` at all.
# bridged now propagates ITS OWN PATH to every worker by default; set `env:`
# only to override that or add more (JAVA_HOME, …). Since the default is the
# daemon's PATH, make sure the daemon is started with a good one — see the PATH
# lines in deploy/dev.ltms.bridged.plist and deploy/bridged.service.
#
# Adapter-owned variables always win over `env:`: ANTHROPIC_BASE_URL and the
# rest of the ANTHROPIC_*/CLAUDE_* wiring are applied after it, so an `env:`
# entry cannot repoint a worker past the SubscriptionGuard — which is checked
# against `baseUrl` alone.
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
workers:
gx10: # ccs profile name (NOT a hostname)
kind: claude-code # which adapter spawns this profile (default; may omit)
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
model: coder
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".mcp.json", ".claude/settings.local.json", ".env", ".envrc"]
ollama:
baseUrl: http://ollama.ltms.dev # local/self-hosted; usually no token
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}"
mcpUrl: http://127.0.0.1:8765/mcp
argv: ["ccs", "ollama"]
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
# The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a
# `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude.
#
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → bridge_send →
# structured bridge_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
# and need NO credentials — check `opencode models` for the current free list, since the names
# change. That also makes the worker off-subscription by construction.
# opencode-free:
# kind: opencode
# model: opencode/north-mini-code-free # `provider/model` selector, injected as `-m`
# placement: tab
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
#
# CB-508: point an opencode profile at your OWN OpenAI-compatible endpoint (local vLLM, llama.cpp,
# LM Studio, TGI…) instead of opencode's gateway. Setting `baseUrl` on a `kind: opencode` profile
# makes the bridge emit a custom `provider` block into the generated opencode.json — opencode has
# no ANTHROPIC_BASE_URL seam, so this is how the endpoint is pinned.
# baseUrl → a bare host:port gets `/v1` appended (where these servers mount the API); a URL that
# already has a path is used verbatim, so a custom mount point still works.
# model → MUST be "<provider>/<model>". The provider half names the generated block; the model
# half must match an id the server reports at /v1/models. One field drives both the
# declaration and the `-m` flag, so they cannot drift apart. A bare model name with a
# baseUrl set is rejected at spawn rather than silently using the default gateway.
# tokenEnv → optional; its value becomes the provider apiKey. Most local servers ignore the key,
# so a placeholder is used when unset (the AI SDK still requires a non-empty one).
# NOTE: no `guard` entry is needed even with a baseUrl set. The SubscriptionGuard exists to stop a
# worker borrowing the primary's Anthropic subscription, and an opencode process has no Anthropic
# credential path at all.
# opencode-local:
# kind: opencode
# baseUrl: http://127.0.0.1:8000
# model: local-vllm/deepseek-v4-flash
# placement: tab
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
defaultWorker: gx10
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
# must carry none. Grounded in ltms-local's real endpoints.
# must carry none. Every profile above must have its host listed here.
guard:
offSubscriptionHosts:
- gx00.gw
- gx01.gw
- ollama.ltms.dev
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
# unknown keys are ignored, so a snake_case key would be silently dropped (default kept).
# spawnReadyTimeoutMs: 20000
# spawnReadyPollMs: 300
# Worktree provisioning root (CB-301-ext). Where per-worker git worktrees are checked out so
# each worker owns an isolated branch instead of sharing the primary's tree. Omit to default
# to a sibling directory of the repo root.
# worktreeRoot: /Users/me/src/.bridged-worktrees
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
# contextCap → force-release a session after this many delegated turns
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
# lifecycle:
# idleTtlSeconds: 300
# contextCap: 10
# drainTimeoutSeconds: 5
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
# Set a broker uri to swap in the AMQP-backed inbox: worker replies with no open send are held
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
# broker:
# uri: amqp://guest:guest@127.0.0.1:5672
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
# pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop
# stops as soon as the primary's inbox is empty.
# terminal → pin the primary's herdr terminal id. Omit to learn it from the connection on
# the first orchestration-side MCP call (the normal case). An off-host or
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
# degrades to pull; the reply is still never lost.
# pushReminders → max nudges before giving up (default 5)
# pushBackoffMs → delay between nudges in ms (default 15000)
# primary:
# terminal: term_65619bd6174568
# pushReminders: 5
# pushBackoffMs: 15000
+142
View File
@@ -0,0 +1,142 @@
# CB-307 — Active push-to-primary + reminder loop (the reliability layer)
**Status:** design (2026-07-19). Builds directly on the shipped durable landing zone
(`AmqpReplyInbox`, main `2bc5f3a`, dogfooded live). gitea #5.
## Why this exists
Stage 2 gave a worker→primary reply a **durable place to wait** when no `bridge_send` is
open: it lands in `agent.<target>.inbox` on the broker and survives a daemon bounce. But
delivery is still **pull** — the primary only sees the reply if it happens to call
`bridge_poll(target)` / `GET /sessions/{id}/replies`. A reply can sit indefinitely while
the primary works on something else.
This layer makes delivery **active**: the bridge *pushes* a nudge to the primary the moment
a reply lands, and keeps reminding (bounded) until the primary drains it. At-least-once,
dedup by `msgId`, and — critically — it never loses the reply even if every push fails,
because the durable inbox is the backstop.
## The hard constraint it works around
The bridge is an MCP **server**; the primary is an MCP **client**. A server cannot call
into a client. So "push to the primary" cannot be an MCP response — it needs a *sideband*
channel. The chosen channel: **inject a synthetic user-turn into the primary's own herdr
terminal pane** — the same mechanism the bridge already uses to deliver tasks to workers,
pointed at the primary's pane instead.
```mermaid
flowchart LR
W["worker"] -->|"bridge_reply (no open send)"| MS["MessageService.reply"]
MS -->|"inbox.publish"| INBOX[("agent.&lt;target&gt;.inbox<br/>(durable, LavinMQ)")]
MS -->|"notify"| LOOP["ReplyPushLoop"]
LOOP -->|"status-gated inject"| PANE["primary's herdr pane"]
PANE -->|"primary drains"| DRAIN["bridge_poll(target)<br/>= peek + ack"]
DRAIN -->|"inbox now empty"| LOOP
LOOP -.->|"still non-empty →<br/>re-inject on backoff"| PANE
classDef store fill:#2c5282,stroke:#1a365d,color:#ffffff;
class INBOX store
```
*Figure 1 — a reply lands in the durable inbox; the push loop nudges the primary's pane;
the primary's drain acks it; a still-full inbox triggers a bounded re-nudge.*
## Three increments
### Increment 1 — learn & store the primary's terminal_id
**Finding (seam map):** `ConnectionIdentity.resolve(remoteAddr, remotePort)` already returns
the caller's herdr `terminal_id` for *every* MCP call, via `PaneLocator.terminalForPid`
(walks `pane.list`, matches the caller PID to a pane's process tree). It is non-null whenever
the caller runs in a herdr pane on this host. Today it's discarded for the primary
(`presence.markPresent` is a no-op on it).
**Plan:** a single-slot `PrimaryRegistry` (thread-safe) holding the primary's `terminal_id`.
Populate it from the **orchestration-side** MCP tools — `bridge_send`, `bridge_spawn`,
`bridge_poll`, `bridge_list`, `bridge_status`, `bridge_profiles` — capturing
`callerTerminal(exchange)` when it is (a) non-null and (b) **not** a registered worker
session in `SessionManager`. That caller is, by construction, the primary. Worker-side tools
(`bridge_reply`, `bridge_ask`) never set it.
- **Config override / pin:** a `primary: { terminal: "<id>" }` block in `BridgedConfig`
(nested record, same shape as `Broker`). Lets an operator pin it, or supply it when
derivation can't (see degrade case).
- **Degrade:** if the primary is off-host or in a non-herdr terminal, `terminalForPid`
returns null and no override is set → **the registry stays empty → the push loop is a
no-op and we fall back to pull** (today's behaviour). The reply is never lost; it's just
not actively pushed. This is a safe, explicit degradation, not a failure.
### Increment 2 — the push loop
A `ReplyPushLoop` component, notified at the single no-waiter call site
(`MessageService.reply` → the `inbox.publish` branch, `MessageService.java:192`).
- **Inject a nudge, not the payload.** The injected turn tells the primary *to drain*
(e.g. "Worker `<target>` returned a reply — run `bridge_poll(target=<target>)` to collect
it"), it does **not** carry the reply text. Rationale: replies can be large/multiline and
terminal injection would mangle them; the drain response is the clean transport. Keeps the
push idempotent — re-nudging is harmless.
- **Ack = drain.** The primary draining (`drainReplies` = peek + ack) is the acknowledgement.
The loop's **stop condition is `inbox.peek(target).isEmpty()`** — the reply is gone from the
inbox because it was acked. No new `bridge_ack` tool needed for v1 (see Increment 3).
- **Status-gated injection (mechanism (b), chosen).** A dedicated lightweight scheduled loop,
**not** the worker `Injector`. It injects via `AgentControl.send(primaryTerminal, nudge)`
(the same herdr `agent.send` = `pane send-text` + submit that delivers to workers) only when
`AgentControl.status(primaryTerminal).injectable()` (IDLE/BLOCKED) — never mid-turn. This keeps
the primary path fully isolated from `WorkerPresence`/`StatusPoller` (which are worker-scoped),
and makes it unit-testable with a fake `AgentControl` + an injected clock (per the CB-306
`LongSupplier` clock + `Runnable` sleeper seam). Rejected (a) reuse-the-Injector: it would force
the primary terminal into the worker poller set and couple to worker-presence semantics — more
integration surface, harder to test, no real gain for a bounded reminder.
- **Bounded reminder / backoff.** While `peek(target)` stays non-empty, re-inject on a
backoff schedule up to a cap (N reminders or a max duration; config
`primary.push_reminders` / `primary.push_backoff_ms`). After the cap, **stop reminding** —
the reply remains in the durable inbox and the next natural poll (or a later worker reply's
nudge) still surfaces it. Bounded so the bridge never spams the primary.
### Increment 3 — optional per-`msgId` `bridge_ack` tool (deferred)
Drain-as-ack is coarse: it clears *all* pending replies for a target at once. If finer
control is ever needed (ack one reply, leave others held), add a `bridge_ack(msgId)` tool
mapping to `inbox.ack(target, msgId)` — the port already supports per-`msgId` ack. Not built
in v1; the stop-on-empty loop is sufficient.
## The two subtleties (decided here)
1. **Which caller is "the primary"?** Connection-derived, not self-reported: the caller whose
resolved terminal is non-null **and not a registered worker session**, seen on an
orchestration-side tool. This never mislabels a worker (workers are in `SessionManager`)
and needs no new env var or argument (identity stays connection-derived, per the existing
`BridgeMcp` invariant).
2. **Readiness-gate mismatch → dedicated loop.** The existing `Injector` gates delivery on
`ready.test(target)` = `WorkerPresence` (the *worker's* MCP connected). The primary is not
in `WorkerPresence`, so reusing `Injector` would mean forcing the primary terminal into the
worker `StatusPoller` set and swapping the `ready` predicate — extra integration surface with
worker-scoped machinery. Decision: **mechanism (b)** — a small dedicated scheduled loop that
calls `AgentControl.status(primaryTerminal).injectable()` then `AgentControl.send(...)`, with
an injected clock. Isolated from worker presence, trivially unit-testable, sufficient for a
bounded reminder. (Verified live: this primary resolves to `term_656c8cc03e1f0b1`, pane
`w2:pY` — the primary genuinely runs in a herdr pane on this host, so the path is exercisable.)
## Boundary note
This is the first time the bridge **writes into the primary's pane** — a new direction of
control. It stays within the communication-bus identity: the injection is a **nudge** (a
synthetic "go drain your replies" turn), **status-gated** so it never interrupts a turn,
**bounded** so it never spams, carries **no env** and **never crosses the subscription
boundary**. The bridge is signalling the primary that it has mail — not driving its work.
## Test plan
- **Unit (hermetic):** `PrimaryRegistry` set/clear/override; the "caller is primary iff
non-null terminal AND not a registered session" predicate; the loop's stop-on-empty and
bounded-reminder logic with an injected clock + a fake injector (no real herdr).
- **Live dogfood (primary-side):** with the daemon on the broker jar + a real worker,
delegate a task, let the worker reply after the `bridge_send` window closes, and observe the
bridge inject a drain nudge into *this* primary pane; confirm draining stops the reminders;
confirm an unreachable primary (registry empty) degrades to pull with no loss.
## Out of scope
Multi-host (CB-308) — the push loop is local-only; a remote primary is reached by its own
local gateway, not cross-host injection. Federation reuses this loop per-gateway.
+123 -4
View File
@@ -17,13 +17,74 @@
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<mainClass>dev.ltms.bridged.Bridged</mainClass>
<jackson.version>2.18.2</jackson.version>
<javalin.version>6.3.0</javalin.version>
<jackson.version>2.19.0</jackson.version>
<javalin.version>6.7.0</javalin.version>
<jetty.version>11.0.25</jetty.version>
<mcp.version>2.0.0</mcp.version>
<slf4j.version>2.0.16</slf4j.version>
<logback.version>1.5.15</logback.version>
<logback.version>1.5.18</logback.version>
<junit.version>5.11.4</junit.version>
<amqp.version>5.22.0</amqp.version>
<testcontainers.version>1.20.4</testcontainers.version>
<commons-compress.version>1.27.1</commons-compress.version>
<commons-lang3.version>3.18.0</commons-lang3.version>
</properties>
<!--
Dependency security (validate with the JetBrains analyzer's Mend.io check on this pom).
Deps are pinned to the latest available versions. Residual advisories with NO upstream fix,
accepted for this loopback-bound daemon that processes no untrusted config:
- jetty-http 11.0.25 (via Javalin): CVE-2026-2332, CVE-2025-11143 — Jetty 11 is EOL;
fixed only in Jetty 12, which needs a Javalin major (6.x rides Jetty 11).
- logback-core 1.5.18: CVE-2025-11226, CVE-2026-1225 — both require a MALICIOUS
logback.xml (attacker with config write already has code execution); ours is trusted.
- jackson-core 2.19.0: WS-2026-0003 — "insufficient information", no fixed version published.
- tools.jackson.core (Jackson 3) 3.0.3 via the MCP SDK: CVE-2026-29062 (nesting-depth
resource exhaustion). The SDK 2.0.0 is pinned to Jackson 3.0.3 + jackson-annotations
3.0-rc5; bumping Jackson 3 to the patched 3.2.x breaks the SDK (annotation mismatch).
Only the loopback /mcp endpoint parses this JSON, from trusted local Claude clients.
The 11.0.23 -> 11.0.25 bump did clear jetty CVE-2024-8184 (5.9) and CVE-2024-6763.
-->
<!-- Force the latest patched Jetty 11.x across all Javalin-pulled Jetty modules (no version
skew). Javalin 6.x rides Jetty 11; a move to Jetty 12 needs a Javalin major. -->
<dependencyManagement>
<dependencies>
<dependency>
<groupId>org.eclipse.jetty</groupId>
<artifactId>jetty-bom</artifactId>
<version>${jetty.version}</version>
<type>pom</type>
<scope>import</scope>
</dependency>
<!-- The MCP SDK (Jackson 3) needs jackson-annotations with JsonFormat.Shape.POJO
(the 3.0 line); it shares the com.fasterxml.jackson.annotation package with our
Jackson 2.19 databind, so both must resolve to the same jar. 3.0 is built to work
with Jackson 2.19 databind too — pin it to reconcile the two. -->
<dependency>
<groupId>com.fasterxml.jackson.core</groupId>
<artifactId>jackson-annotations</artifactId>
<version>3.0-rc5</version>
</dependency>
<!-- Testcontainers 1.20.4 pulls commons-compress 1.24.0 (test scope), which carries
CVE-2024-25710 (8.1) + CVE-2024-26308 — both fixed in 1.26.0. Pin the patched line.
Test-scope only (never shipped in the jar), but bumped per the CVE policy. -->
<dependency>
<groupId>org.apache.commons</groupId>
<artifactId>commons-compress</artifactId>
<version>${commons-compress.version}</version>
</dependency>
<!-- Testcontainers 1.20.4 also pulls commons-lang3 3.16.0 (test scope): CVE-2025-48924
(uncontrolled recursion in ClassUtils), fixed in 3.18.0. Pin the patched line.
Test-scope only (never shipped in the jar), bumped per the CVE policy. -->
<dependency>
<groupId>org.apache.commons</groupId>
<artifactId>commons-lang3</artifactId>
<version>${commons-lang3.version}</version>
</dependency>
</dependencies>
</dependencyManagement>
<dependencies>
<!-- JSON + YAML (config, herdr wire format, REST bodies) -->
<dependency>
@@ -44,6 +105,24 @@
<version>${javalin.version}</version>
</dependency>
<!-- MCP server: the SERVER face. Streamable-HTTP servlet mounted on Javalin's Jetty at
/mcp, exposing bridge_send/bridge_reply/bridge_status as thin adapters over REST. -->
<dependency>
<groupId>io.modelcontextprotocol.sdk</groupId>
<artifactId>mcp</artifactId>
<version>${mcp.version}</version>
</dependency>
<!-- Broker client (CB-307 Stage 2): AMQP 0-9-1. Default deploy targets LavinMQ; this same
client speaks to RabbitMQ unchanged (URI-only swap), so integration tests run against a
stock RabbitMQ container. Only wired when a broker: block is present in config; absent →
the in-memory ReplyInbox. -->
<dependency>
<groupId>com.rabbitmq</groupId>
<artifactId>amqp-client</artifactId>
<version>${amqp.version}</version>
</dependency>
<!-- Logging -->
<dependency>
<groupId>org.slf4j</groupId>
@@ -63,6 +142,22 @@
<version>${junit.version}</version>
<scope>test</scope>
</dependency>
<!-- Testcontainers RabbitMQ: spins a real broker for the @Tag("contract") AMQP integration
test only. Excluded from the default build (contract group), so `mvn clean install`
stays hermetic and green without Docker; run under -Pcontract with Docker present. -->
<dependency>
<groupId>org.testcontainers</groupId>
<artifactId>rabbitmq</artifactId>
<version>${testcontainers.version}</version>
<scope>test</scope>
</dependency>
<dependency>
<groupId>org.testcontainers</groupId>
<artifactId>junit-jupiter</artifactId>
<version>${testcontainers.version}</version>
<scope>test</scope>
</dependency>
</dependencies>
<build>
@@ -74,6 +169,30 @@
<version>3.14.0</version>
</plugin>
<!--
Coverage (CB-509). Build-time tooling only — never a compile/runtime dependency, so
it adds nothing to the shipped jar. Report lands at target/site/jacoco/index.html and
target/site/jacoco/jacoco.csv. No `check` rule / threshold is wired: a coverage gate
rewards writing tests that execute lines, which is the failure mode this project is
trying to avoid, not encourage.
-->
<plugin>
<groupId>org.jacoco</groupId>
<artifactId>jacoco-maven-plugin</artifactId>
<version>0.8.13</version>
<executions>
<execution>
<id>prepare-agent</id>
<goals><goal>prepare-agent</goal></goals>
</execution>
<execution>
<id>report</id>
<phase>test</phase>
<goals><goal>report</goal></goals>
</execution>
</executions>
</plugin>
<!-- Unit tests run by default; contract tests (live herdr) are tag-excluded. -->
<plugin>
<groupId>org.apache.maven.plugins</groupId>
@@ -116,7 +235,7 @@
</profile>
<profile>
<id>contract</id>
<properties><excludedGroups></excludedGroups></properties>
<properties><excludedGroups/></properties>
</profile>
</profiles>
</project>
@@ -3,17 +3,49 @@ package dev.ltms.bridged;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
import dev.ltms.bridged.msg.AmqpReplyInbox;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.msg.ReplyPushLoop;
import dev.ltms.bridged.rest.BridgedApp;
import dev.ltms.bridged.worker.WorkerService;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.session.SessionReaper;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.worker.CompositePeerLauncher;
import dev.ltms.bridged.worker.HerdrPeerLauncher;
import dev.ltms.bridged.worker.OpenCodeLauncher;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
/**
* {@code bridged} entry point. Wires the real herdr socket client to the REST app and
@@ -27,6 +59,10 @@ public final class Bridged {
/** How often the injector samples a busy worker's status while it has queued work. */
private static final long INJECT_POLL_MILLIS = 250;
/** CB-504: how long to wait at startup for herdr's socket before serving degraded. */
private static final long HERDR_WAIT_SECONDS = 30;
private static final long HERDR_WAIT_POLL_MILLIS = 500;
static void main(String[] args) {
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
BridgedConfig cfg = BridgedConfig.load(configPath);
@@ -35,30 +71,236 @@ public final class Bridged {
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
guard.assertPrimaryClean(System.getenv());
// CB-501: refuse to start if the bind is wider than the auth mode can defend. Under
// loopback-trust, "not a known worker" means "the primary" — sound only because the OS
// refuses remote connections to a loopback socket. This throws rather than warns so the
// dangerous configuration cannot be reached by ignoring a log line.
cfg.validateAuthExposure();
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
: UnixSocketHerdrClient.defaultSocketPath();
UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect(socket, new com.fasterxml.jackson.databind.ObjectMapper());
Runtime.getRuntime().addShutdownHook(new Thread(herdr::close));
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
WorkerService workers = new WorkerService(agents, spaces, guard, cfg.worker(), System::getenv);
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
Map<String, BridgedConfig.Worker> claudeProfiles = new LinkedHashMap<>();
Map<String, BridgedConfig.Worker> opencodeProfiles = new LinkedHashMap<>();
cfg.workerProfiles().forEach((name, w) -> {
if (w.isOpenCode()) {
opencodeProfiles.put(name, w);
} else {
claudeProfiles.put(name, w);
}
});
List<HerdrPeerLauncher> adapters = new ArrayList<>();
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
claudeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(agents, spaces,
opencodeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
}
PeerLauncher workers = new CompositePeerLauncher(adapters, cfg.defaultProfile());
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
// with /healthz reporting "degraded" is strictly more useful than exiting.
if (awaitHerdr(herdr)) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
// with the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
} else {
log.warn("herdr did not answer within {}s — starting anyway; /healthz will report "
+ "degraded until it comes up. Orphaned worker panes (if any) were NOT reaped.",
HERDR_WAIT_SECONDS);
}
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
int contextCap = 0;
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
&& cfg.lifecycle().contextCap() > 0) {
contextCap = cfg.lifecycle().contextCap();
}
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()), contextCap);
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
final SessionReaper reaper;
if (cfg.lifecycle() != null
&& cfg.lifecycle().idleTtlSeconds() != null
&& cfg.lifecycle().idleTtlSeconds() > 0) {
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
reaper.start();
} else {
reaper = null;
}
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
Injector injector = new Injector(agents);
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
Rendezvous rendezvous = new Rendezvous();
CompletionResolver completion = new CompletionResolver(agents, rendezvous);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
WorkerPresence presence = sessions.asPresence();
TurnListener turnListener = new TurnListener() {
@Override
public void onTurnComplete(String target) {
completion.onTurnComplete(target);
sessions.onTurnComplete(target);
}
@Override
public void onDelivered(String target) {
completion.onDelivered(target);
sessions.onDelivered(target);
}
@Override
public void onTurnFailed(String target) {
completion.onTurnFailed(target);
sessions.onTurnFailed(target);
}
};
Injector injector = new Injector(agents, turnListener, presence::isPresent, presence::forget);
StatusPoller poller = new StatusPoller(agents, injector, INJECT_POLL_MILLIS);
poller.start();
Runtime.getRuntime().addShutdownHook(new Thread(poller::stop));
Javalin app = new BridgedApp(herdr, workers).build();
// CB-307: reply inbox. A broker: block (with a uri) selects the AMQP-backed durable adapter;
// absent, bridged stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
// connection, so keep the reference to close it in the ordered shutdown hook.
final ReplyInbox replyInbox;
if (cfg.broker() != null && cfg.broker().isConfigured()) {
replyInbox = AmqpReplyInbox.open(cfg.broker().uri());
log.info("reply inbox: AMQP broker (durable) at {}", cfg.broker().uri());
} else {
replyInbox = new InMemoryReplyInbox();
log.info("reply inbox: in-memory (soft-state)");
}
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
PrimaryRegistry primaryRegistry = new PrimaryRegistry(
cfg.primary() != null ? cfg.primary().terminal() : null);
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
// open bridge_send. Uses its own lightweight scheduled executor, separate from the injector.
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
var pushScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-push-").unstarted(r));
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
// CB-512: the push loop takes it too, so nudge outcomes (delivered|exhausted) are counted.
Metrics metrics = BridgedMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, agents, replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
sessions.onRelease(terminal ->
messages.abandon(terminal, "the worker session was released before it replied"));
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
final CallerResolver callers;
if (cfg.auth().tokenMode()) {
String token = System.getenv(cfg.auth().tokenEnv());
if (token == null || token.isBlank()) {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting bridged");
}
callers = new CallerResolver(identity, true, token);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = new CallerResolver(identity);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
BridgeMcp mcp = new BridgeMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics);
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
// last. This replaces the earlier independent hooks that could race and close herdr early.
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
poller.stop();
messages.close();
pushLoop.close();
mcp.close();
if (reaper != null) reaper.stop();
// Release the broker connection last among message resources (no-op for the in-memory inbox).
if (replyInbox instanceof AutoCloseable closeable) {
try {
closeable.close();
} catch (Exception e) {
log.debug("reply inbox close: {}", e.toString());
}
}
herdr.close();
}));
Javalin app = new BridgedApp(herdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics).build();
app.start(cfg.bind().host(), cfg.bind().port());
log.info("bridged listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
*
* @return true if herdr answered, false if it never did
*/
private static boolean awaitHerdr(HerdrClient herdr) {
long deadline = System.nanoTime() + HERDR_WAIT_SECONDS * 1_000_000_000L;
boolean waited = false;
while (true) {
try {
herdr.call("ping");
if (waited) {
log.info("herdr is up");
}
return true;
} catch (HerdrException e) {
if (System.nanoTime() >= deadline) {
return false;
}
if (!waited) {
log.info("waiting up to {}s for the herdr socket…", HERDR_WAIT_SECONDS);
waited = true;
}
try {
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
return false;
}
}
}
}
private Bridged() {
}
}
@@ -0,0 +1,72 @@
package dev.ltms.bridged.auth;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.time.Instant;
import java.time.ZoneOffset;
import java.time.format.DateTimeFormatter;
/**
* Append-only record of privileged actions (CB-505).
*
* <p>Writes JSON lines to a dedicated {@code audit} logger — its own appender, separate from the
* chatty app log — so the trail stays greppable and can later be shipped without dragging debug
* noise along.
*
* <p><strong>Message content is never recorded.</strong> This bridge carries the user's source
* code, diffs, and prompts; an audit trail that quietly accumulated them would be a transcript
* archive wearing a security control's clothing. Records carry <em>who / what / against what /
* outcome</em> and correlation ids only.
*/
public final class AuditLog {
private static final Logger AUDIT = LoggerFactory.getLogger("audit");
private static final DateTimeFormatter TS =
DateTimeFormatter.ofPattern("yyyy-MM-dd'T'HH:mm:ss.SSSXXX").withZone(ZoneOffset.UTC);
private AuditLog() {
}
/** Record an allowed action. */
public static void allowed(Principal caller, Authz.Action action, String target) {
write(caller, action, target, "allowed", null);
}
/** Record a refused action and why. */
public static void denied(Principal caller, Authz.Action action, String target, String reason) {
write(caller, action, target, "denied", reason);
}
/** Record an action that was authorized but then failed downstream (guard, timeout, herdr). */
public static void failed(Principal caller, Authz.Action action, String target, String reason) {
write(caller, action, target, "failed", reason);
}
private static void write(Principal caller, Authz.Action action, String target,
String outcome, String reason) {
Principal c = caller != null ? caller : Principal.anonymous();
StringBuilder sb = new StringBuilder(200);
// The timestamp is built here rather than by the appender pattern: a pattern that wrapped
// literal braces around the message collides with logback's own variable substitution.
sb.append("{\"ts\":\"").append(TS.format(Instant.now())).append('"')
.append(",\"role\":\"").append(c.role()).append('"')
.append(",\"actor\":\"").append(esc(c.describe())).append('"')
.append(",\"pid\":").append(c.pid())
.append(",\"action\":\"").append(action).append('"')
.append(",\"target\":").append(target == null ? "null" : '"' + esc(target) + '"')
.append(",\"outcome\":\"").append(outcome).append('"');
if (reason != null) {
sb.append(",\"reason\":\"").append(esc(reason)).append('"');
}
sb.append('}');
// The appender supplies the timestamp, so it cannot disagree with the app log's clock.
AUDIT.info(sb.toString());
}
/** Minimal JSON string escaping — these values are ids and short reasons, never free text. */
private static String esc(String s) {
return s.replace("\\", "\\\\").replace("\"", "\\\"")
.replace("\n", "\\n").replace("\r", "\\r").replace("\t", "\\t");
}
}
@@ -0,0 +1,72 @@
package dev.ltms.bridged.auth;
/**
* The authorization table (CB-505), stated once and enforced on both entry paths.
*
* <p>Most of these rules are already true de facto — {@code BridgeMcp} derives a worker's identity
* from the connection rather than reading it from an argument, so a worker has never been able to
* reply <em>as</em> another worker over MCP. What was missing is that the REST surface trusted the
* session id in the URL path, and neither surface checked role at all. This class makes the
* invariant explicit and testable rather than emergent.
*/
public final class Authz {
private Authz() {
}
/** A privileged operation, named for the audit trail. */
public enum Action {
/** Spawn a worker peer. */
SPAWN,
/** Tear a worker peer down. */
STOP,
/** Deliver a turn to a session (or answer a worker's question). */
SEND,
/** A worker's terminal reply for its own turn. */
REPLY,
/** A worker's mid-turn question to the primary. */
ASK,
/** Collect held replies from a session's inbox. */
DRAIN,
/** Read-only observation: status, roster, profiles, task polling. */
READ,
/** Scrape the metrics endpoint. */
METRICS
}
/**
* Whether {@code caller} may perform {@code action} against {@code targetSession}.
*
* @param targetSession the session id in the request path; only consulted for the worker-scoped
* actions ({@code REPLY}, {@code ASK}), ignored otherwise, may be
* {@code null}
*/
public static boolean permits(Principal caller, Action action, String targetSession) {
if (caller == null || caller.isAnonymous()) {
return false; // authenticated as nothing ⇒ authorized for nothing
}
return switch (action) {
// Orchestration is the primary's alone. A worker driving spawn/stop/send would be a
// worker escalating into the orchestrator role.
case SPAWN, STOP, SEND, DRAIN -> caller.isPrimary();
// The load-bearing rule: a worker acts only as itself. The primary is deliberately
// excluded — a reply/ask is a worker's own turn output, and letting the primary forge
// one would corrupt the rendezvous correlation it is itself waiting on.
case REPLY, ASK -> caller.ownsSession(targetSession);
// Observation is open to both authenticated roles: a worker legitimately polls its own
// status, and the roster carries no secrets.
case READ, METRICS -> caller.isPrimary() || caller.isWorker();
};
}
/**
* Why a request was refused, for the error body. Distinguishes "you are nobody" from "you are
* somebody, but not the right somebody" — the first is a credential problem (401), the second
* an authorization one (403).
*/
public static boolean isUnauthenticated(Principal caller) {
return caller == null || caller.isAnonymous();
}
}
@@ -0,0 +1,119 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
/**
* Resolves every caller to a {@link Principal}, for both entry paths into the core (CB-501).
*
* <p>There are two of them and they are not layered the way the docs suggest: {@code BridgeMcp}
* calls the service layer directly and is mounted as a raw servlet (so it never passes through a
* Javalin filter), while the REST routes historically resolved no identity at all. Both now
* delegate here, so the authorization rules are stated once instead of drifting apart.
*
* <p><strong>Resolution order</strong> — connection identity first, token second, nothing third:
* <ol>
* <li>A loopback peer PID that maps to a herdr worker pane ⇒ {@link Role#WORKER}. This is
* unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured
* regardless of auth mode, so enabling auth never breaks the fleet.</li>
* <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li>
* <li>Otherwise, under {@code loopback-trust}, a loopback caller ⇒ {@link Role#PRIMARY}
* (the historical behaviour, now an explicit configured choice).</li>
* <li>Otherwise {@link Role#ANONYMOUS}.</li>
* </ol>
*/
public final class CallerResolver {
private final ConnectionIdentity identity;
private final boolean tokenMode;
private final byte[] expectedToken; // null unless tokenMode
/** Loopback-trust resolver: no token required, historical behaviour. */
public CallerResolver(ConnectionIdentity identity) {
this(identity, false, null);
}
/**
* @param identity connection-based worker identification
* @param tokenMode when true, a non-worker caller must present a valid bearer token
* @param token the expected bearer token; required (non-blank) when {@code tokenMode}
*/
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
if (tokenMode && (token == null || token.isBlank())) {
throw new IllegalArgumentException(
"auth.mode=token requires a non-empty token; check that the env var named by "
+ "auth.tokenEnv is exported to the daemon's environment");
}
this.identity = identity;
this.tokenMode = tokenMode;
this.expectedToken = tokenMode ? token.getBytes(StandardCharsets.UTF_8) : null;
}
/**
* Resolve the caller of a request.
*
* @param remoteAddr the connection's remote address
* @param remotePort the connection's remote port (used for the peer-PID lookup)
* @param authorizationHeader the raw {@code Authorization} header, or {@code null}
*/
public Principal resolve(String remoteAddr, int remotePort, String authorizationHeader) {
ConnectionIdentity.Caller c = identity.resolve(remoteAddr, remotePort);
if (c.terminal() != null) {
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
}
if (tokenMode) {
return presentedTokenMatches(authorizationHeader)
? Principal.primary(c.pid())
: Principal.anonymous();
}
// loopback-trust: same-host callers that are not workers are the primary. A non-loopback
// caller is anonymous even here — and startup refuses that combination anyway
// (BridgedConfig.validateAuthExposure), so this is defence in depth, not the control.
return isLoopback(remoteAddr) ? Principal.primary(c.pid()) : Principal.anonymous();
}
/** The working directory of the calling process (CB-112 spawn cwd inheritance), or {@code null}. */
public String cwdForPid(long pid) {
return identity.cwdForPid(pid);
}
/** True when auth requires a bearer token of non-worker callers. */
public boolean tokenMode() {
return tokenMode;
}
private boolean presentedTokenMatches(String authorizationHeader) {
String presented = bearerValue(authorizationHeader);
if (presented == null) {
return false;
}
// Constant-time: MessageDigest.isEqual does not short-circuit on the first differing byte,
// so a token cannot be recovered a byte at a time by timing the response.
return MessageDigest.isEqual(presented.getBytes(StandardCharsets.UTF_8), expectedToken);
}
/** Extract the credential from {@code Authorization: Bearer <token>}, or {@code null}. */
private static String bearerValue(String header) {
if (header == null) {
return null;
}
String h = header.trim();
if (h.length() < 7 || !h.regionMatches(true, 0, "Bearer ", 0, 7)) {
return null;
}
String token = h.substring(7).trim();
return token.isEmpty() ? null : token;
}
private static boolean isLoopback(String remoteAddr) {
if (remoteAddr == null) {
return false;
}
return remoteAddr.equals("127.0.0.1") || remoteAddr.equals("::1")
|| remoteAddr.equals("0:0:0:0:0:0:0:1") || remoteAddr.startsWith("127.");
}
}
@@ -0,0 +1,57 @@
package dev.ltms.bridged.auth;
/**
* A resolved caller: its {@link Role}, and — for a worker — the herdr {@code terminal_id} that
* identifies which worker it is (CB-501).
*
* @param role what this caller is authorized to act as
* @param terminal the worker's herdr terminal id; {@code null} for {@code PRIMARY}/{@code ANONYMOUS}
* @param pid the connecting process id, or {@code -1} when not resolvable (audit context)
*/
public record Principal(Role role, String terminal, long pid) {
/** A caller authenticated as nothing — the default when no check establishes anything else. */
public static Principal anonymous() {
return new Principal(Role.ANONYMOUS, null, -1);
}
/** The orchestrating session. */
public static Principal primary(long pid) {
return new Principal(Role.PRIMARY, null, pid);
}
/** A worker peer, identified by its herdr pane. */
public static Principal worker(String terminal, long pid) {
return new Principal(Role.WORKER, terminal, pid);
}
public boolean isPrimary() {
return role == Role.PRIMARY;
}
public boolean isWorker() {
return role == Role.WORKER;
}
public boolean isAnonymous() {
return role == Role.ANONYMOUS;
}
/**
* Whether this caller may act <em>as</em> {@code sessionId} — the "own session only" rule that
* keeps one worker from replying or asking on another's behalf. Only a worker can own a
* session, and only its own.
*/
public boolean ownsSession(String sessionId) {
return isWorker() && terminal != null && terminal.equals(sessionId);
}
/** Short, non-sensitive description for audit lines and error details. */
public String describe() {
return switch (role) {
case WORKER -> "worker:" + terminal;
case PRIMARY -> "primary";
case ANONYMOUS -> "anonymous";
};
}
}
@@ -0,0 +1,29 @@
package dev.ltms.bridged.auth;
/**
* What a caller is allowed to be on the bus (CB-501).
*
* <p>The ordering matters conceptually: {@link #PRIMARY} is the <em>most</em> privileged role
* (it spawns, stops, sends to any session, and drains any inbox), not the least. Before CB-501
* the daemon reached {@code PRIMARY} by <em>failing</em> every other check — any caller that did
* not resolve to a known worker pane was treated as the primary. That is inverted here:
* {@link #ANONYMOUS} is the fallback, and {@code PRIMARY} must be established.
*/
public enum Role {
/**
* The orchestrating session. Established either by being a loopback caller that is not a
* worker pane (under {@code loopback-trust}) or by presenting a valid bearer token (under
* {@code token} mode).
*/
PRIMARY,
/**
* A worker peer, identified by its herdr pane. Unforgeable: derived from the connection's
* loopback peer PID via herdr's PID→pane map, never from a request argument.
*/
WORKER,
/** Authenticated as nothing. Authorized for nothing but {@code /healthz}. */
ANONYMOUS
}
@@ -8,7 +8,9 @@ import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
/**
@@ -16,23 +18,48 @@ import java.util.Set;
* {@code bridged.example.yaml}). Unknown keys are ignored so config can grow ahead
* of the code.
*
* @param bind REST/MCP listen host:port
* @param herdrSocket path to herdr's Unix socket ({@code null} → client default)
* @param worker worker-spawn settings
* @param guard subscription-boundary allowlist
* @param bind REST/MCP listen host:port
* @param herdrSocket path to herdr's Unix socket ({@code null} → client default)
* @param worker single worker profile (legacy; superseded by {@code workers})
* @param workers named worker profiles, keyed by profile name (multi-backend fleet)
* @param defaultWorker which {@code workers} key a no-argument spawn uses ({@code null} → the
* single {@code worker}, or the sole/first profile)
* @param guard subscription-boundary allowlist
* @param worktreeRoot nullable root directory for provisioned worktrees; defaults to a sibling
* of the repo root
* @param lifecycle session lifecycle limits ({@code null} = all disabled)
* @param spawnReadyTimeoutMs max ms to wait for a spawned worker to reach an injectable state
* ({@code null} / 0 disables the poll gate — legacy non-blocking behaviour)
* @param spawnReadyPollMs poll interval while waiting for the worker to become injectable
* @param broker external AMQP broker for durable reply delivery ({@code null} → in-memory,
* soft-state {@code ReplyInbox}; present → the AMQP-backed adapter, CB-307 Stage 2)
* @param primary optional pinned primary terminal config ({@code null} → derived from connection);
* a non-blank {@code terminal} seeds {@code PrimaryRegistry} and prevents
* connection-derived overrides, CB-307
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
* historical behaviour), CB-501
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record BridgedConfig(
Bind bind,
String herdrSocket,
Worker worker,
Guard guard) {
Map<String, Worker> workers,
String defaultWorker,
Guard guard,
String worktreeRoot,
Lifecycle lifecycle,
Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs,
Broker broker,
Primary primary,
Auth auth) {
@JsonIgnoreProperties(ignoreUnknown = true)
public record Bind(String host, int port) {
public Bind {
if (host == null || host.isBlank()) host = "127.0.0.1";
if (port <= 0) port = 8080;
if (port <= 0) port = 8765;
}
}
@@ -53,17 +80,120 @@ public record BridgedConfig(
* @param tabLabel template for a worker tab's label; {@code {profile}}/{@code {model}}
* and {@code {n}} (per-worker number, to keep sibling tabs distinct)
* are substituted (default {@code "worker: {profile} #{n}"})
* @param mcpUrl bridge MCP URL to provision into the worker's {@code configDir} so it
* can call {@code bridge_reply} ({@code null}/blank → no provisioning; the
* worker won't reply, only the fallback/timeout resolves the send)
* @param cwd fixed working directory for this profile's workers (CB-112 "told otherwise");
* {@code null}/blank → inherit the primary's cwd, else the daemon's
* @param parityOverlay repo-relative paths copied primary→worktree for config parity; null/empty
* defaults to a sensible set of local config files
* @param gitTokenEnv name of the host env var holding the git-forge API token; when set, its
* value is injected as {@code GITEA_TOKEN} so the worker can open its own PR
* at checkpoint (CB-302). {@code null}/blank ⇒ no token is injected
* (minimal-grant default — push over SSH stays free, PR-create is opt-in)
* @param gitHostEnv name of the host env var holding the forge host (default {@code GITEA_HOST});
* injected as {@code GITEA_HOST} <em>only</em> when {@code gitTokenEnv} is set
* @param kind which peer launcher spawns this profile: {@code "claude-code"} (default —
* the {@link dev.ltms.bridged.worker.ClaudeCodeLauncher}) or {@code "opencode"}.
* The {@code CompositePeerLauncher} routes {@code spawn}/reap by this value, so
* each adapter drives only its own kind. Normalised to lower-case; blank ⇒ the
* default. It selects the adapter, not the transport — placement, tabs, cwd, and
* the readiness gate are kind-independent and stay in the shared base.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel) {
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd,
List<String> parityOverlay,
String gitTokenEnv, String gitHostEnv,
String kind,
Map<String, String> env) {
/** Peer kind spawned by {@link dev.ltms.bridged.worker.ClaudeCodeLauncher} (the default). */
public static final String KIND_CLAUDE_CODE = "claude-code";
/** Peer kind spawned by the opencode adapter (CB-402). */
public static final String KIND_OPENCODE = "opencode";
public Worker {
argv = (argv == null || argv.isEmpty()) ? List.of("claude") : List.copyOf(argv);
// A claude-code worker defaults its launch command to `claude`; other kinds carry their own
// argv (e.g. `opencode`) and must not inherit the Claude binary — so only default when unset
// AND this is the claude-code kind.
String k = (kind == null || kind.isBlank()) ? KIND_CLAUDE_CODE : kind.toLowerCase();
argv = (argv == null || argv.isEmpty())
? (KIND_CLAUDE_CODE.equals(k) ? List.of("claude") : List.of(k))
: List.copyOf(argv);
kind = k;
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_WORKER_TOKEN" : tokenEnv;
placement = (placement == null || placement.isBlank()) ? "tab" : placement.toLowerCase();
workspace = (workspace == null || workspace.isBlank()) ? "bridged-workers" : workspace;
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? "worker: {profile} #{n}" : tabLabel;
parityOverlay = (parityOverlay == null || parityOverlay.isEmpty())
? List.of(".mcp.json", ".claude/settings.local.json", ".env", ".envrc")
: List.copyOf(parityOverlay);
// gitTokenEnv stays null when unset (opt-in). gitHostEnv defaults so operators enabling
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
gitHostEnv = (gitHostEnv == null || gitHostEnv.isBlank()) ? "GITEA_HOST" : gitHostEnv;
env = (env == null) ? Map.of() : Map.copyOf(env);
}
/**
* Backward-compatible constructor without the CB-302 git-forge fields — the worker is
* granted no PR-create token (push over SSH is unaffected). Keeps pre-CB-302 call sites
* (and any {@code workers:} YAML that omits the git keys) working unchanged.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, null, null, null);
}
/**
* Backward-compatible constructor with the CB-302 git-forge fields but no explicit peer
* {@code kind} — defaults to {@link #KIND_CLAUDE_CODE}. Keeps pre-CB-402 call sites working.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, null, null);
}
/**
* Backward-compatible constructor without the CB-511 {@code env:} passthrough — the worker
* gets the daemon's PATH and nothing else. Keeps pre-CB-511 call sites working.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, null);
}
/** A copy with {@code profile} set — used to default a profile to its {@code workers} key. */
public Worker withProfile(String p) {
return new Worker(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env);
}
/** True when this profile is served by the Claude Code adapter (the default kind). */
public boolean isClaudeCode() {
return KIND_CLAUDE_CODE.equals(kind);
}
/** True when this profile is served by the opencode adapter (CB-402). */
public boolean isOpenCode() {
return KIND_OPENCODE.equals(kind);
}
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
public boolean hasGitToken() {
return gitTokenEnv != null && !gitTokenEnv.isBlank();
}
/** True when workers should land in their own tab in the worker space. */
@@ -71,6 +201,11 @@ public record BridgedConfig(
return "tab".equals(placement);
}
/** True when the bridge MCP should be mounted into a spawned worker (via launch flags). */
public boolean hasMcp() {
return mcpUrl != null && !mcpUrl.isBlank();
}
/**
* Render {@link #tabLabel} for the {@code n}-th worker (substitutes
* {@code {profile}}/{@code {model}}/{@code {n}}), so sibling worker tabs are distinct.
@@ -83,6 +218,99 @@ public record BridgedConfig(
}
}
/**
* Session lifecycle limits. All knobs are opt-in: {@code null} or {@code 0} disables the
* feature so existing configs keep the previous behaviour.
*
* @param idleTtlSeconds max seconds a {@code READY}/{@code DONE} session may sit idle
* before it is reaped ({@code null} → disabled)
* @param contextCap max delegated turns a session serves before force-release
* ({@code null} → disabled)
* @param drainTimeoutSeconds seconds to wait for {@code BUSY} sessions to finish before
* forced teardown on shutdown (default 5 when unset)
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Lifecycle(Integer idleTtlSeconds, Integer contextCap, Integer drainTimeoutSeconds) {
}
/**
* External AMQP broker for durable, cross-restart reply delivery (CB-307 Stage 2). Its mere
* presence swaps the in-memory {@code ReplyInbox} for the AMQP-backed adapter; absent, bridged
* stays soft-state. Production default is LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1
* and is a URI-only swap.
*
* @param uri AMQP connection URI, e.g. {@code amqp://guest:guest@127.0.0.1:5672/}. Blank/{@code null}
* ⇒ the broker block is treated as absent (in-memory adapter).
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Broker(String uri) {
/** True when a usable broker URI is configured (an empty block does not enable AMQP). */
public boolean isConfigured() {
return uri != null && !uri.isBlank();
}
}
/**
* Optional pinned primary terminal config (CB-307). When present with a non-blank
* {@code terminal}, the bridge uses this as the primary's herdr identity instead of
* deriving it from the MCP connection. Useful when the primary runs off-host or in a
* non-herdr terminal where connection-derived identity is unavailable.
*
* @param terminal the primary's herdr {@code terminal_id} ({@code null}/blank → derive)
* @param pushReminders max reminder nudges before giving up (default 5)
* @param pushBackoffMs delay between reminders in ms (default 15000)
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Primary(String terminal, Integer pushReminders, Integer pushBackoffMs) {
/** @return configured reminder cap, or 5 */
public int remindersOrDefault() {
return pushReminders != null ? pushReminders : 5;
}
/** @return configured backoff in ms, or 15000 */
public long backoffMsOrDefault() {
return pushBackoffMs != null ? pushBackoffMs.longValue() : 15_000L;
}
}
/**
* API authentication (CB-501). Governs how a caller that is <em>not</em> an on-host worker
* pane proves it is the primary.
*
* <p>Worker identity never depends on this block: a loopback peer PID that maps to a herdr
* pane is unforgeable and is always honoured (see
* {@link dev.ltms.bridged.mcp.ConnectionIdentity}). This only decides what happens for
* <em>everyone else</em>.
*
* @param mode {@code "loopback-trust"} (default) — any loopback caller that is not a known
* worker is the primary, no credential needed; this is the historical
* behaviour, now chosen explicitly rather than implied. {@code "token"} — such
* a caller must present {@code Authorization: Bearer <token>} or it is
* {@code ANONYMOUS} and authorized for nothing.
* @param tokenEnv name of the host env var holding the bearer token; the literal value is
* never stored in config. Defaults to {@code BRIDGED_API_TOKEN}. Only read
* when {@code mode} is {@code token}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Auth(String mode, String tokenEnv) {
/** Historical behaviour: loopback non-worker ⇒ primary, no credential. */
public static final String MODE_LOOPBACK_TRUST = "loopback-trust";
/** A non-worker caller must present a valid bearer token to be the primary. */
public static final String MODE_TOKEN = "token";
public Auth {
mode = (mode == null || mode.isBlank()) ? MODE_LOOPBACK_TRUST : mode.toLowerCase();
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_API_TOKEN" : tokenEnv;
}
/** True when a bearer token is required of every non-worker caller. */
public boolean tokenMode() {
return MODE_TOKEN.equals(mode);
}
}
/**
* Subscription boundary. Only these hosts may back a worker's
* {@code ANTHROPIC_BASE_URL}; the primary must carry none.
@@ -100,6 +328,40 @@ public record BridgedConfig(
}
}
/**
* The effective worker profiles, keyed by profile name. Prefers the {@code workers} map (each
* value's {@code profile} defaulted to its key); falls back to the legacy singular {@code worker}
* (keyed by its own profile). Empty if neither is configured.
*/
public Map<String, Worker> workerProfiles() {
if (workers != null && !workers.isEmpty()) {
Map<String, Worker> out = new LinkedHashMap<>();
workers.forEach((name, w) -> out.put(name,
(w.profile() == null || w.profile().isBlank()) ? w.withProfile(name) : w));
return Map.copyOf(out);
}
if (worker != null) {
String name = (worker.profile() == null || worker.profile().isBlank()) ? "default" : worker.profile();
return Map.of(name, worker);
}
return Map.of();
}
/**
* The profile a no-argument spawn uses: {@code defaultWorker} if set, else the legacy single
* {@code worker}'s profile, else the sole/first configured profile, else {@code null}.
*/
public String defaultProfile() {
if (defaultWorker != null && !defaultWorker.isBlank()) {
return defaultWorker;
}
if (worker != null && worker.profile() != null && !worker.profile().isBlank()) {
return worker.profile();
}
Map<String, Worker> p = workerProfiles();
return p.isEmpty() ? null : p.keySet().iterator().next();
}
private static final ObjectMapper YAML = new ObjectMapper(new YAMLFactory());
/** Load and validate config from {@code path}. */
@@ -116,6 +378,46 @@ public record BridgedConfig(
public BridgedConfig withDefaults() {
Bind b = bind != null ? bind : new Bind(null, 0);
Guard g = guard != null ? guard : new Guard(List.of());
return new BridgedConfig(b, herdrSocket, worker, g);
Lifecycle l = lifecycle != null ? lifecycle : new Lifecycle(null, null, null);
Integer timeout = (spawnReadyTimeoutMs != null) ? spawnReadyTimeoutMs : 20000;
Integer pollMs = (spawnReadyPollMs != null) ? spawnReadyPollMs : 300;
Auth a = auth != null ? auth : new Auth(null, null);
// broker is left as-is: null (or an empty/blank uri) keeps the in-memory soft-state inbox.
// primary is left as-is: null defaults to connection-derived identity.
return new BridgedConfig(b, herdrSocket, worker, workers, defaultWorker, g, worktreeRoot, l, timeout, pollMs, broker, primary, a);
}
/**
* Reject a configuration whose network exposure outruns its authentication (CB-501).
*
* <p>{@code loopback-trust} means "any caller that is not a known worker pane is the primary" —
* safe only because the OS refuses non-local connections to a loopback bind. Widen
* {@code bind.host} without switching to {@code token} mode and that sentence becomes "any
* client that can reach this port is the primary", which is the most privileged role on the
* bus. Rather than document the hazard, make it unrepresentable: fail fast at startup.
*
* @throws IllegalStateException when a non-loopback bind is paired with {@code loopback-trust}
*/
public void validateAuthExposure() {
String host = bind().host();
if (isLoopbackBind(host) || auth().tokenMode()) {
return;
}
throw new IllegalStateException(
"refusing to start: bind.host=" + host + " is not loopback, but auth.mode="
+ auth().mode() + ". A non-loopback bind treats every unauthenticated "
+ "caller as the primary (spawn/stop/send/drain on any session). Set "
+ "auth.mode: token (with auth.tokenEnv) before exposing this port, or "
+ "bind to 127.0.0.1 and put a reverse proxy in front.");
}
/** True for the loopback addresses and the unspecified-but-local forms we treat as same-host. */
private static boolean isLoopbackBind(String host) {
if (host == null || host.isBlank()) {
return true; // Bind's own default is 127.0.0.1
}
String h = host.trim().toLowerCase();
return h.equals("127.0.0.1") || h.equals("::1") || h.equals("localhost")
|| h.startsWith("127.");
}
}
@@ -15,6 +15,10 @@ import com.fasterxml.jackson.databind.JsonNode;
* @param agentType detected agent kind, e.g. {@code "claude"}, or {@code null} before herdr
* has detected it (the start-time shape)
* @param status current lifecycle state
* @param name the unique label the agent was started with — for a bridge worker this is
* {@code claude-<profile>-<nonce>-<seq>} (CB-117 keys orphan reaping on the
* nonce); {@code null} for agents the bridge did not start, e.g. a user's own
* Claude session
*/
public record Agent(
String terminalId,
@@ -23,7 +27,8 @@ public record Agent(
String tabId,
String sessionId,
String agentType,
AgentStatus status) {
AgentStatus status,
String name) {
/** Project a herdr {@code agent} node. Tolerates the start-time shape (no session yet). */
public static Agent from(JsonNode a) {
@@ -42,6 +47,7 @@ public record Agent(
a.path("tab_id").asText(null),
sessionId,
type,
AgentStatus.fromWire(a.path("agent_status").asText(null)));
AgentStatus.fromWire(a.path("agent_status").asText(null)),
a.path("name").asText(null));
}
}
@@ -18,6 +18,17 @@ import java.util.Map;
*/
public final class AgentControl {
/**
* The keystroke that submits a prompt in the Claude Code TUI: a carriage return (Enter).
* It must be delivered as its <em>own</em> {@code agent.send} call — herdr delivers a message's
* text as a bracketed paste, and a {@code "\r"} appended to that same text is swallowed as
* literal newline content, not a submit. Sent as a separate keystroke event it lands outside
* the paste and submits. (A bare {@code "\n"} inserts a newline either way.) Verified live
* against Claude Code v2.1.210: an injected task stayed unsubmitted with {@code "text\r"} in
* one call, and submitted the instant a standalone {@code "\r"} was sent.
*/
static final String SUBMIT_KEY = "\r";
private final HerdrClient herdr;
public AgentControl(HerdrClient herdr) {
@@ -36,12 +47,20 @@ public final class AgentControl {
return start(name, argv, env, null);
}
/**
* Spawn an agent into a specific tab. With a non-null {@code tabId} the worker lands
* in that tab (the placement policy's dedicated worker tab); with {@code null} herdr
* splits the currently-focused tab (legacy pane placement).
*/
/** Spawn an agent into {@code tabId} at herdr's default cwd. */
public Agent start(String name, List<String> argv, Map<String, String> env, String tabId) {
return start(name, argv, env, tabId, null);
}
/**
* Spawn an agent. With a non-null {@code tabId} the worker lands in that tab (the placement
* policy's dedicated worker tab); with {@code null} herdr splits the currently-focused tab
* (legacy pane placement). A non-blank {@code cwd} sets the worker process's working directory —
* {@code agent.start} honours {@code cwd} directly (an agent pane does <em>not</em> inherit the
* tab's or workspace's cwd, so this is the only way to root a worker in the primary's directory;
* CB-112).
*/
public Agent start(String name, List<String> argv, Map<String, String> env, String tabId, String cwd) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("name", name);
params.put("argv", argv);
@@ -49,13 +68,31 @@ public final class AgentControl {
if (tabId != null) {
params.put("tab_id", tabId);
}
if (cwd != null && !cwd.isBlank()) {
params.put("cwd", cwd);
}
JsonNode result = herdr.call("agent.start", params);
return Agent.from(result.get("agent"));
}
/** Deliver {@code text} to an agent (its next prompt input). */
/**
* Deliver {@code text} to an agent as its next prompt <em>and submit it</em> — two keystroke
* events: the message (a bracketed paste, so any embedded newlines are preserved verbatim),
* then a standalone {@link #SUBMIT_KEY} (Enter) that actually submits it. Without the second
* event the text just sits in the worker's input box, never processed (see {@link #SUBMIT_KEY}).
*/
public void send(String target, String text) {
herdr.call("agent.send", Map.of("target", target, "text", text));
herdr.call("agent.send", Map.of("target", target, "text", SUBMIT_KEY));
}
/**
* Re-send the submit keystroke (Enter) to {@code target}. The Enter that accompanies a delivery
* can race the paste — especially right as the worker's TUI becomes interactive — leaving the
* text unsubmitted; the injector nudges it with this until the worker actually picks up (CB-113).
*/
public void submit(String target) {
herdr.call("agent.send", Map.of("target", target, "text", SUBMIT_KEY));
}
/**
@@ -2,28 +2,38 @@ package dev.ltms.bridged.herdr;
/**
* A herdr agent's lifecycle state, as reported by {@code agent_status}. Drives the
* status-gated injector: a worker is safe to inject into only when {@link #IDLE} or
* {@link #BLOCKED}, never mid-turn ({@link #WORKING}).
* status-gated injector: a worker is safe to inject into only when {@link #IDLE},
* {@link #BLOCKED}, or {@link #DONE}, never mid-turn ({@link #WORKING}).
*/
public enum AgentStatus {
IDLE,
WORKING,
BLOCKED,
/**
* The worker has finished its turn and is settled at an idle prompt. herdr emits this
* (observed live alongside {@code idle}) as a turn-complete marker; earlier code mapped the
* unrecognized string to {@link #UNKNOWN}, which both wedged delivery (not {@link #injectable})
* and mis-fired the CB-109 stall-failure on a worker that had actually answered. It is a
* turn-boundary equivalent to {@link #IDLE}: injectable, and a {@code working → done} edge is a
* real completion.
*/
DONE,
UNKNOWN;
/** Map herdr's wire string ({@code idle|working|blocked|unknown}) to the enum. */
/** Map herdr's wire string ({@code idle|working|blocked|done|unknown}) to the enum. */
public static AgentStatus fromWire(String s) {
if (s == null) return UNKNOWN;
return switch (s.toLowerCase()) {
case "idle" -> IDLE;
case "working" -> WORKING;
case "blocked" -> BLOCKED;
case "done" -> DONE;
default -> UNKNOWN;
};
}
/** Whether {@code bridged} may inject a message now without stepping on a live turn. */
public boolean injectable() {
return this == IDLE || this == BLOCKED;
return this == IDLE || this == BLOCKED || this == DONE;
}
}
@@ -0,0 +1,59 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import java.util.Map;
/**
* Resolves which herdr pane a process belongs to — the herdr half of connection-based MCP
* identity (CB-105). Given the PID that opened an MCP connection, {@link #terminalForPid} finds
* the agent pane whose process tree contains it, so {@code bridged} can tell <em>which worker</em>
* is calling without the worker sending anything spoofable.
*
* <p>herdr owns the PID→pane truth: {@code pane.process_info} reports each pane's {@code shell_pid}
* and foreground process PIDs. This scans agent panes; a spawn-time {@code pid→terminal} cache is
* the obvious optimization once wired into {@code ClaudeCodeLauncher}.
*/
public final class PaneLocator {
private final HerdrClient herdr;
public PaneLocator(HerdrClient herdr) {
this.herdr = herdr;
}
/**
* The {@code terminal_id} of the agent pane whose process tree contains {@code pid}, or
* {@code null} if no agent pane owns it (e.g. the caller is the primary, or off-host).
*/
public String terminalForPid(long pid) {
if (pid <= 0) {
return null;
}
for (JsonNode pane : herdr.call("pane.list", Map.of()).path("panes")) {
String paneId = pane.path("pane_id").asText(null);
if (paneId != null && paneOwnsPid(paneId, pid)) {
return pane.path("terminal_id").asText(null);
}
}
return null;
}
private boolean paneOwnsPid(String paneId, long pid) {
JsonNode info;
try {
info = herdr.call("pane.process_info", Map.of("pane_id", paneId)).path("process_info");
} catch (HerdrException e) {
return false; // pane vanished mid-scan — just skip it
}
if (info.path("shell_pid").asLong(-1) == pid) {
return true;
}
for (JsonNode p : info.path("foreground_processes")) {
if (p.path("pid").asLong(-1) == pid) {
return true;
}
}
return false;
}
}
@@ -65,13 +65,13 @@ public final class WorkspaceControl {
}
/**
* A brand-new tab in {@code workspaceId} plus the placeholder shell pane herdr seeds
* it with. Start the worker into the tab, then {@code pane.close} the root pane so the
* tab holds only the worker.
* A brand-new tab in {@code workspaceId} plus the placeholder shell pane herdr seeds it with.
* Start the worker into the tab, then {@code pane.close} the root pane so the tab holds only the
* worker. (The worker's own cwd is set on {@code agent.start}, not here — an {@code agent.start}
* pane does not inherit the tab's cwd; see {@code AgentControl.start}.)
*/
public Tab.Created createTab(String workspaceId) {
JsonNode result = herdr.call("tab.create", Map.of("workspace_id", workspaceId));
return Tab.Created.from(result);
return Tab.Created.from(herdr.call("tab.create", Map.of("workspace_id", workspaceId)));
}
/** Give a worker's tab a human label in the tab bar. */
@@ -0,0 +1,256 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.msg.Rendezvous;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
/**
* The CB-106 completion fallback: bridges the {@link Injector}'s turn-completion signal to the
* {@link Rendezvous} so a blocking {@code bridge_send} resolves even when the worker finishes its
* task without ever calling {@code bridge_reply} — the common case for a real delegated coding task.
*
* <p>On a confirmed {@code working → idle} boundary it scrapes the worker's recent transcript and
* resolves the awaiting send with that tail (a {@link Rendezvous.Kind#COMPLETION} resolution, so the
* caller can tell a scrape from a structured reply). It scrapes only when a send is actually waiting
* — a fleet worker's own turns, or a send that already timed out, cost no herdr traffic. An explicit
* {@code bridge_reply} that raced in first wins; {@link Rendezvous#resolveCompletion} is then a no-op.
*
* <p>It also handles the CB-109 stall signal ({@link #onTurnFailed}): a worker that ran a turn then
* wedged in an {@code unknown} state resolves the send as a failure (with the error screen as
* context) rather than leaving it to time out.
*
* <p>The scrape is cleaned to the last {@code ⏺} assistant block (stripping TUI chrome) and guarded
* against misattribution (CB-115): the pane content is baselined on delivery ({@link #onDelivered}),
* and a completion whose scrape is unchanged from that baseline — the previous turn's wind-down
* sampled as this turn's boundary on a rapid back-to-back send — is suppressed rather than resolving
* the send with a stale answer.
*
* <p><strong>Waiter-specific resolution (CB-116).</strong> On delivery we also capture the exact
* {@link Rendezvous} waiter this turn belongs to, and the completion/failure fallbacks resolve
* <em>that</em> waiter — never "whatever send is waiting now". A completion fallback runs on a virtual
* thread and can land after the worker's {@code bridge_reply} already resolved the turn and the
* <em>next</em> send opened its own waiter on the same session; resolving the current waiter would
* then deliver turn N's stale scrape as turn N+1's answer. Targeting the captured waiter makes a late
* completion a harmless no-op (its waiter is already done) instead of a cross-turn stale reply.
*
* <p>Wired as the {@link Injector}'s {@link TurnListener}; the handlers hand off to a virtual thread
* so the scrape's herdr round-trip never stalls the status poller. The captured waiter is read on the
* poller thread (before any next-turn delivery can overwrite it) and passed into the virtual thread.
*/
public final class CompletionResolver implements TurnListener {
private static final Logger log = LoggerFactory.getLogger(CompletionResolver.class);
/**
* herdr {@code agent.read} source for the completion scrape. {@code recent} returns the tail of
* the transcript (the worker's last output), which is what a delegator wants when the worker
* didn't structure a reply.
*/
static final String SCRAPE_SOURCE = "recent";
/** Cap the scraped tail so a long transcript can't return an unbounded blob. */
static final int MAX_SCRAPE_CHARS = 4000;
private final AgentControl agents;
private final Rendezvous rendezvous;
/**
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
* delivering send opened, plus the assistant block present when it was delivered.
*
* <p>The {@code waiter} is what makes a late fallback safe (CB-116): we resolve it, not "whoever
* is waiting now", so a completion that fires after the next send has opened its own waiter is a
* no-op rather than a cross-turn stale reply. The {@code baseline} is the CB-115 staleness
* reference: a completion scrape equal to it means the worker produced no new output (the previous
* turn's wind-down sampled as this boundary), so it is suppressed. Overwritten on each delivery;
* cleared when the turn resolves. Package-private so tests can capture and replay a specific turn.
*/
record InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline) {
}
private final ConcurrentHashMap<String, InFlight> inFlight = new ConcurrentHashMap<>();
public CompletionResolver(AgentControl agents, Rendezvous rendezvous) {
this.agents = agents;
this.rendezvous = rendezvous;
}
@Override
public void onDelivered(String target) {
// Capture the exact waiter this turn belongs to (CB-116) and snapshot the pane's pre-turn
// content — what it shows *before* the just-delivered turn produces output — as the staleness
// reference (CB-115). Done synchronously (like the delivering send itself) so both are in
// place before this turn's completion can fire.
captureBaseline(target);
}
/** Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link #onDelivered}). */
void captureBaseline(String target) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null) {
inFlight.remove(target); // no send is waiting on this delivery — nothing to resolve later
return;
}
String baseline;
try {
// Clip to the same cap resolve() applies to the tail (line ~134): the CB-115 misattribution
// guard compares baseline.equals(tail), so both sides must be the same capped representation.
// An unclipped baseline vs a clipped tail would never match for a >MAX_SCRAPE_CHARS block,
// defeating the guard and letting a stale completion resolve the send.
baseline = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
} catch (RuntimeException e) {
baseline = null; // fail open: no baseline ⇒ no suppression
log.debug("delivery baseline for {} failed: {}", target, e.getMessage());
}
inFlight.put(target, new InFlight(waiter, baseline));
}
/** The turn currently baselined for {@code target}, or {@code null} — a test hook for the captureBaseline path. */
InFlight inFlight(String target) {
return inFlight.get(target);
}
@Override
public void onTurnComplete(String target) {
// Read the in-flight turn on the poller thread — before any next-turn delivery can overwrite
// it — then off-load the scrape (a herdr round-trip we must not block polling on) to a vthread.
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("completion-" + target).start(() -> resolve(target, turn));
}
@Override
public void onTurnFailed(String target) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn));
}
/** Synchronous resolve (the unit-testable core of {@link #onTurnComplete}). */
void resolve(String target, InFlight turn) {
CompletableFuture<Rendezvous.Resolution> waiter = turn == null ? null : turn.waiter();
if (waiter == null || waiter.isDone()) {
// Nobody is blocked on THIS turn (it had no send, or its bridge_reply already won). Skip
// the scrape; resolving the current waiter here would be the CB-116 cross-turn stale reply.
inFlight.remove(target, turn);
return;
}
String tail;
boolean scrapeFailed = false;
try {
tail = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
} catch (RuntimeException e) {
// The worker finished but we couldn't read its screen — still resolve the send so the
// caller unblocks; an empty tail beats hanging until the caller's timeout.
log.warn("completion scrape for {} failed; resolving with an empty tail: {}",
target, e.getMessage());
tail = "";
scrapeFailed = true;
}
// Misattribution guard (CB-115): if the scrape is byte-identical to the pane content at
// delivery, this turn produced no new output — the boundary belongs to the previous turn's
// wind-down (common on rapid back-to-back sends). Suppress rather than resolve the send with
// a stale answer; the real bridge_reply (or a later genuine completion) resolves it instead.
// A scrape that failed to read is exempt — an empty tail there is "couldn't see", not "no change".
String baseline = turn.baseline();
if (!scrapeFailed && baseline != null && baseline.equals(tail)) {
log.debug("suppressing misattributed completion for {} (no output change since delivery)",
target);
return; // keep the in-flight record: a later genuine completion still needs it
}
if (rendezvous.resolveCompletion(waiter, tail)) {
inFlight.remove(target, turn);
log.debug("resolved send to {} via turn-completion fallback ({} chars scraped)",
target, tail.length());
}
}
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
void fail(String target, InFlight turn) {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fall back to the currently-registered waiter (unambiguous — that send never completed, so
// no next turn exists to confuse it with).
CompletableFuture<Rendezvous.Resolution> waiter =
turn != null ? turn.waiter() : rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
inFlight.remove(target, turn); // nobody blocked on this worker — nothing to fail
return;
}
String reason;
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
}
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.debug("failed send to {} via turn-stall fallback", target);
}
}
private static String clip(String s) {
if (s == null) return "";
String trimmed = s.strip();
return trimmed.length() <= MAX_SCRAPE_CHARS
? trimmed
: trimmed.substring(trimmed.length() - MAX_SCRAPE_CHARS);
}
/**
* Extract the last assistant message from a raw Claude Code pane scrape (CB-115). Claude Code
* prefixes each assistant turn with {@code ⏺}; the delegator wants that answer, not the TUI
* chrome around it. Take everything from the final {@code ⏺} onward and stop at the <em>first</em>
* hard interface boundary below it — the spinner/status line, input box, {@code ❯} prompt (which
* may echo the <em>next</em> turn's text), footer, or tips/warnings. Stopping at the first
* boundary (rather than trimming only trailing chrome) is what keeps a following turn's echoed
* prompt out of this reply. Blank lines are not boundaries, so a multi-paragraph answer survives;
* trailing blanks are trimmed at the end. With no {@code ⏺} marker (an unusual render) the whole
* text is scanned the same way, so we never lose the reply.
*
* <p>Package-private and pure so it is unit-testable without herdr.
*/
static String lastAssistantBlock(String raw) {
if (raw == null || raw.isBlank()) return "";
int marker = raw.lastIndexOf('⏺');
String block = marker >= 0 ? raw.substring(marker + 1) : raw;
StringBuilder out = new StringBuilder();
int kept = 0;
for (String line : block.split("\n", -1)) {
if (isBoundary(line)) break; // first TUI boundary ends the assistant message
if (kept++ > 0) out.append('\n');
out.append(line);
}
return out.toString().strip();
}
/**
* A hard TUI boundary line that marks the end of an assistant message and the start of interface
* chrome (input box, prompt, spinner, footer, tips/warnings). Blank lines are <em>not</em>
* boundaries — an answer may contain them — so they are kept and trimmed only if trailing.
*/
private static boolean isBoundary(String line) {
String t = line.strip();
if (t.isEmpty()) return false;
// A horizontal rule / all box-drawing separators (e.g. "──────").
if (t.chars().allMatch(c -> c == '─' || c == '—' || c == '━' || c == '═' || c == '-')) {
return true;
}
String lower = t.toLowerCase();
return t.startsWith("╭") || t.startsWith("│") || t.startsWith("╰") || t.startsWith("┌")
|| t.startsWith("└") || t.startsWith("❯") || t.startsWith("⏵")
|| t.startsWith("⎿") || t.startsWith("⚠")
// Status/spinner lines Claude Code renders below a settled or in-flight turn,
// e.g. "✻ Baked for 21s", "✶ Forming…".
|| t.startsWith("✻") || t.startsWith("✳") || t.startsWith("✽") || t.startsWith("·")
|| t.startsWith("●") || t.startsWith("◐") || t.startsWith("✢") || t.startsWith("✶")
|| lower.contains("auto mode") || lower.contains("for shortcuts")
|| lower.contains("esc to interrupt") || lower.contains("bypass permissions");
}
}
@@ -12,6 +12,8 @@ import java.util.List;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Consumer;
import java.util.function.Predicate;
import java.util.stream.Collectors;
/**
@@ -32,6 +34,14 @@ import java.util.stream.Collectors;
* a transient {@code unknown} — counts as a real pickup, so a detection glitch can't prematurely
* release the latch. Perfectly reliable turn boundaries require a herdr {@code events.subscribe}
* stream; that is the intended upgrade and would replace only the sampling, not this queue.
*
* <p><strong>Turn completion (CB-106).</strong> Beyond delivery, the injector reports when a
* delegated turn <em>finishes</em>: after a delivery is picked up (a real {@code working} sample),
* the next injectable sample is a confirmed {@code working → idle} boundary and fires
* {@link TurnListener#onTurnComplete}. Completion is only ever synthesized from a <em>confirmed</em>
* turn — the pickup-grace path (a turn too fast to sample) unwedges the queue but does not fire
* completion, since without a sampled {@code working} there is no trustworthy "the worker just
* finished the task" signal to act on.
*/
public final class Injector {
@@ -45,11 +55,68 @@ public final class Injector {
*/
private static final int PICKUP_GRACE_POLLS = 8;
/**
* How many consecutive {@code unknown} samples while a delegation is outstanding before we
* declare it stalled and fire {@link TurnListener#onTurnFailed} (CB-109). A worker wedged in a
* state herdr can't classify (e.g. an API-error screen) stays {@code unknown} indefinitely and
* would otherwise never resolve; any {@code working}/{@code idle} sample resets the streak, so a
* transient detection glitch cannot trip it. At the 250ms poll interval this is ~30s — far longer
* than any real detection blip, and still vastly better than the async send's timeout.
*/
private static final int TURN_STALL_GRACE_POLLS = 120;
/**
* How many consecutive injectable samples a queued-but-undelivered message may wait on the
* {@link #ready} gate before we give up and fail it (CB-114). The gate holds a message out of a
* worker's boot window (herdr reports {@code idle} while its Claude is still starting), but a
* worker whose Claude crashes during boot — or never connects the bridge MCP — stays "idle and
* not ready" forever: {@link #ready} never accepts it, the message is never delivered, and the
* target would be polled indefinitely with its caller's future never completing. After this
* grace the queued messages are failed and the target released. At the 250ms poll interval this
* is ~60s — deliberately longer than {@link #TURN_STALL_GRACE_POLLS}, since a first boot (spawn
* + model load + MCP connect) legitimately takes longer than an in-turn detection blip.
*/
private static final int READINESS_GRACE_POLLS = 240;
private final AgentControl agents;
private final TurnListener turnListener;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
private final Consumer<String> forget; // CB-114: clear a gone worker's readiness/presence
private final ConcurrentHashMap<String, Target> targets = new ConcurrentHashMap<>();
/** Delivery only; completion signalling is a no-op and every target is treated as available. */
public Injector(AgentControl agents) {
this(agents, TurnListener.NOOP);
}
/** Delivery plus turn-completion signalling (CB-106); every target is treated as available. */
public Injector(AgentControl agents, TurnListener turnListener) {
this(agents, turnListener, _ -> true);
}
/**
* Delivery, completion signalling (CB-106), and a readiness gate (CB-113): a message is delivered
* only when {@code ready} accepts the target — i.e. the worker's Claude has connected the bridge
* MCP. This holds the first delivery out of the worker's boot window, where herdr already reports
* {@code idle} but the TUI would drop an injected paste.
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready) {
this(agents, turnListener, ready, _ -> {
});
}
/**
* Delivery, completion signalling (CB-106), a readiness gate (CB-113), and readiness cleanup
* (CB-114): {@code forget} is invoked with a target when its worker is gone — dropped
* (pane crash) or timed out on the readiness gate — so its stale presence/readiness is cleared
* and does not linger past the worker's life.
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
this.agents = agents;
this.turnListener = turnListener;
this.ready = ready;
this.forget = forget;
}
/** A pending message and the future that completes when it has been delivered. */
@@ -59,8 +126,12 @@ public final class Injector {
/** Per-worker delivery state, guarded by its own monitor (single writer per worker). */
private static final class Target {
final Deque<Pending> queue = new ArrayDeque<>();
boolean awaitingPickup; // sent a message, waiting for the worker to pick it up
boolean awaitingPickup; // sent a message, waiting for the worker to pick it up
int injectableSincePickup; // consecutive injectable samples while awaitingPickup
boolean awaitingCompletion; // a delivered message's turn is not yet known-complete
boolean turnObserved; // saw a real `working` sample since that delivery (turn ran)
int unknownSinceTurn; // consecutive `unknown` samples while a delegation is outstanding (CB-109)
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
synchronized void add(Pending p) {
queue.add(p);
@@ -99,63 +170,154 @@ public final class Injector {
Pending sent = null;
RuntimeException sendError = null;
boolean turnCompleted = false;
boolean turnFailed = false;
boolean resubmit = false;
List<Pending> notReady = null; // queued messages failed because the worker never became ready
synchronized (t) {
if (status == AgentStatus.WORKING) {
// Definitive pickup: the worker is busy on our last message.
// Definitive pickup: the worker is busy on our last message, and (if a delivery is
// outstanding) a real turn is now confirmed to be running.
t.awaitingPickup = false;
t.injectableSincePickup = 0;
t.unknownSinceTurn = 0;
t.notReadySincePoll = 0;
if (t.awaitingCompletion) t.turnObserved = true;
} else if (status.injectable()) { // IDLE or BLOCKED
if (t.awaitingPickup && ++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
// Pickup edge was never sampled (turn faster than the poll, or status lag).
// The worker has plainly moved on — release the latch rather than wedge.
t.awaitingPickup = false;
t.injectableSincePickup = 0;
t.unknownSinceTurn = 0;
if (t.awaitingPickup) {
if (++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
// Pickup edge was never sampled (turn faster than the poll, or status lag).
// Release the latch rather than wedge — and give up on synthesizing a
// completion for this message, since without a confirmed `working` we cannot
// trust that a task-processing turn actually ran.
t.awaitingPickup = false;
t.injectableSincePickup = 0;
t.awaitingCompletion = false;
t.turnObserved = false;
} else {
// Delivered but still idle → the worker hasn't picked it up; the submit
// keystroke likely raced the paste (esp. right as the TUI became ready).
// Re-nudge Enter (CB-113) until the worker starts (WORKING) or the grace ends.
resubmit = true;
}
}
if (!t.awaitingPickup) {
Pending p = t.queue.peek();
if (p != null) {
try {
agents.send(target, p.text());
t.queue.poll();
t.awaitingPickup = true;
t.injectableSincePickup = 0;
sent = p;
} catch (RuntimeException e) {
// Delivery failed at herdr; drop the poisoned message and surface it
// rather than blocking the queue behind it.
t.queue.poll();
sent = p;
sendError = e;
// A confirmed turn (a `working` sample was seen) that has now returned to idle is
// a trustworthy `working → idle` completion boundary.
if (t.awaitingCompletion && t.turnObserved) {
t.awaitingCompletion = false;
t.turnObserved = false;
turnCompleted = true;
}
// Deliver the next queued message only once the prior turn is fully settled, so a
// completion is never confused with the pickup of the following message — and only
// once the worker is available (CB-113), so we never paste into its boot window.
if (!t.awaitingCompletion) {
Pending p = t.queue.peek();
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
try {
agents.send(target, p.text());
t.queue.poll();
t.awaitingPickup = true;
t.awaitingCompletion = true;
t.turnObserved = false;
t.injectableSincePickup = 0;
sent = p;
} catch (RuntimeException e) {
// Delivery failed at herdr; drop the poisoned message and surface it
// rather than blocking the queue behind it.
t.queue.poll();
sent = p;
sendError = e;
}
} else if (p != null && ++t.notReadySincePoll >= READINESS_GRACE_POLLS) {
// The worker has been idle-but-not-ready for the whole grace: its Claude
// never connected the bridge MCP (crashed during boot, or wedged on a
// startup prompt). The readiness gate would hold this message forever, so
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
t.queue.clear();
t.notReadySincePoll = 0;
}
}
}
} else {
// UNKNOWN (or any other non-injectable, non-working): not a safe window nor a
// reliable pickup signal, so we never deliver or release the pickup latch here. But
// an outstanding delegation whose worker has gone unresponsive — stuck in a state
// herdr can't classify (CB-109) — will never yield a working→idle boundary. After a
// sustained streak, declare it failed so the awaiting send resolves rather than
// riding out the async timeout. (This also frees a delivery that wedged before it
// was ever picked up, which the injectable-only pickup grace could never release.)
if (t.awaitingCompletion && ++t.unknownSinceTurn >= TURN_STALL_GRACE_POLLS) {
t.awaitingPickup = false;
t.awaitingCompletion = false;
t.turnObserved = false;
t.unknownSinceTurn = 0;
turnFailed = true;
}
}
// UNKNOWN (and any other non-injectable, non-working): do nothing — neither a safe
// window nor a reliable pickup signal, so we must not deliver or release the latch.
// Reclaim the entry once the worker is fully quiescent, so the map cannot grow without
// bound across many short-lived workers.
if (t.queue.isEmpty() && !t.awaitingPickup) {
// Reclaim the entry once the worker is fully quiescent (nothing queued, no pickup or
// completion awaited), so the map cannot grow without bound across short-lived workers.
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion) {
targets.remove(target, t);
}
}
// Fire listeners / herdr calls after releasing the monitor so nothing runs on the poller
// thread while it holds the target lock.
if (resubmit) {
try {
agents.submit(target); // nudge a raced Enter so the pending paste submits
} catch (RuntimeException e) {
log.debug("resubmit to {} failed (will retry next poll): {}", target, e.getMessage());
}
}
if (notReady != null) {
// Worker never became available: forget its (never-set) readiness, unblock every queued
// caller, and route the awaiting send through the same failure path as a stalled turn so
// a blocking or async waiter resolves WORKER_FAILED rather than riding out the timeout.
forget.accept(target);
RuntimeException cause = new IllegalStateException(
target + " never became available (no bridge MCP connection within the boot window)");
for (Pending p : notReady) {
p.delivered().completeExceptionally(cause);
}
turnListener.onTurnFailed(target);
}
if (turnCompleted) {
turnListener.onTurnComplete(target);
}
if (turnFailed) {
turnListener.onTurnFailed(target);
}
if (sent != null) {
if (sendError != null) {
log.warn("inject to {} failed, dropped message: {}", target, sendError.getMessage());
sent.delivered().completeExceptionally(sendError);
} else {
// Baseline the pane's pre-turn content so a misattributed completion (no new output)
// can't resolve this send with the previous turn's stale answer (CB-115).
turnListener.onDelivered(target);
sent.delivered().complete(null);
}
}
}
/** Targets the poller must keep sampling: those with a queued message or an awaited pickup. */
/**
* Targets the poller must keep sampling: those with a queued message, an awaited pickup, or an
* awaited turn completion (so the {@code working → idle} boundary is observed).
*/
public Set<String> activeTargets() {
return targets.entrySet().stream()
.filter(e -> {
synchronized (e.getValue()) {
return !e.getValue().queue.isEmpty() || e.getValue().awaitingPickup;
Target t = e.getValue();
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion;
}
})
.map(java.util.Map.Entry::getKey)
@@ -163,20 +325,31 @@ public final class Injector {
}
/**
* Forget a target whose worker is gone, failing every still-queued message so awaiting
* callers unblock instead of hanging forever. Futures are completed after the monitor is
* released.
* Forget a target whose worker is gone, failing every still-queued message so awaiting callers
* unblock instead of hanging forever. If a message had already been <em>delivered</em> but its
* turn was not yet resolved (CB-110 — the worker vanished mid-turn, e.g. its pane crashed), fire
* {@link TurnListener#onTurnFailed} for it: a delivered message is no longer in the queue, so
* failing queued waiters alone would leave that send's rendezvous hanging until the async
* timeout. Futures and listeners are completed after the monitor is released.
*/
public void drop(String target, Throwable cause) {
Target t = targets.remove(target);
if (t == null) return;
List<Pending> pending;
boolean hadDeliveredTurn;
synchronized (t) {
pending = new ArrayList<>(t.queue);
t.queue.clear();
hadDeliveredTurn = t.awaitingCompletion;
t.awaitingCompletion = false;
t.awaitingPickup = false;
}
forget.accept(target); // the worker is gone — clear its readiness/presence too (CB-114)
for (Pending p : pending) {
p.delivered().completeExceptionally(cause);
}
if (hadDeliveredTurn) {
turnListener.onTurnFailed(target);
}
}
}
@@ -23,13 +23,20 @@ public final class StatusPoller {
private final AgentControl agents;
private final Injector injector;
private final StatusRefiner refiner;
private final long intervalMillis;
private volatile boolean running;
private Thread thread;
public StatusPoller(AgentControl agents, Injector injector, long intervalMillis) {
this(agents, injector, new StatusRefiner(agents), intervalMillis);
}
public StatusPoller(AgentControl agents, Injector injector, StatusRefiner refiner,
long intervalMillis) {
this.agents = agents;
this.injector = injector;
this.refiner = refiner;
this.intervalMillis = intervalMillis;
}
@@ -47,7 +54,9 @@ public final class StatusPoller {
for (String target : active) {
if (!running) return;
try {
AgentStatus status = agents.status(target);
// herdr's agent_status can misreport a settled worker as `unknown`; refine it
// against the pane content before it drives delivery/completion (CB-115).
AgentStatus status = refiner.refine(target, agents.status(target));
injector.onStatus(target, status);
} catch (HerdrException e) {
// The worker's agent is gone — stop trying and unblock its waiters.
@@ -0,0 +1,89 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
/**
* Refines an unreliable {@link AgentStatus#UNKNOWN} into a real state by reading the worker's
* terminal content (CB-115).
*
* <p>Some workers' panes are misclassified by herdr as {@code unknown} even when the worker is
* plainly settled at an idle prompt (empty {@code ❯}, "auto mode on" footer, a completed
* {@code ⏺} answer above). Left as {@code UNKNOWN} that both <em>wedges delivery</em> — the
* status-gated {@link Injector} only injects into an {@link AgentStatus#injectable} worker — and
* <em>mis-fires the CB-109 stall failure</em> on a worker that has actually answered. herdr's
* {@code agent_status} is a heuristic; the pane content is the ground truth.
*
* <p>The refinement only ever runs on a raw {@code UNKNOWN} sample (every other status is trusted
* as-is), so a healthy worker adds zero extra herdr traffic; a persistently-{@code unknown} worker
* costs one extra {@code agent.read} per poll while it has work outstanding. Classification is
* deliberately conservative — it upgrades {@code UNKNOWN} to {@link AgentStatus#WORKING} or
* {@link AgentStatus#IDLE} only on a clear signal, and leaves a genuinely unclassifiable screen
* (e.g. a wedged error state) as {@code UNKNOWN} so the CB-109 stall path can still fail it.
*/
public final class StatusRefiner {
private static final Logger log = LoggerFactory.getLogger(StatusRefiner.class);
/**
* herdr {@code agent.read} source used to inspect the pane. {@code detection} is the region
* herdr itself uses for status detection (the prompt/footer tail), which is exactly what we
* need to tell "idle at prompt" from "mid-turn".
*/
static final String PROBE_SOURCE = "detection";
private final AgentControl agents;
public StatusRefiner(AgentControl agents) {
this.agents = agents;
}
/**
* Return a trustworthy status for {@code target}. Any non-{@code UNKNOWN} {@code raw} is returned
* unchanged; an {@code UNKNOWN} triggers a pane read and content classification. A read failure
* leaves it {@code UNKNOWN} (the safe default: no delivery, and the stall path still applies).
*/
public AgentStatus refine(String target, AgentStatus raw) {
if (raw != AgentStatus.UNKNOWN) return raw;
String pane;
try {
pane = agents.read(target, PROBE_SOURCE);
} catch (RuntimeException e) {
log.debug("status refine read for {} failed; leaving UNKNOWN: {}", target, e.getMessage());
return AgentStatus.UNKNOWN;
}
AgentStatus refined = classify(pane);
if (refined != AgentStatus.UNKNOWN) {
log.debug("refined {} from UNKNOWN to {} via pane content", target, refined);
}
return refined;
}
/**
* Classify a Claude Code TUI pane tail. Package-private and pure so it is unit-testable without
* herdr.
*
* <ul>
* <li>An active-generation marker ({@code esc to interrupt}) ⇒ {@link AgentStatus#WORKING} —
* never inject here.</li>
* <li>Otherwise, an interactive input prompt with no active-turn marker ({@code ❯}, the
* {@code │ >} input box, or the idle {@code auto mode} / shortcuts footer) ⇒
* {@link AgentStatus#IDLE} — settled and safe to inject / a completed turn.</li>
* <li>Anything else (blank, or an unrecognizable screen) ⇒ {@link AgentStatus#UNKNOWN}.</li>
* </ul>
*/
static AgentStatus classify(String pane) {
if (pane == null || pane.isBlank()) return AgentStatus.UNKNOWN;
String lower = pane.toLowerCase();
// Claude Code shows "(esc to interrupt)" only while a turn is actively generating.
if (lower.contains("esc to interrupt")) return AgentStatus.WORKING;
// A settled, ready input prompt with no active-turn marker = idle-at-prompt.
boolean readyPrompt = pane.contains("❯")
|| pane.contains("│ >")
|| lower.contains("auto mode on")
|| lower.contains("? for shortcuts");
return readyPrompt ? AgentStatus.IDLE : AgentStatus.UNKNOWN;
}
}
@@ -0,0 +1,39 @@
package dev.ltms.bridged.inject;
/**
* Notified when a worker's delegated turn is observed to complete — a confirmed
* {@code WORKING → IDLE} transition after a delivery. This is the CB-106 completion signal the
* {@code CompletionResolver} uses to resolve a blocked send whose worker never called
* {@code bridge_reply}. Kept as a seam so the {@link Injector} needs no dependency on the message
* layer and stays unit-testable with a capturing fake.
*/
@FunctionalInterface
public interface TurnListener {
/** A worker's delegated turn finished (worker returned to idle after visibly working). */
void onTurnComplete(String target);
/**
* A worker that visibly ran a delegated turn then wedged in a non-idle, non-working state
* (CB-109) — e.g. an error screen herdr classifies as {@code unknown} — so no
* {@code working → idle} completion boundary will ever arrive. A default no-op keeps this a
* functional interface; the completion resolver overrides it to fail the awaiting send.
*/
default void onTurnFailed(String target) {
}
/**
* A message was just delivered into {@code target}'s pane (CB-115). Fired so the completion
* resolver can snapshot the pane's pre-turn content: a later {@link #onTurnComplete} whose
* scrape is unchanged from this baseline is a <em>misattributed</em> boundary (e.g. the prior
* turn's wind-down sampled as this turn's completion on rapid back-to-back sends) and must not
* resolve the send with the previous turn's stale answer. A default no-op keeps the interface
* functional for callers that don't scrape.
*/
default void onDelivered(String target) {
}
/** No-op default for callers that only need delivery, not completion signalling. */
TurnListener NOOP = _ -> {
};
}
@@ -0,0 +1,38 @@
package dev.ltms.bridged.inject;
import java.util.concurrent.ConcurrentHashMap;
import java.util.Set;
/**
* Tracks which workers are <em>available</em> — their Claude has booted and connected its MCP client
* to the bridge (CB-113). This is the reliable readiness signal, unlike herdr's {@code agent_status},
* which reports {@code idle} for a worker whose Claude is still booting. Delivering into that boot
* window pastes into a not-yet-ready TUI (the text is lost) and wedges the worker's delivery state,
* so the {@link Injector} holds the first delivery until the worker is present here.
*
* <p>Populated from the MCP transport: any MCP request whose connection resolves to a worker terminal
* marks that worker present (its {@code initialize} is the first such contact). A worker that never
* mounts the bridge MCP is never marked present — its sends stay queued until they time out, which is
* correct (it could not have replied anyway).
*/
public class WorkerPresence {
private final Set<String> present = ConcurrentHashMap.newKeySet();
/** Record that {@code terminal}'s worker has connected its MCP client (is available). */
public void markPresent(String terminal) {
if (terminal != null && !terminal.isBlank()) {
present.add(terminal);
}
}
/** Whether {@code terminal}'s worker is available (has been seen on the bridge MCP). */
public boolean isPresent(String terminal) {
return present.contains(terminal);
}
/** Forget a torn-down worker so its terminal id does not linger as "present". */
public void forget(String terminal) {
present.remove(terminal);
}
}
@@ -0,0 +1,793 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.AuditLog;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.auth.Role;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.peer.PeerLauncher;
import io.modelcontextprotocol.common.McpTransportContext;
import io.modelcontextprotocol.json.McpJsonMapper;
import io.modelcontextprotocol.json.jackson3.JacksonMcpJsonMapperSupplier;
import io.modelcontextprotocol.server.McpServer;
import io.modelcontextprotocol.server.McpSyncServer;
import io.modelcontextprotocol.server.McpSyncServerExchange;
import io.modelcontextprotocol.server.transport.HttpServletStreamableServerTransportProvider;
import io.modelcontextprotocol.spec.McpSchema;
import com.fasterxml.jackson.databind.ObjectMapper;
import jakarta.servlet.http.HttpServlet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.function.Function;
import java.util.stream.Collectors;
/**
* The MCP SERVER face (CB-105): a Streamable-HTTP MCP server whose tools are <em>thin adapters</em>
* over the same {@link MessageService}/{@link Rendezvous} the REST routes use — so the two are
* validated by parity, not by re-implementing behaviour. The primary Opus calls {@code bridge_send}
* / {@code bridge_status}; the worker calls {@code bridge_reply}.
*
* <p>Beyond delegation the primary also manages the fleet here (CB-108): {@code bridge_spawn} /
* {@code bridge_list} / {@code bridge_stop} drive the {@link PeerLauncher} SPI so a worker's whole
* lifecycle is managed through MCP, with each adapter's subscription boundary enforced inside it.
*
* <p>The tool <em>logic</em> lives in package-private static methods returning a
* {@link McpSchema.CallToolResult}, so it is unit-testable without standing up the HTTP transport;
* the SDK owns the wire protocol. Mount {@link #servlet()} at {@code /mcp} on the daemon's Jetty.
*/
public final class BridgeMcp {
private static final long DEFAULT_TIMEOUT_MS = 25_000;
private static final long MAX_TIMEOUT_MS = 120_000;
// bridge_ask blocks the WORKER's own MCP call, which its client caps near 60s — default under
// that so the bridge returns a clean timeout before the client severs the call (CB-205).
private static final long ASK_DEFAULT_TIMEOUT_MS = 55_000;
private static final long ASK_MAX_TIMEOUT_MS = 115_000;
private static final ObjectMapper MAPPER = new ObjectMapper(); // worker-view JSON projections
/** Transport-context key under which the extractor stashes the resolved caller identity. */
static final String CALLER_TERMINAL = "callerTerminal";
/** Transport-context key under which the extractor stashes the caller's PID (for cwd inherit). */
static final String CALLER_PID = "callerPid";
/** Transport-context key under which the extractor stashes the resolved {@link Role} (CB-501). */
static final String CALLER_ROLE = "callerRole";
private final HttpServletStreamableServerTransportProvider transport;
private final McpSyncServer server;
private final CallerResolver authz; // CB-501: null → authorization not enforced (legacy)
private final Metrics metrics; // CB-502: null → auth failures not counted
/**
* Legacy constructor — no authorization. Retained so existing tests exercise tool behaviour
* without an auth fixture.
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
PrimaryRegistry primaryRegistry) {
this(messages, workers, sessions, identity, presence, primaryRegistry, null, null);
}
/**
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
* Jetty's context handler and never passes through Javalin's {@code before}
* filter, so the REST guard does not cover it.
* @param metrics registry for auth-failure counting; may be {@code null}
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
PrimaryRegistry primaryRegistry, CallerResolver callers, Metrics metrics) {
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
.jsonMapper(json)
.mcpEndpoint("/mcp")
// Resolve the caller from the connection (peer PID → herdr pane) in one lookup: the
// worker terminal for bridge_reply (no spoofable arg), and the PID so bridge_spawn can
// inherit the primary's cwd (CB-112). Any contact from a worker marks it available
// (CB-113) — its MCP initialize is the reliable "the agent is up" signal.
.contextExtractor(req -> {
// One resolution per call, shared with the REST surface via CallerResolver so
// the two paths cannot drift on who a caller is.
Principal p = callers != null
? callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization"))
: legacyPrincipal(identity, req.getRemoteAddr(), req.getRemotePort());
presence.markPresent(p.terminal()); // no-op for the primary (null terminal)
return McpTransportContext.create(Map.of(
CALLER_TERMINAL, orEmpty(p.terminal()),
CALLER_PID, Long.toString(p.pid()),
CALLER_ROLE, p.role().name()));
})
.build();
this.server = McpServer.sync(transport)
.serverInfo("bridge", "0.1.0")
.capabilities(McpSchema.ServerCapabilities.builder().tools(true).build())
.toolCall(sendTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SEND,
str(req.arguments(), "sessionId"));
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
Map<String, Object> a = req.arguments();
String turnId = str(a, "turnId");
if (turnId != null && !turnId.isBlank()) {
// Answering a worker's bridge_ask (CB-205): resolve its blocked question and
// block for the worker's reply as it resumes the same turn.
return answer(messages, turnId, str(a, "content"), timeoutMs(a));
}
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, str(a, "sessionId"), str(a, "content"))
: send(messages, str(a, "sessionId"), str(a, "content"), timeoutMs(a));
})
// bridge_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
.toolCall(replyTool(), (exchange, req) -> {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.REPLY, self);
if (denied != null) return denied;
return reply(messages, self, str(req.arguments(), "content"));
})
// bridge_ask (CB-205): a worker's mid-turn question — identity from the CONNECTION.
.toolCall(askTool(), (exchange, req) -> {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.ASK, self);
if (denied != null) return denied;
return ask(messages, self, str(req.arguments(), "question"), timeoutMs(req.arguments()));
})
.toolCall(statusTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return status(messages, str(req.arguments(), "sessionId"));
})
.toolCall(pollTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
Map<String, Object> a = req.arguments();
return poll(messages, str(a, "ticket"), str(a, "target"));
})
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
// Acking removes a reply from the inbox, so it is a drain, not a read.
.toolCall(ackTool(), (exchange, req) -> {
Map<String, Object> a = req.arguments();
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.DRAIN, str(a, "target"));
if (denied != null) return denied;
return ack(messages, str(a, "target"), str(a, "msgId"));
})
// Fleet management (CB-108): spawn/list/stop over ClaudeCodeLauncher.
.toolCall(spawnTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SPAWN, null);
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
Map<String, Object> a = req.arguments();
// CB-112: worker inherits the primary's cwd unless the call pins one.
// CB-301: carry the caller's identity as the session owner (null for the primary).
// CB-301-ext: optional isolated worktree for parallel implementers.
String callerCwd = identity.cwdForPid(callerPid(exchange));
return spawn(sessions, str(a, "profile"), str(a, "cwd"), callerCwd,
callerTerminal(exchange), worktreeRequest(a));
})
.toolCall(listTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listWorkers(workers, sessions);
})
.toolCall(stopTool(), (exchange, req) -> {
String paneId = str(req.arguments(), "paneId");
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.STOP, paneId);
if (denied != null) return denied;
return stop(sessions, paneId);
})
.toolCall(profilesTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return profiles(workers);
})
.toolCall(whoamiTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return whoami(principal(exchange), sessions);
})
.build();
this.authz = callers;
this.metrics = metrics;
}
/**
* Pre-CB-501 identity: worker if the connection maps to a pane, otherwise the primary. Used
* only by the legacy constructor, where authorization is not enforced anyway.
*/
private static Principal legacyPrincipal(ConnectionIdentity identity, String addr, int port) {
ConnectionIdentity.Caller c = identity.resolve(addr, port);
return c.terminal() != null
? Principal.worker(c.terminal(), c.pid())
: Principal.primary(c.pid());
}
/** The caller reconstructed from the transport context. */
private static Principal principal(McpSyncServerExchange exchange) {
return principalFrom(exchange.transportContext().get(CALLER_ROLE),
callerTerminal(exchange), callerPid(exchange));
}
/**
* Rebuild a {@link Principal} from the three values the context extractor stashed.
*
* <p>Split out from {@link #principal(McpSyncServerExchange)} so the identity rules are
* reachable without an {@code McpSyncServerExchange} — that is an SDK type this project has no
* mocking library to fabricate, which is why this logic had no test at all until CB-513.
*
* @param role the stashed {@link Role} name, or {@code null} on the legacy path
* @param terminal the worker terminal, or {@code null} for a non-worker
* @param pid the calling pid, or {@code -1}
*/
static Principal principalFrom(Object role, String terminal, long pid) {
if (role == null) {
// No role stashed (legacy path): fall back to the historical interpretation.
return terminal != null ? Principal.worker(terminal, pid) : Principal.primary(pid);
}
return new Principal(Role.valueOf(role.toString()), terminal, pid);
}
/**
* Gate a tool call on the CB-505 table. Returns {@code null} when the call may proceed, or the
* error result to return when it may not.
*/
private McpSchema.CallToolResult deny(McpSyncServerExchange exchange, Authz.Action action,
String target) {
return denyFor(principal(exchange), action, target);
}
/**
* The policy half of {@link #deny}: everything except pulling the caller out of the MCP
* exchange. Kept separate so the authorization decision — the actual control — is unit-testable
* without fabricating an SDK {@code McpSyncServerExchange}.
*
* <p>This surface exists because the enforcement was previously unreachable from a test: no
* test constructs a {@code BridgeMcp}, so the whole MCP-side gate ran zero times in the suite
* while the REST-side equivalent had ten tests. A security control nothing exercises is a
* claim, not a control.
*
* @return {@code null} when the call may proceed, or the error result to return when it may not
*/
McpSchema.CallToolResult denyFor(Principal caller, Authz.Action action, String target) {
// The enforcement switch lives HERE rather than in the exchange-facing wrapper: any future
// tool that calls this directly must not be able to skip the gate by accident.
if (authz == null) {
return null; // legacy constructor: authorization not enforced
}
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
return null;
}
String reason = Authz.isUnauthenticated(caller) ? "unauthenticated" : "forbidden";
AuditLog.denied(caller, action, target, reason);
if (metrics != null) {
metrics.inc(BridgedMetrics.AUTH_FAILURES, "reason", reason);
}
return error(reason + ": " + caller.describe() + " may not " + action);
}
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
private static String callerTerminal(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_TERMINAL);
String s = v == null ? null : v.toString();
return (s == null || s.isBlank()) ? null : s;
}
/** The caller's PID resolved from this call's connection, or {@code -1} if unknown. */
private static long callerPid(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_PID);
try {
return v == null ? -1 : Long.parseLong(v.toString());
} catch (NumberFormatException e) {
return -1;
}
}
private static String orEmpty(String s) {
return s == null ? "" : s;
}
/** The Streamable-HTTP servlet to mount at {@code /mcp} on the daemon's Jetty. */
public HttpServlet servlet() {
return transport;
}
/** Graceful shutdown of the MCP server. */
public void close() {
server.closeGracefully();
}
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
/** {@code bridge_send}: delegate {@code content} to a worker session and block for its reply. */
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content, Long timeoutMs) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
try {
return formatReply(messages.send(sessionId, content, timeout), timeout);
} catch (HerdrException e) {
return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
}
}
/**
* {@code bridge_send} carrying a {@code turnId}: the primary's answer to a worker's
* {@code bridge_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as
* it resumes the same turn — surfaced to the primary identically to a normal send.
*/
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs) {
if (isBlank(turnId) || isBlank(content)) {
return error("turnId and content are required to answer a worker's question");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
return formatReply(messages.answer(turnId, content, timeout), timeout);
}
/**
* {@code bridge_ask} (CB-205): a worker pauses its delegated turn to ask the primary, blocking
* until the primary answers. The worker is identified by its connection ({@code callerTerminal}),
* never an argument — a {@code null} means the caller is not a known worker.
*/
static McpSchema.CallToolResult ask(MessageService messages, String callerTerminal, String question, Long timeoutMs) {
if (callerTerminal == null) {
return error("bridge_ask is for workers only — could not identify the calling worker "
+ "from the connection");
}
if (isBlank(question)) {
return error("question is required");
}
long timeout = Math.clamp(timeoutMs == null ? ASK_DEFAULT_TIMEOUT_MS : timeoutMs, 1, ASK_MAX_TIMEOUT_MS);
MessageService.AskResult r = messages.ask(callerTerminal, question, timeout);
return switch (r.outcome()) {
case ANSWERED -> text(r.answer());
case NO_WAITER -> error("no primary is awaiting this turn — bridge_ask only works while a "
+ "bridge_send delegation is open to answer it");
case TIMED_OUT -> text("[no answer within " + timeout + "ms — the primary did not respond; "
+ "proceed on your best judgement, then call bridge_reply to end the turn]");
};
}
/** Render a {@link MessageService.Reply} as a tool result — shared by {@link #send} and {@link #answer}. */
private static McpSchema.CallToolResult formatReply(MessageService.Reply r, long timeout) {
return switch (r.outcome()) {
case REPLIED -> text(r.text());
// The worker's turn finished but it never called bridge_reply — hand back the scraped
// transcript tail, flagged so the primary knows it isn't a structured reply.
case COMPLETED_UNREPLIED -> text(
"[worker finished without a structured bridge_reply — transcript tail follows]\n" + r.text());
// The worker ran the turn then wedged (CB-109) — surface the error context.
case WORKER_FAILED -> text("[worker failed — turn ended in an unrecoverable state]\n" + r.text());
// The worker paused mid-turn to ask (CB-205) — tell the primary how to answer in-turn.
case QUESTION -> text("[question] the worker paused to ask before it can finish:\n" + r.text()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + r.turnId()
+ "\" and content set to your answer; the worker resumes the same turn.");
case STALE_TURN -> error("that question is no longer open — it timed out or was already "
+ "answered (turnId stale)");
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> text("[no reply within " + timeout + "ms — worker "
+ r.outcome().name().toLowerCase().replace("timed_out_", "") + "; retry or poll status]");
};
}
/**
* {@code bridge_send} with {@code wait:false}: delegate {@code content} and return a ticket
* immediately (fire-and-poll), so a long task isn't cut off by the caller's MCP call timeout.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
String ticket = messages.sendAsync(sessionId, content);
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
}
/** {@code bridge_poll}: check an async delegation by ticket, or drain a worker's inbox by target. */
static McpSchema.CallToolResult poll(MessageService messages, String ticket, String target) {
if (!isBlank(target)) {
var replies = messages.drainReplies(target);
if (replies.isEmpty()) {
return text("[]");
}
return text(json(replies));
}
if (isBlank(ticket)) {
return error("ticket (or target) is required");
}
MessageService.TaskView v = messages.poll(ticket);
if (v == null) {
return error("unknown ticket: " + ticket + " (never issued, or expired)");
}
return switch (v.phase()) {
case DONE -> text(v.replySource() != null && v.replySource().equals("transcript")
? "[done — worker finished without a structured bridge_reply; transcript tail follows]\n" + v.reply()
: v.reply());
case PENDING -> text("[pending — " + v.detail() + "]");
case FAILED -> text("[failed — " + v.detail() + "]");
};
}
/**
* {@code bridge_reply}: the worker returns its structured answer, resolving the awaiting send
* or — when no send is open — queueing the reply in the inbox for later drain (CB-307).
* {@code callerTerminal} is resolved from the connection (never an argument); a {@code null}
* means the caller is not a known worker (e.g. the primary called it by mistake).
*/
static McpSchema.CallToolResult reply(MessageService messages, String callerTerminal, String content) {
if (callerTerminal == null) {
return error("bridge_reply is for workers only — could not identify the calling worker "
+ "from the connection");
}
if (content == null) {
return error("content is required");
}
messages.reply(callerTerminal, content);
return text("delivered");
}
/** {@code bridge_ack}: acknowledge (remove) a specific reply from the inbox. */
static McpSchema.CallToolResult ack(MessageService messages, String target, String msgId) {
if (isBlank(target) || isBlank(msgId)) {
return error("target and msgId are required");
}
messages.ackReply(target, msgId);
return text("acknowledged " + msgId);
}
/** {@code bridge_status}: the live lifecycle status of a worker session. */
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
if (isBlank(sessionId)) {
return error("sessionId is required");
}
try {
return text(messages.status(sessionId).name().toLowerCase());
} catch (HerdrException e) {
return error("herdr error for session " + sessionId + ": " + e.getMessage());
}
}
/**
* {@code bridge_whoami}: the caller's own identity, as the daemon already resolved it.
*
* <p>Every other tool <em>consumes</em> this identity — the authorization gate, the reply
* rendezvous, the cwd inherit — but none reported it, so an agent had to infer its own role
* from side channels the daemon does not control: a charter string in its system prompt, the
* name its MCP mount happens to carry, or {@code ANTHROPIC_BASE_URL} (which Claude-model
* workers do not set). The failure mode of guessing is asymmetric and silent: a primary that
* mistakes itself for a worker is refused by {@link Authz} and learns immediately, while a
* worker that mistakes itself for the primary ends its turn without {@code bridge_reply} and
* the sender simply receives nothing. This tool removes the guess.
*
* <p>For a worker the session registry adds what it knows about that session. A worker the
* registry has no record of — one that outlived a daemon restart — still gets its role and
* {@code sessionId}, which is the load-bearing part.
*/
static McpSchema.CallToolResult whoami(Principal caller, SessionManager sessions) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("role", caller.role().name().toLowerCase());
if (!caller.isWorker()) {
return text(json(m));
}
m.put("sessionId", caller.terminal());
sessions.roster().stream()
.filter(s -> caller.terminal().equals(s.terminalId()))
.findFirst()
.ifPresent(s -> {
m.put("paneId", s.paneId());
m.put("profile", s.profile());
m.put("state", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
if (s.ownerTerminal() != null) {
m.put("owner", s.ownerTerminal());
}
});
return text(json(m));
}
// --- fleet management logic (CB-108 / CB-301) --------------------------------------------
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
return spawn(sessions, profile, null, null, null, null);
}
/**
* {@code bridge_spawn}: launch a guard-checked worker for {@code profile} (blank → the default
* profile) and return its session id + pane id. The worker's cwd is {@code requestedCwd} if given,
* else the profile's config, else {@code callerCwd} (the primary's directory), else the daemon's.
* CB-301: the session is registered with {@code ownerTerminal} as its owner.
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
*/
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile,
String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest worktreeRequest) {
try {
WorkerSession worker = sessions.acquire(isBlank(profile) ? null : profile,
requestedCwd, callerCwd, ownerTerminal, worktreeRequest);
return text(json(workerView(worker)));
} catch (GuardException e) {
return error("subscription boundary: " + e.getMessage());
} catch (IllegalArgumentException e) {
return error(e.getMessage()); // unknown / no-default profile
} catch (PeerUnreachableException e) {
return error("spawn timed out — worker pane never reached injectable state: " + e.getMessage());
} catch (HerdrException e) {
return error("herdr error spawning worker: " + e.getMessage());
}
}
/** Build a {@link WorktreeRequest} from {@code bridge_spawn}'s optional {@code worktree}/{@code ticket} args. */
private static WorktreeRequest worktreeRequest(Map<String, Object> a) {
Object w = a.get("worktree");
if (w == null || Boolean.FALSE.equals(w)) {
return null;
}
String ticket = str(a, "ticket");
if (w instanceof String s) {
if (s.isBlank() || "false".equalsIgnoreCase(s)) {
return null;
}
if ("true".equalsIgnoreCase(s)) {
if (isBlank(ticket)) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return new WorktreeRequest(s, null);
}
if (w instanceof Boolean b && b) {
if (isBlank(ticket)) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return null;
}
/** {@code bridge_profiles}: the configured worker profiles and the default. */
static McpSchema.CallToolResult profiles(PeerLauncher workers) {
return text(json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
}
/** {@code bridge_list}: bridge-owned roster merged with live herdr status by paneId. */
static McpSchema.CallToolResult listWorkers(PeerLauncher workers, SessionManager sessions) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.paneId() != null)
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.paneId())))
.toList();
return text(json(Map.of("workers", out)));
} catch (HerdrException e) {
return error("herdr error listing workers: " + e.getMessage());
}
}
/** {@code bridge_stop}: tear a worker down by its pane id. */
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
if (isBlank(paneId)) {
return error("paneId is required");
}
try {
sessions.release(paneId);
return text("stopped " + paneId);
} catch (HerdrException e) {
return error("herdr error stopping " + paneId + ": " + e.getMessage());
}
}
/** CB-301 projection from the authoritative session registry. */
private static Map<String, Object> workerView(WorkerSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", s.terminalId());
m.put("paneId", s.paneId());
m.put("status", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
return m;
}
private static String json(Object o) {
try {
return MAPPER.writeValueAsString(o);
} catch (Exception e) {
return String.valueOf(o);
}
}
// --- tool schemas --------------------------------------------------------------------------
private static McpSchema.Tool sendTool() {
return tool("bridge_send",
"Delegate a task to a worker session. By default blocks until the worker replies and "
+ "returns its reply (or a 'still working / queued' note on timeout). Pass wait:false "
+ "for a long task to return a ticket immediately, then poll it with bridge_poll. To "
+ "answer a worker's bridge_ask, pass its turnId (with content) instead of sessionId.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id (herdr terminal_id) to delegate to"),
"content", stringProp("The task/message to send to the worker (or your answer, with turnId)"),
"timeoutMs", Map.of("type", "integer", "description", "Max ms to wait for a reply (blocking mode)"),
"wait", Map.of("type", "boolean",
"description", "Block for the reply (default true); false returns a ticket to poll"),
"turnId", stringProp("When answering a worker's bridge_ask, its question turnId — "
+ "routes your answer back into the same turn (omit for a normal delegation)")),
List.of("content")));
}
private static McpSchema.Tool askTool() {
// No target/session arg — the worker's identity is resolved from the connection.
return tool("bridge_ask",
"Pause your current delegated turn to ask the primary a question, blocking until it "
+ "answers — then resume the same turn with the answer. Use this when only the "
+ "primary has a decision or detail you need to continue. You do not address the "
+ "primary; identity is resolved from your connection.",
objectSchema(Map.of(
"question", stringProp("The question to put to the primary"),
"timeoutMs", Map.of("type", "integer",
"description", "Max ms to wait for the primary's answer")),
List.of("question")));
}
private static McpSchema.Tool pollTool() {
return tool("bridge_poll",
"Check an async delegation (a bridge_send with wait:false) by its ticket: "
+ "pending, done (with the worker's reply), or failed. When target (a worker "
+ "session id) is present instead of ticket, drain that worker's inbox of "
+ "replies delivered when no send was open.",
objectSchema(Map.of(
"ticket", stringProp("The ticket returned by bridge_send wait:false"),
"target", stringProp("Worker session id to drain pending replies from (optional)")),
List.of()));
}
private static McpSchema.Tool ackTool() {
return tool("bridge_ack",
"Acknowledge (remove) a specific reply from a worker's inbox. Use when the primary "
+ "has processed a reply and wants to confirm it, leaving other pending replies "
+ "in the inbox for later drain.",
objectSchema(Map.of(
"target", stringProp("Worker session id whose inbox to ack from"),
"msgId", stringProp("The message id to acknowledge")),
List.of("target", "msgId")));
}
private static McpSchema.Tool spawnTool() {
return tool("bridge_spawn",
"Spawn a new off-subscription worker session. Pass a profile (from bridge_profiles) to "
+ "pick the backend, or omit it for the default. The worker opens your current "
+ "directory by default; pass cwd to pin a different one. Pass worktree:true (with "
+ "ticket) or worktree:<ticket-slug> to provision an isolated git worktree. "
+ "Returns the worker's sessionId (use with bridge_send) and paneId (use with bridge_stop).",
objectSchema(Map.of(
"profile", stringProp("Worker profile to spawn (omit for the default profile)"),
"cwd", stringProp("Working directory for the worker (omit to inherit yours)"),
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
"ticket", stringProp("Ticket slug when worktree:true")),
List.of()));
}
private static McpSchema.Tool profilesTool() {
return tool("bridge_profiles",
"List the configured worker profiles (backends) and which one bridge_spawn uses by default.",
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool listTool() {
return tool("bridge_list",
"List the worker sessions the bridge tracks — each with its sessionId, paneId, profile, "
+ "state, optional worktree/branch/owner, and live herdr status.",
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool stopTool() {
return tool("bridge_stop",
"Tear down a worker session by its paneId (from bridge_spawn or bridge_list).",
objectSchema(Map.of(
"paneId", stringProp("The worker's paneId to stop")),
List.of("paneId")));
}
private static McpSchema.Tool replyTool() {
// No session/target arg — the worker's identity is resolved from the connection.
return tool("bridge_reply",
"Return your structured answer for the task you were delegated, "
+ "resolving the caller's blocked bridge_send.",
objectSchema(Map.of(
"content", stringProp("Your reply/answer")),
List.of("content")));
}
private static McpSchema.Tool statusTool() {
return tool("bridge_status",
"Get the live lifecycle status (idle/working/blocked/unknown) of a worker session.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id to query")),
List.of("sessionId")));
}
private static McpSchema.Tool whoamiTool() {
return tool("bridge_whoami",
"Report who YOU are on the bridge — your role is resolved from your connection "
+ "(unforgeable), never from anything you claim. Returns role 'primary' (you "
+ "orchestrate: spawn/send/stop, and you must never call bridge_reply) or "
+ "'worker' (you were delegated to: you must end every turn with exactly one "
+ "bridge_reply, and cannot spawn or send), plus your own sessionId, profile, "
+ "worktree and branch when you are a worker. Call this first when following "
+ "role-conditional instructions rather than guessing your role.",
objectSchema(Map.of(), List.of()));
}
// --- small helpers -------------------------------------------------------------------------
// The SDK 2.0.0 deprecates its own Tool builders without a stable replacement — isolate it here.
@SuppressWarnings("deprecation")
private static McpSchema.Tool tool(String name, String description, Map<String, Object> inputSchema) {
return McpSchema.Tool.builder(name).description(description).inputSchema(inputSchema).build();
}
private static Map<String, Object> objectSchema(Map<String, Object> properties, List<String> required) {
return Map.of("type", "object", "properties", properties, "required", required);
}
private static Map<String, Object> stringProp(String description) {
return Map.of("type", "string", "description", description);
}
private static McpSchema.CallToolResult text(String s) {
return McpSchema.CallToolResult.builder().addTextContent(s == null ? "" : s).build();
}
private static McpSchema.CallToolResult error(String s) {
return McpSchema.CallToolResult.builder().addTextContent(s).isError(true).build();
}
private static String str(Map<String, Object> args, String key) {
Object v = args.get(key);
return v == null ? null : v.toString();
}
private static Long timeoutMs(Map<String, Object> args) {
Object v = args.get("timeoutMs");
return v instanceof Number n ? n.longValue() : null;
}
private static long clamp(long ms) {
return Math.clamp(ms, 1, MAX_TIMEOUT_MS);
}
private static boolean isBlank(String s) {
return s == null || s.isBlank();
}
}
@@ -0,0 +1,66 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.herdr.PaneLocator;
/**
* Resolves <em>who is calling</em> an MCP tool from the connection alone — the anti-spoofing
* identity model of the MCP contract. It ties the connection's loopback peer PID (from the OS)
* to a herdr agent pane (from herdr), yielding the caller's worker {@code terminal_id}. A caller
* that maps to no worker pane — the primary, or an off-host client — resolves to {@code null}.
*
* <p>Both sources are authoritative and unforgeable: the OS reports the real connecting PID, and
* herdr owns the PID→pane mapping. A worker cannot claim to be another worker, nor the primary.
* Single-host only (the herd shares the {@code bridged} host); the token path is the split-host
* fallback.
*/
public final class ConnectionIdentity {
private final PaneLocator panes;
private final PeerPidLookup pids;
private final ProcessCwdLookup cwds;
/** Identity only (no cwd resolution — {@link #cwdForPid} returns {@code null}). */
public ConnectionIdentity(PaneLocator panes, PeerPidLookup pids) {
this(panes, pids, _ -> null);
}
/** Identity plus cwd resolution (CB-112 — inherit the primary's directory on spawn). */
public ConnectionIdentity(PaneLocator panes, PeerPidLookup pids, ProcessCwdLookup cwds) {
this.panes = panes;
this.pids = pids;
this.cwds = cwds;
}
/**
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
* primary / an off-host client) and its {@code pid} (or {@code -1} if not resolvable).
*/
public record Caller(String terminal, long pid) {
}
/** Resolve the caller's terminal and PID from one peer-PID lookup. */
public Caller resolve(String remoteAddr, int remotePort) {
if (!isLoopback(remoteAddr)) {
return new Caller(null, -1); // only same-host callers can be workers
}
long pid = pids.pidForLocalPort(remotePort);
return new Caller(panes.terminalForPid(pid), pid);
}
/**
* The calling worker's {@code terminal_id}, or {@code null} if the caller is not a known
* on-host worker (treat as the primary).
*/
public String callerTerminal(String remoteAddr, int remotePort) {
return resolve(remoteAddr, remotePort).terminal();
}
/** The working directory of {@code pid} (the primary's cwd on an MCP spawn), or {@code null}. */
public String cwdForPid(long pid) {
return pid > 0 ? cwds.cwdForPid(pid) : null;
}
private static boolean isLoopback(String addr) {
return "127.0.0.1".equals(addr) || "::1".equals(addr) || "0:0:0:0:0:0:0:1".equals(addr);
}
}
@@ -0,0 +1,58 @@
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.BufferedReader;
import java.io.InputStreamReader;
import java.nio.charset.StandardCharsets;
import java.util.concurrent.TimeUnit;
/**
* {@link PeerPidLookup} via {@code lsof} (present on macOS and Linux). For a loopback TCP source
* {@code port}, both the client and this daemon appear on that port — so we exclude our own PID
* and take the other end, which is the calling process.
*/
public final class LsofPeerPidLookup implements PeerPidLookup {
private static final Logger log = LoggerFactory.getLogger(LsofPeerPidLookup.class);
private final long selfPid = ProcessHandle.current().pid();
@Override
public long pidForLocalPort(int port) {
try {
Process p = new ProcessBuilder("lsof", "-nP", "-FpP", "-iTCP:" + port)
.redirectErrorStream(true).start();
long found = -1;
try (BufferedReader r = new BufferedReader(
new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
long current = -1;
String line;
// -Fp emits records: a 'p<pid>' line, then the ports/files under that pid.
while ((line = r.readLine()) != null) {
if (line.startsWith("p")) {
current = parse(line.substring(1));
} else if (current > 0 && current != selfPid) {
found = current; // first process on this port that isn't us = the client
}
}
}
if (!p.waitFor(2, TimeUnit.SECONDS)) {
p.destroyForcibly();
}
return found;
} catch (Exception e) {
log.debug("lsof peer-pid lookup for port {} failed: {}", port, e.getMessage());
return -1;
}
}
private static long parse(String s) {
try {
return Long.parseLong(s.trim());
} catch (NumberFormatException e) {
return -1;
}
}
}
@@ -0,0 +1,48 @@
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.BufferedReader;
import java.io.InputStreamReader;
import java.nio.charset.StandardCharsets;
import java.util.concurrent.TimeUnit;
/**
* {@link ProcessCwdLookup} via {@code lsof} (present on macOS and Linux): {@code lsof -a -p <pid>
* -d cwd -Fn} prints the process's cwd on the {@code n…} line. Used to inherit the primary's
* working directory for a spawned worker (CB-112).
*/
public final class LsofProcessCwdLookup implements ProcessCwdLookup {
private static final Logger log = LoggerFactory.getLogger(LsofProcessCwdLookup.class);
@Override
public String cwdForPid(long pid) {
if (pid <= 0) {
return null;
}
try {
Process p = new ProcessBuilder("lsof", "-a", "-p", Long.toString(pid), "-d", "cwd", "-Fn")
.redirectErrorStream(true).start();
String cwd = null;
try (BufferedReader r = new BufferedReader(
new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
String line;
while ((line = r.readLine()) != null) {
if (line.startsWith("n")) { // 'n<path>' is the file-name field for the cwd fd
cwd = line.substring(1);
break;
}
}
}
if (!p.waitFor(2, TimeUnit.SECONDS)) {
p.destroyForcibly();
}
return (cwd == null || cwd.isBlank()) ? null : cwd;
} catch (Exception e) {
log.debug("lsof cwd lookup for pid {} failed: {}", pid, e.getMessage());
return null;
}
}
}
@@ -0,0 +1,13 @@
package dev.ltms.bridged.mcp;
/**
* Resolves the OS PID that owns a loopback TCP source port — the OS half of connection-based MCP
* identity. Java exposes no peer PID for a TCP socket, so this shells out. Injectable so
* {@link ConnectionIdentity} is testable without a real connection.
*/
@FunctionalInterface
public interface PeerPidLookup {
/** The PID whose socket has local (source) {@code port} on loopback, or {@code -1} if unknown. */
long pidForLocalPort(int port);
}
@@ -0,0 +1,68 @@
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Optional;
import java.util.concurrent.atomic.AtomicReference;
/**
* Single-slot, thread-safe registry for the primary's herdr {@code terminal_id}.
*
* <p>Populated from the caller terminal of orchestration-side MCP tools
* ({@code bridge_send}, {@code bridge_spawn}) — tools that only the primary calls.
* A pinned terminal (from config) seeds the registry at construction and makes
* subsequent {@link #record(String)} calls no-ops.
*
* <p>The push loop ({@code ReplyPushLoop}) uses {@link #isKnown()} to decide
* whether active nudging is possible; an empty registry means the primary is
* off-host or non-herdr and delivery falls back to pull.
*/
public final class PrimaryRegistry {
private static final Logger log = LoggerFactory.getLogger(PrimaryRegistry.class);
private final AtomicReference<String> terminal = new AtomicReference<>();
private final boolean pinned;
/**
* @param pinnedTerminal an optional pinned terminal from config ({@code null}/blank = unpinned)
*/
public PrimaryRegistry(String pinnedTerminal) {
if (pinnedTerminal != null && !pinnedTerminal.isBlank()) {
this.terminal.set(pinnedTerminal);
this.pinned = true;
log.info("primary terminal pinned: {}", pinnedTerminal);
} else {
this.pinned = false;
}
}
/**
* Record a terminal_id. No-op when:
* <ul>
* <li>the registry is pinned (config override),
* <li>{@code terminalId} is {@code null} or blank (non-herdr caller).
* </ul>
*/
public void record(String terminalId) {
if (pinned) return;
if (terminalId == null || terminalId.isBlank()) return;
String prev = terminal.getAndSet(terminalId);
if (prev == null) {
log.debug("primary terminal learned: {}", terminalId);
} else if (!prev.equals(terminalId)) {
log.debug("primary terminal changed: {} -> {}", prev, terminalId);
}
}
/** The known primary terminal, or empty if not yet learned (and not pinned). */
public Optional<String> primaryTerminal() {
return Optional.ofNullable(terminal.get());
}
/** {@code true} once a terminal has been recorded (or was pinned at construction). */
public boolean isKnown() {
return terminal.get() != null;
}
}
@@ -0,0 +1,13 @@
package dev.ltms.bridged.mcp;
/**
* Resolves a process's current working directory from its PID — the OS half of CB-112's
* "a worker inherits the primary's directory." Injectable so {@link ConnectionIdentity} stays
* testable without shelling out.
*/
@FunctionalInterface
public interface ProcessCwdLookup {
/** The working directory of {@code pid}, or {@code null} if unknown. */
String cwdForPid(long pid);
}
@@ -0,0 +1,96 @@
package dev.ltms.bridged.metrics;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import java.util.LinkedHashMap;
import java.util.Map;
/**
* The daemon's metric definitions (CB-502) — one place where every series is named, described, and
* (for gauges) bound to live state.
*
* <p>The set is deliberately small: each series maps to a failure mode this project has actually
* hit, not to whatever was easy to count. The two worth watching in practice are
* {@code bridged_sends_total{outcome="completion_fallback"}} — a rising share means turn detection
* is degrading, the CB-115/116/118 failure family — and
* {@code bridged_push_nudges_total{outcome="exhausted"}}, which means the primary stopped draining
* its inbox and CB-307's active push gave up.
*/
public final class BridgedMetrics {
/** Counter: delegated sends by terminal outcome. */
public static final String SENDS = "bridged_sends_total";
/** Counter: worker replies by the path that carried them (rendezvous vs stranded-to-inbox). */
public static final String REPLIES = "bridged_replies_total";
/** Counter: push-loop nudges to the primary, by outcome. */
public static final String PUSH_NUDGES = "bridged_push_nudges_total";
/** Counter: spawn attempts by peer kind and outcome. */
public static final String SPAWNS = "bridged_spawns_total";
/** Counter: herdr socket calls by method and outcome. */
public static final String HERDR_CALLS = "bridged_herdr_calls_total";
/** Counter: rejected requests by reason (CB-501). */
public static final String AUTH_FAILURES = "bridged_auth_failures_total";
/** Gauge: session census by lifecycle state. */
public static final String SESSIONS = "bridged_sessions";
/** Gauge: undrained replies held per target. */
public static final String INBOX_DEPTH = "bridged_inbox_depth";
private BridgedMetrics() {
}
/**
* Build the registry with its help text and live gauges bound.
*
* @param sessions the authoritative session registry (census gauge)
* @param inbox the reply inbox; only used for a depth gauge when it can be inspected
*/
public static Metrics create(SessionManager sessions, ReplyInbox inbox) {
Metrics m = new Metrics();
m.describe(SENDS, "counter",
"Delegated sends by terminal outcome (replied|completion_fallback|timeout|failed).");
m.describe(REPLIES, "counter",
"Worker replies by delivery path (rendezvous=resolved an open send, inbox=stranded and held).");
m.describe(PUSH_NUDGES, "counter",
"CB-307 push-loop nudges to the primary (delivered|exhausted).");
m.describe(SPAWNS, "counter",
"Worker spawn attempts by peer kind and outcome (ready|timeout|guard_rejected).");
m.describe(HERDR_CALLS, "counter",
"herdr socket calls by method and outcome — the dependency everything else rests on.");
m.describe(AUTH_FAILURES, "counter",
"Requests refused by CB-501/505 (unauthenticated|forbidden).");
m.describe(SESSIONS, "gauge",
"Registered worker sessions by lifecycle state.");
m.describe(INBOX_DEPTH, "gauge",
"Replies held for a target that the primary has not drained. Steady state is 0; "
+ "a target stuck above 0 means CB-307 delivery is not completing.");
// One gauge per state so a scrape shows the whole census even when a state is empty —
// an absent series and a zero series read very differently on a dashboard.
for (WorkerSession.State state : WorkerSession.State.values()) {
String label = state.name().toLowerCase();
m.gauge(SESSIONS, () -> countIn(sessions, state), "state", label);
}
// Depth is per live session, so the label set is only known at scrape time. peek() is the
// port's non-destructive read — scraping metrics must never ack a reply out of the inbox.
m.collector(INBOX_DEPTH, "target", () -> {
Map<String, Number> depths = new LinkedHashMap<>();
for (WorkerSession s : sessions.roster()) {
String target = s.terminalId();
if (target == null) {
continue;
}
depths.put(target, inbox.peek(target).size());
}
return depths;
});
return m;
}
private static long countIn(SessionManager sessions, WorkerSession.State state) {
return sessions.roster().stream().filter(s -> s.state() == state).count();
}
}
@@ -0,0 +1,181 @@
package dev.ltms.bridged.metrics;
import java.util.Map;
import java.util.NavigableMap;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ConcurrentSkipListMap;
import java.util.concurrent.atomic.LongAdder;
import java.util.function.Supplier;
/**
* The daemon's metric registry and Prometheus text renderer (CB-502).
*
* <p>Deliberately dependency-free. The roadmap's tech-stack table specified Micrometer, but this
* pom already carries an unusual reconciliation burden (a hand-pinned {@code jackson-annotations}
* to make the MCP SDK's Jackson 3 coexist with our Jackson 2, a Jetty BOM import to stop version
* skew, and four documented accepted-CVE advisories), and the dependency CVE gate this project
* mandates could not be run when this landed. The metric set is small and fully known, and
* Prometheus text exposition is a stable, well-specified format — so the registry is ~100 lines
* here instead of a new transitive tree. {@code GET /metrics} is the swap seam if Micrometer's
* ecosystem is ever wanted.
*
* <p>Thread-safe: counters are {@link LongAdder} (built for contended increment), gauges are
* supplier-backed so they read live state at scrape time rather than needing to be pushed.
*/
public final class Metrics {
/** Counter series, keyed by the fully-rendered {@code name{labels}} sample id. */
private final NavigableMap<String, LongAdder> counters = new ConcurrentSkipListMap<>();
/** Gauge series, evaluated at scrape time. */
private final NavigableMap<String, Supplier<Number>> gauges = new ConcurrentSkipListMap<>();
/** Gauge families whose label set is only known at scrape time, keyed by metric name. */
private final NavigableMap<String, Collector> collectors = new ConcurrentSkipListMap<>();
/** HELP/TYPE metadata, keyed by bare metric name. */
private final Map<String, String[]> meta = new ConcurrentHashMap<>();
/** A gauge family whose series are discovered per scrape (one label, many values). */
private record Collector(String labelName, Supplier<Map<String, Number>> samples) {
}
/** Declare a metric's help text and type once, so the exposition carries HELP/TYPE lines. */
public Metrics describe(String name, String type, String help) {
meta.put(name, new String[]{type, help});
return this;
}
/** Increment a counter by one. */
public void inc(String name, String... labelPairs) {
add(name, 1, labelPairs);
}
/** Increment a counter by {@code delta}. */
public void add(String name, long delta, String... labelPairs) {
counters.computeIfAbsent(sample(name, labelPairs), _ -> new LongAdder()).add(delta);
}
/**
* Register a live gauge. The supplier is called at scrape time, so it reflects current state
* (session census, inbox depth) without anything having to remember to update it.
*/
public void gauge(String name, Supplier<Number> value, String... labelPairs) {
gauges.put(sample(name, labelPairs), value);
}
/**
* Register a gauge family whose label values are not known up front — inbox depth per target,
* for instance, where the set of targets changes as workers come and go. The supplier returns
* {@code labelValue → value} and is evaluated once per scrape.
*/
public void collector(String name, String labelName, Supplier<Map<String, Number>> samples) {
collectors.put(name, new Collector(labelName, samples));
}
/** Current value of a counter series — for assertions in tests. */
public long count(String name, String... labelPairs) {
LongAdder a = counters.get(sample(name, labelPairs));
return a == null ? 0 : a.sum();
}
/**
* Render the Prometheus text exposition format (version 0.0.4): optional {@code # HELP} and
* {@code # TYPE} lines per metric family, then one line per sample.
*/
public String render() {
StringBuilder out = new StringBuilder(1024);
String lastFamily = null;
for (Map.Entry<String, LongAdder> e : counters.entrySet()) {
lastFamily = emitHeader(out, e.getKey(), lastFamily);
out.append(e.getKey()).append(' ').append(e.getValue().sum()).append('\n');
}
for (Map.Entry<String, Supplier<Number>> e : gauges.entrySet()) {
lastFamily = emitHeader(out, e.getKey(), lastFamily);
Number v;
try {
v = e.getValue().get();
} catch (RuntimeException ex) {
continue; // a broken gauge must never break the whole scrape
}
if (v == null) {
continue;
}
out.append(e.getKey()).append(' ').append(format(v)).append('\n');
}
for (Map.Entry<String, Collector> e : collectors.entrySet()) {
Map<String, Number> samples;
try {
samples = e.getValue().samples().get();
} catch (RuntimeException ex) {
continue; // a broken collector must never break the whole scrape
}
if (samples == null || samples.isEmpty()) {
continue;
}
lastFamily = emitHeader(out, e.getKey(), lastFamily);
// Sort so repeated scrapes are byte-stable and diffable.
new java.util.TreeMap<>(samples).forEach((label, v) -> {
if (v != null) {
out.append(sample(e.getKey(), e.getValue().labelName(), label))
.append(' ').append(format(v)).append('\n');
}
});
}
return out.toString();
}
/** Emit HELP/TYPE when the sample starts a new metric family; returns the current family. */
private String emitHeader(StringBuilder out, String sampleId, String lastFamily) {
String family = familyOf(sampleId);
if (family.equals(lastFamily)) {
return lastFamily;
}
String[] m = meta.get(family);
if (m != null) {
out.append("# HELP ").append(family).append(' ').append(m[1]).append('\n');
out.append("# TYPE ").append(family).append(' ').append(m[0]).append('\n');
}
return family;
}
private static String familyOf(String sampleId) {
int brace = sampleId.indexOf('{');
return brace < 0 ? sampleId : sampleId.substring(0, brace);
}
/** Whole numbers render without a decimal point; everything else as-is. */
private static String format(Number v) {
double d = v.doubleValue();
return (d == Math.rint(d) && !Double.isInfinite(d))
? Long.toString((long) d)
: Double.toString(d);
}
/** Build the {@code name{k="v",k2="v2"}} sample id; labels are sorted for stable output. */
private static String sample(String name, String... labelPairs) {
if (labelPairs == null || labelPairs.length == 0) {
return name;
}
if (labelPairs.length % 2 != 0) {
throw new IllegalArgumentException("labels must be key/value pairs, got " + labelPairs.length);
}
NavigableMap<String, String> sorted = new java.util.TreeMap<>();
for (int i = 0; i < labelPairs.length; i += 2) {
sorted.put(labelPairs[i], labelPairs[i + 1] == null ? "" : labelPairs[i + 1]);
}
StringBuilder sb = new StringBuilder(name.length() + 16 * sorted.size());
sb.append(name).append('{');
boolean first = true;
for (Map.Entry<String, String> e : sorted.entrySet()) {
if (!first) {
sb.append(',');
}
first = false;
sb.append(e.getKey()).append("=\"").append(escapeLabel(e.getValue())).append('"');
}
return sb.append('}').toString();
}
/** Label values are escaped per the exposition format: backslash, quote, newline. */
private static String escapeLabel(String v) {
return v.replace("\\", "\\\\").replace("\"", "\\\"").replace("\n", "\\n");
}
}
@@ -0,0 +1,233 @@
package dev.ltms.bridged.msg;
import com.rabbitmq.client.AMQP;
import com.rabbitmq.client.Channel;
import com.rabbitmq.client.Connection;
import com.rabbitmq.client.ConnectionFactory;
import com.rabbitmq.client.DeliverCallback;
import com.rabbitmq.client.Recoverable;
import com.rabbitmq.client.RecoveryListener;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
/**
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
* port {@link InMemoryReplyInbox} implements as soft state.
*
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target owns a durable
* queue {@code agent.<target>.inbox}. A manual-ack consumer pulls persistent messages off that queue
* into an in-memory <em>held</em> map (keyed by {@code msgId}) but does <em>not</em> ack them.
* {@link #peek} returns that snapshot; {@link #ack} acks the broker delivery-tag and drops the entry.
* Because messages stay unacked until the primary actually drains them, a crash (or a {@code java -jar}
* bounce) before caller-ack leaves them on the broker — it redelivers on reconnect. That is the
* durability the in-memory adapter cannot give, with the port contract preserved.
*
* <p><strong>Dedup.</strong> The consumer keys the held map by {@code msgId}; a redelivered duplicate
* (at-least-once, or a producer double-publish) is acked-and-dropped on arrival, so it never
* double-queues. {@link #publish} additionally short-circuits an already-held {@code msgId} — a
* fast path; the consumer-side check is the real guarantee.
*
* <p><strong>Visibility.</strong> Unlike the in-memory adapter, publish → broker → consumer is
* asynchronous, so a {@link #peek} immediately after {@link #publish} may not yet see the message
* (broker delivery latency). Callers that need the reply drained poll (as the primary already does);
* the contract test waits for visibility. This is inherent to broker-backed delivery, not a defect.
*
* <p>The default deploy targets LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1 (URI-only swap),
* so the {@code @Tag("contract")} integration test runs against a RabbitMQ container.
*/
public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
private static final Logger log = LoggerFactory.getLogger(AmqpReplyInbox.class);
private static final String QUEUE_PREFIX = "agent.";
private static final String QUEUE_SUFFIX = ".inbox";
private final Connection connection;
private final Channel channel;
/** All channel operations (publish/declare/ack) serialize on this — a Channel is not thread-safe. */
private final Object channelLock = new Object();
/** target → (msgId → held delivery). Per-target map is guarded by synchronizing on itself. */
private final ConcurrentHashMap<String, LinkedHashMap<String, Held>> held = new ConcurrentHashMap<>();
/** Targets whose queue is declared and consumer is running. */
private final Set<String> consuming = ConcurrentHashMap.newKeySet();
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
private record Held(long deliveryTag, InboxMessage message) {}
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) and open the inbox. */
public static AmqpReplyInbox open(String uri) {
try {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"));
} catch (Exception e) {
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
}
}
/** Wrap an already-open connection (injection seam for the contract test). */
AmqpReplyInbox(Connection connection) {
this.connection = connection;
try {
this.channel = connection.createChannel();
} catch (IOException e) {
throw new IllegalStateException("cannot open AMQP channel", e);
}
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
// repopulates it with valid tags (dedup by msgId still prevents any double-queue).
if (connection instanceof Recoverable recoverable) {
recoverable.addRecoveryListener(new RecoveryListener() {
@Override
public void handleRecovery(Recoverable recoverable) {
held.clear();
log.info("AMQP connection recovered; cleared held replies for fresh redelivery");
}
@Override
public void handleRecoveryStarted(Recoverable recoverable) {
// no-op: we act once recovery completes
}
});
}
}
@Override
public void publish(String target, String msgId, String content) {
ensureConsuming(target);
var perTarget = held.get(target);
if (perTarget != null) {
synchronized (perTarget) {
if (perTarget.containsKey(msgId)) {
return; // already held — producer-side fast dedup
}
}
}
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
.messageId(msgId)
.deliveryMode(2) // persistent — survives a broker restart
.contentType("text/plain")
.build();
try {
synchronized (channelLock) {
channel.basicPublish("", queueName(target), props, content.getBytes(StandardCharsets.UTF_8));
}
} catch (IOException e) {
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
}
}
@Override
public List<InboxMessage> peek(String target) {
ensureConsuming(target);
var perTarget = held.get(target);
if (perTarget == null) {
return List.of();
}
synchronized (perTarget) {
return perTarget.values().stream().map(Held::message).toList();
}
}
@Override
public void ack(String target, String msgId) {
var perTarget = held.get(target);
if (perTarget == null) {
return;
}
Held h;
synchronized (perTarget) {
h = perTarget.remove(msgId);
}
if (h == null) {
return; // never held (or already acked) — no-op
}
try {
synchronized (channelLock) {
channel.basicAck(h.deliveryTag(), false);
}
} catch (IOException e) {
// Ack didn't reach the broker: restore the entry so a later ack (or a redelivery after
// reconnect) can retry. Keeps the at-least-once contract — a reply is never silently lost.
synchronized (perTarget) {
perTarget.putIfAbsent(msgId, h);
}
throw new IllegalStateException("cannot ack reply " + msgId + " on " + queueName(target), e);
}
}
/** Declare the durable per-target queue and start its manual-ack consumer, once per target. */
private void ensureConsuming(String target) {
if (consuming.contains(target)) {
return;
}
synchronized (channelLock) {
if (!consuming.add(target)) {
return; // another thread just set it up
}
String queue = queueName(target);
try {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
} catch (IOException e) {
consuming.remove(target);
throw new IllegalStateException("cannot consume queue " + queue, e);
}
}
}
private DeliverCallback deliverCallback(String target) {
return (_, delivery) -> {
String msgId = delivery.getProperties().getMessageId();
long tag = delivery.getEnvelope().getDeliveryTag();
if (msgId == null || msgId.isBlank()) {
msgId = Long.toHexString(tag); // synthesize an id so dedup still has a key
}
String content = new String(delivery.getBody(), StandardCharsets.UTF_8);
var perTarget = held.computeIfAbsent(target, _ -> new LinkedHashMap<>());
boolean duplicate;
synchronized (perTarget) {
if (perTarget.containsKey(msgId)) {
duplicate = true;
} else {
perTarget.put(msgId, new Held(tag, new InboxMessage(msgId, target, content)));
duplicate = false;
}
}
if (duplicate) {
// Redelivered duplicate: ack the new tag and drop it so the broker stops resending.
synchronized (channelLock) {
channel.basicAck(tag, false);
}
}
};
}
private static String queueName(String target) {
return QUEUE_PREFIX + target + QUEUE_SUFFIX;
}
@Override
public void close() {
try {
channel.close();
} catch (Exception e) {
log.debug("AMQP channel close: {}", e.toString());
}
try {
connection.close();
} catch (Exception e) {
log.debug("AMQP connection close: {}", e.toString());
}
}
}
@@ -0,0 +1,50 @@
package dev.ltms.bridged.msg;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.concurrent.ConcurrentHashMap;
/**
* Soft-state {@link ReplyInbox} backed by a {@link ConcurrentHashMap} keyed by target session.
* Per-target FIFO ordering (insertion order via {@link LinkedHashMap}). Dedup by {@code msgId}
* within a target. Thread-safe for concurrent publish vs. drain.
*
* <p><strong>This is soft-state, NOT persistence.</strong> Lost on a {@code java -jar} bounce — that
* is correct and consistent with "bridged stays soft-state." The Stage-2 AMQP adapter replaces this.
*/
public final class InMemoryReplyInbox implements ReplyInbox {
private final ConcurrentHashMap<String, LinkedHashMap<String, InboxMessage>> store = new ConcurrentHashMap<>();
@Override
public void publish(String target, String msgId, String content) {
var perTarget = store.computeIfAbsent(target, _ -> new LinkedHashMap<>());
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
perTarget.putIfAbsent(msgId, new InboxMessage(msgId, target, content));
}
}
@Override
public List<InboxMessage> peek(String target) {
var perTarget = store.get(target);
if (perTarget == null) {
return List.of();
}
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
return List.copyOf(perTarget.values());
}
}
@Override
public void ack(String target, String msgId) {
var perTarget = store.get(target);
if (perTarget != null) {
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
perTarget.remove(msgId);
}
}
}
}
@@ -0,0 +1,527 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.UUID;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.CompletionException;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
import java.util.concurrent.atomic.AtomicLong;
import java.util.concurrent.locks.ReentrantLock;
/**
* The blocking delegation feature (CB-104): deliver {@code content} into a worker and block until
* the worker returns a <em>structured reply</em> via {@code bridge_reply} (the {@link Rendezvous}),
* then hand that reply back. Delivery is the {@link Injector}'s job (the background poller sends it
* when the worker is injectable); this service never drives the injector or scrapes the terminal —
* completion is the worker's explicit reply, not a guess about {@code agent_status}.
*
* <p>Sends are serialized per session so exactly one reply can be outstanding per worker, which is
* what lets a reply map unambiguously to its send (no cross-talk between concurrent callers).
*
* <p>If the worker never replies within the timeout, the caller gets a typed "still working" /
* "queued" outcome — the message may still be mid-flight. A finished-but-unreplied turn is caught
* by the CB-106 completion fallback (see {@link Rendezvous#resolveCompletion}).
*
* <p><strong>Async fire-and-poll (CB-107).</strong> A caller's MCP client caps a blocking call at
* ~60s, but a real delegated task runs for minutes. {@link #sendAsync} therefore runs the same
* blocking {@link #send} on a background virtual thread and hands back a <em>ticket</em> the caller
* polls with {@link #poll}. The blocking and async paths share one code path (and the same per-target
* serialization), so async inherits the reply + completion resolution behaviour for free.
*/
public final class MessageService {
private static final Logger log = LoggerFactory.getLogger(MessageService.class);
/**
* The window a fire-and-poll send waits for resolution — generous, since no caller is blocked on
* it; a real delegated task resolves (reply or completion) well within this, and only a genuinely
* hung worker rides it out.
*/
private static final long ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L;
/** How long a finished (terminal) ticket is retained for polling before it is pruned. */
private static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
/** Outcome of a blocking send. */
public enum Outcome {
/** The worker called {@code bridge_reply}; {@code text} holds the structured answer. */
REPLIED,
/**
* The worker's delegated turn finished without a {@code bridge_reply} (CB-106 fallback);
* {@code text} is the scraped transcript tail rather than a structured answer.
*/
COMPLETED_UNREPLIED,
/**
* The worker ran the turn then wedged in an unrecoverable state (CB-109); {@code text} is the
* failure context (e.g. the error screen). Terminal, but not a successful completion.
*/
WORKER_FAILED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with
* {@link #answer(String, String, long)} and the turn resumes.
*/
QUESTION,
/** Timed out after the message was delivered — the worker is still working. */
TIMED_OUT_WORKING,
/** Timed out before delivery — the message is still queued for the worker. */
TIMED_OUT_QUEUED,
/** Another send to this session was in flight for the whole window. */
BUSY,
/**
* An answer ({@link #answer(String, String, long)}) referenced a {@code turnId} that is no
* longer open — the worker's {@code bridge_ask} already timed out or was answered.
*/
STALE_TURN
}
/**
* @param outcome how the send ended (or paused)
* @param text the worker's answer when {@link #completed()} (a structured {@code bridge_reply}
* for {@link Outcome#REPLIED}, a scraped transcript tail for
* {@link Outcome#COMPLETED_UNREPLIED}), or the question for {@link Outcome#QUESTION},
* else {@code null}
* @param turnId correlation id for a {@link Outcome#QUESTION} (answered via
* {@link #answer(String, String, long)}), else {@code null}
*/
public record Reply(Outcome outcome, String text, String turnId) {
/** A reply with no correlation id (the common terminal outcomes). */
public Reply(Outcome outcome, String text) {
this(outcome, text, null);
}
/** Whether the worker's turn actually finished with an answer (replied or scraped). */
public boolean completed() {
return outcome == Outcome.REPLIED || outcome == Outcome.COMPLETED_UNREPLIED;
}
}
/** How a worker's {@code bridge_ask} (CB-205) resolved. */
public enum AskOutcome {
/** The primary answered; {@link AskResult#answer} carries it. */
ANSWERED,
/** No delegation was open to surface the question to — the worker has no one to ask. */
NO_WAITER,
/** The primary did not answer within the window. */
TIMED_OUT
}
/** The outcome of a worker's {@code bridge_ask}: how it resolved and (if answered) the answer. */
public record AskResult(AskOutcome outcome, String answer) {
}
/** Lifecycle phase of an async delegation ticket. */
public enum Phase {
/** Delegated and in flight — queued for the worker or being worked. */
PENDING,
/** The worker's turn finished; {@link TaskView#reply} holds the answer. */
DONE,
/** The delegation could not complete (timed out, worker gone, or busy). */
FAILED
}
/**
* A poll snapshot of an async delegation.
*
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, else {@code null}
* @param replySource {@code "reply"} (structured {@code bridge_reply}) or {@code "transcript"}
* (completion scrape) when {@link Phase#DONE}, else {@code null}
* @param detail a human note (live worker status while pending, or the failure reason)
*/
public record TaskView(String ticket, Phase phase, String reply, String replySource, String detail) {
}
/** An in-flight or finished async delegation, keyed by its ticket. */
private record Task(String target, CompletableFuture<Reply> future, long createdNanos) {
}
private final AgentControl agents;
private final Injector injector;
private final Rendezvous rendezvous;
private final ReplyInbox inbox;
private final ReplyPushLoop pushLoop;
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
Thread.ofVirtual().name("bridge-async-", 0).factory());
/**
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
*
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
* branch ({@link #reply}) so it can nudge the primary to drain the inbox
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop) {
this(agents, injector, rendezvous, inbox, pushLoop, null);
}
/**
* As above, with a metric registry (CB-502). Instrumenting here rather than at the REST and MCP
* edges means both surfaces are counted by one piece of code and cannot drift.
*
* @param metrics nullable — when null, nothing is recorded
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop, Metrics metrics) {
this.agents = agents;
this.injector = injector;
this.rendezvous = rendezvous;
this.inbox = inbox;
this.pushLoop = pushLoop;
this.metrics = metrics;
}
/** Create with an explicit {@link ReplyInbox} and no push loop. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox) {
this(agents, injector, rendezvous, inbox, null);
}
/** Backward-compatible constructor that uses a default {@link InMemoryReplyInbox}. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous) {
this(agents, injector, rendezvous, new InMemoryReplyInbox());
}
/** Current lifecycle status of a worker (the {@code GET /sessions/{id}/status} surface). */
public AgentStatus status(String target) {
return agents.status(target);
}
/**
* Route a worker's explicit {@code bridge_reply}: resolve an open send, or queue it in the
* inbox if no send is currently open. Unlike the bare {@link Rendezvous#resolve}, a no-waiter
* result is <em>not</em> a failure — the reply is held for later drain.
*
* <p><strong>Do NOT use this for mid-turn questions.</strong> {@code bridge_ask} /
* {@link Rendezvous#resolveQuestion} must keep today's {@code NO_WAITER} behaviour — questions
* are interactive and must never be queued.
*
* @return always {@code true} — the reply either resolved a live send or was queued
*/
public boolean reply(String session, String content) {
if (rendezvous.resolve(session, content)) {
count(BridgedMetrics.REPLIES, "path", "rendezvous");
return true; // a live send took it — unchanged fast path
}
inbox.publish(session, UUID.randomUUID().toString(), content);
// A rising inbox share is the signal CB-307 exists to make visible: the worker finished but
// nobody was waiting, so delivery now depends on the push loop and a drain.
count(BridgedMetrics.REPLIES, "path", "inbox");
if (pushLoop != null) {
pushLoop.onReplyQueued(session);
}
return true; // held, not lost
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
if (metrics != null) {
metrics.inc(name, labels);
}
}
/** Count a send's terminal outcome and pass the reply through unchanged. */
private Reply recorded(Reply r) {
String label = sendOutcomeLabel(r.outcome());
if (label != null) {
count(BridgedMetrics.SENDS, "outcome", label);
}
return r;
}
/** Map a terminal send outcome to its metric label, or {@code null} for non-terminal ones. */
private static String sendOutcomeLabel(Outcome o) {
return switch (o) {
case REPLIED -> "replied";
case COMPLETED_UNREPLIED -> "completion_fallback";
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> "timeout";
case WORKER_FAILED -> "failed";
case STALE_TURN, QUESTION -> null; // not a completed delegation
};
}
/**
* Abandon any send still waiting on {@code target} because its session has gone away (CB-516).
*
* <p>Without this, tearing a worker down left its rendezvous waiter open: a blocking
* {@code bridge_send} kept blocking, and an async one kept reporting {@code PENDING} until
* {@link #ASYNC_TIMEOUT_MS} — thirty minutes — even though the worker provably no longer
* existed and the delegation could never complete. Worse, {@code poll} already had the evidence
* (it calls {@code liveStatus} to build its detail string and gets back {@code "unknown"}) and
* reported {@code PENDING} anyway.
*
* <p>Resolving the waiter as a failure — rather than letting it time out — also means the
* outcome is counted, so a torn-down delegation stops being invisible to {@code /metrics}.
*
* @return true if a live waiter was failed
*/
public boolean abandon(String target, String reason) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
return false; // nobody is blocked on this worker — nothing to abandon
}
boolean failed = rendezvous.resolveFailure(waiter, reason);
if (failed) {
log.debug("abandoned send to {}: {}", target, reason);
}
return failed;
}
/**
* Acknowledge a specific reply by {@code msgId} for {@code target}. Removes it from the inbox
* so that a subsequent drain or peek no longer returns it.
*/
public void ackReply(String target, String msgId) {
inbox.ack(target, msgId);
}
/**
* Drain (peek + ack) all pending inbox replies for {@code target}. At-least-once: returns the
* messages and acknowledges them; an in-flight failure between returning and the caller
* processing them re-surfaces them on a subsequent drain (the ack is local).
*
* @return the drained messages, newest last (FIFO); empty list if none
*/
public List<ReplyInbox.InboxMessage> drainReplies(String target) {
var messages = inbox.peek(target);
for (var msg : messages) {
inbox.ack(target, msg.msgId());
}
return messages;
}
/**
* Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
*/
public Reply send(String target, String content, long timeoutMillis) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
}
try {
CompletableFuture<Void> delivered = injector.enqueue(target, content);
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
} finally {
rendezvous.close(target, reply);
}
} finally {
lock.unlock();
}
}
/**
* A worker's mid-turn question (CB-205 reverse rendezvous): surface {@code question} to the
* primary by resolving its open blocking {@code bridge_send}, then block this (worker) call until
* the primary answers via {@link #answer} or {@code timeoutMillis} elapses. Identity is the
* worker's own session — it does not address the primary.
*
* <p>Returns {@link AskOutcome#NO_WAITER} when no delegation is open to surface the question to
* (nothing to answer it), {@link AskOutcome#ANSWERED} with the primary's answer, or
* {@link AskOutcome#TIMED_OUT} if the primary stayed silent. The worker resumes its turn either
* way — an answered ask hands back the answer; an unanswered one leaves it to proceed alone.
*/
public AskResult ask(String workerSession, String question, long timeoutMillis) {
Rendezvous.AskTicket ticket = rendezvous.openAsk(workerSession);
// Only the freshly-opening caller surfaces the question; a coalesced duplicate simply blocks on
// the shared answer future that the fresh owner is already responsible for.
if (ticket.fresh()) {
// Register the reverse waiter first, then surface the question — so the answer, which can
// arrive the instant the primary reacts, always finds an open waiter to resolve.
if (!rendezvous.resolveQuestion(workerSession, question, ticket.turnId())) {
rendezvous.closeAsk(ticket.turnId());
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
}
}
try {
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
return new AskResult(AskOutcome.ANSWERED, answer);
} catch (TimeoutException e) {
log.debug("bridge_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
return new AskResult(AskOutcome.TIMED_OUT, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting the primary's answer for " + workerSession, e);
} finally {
// Only the fresh owner tears down the shared turn; a duplicate must leave it open.
if (ticket.fresh()) {
rendezvous.closeAsk(ticket.turnId());
}
}
}
/**
* The primary's answer to a worker's {@code bridge_ask} (CB-205): resolve the worker's blocked
* question identified by {@code turnId}, then — like a fresh {@link #send} — block for the worker's
* eventual {@code bridge_reply} as it finishes the resumed turn. The worker session is derived from
* {@code turnId}, never a caller argument.
*
* <p>Unlike {@link #send} this does not re-inject through the {@link Injector}: the worker is
* mid-turn (already picked up), so the answer flows back through its own open {@code bridge_ask}
* call, not a new status-gated delivery. The forward waiter is opened <em>before</em> the worker
* is unblocked so a reply that lands the instant it resumes is not lost.
*/
public Reply answer(String turnId, String content, long timeoutMillis) {
String workerSession = rendezvous.askSession(turnId);
if (workerSession == null) {
return new Reply(Outcome.STALE_TURN, null); // the ask lapsed (timed out or already answered)
}
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(workerSession, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null);
}
try {
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession);
if (!rendezvous.answerAsk(turnId, content)) {
rendezvous.close(workerSession, reply);
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
}
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
} catch (TimeoutException e) {
// The worker resumed but hasn't replied yet — no completion fallback arms an answered
// turn (it never re-entered the injector), so a silent worker rides out the window.
return new Reply(Outcome.TIMED_OUT_WORKING, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + workerSession, e);
} finally {
rendezvous.close(workerSession, reply);
}
} finally {
lock.unlock();
}
}
/**
* Fire-and-poll variant of {@link #send}: deliver {@code content} to {@code target} on a
* background virtual thread and return immediately with a ticket to {@link #poll}. This is how a
* long task is delegated without tripping the caller's MCP client call timeout.
*
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content) {
String ticket = "task-" + ticketSeq.incrementAndGet();
CompletableFuture<Reply> future =
CompletableFuture.supplyAsync(() -> send(target, content, ASYNC_TIMEOUT_MS), asyncExecutor);
tasks.put(ticket, new Task(target, future, System.nanoTime()));
pruneTerminalTickets();
log.debug("async send {} -> {}", ticket, target);
return ticket;
}
/**
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket;
* otherwise a {@link Phase#PENDING} view (with the live worker status as detail), a
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason.
*/
public TaskView poll(String ticket) {
Task task = tasks.get(ticket);
if (task == null) {
return null;
}
CompletableFuture<Reply> f = task.future();
if (!f.isDone()) {
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target()));
}
Reply r;
try {
r = f.getNow(null);
} catch (CompletionException | java.util.concurrent.CancellationException e) {
Throwable cause = (e instanceof CompletionException ce && ce.getCause() != null) ? ce.getCause() : e;
return new TaskView(ticket, Phase.FAILED, null, null, cause.getMessage());
}
if (r.completed()) {
String source = r.outcome() == Outcome.REPLIED ? "reply" : "transcript";
return new TaskView(ticket, Phase.DONE, r.text(), source, null);
}
// A wedged worker (CB-109) carries the error context as its reason; the timeout/busy
// outcomes carry none, so fall back to the outcome name.
String detail = r.outcome() == Outcome.WORKER_FAILED && r.text() != null
? r.text()
: "no reply — " + r.outcome().name().toLowerCase();
return new TaskView(ticket, Phase.FAILED, null, null, detail);
}
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
private String liveStatus(String target) {
try {
return agents.status(target).name().toLowerCase();
} catch (RuntimeException e) {
return "unknown";
}
}
/** Drop finished tickets older than the TTL so the registry cannot grow without bound. */
private void pruneTerminalTickets() {
long cutoff = System.nanoTime() - TICKET_TTL_NANOS;
tasks.values().removeIf(t -> t.future().isDone() && t.createdNanos() < cutoff);
}
/** Release the async executor. */
public void close() {
asyncExecutor.shutdown();
}
/** Map a rendezvous {@link Rendezvous.Kind} onto its send {@link Outcome} (shared by send/answer). */
private static Outcome outcomeOf(Rendezvous.Kind kind) {
return switch (kind) {
case REPLY -> Outcome.REPLIED;
case COMPLETION -> Outcome.COMPLETED_UNREPLIED;
case FAILED -> Outcome.WORKER_FAILED;
case QUESTION -> Outcome.QUESTION;
};
}
private static boolean tryLock(ReentrantLock lock, long millis) {
try {
return lock.tryLock(Math.max(0, millis), TimeUnit.MILLISECONDS);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting the session send lock", e);
}
}
private static long remainingMillis(long deadlineNanos) {
return (deadlineNanos - System.nanoTime()) / 1_000_000L;
}
}
@@ -0,0 +1,216 @@
package dev.ltms.bridged.msg;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicLong;
/**
* The reply rendezvous: where a blocking {@code bridge_send} awaits how the worker's delegated turn
* ends. The sending (primary) request thread {@link #open}s a waiter; it is resolved either by the
* worker's explicit {@code bridge_reply} ({@link #resolve}, arriving on a different thread via
* {@code POST /sessions/{id}/reply}) or — the CB-106 fallback — by the injector observing the
* worker's delegated turn return to idle without a reply ({@link #resolveCompletion}).
*
* <p>At most one waiter per session — {@link MessageService} serializes sends per session, so a
* resolution maps unambiguously to the one outstanding send and cannot be captured by another.
*
* <p><strong>Waiter identity (CB-116).</strong> The completion/failure fallbacks run asynchronously
* and can fire <em>after</em> the turn they belong to has already been resolved by an explicit reply
* and a <em>next</em> send has opened its own waiter on the same session. Resolving "whatever waiter
* is registered now" would then land turn N's stale scrape on turn N+1's send. So those fallbacks
* resolve a <em>specific</em> {@link CompletableFuture} captured when their turn was delivered
* ({@link #resolveCompletion(CompletableFuture, String)} /
* {@link #resolveFailure(CompletableFuture, String)}): a no-op if that waiter was already resolved,
* and it can never touch a later send's waiter.
*/
public final class Rendezvous {
/** How a delegated turn ended (or paused). */
public enum Kind {
/** The worker called {@code bridge_reply} with a structured answer. */
REPLY,
/** The worker's turn finished without a {@code bridge_reply}; {@code text} is a scrape. */
COMPLETION,
/** The worker ran the turn then wedged (CB-109); {@code text} is the failure context. */
FAILED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205 reverse rendezvous);
* {@code text} is the question and {@code turnId} correlates the primary's answer back to
* the worker's blocked {@code bridge_ask}. Not terminal — the turn resumes after the answer.
*/
QUESTION
}
/**
* The resolved outcome of a send: its {@link Kind}, the associated text, and — only for
* {@link Kind#QUESTION} — the {@code turnId} the primary answers with (else {@code null}).
*/
public record Resolution(Kind kind, String text, String turnId) {
/** A terminal resolution (reply / completion / failure) with no correlation id. */
public Resolution(Kind kind, String text) {
this(kind, text, null);
}
}
/** A worker's open mid-turn question: the worker session it belongs to and the answer future. */
private record AskWaiter(String session, CompletableFuture<String> answer) {
}
/**
* Handle to a reverse-rendezvous turn: the {@code turnId}, its answer future, and whether this
* call freshly opened it (versus coalescing onto an already-open ask).
*/
public record AskTicket(String turnId, CompletableFuture<String> answer, boolean fresh) {
}
private final ConcurrentHashMap<String, CompletableFuture<Resolution>> waiters = new ConcurrentHashMap<>();
/** Reverse rendezvous (CB-205): worker questions awaiting the primary's answer, keyed by {@code turnId}. */
private final ConcurrentHashMap<String, AskWaiter> asks = new ConcurrentHashMap<>();
private final AtomicLong askSeq = new AtomicLong();
/** Per-session index of the currently-open ask, so duplicate bridge_ask calls coalesce onto one turn. */
private final ConcurrentHashMap<String, String> openAsksBySession = new ConcurrentHashMap<>();
/**
* Register a waiter for {@code session} — the await side of the public {@code resolve*} methods.
* The caller must hold that session's send lock.
*/
public CompletableFuture<Resolution> open(String session) {
CompletableFuture<Resolution> waiter = new CompletableFuture<>();
waiters.put(session, waiter);
return waiter;
}
/** Remove {@code waiter} for {@code session} (only if it is still the registered one). */
void close(String session, CompletableFuture<Resolution> waiter) {
waiters.remove(session, waiter);
}
/** Whether a send is currently awaiting a resolution for {@code session}. */
public boolean isWaiting(String session) {
return waiters.containsKey(session);
}
/**
* The waiter currently registered for {@code session}, or {@code null} if none is waiting. The
* completion/failure fallbacks capture this at delivery time so they can later resolve that exact
* send (see the CB-116 note above) rather than whichever send happens to be waiting when they fire.
*/
public CompletableFuture<Resolution> currentWaiter(String session) {
return waiters.get(session);
}
/**
* Resolve the send awaiting on {@code session} with the worker's explicit reply {@code content}.
*
* @return {@code true} if a waiter was resolved; {@code false} if none was waiting (a late or
* spurious reply — e.g. the send already timed out)
*/
public boolean resolve(String session, String content) {
return complete(session, new Resolution(Kind.REPLY, content));
}
// --- reverse rendezvous (CB-205 bridge_ask) ------------------------------------------------
/**
* Open a reverse-rendezvous waiter for a worker's mid-turn question. If {@code session} already has
* an open ask, coalesce onto it (same {@code turnId}, same answer future). Otherwise atomically mint
* a fresh {@code turnId}, register it in both the per-turn and per-session indexes, and hand it back
* marked fresh. The caller then {@link #resolveQuestion surfaces the question} to the primary and
* blocks on the returned future until the primary {@link #answerAsk answers}.
*/
public AskTicket openAsk(String session) {
while (true) {
AskWaiter[] minted = { null };
String turnId = openAsksBySession.computeIfAbsent(session, _ -> {
String newTurnId = session + "#" + askSeq.incrementAndGet();
CompletableFuture<String> answer = new CompletableFuture<>();
AskWaiter waiter = new AskWaiter(session, answer);
asks.put(newTurnId, waiter);
minted[0] = waiter;
return newTurnId;
});
if (minted[0] != null) {
return new AskTicket(turnId, minted[0].answer(), true);
}
AskWaiter existing = asks.get(turnId);
if (existing != null) {
return new AskTicket(turnId, existing.answer(), false);
}
// A close raced and removed the waiter after we read the turnId; clear the stale index entry
// and retry so a fresh ask is always backed by a registered waiter.
openAsksBySession.remove(session, turnId);
}
}
/**
* Surface a worker's mid-turn {@code question} by resolving the primary's open {@code bridge_send}
* with a {@link Kind#QUESTION} carrying {@code turnId}. Same session-keyed semantics as
* {@link #resolve}: the one outstanding send for {@code session} unblocks with the question.
*
* @return {@code true} if a send was awaiting (the question reached the primary); {@code false}
* if none was (no delegation is open to answer it)
*/
public boolean resolveQuestion(String session, String question, String turnId) {
return complete(session, new Resolution(Kind.QUESTION, question, turnId));
}
/** The worker session an outstanding ask {@code turnId} belongs to, or {@code null} if unknown/lapsed. */
public String askSession(String turnId) {
AskWaiter w = asks.get(turnId);
return w == null ? null : w.session();
}
/**
* Resolve a worker's blocked {@code bridge_ask} with the primary's {@code answer}, unblocking it
* to resume its turn.
*
* @return {@code true} if the ask was still open and got the answer; {@code false} if the
* {@code turnId} is unknown or the ask already lapsed (timed out / was answered)
*/
public boolean answerAsk(String turnId, String answer) {
AskWaiter w = asks.get(turnId);
return w != null && w.answer().complete(answer);
}
/** Drop a reverse-rendezvous turn once its {@code bridge_ask} has resolved (answered or lapsed). */
public void closeAsk(String turnId) {
AskWaiter w = asks.get(turnId);
if (w == null) {
return;
}
// Remove the session index first and only if it still points to this turn, so a concurrent
// fresh ask cannot inherit a waiter we are about to drop.
openAsksBySession.remove(w.session(), turnId);
asks.remove(turnId);
}
/**
* Resolve a specific captured {@code waiter} as a completion (the delegated turn finished with no
* {@code bridge_reply}); {@code text} is the scraped transcript tail. The waiter is the one
* captured when this turn was delivered, so a late completion for turn N cannot land on turn N+1's
* send (CB-116). A no-op if that waiter was already resolved — a raced {@code bridge_reply} wins.
*
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
*/
public boolean resolveCompletion(CompletableFuture<Resolution> waiter, String text) {
return waiter != null && waiter.complete(new Resolution(Kind.COMPLETION, text));
}
/**
* Resolve a specific captured {@code waiter} as a failure — the worker ran the turn but wedged in
* an unrecoverable state (CB-109); {@code reason} is the failure context (e.g. the error screen).
* Like {@link #resolveCompletion(CompletableFuture, String)} it targets the exact captured send
* (CB-116). A no-op if that waiter was already resolved — first resolution wins.
*
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
*/
public boolean resolveFailure(CompletableFuture<Resolution> waiter, String reason) {
return waiter != null && waiter.complete(new Resolution(Kind.FAILED, reason));
}
private boolean complete(String session, Resolution resolution) {
CompletableFuture<Resolution> waiter = waiters.get(session);
return waiter != null && waiter.complete(resolution);
}
}
@@ -0,0 +1,31 @@
package dev.ltms.bridged.msg;
import java.util.List;
/**
* Holds terminal worker→primary replies that arrive with no live send to resolve, keyed by worker
* session (target), until the primary drains them. Soft-state in Stage 1 (in-memory, lost on restart);
* the Stage 2 AMQP adapter implements the same contract with cross-restart durability.
*
* <p><strong>This interface is the port.</strong> {@link InMemoryReplyInbox} is the Stage-1 adapter;
* an AMQP-backed adapter (Stage 2) must implement the same contract (idempotent publish, FIFO peek,
* at-least-once ack).
*/
public interface ReplyInbox {
/** A queued reply: an idempotency id, the worker session it came from, and the reply text. */
record InboxMessage(String msgId, String target, String content) {}
/**
* Queue {@code content} from worker {@code target} under {@code msgId}. Idempotent: publishing an
* already-present {@code msgId} for {@code target} is a no-op (dedup), so an at-least-once Stage-2
* redelivery cannot double-queue.
*/
void publish(String target, String msgId, String content);
/** Non-destructive snapshot of pending replies for {@code target} (FIFO), empty list if none. */
List<InboxMessage> peek(String target);
/** Remove the reply {@code msgId} for {@code target} once the primary has taken it. No-op if absent. */
void ack(String target, String msgId);
}
@@ -0,0 +1,176 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
/**
* Mechanism (b) of CB-307: a dedicated, status-gated push loop that nudges the primary's own
* herdr pane when a worker reply lands with no live {@code bridge_send} to resolve it.
*
* <p>The loop is triggered by {@link #onReplyQueued(String)} (called from
* {@link MessageService#reply} after the durable inbox publish). It checks four conditions
* at each tick via {@link #decide(String, int)}, then either injects a drain nudge,
* waits for the primary to become injectable, or stops reminding.
*
* <p>Bounded: at most {@link #maxReminders} nudges per target, with a configurable backoff
* between them. The reply is never lost — the durable inbox is the backstop.
*/
public final class ReplyPushLoop {
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
private final ReplyInbox inbox;
private final ScheduledExecutorService scheduler;
private final int maxReminders;
private final long backoffMs;
private final Metrics metrics; // CB-512: nullable — no registry in unit tests
/** Track targets that have an active schedule. */
private final ConcurrentHashMap<String, Boolean> activeTargets = new ConcurrentHashMap<>();
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
int maxReminders, long backoffMs) {
this(primaryRegistry, agents, inbox, scheduler, maxReminders, backoffMs, null);
}
/** As above, with a metric registry (CB-512) so push outcomes are counted. */
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
int maxReminders, long backoffMs, Metrics metrics) {
this.primaryRegistry = primaryRegistry;
this.agents = agents;
this.inbox = inbox;
this.scheduler = scheduler;
this.maxReminders = maxReminders;
this.backoffMs = backoffMs;
this.metrics = metrics;
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
if (metrics != null) {
metrics.inc(name, labels);
}
}
// --- decision logic (package-private for unit-testing) -------------------------------------
/** The action the loop should take for a target at the given reminder count. */
enum Action { INJECT, WAIT_BUSY, STOP }
/**
* Pure decision function: examine the current state and return what the loop should do.
*
* @param target the worker session (target terminal id)
* @param reminderCount how many nudges have been sent so far for this target
* @return the action the caller should take
*/
Action decide(String target, int reminderCount) {
if (!primaryRegistry.isKnown()) {
log.debug("push: primary unknown, stopping reminder for {}", target);
return Action.STOP;
}
if (inbox.peek(target).isEmpty()) {
log.debug("push: inbox empty for {}, stopping reminder", target);
return Action.STOP;
}
if (reminderCount >= maxReminders) {
log.debug("push: reminder cap ({}) reached for {}, stopping", maxReminders, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted");
return Action.STOP;
}
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
AgentStatus status;
try {
status = agents.status(primaryTerminal);
} catch (RuntimeException e) {
log.debug("push: status check failed for primary {}, will retry", primaryTerminal, e);
return Action.WAIT_BUSY;
}
if (status.injectable()) {
return Action.INJECT;
}
log.debug("push: primary {} is {} (not injectable), waiting", primaryTerminal, status);
return Action.WAIT_BUSY;
}
// --- public entrypoint ---------------------------------------------------------------------
/**
* Called when a reply is queued for {@code target}. Idempotent per target: a second call while
* a schedule is active is a no-op. The schedule nudges the primary, then schedules a follow-up
* check (reminder on backoff, or re-check on WAIT_BUSY), until the inbox is empty or the cap
* is reached.
*/
public void onReplyQueued(String target) {
if (activeTargets.putIfAbsent(target, Boolean.TRUE) != null) {
log.debug("push: already active for {}, ignoring duplicate trigger", target);
return; // already scheduled
}
log.debug("push: starting reminder loop for {}", target);
scheduleNext(target, 0);
}
/** Execute one loop tick — called on the scheduler thread. */
private void tick(String target, int reminderCount) {
var action = decide(target, reminderCount);
switch (action) {
case INJECT -> {
injectNudge(target, reminderCount);
scheduleNext(target, reminderCount + 1);
}
// Re-check after the configured backoff; the primary may become injectable soon.
case WAIT_BUSY -> scheduleNext(target, reminderCount);
case STOP -> {
activeTargets.remove(target);
log.debug("push: reminder loop ended for {}", target);
}
}
}
/** Send the nudge and log the event. */
private void injectNudge(String target, int reminderCount) {
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
String nudge = NUDGE_FORMAT.formatted(target, target);
try {
agents.send(primaryTerminal, nudge);
log.debug("push: nudge {}/{} sent to primary {} for target {}",
reminderCount + 1, maxReminders, primaryTerminal, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered");
} catch (RuntimeException e) {
log.warn("push: failed to nudge primary {} for target {} (reminder {}/{}): {}",
primaryTerminal, target, reminderCount + 1, maxReminders, e.toString());
}
}
/** Schedule the next tick on the scheduler thread pool. */
private void scheduleNext(String target, int nextReminderCount) {
scheduler.schedule(() -> tick(target, nextReminderCount), backoffMs, TimeUnit.MILLISECONDS);
}
// --- lifecycle -----------------------------------------------------------------------------
/** Shut down the scheduler. Outstanding reminders are cancelled. */
public void stop() {
scheduler.shutdownNow();
activeTargets.clear();
}
/** @see #stop() */
public void close() {
stop();
}
}
@@ -0,0 +1,35 @@
package dev.ltms.bridged.peer;
/**
* Declared capabilities of a {@link PeerLauncher}. The protocol is the union across all
* configured launchers; a verb invoked against a peer that lacks the capability returns a clean
* "unsupported for this peer" rather than a crash. Capabilities keep the protocol honest as peers
* diversify and prevent the core from assuming "every peer is a Claude in a worktree."
*/
public enum Capability {
/**
* The peer supports {@code bridge_ask} rendezvous — pausing its delegated turn to ask
* the primary a question, then resuming once answered. All Claude Code peers support this.
*/
MID_TURN_ASK,
/**
* The peer can open its own PR at the end of an implementation turn (CB-302). Opt-in per
* profile: granted only when the profile carries a git-forge token ({@code gitTokenEnv}).
*/
SELF_PR,
/**
* The peer can run inside a provisioned isolated git worktree. All CLI-based peers support
* this since their cwd is set at spawn time.
*/
WORKTREE,
/**
* The spawner can reconcile orphaned peers on boot — workers that outlived a prior daemon
* process and whose pane ids died with it (CB-117). Claude Code over herdr supports this
* via name-based matching against the herdr agent list.
*/
ORPHAN_REAP
}
@@ -0,0 +1,29 @@
package dev.ltms.bridged.peer;
/**
* An opaque handle returned by {@link PeerLauncher#spawn(SpawnRequest)}. The core routes on
* {@link #id()} (the registry/routing key) and uses {@link #terminalId()} for session tracking;
* launcher-private coordinates beyond these are reachable through the concrete implementation.
*
* <p>A {@link PeerHandle} is returned <em>after</em> the peer process is live — the launcher
* has already completed subscription-guarded env/vfs setup, process start, and placement. The
* handle is a ticket the core exchanges for the running peer, not a lazy/delayed reference.
*/
public interface PeerHandle {
/**
* The registry/routing key — an opaque, launcher-assigned identifier. For the herdr-backed
* launcher this is the herdr pane id; for other launchers it is whatever their transport
* uses. Guaranteed to be non-null and unique among live peers within a single daemon process.
*/
String id();
/**
* The transport-level session identifier used for message routing and presence tracking.
* For the herdr launcher this is the herdr terminal UUID. A non-herdr launcher may return
* its own analogous identifier, or {@code null} if the concept does not apply.
*/
default String terminalId() {
return null;
}
}
@@ -0,0 +1,85 @@
package dev.ltms.bridged.peer;
import java.util.List;
import java.util.Set;
/**
* SPI for materializing a connected peer — the only way the bridge core creates or tears down
* a peer process. Every launcher is a first-party, in-tree adapter selected by (future) profile
* config; today's single adapter is the {@code ClaudeCodeLauncher} / Claude Code over herdr.
*
* <p>The core delegates spawn and teardown to this interface without knowing how the peer is set
* up. Environment variables, CLI flags, subscription guards, transport (herdr tab/pane) layout,
* and naming conventions are all adapter-private — the core sees only the returned
* {@link PeerHandle} whose {@code id()} is the registry/routing key.
*
* <p>The interface is a superset of what {@code SessionManager} and {@code Bridged.main} call
* on the concrete launcher today.
*/
public interface PeerLauncher {
/**
* The set of {@link Capability capabilities} this launcher declares. A peer whose profile
* opts into a git-forge token should include {@link Capability#SELF_PR}; the base set for
* the Claude Code herdr adapter is always {@code MID_TURN_ASK, WORKTREE, ORPHAN_REAP}.
*/
Set<Capability> capabilities();
/**
* {@code profileName}/requestedCwd null/blank → default resolution. Returns after the peer
* process is live (env + argv + placement complete). Never returns {@code null}.
*
* @param req the spawn parameters (profile, requested cwd, caller cwd)
* @return a handle whose {@link PeerHandle#id()} is the registry/routing key
* @throws IllegalArgumentException if the profile is unknown and no default is configured
*/
PeerHandle spawn(SpawnRequest req);
/**
* The configured worker profile names — the set of names {@code spawn(profileName)} accepts.
*/
Set<String> profiles();
/**
* The profile a no-argument {@link #spawn(SpawnRequest)} uses, or {@code null} if none is configured.
*/
String defaultProfile();
/**
* Resolve the effective working directory for a spawn {@code req} without actually spawning.
* Resolution order: requestedCwd → profile cwd → callerCwd → daemon cwd.
*
* @return the resolved absolute path, never null/blank
*/
String effectiveCwd(SpawnRequest req);
/**
* The parity-overlay file list for {@code profileName} (default list when unset). Used by
* worktree provisioning to copy config files into the isolated checkout before spawning.
*/
List<String> parityOverlay(String profileName);
/**
* The set of all agents this launcher currently tracks, transport-specific. Each element
* exposes at minimum a pane-like {@code id()} matching this launcher's {@link PeerHandle}
* scheme, plus transport-level status. Callers merge this set with the session registry to
* build a live roster view.
*/
List<?> list();
/**
* Reap orphaned peers left behind by a prior daemon process. Only peers whose naming scheme
* matches this launcher's and whose nonce differs from the current process are eligible.
* Best-effort: a failure to list or to stop any one peer is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
int reapOrphanWorkers();
/**
* Tear a peer down by its registry/routing key ({@link PeerHandle#id()}). Tolerates an
* already-gone peer. Also cleans up launcher-private resources (e.g. empty dedicated tabs)
* when safe to do so.
*/
void stop(String id);
}
@@ -0,0 +1,18 @@
package dev.ltms.bridged.peer;
/**
* Thrown when a {@link PeerLauncher} starts a peer process but the peer
* does not reach an injectable (ready-to-receive) state within the configured
* timeout. The launcher MUST clean up any resources it created (pane, tab)
* before throwing — no orphaned peer or pane is left behind.
*
* <p>This is a spawn-time failure, distinct from a post-spawn disconnect.
* Callers treat this as a clean spawn error (the peer never materialized
* into a usable session), not a mid-life session fault.
*/
public final class PeerUnreachableException extends RuntimeException {
public PeerUnreachableException(String message) {
super(message);
}
}
@@ -0,0 +1,13 @@
package dev.ltms.bridged.peer;
/**
* Parameters for a {@link PeerLauncher#spawn(SpawnRequest)} call — the peer-neutral
* aggregation of what the core knows at delegation time: which profile to use, the caller's
* requested working directory, and the caller's own cwd (to inherit when no other cwd is set).
*
* <p>A null or blank {@code profileName} means "use the launcher's default profile."
* A null or blank {@code requestedCwd} means "inherit from config or caller."
* A null {@code callerCwd} means "the request came from the daemon itself (not a primary)."
*/
public record SpawnRequest(String profileName, String requestedCwd, String callerCwd) {
}
@@ -1,18 +1,34 @@
package dev.ltms.bridged.rest;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.bridged.auth.AuditLog;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.worker.WorkerService;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.peer.PeerLauncher;
import io.javalin.Javalin;
import io.javalin.http.Context;
import jakarta.servlet.http.HttpServlet;
import org.eclipse.jetty.servlet.ServletHolder;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.function.Function;
import java.util.stream.Collectors;
/**
* The REST surface — {@code bridged}'s contract and its testability seam. Every
@@ -25,25 +41,141 @@ import java.util.Map;
*/
public final class BridgedApp {
private final HerdrClient herdr;
private final WorkerService workers;
/** Default blocking window for a message; kept under typical HTTP idle timeouts. */
private static final long DEFAULT_MESSAGE_TIMEOUT_MS = 25_000;
private static final long MAX_MESSAGE_TIMEOUT_MS = 120_000;
/** Blocking window for a worker's bridge_ask (CB-205); the worker's MCP client caps its own call. */
private static final long DEFAULT_ASK_TIMEOUT_MS = 55_000;
private static final long MAX_ASK_TIMEOUT_MS = 115_000;
public BridgedApp(HerdrClient herdr, WorkerService workers) {
/** Context attribute under which the resolved caller is stashed by the auth filter. */
private static final String CALLER = "bridged.caller";
private final HerdrClient herdr;
private final PeerLauncher workers;
private final SessionManager sessions; // CB-301: authoritative session registry
private final MessageService messages;
private final WorkerPresence presence; // CB-113: which workers are MCP-connected (available)
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
private final CallerResolver auth; // CB-501: null → authz not enforced (legacy behaviour)
private final Metrics metrics; // CB-502: null → /metrics not exposed
private final ObjectMapper mapper = new ObjectMapper();
/**
* Legacy constructor — no identity resolution and no authorization, exactly as the REST surface
* behaved before CB-501. Retained so existing acceptance tests keep exercising handler
* behaviour without each needing an auth fixture.
*/
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
HttpServlet mcpServlet) {
this(herdr, workers, sessions, messages, presence, mcpServlet, null, null);
}
/**
* @param auth resolves each request's {@link Principal}; {@code null} disables authorization
* entirely (legacy). {@code main} always supplies one.
* @param metrics registry to instrument and expose at {@code GET /metrics}; {@code null} omits
* the endpoint
*/
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics) {
this.herdr = herdr;
this.workers = workers;
this.sessions = sessions;
this.messages = messages;
this.presence = presence;
this.mcpServlet = mcpServlet;
this.auth = auth;
this.metrics = metrics;
}
/** Wire routes onto a fresh, unstarted Javalin instance. Caller starts it. */
public Javalin build() {
Javalin app = Javalin.create(cfg -> cfg.showJavalinBanner = false);
Javalin app = Javalin.create(cfg -> {
cfg.showJavalinBanner = false;
if (mcpServlet != null) {
// The MCP server shares the daemon's port; Jetty routes /mcp to its servlet.
cfg.jetty.modifyServletContextHandler(h ->
h.addServlet(new ServletHolder(mcpServlet), "/mcp"));
}
});
// CB-501: resolve identity once per request, before any handler. /mcp does NOT pass through
// here — it is a raw servlet on Jetty's context handler — so BridgeMcp enforces separately
// against the same CallerResolver. Any check that lives in only one place is not a control.
if (auth != null) {
app.before(ctx -> ctx.attribute(CALLER,
auth.resolve(ctx.req().getRemoteAddr(), ctx.req().getRemotePort(),
ctx.header("Authorization"))));
}
app.get("/healthz", this::healthz);
if (metrics != null) {
app.get("/metrics", this::metrics);
}
app.get("/sessions", this::sessions);
app.get("/agents", this::agents);
app.post("/workers", this::spawnWorker);
app.get("/workers", this::listWorkers); // CB-304: registry roster + live herdr status
app.get("/profiles", this::profiles); // configured worker profiles
app.post("/workers", this::spawnWorker); // optional ?profile= or {"profile":…}
app.delete("/workers/{paneId}", this::stopWorker);
app.post("/sessions/{id}/message", this::sendMessage); // bridge_send (primary; blocking, wait:false, or answer via turnId)
app.post("/sessions/{id}/reply", this::replyMessage); // bridge_reply (worker)
app.get("/sessions/{id}/replies", this::drainReplies); // drain reply inbox (CB-307)
app.post("/sessions/{id}/ask", this::askMessage); // bridge_ask (worker → primary, CB-205)
app.get("/sessions/{id}/status", this::sessionStatus); // bridge_status
app.get("/tasks/{ticket}", this::taskStatus); // poll an async (wait:false) send
return app;
}
/**
* Gate a handler on the CB-505 authorization table. Returns {@code true} when the request may
* proceed; otherwise writes the error response and returns {@code false}.
*
* <p>401 vs 403 is a real distinction here: 401 means "you presented no usable identity" (a
* credential problem the caller can fix), 403 means "you are authenticated, but this is not
* yours" (a worker reaching for another worker's session, or for orchestration).
*/
private boolean allow(Context ctx, Authz.Action action, String target) {
if (auth == null) {
return true; // legacy: authorization not enforced
}
Principal caller = ctx.attribute(CALLER);
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ && action != Authz.Action.METRICS) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
return true;
}
if (Authz.isUnauthenticated(caller)) {
AuditLog.denied(caller, action, target, "unauthenticated");
countAuthFailure("unauthenticated");
ctx.status(401).json(Map.of("error", "unauthenticated",
"detail", "present Authorization: Bearer <token>"));
} else {
AuditLog.denied(caller, action, target, "forbidden");
countAuthFailure("forbidden");
ctx.status(403).json(Map.of("error", "forbidden",
"detail", caller.describe() + " may not " + action + " on "
+ (target == null ? "this resource" : target)));
}
return false;
}
private void countAuthFailure(String reason) {
if (metrics != null) {
metrics.inc("bridged_auth_failures_total", "reason", reason);
}
}
/** Prometheus scrape endpoint (CB-502). */
private void metrics(Context ctx) {
if (!allow(ctx, Authz.Action.METRICS, null)) {
return;
}
ctx.status(200).contentType("text/plain; version=0.0.4; charset=utf-8").result(metrics.render());
}
/** Liveness + herdr reachability. 200 when herdr answers ping, 503 otherwise. */
private void healthz(Context ctx) {
try {
@@ -63,6 +195,9 @@ public final class BridgedApp {
/** Sessions view derived from herdr {@code workspace.list} (one workspace → one row). */
private void sessions(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
JsonNode result = herdr.call("workspace.list");
List<Map<String, Object>> out = new ArrayList<>();
for (JsonNode w : result.path("workspaces")) {
@@ -78,25 +213,322 @@ public final class BridgedApp {
/** Discovery: every agent herdr tracks, keyed by its Claude session UUID. */
private void agents(Context ctx) {
ctx.status(200).json(Map.of("agents", workers.list().stream().map(BridgedApp::view).toList()));
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
ctx.status(200).json(Map.of("agents",
workers.list().stream().map(Agent.class::cast).map(BridgedApp::view).toList()));
}
/** Spawn a guard-checked worker. 403 if the base_url would breach the subscription boundary. */
/** CB-304: bridge-owned roster merged with live herdr status by paneId. */
private void listWorkers(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.paneId() != null)
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.paneId())))
.toList();
ctx.status(200).json(Map.of("workers", out));
}
/** The configured worker profiles and which one a no-argument spawn uses. */
private void profiles(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
ctx.status(200).json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile()));
}
/**
* Spawn a guard-checked worker. An optional {@code profile} (query param or {@code {"profile":…}}
* body) picks which configured profile; omitted → the default. 403 if the base_url would breach
* the subscription boundary, 400 for an unknown profile.
*/
private void spawnWorker(Context ctx) {
if (!allow(ctx, Authz.Action.SPAWN, null)) {
return;
}
String profile = ctx.queryParam("profile");
String cwd = ctx.queryParam("cwd");
String worktree = ctx.queryParam("worktree");
String ticket = ctx.queryParam("ticket");
if (profile == null || profile.isBlank() || cwd == null || cwd.isBlank()
|| worktree == null || worktree.isBlank()) {
try {
String body = ctx.body();
if (!body.isBlank()) {
JsonNode b = mapper.readTree(body);
if (profile == null || profile.isBlank()) profile = b.path("profile").asText(null);
if (cwd == null || cwd.isBlank()) cwd = b.path("cwd").asText(null);
if (worktree == null || worktree.isBlank()) worktree = b.path("worktree").asText(null);
if (ticket == null || ticket.isBlank()) ticket = b.path("ticket").asText(null);
}
} catch (Exception ignored) {
// A malformed/empty body just means "no overrides" → fall through to defaults.
}
}
WorktreeRequest wt = worktreeRequest(worktree, ticket);
try {
Agent worker = workers.spawn();
// No MCP caller over REST, so callerCwd and ownerTerminal are null.
WorkerSession worker = sessions.acquire(blankToNull(profile), blankToNull(cwd), null, null, wt);
ctx.status(201).json(view(worker));
} catch (GuardException e) {
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
} catch (IllegalArgumentException e) {
ctx.status(400).json(Map.of("error", "unknown_profile", "detail", e.getMessage()));
} catch (PeerUnreachableException e) {
ctx.status(502).json(Map.of("error", "spawn_timeout", "detail", e.getMessage()));
}
}
private static WorktreeRequest worktreeRequest(String worktree, String ticket) {
if (worktree == null || worktree.isBlank() || "false".equalsIgnoreCase(worktree)) {
return null;
}
if ("true".equalsIgnoreCase(worktree)) {
if (ticket == null || ticket.isBlank()) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return new WorktreeRequest(worktree, null);
}
private static String blankToNull(String s) {
return (s == null || s.isBlank()) ? null : s;
}
/** Tear a worker down by pane id. */
private void stopWorker(Context ctx) {
workers.stop(ctx.pathParam("paneId"));
String paneId = ctx.pathParam("paneId");
if (!allow(ctx, Authz.Action.STOP, paneId)) {
return;
}
sessions.release(paneId);
ctx.status(204);
}
/**
* The blocking delegation call (CB-104): inject {@code content} into the worker via the
* status-gated injector and block until the worker returns a structured {@code bridge_reply}.
* Times out with a typed 202 (working / queued / busy) rather than an error — the message may
* still land.
*/
private void sendMessage(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.SEND, id)) {
return;
}
String content;
String turnId;
long timeout;
boolean wait;
try {
JsonNode body = mapper.readTree(ctx.body());
content = body.path("content").asText("");
turnId = body.path("turnId").asText(null);
timeout = body.path("timeoutMs").asLong(DEFAULT_MESSAGE_TIMEOUT_MS);
wait = body.path("wait").asBoolean(true); // default: block for the reply (CB-104)
} catch (Exception e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
if (content.isBlank()) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "content is required"));
return;
}
timeout = Math.clamp(timeout, 1, MAX_MESSAGE_TIMEOUT_MS);
// Answering a worker's bridge_ask (CB-205): always blocks, and derives the worker from turnId.
if (turnId != null && !turnId.isBlank()) {
writeReply(ctx, id, messages.answer(turnId, content, timeout), timeout);
return;
}
if (!wait) {
// Fire-and-poll (CB-107): return a ticket immediately; the caller polls GET /tasks/{ticket}.
String ticket = messages.sendAsync(id, content);
ctx.status(202).json(Map.of("sessionId", id, "ticket", ticket, "status", "accepted"));
return;
}
try {
writeReply(ctx, id, messages.send(id, content, timeout), timeout);
} catch (HerdrException e) {
herdrError(ctx, e);
}
}
/**
* Render a {@link MessageService.Reply} onto the response — shared by a normal send and a
* bridge_ask answer. A structured/scraped completion is 200; a worker's mid-turn question a 202
* (with its {@code turnId}); a stale answer a 409; every other non-terminal outcome a typed 202.
*/
private void writeReply(Context ctx, String id, MessageService.Reply reply, long timeout) {
switch (reply.outcome()) {
case QUESTION -> ctx.status(202).json(Map.of(
"sessionId", id, "status", "question",
"question", reply.text(), "turnId", reply.turnId()));
case STALE_TURN -> ctx.status(409).json(Map.of(
"sessionId", id, "error", "stale_turn",
"detail", "that question is no longer open (timed out or already answered)"));
case REPLIED, COMPLETED_UNREPLIED -> {
// replySource distinguishes a structured bridge_reply from the CB-106 completion
// fallback (a scrape of the worker's transcript when it finished without replying).
String source = reply.outcome() == MessageService.Outcome.REPLIED ? "reply" : "transcript";
ctx.status(200).json(Map.of("sessionId", id, "reply", reply.text(), "replySource", source));
}
default -> ctx.status(202).json(Map.of(
"sessionId", id,
"status", switch (reply.outcome()) {
case TIMED_OUT_WORKING -> "working";
case TIMED_OUT_QUEUED -> "queued";
case BUSY -> "busy";
case WORKER_FAILED -> "failed";
default -> "done"; // unreachable (terminal outcomes handled above)
},
"detail", reply.outcome() == MessageService.Outcome.WORKER_FAILED && reply.text() != null
? reply.text()
: "no reply within " + timeout + "ms; poll status or retry"));
}
}
/**
* A worker's mid-turn question ({@code bridge_ask}, CB-205) — surfaces to the primary's open
* blocking send and blocks until it answers. 200 with the answer, 409 if no delegation is open,
* 202 if the primary stayed silent.
*/
private void askMessage(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.ASK, id)) {
return;
}
String question;
long timeout;
try {
JsonNode body = mapper.readTree(ctx.body());
question = body.path("question").asText("");
timeout = body.path("timeoutMs").asLong(DEFAULT_ASK_TIMEOUT_MS);
} catch (Exception e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
if (question.isBlank()) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "question is required"));
return;
}
timeout = Math.clamp(timeout, 1, MAX_ASK_TIMEOUT_MS);
MessageService.AskResult r = messages.ask(id, question, timeout);
switch (r.outcome()) {
case ANSWERED -> ctx.status(200).json(Map.of("sessionId", id, "answered", true, "answer", r.answer()));
case NO_WAITER -> ctx.status(409).json(Map.of(
"sessionId", id, "error", "no_pending_send",
"detail", "no primary is awaiting this turn to answer a question"));
case TIMED_OUT -> ctx.status(202).json(Map.of(
"sessionId", id, "status", "no_answer",
"detail", "the primary did not answer within " + timeout + "ms"));
}
}
/**
* The worker's structured reply ({@code bridge_reply}) — resolves the blocking send awaiting
* on this session, or queues the reply in the inbox when no send is open (CB-307).
*/
private void replyMessage(Context ctx) {
String id = ctx.pathParam("id");
// The rule that matters: a worker may reply only as itself. Over MCP this was already true
// structurally (identity comes from the connection, never an argument); over REST the path
// id was simply trusted, so this is where the invariant actually gets enforced.
if (!allow(ctx, Authz.Action.REPLY, id)) {
return;
}
String content;
try {
content = mapper.readTree(ctx.body()).path("content").asText("");
} catch (Exception e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
messages.reply(id, content);
ctx.status(200).json(Map.of("sessionId", id, "delivered", true));
}
/**
* Drain the reply inbox for a worker session — peek + ack any replies that arrived when no send
* was open. At-least-once: draining removes them from the inbox so a subsequent read returns
* nothing; an in-flight failure between the drain and the caller's processing re-surfaces them.
*/
private void drainReplies(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.DRAIN, id)) {
return;
}
var replies = messages.drainReplies(id);
ctx.status(200).json(Map.of("sessionId", id, "replies",
replies.stream().map(m -> Map.of(
"msgId", m.msgId(),
"content", m.content())).toList()));
}
/**
* Live lifecycle status of a worker (MCP `bridge_status` wraps this in CB-105), plus its
* <em>readiness</em> (CB-113): {@code ready} is true once the worker's Claude has connected the
* bridge MCP — the reliable "available to receive a task" signal, unlike bare {@code idle}, which
* is also true during boot.
*/
private void sessionStatus(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.READ, id)) {
return;
}
try {
ctx.status(200).json(Map.of(
"sessionId", id,
"status", messages.status(id).name().toLowerCase(),
"ready", presence.isPresent(id)));
} catch (HerdrException e) {
herdrError(ctx, e);
}
}
/** Poll an async (wait:false) delegation by ticket. 404 for an unknown/expired ticket. */
private void taskStatus(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"));
if (v == null) {
ctx.status(404).json(Map.of("error", "unknown_ticket", "detail", "no such task (or it has expired)"));
return;
}
Map<String, Object> body = new LinkedHashMap<>();
body.put("ticket", v.ticket());
body.put("phase", v.phase().name().toLowerCase());
if (v.reply() != null) {
body.put("reply", v.reply());
body.put("replySource", v.replySource());
}
if (v.detail() != null) {
body.put("detail", v.detail());
}
ctx.status(200).json(body);
}
/** Map a herdr failure: unknown target → 404, anything else → 502 (herdr is upstream). */
private static void herdrError(Context ctx, HerdrException e) {
if (e.code() != null && e.code().endsWith("_not_found")) {
ctx.status(404).json(Map.of("error", "session_not_found", "detail", e.getMessage()));
} else {
ctx.status(502).json(Map.of("error", "herdr_error", "detail", e.getMessage()));
}
}
/** Stable JSON projection of an agent (null-safe for the start-time shape). */
private static Map<String, Object> view(Agent a) {
Map<String, Object> m = new LinkedHashMap<>();
@@ -109,4 +541,22 @@ public final class BridgedApp {
m.put("status", a.status().name().toLowerCase());
return m;
}
/** CB-301 projection of an authoritative bridge-owned session. */
private static Map<String, Object> view(WorkerSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("terminalId", s.terminalId());
m.put("paneId", s.paneId());
m.put("profile", s.profile());
m.put("cwd", s.cwd());
m.put("ownerTerminal", s.ownerTerminal());
m.put("state", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
return m;
}
}
@@ -0,0 +1,179 @@
package dev.ltms.bridged.session;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.BufferedReader;
import java.io.IOException;
import java.io.InputStreamReader;
import java.io.UncheckedIOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.security.SecureRandom;
import java.util.List;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.stream.Collectors;
/**
* Production {@link Worktrees} implementation that shells {@code git} via {@link ProcessBuilder}.
* Non-zero exits become {@link WorktreeException}. Worktree directories live under a configurable
* root (default: a sibling {@code .bridged-worktrees} of the repo root) so they are never nested
* inside the primary working tree.
*/
public final class GitWorktrees implements Worktrees {
private static final Logger log = LoggerFactory.getLogger(GitWorktrees.class);
private final String configuredRoot;
private final SecureRandom random = new SecureRandom();
private final AtomicLong seq = new AtomicLong();
/** Default constructor: worktree root is derived per-repo as {@code <repoRoot>/../.bridged-worktrees}. */
public GitWorktrees() {
this(null);
}
/** @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of the repo root. */
public GitWorktrees(String configuredRoot) {
this.configuredRoot = configuredRoot;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
String base = (baseRef == null || baseRef.isBlank()) ? "HEAD" : baseRef;
String nonce = nonce();
Path root = resolveRoot(repoRoot);
Path path = root.resolve(nonce);
try {
Files.createDirectories(root);
} catch (IOException e) {
throw new WorktreeException("cannot create worktree root " + root + ": " + e.getMessage(), e);
}
String wt = path.toAbsolutePath().toString();
log.info("adding worktree branch={} path={} base={}", branch, wt, base);
exec("git", "-C", repoRoot, "worktree", "add", wt, "-b", branch, base);
return wt;
}
@Override
public void remove(String repoRoot, String worktreePath) {
Path p = Path.of(worktreePath);
if (!Files.exists(p)) {
log.debug("worktree {} already gone — nothing to remove", worktreePath);
return;
}
log.info("removing worktree {}", worktreePath);
exec("git", "-C", repoRoot, "worktree", "remove", "--force", worktreePath);
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
if (overlay == null || overlay.isEmpty()) {
return;
}
Path srcRoot = Path.of(repoRoot).toAbsolutePath().normalize();
Path dstRoot = Path.of(worktreePath).toAbsolutePath().normalize();
for (String rel : overlay) {
Path src = srcRoot.resolve(rel).normalize();
if (!Files.exists(src)) {
log.debug("parity overlay source missing — skipping {}", rel);
continue;
}
Path dst = dstRoot.resolve(rel).normalize();
try {
Files.createDirectories(dst.getParent());
Files.copy(src, dst, StandardCopyOption.REPLACE_EXISTING, StandardCopyOption.COPY_ATTRIBUTES);
log.debug("copied parity overlay {}", rel);
} catch (IOException e) {
throw new WorktreeException("cannot copy overlay " + rel + ": " + e.getMessage(), e);
}
if (isTracked(dstRoot, rel)) {
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", rel);
log.debug("marked overlay --skip-worktree {}", rel);
}
}
}
@Override
public String repoRoot(String cwd) {
String out = exec("git", "-C", cwd, "rev-parse", "--show-toplevel");
return Path.of(out.trim()).toAbsolutePath().normalize().toString();
}
/** Resolve the directory that will hold per-session worktree checkouts. */
private Path resolveRoot(String repoRoot) {
if (configuredRoot != null && !configuredRoot.isBlank()) {
return Path.of(configuredRoot).toAbsolutePath().normalize();
}
Path repo = Path.of(repoRoot).toAbsolutePath().normalize();
return repo.resolveSibling(".bridged-worktrees");
}
private String nonce() {
return String.format("%06x", random.nextInt(1 << 24)) + "-" + seq.incrementAndGet();
}
private boolean isTracked(Path worktreeRoot, String rel) {
return exitCode("git", "-C", worktreeRoot.toString(), "ls-files", "--error-unmatch", rel) == 0;
}
/**
* Run a command and return its stdout. Non-zero exit → {@link WorktreeException} with both
* stdout and stderr (merged by redirectErrorStream).
*/
private String exec(String... command) {
String out;
int code;
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
try (BufferedReader r = new BufferedReader(new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
out = r.lines().collect(Collectors.joining("\n"));
} catch (IOException e) {
p.destroyForcibly();
throw new UncheckedIOException(e);
}
try {
if (!p.waitFor(30, TimeUnit.SECONDS)) {
p.destroyForcibly();
throw new WorktreeException("command timed out: " + String.join(" ", command) + "\n" + out);
}
code = p.exitValue();
} catch (InterruptedException e) {
p.destroyForcibly();
Thread.currentThread().interrupt();
throw new WorktreeException("interrupted waiting for command: " + String.join(" ", command), e);
}
if (code != 0) {
throw new WorktreeException("exit " + code + " for: " + String.join(" ", command)
+ (out.isBlank() ? "" : "\n" + out));
}
return out;
}
private int exitCode(String... command) {
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
try {
if (!p.waitFor(30, TimeUnit.SECONDS)) {
p.destroyForcibly();
throw new WorktreeException("command timed out: " + String.join(" ", command));
}
return p.exitValue();
} catch (InterruptedException e) {
p.destroyForcibly();
Thread.currentThread().interrupt();
throw new WorktreeException("interrupted waiting for command: " + String.join(" ", command), e);
}
}
}
@@ -0,0 +1,444 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Consumer;
import java.util.function.LongSupplier;
/**
* Authoritative in-daemon registry of the worker sessions this {@code bridged} process spawned.
* Delegates spawn/teardown to a {@link PeerLauncher} (which performs subscription-guarded env
* setup and process/materialization) and adds lifecycle tracking, ownership, and deterministic
* teardown on top.
*
* <p>The state machine is intentionally one-shot / no-reuse: every acquired worker is fresh,
* and a finished or released worker is torn down, never pooled. {@link #recycle} is a convenience
* for {@code release + acquire} with a new distinct pane id.
*
* <p>The manager implements {@link TurnListener} so the injector's turn boundaries drive
* {@code READY → BUSY → DONE} (or {@code FAILED}). It exposes a {@link WorkerPresence} view via
* {@link #asPresence()}: any MCP contact from a worker marks it present and simultaneously
* transitions the session {@code SPAWNING → READY}.
*/
public final class SessionManager implements TurnListener {
private static final Logger log = LoggerFactory.getLogger(SessionManager.class);
private final PeerLauncher launcher;
private final Worktrees worktrees;
private final ConcurrentHashMap<String /*paneId*/, WorkerSession> registry = new ConcurrentHashMap<>();
private final WorkerPresence presence;
private final SecureRandom nonceRandom = new SecureRandom();
private final AtomicLong nonceSeq = new AtomicLong();
private final LongSupplier nowNanos;
private final int contextCap;
/** CB-516: notified with a terminalId on every release; no-op until wired. */
private volatile Consumer<String> releaseListener = _ -> { };
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
public SessionManager(PeerLauncher launcher) {
this(launcher, new GitWorktrees(), System::nanoTime, 0);
}
/** Backward-compatible constructor with an injectable worktree seam. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees) {
this(launcher, worktrees, System::nanoTime, 0);
}
/** Test constructor with an injectable clock. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos) {
this(launcher, worktrees, nowNanos, 0);
}
/** Production constructor with a configured context turn cap. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, int contextCap) {
this(launcher, worktrees, System::nanoTime, contextCap);
}
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos,
int contextCap) {
this.launcher = launcher;
this.worktrees = worktrees;
this.presence = new PresenceBridge(this);
this.nowNanos = nowNanos;
this.contextCap = contextCap;
}
/**
* The single {@link WorkerPresence} view of this manager: it records availability and forwards
* the signal to the {@code SPAWNING → READY} transition. Pass this to the {@code Injector} and
* {@code BridgeMcp} where they previously accepted a plain {@link WorkerPresence}. The same
* instance is returned every call — presence is shared state, so a fresh bridge per call would
* fragment the {@code present} set and lose signals across callers.
*/
public WorkerPresence asPresence() {
return presence;
}
/**
* Spawn a worker and register it as {@link WorkerSession.State#SPAWNING}. The caller's
* identity is recorded as {@code ownerTerminal} ({@code null} for daemon/anon callers).
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal) {
return acquire(profile, requestedCwd, callerCwd, ownerTerminal, null);
}
/**
* Spawn a worker, optionally inside a fresh git worktree. When {@code wt} is non-null the
* worktree is provisioned, parity-overlaid, and its path becomes the worker's cwd. On any
* failure before registration the worktree is removed so no dangling checkout is left.
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
if (wt == null) {
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd);
PeerHandle handle = launcher.spawn(req);
String resolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
String cwd = launcher.effectiveCwd(req);
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
cwd,
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
null,
null);
registry.put(handle.id(), session);
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
return session;
}
return acquireWithWorktree(profile, requestedCwd, callerCwd, ownerTerminal, wt);
}
/** Tear a worker down by pane id and remove it from the registry. Idempotent. */
public void release(String paneId) {
WorkerSession removed = registry.remove(paneId);
if (removed != null) {
log.debug("releasing session pane={} terminal={} state={}",
removed.paneId(), removed.terminalId(), removed.state());
// CB-516: a send still waiting on this worker can never be answered now. Tell the
// listener BEFORE the pane is torn down, so a blocked caller fails fast with a real
// reason instead of sitting on a rendezvous nothing will ever resolve.
notifyReleased(removed.terminalId());
}
launcher.stop(paneId);
if (removed != null && removed.worktree() != null) {
worktrees.remove(worktrees.repoRoot(removed.cwd()), removed.worktree());
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is released
* (CB-516). Every teardown path funnels through {@link #release}, so one hook covers the REST
* and MCP stop tools, the idle-TTL reaper, {@code recycle}, and shutdown drain alike.
*
* <p>Set rather than injected because {@code MessageService} — the intended listener — is
* constructed after this manager (it needs the injector and rendezvous, which need the session
* presence view this manager exposes). Wiring it at construction would require breaking that
* cycle for one callback.
*/
public void onRelease(Consumer<String> listener) {
this.releaseListener = (listener == null) ? _ -> { } : listener;
}
/** A listener failure must never prevent the teardown it is reacting to. */
private void notifyReleased(String terminalId) {
if (terminalId == null) {
return;
}
try {
releaseListener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
}
}
private WorkerSession acquireWithWorktree(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
String resolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
// nor a caller cwd, so taking the first non-blank of those two yielded null and put
// `git -C null` on the command line — an NPE out of ProcessBuilder, surfacing as HTTP 500.
// The non-worktree path always used this chain; only this branch was missed.
String repoRoot = worktrees.repoRoot(
launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd)));
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
String path = null;
PeerHandle handle;
try {
path = worktrees.add(repoRoot, branch, wt.baseRef());
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(resolvedProfile));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd));
} catch (RuntimeException e) {
if (path != null) {
try {
worktrees.remove(repoRoot, path);
} catch (RuntimeException cleanup) {
log.warn("failed to clean up worktree {} after spawn error: {}", path, cleanup.getMessage());
}
}
throw e;
}
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
resolveCwd(path, profile, callerCwd),
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
path,
branch);
registry.put(handle.id(), session);
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
return session;
}
private String slug(String raw) {
return raw == null ? "ticket" : raw.toLowerCase().replaceAll("[^a-z0-9]+", "-").replaceAll("^-+|-+$", "");
}
private String nonce() {
return String.format("%06x", nonceRandom.nextInt(1 << 24)) + "-" + nonceSeq.incrementAndGet();
}
/**
* Release the old session and acquire a fresh one with the same profile and working directory.
* The new session is guaranteed to have a pane id distinct from the old one (no-reuse invariant).
*/
public WorkerSession recycle(String paneId) {
WorkerSession old = registry.get(paneId);
if (old == null) {
throw new IllegalArgumentException("no session for paneId " + paneId);
}
release(paneId);
return acquire(old.profile(), old.cwd(), old.cwd(), old.ownerTerminal());
}
/** The session for {@code paneId}, if it is still registered and not released. */
public Optional<WorkerSession> get(String paneId) {
return Optional.ofNullable(registry.get(paneId));
}
/** Bridge-owned roster: all registered sessions (acquired minus released). */
public List<WorkerSession> roster() {
return List.copyOf(registry.values());
}
/**
* CB-304 merged roster+live view. The registry is authoritative for worktree, branch,
* profile, owner, and state; the optional live agent supplies the herdr-reported status.
*/
public static Map<String, Object> rosterView(WorkerSession session, Agent live) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", session.terminalId());
m.put("paneId", session.paneId());
m.put("profile", session.profile());
m.put("state", session.state().name().toLowerCase());
if (session.worktree() != null) {
m.put("worktree", session.worktree());
}
if (session.branch() != null) {
m.put("branch", session.branch());
}
if (session.ownerTerminal() != null) {
m.put("owner", session.ownerTerminal());
}
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
return m;
}
/** Lifecycle hook: worker became available on the bridge MCP. */
void onReady(String terminalId) {
transitionByTerminal(terminalId, WorkerSession.State.SPAWNING, WorkerSession.State.READY);
}
/**
* Lifecycle hook: a message was delivered into the worker — it is now busy on a turn.
* The turn count is bumped and the activity timestamp is refreshed. A {@code DONE} session
* can be re-delivered for multi-turn reuse until it is released.
*/
@Override
public void onDelivered(String target) {
WorkerSession current = findByTerminal(target);
if (current == null) return;
if (current.state() != WorkerSession.State.READY && current.state() != WorkerSession.State.DONE) {
return;
}
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.BUSY).bumpTurn(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} {} -> BUSY turn={}",
target, current.paneId(), current.state(), updated.turnCount());
}
}
/** Lifecycle hook: the worker's delegated turn completed successfully. */
@Override
public void onTurnComplete(String target) {
WorkerSession current = findByTerminal(target);
if (current == null || current.state() != WorkerSession.State.BUSY) return;
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.DONE).withActivity(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} BUSY -> DONE turn={}",
target, current.paneId(), updated.turnCount());
}
if (contextCap > 0 && updated.turnCount() >= contextCap) {
release(current.paneId());
}
}
/** Lifecycle hook: the worker's delegated turn failed. */
@Override
public void onTurnFailed(String target) {
onFailed(target);
}
/** Lifecycle hook: the worker vanished or was dropped mid-life. */
void onFailed(String target) {
WorkerSession current = findByTerminal(target);
if (current == null) return;
if (current.state() == WorkerSession.State.RELEASED) return;
if (replace(current, current.withState(WorkerSession.State.FAILED))) {
log.debug("session marked failed terminal={} pane={}", target, current.paneId());
}
}
/**
* Best-effort reap of sessions that have been idle longer than {@code idleTtlNanos}. Only
* {@code READY} and {@code DONE} sessions are eligible — never a {@code SPAWNING} or
* {@code BUSY} worker. Returns the number of sessions released.
*/
int reapIdle(long idleTtlNanos) {
long now = nowNanos.getAsLong();
int reaped = 0;
for (WorkerSession s : roster()) {
if (s.state() != WorkerSession.State.READY && s.state() != WorkerSession.State.DONE) {
continue;
}
if (now - s.lastActivityAtNanos() > idleTtlNanos) {
release(s.paneId());
reaped++;
}
}
return reaped;
}
/**
* Gracefully drain all registered sessions. For each session that is {@code BUSY}, poll up to
* {@code timeoutNanos} for it to leave {@code BUSY}, then release it regardless. Non-busy
* sessions are released immediately. A failure releasing one session is logged and does not
* abort the rest.
*/
void drainAll(long timeoutNanos) {
long deadline = System.nanoTime() + timeoutNanos;
for (WorkerSession s : roster()) {
try {
if (s.state() == WorkerSession.State.BUSY) {
while (System.nanoTime() < deadline) {
WorkerSession current = registry.get(s.paneId());
if (current == null || current.state() != WorkerSession.State.BUSY) {
break;
}
try {
long remaining = deadline - System.nanoTime();
Thread.sleep(Math.min(TimeUnit.NANOSECONDS.toMillis(remaining), 50));
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
break;
}
}
}
release(s.paneId());
} catch (RuntimeException e) {
log.warn("drain failed for pane={}; continuing with remaining sessions", s.paneId(), e);
}
}
}
/**
* Close this manager by draining all sessions. The timeout comes from configuration when set,
* otherwise a sensible default.
*/
public void close(Integer drainTimeoutSeconds) {
int seconds = (drainTimeoutSeconds != null && drainTimeoutSeconds > 0) ? drainTimeoutSeconds : 5;
drainAll(TimeUnit.SECONDS.toNanos(seconds));
}
/** Number of sessions currently registered. */
public int size() {
return registry.size();
}
private WorkerSession findByTerminal(String terminalId) {
for (WorkerSession s : registry.values()) {
if (terminalId.equals(s.terminalId())) return s;
}
return null;
}
private void transitionByTerminal(String terminalId, WorkerSession.State from,
WorkerSession.State to) {
WorkerSession current = findByTerminal(terminalId);
if (current == null || current.state() != from) return;
long now = nowNanos.getAsLong();
if (replace(current, current.withState(to).withActivity(now))) {
log.debug("session transitioned terminal={} pane={} {} -> {}",
terminalId, current.paneId(), from, to);
}
}
private boolean replace(WorkerSession expected, WorkerSession updated) {
return registry.replace(expected.paneId(), expected, updated);
}
private String resolveCwd(String requestedCwd, String profileName, String callerCwd) {
return launcher.effectiveCwd(new SpawnRequest(profileName, requestedCwd, callerCwd));
}
/** WorkerPresence bridge that also drives the manager's READY transition. */
private static final class PresenceBridge extends WorkerPresence {
private final SessionManager sessions;
PresenceBridge(SessionManager sessions) {
this.sessions = sessions;
}
@Override
public void markPresent(String terminal) {
super.markPresent(terminal);
sessions.onReady(terminal);
}
}
}
@@ -0,0 +1,70 @@
package dev.ltms.bridged.session;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.TimeUnit;
/**
* Periodic virtual-thread reaper that tears down {@code READY}/{@code DONE} sessions which have
* exceeded their idle TTL. Modeled on {@link dev.ltms.bridged.inject.StatusPoller}: a single
* virtual-thread loop, idempotent start/stop, and no {@code ScheduledExecutorService}.
*/
public final class SessionReaper {
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
private final SessionManager sessions;
private final long idleTtlNanos;
private final long intervalMillis;
private volatile boolean running;
private Thread thread;
/** Construct a reaper with the default 5-second polling interval. */
public SessionReaper(SessionManager sessions, long idleTtlSeconds) {
this(sessions, idleTtlSeconds, DEFAULT_INTERVAL_MILLIS);
}
/** Construct a reaper with an explicit polling interval (useful for tests). */
public SessionReaper(SessionManager sessions, long idleTtlSeconds, long intervalMillis) {
this.sessions = sessions;
this.idleTtlNanos = TimeUnit.SECONDS.toNanos(idleTtlSeconds);
this.intervalMillis = intervalMillis;
}
/** Start the reaper loop on a virtual thread. Idempotent. */
public synchronized void start() {
if (running) return;
running = true;
thread = Thread.ofVirtual().name("session-reaper").start(this::loop);
log.info("session reaper started (idle ttl {}s, interval {}ms)",
TimeUnit.NANOSECONDS.toSeconds(idleTtlNanos), intervalMillis);
}
private void loop() {
while (running) {
try {
sessions.reapIdle(idleTtlNanos);
} catch (RuntimeException e) {
log.warn("session reaper iteration failed; continuing", e);
}
sleep();
}
}
private void sleep() {
try {
Thread.sleep(intervalMillis);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
running = false;
}
}
/** Stop the reaper loop. Idempotent. */
public synchronized void stop() {
running = false;
if (thread != null) thread.interrupt();
}
}
@@ -0,0 +1,58 @@
package dev.ltms.bridged.session;
/**
* A bridge-owned worker session — the authoritative in-daemon record of a worker this
* process spawned. Immutable; state transitions are performed by replacing the record in
* {@link SessionManager}'s registry.
*
* @param paneId herdr pane handle — the registry key and the argument to teardown
* @param terminalId herdr terminal handle — the {@code target} for send/read/status
* @param profile the worker profile name that spawned this session
* @param cwd the resolved working directory the worker started in
* @param ownerTerminal the caller that requested this worker ({@code null} = daemon/anon)
* @param spawnedAtNanos {@link System#nanoTime()} when the session was registered
* @param lastActivityAtNanos {@link System#nanoTime()} of the most recent lifecycle event
* @param turnCount number of delegated turns that have been delivered to this session
* @param state current lifecycle state in the one-shot FSM
*/
public record WorkerSession(
String paneId,
String terminalId,
String profile,
String cwd,
String ownerTerminal,
long spawnedAtNanos,
long lastActivityAtNanos,
int turnCount,
State state,
String worktree,
String branch) {
/** One-shot worker lifecycle states. */
public enum State {
SPAWNING,
READY,
BUSY,
DONE,
FAILED,
RELEASED
}
/** Return a copy of this session in {@code state}. */
public WorkerSession withState(State state) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch);
}
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
public WorkerSession withActivity(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount, state, worktree, branch);
}
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
public WorkerSession bumpTurn(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount + 1, state, worktree, branch);
}
}
@@ -0,0 +1,12 @@
package dev.ltms.bridged.session;
/** Non-zero exit or I/O failure from a git worktree operation. */
public final class WorktreeException extends RuntimeException {
public WorktreeException(String message) {
super(message);
}
public WorktreeException(String message, Throwable cause) {
super(message, cause);
}
}
@@ -0,0 +1,6 @@
package dev.ltms.bridged.session;
/** Ask {@link SessionManager#acquire} to provision an isolated worktree. null ⇒ run in the shared primary tree. */
public record WorktreeRequest(String ticketSlug, String baseRef) {
// ticketSlug seeds the branch name; baseRef null/blank ⇒ current HEAD of the repo.
}
@@ -0,0 +1,18 @@
package dev.ltms.bridged.session;
import java.util.List;
/** Seam between {@link SessionManager} and git worktree operations. Tests use a recording fake. */
public interface Worktrees {
/** git -C <repoRoot> worktree add <path> -b <branch> <baseRef|HEAD>. Returns the worktree path. */
String add(String repoRoot, String branch, String baseRef);
/** git -C <repoRoot> worktree remove --force <path>. Idempotent (already-gone tolerated). */
void remove(String repoRoot, String worktreePath);
/** Copy each existing overlay path repoRoot→worktree; mark tracked ones --skip-worktree. */
void overlayParity(String repoRoot, String worktreePath, List<String> overlay);
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
String repoRoot(String cwd);
}
@@ -0,0 +1,195 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>Claude Code</strong> — the safe path from a
* delegation request to a running off-subscription Claude.
*
* <p>Everything transport-related (tab/pane placement, the CB-306 spawn-readiness gate, unique
* naming, CB-117 orphan reap, teardown, listing, cwd resolution) lives in the base. This class
* supplies only the two Claude-specific seams:
* <ul>
* <li>the {@code claude} name prefix (so reap matches {@code claude-*} panes, never another
* adapter's), and</li>
* <li>{@link #buildLaunch}, which encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em> touching
* herdr, and mount the bridge MCP + reply charter as inline launch flags. A worker's base_url
* lives in the env map handed to herdr and nowhere else; {@code bridged}'s own environment is
* never mutated, and nothing is written to the worker's profile.</li>
* </ul>
*/
public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "claude";
private final SubscriptionGuard guard;
/**
* Standing instruction appended to the worker's system prompt so it returns its result via
* {@code bridge_reply}. Injected as a launch flag, so nothing is written to the worker's
* profile — it is guidance, and a worker that never replies is caught by the send's timeout.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet. Every message you "
+ "receive arrives through the bridge, and the ONLY channel back to the sender is the "
+ "bridge_reply MCP tool. Text you write in your terminal is NOT sent anywhere — the "
+ "sender cannot see your screen, so an in-terminal answer is silently discarded. "
+ "Therefore you MUST end EVERY turn by calling bridge_reply with `content` set to your "
+ "complete response. This holds for every message without exception — tasks, questions, "
+ "clarifications, acknowledgements, and ordinary back-and-forth conversation. Call "
+ "bridge_reply exactly once, as the final action of your turn, with your full answer in "
+ "`content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so
* existing deployments and tests keep the legacy non-blocking spawn semantics.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, guard, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300));
}
/**
* Production constructor with spawn-ready gate enabled. The gate polls {@code agents.status()}
* until the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs));
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
* {@code sleeper} is never called when the gate is disabled ({@code spawnReadyTimeoutMs == 0}).
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param guard subscription-boundary guard (checked before spawning)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (e.g. {@code () -> Thread.sleep(pollMs)}); it
* already encodes the poll interval, so the 8th positional argument
* (poll ms) is accepted for API symmetry but otherwise unused here
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.guard = guard;
}
/**
* {@inheritDoc}
*
* <p>The spawn sequence encodes the subscription boundary: assert the profile's base_url is on
* the allowlist <em>before</em> any herdr call, then build the worker env with
* {@code ANTHROPIC_*}, the parity-neutral git-forge grant, and the bridge MCP + reply charter
* mounted as inline launch flags.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
String baseUrl = cfg.baseUrl();
guard.assertWorker(baseUrl); // hard stop before we spawn anything
Map<String, String> workerEnv = baseEnv(cfg);
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithBridge(cfg));
}
/**
* The launch argv, plus — when {@code worker.mcpUrl} is set — inline {@code --mcp-config} for
* the bridge server and {@code --append-system-prompt} for the {@link #REPLY_CHARTER}. Neither
* touches the profile's config; both are pure command-line flags. This inline-flag mount is
* Claude Code specific — other adapters mount MCP and instructions their own way.
*/
private List<String> argvWithBridge(BridgedConfig.Worker cfg) {
if (!cfg.hasMcp()) {
return cfg.argv();
}
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
List<String> argv = mutableArgv(cfg.argv());
argv.add("--mcp-config");
argv.add(mcpJson);
argv.add("--append-system-prompt");
argv.add(REPLY_CHARTER);
return argv;
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
// --- CB-117 reap predicate (Claude prefix), kept for direct unit testing -------------------
/**
* Whether {@code name} is a Claude Code bridge worker started by a <em>different</em> process
* than {@code currentNonce}. A thin {@code claude}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -0,0 +1,160 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.EnumSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
/**
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
* router in front of one {@link HerdrPeerLauncher} per peer {@code kind} (Claude Code, opencode, …).
* It owns no transport of its own; it dispatches each SPI call to the delegate that owns the profile
* involved, and fans the fleet-wide queries (list/reap/caps/profiles) across all delegates.
*
* <p>Routing rules:
* <ul>
* <li><strong>By profile</strong> — {@link #spawn}, {@link #effectiveCwd}, {@link #parityOverlay}
* resolve the profile (a null/blank name → the global {@link #defaultProfile}) and delegate to
* the single adapter that declares it. Profiles partition cleanly across adapters: the
* constructor rejects a name claimed by two.</li>
* <li><strong>By pane id</strong> — {@link #stop} routes to the adapter that spawned that pane
* (recorded at spawn time). A pane the composite never spawned (only real for a caller that
* hand-rolls an id) falls back to the first delegate; teardown is pane-id addressed and
* tab cleanup is single-occupant guarded, so it is safe either way.</li>
* <li><strong>Fleet-wide</strong> — {@link #reapOrphanWorkers} and {@link #capabilities} fan out
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
* shares one herdr connection and so reports the same global agent set.</li>
* </ul>
*/
public final class CompositePeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(CompositePeerLauncher.class);
private final List<HerdrPeerLauncher> delegates;
private final Map<String, HerdrPeerLauncher> byProfile;
private final String defaultProfile;
/** paneId → the delegate that spawned it, so {@link #stop} tears down through the right adapter. */
private final Map<String, HerdrPeerLauncher> spawnedBy = new ConcurrentHashMap<>();
/**
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to (may be null)
* @throws IllegalArgumentException if {@code delegates} is empty or two adapters claim one profile
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates, String defaultProfile) {
if (delegates.isEmpty()) {
throw new IllegalArgumentException("at least one peer adapter must be configured");
}
this.delegates = List.copyOf(delegates);
this.defaultProfile = defaultProfile;
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
HerdrPeerLauncher prev = index.putIfAbsent(profile, d);
if (prev != null) {
throw new IllegalArgumentException(
"worker profile '" + profile + "' is claimed by two peer adapters");
}
}
}
this.byProfile = Map.copyOf(index);
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (resolved == null) {
// No profile and no default configured — hand to the first delegate so it raises the
// same "no default" error it would on its own; keeps the SPI contract single-sourced.
return delegates.getFirst();
}
HerdrPeerLauncher d = byProfile.get(resolved);
if (d == null) {
throw new IllegalArgumentException("unknown worker profile: " + resolved);
}
return d;
}
@Override
public PeerHandle spawn(SpawnRequest req) {
HerdrPeerLauncher d = route(req.profileName());
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
}
@Override
public String effectiveCwd(SpawnRequest req) {
return route(req.profileName()).effectiveCwd(req);
}
@Override
public List<String> parityOverlay(String profileName) {
return route(profileName).parityOverlay(profileName);
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.remove(id);
if (d == null) {
log.debug("stop({}) — no recorded owner, routing to the first adapter (pane-addressed)", id);
d = delegates.getFirst();
}
d.stop(id);
}
@Override
public Set<String> profiles() {
return byProfile.keySet();
}
@Override
public String defaultProfile() {
return defaultProfile;
}
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
Map<String, Agent> byPane = new LinkedHashMap<>();
for (HerdrPeerLauncher d : delegates) {
for (Agent a : d.list()) {
if (a.paneId() != null) {
byPane.putIfAbsent(a.paneId(), a);
}
}
}
return List.copyOf(byPane.values());
}
@Override
public int reapOrphanWorkers() {
int reaped = 0;
for (HerdrPeerLauncher d : delegates) {
reaped += d.reapOrphanWorkers();
}
return reaped;
}
/** The union of every adapter's capabilities — a capability any adapter offers, the fleet offers. */
@Override
public Set<Capability> capabilities() {
EnumSet<Capability> caps = EnumSet.noneOf(Capability.class);
for (HerdrPeerLauncher d : delegates) {
caps.addAll(d.capabilities());
}
return Set.copyOf(caps);
}
}
@@ -0,0 +1,546 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.Collection;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* Abstract base for {@link PeerLauncher} adapters that materialize a peer as a <em>herdr</em>
* agent (a CLI coding agent running in a herdr tab/pane). It owns everything that is the same
* regardless of <em>which</em> coding agent runs: tab/pane placement, the CB-306 spawn-readiness
* gate, unique naming, CB-117 orphan reap, teardown, {@link #list() listing}, and cwd resolution.
*
* <p>Two seams are peer-specific and supplied by the concrete adapter:
* <ul>
* <li>{@code namePrefix} (constructor arg) — the label prefix ({@code claude}, {@code opencode})
* that drives both unique naming and the orphan-reap pattern, so each adapter reaps only its
* own kind of pane and never another's.</li>
* <li>{@link #buildLaunch(BridgedConfig.Worker)} — the peer-specific env map + argv, including any
* subscription/guard check, MCP mount, and instruction injection. The base never sees how the
* peer is configured; it only places and starts the returned {@link Launch}.</li>
* </ul>
*
* <p>Placement: in the default {@code tab} policy a peer lands in its own tab inside a dedicated
* worker space (found-or-created once, then shared), so peers never split or clutter the user's
* real work spaces. Teardown removes the peer's pane <em>and</em> its now-empty tab, tolerating an
* already-gone peer so a repeated DELETE is harmless.
*/
public abstract class HerdrPeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(HerdrPeerLauncher.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
private final String namePrefix; // label prefix: naming + reap scheme
private final AgentControl agents;
private final WorkspaceControl spaces;
private final Map<String, BridgedConfig.Worker> profiles; // profile name → spawn settings
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
/** Host env lookup (injectable for tests); adapters read it in {@link #buildLaunch}. */
protected final Function<String, String> env;
private final AtomicLong nameSeq = new AtomicLong(); // per-peer counter (also the tab #)
private final long spawnReadyTimeoutMs; // 0 = disable gate (legacy non-blocking spawn)
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
private final Runnable sleeper; // sleep/wait hook (injectable for tests; never real-sleep in unit tests)
// Per-process token mixed into each peer name so a fresh process (nameSeq back at 0) cannot
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
/**
* @param namePrefix label prefix for this peer kind (drives naming and reap)
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param profiles configured peer profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (never called when the gate is disabled); the poll
* interval is baked into this hook, so the base needs no poll field
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
this.namePrefix = namePrefix;
this.agents = agents;
this.spaces = spaces;
this.profiles = Map.copyOf(profiles);
this.defaultProfile = defaultProfile;
this.env = env;
this.spawnReadyTimeoutMs = spawnReadyTimeoutMs;
this.nowMillis = nowMillis;
this.sleeper = sleeper;
}
// --- adapter seams -------------------------------------------------------------------------
/**
* Build the peer-specific launch for {@code cfg}: the environment map and argv handed to herdr.
* Any subscription/guard check, MCP mount, and instruction injection happen here. The env map
* and argv are adapter-private; the base only places and starts what is returned.
*/
protected abstract Launch buildLaunch(BridgedConfig.Worker cfg);
/** A peer-specific launch: the herdr {@code env} map and {@code argv}. */
protected record Launch(Map<String, String> env, List<String> argv) {
}
// --- profile surface -----------------------------------------------------------------------
/** The configured peer profile names (what {@code spawn(profile)} accepts). */
@Override
public Set<String> profiles() {
return profiles.keySet();
}
/** The parity-overlay file list for {@code profileName} (default list when unset). */
@Override
public List<String> parityOverlay(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
return List.of();
}
BridgedConfig.Worker cfg = profiles.get(name);
return cfg == null ? List.of() : cfg.parityOverlay();
}
/** The profile a no-argument spawn uses, or {@code null} if none is configured. */
@Override
public String defaultProfile() {
return defaultProfile;
}
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
protected Collection<BridgedConfig.Worker> profileConfigs() {
return profiles.values();
}
/** Resolve {@code profileName} (null/blank → default) to its config, or throw with the options. */
protected BridgedConfig.Worker requireProfile(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
throw new IllegalArgumentException("no default worker profile is configured — "
+ "pass a profile; configured: " + profiles.keySet());
}
BridgedConfig.Worker cfg = profiles.get(name);
if (cfg == null) {
throw new IllegalArgumentException("unknown worker profile '" + name
+ "' — configured: " + profiles.keySet());
}
return cfg;
}
// --- spawn ---------------------------------------------------------------------------------
/**
* Spawn a peer. {@code profileName} null/blank → the default profile. The working directory
* (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd}, else the
* profile's configured {@code cwd}, else {@code callerCwd} (the primary's cwd, when the spawn
* came from the primary over MCP), else the daemon's cwd — never assumed to be {@code $HOME}.
* The adapter's {@link #buildLaunch} runs before any herdr call.
*/
protected Agent spawnInternal(String profileName, String requestedCwd, String callerCwd) {
BridgedConfig.Worker cfg = requireProfile(profileName);
Launch launch = buildLaunch(cfg);
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
return cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd);
}
/**
* {@inheritDoc}
*
* <p>Delegates to {@link #spawnInternal} and wraps the resulting herdr {@link Agent} in a
* {@link WorkerHandle} whose {@link PeerHandle#id()} equals the agent's paneId. When
* {@code spawnReadyTimeoutMs > 0}, blocks until the peer's herdr status is injectable or the
* timeout elapses; on timeout the pane is closed (no orphan) and a
* {@link PeerUnreachableException} is thrown.
*/
@Override
public PeerHandle spawn(SpawnRequest req) {
Agent agent = spawnInternal(req.profileName(), req.requestedCwd(), req.callerCwd());
String paneId = agent.paneId();
if (spawnReadyTimeoutMs > 0) {
waitUntilInjectableOrThrow(paneId);
}
return new WorkerHandle(paneId, agent.terminalId());
}
@Override
public String effectiveCwd(SpawnRequest req) {
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
}
/**
* CB-301: the effective working directory a spawn for {@code profileName} would use, without
* actually spawning.
*/
private String effectiveCwd(String profileName, String requestedCwd, String callerCwd) {
return resolveCwd(requestedCwd, requireProfile(profileName), callerCwd);
}
/**
* CB-112 cwd resolution: spawn arg → profile config → the primary's cwd → the daemon's cwd.
* Never returns {@code null}/blank: {@code "."} (the daemon's own working directory) is the
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
*/
private static String resolveCwd(String requestedCwd, BridgedConfig.Worker cfg, String callerCwd) {
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
}
private static String firstNonBlank(String... values) {
for (String v : values) {
if (v != null && !v.isBlank()) return v;
}
return null;
}
/** Dedicated worker space → own tab → start the peer (rooted at {@code cwd}) → drop the shell. */
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId());
log.info("spawning {} profile={} space={} tab={} cwd={}",
namePrefix, cfg.profile(), space.workspaceId(), tab.tab().tabId(), cwd);
Started started;
try {
started = startUniquelyNamed(cfg, workerEnv, argv, tab.tab().tabId(), cwd);
} catch (RuntimeException e) {
// The peer never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("failed to close orphaned tab {} after spawn error: {}",
tab.tab().tabId(), cleanup.getMessage());
}
throw e;
}
// The peer is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell so the
// tab holds only the peer; label the tab). They must not fail the spawn or orphan the
// running peer — on error we log and still return it so the caller gets its paneId and can
// tear it down.
if (tab.rootPaneId() != null) {
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
} else {
log.warn("tab {} had no seed pane in the create response; peer tab may hold an extra pane",
tab.tab().tabId());
}
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
log.info("{} started pane={} tab={} terminal={}",
namePrefix, started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
return started.agent();
}
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
private void tidy(String what, Runnable step) {
try {
step.run();
} catch (RuntimeException e) {
log.warn("post-start step failed ({}) — peer is running regardless: {}", what, e.getMessage());
}
}
/** Legacy placement: herdr splits the currently-focused tab; the peer still starts in {@code cwd}. */
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
namePrefix, cfg.profile(), cwd, argv);
Agent peer = startUniquelyNamed(cfg, workerEnv, argv, null, cwd).agent();
log.info("{} started pane={} terminal={}", namePrefix, peer.paneId(), peer.terminalId());
return peer;
}
/** A started peer together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
/**
* Start the peer under a unique herdr agent name. herdr requires each running agent's
* {@code name} to be distinct (a 2nd identical {@code name} fails {@code agent_name_taken}) —
* the exact case that makes multiple peers useful. The name is
* {@code <prefix>-<profile>-<nonce>-<seq>}: {@code seq} distinguishes peers within this process,
* and the per-process {@code nonce} keeps a fresh process (whose {@code seq} restarts at 0) from
* colliding with same-profile peers that outlived a restart. The retry is a belt-and-braces
* backstop for the astronomically unlikely nonce+seq clash; the name is a label only — herdr
* detects kind and status from terminal output, not from it.
*/
private Started startUniquelyNamed(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String tabId, String cwd) {
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(agents.start(name, argv, workerEnv, tabId, cwd), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("peer name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
// --- discovery + reap ----------------------------------------------------------------------
/** All herdr-tracked agents — discovery for "what peers exist". */
@Override
public List<Agent> list() {
return agents.list();
}
/**
* Reap peer panes left behind by an earlier daemon process (CB-117). herdr keeps a peer's pane
* alive across a daemon restart <em>by design</em>, and that pane's id is held only by its
* spawner — so a peer whose owning process exited before issuing the matching teardown leaks
* with nothing tracking it. On boot we scan herdr for agents whose name matches our
* {@code <prefix>-<profile>-<nonce>-<seq>} scheme with a nonce <em>other</em> than this
* process's {@link #nameNonce}, and tear each one down (its pane and, via {@link #stop}, its
* now-empty dedicated tab). A current-nonce peer is ours and live, so it is left running; a
* user's own session carries no such name and is never touched. A peer from a <em>different</em>
* adapter (different prefix) is likewise never touched. Best-effort: a failed listing, or a
* failure to stop any one peer, is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
@Override
public int reapOrphanWorkers() {
List<Agent> all;
try {
all = agents.list();
} catch (RuntimeException e) {
log.warn("orphan-peer reap skipped — agent.list failed: {}", e.getMessage());
return 0;
}
int reaped = 0;
for (Agent a : all) {
if (!isForeignWorker(namePrefix, a.name(), nameNonce)) continue;
try {
stop(a.paneId());
reaped++;
log.info("reaped orphan {} {} (pane={} tab={}) left by a prior daemon",
namePrefix, a.name(), a.paneId(), a.tabId());
} catch (RuntimeException e) {
log.warn("could not reap orphan {} {} (pane={}): {}",
namePrefix, a.name(), a.paneId(), e.getMessage());
}
}
if (reaped > 0) {
log.info("orphan-peer reap complete — {} stale {} peer(s) removed at startup", reaped, namePrefix);
}
return reaped;
}
/** The {@code <prefix>-<profile>-<nonce>-<seq>} name pattern; group 1 captures the 6-hex nonce. */
static Pattern workerNamePattern(String prefix) {
return Pattern.compile(prefix + "-.*-([0-9a-f]{6})-\\d+");
}
/**
* Whether {@code name} is a peer of kind {@code prefix} started by a <em>different</em> process
* than {@code currentNonce} — the reap predicate (CB-117). True only for the prefix's naming
* scheme with a foreign nonce: a non-peer name, a different adapter's name, or our own live
* nonce is excluded. Pure and package-private so the decision is unit-testable without herdr.
*/
static boolean isForeignWorker(String prefix, String name, String currentNonce) {
String nonce = workerNonce(prefix, name);
return nonce != null && !nonce.equals(currentNonce);
}
/** The 6-hex nonce embedded in a {@code prefix} peer name, or {@code null} if not one. */
static String workerNonce(String prefix, String name) {
if (name == null) return null;
Matcher m = workerNamePattern(prefix).matcher(name);
return m.matches() ? m.group(1) : null;
}
/** This process's peer-name nonce (a label component only; exposed for reaper tests). */
String nameNonce() {
return nameNonce;
}
// --- teardown ------------------------------------------------------------------------------
/**
* Tear a peer down by pane id: close the pane, and close its tab <em>only</em> when the peer is
* that tab's sole occupant. The single-pane check is what makes this safe regardless of how the
* peer was placed (or a placement-config change across a restart): a pane-placement peer sitting
* in one of the user's shared tabs has siblings, so its tab is never closed — we only ever
* remove a tab we created to hold one peer.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed peer) is treated as success; any other failure propagates so a
* genuinely failed teardown is not reported as done.
*/
@Override
public void stop(String paneId) {
// Teardown knows only the paneId, not which profile spawned it. Attempt tab cleanup when any
// profile uses tab placement (so the bridge may have created a dedicated peer tab); the
// single-occupant check below is what actually protects the user's shared tabs.
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) throw e;
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
spaces.closeTab(loc.tabId());
} else if (loc != null) {
log.debug("not closing tab {} — it holds {} panes (not a dedicated peer tab)",
loc.tabId(), loc.tabPaneCount());
}
}
/** Whether any configured profile places peers in their own tab (so tabs may need cleanup). */
private boolean usesTabPlacement() {
return profiles.values().stream().anyMatch(BridgedConfig.Worker::tabPlacement);
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
// --- spawn-readiness gate (CB-306) ---------------------------------------------------------
/**
* Poll {@link AgentControl#status} until the pane reports an injectable state or the configured
* timeout elapses. On timeout, close the pane (self-reap) and throw.
*/
private void waitUntilInjectableOrThrow(String paneId) {
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
while (nowMillis.getAsLong() < deadline) {
if (agents.status(paneId).injectable()) {
log.debug("peer pane={} reached injectable state", paneId);
return;
}
sleeper.run();
}
log.warn("peer pane={} did not become injectable within {}ms — closing", paneId, spawnReadyTimeoutMs);
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " did not reach injectable state within "
+ spawnReadyTimeoutMs + "ms");
}
/** A concrete {@link PeerHandle} wrapping herdr agent coordinates. */
private record WorkerHandle(String id, String terminalId) implements PeerHandle {
}
// --- shared helpers ------------------------------------------------------------------------
/** Put {@code k → v} only when {@code v} is present (non-null, non-blank). */
protected static void putIfPresent(Map<String, String> m, String k, String v) {
if (v != null && !v.isBlank()) {
m.put(k, v);
}
}
/** Host env lookup that tolerates an unconfigured (null/blank) var name — returns null then. */
protected String resolveEnv(String name) {
return (name == null || name.isBlank()) ? null : env.apply(name);
}
/**
* The parity-neutral git-forge token grant (CB-302): when {@code cfg} opts in via
* {@code gitTokenEnv} and the token resolves, inject {@code GITEA_TOKEN} plus its paired
* {@code GITEA_HOST}. Push over SSH is unaffected; the only incremental grant is PR-create.
* Peer-neutral, so every herdr adapter reuses it unchanged.
*/
protected void applyGitToken(Map<String, String> workerEnv, BridgedConfig.Worker cfg) {
if (!cfg.hasGitToken()) {
return;
}
String gitToken = resolveEnv(cfg.gitTokenEnv());
if (gitToken != null) {
workerEnv.put("GITEA_TOKEN", gitToken);
putIfPresent(workerEnv, "GITEA_HOST", resolveEnv(cfg.gitHostEnv()));
}
}
/** A fresh mutable env map — the conventional starting point for {@link #buildLaunch}. */
/**
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
* {@code env:} entries.
*
* <p>Why this exists: bridged passes herdr an explicit env map, and herdr merges it into
* <em>its own</em> process environment. So before this, a worker inherited whatever PATH the
* herdr server happened to be started with — on this host, one from weeks earlier with no JDK
* and no Maven, which left workers unable to run the build they were being asked to run. The
* worker's toolchain must follow from configuration, not from how a long-lived daemon was
* launched.
*
* <p>Adapter-specific variables are layered on top of this by {@code buildLaunch} and therefore
* win. That ordering is deliberate and load-bearing: it stops a profile's {@code env:} from
* overriding {@code ANTHROPIC_BASE_URL} and slipping past {@link
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
* {@code baseUrl} and nothing else.
*/
protected Map<String, String> baseEnv(BridgedConfig.Worker cfg) {
Map<String, String> workerEnv = new LinkedHashMap<>();
String path = env.apply("PATH");
if (path != null && !path.isBlank()) {
workerEnv.put("PATH", path);
}
if (cfg != null && cfg.env() != null) {
workerEnv.putAll(cfg.env());
}
return workerEnv;
}
/** Defensive copy of {@code argv} plus room to append launch flags. */
protected static List<String> mutableArgv(List<String> argv) {
return new ArrayList<>(argv);
}
/**
* Uninterruptible sleep — the production {@link #sleeper}. Tests supply their own no-op /
* fast-faking sleeper so they never real-sleep.
*/
protected static void sleepUninterruptibly(long ms) {
try {
Thread.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
// preserve the interrupt flag but continue — poll loops should not be aborted by an
// interrupt that was not meant for them.
}
}
}
@@ -0,0 +1,316 @@
package dev.ltms.bridged.worker;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.node.ObjectNode;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>opencode</strong> — an open-source,
* provider-agnostic terminal coding agent. Its whole reason for existing is to prove the
* {@code PeerLauncher} SPI is genuinely provider-neutral: opencode shares none of Claude Code's
* private launch seams, yet reuses every line of shared transport in the base (tab/pane placement,
* the CB-306 readiness gate, unique naming + CB-117 reap, teardown, listing, cwd).
*
* <p>The divergences from {@link ClaudeCodeLauncher}, all confined to {@link #buildLaunch}:
* <ul>
* <li><strong>No subscription boundary.</strong> opencode carries no {@code ANTHROPIC_BASE_URL}
* and there is no {@link dev.ltms.bridged.guard.SubscriptionGuard} — the guard is a
* Claude-private concern, not part of the SPI. opencode reads the operator's own provider
* credentials from its global {@code auth.json}; the bridge injects none.</li>
* <li><strong>File-based MCP mount + instructions.</strong> opencode has no inline
* {@code --mcp-config}/{@code --append-system-prompt}. Instead the bridge writes an ephemeral
* {@code opencode.json} that declares the bridge as a {@code remote} MCP server and lists a
* reply-charter file under {@code instructions}, then points the worker at it with
* {@code OPENCODE_CONFIG}. This is the one place the launcher touches disk — Claude never did.</li>
* <li><strong>Model as a flag.</strong> the {@code provider/model} selector is passed as
* {@code -m}, not an env var.</li>
* <li><strong>{@code opencode} name prefix</strong> so reap matches {@code opencode-*} panes and
* never another adapter's.</li>
* </ul>
*/
public final class OpenCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "opencode";
/** Writer for the generated {@code opencode.json}. */
private static final ObjectMapper JSON = new ObjectMapper();
/**
* Standing instruction written to the charter file and mounted via the config's
* {@code instructions} so the worker returns its result through {@code bridge_reply}. Kept on
* disk (not a launch flag) because opencode's {@code instructions} takes file paths, not inline
* text — the file is regenerated per spawn and never touches the worker's own profile.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet, running under opencode. "
+ "Every message you receive arrives through the bridge, and the ONLY channel back to the "
+ "sender is the bridge_reply MCP tool. Text you write in your terminal is NOT sent "
+ "anywhere — the sender cannot see your screen, so an in-terminal answer is silently "
+ "discarded. Therefore you MUST end EVERY turn by calling bridge_reply with `content` set "
+ "to your complete response. This holds for every message without exception — tasks, "
+ "questions, clarifications, acknowledgements, and ordinary back-and-forth conversation. "
+ "Call bridge_reply exactly once, as the final action of your turn, with your full answer "
+ "in `content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/** Root under which per-spawn opencode config dirs are created (injectable for tests). */
private final Path configRoot;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so it
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300),
defaultConfigRoot());
}
/**
* Production constructor with the spawn-ready gate enabled. Polls {@code agents.status()} until
* the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot());
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply a
* fake clock ({@code nowMillis}), poll-loop wait ({@code sleeper}), and a temp {@code configRoot}
* they can inspect the generated {@code opencode.json}/charter under.
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (encodes the poll interval; never called when the
* gate is disabled)
* @param configRoot existing directory under which per-spawn config dirs are created
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper, Path configRoot) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.configRoot = configRoot;
}
private static Path defaultConfigRoot() {
return Path.of(System.getProperty("java.io.tmpdir"));
}
/**
* {@inheritDoc}
*
* <p>Builds the opencode launch: no {@code ANTHROPIC_*} and no guard (opencode reads its own
* provider credentials); when the profile mounts the bridge MCP, generate an ephemeral
* {@code opencode.json} (remote MCP server + reply-charter instructions) and point the worker at
* it via {@code OPENCODE_CONFIG}; carry the parity-neutral git-forge grant; and select the model
* with {@code -m}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
Map<String, String> workerEnv = baseEnv(cfg);
// A config file is needed for the bridge MCP mount, for a pinned endpoint (CB-508), or both.
if (cfg.hasMcp() || hasCustomProvider(cfg)) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg).toString());
}
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithModel(cfg));
}
/**
* True when this profile pins its own OpenAI-compatible endpoint (CB-508) rather than using
* whatever provider opencode resolves by default.
*
* <p>Note this reuses {@code baseUrl}, the same field the Claude adapter injects as
* {@code ANTHROPIC_BASE_URL} — but it does <em>not</em> go through {@code SubscriptionGuard}.
* That asymmetry is deliberate and safe: the guard exists to stop a worker borrowing the
* primary's Anthropic subscription, and an opencode process has no Anthropic credential path
* at all. Pointing it at a local vLLM cannot leak the subscription.
*/
private static boolean hasCustomProvider(BridgedConfig.Worker cfg) {
return cfg.baseUrl() != null && !cfg.baseUrl().isBlank();
}
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
private List<String> argvWithModel(BridgedConfig.Worker cfg) {
List<String> argv = mutableArgv(cfg.argv());
if (cfg.model() != null && !cfg.model().isBlank()) {
argv.add("-m");
argv.add(cfg.model());
}
return argv;
}
/**
* Write an ephemeral {@code opencode.json} (and the reply-charter file it references) into a
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
*/
private Path writeConfig(BridgedConfig.Worker cfg) {
try {
Path dir = Files.createTempDirectory(configRoot, "bridged-opencode-");
dir.toFile().deleteOnExit();
ObjectNode root = JSON.createObjectNode();
root.put("$schema", "https://opencode.ai/config.json");
if (cfg.hasMcp()) {
Path charter = dir.resolve("reply-charter.md");
Files.writeString(charter, REPLY_CHARTER);
charter.toFile().deleteOnExit();
ObjectNode bridge = root.putObject("mcp").putObject("bridge");
bridge.put("type", "remote");
bridge.put("url", cfg.mcpUrl());
bridge.put("enabled", true);
root.putArray("instructions").add(charter.toAbsolutePath().toString());
}
if (hasCustomProvider(cfg)) {
addCustomProvider(root, cfg);
}
Path cfgFile = dir.resolve("opencode.json");
// Built with Jackson rather than string concatenation: the provider block is nested and
// carries operator-supplied values (URL, model id, api key), so escaping must be real.
Files.writeString(cfgFile, JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root));
cfgFile.toFile().deleteOnExit();
return cfgFile;
} catch (IOException e) {
throw new UncheckedIOException(
"cannot write opencode config for profile " + cfg.profile(), e);
}
}
/**
* Declare a custom OpenAI-compatible provider so the worker talks to a pinned endpoint (a local
* vLLM, say) instead of opencode's default gateway (CB-508).
*
* <p>The provider id comes from the {@code provider/model} selector in {@code model:}, so one
* field drives both the declaration and the {@code -m} flag and they cannot drift apart.
*/
private void addCustomProvider(ObjectNode root, BridgedConfig.Worker cfg) {
String[] parts = splitModelSelector(cfg);
String providerId = parts[0];
String modelId = parts[1];
ObjectNode provider = root.putObject("provider").putObject(providerId);
provider.put("npm", "@ai-sdk/openai-compatible");
provider.put("name", providerId + " (bridged)");
ObjectNode options = provider.putObject("options");
options.put("baseURL", openAiBaseUrl(cfg.baseUrl()));
// vLLM and friends usually ignore the key, but the AI SDK still requires a non-empty one.
String token = resolveEnv(cfg.tokenEnv());
options.put("apiKey", (token == null || token.isBlank()) ? "bridged-local-noauth" : token);
provider.putObject("models").putObject(modelId).put("name", modelId);
}
/**
* Split {@code model:} into its {@code provider} and {@code model} halves. A pinned endpoint
* needs both, so a bare model name is rejected loudly rather than silently falling back to the
* default gateway — a worker quietly talking to the wrong endpoint is the failure this avoids.
*/
private static String[] splitModelSelector(BridgedConfig.Worker cfg) {
String model = cfg.model();
int slash = model == null ? -1 : model.indexOf('/');
if (model == null || model.isBlank() || slash <= 0 || slash == model.length() - 1) {
throw new IllegalArgumentException(
"profile " + cfg.profile() + " sets baseUrl (a pinned opencode endpoint) so"
+ " model: must be \"<provider>/<model>\", e.g."
+ " \"local-vllm/deepseek-v4-flash\"; got "
+ (model == null ? "null" : '"' + model + '"'));
}
return new String[]{model.substring(0, slash), model.substring(slash + 1)};
}
/**
* The OpenAI-compatible base URL for {@code baseUrl}. A bare {@code host:port} gets {@code /v1}
* appended (where these servers put the API); a URL that already carries a path is taken as-is,
* so an endpoint mounted somewhere unusual is still reachable.
*/
private static String openAiBaseUrl(String baseUrl) {
String trimmed = baseUrl.trim();
while (trimmed.endsWith("/")) {
trimmed = trimmed.substring(0, trimmed.length() - 1);
}
int schemeEnd = trimmed.indexOf("://");
String afterScheme = schemeEnd < 0 ? trimmed : trimmed.substring(schemeEnd + 3);
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
// --- CB-117 reap predicate (opencode prefix), kept for direct unit testing -----------------
/**
* Whether {@code name} is an opencode bridge worker started by a <em>different</em> process than
* {@code currentNonce}. A thin {@code opencode}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -1,206 +0,0 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
/**
* Spawns and lists worker sessions — the safe path from a delegation request to a
* running off-subscription Claude.
*
* <p>The spawn sequence encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em>
* touching herdr, and only then {@code agent.start}. A worker's base_url lives in the
* env map handed to herdr and nowhere else; {@code bridged}'s own environment is never
* mutated.
*
* <p>Placement: in the default {@code tab} policy a worker lands in its own tab inside a
* dedicated worker space (found-or-created once, then shared), so workers never split or
* clutter the user's real work spaces. Teardown removes the worker's pane <em>and</em> its
* now-empty tab, tolerating an already-gone worker so a repeated DELETE is harmless.
*/
public final class WorkerService {
private static final Logger log = LoggerFactory.getLogger(WorkerService.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
private final AgentControl agents;
private final WorkspaceControl spaces;
private final SubscriptionGuard guard;
private final BridgedConfig.Worker cfg;
private final Function<String, String> env; // host env lookup (injectable for tests)
private final AtomicLong nameSeq = new AtomicLong(); // per-worker counter (also the tab #)
// Per-process token mixed into each worker name so a fresh process (nameSeq back at 0)
// cannot collide with same-profile workers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
public WorkerService(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
BridgedConfig.Worker cfg, Function<String, String> env) {
this.agents = agents;
this.spaces = spaces;
this.guard = guard;
this.cfg = cfg;
this.env = env;
}
/** Spawn a worker for the configured profile. Guard runs before any herdr call. */
public Agent spawn() {
String baseUrl = cfg.baseUrl();
guard.assertWorker(baseUrl); // hard stop before we spawn anything
Map<String, String> workerEnv = new LinkedHashMap<>();
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
String token = env.apply(cfg.tokenEnv());
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", token);
return cfg.tabPlacement() ? spawnInTab(workerEnv) : spawnAsPane(workerEnv);
}
/** Dedicated worker space → own tab → drop the placeholder shell so only the worker remains. */
private Agent spawnInTab(Map<String, String> workerEnv) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId());
log.info("spawning worker profile={} base_url={} space={} tab={}",
cfg.profile(), cfg.baseUrl(), space.workspaceId(), tab.tab().tabId());
Started started;
try {
started = startUniquelyNamed(workerEnv, tab.tab().tabId());
} catch (RuntimeException e) {
// The worker never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("failed to close orphaned tab {} after spawn error: {}",
tab.tab().tabId(), cleanup.getMessage());
}
throw e;
}
// The worker is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell
// so the tab holds only the worker; label the tab). They must not fail the spawn or
// orphan the running worker — on error we log and still return it so the caller gets
// its paneId and can tear it down.
if (tab.rootPaneId() != null) {
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
} else {
log.warn("tab {} had no seed pane in the create response; worker tab may hold an extra pane",
tab.tab().tabId());
}
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
log.info("worker started pane={} tab={} terminal={}",
started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
return started.agent();
}
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
private void tidy(String what, Runnable step) {
try {
step.run();
} catch (RuntimeException e) {
log.warn("post-start step failed ({}) — worker is running regardless: {}", what, e.getMessage());
}
}
/** Legacy placement: herdr splits the currently-focused tab. */
private Agent spawnAsPane(Map<String, String> workerEnv) {
log.info("spawning worker (pane placement) profile={} base_url={} argv={}",
cfg.profile(), cfg.baseUrl(), cfg.argv());
Agent worker = startUniquelyNamed(workerEnv, null).agent();
log.info("worker started pane={} terminal={}", worker.paneId(), worker.terminalId());
return worker;
}
/** A started worker together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
/**
* Start the worker under a unique herdr agent name. herdr requires each running
* agent's {@code name} to be distinct (a 2nd {@code name:"claude"} fails
* {@code agent_name_taken}) — the exact case that makes multiple workers useful. The name
* is {@code claude-<profile>-<nonce>-<seq>}: {@code seq} distinguishes workers within this
* process, and the per-process {@code nonce} keeps a fresh process (whose {@code seq}
* restarts at 0) from colliding with same-profile workers that outlived a restart. The
* retry is a belt-and-braces backstop for the astronomically unlikely nonce+seq clash;
* the name is a label only — herdr detects kind and status from terminal output, not it.
*/
private Started startUniquelyNamed(Map<String, String> workerEnv, String tabId) {
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = "claude-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(agents.start(name, cfg.argv(), workerEnv, tabId), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("worker name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
/** All herdr-tracked agents — discovery for "what workers exist". */
public List<Agent> list() {
return agents.list();
}
/**
* Tear a worker down by pane id: close the pane, and close its tab <em>only</em> when the
* worker is that tab's sole occupant. The single-pane check is what makes this safe
* regardless of how the worker was placed (or a placement-config change across a restart):
* a pane-placement worker sitting in one of the user's shared tabs has siblings, so its
* tab is never closed — we only ever remove a tab we created to hold one worker.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed worker) is treated as success; any other failure propagates so
* a genuinely failed teardown is not reported as done.
*/
public void stop(String paneId) {
WorkspaceControl.PaneLocation loc = cfg.tabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) throw e;
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
spaces.closeTab(loc.tabId());
} else if (loc != null) {
log.debug("not closing tab {} — it holds {} panes (not a dedicated worker tab)",
loc.tabId(), loc.tabPaneCount());
}
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
private static void putIfPresent(Map<String, String> m, String k, String v) {
if (v != null && !v.isBlank()) {
m.put(k, v);
}
}
}
+29
View File
@@ -5,10 +5,39 @@
</encoder>
</appender>
<!--
CB-505 audit trail. Its own file, deliberately not the app log: privileged actions
(spawn/stop/send/reply/drain) must stay greppable and shippable without dragging DEBUG noise
along. AuditLog emits a complete JSON object including its own ISO-8601 "ts" field, so the
pattern is a bare %msg — a pattern that spliced literal braces around the message would
collide with logback's own variable substitution. Rolls daily, 30 days retained, 100MB cap.
NOTE: records carry who/what/target/outcome only. Message CONTENT is never written here —
this bridge carries source code and prompts, and an audit log that accumulated them would be
a transcript archive rather than a control.
-->
<appender name="AUDIT" class="ch.qos.logback.core.rolling.RollingFileAppender">
<file>logs/audit.log</file>
<rollingPolicy class="ch.qos.logback.core.rolling.SizeAndTimeBasedRollingPolicy">
<fileNamePattern>logs/audit.%d{yyyy-MM-dd}.%i.log</fileNamePattern>
<maxFileSize>10MB</maxFileSize>
<maxHistory>30</maxHistory>
<totalSizeCap>100MB</totalSizeCap>
</rollingPolicy>
<encoder>
<pattern>%msg%n</pattern>
</encoder>
</appender>
<logger name="dev.ltms.bridged" level="DEBUG"/>
<logger name="io.javalin" level="INFO"/>
<logger name="org.eclipse.jetty" level="WARN"/>
<!-- additivity=false keeps the audit stream out of stdout; it is its own record. -->
<logger name="audit" level="INFO" additivity="false">
<appender-ref ref="AUDIT"/>
</logger>
<root level="INFO">
<appender-ref ref="STDOUT"/>
</root>
@@ -0,0 +1,100 @@
package dev.ltms.bridged.auth;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-505 — the audit record's shape.
*
* <p>These exist because the first cut of this feature emitted lines that were <em>not</em> valid
* JSON: the timestamp was spliced on by a logback pattern whose literal braces collided with
* logback's variable substitution. The appender failed to parse, and nothing in the build noticed.
* An audit trail that silently stops being machine-readable is worse than none.
*/
class AuditLogTest {
private final ObjectMapper mapper = new ObjectMapper();
private ListAppender<ILoggingEvent> appender;
private ch.qos.logback.classic.Logger auditLogger;
@BeforeEach
void attach() {
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
auditLogger = ctx.getLogger("audit");
appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
auditLogger.addAppender(appender);
auditLogger.setLevel(Level.INFO);
}
@AfterEach
void detach() {
auditLogger.detachAppender(appender);
}
private JsonNode onlyRecord() throws Exception {
assertEquals(1, appender.list.size(), "exactly one audit line expected");
String line = appender.list.getFirst().getFormattedMessage();
return mapper.readTree(line); // throws if the line is not valid JSON
}
@Test
void anAllowedActionIsRecordedAsValidJson() throws Exception {
AuditLog.allowed(Principal.primary(4242), Authz.Action.SPAWN, "term_a");
JsonNode r = onlyRecord();
assertEquals("PRIMARY", r.path("role").asText());
assertEquals("primary", r.path("actor").asText());
assertEquals(4242, r.path("pid").asLong());
assertEquals("SPAWN", r.path("action").asText());
assertEquals("term_a", r.path("target").asText());
assertEquals("allowed", r.path("outcome").asText());
assertFalse(r.path("ts").asText().isBlank(), "every record carries its own timestamp");
}
@Test
void aDenialRecordsTheReason() throws Exception {
AuditLog.denied(Principal.worker("term_b", 7), Authz.Action.REPLY, "term_a", "forbidden");
JsonNode r = onlyRecord();
assertEquals("WORKER", r.path("role").asText());
assertEquals("worker:term_b", r.path("actor").asText());
assertEquals("denied", r.path("outcome").asText());
assertEquals("forbidden", r.path("reason").asText());
}
@Test
void aNullCallerIsRecordedAsAnonymousRatherThanCrashing() throws Exception {
AuditLog.failed(null, Authz.Action.SEND, null, "herdr unreachable");
JsonNode r = onlyRecord();
assertEquals("ANONYMOUS", r.path("role").asText());
assertTrue(r.path("target").isNull(), "an absent target is JSON null, not the string \"null\"");
assertEquals("failed", r.path("outcome").asText());
}
@Test
void hostileValuesAreEscapedAndCannotForgeAnExtraRecord() throws Exception {
// A target id containing a quote and a newline must not be able to terminate the JSON
// object early and inject a second, attacker-shaped audit line.
AuditLog.denied(Principal.worker("term_a", 1), Authz.Action.REPLY,
"evil\",\"outcome\":\"allowed\"}\n{\"forged\":true", "forbidden");
JsonNode r = onlyRecord();
assertEquals("denied", r.path("outcome").asText(),
"the injected outcome must not override the real one");
assertTrue(r.path("target").asText().contains("forged"),
"the hostile text survives as inert data inside the target field");
}
}
@@ -0,0 +1,80 @@
package dev.ltms.bridged.auth;
import org.junit.jupiter.api.Test;
import static dev.ltms.bridged.auth.Authz.Action.*;
import static org.junit.jupiter.api.Assertions.*;
/** CB-505 — the authorization table, pinned so it cannot drift silently. */
class AuthzTest {
private static final Principal PRIMARY = Principal.primary(100);
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal WORKER_B = Principal.worker("term_b", 300);
private static final Principal ANON = Principal.anonymous();
@Test
void anonymousIsAuthorizedForNothing() {
for (Authz.Action a : Authz.Action.values()) {
assertFalse(Authz.permits(ANON, a, "term_a"),
a + " must be refused to an unauthenticated caller");
}
}
@Test
void aNullCallerIsTreatedAsAnonymous() {
assertFalse(Authz.permits(null, READ, null));
assertTrue(Authz.isUnauthenticated(null));
}
@Test
void orchestrationBelongsToThePrimaryAlone() {
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, SEND, DRAIN}) {
assertTrue(Authz.permits(PRIMARY, a, "term_a"), "the primary orchestrates: " + a);
assertFalse(Authz.permits(WORKER_A, a, "term_a"),
"a worker performing " + a + " would be escalating into the orchestrator role");
}
}
@Test
void aWorkerMayReplyAndAskOnlyAsItself() {
assertTrue(Authz.permits(WORKER_A, REPLY, "term_a"));
assertTrue(Authz.permits(WORKER_A, ASK, "term_a"));
assertFalse(Authz.permits(WORKER_A, REPLY, "term_b"),
"worker A must not be able to reply on worker B's session");
assertFalse(Authz.permits(WORKER_B, ASK, "term_a"),
"worker B must not be able to ask as worker A");
}
@Test
void thePrimaryMayNotForgeAWorkersReply() {
// Not a hypothetical nicety: a forged reply would resolve the rendezvous the primary is
// itself blocked on, corrupting the correlation between a turn and its answer.
assertFalse(Authz.permits(PRIMARY, REPLY, "term_a"));
assertFalse(Authz.permits(PRIMARY, ASK, "term_a"));
}
@Test
void aWorkerWithNoTargetCannotReply() {
assertFalse(Authz.permits(WORKER_A, REPLY, null),
"an absent session id must not satisfy the own-session rule");
}
@Test
void observationIsOpenToBothAuthenticatedRoles() {
assertTrue(Authz.permits(PRIMARY, READ, null));
assertTrue(Authz.permits(WORKER_A, READ, null));
assertTrue(Authz.permits(PRIMARY, METRICS, null));
assertTrue(Authz.permits(WORKER_A, METRICS, null));
}
@Test
void unauthenticatedIsDistinguishedFromMerelyForbidden() {
// Drives the 401-vs-403 split: a missing credential is fixable by the caller, a wrong role
// is not.
assertTrue(Authz.isUnauthenticated(ANON));
assertFalse(Authz.isUnauthenticated(WORKER_A));
assertFalse(Authz.isUnauthenticated(PRIMARY));
}
}
@@ -0,0 +1,106 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-501. The behaviour under test is the inversion of the pre-CB-501 default: failing every
* identity check must yield {@link Role#ANONYMOUS}, not {@code PRIMARY}.
*/
class CallerResolverTest {
private final FakeHerdr herdr = new FakeHerdr();
/** Identity resolving the canned worker pane, keyed off a faked peer-PID lookup. */
private ConnectionIdentity identity(long pid) {
return new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
}
/** A PID that owns a worker pane in the fake. */
private ConnectionIdentity workerIdentity() {
return identity(FakeHerdr.WORKER_PID);
}
/** A PID that owns no pane — i.e. the primary, or any other local process. */
private ConnectionIdentity nonWorkerIdentity() {
return identity(999_999);
}
@Test
void aLoopbackWorkerPaneResolvesToWorkerRegardlessOfAuthMode() {
Principal underTrust = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
Principal underToken = new CallerResolver(workerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, underTrust.role());
assertEquals("term_a", underTrust.terminal());
assertEquals(Role.WORKER, underToken.role(),
"worker identity is unforgeable and must never be token-gated — otherwise enabling "
+ "auth would lock the whole fleet out of bridge_reply");
assertEquals("term_a", underToken.terminal());
}
@Test
void loopbackTrustTreatsANonWorkerLoopbackCallerAsThePrimary() {
Principal p = new CallerResolver(nonWorkerIdentity()).resolve("127.0.0.1", 99, null);
assertEquals(Role.PRIMARY, p.role(), "the historical behaviour, now an explicit choice");
}
@Test
void tokenModeRefusesANonWorkerCallerThatPresentsNoToken() {
Principal p = new CallerResolver(nonWorkerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 99, null);
assertEquals(Role.ANONYMOUS, p.role(),
"no credential must mean NOTHING, not the most privileged role on the bus");
}
@Test
void tokenModeAcceptsAValidBearerTokenAsThePrimary() {
Principal p = new CallerResolver(nonWorkerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 99, "Bearer s3cret");
assertEquals(Role.PRIMARY, p.role());
}
@Test
void tokenModeRejectsAWrongOrMalformedCredential() {
CallerResolver r = new CallerResolver(nonWorkerIdentity(), true, "s3cret");
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "Bearer wrong").role());
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "s3cret").role(), "scheme required");
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "Bearer ").role(), "empty credential");
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "Basic s3cret").role(), "wrong scheme");
}
@Test
void theBearerSchemeIsCaseInsensitivePerRfc7235() {
CallerResolver r = new CallerResolver(nonWorkerIdentity(), true, "s3cret");
assertEquals(Role.PRIMARY, r.resolve("127.0.0.1", 99, "bearer s3cret").role());
assertEquals(Role.PRIMARY, r.resolve("127.0.0.1", 99, "BEARER s3cret").role());
}
@Test
void aNonLoopbackCallerIsNeverThePrimaryUnderLoopbackTrust() {
// Defence in depth: startup already refuses this pairing (validateAuthExposure), but if a
// proxy ever forwards a remote peer onto the loopback listener, the resolver must not
// hand it the primary role.
Principal p = new CallerResolver(nonWorkerIdentity()).resolve("10.0.0.7", 99, null);
assertEquals(Role.ANONYMOUS, p.role());
}
@Test
void tokenModeRequiresANonEmptyConfiguredToken() {
ConnectionIdentity id = nonWorkerIdentity();
assertThrows(IllegalArgumentException.class, () -> new CallerResolver(id, true, null));
assertThrows(IllegalArgumentException.class, () -> new CallerResolver(id, true, " "));
}
}
@@ -5,6 +5,7 @@ import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
@@ -46,10 +47,322 @@ class BridgedConfigTest {
assertTrue(cfg.guard().offSubscriptionHosts().isEmpty());
}
@Test
void singleWorkerBecomesAOneEntryProfileMapWithItselfAsDefault(@TempDir Path dir) throws Exception {
Path f = dir.resolve("single.yaml");
Files.writeString(f, """
worker:
profile: ltms-local
baseUrl: http://gx00.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(Set.of("ltms-local"), cfg.workerProfiles().keySet(), "legacy worker → one profile");
assertEquals("ltms-local", cfg.defaultProfile());
}
@Test
void loadsMultipleWorkerProfilesWithADefault(@TempDir Path dir) throws Exception {
Path f = dir.resolve("multi.yaml");
Files.writeString(f, """
workers:
gx10:
baseUrl: http://gx10.gw:8000
argv: ["ccs", "gx10"]
ollama:
baseUrl: http://ollama.ltms.dev
argv: ["ccs", "ollama"]
defaultWorker: gx10
guard:
offSubscriptionHosts: [gx10.gw, ollama.ltms.dev]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(Set.of("gx10", "ollama"), cfg.workerProfiles().keySet());
assertEquals("gx10", cfg.defaultProfile());
assertEquals("ollama", cfg.workerProfiles().get("ollama").profile(), "profile defaults to its map key");
assertEquals("http://gx10.gw:8000", cfg.workerProfiles().get("gx10").baseUrl());
}
@Test
void ignoresUnknownKeys(@TempDir Path dir) throws Exception {
Path f = dir.resolve("future.yaml");
Files.writeString(f, "bind:\n port: 8080\nfutureFeature:\n enabled: true\n");
assertDoesNotThrow(() -> BridgedConfig.load(f));
}
@Test
void absentBrokerBlockLeavesInboxSoftState(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-broker.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNull(cfg.broker(), "no broker: block → null → in-memory inbox is selected");
}
@Test
void brokerBlockWithUriEnablesAmqpAdapter(@TempDir Path dir) throws Exception {
Path f = dir.resolve("broker.yaml");
Files.writeString(f, """
bind:
port: 8080
broker:
uri: amqp://guest:guest@127.0.0.1:5672/
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.broker());
assertTrue(cfg.broker().isConfigured(), "a non-blank uri enables the AMQP adapter");
assertEquals("amqp://guest:guest@127.0.0.1:5672/", cfg.broker().uri());
}
@Test
void brokerBlockWithBlankUriStaysSoftState(@TempDir Path dir) throws Exception {
Path f = dir.resolve("broker-blank.yaml");
Files.writeString(f, "bind:\n port: 8080\nbroker:\n uri: \"\"\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.broker());
assertFalse(cfg.broker().isConfigured(), "an empty uri must not enable AMQP");
}
@Test
void absentPrimaryBlockLeavesPrimaryNull(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-primary.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNull(cfg.primary(), "no primary: block → null → connection-derived identity");
}
@Test
void primaryBlockWithTerminalPinsIdentity(@TempDir Path dir) throws Exception {
Path f = dir.resolve("primary-pinned.yaml");
Files.writeString(f, """
bind:
port: 8080
primary:
terminal: term_fixed
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.primary());
assertEquals("term_fixed", cfg.primary().terminal());
}
@Test
void primaryBlockWithBlankTerminalDefaultsToDerived(@TempDir Path dir) throws Exception {
Path f = dir.resolve("primary-blank.yaml");
Files.writeString(f, "bind:\n port: 8080\nprimary:\n terminal: \"\"\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.primary());
assertTrue(cfg.primary().terminal() == null || cfg.primary().terminal().isBlank(),
"a blank terminal in yaml should be treated as absent — null or empty are equivalent");
}
// --- CB-402: peer kind discriminator -------------------------------------------------------
@Test
void workerKindDefaultsToClaudeCodeWhenOmitted(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-absent.yaml");
Files.writeString(f, """
workers:
gx10:
baseUrl: http://gx10.gw:8000
argv: ["ccs", "gx10"]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(BridgedConfig.Worker.KIND_CLAUDE_CODE, cfg.workerProfiles().get("gx10").kind(),
"a worker with no kind: is a claude-code worker (backward compatible)");
}
@Test
void opencodeKindIsNormalizedToLowerCase(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-opencode.yaml");
Files.writeString(f, """
workers:
gemini:
kind: OpenCode
model: google/gemini-2.5-pro
argv: ["opencode"]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(BridgedConfig.Worker.KIND_OPENCODE, cfg.workerProfiles().get("gemini").kind(),
"kind is normalised to lower-case so YAML casing does not matter");
}
@Test
void kindPredicatesReflectTheResolvedKind(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-predicates.yaml");
Files.writeString(f, """
workers:
claude:
baseUrl: http://gx10.gw:8000
gemini:
kind: opencode
model: google/gemini-2.5-pro
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Worker claude = cfg.workerProfiles().get("claude");
BridgedConfig.Worker gemini = cfg.workerProfiles().get("gemini");
assertTrue(claude.isClaudeCode(), "the default-kind worker is claude-code");
assertFalse(claude.isOpenCode(), "a claude-code worker is not opencode");
assertTrue(gemini.isOpenCode(), "the kind: opencode worker is opencode");
assertFalse(gemini.isClaudeCode(), "an opencode worker is not claude-code");
}
@Test
void argvDefaultsToTheKindBinaryWhenUnset(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-argv.yaml");
Files.writeString(f, """
workers:
claude:
baseUrl: http://gx10.gw:8000
gemini:
kind: opencode
model: google/gemini-2.5-pro
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(java.util.List.of("claude"), cfg.workerProfiles().get("claude").argv(),
"a claude-code worker with no argv defaults to the claude binary");
assertEquals(java.util.List.of("opencode"), cfg.workerProfiles().get("gemini").argv(),
"an opencode worker with no argv defaults to the opencode binary, never claude");
}
@Test
void authDefaultsToLoopbackTrustSoExistingConfigsBehaveAsBefore(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-auth-block.yaml");
Files.writeString(f, "bind:\n host: 127.0.0.1\n port: 8765\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.auth(), "auth must default rather than be null");
assertFalse(cfg.auth().tokenMode());
assertEquals("BRIDGED_API_TOKEN", cfg.auth().tokenEnv(), "documented default env var");
assertDoesNotThrow(cfg::validateAuthExposure, "loopback + loopback-trust is the safe pairing");
}
/**
* CB-501's highest-value check. Under loopback-trust, "not a known worker" means "the primary" —
* sound only while the OS refuses remote connections. Widening the bind without token mode
* would silently promote every reachable client to the most privileged role on the bus.
*/
@Test
void aNonLoopbackBindWithoutTokenModeIsRefusedAtStartup(@TempDir Path dir) throws Exception {
Path f = dir.resolve("exposed.yaml");
Files.writeString(f, "bind:\n host: 0.0.0.0\n port: 8765\n");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAuthExposure);
assertTrue(e.getMessage().contains("auth.mode: token"),
"the error must say how to fix it, not just that it refused");
}
@Test
void aNonLoopbackBindIsAllowedOnceTokenModeIsOn(@TempDir Path dir) throws Exception {
Path f = dir.resolve("exposed-with-token.yaml");
Files.writeString(f, """
bind:
host: 0.0.0.0
port: 8765
auth:
mode: token
tokenEnv: MY_TOKEN
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertTrue(cfg.auth().tokenMode());
assertEquals("MY_TOKEN", cfg.auth().tokenEnv());
assertDoesNotThrow(cfg::validateAuthExposure);
}
@Test
void loopbackFormsAreAllRecognised(@TempDir Path dir) throws Exception {
for (String host : new String[]{"127.0.0.1", "localhost", "::1", "127.0.0.53"}) {
Path f = dir.resolve("lb-" + host.replace(':', '_') + ".yaml");
Files.writeString(f, "bind:\n host: \"" + host + "\"\n port: 8765\n");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateAuthExposure(),
host + " is loopback and must not trip the exposure guard");
}
}
/**
* The shipped {@code bridged.example.yaml} must actually parse. Config binds through a plain
* Jackson mapper with {@code ignoreUnknown = true}, so a misspelled key in the example is
* silently dropped and the operator gets a default they did not ask for — exactly how a
* {@code spawn_ready_timeout_ms} typo survived in the example until the CB-5xx wrap-up.
*/
@Test
void shippedExampleConfigParses() {
Path example = Path.of("bridged.example.yaml");
assertTrue(Files.exists(example), "bridged.example.yaml must ship next to the pom");
BridgedConfig cfg = BridgedConfig.load(example);
assertEquals(8765, cfg.bind().port(), "example binds the documented default port");
assertTrue(cfg.workerProfiles().containsKey("gx10"), "example documents the gx10 profile");
assertEquals("gx10", cfg.defaultProfile(), "example's defaultWorker resolves");
assertTrue(cfg.guard().hostSet().contains("gx01.gw"),
"every example profile's base_url host must be in the example allowlist");
}
/**
* Every optional knob the example documents must bind under the exact spelling used there.
* Keep this list in step with {@code bridged.example.yaml}: a rename that updates the record
* but not the example (or vice versa) fails here instead of silently no-op'ing in production.
*/
@Test
void everyOptionalKnobDocumentedInTheExampleBinds(@TempDir Path dir) throws Exception {
Path f = dir.resolve("all-knobs.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
spawnReadyTimeoutMs: 25000
spawnReadyPollMs: 400
worktreeRoot: /tmp/bridged-worktrees
workers:
gx10:
kind: claude-code
baseUrl: http://gx01.gw:8000
configDir: /tmp/ccs/gx10
cwd: /tmp/repo
parityOverlay: [".mcp.json", ".env"]
gitTokenEnv: GITEA_TOKEN
gitHostEnv: GITEA_HOST
lifecycle:
idleTtlSeconds: 300
contextCap: 10
drainTimeoutSeconds: 5
broker:
uri: amqp://guest:guest@127.0.0.1:5672
primary:
terminal: term_abc123
pushReminders: 5
pushBackoffMs: 15000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(25000, cfg.spawnReadyTimeoutMs(), "spawnReadyTimeoutMs is camelCase, not snake_case");
assertEquals(400, cfg.spawnReadyPollMs(), "spawnReadyPollMs is camelCase, not snake_case");
assertEquals("/tmp/bridged-worktrees", cfg.worktreeRoot());
BridgedConfig.Worker w = cfg.workerProfiles().get("gx10");
assertEquals("/tmp/ccs/gx10", w.configDir());
assertEquals("/tmp/repo", w.cwd());
assertEquals(java.util.List.of(".mcp.json", ".env"), w.parityOverlay());
assertTrue(w.hasGitToken(), "gitTokenEnv binds and enables the CB-302 PR grant");
assertEquals("GITEA_HOST", w.gitHostEnv());
assertEquals(300, cfg.lifecycle().idleTtlSeconds());
assertEquals(10, cfg.lifecycle().contextCap());
assertEquals(5, cfg.lifecycle().drainTimeoutSeconds());
assertEquals("amqp://guest:guest@127.0.0.1:5672", cfg.broker().uri());
assertEquals("term_abc123", cfg.primary().terminal());
assertEquals(5, cfg.primary().remindersOrDefault());
assertEquals(15000L, cfg.primary().backoffMsOrDefault());
}
}
@@ -0,0 +1,40 @@
package dev.ltms.bridged.herdr;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
/** Unit-level behaviour of {@link AgentControl} over a fake herdr. */
class AgentControlTest {
/** The {@code text} of every agent.send, in call order. */
@SuppressWarnings("unchecked")
private static List<String> sendTexts(FakeHerdr herdr) {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.toList();
}
@Test
void sendDeliversThePayloadThenAStandaloneSubmitKey() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_x", "do the thing");
// The Enter must be its own event — appended to the paste it would be swallowed as text.
assertEquals(List.of("do the thing", "\r"), sendTexts(herdr),
"payload paste first, then a separate carriage-return keystroke to submit it");
}
@Test
void sendPreservesEmbeddedNewlinesAndSubmitsOnlyOnce() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_x", "line1\nline2");
assertEquals(List.of("line1\nline2", "\r"), sendTexts(herdr),
"multiline content is delivered verbatim; a single trailing Enter submits it");
}
}
@@ -0,0 +1,43 @@
package dev.ltms.bridged.herdr;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Wire mapping and injectability of {@link AgentStatus}, including the CB-115 {@code done} state. */
class AgentStatusTest {
@Test
void mapsTheKnownWireStrings() {
assertEquals(AgentStatus.IDLE, AgentStatus.fromWire("idle"));
assertEquals(AgentStatus.WORKING, AgentStatus.fromWire("working"));
assertEquals(AgentStatus.BLOCKED, AgentStatus.fromWire("blocked"));
assertEquals(AgentStatus.DONE, AgentStatus.fromWire("done"));
}
@Test
void mapsDoneCaseInsensitively() {
assertEquals(AgentStatus.DONE, AgentStatus.fromWire("DONE"));
assertEquals(AgentStatus.DONE, AgentStatus.fromWire("Done"));
}
@Test
void unknownAndNullFallToUnknown() {
assertEquals(AgentStatus.UNKNOWN, AgentStatus.fromWire("unknown"));
assertEquals(AgentStatus.UNKNOWN, AgentStatus.fromWire("something-else"));
assertEquals(AgentStatus.UNKNOWN, AgentStatus.fromWire(null));
}
@Test
void doneIsInjectableLikeIdle() {
// The whole point of CB-115: a finished worker herdr reports as `done` must be deliverable,
// not treated as UNKNOWN (which wedged delivery and mis-fired the stall failure).
assertTrue(AgentStatus.DONE.injectable());
assertTrue(AgentStatus.IDLE.injectable());
assertTrue(AgentStatus.BLOCKED.injectable());
assertFalse(AgentStatus.WORKING.injectable());
assertFalse(AgentStatus.UNKNOWN.injectable());
}
}
@@ -16,15 +16,20 @@ public final class FakeHerdr implements HerdrClient {
public record Call(String method, Object params) {
}
/** The foreground PID of the one agent pane (term_a) in the canned {@code pane.process_info}. */
public static final long WORKER_PID = 4242;
private final ObjectMapper mapper = new ObjectMapper();
public final List<Call> calls = new ArrayList<>();
private boolean healthy = true;
private final List<String> extraWorkspaces = new ArrayList<>();
private final List<String> extraAgents = new ArrayList<>();
private int agentNameTakenFor = 0;
private int workerTabPaneCount = 1;
private String paneCloseErrorCode = null;
private String agentSendErrorCode = null;
private volatile String agentStatus = "idle"; // what agent.get reports
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String readText = "worker transcript tail"; // canned agent.read output
public FakeHerdr healthy(boolean h) {
this.healthy = h;
@@ -55,12 +60,32 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/** The text {@code agent.read} returns (the CB-106 completion scrape). */
public FakeHerdr readText(String text) {
this.readText = text;
return this;
}
/** Make {@code agent.send} fail with this herdr error code. */
public FakeHerdr agentSendFailsWith(String code) {
this.agentSendErrorCode = code;
return this;
}
/**
* Seed a named agent into {@code agent.list} (e.g. an orphaned worker for CB-117 reaper tests).
* The {@code name} carries the worker label the reaper keys on; {@code paneId}/{@code tabId}
* locate its pane for teardown.
*/
public FakeHerdr withAgent(String name, String terminalId, String paneId, String tabId) {
extraAgents.add(("{\"terminal_id\":\"%s\",\"agent\":\"claude\",\"agent_status\":\"idle\","
+ "\"name\":\"%s\",\"agent_session\":{\"kind\":\"id\",\"value\":\"sess-%s\"},"
+ "\"workspace_id\":\"wQ\",\"tab_id\":\"%s\",\"pane_id\":\"%s\"}")
.formatted(terminalId, name, terminalId, tabId, paneId));
return this;
}
/** Seed an additional workspace into {@code workspace.list} (e.g. a pre-existing worker space). */
public FakeHerdr withWorkspace(String id, String label) {
extraWorkspaces.add(("{\"workspace_id\":\"%s\",\"label\":\"%s\",\"focused\":false,"
@@ -91,11 +116,12 @@ public final class FakeHerdr implements HerdrClient {
{"workspace_id":"w1","label":"dev-mgnl","focused":true,"pane_count":7,"agent_status":"unknown"},
{"workspace_id":"w2","label":"ltms","focused":false,"pane_count":5,"agent_status":"done"}%s]}""")
.formatted(extraWorkspaces.isEmpty() ? "" : "," + String.join(",", extraWorkspaces)));
case "agent.list" -> mapper.readTree("""
case "agent.list" -> mapper.readTree(("""
{"type":"agent_list","agents":[
{"terminal_id":"term_a","agent":"claude","agent_status":"idle",
"agent_session":{"kind":"id","value":"sess-1111"},
"workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}]}""");
"workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}%s]}""")
.formatted(extraAgents.isEmpty() ? "" : "," + String.join(",", extraAgents)));
case "agent.send" -> {
if (agentSendErrorCode != null) {
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.send failed",
@@ -107,6 +133,8 @@ public final class FakeHerdr implements HerdrClient {
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":"claude",
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentStatus));
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
case "agent.start" -> {
long starts = calls.stream().filter(c -> c.method().equals("agent.start")).count();
if (starts <= agentNameTakenFor) {
@@ -114,10 +142,12 @@ public final class FakeHerdr implements HerdrClient {
"herdr error [agent_name_taken]: agent name already used",
"agent_name_taken", null);
}
yield mapper.readTree("""
long n = starts - agentNameTakenFor;
yield mapper.readTree(("""
{"type":"agent_started","agent":{
"terminal_id":"term_new","name":"claude","agent_status":"unknown",
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"w9:pW"}}""");
"terminal_id":"term_new_%d","name":"claude","agent_status":"unknown",
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"w9:pW_%d"}}""")
.formatted(n, n));
}
case "workspace.create" -> mapper.readTree("""
{"type":"workspace_created",
@@ -141,6 +171,21 @@ public final class FakeHerdr implements HerdrClient {
case "pane.get" -> mapper.readTree("""
{"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9",
"tab_id":"w9:t2","agent_status":"idle"}}""");
case "pane.list" -> mapper.readTree("""
{"type":"pane_list","panes":[
{"pane_id":"w2:p7","terminal_id":"term_a","workspace_id":"w2","tab_id":"w2:t7","agent":"claude"},
{"pane_id":"w2:p9","terminal_id":"term_shell","workspace_id":"w2","tab_id":"w2:t8"}]}""");
case "pane.process_info" -> {
Object paneId = params instanceof java.util.Map<?, ?> m ? m.get("pane_id") : null;
yield "w2:p7".equals(paneId)
? mapper.readTree(("""
{"type":"pane_process_info","process_info":{"pane_id":"w2:p7","shell_pid":%d,
"foreground_processes":[{"pid":%d,"name":"node","argv0":"claude"}]}}""")
.formatted(WORKER_PID, WORKER_PID))
: mapper.readTree("""
{"type":"pane_process_info","process_info":{"pane_id":"w2:p9","shell_pid":9001,
"foreground_processes":[]}}""");
}
case "pane.close" -> {
if (paneCloseErrorCode != null) {
throw new HerdrException("herdr error [" + paneCloseErrorCode + "]: pane.close failed",
@@ -0,0 +1,43 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import java.nio.file.Files;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the herdr half of connection-based identity against a REAL herdr: spawn a
* harmless probe, read its actual {@code shell_pid} from {@code pane.process_info}, and confirm
* {@link PaneLocator} resolves that PID back to the probe's own {@code terminal_id}.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@Tag("contract")
class PaneLocatorContractTest {
@Test
void resolvesTheTerminalOwningARealProcessPid() throws Exception {
assumeTrue(Files.exists(UnixSocketHerdrClient.defaultSocketPath()), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
AgentControl agents = new AgentControl(herdr);
Agent probe = agents.start("__pidprobe__", List.of("bash", "-c", "sleep 20"), Map.of());
try {
JsonNode info = herdr.call("pane.process_info", Map.of("pane_id", probe.paneId()))
.path("process_info");
long shellPid = info.path("shell_pid").asLong(-1);
assertTrue(shellPid > 0, "probe pane should report a shell pid");
assertEquals(probe.terminalId(), new PaneLocator(herdr).terminalForPid(shellPid),
"a real PID must resolve back to its own pane's terminal_id");
} finally {
agents.close(probe.paneId());
}
}
}
}
@@ -0,0 +1,27 @@
package dev.ltms.bridged.herdr;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/** Unit tests for PID → pane resolution (the herdr half of connection-based MCP identity). */
class PaneLocatorTest {
private final PaneLocator loc = new PaneLocator(new FakeHerdr());
@Test
void resolvesTerminalForAForegroundPid() {
assertEquals("term_a", loc.terminalForPid(FakeHerdr.WORKER_PID));
}
@Test
void nullForAPidInNoPane() {
assertNull(loc.terminalForPid(999_999));
}
@Test
void nullForNonPositivePid() {
assertNull(loc.terminalForPid(0));
assertNull(loc.terminalForPid(-1));
}
}
@@ -0,0 +1,287 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.msg.Rendezvous;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Unit behaviour of the CB-106 completion resolver in isolation from the injector. */
class CompletionResolverTest {
@Test
void skipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
resolver.resolve("term_a", null); // no in-flight turn captured for this target
assertFalse(herdr.called("agent.read"),
"a turn nobody is blocked on must not cost a transcript scrape");
}
@Test
void failSkipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
resolver.fail("term_a", null); // no in-flight turn, and no registered waiter to fall back to
assertFalse(herdr.called("agent.read"),
"a wedge nobody is blocked on must not cost a transcript scrape");
}
@Test
void captureBaselineSkipsTheReadWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ X\n❯ ");
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
resolver.captureBaseline("term_a"); // no send to attribute a later completion to
assertFalse(herdr.called("agent.read"),
"with no waiting send there is no turn to baseline — skip the scrape");
}
// --- CB-115 clean scrape: extract the last assistant block ----------------
@Test
void extractsTheLastAssistantBlockStrippingChrome() {
String raw = """
⏺ Reading the file…
⏺ Done. The bug was an off-by-one in the loop bound.
╭──────────────────────────────────────╮
│ > │
╰──────────────────────────────────────╯
⏵⏵ auto mode on · ? for shortcuts
""";
assertEquals("Done. The bug was an off-by-one in the loop bound.",
CompletionResolver.lastAssistantBlock(raw));
}
@Test
void keepsMultiLineAssistantContent() {
String raw = "⏺ Line one.\nLine two.\n❯ ";
assertEquals("Line one.\nLine two.", CompletionResolver.lastAssistantBlock(raw));
}
@Test
void fallsBackToRawTextWhenThereIsNoMarker() {
String raw = "plain worker output with no glyph";
assertEquals("plain worker output with no glyph", CompletionResolver.lastAssistantBlock(raw));
}
@Test
void blankScrapeYieldsEmpty() {
assertTrue(CompletionResolver.lastAssistantBlock("").isEmpty());
assertTrue(CompletionResolver.lastAssistantBlock(null).isEmpty());
}
@Test
void stripsSpinnerAndRuleChrome() {
String raw = """
⏺ Channel check confirmed — your message got through.
✻ Brewed for 11s
─────────────────────────────────────
""";
assertEquals("Channel check confirmed — your message got through.",
CompletionResolver.lastAssistantBlock(raw));
}
@Test
void cutsANextTurnPromptEchoAndTrailingTipsFromTheBlock() {
// The exact turn-2 leak: the scrape captured the settled answer, then a "✻ Cooked" spinner,
// then the NEXT turn's echoed prompt, then a "✶ Forming…" spinner and trailing tips/warnings
// whose lines (⎿, ⚠) are not themselves chrome-terminated. Stopping at the first boundary
// (the ✻ spinner) is what keeps every one of those interface lines out of the reply.
String raw = """
⏺ Channel confirmed — the bridge reply delivered successfully.
✻ Cooked for 9s
❯ Thanks. Now a small task: what is 17 * 23? Show just the number.
✶ Forming…
⎿ Tip: Name your conversations with /rename
⚠ claude.ai connectors are disabled because ANTHROPIC_API_KEY is set
""";
assertEquals("Channel confirmed — the bridge reply delivered successfully.",
CompletionResolver.lastAssistantBlock(raw));
}
// --- CB-115 misattribution guard: suppress a stale (unchanged) completion -------
@Test
void suppressesACompletionWhoseScrapeIsUnchangedFromDelivery() {
// Rapid back-to-back turn: the pane still shows the PREVIOUS turn's answer when this turn's
// (misattributed) completion boundary fires. The scrape == the delivery baseline, so the
// send must NOT be resolved with the stale answer.
FakeHerdr herdr = new FakeHerdr().readText("⏺ 391\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
// The turn as captured at delivery: its waiter, and the previous turn's answer still on screen.
var turn = new CompletionResolver.InFlight(waiter, "391");
resolver.resolve("term_a", turn); // scrape still "391" == baseline → suppress
assertFalse(waiter.isDone(), "a completion with no output change must not resolve the send");
assertTrue(rendezvous.isWaiting("term_a"), "the send stays waiting for a real reply");
}
@Test
void resolvesACompletionWhoseScrapeChangedSinceDelivery() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ No, 391 = 17 × 23.\n❯ "); // the worker's real answer
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
// Delivery baseline was the previous turn's "391"; the scrape now differs → resolve.
var turn = new CompletionResolver.InFlight(waiter, "391");
resolver.resolve("term_a", turn);
assertTrue(waiter.isDone(), "a completion with new output must resolve the send");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
assertEquals("No, 391 = 17 × 23.", waiter.getNow(null).text());
}
@Test
void suppressesAnUnchangedCompletionEvenWhenTheBlockExceedsTheScrapeCap() {
// The fan-out issue-hunt finding: captureBaseline once stored the RAW (unclipped) assistant
// block while resolve compares against a clip()'d tail. For a block longer than MAX_SCRAPE_CHARS
// the two capped representations differ even when the pane never changed, so the CB-115
// byte-identical guard failed to fire and a stale completion could resolve the send. Both sides
// must clip identically; here an unchanged >cap block on rapid back-to-back turns stays suppressed.
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
resolver.captureBaseline("term_a"); // baseline is the clipped >cap block
var turn = resolver.inFlight("term_a");
assertEquals(CompletionResolver.MAX_SCRAPE_CHARS, turn.baseline().length(),
"the delivery baseline is clipped to the same cap resolve() applies to the tail");
resolver.resolve("term_a", turn); // scrape unchanged → clipped tail == baseline → suppress
assertFalse(waiter.isDone(),
"an unchanged >cap block must still be recognised as stale and suppressed");
assertTrue(rendezvous.isWaiting("term_a"), "the send stays waiting for a real reply");
}
@Test
void resolvesWhenThereIsNoBaseline() {
// No delivery baseline (e.g. the pre-turn read failed) ⇒ never suppress; the completion resolves.
FakeHerdr herdr = new FakeHerdr().readText("⏺ hello\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "with no baseline a completion resolves as before");
assertEquals("hello", waiter.getNow(null).text());
}
@Test
void resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent() {
// The most important branch of the CB-115 guard: a failed read means the resolver could not
// SEE the screen — "couldn't see", not "no change". It must still resolve the send (an empty
// tail beats hanging until the caller's timeout), even though a baseline was captured. The
// baseline here is "" (an empty pane at delivery), so without the !scrapeFailed clause the
// byte-identical guard would wrongly match the empty tail and suppress.
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.read throws HerdrException
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, ""); // empty pane baselined at delivery
resolver.resolve("term_a", turn);
assertTrue(waiter.isDone(),
"a failed scrape must still resolve the send, not hang until the caller's timeout");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
assertEquals("", waiter.getNow(null).text(), "the tail is empty because the screen was unreadable");
}
// --- CB-115/CB-116 fail guard: an already-done or absent waiter is left alone ---------
@Test
void failLeavesAnAlreadyResolvedWaiterUntouchedAndSkipsTheScrape() {
// The send was already resolved (e.g. by the worker's explicit reply) before fail fired.
// fail must not overwrite that value, and must not even scrape the worker — nobody needs it.
FakeHerdr herdr = new FakeHerdr().readText("an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null);
assertTrue(rendezvous.resolveCompletion(waiter, "already replied"));
resolver.fail("term_a", turn);
assertFalse(herdr.called("agent.read"),
"fail must not scrape a waiter that is already done");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"fail must not overwrite the existing resolution");
assertEquals("already replied", waiter.getNow(null).text());
}
@Test
void failFallsBackToTheRegisteredWaiterWhenThereIsNoInFlightTurn() {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fail falls back to the waiter currently registered on the Rendezvous and fails it.
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a"); // send registered, but no captureBaseline ever ran
resolver.fail("term_a", null); // no in-flight turn → fall back to the registered waiter
assertTrue(waiter.isDone(), "fail falls back to the registered waiter when no turn is in flight");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
assertEquals("stuck on an error screen", waiter.getNow(null).text());
}
// --- CB-116 waiter identity: a late completion never crosses into the next turn ---------
@Test
void aLateCompletionForOneTurnNeverResolvesTheNextTurnsWaiter() {
// The cross-turn stale reply the conversation test surfaced: turn N's completion fallback
// fires AFTER turn N was resolved by an explicit bridge_reply and turn N+1 has opened its own
// waiter on the same session. Resolving "whatever is waiting now" would hand turn N's stale
// scrape to turn N+1; targeting turn N's captured waiter makes the late completion a no-op.
FakeHerdr herdr = new FakeHerdr().readText("⏺ turn N answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiterN = rendezvous.open("term_a"); // turn N's send
// The turn as the injector captured it at delivery (waiter + pre-turn baseline).
var turnN = new CompletionResolver.InFlight(waiterN, "an earlier answer");
// Turn N is resolved by the worker's explicit reply.
assertTrue(rendezvous.resolve("term_a", "N replied"));
// Turn N+1's send opens its own waiter on the same session (replacing the registered one).
var waiterN1 = rendezvous.open("term_a");
resolver.resolve("term_a", turnN); // turn N's completion fallback finally fires
assertFalse(waiterN1.isDone(), "turn N's late completion must not resolve turn N+1's waiter");
assertEquals(Rendezvous.Kind.REPLY, waiterN.getNow(null).kind(),
"turn N stays resolved by its own reply");
assertTrue(rendezvous.isWaiting("term_a"), "turn N+1 is still awaiting its own resolution");
}
}
@@ -6,8 +6,10 @@ import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.TimeUnit;
@@ -25,12 +27,18 @@ class InjectorTest {
private final FakeHerdr herdr = new FakeHerdr();
private final Injector injector = new Injector(new AgentControl(herdr));
/** Text of every agent.send, in order. */
/**
* The logical messages delivered, in order. AgentControl.send emits each delivery as two
* agent.send calls — the payload, then a standalone Enter keystroke ({@code "\r"}) to submit
* it; these tests assert delivery ordering/gating, not the submit event, so drop the bare
* carriage returns.
*/
@SuppressWarnings("unchecked")
private List<String> sent() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.filter(t -> !t.equals("\r"))
.toList();
}
@@ -43,6 +51,47 @@ class InjectorTest {
assertEquals(List.of("hello"), sent());
}
@Test
void holdsDeliveryUntilTheWorkerIsAvailable() {
// CB-113: idle alone is not enough — hold until the worker's MCP is connected (ready).
java.util.Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // idle but not yet available → held out of the boot window
assertEquals(List.of(), sent(), "must not deliver into a not-yet-available worker");
ready.add(T); // the worker's Claude connects the bridge MCP
inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("task"), sent(), "delivers once the worker is available");
}
@SuppressWarnings("unchecked")
private long enterKeystrokes() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> "\r".equals(((Map<String, Object>) c.params()).get("text")))
.count();
}
@Test
void resubmitsEnterWhenADeliveredMessageIsNotPickedUp() {
// CB-113: the Enter at delivery can race the paste; while the worker stays idle (not picked
// up), the injector re-nudges Enter so the pending paste submits.
injector.enqueue(T, "task");
injector.onStatus(T, AgentStatus.IDLE); // deliver: paste + one Enter
long afterDeliver = enterKeystrokes();
injector.onStatus(T, AgentStatus.IDLE); // still idle → re-nudge Enter
injector.onStatus(T, AgentStatus.IDLE); // and again
assertTrue(enterKeystrokes() > afterDeliver, "an unpicked-up delivery re-nudges Enter");
injector.onStatus(T, AgentStatus.WORKING); // worker finally starts
long atPickup = enterKeystrokes();
injector.onStatus(T, AgentStatus.WORKING);
assertEquals(atPickup, enterKeystrokes(), "no more nudges once the worker has picked up");
}
@Test
void holdsWhileWorkingThenDeliversOnIdle() {
injector.enqueue(T, "later");
@@ -120,17 +169,97 @@ class InjectorTest {
}
@Test
void activeWhileQueuedOrInFlightThenQuietAfterPickup() {
void activeWhileQueuedOrInFlightThenQuietAfterTurnCompletes() {
assertTrue(injector.activeTargets().isEmpty());
injector.enqueue(T, "x");
assertEquals(java.util.Set.of(T), injector.activeTargets(), "active while a message is queued");
assertEquals(Set.of(T), injector.activeTargets(), "active while a message is queued");
injector.onStatus(T, AgentStatus.IDLE); // delivers; still in-flight (awaiting pickup)
assertEquals(java.util.Set.of(T), injector.activeTargets(),
injector.onStatus(T, AgentStatus.IDLE); // delivers; awaiting pickup
assertEquals(Set.of(T), injector.activeTargets(),
"stays active so the poller can observe the worker pick the message up");
injector.onStatus(T, AgentStatus.WORKING); // pickup observed → in-flight cleared
assertTrue(injector.activeTargets().isEmpty(), "quiet once queue is empty and pickup is seen");
injector.onStatus(T, AgentStatus.WORKING); // pickup observed; now awaiting turn completion
assertEquals(Set.of(T), injector.activeTargets(),
"stays active after pickup so the working→idle completion boundary is observed");
injector.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete
assertTrue(injector.activeTargets().isEmpty(), "quiet once the delegated turn has completed");
}
@Test
void firesTurnCompleteOnAConfirmedWorkingThenIdle() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // pickup + turn running
assertEquals(List.of(), completed, "no completion until the turn returns to idle");
inj.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete
assertEquals(List.of(T), completed, "a confirmed working→idle fires exactly one completion");
}
@Test
void doesNotSynthesizeCompletionFromAnUnconfirmedTurn() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
// Deliver, then only ever idle — a `working` sample is never seen. The pickup grace unwedges
// the queue but must NOT invent a completion: without a sampled turn there is no trustworthy
// "the worker finished the task" signal, so the send should fall through to its timeout.
for (int i = 0; i < 15; i++) inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of(), completed, "no completion is synthesized from an unconfirmed turn");
}
/** Captures both turn-lifecycle callbacks so the CB-109 stall path can be asserted. */
private static final class Captor implements TurnListener {
final List<String> completed = new ArrayList<>();
final List<String> failed = new ArrayList<>();
@Override
public void onTurnComplete(String target) {
completed.add(target);
}
@Override
public void onTurnFailed(String target) {
failed.add(target);
}
}
// ~30s of unknown at the 250ms prod poll interval; enough onStatus samples to trip the stall.
private static final int STALL_SAMPLES = 130;
@Test
void failsAnOutstandingDelegationWhoseWorkerWedgesInUnknown() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // worker starts the turn
for (int i = 0; i < STALL_SAMPLES; i++) inj.onStatus(T, AgentStatus.UNKNOWN); // then wedges
assertEquals(List.of(T), cap.failed, "a sustained unknown streak fails the outstanding send");
assertEquals(List.of(), cap.completed, "a wedge is a failure, not a completion");
assertTrue(inj.activeTargets().isEmpty(), "the wedged target is reclaimed, not polled forever");
}
@Test
void aTransientUnknownGlitchNeitherFailsNorBlocksCompletion() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // confirmed turn
for (int i = 0; i < 10; i++) inj.onStatus(T, AgentStatus.UNKNOWN); // brief glitch, well under grace
inj.onStatus(T, AgentStatus.IDLE); // working → idle: the real completion
assertEquals(List.of(), cap.failed, "a short unknown blip must not fail the turn");
assertEquals(List.of(T), cap.completed, "the streak reset, so the turn still completes");
}
@Test
@@ -151,6 +280,86 @@ class InjectorTest {
assertTrue(f.isCompletedExceptionally(), "queued waiters unblock when the worker vanishes");
}
@Test
void dropFailsTheTurnOfADeliveredMessageWhenTheWorkerVanishes() {
// CB-110: the message was delivered (no longer queued), so failing queued waiters alone would
// leave its send hanging. A vanished worker must fail that in-flight turn too.
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // turn running
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), cap.failed, "a worker that vanishes mid-turn fails its in-flight send");
}
@Test
void dropFailsADeliveredTurnThatVanishesBeforePickupIsConfirmed() {
// Delivered but no WORKING sampled yet (awaitingCompletion=true, awaitingPickup still true,
// turnObserved=false) — a distinct state the other two drop tests don't cover. (Gap surfaced
// by an off-sub worker's review of CB-110, delegated through the bridge.)
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver; pickup never confirmed
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), cap.failed, "a delivery that vanishes before pickup still fails its send");
}
// ~60s of idle-but-not-ready at the 250ms prod poll interval; enough to trip the readiness grace.
private static final int READINESS_SAMPLES = 245;
@Test
void failsAQueuedMessageWhoseWorkerNeverBecomesReady() {
// CB-114: herdr keeps reporting the worker idle, but its Claude never connects the bridge MCP,
// so the readiness gate never opens. The message must not be held (and the target polled)
// forever — after the grace it fails, the caller unblocks via the worker-failure path, the
// target is reclaimed, and the never-set presence is cleared.
Captor cap = new Captor();
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), cap, _ -> false, forgotten::add);
CompletableFuture<Void> f = inj.enqueue(T, "task");
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of(), sent(), "a never-ready worker is never delivered to");
assertTrue(f.isCompletedExceptionally(), "the caller's future fails instead of hanging forever");
assertEquals(List.of(T), cap.failed, "the awaiting send resolves through the worker-failure path");
assertEquals(List.of(), cap.completed, "a never-ready worker is a failure, not a completion");
assertEquals(List.of(T), forgotten, "the never-ready worker's presence is cleared");
assertTrue(inj.activeTargets().isEmpty(), "the target is reclaimed, not polled forever");
}
@Test
void aWorkerThatBecomesReadyWithinTheGraceIsDeliveredNormally() {
// The readiness grace must not fail a worker that is merely slow to boot: once it becomes
// available before the grace elapses, delivery proceeds as usual (the counter resets).
Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains, _ -> {
});
inj.enqueue(T, "task");
for (int i = 0; i < 100; i++) inj.onStatus(T, AgentStatus.IDLE); // still booting, well under grace
assertEquals(List.of(), sent());
ready.add(T); // MCP connects before the grace elapses
inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("task"), sent(), "a worker that connects within the grace is delivered to");
}
@Test
void dropClearsWorkerPresence() {
// CB-114 (finding #1): a vanished worker's readiness must be forgotten so a stale entry cannot
// linger past the worker's life (WorkerPresence.forget had no caller before this).
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, forgotten::add);
inj.enqueue(T, "orphan");
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), forgotten, "drop clears the gone worker's presence");
}
@Test
void pollerDeliversToAnIdleWorker() throws Exception {
// End-to-end through the poller: idle worker → message delivered without manual onStatus.
@@ -170,7 +379,9 @@ class InjectorTest {
@SuppressWarnings("unchecked")
Map<String, Object> p = (Map<String, Object>) c.params();
return p.get("text").toString();
}).toList());
})
.filter(t -> !t.equals("\r")) // drop the standalone submit keystroke
.toList());
}
@Test
@@ -0,0 +1,88 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Content-based refinement of an unreliable {@code UNKNOWN} status (CB-115). */
class StatusRefinerTest {
// --- pure classification -------------------------------------------------
@Test
void classifiesAnIdlePromptAsIdle() {
String pane = """
⏺ All done — the file compiles cleanly.
╭──────────────────────────────────────╮
│ > │
╰──────────────────────────────────────╯
⏵⏵ auto mode on (shift+tab to cycle)
""";
assertEquals(AgentStatus.IDLE, StatusRefiner.classify(pane));
}
@Test
void classifiesABarePromptGlyphAsIdle() {
assertEquals(AgentStatus.IDLE, StatusRefiner.classify("some output\n❯ "));
}
@Test
void classifiesActiveGenerationAsWorking() {
String pane = """
⏺ Working on it…
✳ Thinking… (12s · esc to interrupt)
""";
assertEquals(AgentStatus.WORKING, StatusRefiner.classify(pane));
}
@Test
void anEscToInterruptScreenIsWorkingEvenWithAPromptBox() {
// "esc to interrupt" wins over a prompt box: the turn is still generating.
String pane = "│ > │\n esc to interrupt";
assertEquals(AgentStatus.WORKING, StatusRefiner.classify(pane));
}
@Test
void anUnrecognizableScreenStaysUnknown() {
assertEquals(AgentStatus.UNKNOWN, StatusRefiner.classify("garbled ansi noise with no prompt"));
assertEquals(AgentStatus.UNKNOWN, StatusRefiner.classify(""));
assertEquals(AgentStatus.UNKNOWN, StatusRefiner.classify(null));
}
// --- refine() wiring -----------------------------------------------------
@Test
void refinePassesNonUnknownStatusesThroughWithoutReading() {
FakeHerdr herdr = new FakeHerdr();
StatusRefiner refiner = new StatusRefiner(new AgentControl(herdr));
assertEquals(AgentStatus.WORKING, refiner.refine("term_a", AgentStatus.WORKING));
assertEquals(AgentStatus.IDLE, refiner.refine("term_a", AgentStatus.IDLE));
assertFalse(herdr.called("agent.read"),
"a trusted status must not cost a pane read");
}
@Test
void refineUpgradesUnknownToIdleFromPaneContent() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ answer\n❯ ");
StatusRefiner refiner = new StatusRefiner(new AgentControl(herdr));
assertEquals(AgentStatus.IDLE, refiner.refine("term_a", AgentStatus.UNKNOWN));
assertTrue(herdr.called("agent.read"), "an UNKNOWN must trigger a pane read");
}
@Test
void refineLeavesUnknownWhenContentIsUnclassifiable() {
FakeHerdr herdr = new FakeHerdr().readText("nothing recognizable here");
StatusRefiner refiner = new StatusRefiner(new AgentControl(herdr));
assertEquals(AgentStatus.UNKNOWN, refiner.refine("term_a", AgentStatus.UNKNOWN));
}
}
@@ -0,0 +1,33 @@
package dev.ltms.bridged.inject;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** The CB-113 worker-availability registry. */
class WorkerPresenceTest {
@Test
void tracksPresenceAndForgets() {
WorkerPresence p = new WorkerPresence();
assertFalse(p.isPresent("term_a"), "unseen worker is not available");
p.markPresent("term_a");
assertTrue(p.isPresent("term_a"), "a worker seen on the MCP is available");
p.forget("term_a");
assertFalse(p.isPresent("term_a"), "a torn-down worker is no longer available");
}
@Test
void nullOrBlankMarkIsANoOp() {
WorkerPresence p = new WorkerPresence();
assertDoesNotThrow(() -> {
p.markPresent(null);
p.markPresent(" ");
});
assertFalse(p.isPresent(""), "blank/null contacts (the primary) are never present");
}
}
@@ -0,0 +1,168 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.auth.Role;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-513 — the CB-505 authorization gate on the <strong>MCP</strong> entry path.
*
* <p>Why this file exists: CB-505 claimed authorization is "enforced on both entry paths", and it
* is — but only REST was ever tested ({@code BridgedAppAuthTest}). Coverage showed
* {@code BridgeMcp.deny()}, {@code principal()} and every tool-registration lambda at <em>zero</em>
* executed lines, because no test had ever constructed a {@code BridgeMcp} — the existing
* {@code BridgeMcpTest} calls only the static handler methods. An unexercised security control is
* a claim, not a control.
*
* <p>These tests construct a real {@code BridgeMcp} (which also exercises the constructor and the
* tool wiring) and drive the policy half of the gate directly.
*/
class BridgeMcpAuthzTest {
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private Metrics metrics;
private BridgeMcp mcp;
@AfterEach
void close() {
if (mcp != null) mcp.close();
}
/** A fully wired BridgeMcp on fakes — constructing it is itself part of what is under test. */
private BridgeMcp mcp(boolean enforce) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "tok");
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(),
new InMemoryReplyInbox());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
metrics = BridgedMetrics.create(sessions, new InMemoryReplyInbox());
mcp = new BridgeMcp(messages, workers, sessions, identity, sessions.asPresence(),
new PrimaryRegistry(null),
enforce ? new CallerResolver(identity) : null,
metrics);
return mcp;
}
private static final Principal PRIMARY = Principal.primary(100);
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal ANON = Principal.anonymous();
// --- the table, enforced on THIS path too ---------------------------------------------------
@Test
void primaryMayOrchestrate() {
BridgeMcp m = mcp(true);
for (Authz.Action a : new Authz.Action[]{Authz.Action.SPAWN, Authz.Action.STOP,
Authz.Action.SEND, Authz.Action.DRAIN, Authz.Action.READ}) {
assertNull(m.denyFor(PRIMARY, a, "term_a"), a + " is the primary's to perform");
}
}
@Test
void aWorkerMayNotOrchestrateOverMcp() {
BridgeMcp m = mcp(true);
for (Authz.Action a : new Authz.Action[]{Authz.Action.SPAWN, Authz.Action.STOP,
Authz.Action.SEND, Authz.Action.DRAIN}) {
McpSchema.CallToolResult denied = m.denyFor(WORKER_A, a, "term_a");
assertNotNull(denied, a + " must be refused to a worker");
assertTrue(denied.isError(), "a refusal is returned as an MCP tool error");
}
}
@Test
void aWorkerMayReplyAndAskOnlyAsItself() {
BridgeMcp m = mcp(true);
assertNull(m.denyFor(WORKER_A, Authz.Action.REPLY, "term_a"), "its own session is allowed");
assertNull(m.denyFor(WORKER_A, Authz.Action.ASK, "term_a"));
assertNotNull(m.denyFor(WORKER_A, Authz.Action.REPLY, "term_b"),
"worker A must not reply on worker B's session");
assertNotNull(m.denyFor(WORKER_A, Authz.Action.ASK, "term_b"));
}
@Test
void thePrimaryMayNotForgeAWorkerReplyOverMcp() {
BridgeMcp m = mcp(true);
// A forged reply would resolve the very rendezvous the primary is blocked on.
assertNotNull(m.denyFor(PRIMARY, Authz.Action.REPLY, "term_a"));
assertNotNull(m.denyFor(PRIMARY, Authz.Action.ASK, "term_a"));
}
@Test
void anonymousIsRefusedEverythingAndCountedAsUnauthenticated() {
BridgeMcp m = mcp(true);
McpSchema.CallToolResult denied = m.denyFor(ANON, Authz.Action.READ, null);
assertNotNull(denied, "authenticated as nothing ⇒ authorized for nothing");
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "unauthenticated"));
assertEquals(0, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "forbidden"),
"a missing credential is 401-shaped, not 403-shaped");
}
@Test
void aWrongRoleIsCountedAsForbiddenNotUnauthenticated() {
BridgeMcp m = mcp(true);
assertNotNull(m.denyFor(WORKER_A, Authz.Action.SPAWN, null));
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "forbidden"));
assertEquals(0, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "unauthenticated"),
"the caller IS authenticated — it is just not the right role");
}
@Test
void theLegacyConstructorLeavesTheGateOpen() {
// The 22 pre-existing BridgeMcpTest cases rely on no authorization being enforced.
BridgeMcp m = mcp(false);
assertNull(m.denyFor(ANON, Authz.Action.SPAWN, null),
"no CallerResolver supplied ⇒ authorization not enforced (legacy behaviour)");
}
// --- identity reconstruction from the transport context ------------------------------------
@Test
void principalIsRebuiltFromTheStashedRole() {
assertEquals(Role.WORKER, BridgeMcp.principalFrom("WORKER", "term_a", 7).role());
assertEquals("term_a", BridgeMcp.principalFrom("WORKER", "term_a", 7).terminal());
assertEquals(Role.PRIMARY, BridgeMcp.principalFrom("PRIMARY", null, 7).role());
assertEquals(Role.ANONYMOUS, BridgeMcp.principalFrom("ANONYMOUS", null, -1).role());
}
@Test
void aMissingRoleFallsBackToTheHistoricalInterpretation() {
// Legacy path: no role stashed. A terminal means worker; its absence meant "the primary",
// which is exactly the pre-CB-501 default CB-501 inverted — preserved only here.
assertEquals(Role.WORKER, BridgeMcp.principalFrom(null, "term_a", 7).role());
assertEquals(Role.PRIMARY, BridgeMcp.principalFrom(null, null, 7).role());
}
}
@@ -0,0 +1,401 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
/**
* Parity tests for the MCP tool adapters — they must produce the same outcomes as the REST routes,
* since both drive the same {@link MessageService}/{@link Rendezvous}. The MCP wire protocol itself
* is the SDK's concern; here we test the thin adapter logic directly.
*/
class BridgeMcpTest {
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final MessageService messages = new MessageService(agents, new Injector(agents), rendezvous);
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
private static ClaudeCodeLauncher workerService(FakeHerdr h, String baseUrl, Set<String> allow) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", baseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(allow), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> "tok");
}
private static SessionManager sessionManager(FakeHerdr h, String baseUrl, Set<String> allow) {
return new SessionManager(workerService(h, baseUrl, allow));
}
@Test
void sendThenReplyRoundTrips() throws Exception {
// bridge_send blocks; bridge_reply resolves it with the worker's structured answer.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> BridgeMcp.send(messages, "term_a", "review this", 4000L));
// Wait until the send has opened its waiter so the reply resolves it (CB-307: reply now
// queues in the inbox if no waiter is open, which would break the round-trip).
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "send should have opened its waiter");
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "LGTM");
assertEquals("delivered", textOf(reply));
McpSchema.CallToolResult res = send.get(6, TimeUnit.SECONDS);
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("LGTM", textOf(res));
}
@Test
void asyncSendReturnsATicketThenPollReportsTheReply() throws Exception {
// wait:false parity — a ticket is issued, resolved by a reply, and surfaced by bridge_poll.
McpSchema.CallToolResult accepted = BridgeMcp.sendAsync(messages, "term_a", "do it");
assertNotEquals(Boolean.TRUE, accepted.isError());
String out = textOf(accepted);
assertTrue(out.contains("ticket="), out);
String ticket = out.substring(out.indexOf("ticket=") + "ticket=".length()).trim();
// Wait until the send has opened its waiter before replying (CB-307: reply never errors,
// so the old retry-on-error pattern no longer works — it would queue instead of resolve).
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "send should have opened its waiter");
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "async LGTM");
assertEquals("delivered", textOf(reply));
// Poll until the async send completes and reports the reply.
McpSchema.CallToolResult polled = BridgeMcp.poll(messages, ticket, null);
deadline = System.currentTimeMillis() + 3000;
while (!textOf(polled).contains("async LGTM") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(10);
polled = BridgeMcp.poll(messages, ticket, null);
}
assertEquals("async LGTM", textOf(polled));
}
@Test
void pollUnknownTicketIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.poll(messages, "task-999", null);
assertTrue(res.isError());
assertTrue(textOf(res).contains("unknown ticket"));
}
@Test
void sendTimesOutWithAWorkingNote() {
McpSchema.CallToolResult res = BridgeMcp.send(messages, "term_a", "hi", 120L);
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res));
}
@Test
void sendRejectsMissingArgs() {
assertTrue(BridgeMcp.send(messages, null, "hi", null).isError());
assertTrue(BridgeMcp.send(messages, "term_a", " ", null).isError());
}
@Test
void replyWithNoPendingSendIsQueuedNotError() {
// CB-307: a reply with no open send is now queued in the inbox, not an error.
McpSchema.CallToolResult res = BridgeMcp.reply(messages, "term_a", "orphan");
assertNotEquals(Boolean.TRUE, res.isError(), "a queued reply is not an error");
assertEquals("delivered", textOf(res));
// The reply is drainable by target.
var drained = messages.drainReplies("term_a");
assertEquals(1, drained.size());
assertEquals("orphan", drained.getFirst().content());
}
@Test
void bridgePollWithTargetDrainsReplies() {
// A reply with no open send queues it in the inbox.
BridgeMcp.reply(messages, "term_a", "queued-msg");
// bridge_poll with target drains the inbox.
McpSchema.CallToolResult res = BridgeMcp.poll(messages, null, "term_a");
assertNotEquals(Boolean.TRUE, res.isError());
String text = textOf(res);
assertTrue(text.contains("queued-msg"), "the drained reply should appear in the result");
// Second drain returns empty.
McpSchema.CallToolResult empty = BridgeMcp.poll(messages, null, "term_a");
assertEquals("[]", textOf(empty));
}
@Test
void askThenAnswerRoundTrips() throws Exception {
// The primary delegates and blocks; wait until its waiter is open before the worker asks.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> BridgeMcp.send(messages, "term_a", "do X", 5000L));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "the send must be waiting for the ask to surface to");
// The worker asks mid-turn; the call blocks for the primary's answer.
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
() -> BridgeMcp.ask(messages, "term_a", "which config?", 5000L));
// The primary's send unblocks with the question and a turnId to answer on.
McpSchema.CallToolResult q = send.get(6, TimeUnit.SECONDS);
assertNotEquals(Boolean.TRUE, q.isError());
String qt = textOf(q);
assertTrue(qt.contains("[question]"), qt);
String afterMarker = qt.substring(qt.indexOf("turnId=\"") + "turnId=\"".length());
String turnId = afterMarker.substring(0, afterMarker.indexOf('"'));
// The primary answers via bridge_send(turnId); this blocks again for the worker's reply.
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> BridgeMcp.answer(messages, turnId, "config.yaml", 5000L));
// The worker's ask returns the answer — it resumes the same turn.
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
// The resumed worker replies, resolving the answering send (wait for the reopened waiter).
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "the answer should have reopened a waiter");
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "done");
assertEquals("delivered", textOf(reply));
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
}
@Test
void askFromANonWorkerConnectionIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.ask(messages, null, "which config?", 500L);
assertTrue(res.isError());
assertTrue(textOf(res).contains("workers only"), textOf(res));
}
@Test
void answerToAStaleTurnIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.answer(messages, "term_a#999", "too late", 500L);
assertTrue(res.isError());
assertTrue(textOf(res).contains("no longer open"), textOf(res));
}
@Test
void spawnReturnsTheNewWorkersSessionAndPane() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"sessionId\":\"term_new_1\""), out);
assertTrue(out.contains("\"paneId\":\"w9:pW_1\""), out);
assertTrue(out.contains("\"status\":\"spawning\""), out);
}
@Test
void spawnRejectsAnOffAllowlistProfileWithoutTouchingHerdr() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res =
BridgeMcp.spawn(sessionManager(h, "https://api.anthropic.com", Set.of("gx00.gw")), null);
assertTrue(res.isError());
assertTrue(textOf(res).contains("subscription boundary"));
assertFalse(h.called("agent.start"), "the guard must block before any spawn");
}
@Test
void spawnRejectsAnUnknownProfileAsAnError() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res =
BridgeMcp.spawn(sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), "nope");
assertTrue(res.isError());
assertTrue(textOf(res).contains("unknown worker profile"), textOf(res));
}
@Test
void spawnPassesTheRequestedCwdToTheWorker() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "/req/dir", null, null, null);
assertNotEquals(Boolean.TRUE, res.isError());
@SuppressWarnings("unchecked")
Map<String, Object> start = (Map<String, Object>) h.lastCall("agent.start").params();
assertEquals("/req/dir", start.get("cwd"));
}
@Test
void profilesListsConfiguredProfilesAndDefault() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.profiles(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("ltms-local"), out);
assertTrue(out.contains("\"default\":\"ltms-local\""), out);
}
@Test
void listReportsTrackedWorkers() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-304", null));
McpSchema.CallToolResult res = BridgeMcp.listWorkers(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"sessionId\":\"" + s.terminalId() + "\""), out);
assertTrue(out.contains("\"paneId\":\"" + s.paneId() + "\""), out);
assertTrue(out.contains("\"profile\":\"ltms-local\""), out);
assertTrue(out.contains("\"state\":\"spawning\""), out);
assertTrue(out.contains("\"worktree\":\"" + s.worktree() + "\""), out);
assertTrue(out.contains("\"branch\":\"" + s.branch() + "\""), out);
assertTrue(out.contains("\"owner\":\"term_primary\""), out);
assertTrue(out.contains("\"liveStatus\":\"unknown\""), out);
}
@Test
void stopTearsDownAWorkerByPane() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.stop(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), "w9:pW");
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("stopped w9:pW", textOf(res));
assertTrue(h.called("pane.close"));
}
@Test
void stopRequiresAPaneId() {
FakeHerdr h = new FakeHerdr();
assertTrue(BridgeMcp.stop(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), " ").isError());
}
@Test
void bridgeAckReturnsConfirmationForValidArgs() {
McpSchema.CallToolResult res = BridgeMcp.ack(messages, "term_a", "msg-1");
assertNotEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("msg-1"), "response should mention the msgId");
}
@Test
void bridgeAckRejectsMissingArgs() {
assertTrue(BridgeMcp.ack(messages, null, "msg-1").isError());
assertTrue(BridgeMcp.ack(messages, "term_a", null).isError());
assertTrue(BridgeMcp.ack(messages, " ", "msg-1").isError());
}
@Test
void bridgeAckRemovesSpecificReply() {
// Queue a reply and capture its msgId.
BridgeMcp.reply(messages, "term_a", "orphan");
var before = messages.drainReplies("term_a");
assertEquals(1, before.size(), "one reply in the inbox");
String msgId = before.getFirst().msgId();
// Publish the same reply again and ack it via bridge_ack surface.
BridgeMcp.reply(messages, "term_a", "orphan-again");
var peeked = messages.drainReplies("term_a");
assertEquals(1, peeked.size(), "one fresh reply in the inbox");
// ackReply works (no-op since published with a different UUID, but callable).
assertDoesNotThrow(() -> messages.ackReply("term_a", msgId));
}
@Test
void statusReportsLiveAgentStatus() {
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
AgentControl blockedAgents = new AgentControl(blocked);
McpSchema.CallToolResult res = BridgeMcp.status(
new MessageService(blockedAgents, new Injector(blockedAgents), rendezvous), "term_a");
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("blocked", textOf(res));
}
// --- bridge_whoami: the caller's own identity, so an agent never has to guess its role -------
@Test
void whoamiReportsThePrimaryAsPrimaryAndNothingElse() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.whoami(
Principal.primary(100), sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"primary\""), out);
// The primary owns no session — leaking a sessionId here would invite it to reply as one.
assertFalse(out.contains("sessionId"), out);
}
@Test
void whoamiReportsAWorkerWithItsRegisteredSession() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-517", null));
McpSchema.CallToolResult res = BridgeMcp.whoami(Principal.worker(s.terminalId(), 200), sessions);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"worker\""), out);
assertTrue(out.contains("\"sessionId\":\"" + s.terminalId() + "\""), out);
assertTrue(out.contains("\"profile\":\"ltms-local\""), out);
assertTrue(out.contains("\"worktree\":\"" + s.worktree() + "\""), out);
assertTrue(out.contains("\"branch\":\"" + s.branch() + "\""), out);
assertTrue(out.contains("\"owner\":\"term_primary\""), out);
}
/**
* A worker the registry has no record of — it outlived a daemon restart — must still learn the
* load-bearing fact. Degrading to "I don't know who you are" would put it back to guessing,
* which is the failure this tool exists to remove.
*/
@Test
void whoamiStillReportsWorkerRoleWhenTheSessionIsUnregistered() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.whoami(Principal.worker("term_orphan", 200),
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"worker\""), out);
assertTrue(out.contains("\"sessionId\":\"term_orphan\""), out);
assertFalse(out.contains("profile"), out); // nothing invented for a session we don't track
}
}
@@ -0,0 +1,46 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/** Connection → caller-identity resolution, with the OS peer-PID lookup faked. */
class ConnectionIdentityTest {
private final FakeHerdr herdr = new FakeHerdr();
private ConnectionIdentity with(PeerPidLookup pids) {
return new ConnectionIdentity(new PaneLocator(herdr), pids);
}
@Test
void resolvesWorkerFromLoopbackPeerPid() {
assertEquals("term_a", with(_ -> FakeHerdr.WORKER_PID).callerTerminal("127.0.0.1", 55555));
}
@Test
void nullForOffHostCaller() {
// A non-loopback peer can't be an on-host worker → treat as primary/unknown.
assertNull(with(_ -> FakeHerdr.WORKER_PID).callerTerminal("10.0.0.9", 55555));
}
@Test
void nullWhenPidOwnsNoPane() {
// e.g. the primary — its PID maps to no worker pane.
assertNull(with(_ -> 999_999).callerTerminal("127.0.0.1", 55555));
}
@Test
void resolvesTheCallersPidAndCwd() {
// CB-112: the primary maps to no pane, but its PID and cwd are still readable.
ConnectionIdentity id = new ConnectionIdentity(
new PaneLocator(herdr), _ -> 999_999, pid -> pid == 999_999 ? "/main/project" : null);
ConnectionIdentity.Caller c = id.resolve("127.0.0.1", 55555);
assertNull(c.terminal(), "the primary owns no worker pane");
assertEquals(999_999, c.pid());
assertEquals("/main/project", id.cwdForPid(c.pid()), "the primary's cwd is resolvable from its PID");
assertNull(id.cwdForPid(-1), "no cwd for an unresolved PID");
}
}
@@ -0,0 +1,108 @@
package dev.ltms.bridged.mcp;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for {@link PrimaryRegistry}: pin vs record, isKnown transitions,
* null/blank guard.
*/
class PrimaryRegistryTest {
@Test
void unpinnedInitiallyUnknown() {
var reg = new PrimaryRegistry(null);
assertFalse(reg.isKnown());
assertTrue(reg.primaryTerminal().isEmpty());
}
@Test
void unpinnedAcceptsBlankAsAbsent() {
var reg = new PrimaryRegistry("");
assertFalse(reg.isKnown());
assertTrue(reg.primaryTerminal().isEmpty());
}
@Test
void pinnedFromConstruction() {
var reg = new PrimaryRegistry("term_fixed");
assertTrue(reg.isKnown());
assertEquals("term_fixed", reg.primaryTerminal().orElseThrow());
}
@Test
void recordWhenUnpinnedSetsTheTerminal() {
var reg = new PrimaryRegistry(null);
reg.record("term_abc");
assertTrue(reg.isKnown());
assertEquals("term_abc", reg.primaryTerminal().orElseThrow());
}
@Test
void recordWithNullDoesNothingWhenUnpinned() {
var reg = new PrimaryRegistry(null);
reg.record(null);
assertFalse(reg.isKnown());
}
@Test
void recordWithBlankDoesNothingWhenUnpinned() {
var reg = new PrimaryRegistry(null);
reg.record(" ");
assertFalse(reg.isKnown());
}
@Test
void recordOverwritesWhenUnpinned() {
var reg = new PrimaryRegistry(null);
reg.record("term_first");
assertEquals("term_first", reg.primaryTerminal().orElseThrow());
reg.record("term_second");
assertEquals("term_second", reg.primaryTerminal().orElseThrow());
}
@Test
void recordIsIgnoredWhenPinned() {
var reg = new PrimaryRegistry("term_pinned");
reg.record("term_other");
assertEquals("term_pinned", reg.primaryTerminal().orElseThrow(), "pinned value must survive record");
}
@Test
void nullRecordIsIgnoredWhenPinned() {
var reg = new PrimaryRegistry("term_pinned");
reg.record(null);
assertTrue(reg.isKnown());
assertEquals("term_pinned", reg.primaryTerminal().orElseThrow());
}
@Test
void blankRecordIsIgnoredWhenPinned() {
var reg = new PrimaryRegistry("term_pinned");
reg.record(" ");
assertTrue(reg.isKnown());
assertEquals("term_pinned", reg.primaryTerminal().orElseThrow());
}
@Test
void isKnownFalseAfterConstructionWithNull() {
var reg = new PrimaryRegistry(null);
assertFalse(reg.isKnown());
}
@Test
void isKnownAfterRecord() {
var reg = new PrimaryRegistry(null);
reg.record("term_x");
assertTrue(reg.isKnown());
}
@Test
void primaryTerminalRoundTrip() {
var reg = new PrimaryRegistry(null);
assertTrue(reg.primaryTerminal().isEmpty());
reg.record("term_found");
assertEquals("term_found", reg.primaryTerminal().get());
}
}
@@ -0,0 +1,127 @@
package dev.ltms.bridged.metrics;
import org.junit.jupiter.api.Test;
import java.util.LinkedHashMap;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/** CB-502 — the zero-dependency Prometheus text renderer. */
class MetricsTest {
@Test
void countersAccumulatePerLabelSet() {
Metrics m = new Metrics();
m.inc("bridged_sends_total", "outcome", "replied");
m.inc("bridged_sends_total", "outcome", "replied");
m.inc("bridged_sends_total", "outcome", "timeout");
assertEquals(2, m.count("bridged_sends_total", "outcome", "replied"));
assertEquals(1, m.count("bridged_sends_total", "outcome", "timeout"));
assertEquals(0, m.count("bridged_sends_total", "outcome", "failed"),
"an untouched series reads as zero, not an error");
}
@Test
void rendersHelpAndTypeOncePerFamily() {
Metrics m = new Metrics();
m.describe("bridged_sends_total", "counter", "Delegated sends by outcome.");
m.inc("bridged_sends_total", "outcome", "replied");
m.inc("bridged_sends_total", "outcome", "timeout");
String out = m.render();
assertEquals(1, countOccurrences(out, "# HELP bridged_sends_total"),
"HELP is per family, not per series");
assertEquals(1, countOccurrences(out, "# TYPE bridged_sends_total counter"));
assertTrue(out.contains("bridged_sends_total{outcome=\"replied\"} 1"));
assertTrue(out.contains("bridged_sends_total{outcome=\"timeout\"} 1"));
}
@Test
void labelsAreSortedSoScrapesAreByteStable() {
Metrics a = new Metrics();
a.inc("m", "b", "2", "a", "1");
Metrics b = new Metrics();
b.inc("m", "a", "1", "b", "2");
assertEquals(a.render(), b.render(), "label order in the call must not change the output");
assertTrue(a.render().contains("m{a=\"1\",b=\"2\"}"));
}
@Test
void gaugesAreEvaluatedAtScrapeTimeNotRegistrationTime() {
Metrics m = new Metrics();
int[] live = {1};
m.gauge("bridged_sessions", () -> live[0], "state", "ready");
assertTrue(m.render().contains("bridged_sessions{state=\"ready\"} 1"));
live[0] = 5;
assertTrue(m.render().contains("bridged_sessions{state=\"ready\"} 5"),
"the gauge must read current state on every scrape");
}
@Test
void aThrowingGaugeDoesNotBreakTheWholeScrape() {
Metrics m = new Metrics();
m.inc("good_total");
m.gauge("bad_gauge", () -> {
throw new IllegalStateException("herdr is down");
});
String out = assertDoesNotThrow(m::render);
assertTrue(out.contains("good_total 1"), "healthy series must still be exported");
assertFalse(out.contains("bad_gauge"), "the broken series is simply absent");
}
@Test
void collectorsDiscoverTheirLabelSetPerScrape() {
Metrics m = new Metrics();
Map<String, Number> depths = new LinkedHashMap<>();
m.collector("bridged_inbox_depth", "target", () -> depths);
assertFalse(m.render().contains("bridged_inbox_depth"), "no targets yet ⇒ no series");
depths.put("term_a", 2);
depths.put("term_b", 0);
String out = m.render();
assertTrue(out.contains("bridged_inbox_depth{target=\"term_a\"} 2"));
assertTrue(out.contains("bridged_inbox_depth{target=\"term_b\"} 0"));
}
@Test
void labelValuesAreEscaped() {
Metrics m = new Metrics();
m.inc("m", "detail", "he said \"hi\"\nand \\left");
String out = m.render();
assertTrue(out.contains("\\\""), "quotes escaped");
assertTrue(out.contains("\\n"), "newlines escaped — a raw one would corrupt the exposition");
assertTrue(out.contains("\\\\"), "backslashes escaped");
}
@Test
void wholeNumberGaugesRenderWithoutADecimalPoint() {
Metrics m = new Metrics();
m.gauge("whole", () -> 3.0);
m.gauge("fractional", () -> 1.5);
String out = m.render();
assertTrue(out.contains("whole 3"), "3.0 should not render as 3.0");
assertTrue(out.contains("fractional 1.5"));
}
@Test
void oddLabelCountIsRejected() {
Metrics m = new Metrics();
assertThrows(IllegalArgumentException.class, () -> m.inc("m", "dangling"));
}
private static int countOccurrences(String haystack, String needle) {
int n = 0;
for (int i = haystack.indexOf(needle); i >= 0; i = haystack.indexOf(needle, i + 1)) {
n++;
}
return n;
}
}
@@ -0,0 +1,113 @@
package dev.ltms.bridged.msg;
import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import org.testcontainers.containers.RabbitMQContainer;
import org.testcontainers.junit.jupiter.Container;
import org.testcontainers.junit.jupiter.Testcontainers;
import org.testcontainers.utility.DockerImageName;
import java.util.List;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Contract test for {@link AmqpReplyInbox} against a REAL broker (a RabbitMQ container — the same
* AMQP 0-9-1 the production LavinMQ deploy speaks, URI-only swap). Tagged {@code contract} so it is
* excluded from {@code mvn test}/{@code mvn clean install} (which stay hermetic and need no Docker);
* run it with Docker present via {@code mvn test -Pcontract}.
*
* <p>It proves the port contract on genuine infrastructure: eventual visibility of a published reply,
* ack removal, msgId dedup, and — the reason Stage 2 exists — cross-restart durability: an unacked
* reply survives closing the inbox and is redelivered to a fresh connection.
*/
@Tag("contract")
@Testcontainers
class AmqpReplyInboxContractTest {
@Container
static final RabbitMQContainer BROKER =
new RabbitMQContainer(DockerImageName.parse("rabbitmq:3.13-management"));
private static String uri() {
// guest/guest against the mapped AMQP port. No trailing slash: an empty path is vhost "",
// which does not exist — omitting it selects the default vhost "/".
return "amqp://guest:guest@" + BROKER.getHost() + ":" + BROKER.getAmqpPort();
}
@Test
void publishThenPeekThenAck() throws Exception {
String target = "worker-pub-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.publish(target, "m1", "hello primary");
List<ReplyInbox.InboxMessage> got = awaitPeek(inbox, target);
assertEquals(1, got.size(), "the published reply should be held for drain");
assertEquals("m1", got.getFirst().msgId());
assertEquals(target, got.getFirst().target());
assertEquals("hello primary", got.getFirst().content());
inbox.ack(target, "m1");
assertTrue(inbox.peek(target).isEmpty(), "an acked reply is dropped");
}
}
@Test
void duplicateMsgIdIsNotDoubleQueued() throws Exception {
String target = "worker-dedup-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.publish(target, "dup", "first");
awaitPeek(inbox, target);
inbox.publish(target, "dup", "second"); // same msgId — must be a no-op
// Give any erroneous second delivery time to land, then assert still exactly one.
Thread.sleep(500);
List<ReplyInbox.InboxMessage> got = inbox.peek(target);
assertEquals(1, got.size(), "a repeated msgId must not double-queue");
assertEquals("first", got.getFirst().content(), "the first payload wins");
}
}
@Test
void unackedReplySurvivesRestartAndIsRedelivered() throws Exception {
String target = "worker-durable-" + System.nanoTime();
// First "process life": publish, see it held, but crash before acking.
try (AmqpReplyInbox first = AmqpReplyInbox.open(uri())) {
first.publish(target, "persist-1", "survive me");
assertEquals(1, awaitPeek(first, target).size());
// no ack — simulate a java -jar bounce with the reply still pending
}
// Second "process life": a fresh connection to the same broker must be redelivered the reply.
try (AmqpReplyInbox second = AmqpReplyInbox.open(uri())) {
List<ReplyInbox.InboxMessage> got = awaitPeek(second, target);
assertEquals(1, got.size(), "an unacked persistent reply is redelivered after restart");
assertEquals("persist-1", got.getFirst().msgId());
assertEquals("survive me", got.getFirst().content());
second.ack(target, "persist-1");
}
// Third life: once acked, it is gone for good — durability is not endless replay.
try (AmqpReplyInbox third = AmqpReplyInbox.open(uri())) {
Thread.sleep(500);
assertTrue(third.peek(target).isEmpty(), "an acked reply does not come back on the next restart");
}
}
/** Poll peek (broker delivery is async) until a reply for {@code target} appears or ~10s elapse. */
@SuppressWarnings("BusyWait") // deliberate poll for async broker delivery, bounded by the deadline
private static List<ReplyInbox.InboxMessage> awaitPeek(AmqpReplyInbox inbox, String target)
throws InterruptedException {
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(10);
List<ReplyInbox.InboxMessage> msgs = inbox.peek(target);
while (msgs.isEmpty() && System.nanoTime() < deadline) {
Thread.sleep(50);
msgs = inbox.peek(target);
}
return msgs;
}
}
@@ -0,0 +1,145 @@
package dev.ltms.bridged.msg;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for {@link InMemoryReplyInbox}: publish, peek, ack, dedup, FIFO ordering, and thread
* safety under concurrent publish vs. drain.
*/
class InMemoryReplyInboxTest {
private final ReplyInbox inbox = new InMemoryReplyInbox();
@Test
void publishThenPeekReturnsTheMessage() {
inbox.publish("term_a", "m1", "hello");
var msgs = inbox.peek("term_a");
assertEquals(1, msgs.size());
assertEquals("m1", msgs.getFirst().msgId());
assertEquals("term_a", msgs.getFirst().target());
assertEquals("hello", msgs.getFirst().content());
}
@Test
void peekForUnknownTargetReturnsEmpty() {
assertTrue(inbox.peek("no-such-target").isEmpty());
}
@Test
void ackRemovesTheMessage() {
inbox.publish("term_a", "m1", "hello");
inbox.ack("term_a", "m1");
assertTrue(inbox.peek("term_a").isEmpty(), "after ack, the message is gone");
}
@Test
void ackForUnknownMsgIdIsNoOp() {
inbox.publish("term_a", "m1", "hello");
inbox.ack("term_a", "no-such-id"); // no-op
assertEquals(1, inbox.peek("term_a").size(), "the published message is still there");
}
@Test
void ackForUnknownTargetIsNoOp() {
inbox.ack("no-such-target", "m1"); // no-op, should not throw
}
@Test
void dedupByIdempotentMsgId() {
inbox.publish("term_a", "m1", "first");
inbox.publish("term_a", "m1", "second"); // same msgId, different content
var msgs = inbox.peek("term_a");
assertEquals(1, msgs.size(), "dedup: second publish with same msgId is a no-op");
assertEquals("first", msgs.getFirst().content(), "the original content is retained");
}
@Test
void publishesWithDifferentMsgIdsBothAppear() {
inbox.publish("term_a", "m1", "first");
inbox.publish("term_a", "m2", "second");
var msgs = inbox.peek("term_a");
assertEquals(2, msgs.size());
assertEquals("m1", msgs.get(0).msgId());
assertEquals("m2", msgs.get(1).msgId());
}
@Test
void perTargetIsolation() {
inbox.publish("term_a", "m1", "for-a");
inbox.publish("term_b", "m2", "for-b");
assertEquals(1, inbox.peek("term_a").size());
assertEquals(1, inbox.peek("term_b").size());
}
@Test
void fifoOrderIsPreserved() {
inbox.publish("term_a", "m1", "first");
inbox.publish("term_a", "m2", "second");
inbox.publish("term_a", "m3", "third");
var msgs = inbox.peek("term_a");
assertEquals(3, msgs.size());
assertEquals("m1", msgs.get(0).msgId());
assertEquals("m2", msgs.get(1).msgId());
assertEquals("m3", msgs.get(2).msgId());
}
@Test
void peekReturnsAnImmutableCopy() {
inbox.publish("term_a", "m1", "hello");
var msgs = inbox.peek("term_a");
assertThrows(UnsupportedOperationException.class, () -> msgs.add(
new ReplyInbox.InboxMessage("x", "term_a", "x")));
}
@Test
void ackRemovesOneMessageLeavesOthers() {
inbox.publish("term_a", "m1", "first");
inbox.publish("term_a", "m2", "second");
inbox.ack("term_a", "m1");
var msgs = inbox.peek("term_a");
assertEquals(1, msgs.size());
assertEquals("m2", msgs.getFirst().msgId());
}
@Test
void concurrentPublishAndDrain() throws Exception {
int msgCount = 100;
ExecutorService exec = Executors.newVirtualThreadPerTaskExecutor();
try {
// Concurrent publishers
var pubDone = new CountDownLatch(msgCount);
for (int i = 0; i < msgCount; i++) {
final int id = i;
exec.submit(() -> {
inbox.publish("term_a", "m" + id, "content-" + id);
pubDone.countDown();
});
}
// Concurrent drainer
AtomicReference<Exception> drainError = new AtomicReference<>();
exec.submit(() -> {
try {
pubDone.await();
for (int i = 0; i < 50; i++) {
var peeked = inbox.peek("term_a");
for (var msg : peeked) {
inbox.ack("term_a", msg.msgId());
}
}
} catch (Exception e) {
drainError.set(e);
}
}).get();
assertNull(drainError.get(), "concurrent drain should not throw");
} finally {
exec.shutdown();
}
}
}
@@ -0,0 +1,467 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.Injector;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* The message layer's resolution paths (CB-104 reply + CB-106 completion fallback). The turn is
* driven deterministically by feeding {@code onStatus} rather than running a real poller.
*/
class MessageServiceTest {
private static final String T = "term_a";
private final FakeHerdr herdr = new FakeHerdr().readText("BUILD GREEN: 391 files");
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final CompletionResolver completion = new CompletionResolver(agents, rendezvous);
private final Injector injector = new Injector(agents, completion);
private final MessageService messages = new MessageService(agents, injector, rendezvous);
/** Run {@code send} on a background thread; the current thread drives the worker's turn. */
private CompletableFuture<MessageService.Reply> sendAsync() {
return CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 5000));
}
private void awaitWaiting() throws InterruptedException {
long deadline = System.currentTimeMillis() + 2000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T), "send should have opened its rendezvous waiter");
}
@Test
void completionFallbackResolvesATurnThatNeverCalledBridgeReply() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
herdr.readText("$ prompt"); // pre-turn pane: no answer yet (baseline reference)
injector.onStatus(T, AgentStatus.IDLE); // deliver the task (baselines the pre-turn content)
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up and works
herdr.readText("BUILD GREEN: 391 files"); // the worker's turn produced new output
injector.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete, no bridge_reply
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, reply.outcome(),
"an unreplied but finished turn resolves via the completion fallback");
assertEquals("BUILD GREEN: 391 files", reply.text(), "the scraped transcript tail is returned");
assertTrue(reply.completed(), "a scraped completion still counts as completed");
}
@Test
void explicitBridgeReplyResolvesAsReplied() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker working
assertTrue(rendezvous.resolve(T, "LGTM ship it"), "an explicit reply resolves the send");
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, reply.outcome());
assertEquals("LGTM ship it", reply.text());
}
@Test
void aWedgedWorkerResolvesTheSendAsFailedWithTheErrorContext() throws Exception {
herdr.readText("API Error: Unable to connect to API (ENOTFOUND)");
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker starts the turn
for (int i = 0; i < 130; i++) injector.onStatus(T, AgentStatus.UNKNOWN); // then wedges (CB-109)
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome());
assertFalse(reply.completed(), "a wedge is terminal but not a successful completion");
assertTrue(reply.text().contains("ENOTFOUND"), "the error screen is carried as the failure reason");
}
@Test
void aWorkerThatVanishesMidTurnResolvesTheSendAsFailed() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker starts the turn
// The worker's pane crashes — the poller sees a *_not_found and drops it (CB-110).
injector.drop(T, new HerdrException("worker gone", "pane_not_found", null));
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome(),
"a delivered send whose worker vanishes fails instead of hanging to the timeout");
assertFalse(reply.completed());
}
// --- bridge_ask reverse rendezvous (CB-205) ------------------------------------------------
@Test
void askSurfacesAsAQuestionAndTheAnswerResumesTheSameTurn() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up, then pauses to ask
// The worker asks mid-turn on its own thread; the call blocks for the primary's answer.
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
// The primary's blocking send unblocks with the question and a turnId to answer on.
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertEquals("which config file?", q.text());
assertNotNull(q.turnId(), "a question carries a turnId to answer on");
// The primary answers via bridge_send(turnId); this blocks again for the worker's reply.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
// The worker's ask returns the answer — it resumes the same turn.
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome());
assertEquals("config.yaml", a.answer());
// The resumed worker finishes with a structured reply, resolving the answering send.
awaitWaiting(); // the answering send has (re)opened its forward waiter
assertTrue(rendezvous.resolve(T, "done"), "the worker's final reply resolves the answering send");
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
assertEquals("done", done.text());
}
@Test
void duplicateAsksFromTheSameSessionCoalesceToOneTurn() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up, then pauses to ask
// A transport retry: two concurrent bridge_ask calls from the same worker session.
CompletableFuture<MessageService.AskResult> ask1 =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
CompletableFuture<MessageService.AskResult> ask2 =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
// The primary's single blocked send surfaces exactly ONE question (one turnId).
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertEquals("which config file?", q.text());
assertNotNull(q.turnId(), "only one turnId should be minted");
// The primary answers that one turnId; both asks unblock with the same answer.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
MessageService.AskResult a1 = ask1.get(5, TimeUnit.SECONDS);
MessageService.AskResult a2 = ask2.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a1.outcome());
assertEquals("config.yaml", a1.answer());
assertEquals(MessageService.AskOutcome.ANSWERED, a2.outcome());
assertEquals("config.yaml", a2.answer());
// The resumed worker finishes with a structured reply, resolving the answering send.
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"), "the worker's final reply resolves the answering send");
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
assertEquals("done", done.text());
}
@Test
void askWithNoOpenDelegationReturnsNoWaiter() {
MessageService.AskResult r = messages.ask(T, "anyone listening?", 500);
assertEquals(MessageService.AskOutcome.NO_WAITER, r.outcome(),
"a question with no blocked send has no primary to answer it");
}
@Test
void askTimesOutWhenThePrimaryNeverAnswers() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
MessageService.AskResult r = messages.ask(T, "still there?", 200); // primary never answers
assertEquals(MessageService.AskOutcome.TIMED_OUT, r.outcome());
// The send itself already unblocked with the question the instant the ask surfaced.
MessageService.Reply q = send.get(2, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
}
@Test
void answeringAnUnknownTurnIsStale() {
MessageService.Reply r = messages.answer(T + "#999", "too late", 500);
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
"an answer to a turn that never existed (or already lapsed) is stale, not a hang");
}
// --- timeout, answer, poll, and lock-contention edges ----------------------------------
@Test
void sendTimesOutBeforeDeliveryIsQueuedNotWorking() {
// Nothing ever delivers the message and nothing resolves the send, so the reply future
// times out with delivery still incomplete — the message is still queued for the worker.
MessageService.Reply r = messages.send(T, "never delivered", 50);
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome(),
"an undelivered send that times out is still queued, not working");
assertNull(r.text());
}
@Test
void sendTimesOutAfterDeliveryIsStillWorking() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 300));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver — the delivered future now completes
injector.onStatus(T, AgentStatus.WORKING); // worker starts but never replies
// No rendezvous.resolve(T, ...) — the reply future rides out its short timeout.
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, r.outcome(),
"a delivered send whose worker never replies times out as still working");
}
@Test
void answerTimesOutWhenTheResumedWorkerNeverReplies() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertNotNull(q.turnId());
// The primary answers, unblocking the worker; but the worker never sends the follow-up
// bridge_reply, so the answering send rides out its short window as still-working.
MessageService.Reply answer = messages.answer(q.turnId(), "config.yaml", 200);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answer.outcome(),
"an answered worker that never replies times out as still working");
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome());
assertEquals("config.yaml", a.answer());
}
@Test
void pollReturnsNullForAnUnknownTicket() {
assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown");
}
@Test
void pollReportsACompletedTicket() throws Exception {
String ticket = messages.sendAsync(T, "long task");
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker works
assertTrue(rendezvous.resolve(T, "async result"), "a reply resolves the async send");
// Wait for the background send to finish and publish a DONE view.
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 2000;
while (view == null || view.phase() != MessageService.Phase.DONE) {
if (System.currentTimeMillis() >= deadline) break;
view = messages.poll(ticket);
//noinspection BusyWait
Thread.sleep(5);
}
assertNotNull(view, "a resolved async send must become DONE");
assertEquals(MessageService.Phase.DONE, view.phase());
assertEquals("async result", view.reply(), "the completed ticket reports the reply");
assertEquals("reply", view.replySource(), "a structured bridge_reply is sourced from 'reply'");
}
@Test
void concurrentSendToSameSessionWhileFirstHoldsItIsBusy() throws Exception {
CompletableFuture<MessageService.Reply> first =
CompletableFuture.supplyAsync(() -> messages.send(T, "first", 5000));
awaitWaiting(); // the first send now holds the session lock, blocked on its reply
// A second send to the SAME session cannot take the lock within its short window.
MessageService.Reply busy = messages.send(T, "second", 100);
assertEquals(MessageService.Outcome.BUSY, busy.outcome(),
"a second send while another holds the session is busy, not a hang");
assertNull(busy.text());
// Release the first send so it resolves cleanly and the test thread is not left pinned.
injector.onStatus(T, AgentStatus.IDLE); // deliver the first message
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up
assertTrue(rendezvous.resolve(T, "first done"), "the first send resolves with a reply");
MessageService.Reply firstReply = first.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, firstReply.outcome());
assertEquals("first done", firstReply.text());
}
// --- CB-307 reply inbox ----------------------------------------------------------------
@Test
void replyQueuesInInboxWhenNoSendIsOpen() {
// No send is open for this session — reply should queue in the inbox.
assertTrue(messages.reply(T, "queued-text"), "reply should succeed (queued)");
var drained = messages.drainReplies(T);
assertEquals(1, drained.size());
assertEquals("queued-text", drained.getFirst().content());
}
@Test
void replyResolvesOpenSendDoesNotQueue() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitUninterruptibly(T);
// An explicit reply resolves the open send.
assertTrue(messages.reply(T, "send-resolved"), "reply should succeed (resolved live send)");
// The inbox should be empty — the reply went to the send, not the inbox.
assertTrue(messages.drainReplies(T).isEmpty(), "no reply in the inbox");
MessageService.Reply r = send.get(3, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, r.outcome());
assertEquals("send-resolved", r.text());
}
@Test
void drainRepliesReturnsAllPendingThenEmptyOnNextCall() {
messages.reply(T, "msg-1");
messages.reply(T, "msg-2");
var first = messages.drainReplies(T);
assertEquals(2, first.size());
var second = messages.drainReplies(T);
assertTrue(second.isEmpty(), "second drain should be empty (acked)");
}
@Test
void aQuestionIsNeverQueuedInTheInbox() {
// No send is open — bridge_ask with no delegation returns NO_WAITER,
// and the question text MUST NOT appear in the reply inbox.
// The inbox is only fed by MessageService.reply(), not by bridge_ask.
MessageService.AskResult r = messages.ask(T, "anyone there?", 500);
assertEquals(MessageService.AskOutcome.NO_WAITER, r.outcome(),
"bridge_ask with no open delegation must return NO_WAITER, never queued");
assertTrue(messages.drainReplies(T).isEmpty(), "questions must never be queued");
}
@Test
void completionFallbackIsNeverQueued() throws Exception {
// The fallback resolves a captured waiter, never the inbox.
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitUninterruptibly(T);
injectDelivery();
// The worker never sends bridge_reply, but the turn completes.
herdr.readText("done-scraped");
completion.onTurnComplete(T); // The fallback arms and resolves the captured waiter.
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, r.outcome());
// The inbox should be empty — the reply went to the captured waiter.
assertTrue(messages.drainReplies(T).isEmpty(), "completion fallback must not queue");
}
// --- helpers ---------------------------------------------------------------------------
/** Like {@link #awaitWaiting()} but rethrows as unchecked. */
private void awaitUninterruptibly(String session) {
try {
awaitWaiting();
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException(e);
}
}
/** Set up a delivered turn so the worker is working, ready for an ask or completion. */
private void injectDelivery() {
herdr.readText("$ prompt"); // pre-turn content baseline
injector.onStatus(T, AgentStatus.IDLE); // deliver the task
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up
}
// --- CB-516: a released session must not leave a send hanging ------------------------------
/**
* The bug this fixes: tearing a worker down left its rendezvous waiter open, so a blocking send
* kept blocking and an async one kept reporting PENDING until the 30-minute async timeout —
* even though the worker provably no longer existed.
*/
@Test
void abandonFailsASendThatIsStillWaitingOnAReleasedSession() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
awaitWaiting();
assertTrue(messages.abandon(T, "session released"), "a live waiter is abandoned");
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, r.outcome(),
"an abandoned send fails rather than riding out its timeout");
assertEquals("session released", r.text(), "the caller is told why");
}
@Test
void abandonIsANoOpWhenNobodyIsWaiting() {
assertFalse(messages.abandon(T, "session released"),
"no open send ⇒ nothing to abandon");
}
@Test
void abandonDoesNotOverwriteAnAlreadyResolvedSend() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
awaitWaiting();
assertTrue(rendezvous.resolve(T, "the real answer"));
assertFalse(messages.abandon(T, "session released"),
"a send already answered by the worker must not be clobbered");
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals("the real answer", r.text());
}
/** The async path is the one that hung: poll must report FAILED, not PENDING forever. */
@Test
void anAbandonedAsyncTaskPollsAsFailedNotPending() throws Exception {
String ticket = messages.sendAsync(T, "long task");
awaitWaiting();
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase());
messages.abandon(T, "session released");
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 3000;
while (System.currentTimeMillis() < deadline) {
view = messages.poll(ticket);
if (view.phase() != MessageService.Phase.PENDING) break;
Thread.sleep(10);
}
assertNotNull(view);
assertEquals(MessageService.Phase.FAILED, view.phase(),
"a delegation whose worker is gone must not keep reporting PENDING");
assertTrue(view.detail() != null && view.detail().contains("released"),
"and the detail says why, rather than 'worker unknown'");
}
}
@@ -0,0 +1,102 @@
package dev.ltms.bridged.msg;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CompletableFuture;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* The reverse rendezvous (CB-205): the {@code bridge_ask} registry that lets a worker pause mid-turn
* to ask the primary. Unit-level — the message-layer round-trip is covered in {@link MessageServiceTest}.
*/
class RendezvousTest {
private static final String W = "term_a";
private final Rendezvous rendezvous = new Rendezvous();
@Test
void openAskMintsAUniqueTurnScopedToItsSessionAndCoalescesDuplicates() {
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
Rendezvous.AskTicket t2 = rendezvous.openAsk(W);
assertEquals(t1.turnId(), t2.turnId(), "duplicate asks from the same session coalesce onto one turn");
assertTrue(t1.fresh(), "the first ask freshly opens the turn");
assertFalse(t2.fresh(), "the coalesced ask rides the existing turn");
assertTrue(t1.turnId().startsWith(W + "#"), "the turnId is scoped to the worker session");
assertEquals(W, rendezvous.askSession(t1.turnId()));
}
@Test
void openAskAfterCloseMintsANewTurn() {
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
rendezvous.closeAsk(t1.turnId());
Rendezvous.AskTicket t2 = rendezvous.openAsk(W);
assertNotEquals(t1.turnId(), t2.turnId(), "after closing, a new ask gets a fresh turnId");
assertTrue(t2.fresh(), "the reopened ask is fresh");
assertEquals(W, rendezvous.askSession(t2.turnId()));
}
@Test
void answerAskCompletesTheWaitersFuture() {
Rendezvous.AskTicket t = rendezvous.openAsk(W);
assertTrue(rendezvous.answerAsk(t.turnId(), "config.yaml"), "answering an open ask succeeds");
assertEquals("config.yaml", t.answer().getNow(null), "the answer reaches the blocked worker");
}
@Test
void answerAskOnAnUnknownTurnIsFalse() {
assertFalse(rendezvous.answerAsk("no-such#1", "x"), "an answer to an unknown turn is a no-op");
}
@Test
void resolveQuestionResolvesAnOpenSendWithTheQuestionKindAndTurnId() {
CompletableFuture<Rendezvous.Resolution> send = rendezvous.open(W);
assertTrue(rendezvous.resolveQuestion(W, "which config?", W + "#7"),
"the question resolves the primary's open send");
Rendezvous.Resolution r = send.getNow(null);
assertEquals(Rendezvous.Kind.QUESTION, r.kind());
assertEquals("which config?", r.text());
assertEquals(W + "#7", r.turnId(), "the turnId rides along so the primary can answer");
}
@Test
void resolveQuestionWithNoOpenSendIsFalse() {
assertFalse(rendezvous.resolveQuestion(W, "anyone?", W + "#1"),
"no blocked send means no primary to surface the question to");
}
@Test
void resolveCompletionTwiceIsANoOpTheSecondTime() {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(W);
assertTrue(rendezvous.resolveCompletion(waiter, "first scrape"), "the first completion resolves");
assertFalse(rendezvous.resolveCompletion(waiter, "second scrape"),
"a second completion on an already-resolved waiter returns false");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
assertEquals("first scrape", waiter.getNow(null).text(),
"the first resolution wins; the stored value is unchanged");
}
@Test
void resolveFailureTwiceIsANoOpTheSecondTime() {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(W);
assertTrue(rendezvous.resolveFailure(waiter, "first reason"), "the first failure resolves");
assertFalse(rendezvous.resolveFailure(waiter, "second reason"),
"a second failure on an already-resolved waiter returns false");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
assertEquals("first reason", waiter.getNow(null).text(),
"the first resolution wins; the stored value is unchanged");
}
@Test
void closeAskRemovesTheTurn() {
Rendezvous.AskTicket t = rendezvous.openAsk(W);
rendezvous.closeAsk(t.turnId());
assertNull(rendezvous.askSession(t.turnId()), "a closed ask is forgotten");
assertFalse(rendezvous.answerAsk(t.turnId(), "late"), "a closed ask can no longer be answered");
}
}
@@ -0,0 +1,303 @@
package dev.ltms.bridged.msg;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.Collections;
import java.util.List;
import java.util.Map;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for {@link ReplyPushLoop}: decision logic, nudge injection, idempotency,
* bounded reminders, and stop conditions.
*
* <p>Uses a {@link RecordingHerdrClient} that synchronizes access to its call list so the
* scheduler thread and test thread never have memory ordering issues. The {@code decide()}
* tests use a simple client with no concurrency concern.
*/
class ReplyPushLoopTest {
private static final String PRIMARY = "term_primary";
private static final String WORKER = "term_worker";
private static final ObjectMapper MAPPER = new ObjectMapper();
private PrimaryRegistry registry;
private AgentControl agents;
private InMemoryReplyInbox inbox;
private ScheduledExecutorService scheduler;
@BeforeEach
void setUp() {
registry = new PrimaryRegistry(PRIMARY);
inbox = new InMemoryReplyInbox();
scheduler = Executors.newSingleThreadScheduledExecutor();
}
@AfterEach
void tearDown() {
scheduler.shutdownNow();
}
// --- decide() logic ------------------------------------------------------------------------
@Test
void decideWithoutPrimaryIsStop() {
agents = agentWithStatus("idle");
var loop = new ReplyPushLoop(
new PrimaryRegistry(null), agents, inbox, scheduler, 5, 100);
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(WORKER, 0));
}
@Test
void decideWithEmptyInboxIsStop() {
agents = agentWithStatus("idle");
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
}
@Test
void decideAtCapIsStop() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100).decide(WORKER, 2));
}
@Test
void decideUnderCapWithInjectablePrimaryIsInject() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
}
@Test
void decideUnderCapWithBlockedPrimaryIsInject() {
agents = agentWithStatus("blocked");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
"BLOCKED is injectable");
}
@Test
void decideUnderCapWithDonePrimaryIsInject() {
agents = agentWithStatus("done");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
"DONE is injectable");
}
@Test
void decideUnderCapWithBusyPrimaryIsWaitBusy() {
agents = agentWithStatus("working");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
}
@Test
void decideUnderCapWithUnknownPrimaryIsWaitBusy() {
agents = agentWithStatus("unknown");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
}
@Test
void decideStopsAfterInboxIsEmptied() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
inbox.ack(WORKER, "m1");
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
}
// --- onReplyQueued integration -------------------------------------------------------------
@Test
void injectablePrimaryCausesExactlyOneNudge() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
loop(1, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"one nudge (2 agent.send calls) should have been sent");
// Exactly one nudge = exactly 2 agent.send calls (text + submit)
assertEquals(2, rec.sendCount());
assertTrue(rec.sentParams().stream()
.anyMatch(e -> e.getValue().toString().contains("bridge_poll")),
"nudge text should contain bridge_poll");
}
@Test
void onReplyQueuedIsIdempotentPerTarget() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
var loop = loop(1, 100);
loop.onReplyQueued(WORKER);
loop.onReplyQueued(WORKER); // second call — should be a no-op
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"expected exactly one nudge (2 sends)");
Thread.sleep(200);
assertEquals(2, rec.sendCount(),
"second onReplyQueued must not trigger another nudge");
}
@Test
void sendsUpToCapThenStops() throws Exception {
int cap = 2;
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
rec.sendLatch = new CountDownLatch(cap * 2);
loop(cap, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(5, TimeUnit.SECONDS),
cap + " nudges (" + (cap * 2) + " sends) should have fired");
Thread.sleep(300);
assertEquals(cap * 2, rec.sendCount(),
"exactly " + (cap * 2) + " agent.send calls (cap=" + cap + ")");
}
// --- nudge format --------------------------------------------------------------------------
@Test
void nudgeFormatIsCorrect() {
String nudge = ReplyPushLoop.NUDGE_FORMAT.formatted(WORKER, WORKER);
assertTrue(nudge.contains("Worker term_worker"));
assertTrue(nudge.contains("bridge_poll(target=term_worker)"));
}
// --- metrics (CB-512) ----------------------------------------------------------------------
@Test
void successfulNudgeIncrementsDelivered() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
Metrics metrics = new Metrics();
loop(1, 50, metrics).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"one nudge (2 agent.send calls) should have been sent");
// The delivered count is bumped on the scheduler thread right after the send that releases
// the latch — settle briefly so the counter is published before we read it.
Thread.sleep(200);
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"),
"a successfully sent nudge must count as delivered");
}
@Test
void reminderCapIncrementsExhausted() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
Metrics metrics = new Metrics();
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100, metrics).decide(WORKER, 2));
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
"hitting the reminder cap must count as exhausted");
assertEquals(0, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"));
}
// --- helpers -------------------------------------------------------------------------------
private ReplyPushLoop loop() {
return loop(5, 100);
}
private ReplyPushLoop loop(int maxReminders, long backoffMs) {
return new ReplyPushLoop(registry, agents, inbox, scheduler, maxReminders, backoffMs);
}
private ReplyPushLoop loop(int maxReminders, long backoffMs, Metrics metrics) {
return new ReplyPushLoop(registry, agents, inbox, scheduler, maxReminders, backoffMs, metrics);
}
private static AgentControl agentWithStatus(String status) {
return new AgentControl(new FakeHerdrClient(status));
}
/** Non-recording (single-threaded) fake — safe for decide() tests. */
private static final class FakeHerdrClient implements HerdrClient {
private final String agentStatus;
FakeHerdrClient(String agentStatus) {
this.agentStatus = agentStatus;
}
@Override
public JsonNode call(String method, Object params) {
if ("agent.get".equals(method)) {
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", PRIMARY)
.put("agent_status", agentStatus));
}
return MAPPER.createObjectNode();
}
@Override
public void close() {
}
}
/**
* Thread-safe recording fake that counts agent.send calls. Uses synchronized access
* so the scheduler thread and test thread never race.
*/
private static final class RecordingHerdrClient implements HerdrClient {
private final List<Map.Entry<String, Object>> calls =
Collections.synchronizedList(new ArrayList<>());
volatile CountDownLatch sendLatch = new CountDownLatch(2);
@Override
public JsonNode call(String method, Object params) {
if ("agent.get".equals(method)) {
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", PRIMARY)
.put("agent_status", "idle")); // recording double is always injectable
}
if ("agent.send".equals(method)) {
calls.add(Map.entry(method, params));
sendLatch.countDown();
}
return MAPPER.createObjectNode();
}
long sendCount() {
return calls.size();
}
List<Map.Entry<String, Object>> sentParams() {
return List.copyOf(calls);
}
@Override
public void close() {
}
}
private static RecordingHerdrClient recordingClient() {
return new RecordingHerdrClient();
}
}
@@ -0,0 +1,225 @@
package dev.ltms.bridged.rest;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-501/505 enforcement over real HTTP. The unit tests pin the policy; these pin that the policy
* is actually reached from a request — a rule enforced nowhere is not a control.
*/
class BridgedAppAuthTest {
private final HttpClient http = HttpClient.newHttpClient();
private Javalin app;
private Metrics metrics;
@AfterEach
void stop() {
if (app != null) app.stop();
}
/**
* Start the app with the given identity/auth wiring.
*
* @param pid the PID every connection resolves to — {@link FakeHerdr#WORKER_PID} makes the
* caller worker {@code term_a}, anything else makes it a non-worker
*/
private int start(long pid, boolean tokenMode, String token) {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
k -> "BRIDGED_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, new Rendezvous());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
CallerResolver callers = tokenMode
? new CallerResolver(identity, true, token)
: new CallerResolver(identity);
metrics = BridgedMetrics.create(sessions, new dev.ltms.bridged.msg.InMemoryReplyInbox());
app = new BridgedApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
callers, metrics).build().start("127.0.0.1", 0);
return app.port();
}
private HttpResponse<String> send(int port, String method, String path, String body, String auth)
throws Exception {
HttpRequest.Builder b = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path))
.header("Content-Type", "application/json");
if (auth != null) {
b.header("Authorization", auth);
}
b = switch (method) {
case "POST" -> b.POST(body == null
? HttpRequest.BodyPublishers.noBody()
: HttpRequest.BodyPublishers.ofString(body));
case "DELETE" -> b.DELETE();
default -> b.GET();
};
return http.send(b.build(), HttpResponse.BodyHandlers.ofString());
}
// --- loopback-trust: the caller is the primary -------------------------------------------
@Test
void thePrimaryMayOrchestrateButMayNotForgeAWorkerReply() throws Exception {
int port = start(999_999, false, null); // no pane ⇒ primary
HttpResponse<String> read = send(port, "GET", "/profiles", null, null);
assertEquals(200, read.statusCode(), "the primary may observe");
HttpResponse<String> reply = send(port, "POST", "/sessions/term_a/reply",
"{\"content\":\"forged\"}", null);
assertEquals(403, reply.statusCode(),
"a forged reply would resolve the rendezvous the primary is itself waiting on");
assertTrue(reply.body().contains("forbidden"));
}
// --- loopback-trust: the caller is a worker ------------------------------------------------
@Test
void aWorkerMayReplyAsItselfButNotAsAnother() throws Exception {
int port = start(FakeHerdr.WORKER_PID, false, null); // resolves to term_a
HttpResponse<String> own = send(port, "POST", "/sessions/term_a/reply",
"{\"content\":\"done\"}", null);
assertEquals(200, own.statusCode(), "a worker replies on its own session");
HttpResponse<String> other = send(port, "POST", "/sessions/term_b/reply",
"{\"content\":\"not mine\"}", null);
assertEquals(403, other.statusCode(),
"REST trusted the path id before CB-505; this is the hole being closed");
}
@Test
void aWorkerMayNotOrchestrate() throws Exception {
int port = start(FakeHerdr.WORKER_PID, false, null);
assertEquals(403, send(port, "POST", "/workers", null, null).statusCode(),
"a worker spawning workers would be escalating into the orchestrator role");
assertEquals(403, send(port, "DELETE", "/workers/w2:p7", null, null).statusCode());
assertEquals(403, send(port, "POST", "/sessions/term_b/message",
"{\"content\":\"hi\"}", null).statusCode());
assertEquals(403, send(port, "GET", "/sessions/term_a/replies", null, null).statusCode(),
"draining an inbox is the primary's collection step");
}
// --- token mode ---------------------------------------------------------------------------
@Test
void tokenModeRejectsAnUncredentialedNonWorkerWith401() throws Exception {
int port = start(999_999, true, "s3cret");
HttpResponse<String> res = send(port, "GET", "/profiles", null, null);
assertEquals(401, res.statusCode(), "no credential ⇒ authenticated as nothing");
assertTrue(res.body().contains("unauthenticated"));
}
@Test
void tokenModeAcceptsAValidBearerToken() throws Exception {
int port = start(999_999, true, "s3cret");
assertEquals(200, send(port, "GET", "/profiles", null, "Bearer s3cret").statusCode());
}
@Test
void tokenModeStillHonoursConnectionDerivedWorkerIdentity() throws Exception {
// The fleet must keep working when auth is switched on: a worker presents no token, and
// must still be able to reply.
int port = start(FakeHerdr.WORKER_PID, true, "s3cret");
assertEquals(200, send(port, "POST", "/sessions/term_a/reply",
"{\"content\":\"done\"}", null).statusCode());
}
// --- health, metrics ----------------------------------------------------------------------
@Test
void healthzStaysOpenWithoutCredentials() throws Exception {
int port = start(999_999, true, "s3cret");
assertEquals(200, send(port, "GET", "/healthz", null, null).statusCode(),
"a supervisor must be able to probe liveness before any credential is configured");
}
@Test
void metricsRequireAuthenticationAndRenderPrometheusText() throws Exception {
int port = start(999_999, true, "s3cret");
assertEquals(401, send(port, "GET", "/metrics", null, null).statusCode());
HttpResponse<String> ok = send(port, "GET", "/metrics", null, "Bearer s3cret");
assertEquals(200, ok.statusCode());
assertTrue(ok.headers().firstValue("Content-Type").orElse("").startsWith("text/plain"));
assertTrue(ok.body().contains("bridged_sessions{state=\"ready\"}"),
"the session census gauge is exported even when empty");
}
@Test
void refusalsAreCounted() throws Exception {
int port = start(999_999, true, "s3cret");
send(port, "GET", "/profiles", null, null); // 401
send(port, "POST", "/sessions/term_a/reply", "{}", "Bearer s3cret"); // 403
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "unauthenticated"));
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "forbidden"));
}
// --- legacy constructor -------------------------------------------------------------------
@Test
void theLegacyConstructorLeavesAuthorizationOff() throws Exception {
// The 29 pre-existing acceptance tests rely on this: no auth fixture, no enforcement.
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
Map.of(wcfg.profile(), wcfg), wcfg.profile(), _ -> "tok");
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous());
app = new BridgedApp(herdr, workers, sessions, messages, sessions.asPresence(), null)
.build().start("127.0.0.1", 0);
assertEquals(200, send(app.port(), "POST", "/sessions/term_a/reply",
"{\"content\":\"x\"}", null).statusCode());
assertEquals(404, send(app.port(), "GET", "/metrics", null, null).statusCode(),
"no registry supplied ⇒ the endpoint is not mounted at all");
}
}
@@ -7,7 +7,16 @@ import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.WorkerService;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.Worktrees;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
@@ -32,10 +41,13 @@ class BridgedAppTest {
private final ObjectMapper mapper = new ObjectMapper();
private final HttpClient http = HttpClient.newHttpClient();
private WorkerPresence presence;
private Javalin app;
private StatusPoller poller;
@AfterEach
void stop() {
if (poller != null) poller.stop();
if (app != null) app.stop();
}
@@ -44,13 +56,27 @@ class BridgedAppTest {
}
private int start(FakeHerdr herdr, String workerBaseUrl, Set<String> allow, String placement) {
return start(herdr, workerBaseUrl, allow, placement, new GitWorktrees());
}
private int start(FakeHerdr herdr, String workerBaseUrl, Set<String> allow, String placement, Worktrees worktrees) {
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
"ltms-local", workerBaseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
placement, "bridged-workers", "worker: {profile} #{n}");
WorkerService workers = new WorkerService(
new AgentControl(herdr), new WorkspaceControl(herdr), new SubscriptionGuard(allow), wcfg,
placement, "bridged-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(allow),
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
k -> "BRIDGED_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
app = new BridgedApp(herdr, workers).build().start("127.0.0.1", 0);
SessionManager sessions = new SessionManager(workers, worktrees);
this.presence = sessions.asPresence();
Injector injector = new Injector(agents);
poller = new StatusPoller(agents, injector, 5); // delivers when the fake reports idle
poller.start();
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
app = new BridgedApp(herdr, workers, sessions, messages, this.presence, null)
.build().start("127.0.0.1", 0);
return app.port();
}
@@ -58,6 +84,18 @@ class BridgedAppTest {
return start(new FakeHerdr(), "http://gx00.gw:8000", Set.of("gx00.gw"));
}
private HttpResponse<String> postMessage(int port, String json) throws Exception {
return postJson(port, "/sessions/term_a/message", json);
}
private HttpResponse<String> postJson(int port, String path, String json) throws Exception {
HttpRequest r = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path))
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(json)).build();
return http.send(r, HttpResponse.BodyHandlers.ofString());
}
private HttpResponse<String> req(int port, String method, String path) throws Exception {
HttpRequest.Builder b = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path));
b = switch (method) {
@@ -85,7 +123,8 @@ class BridgedAppTest {
@Test
void healthzDegradedWhenHerdrDown() throws Exception {
int port = start(new FakeHerdr().healthy(false), "http://gx00.gw:8000", Set.of("gx00.gw"));
FakeHerdr down = new FakeHerdr().healthy(false);
int port = start(down, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "GET", "/healthz");
assertEquals(503, res.statusCode());
assertEquals("degraded", mapper.readTree(res.body()).get("status").asText());
@@ -117,8 +156,8 @@ class BridgedAppTest {
HttpResponse<String> res = req(port, "POST", "/workers");
assertEquals(201, res.statusCode());
JsonNode body = mapper.readTree(res.body());
assertEquals("w9:pW", body.get("paneId").asText());
assertEquals("w9:t2", body.get("tabId").asText());
assertEquals("w9:pW_1", body.get("paneId").asText());
assertEquals("spawning", body.get("state").asText());
// Subscription boundary: agent.start carried base_url + token in its env map.
Map<String, Object> start = params(herdr, "agent.start");
@@ -137,6 +176,58 @@ class BridgedAppTest {
"tab label carries the worker number so siblings stay distinct");
}
@Test
void profilesEndpointListsConfiguredProfilesAndDefault() throws Exception {
int port = startHealthy();
JsonNode body = mapper.readTree(req(port, "GET", "/profiles").body());
assertEquals("ltms-local", body.get("default").asText());
assertEquals("ltms-local", body.get("profiles").get(0).asText());
}
@Test
void workersEndpointReturnsRegistryRosterWithLiveStatus() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"), "tab",
new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt"));
HttpResponse<String> spawn = req(port, "POST", "/workers?worktree=true&ticket=cb-304");
assertEquals(201, spawn.statusCode());
JsonNode spawned = mapper.readTree(spawn.body());
String paneId = spawned.get("paneId").asText();
HttpResponse<String> res = req(port, "GET", "/workers");
assertEquals(200, res.statusCode());
JsonNode workers = mapper.readTree(res.body()).get("workers");
assertEquals(1, workers.size());
JsonNode w = workers.get(0);
assertEquals(spawned.get("terminalId").asText(), w.get("sessionId").asText());
assertEquals(paneId, w.get("paneId").asText());
assertEquals("ltms-local", w.get("profile").asText());
assertEquals("spawning", w.get("state").asText());
assertTrue(w.has("worktree"), "worktree-backed session exposes worktree");
assertTrue(w.has("branch"), "worktree-backed session exposes branch");
assertEquals("unknown", w.get("liveStatus").asText(),
"liveStatus is unknown when herdr has no matching pane");
}
@Test
void spawnWithACwdParamRootsTheWorkerThere() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(201, req(port, "POST", "/workers?cwd=/tmp/proj").statusCode());
assertEquals("/tmp/proj", params(herdr, "agent.start").get("cwd"), "the worker starts in cwd");
}
@Test
void spawnWithAnUnknownProfileIs400() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "POST", "/workers?profile=nope");
assertEquals(400, res.statusCode());
assertEquals("unknown_profile", mapper.readTree(res.body()).get("error").asText());
assertFalse(herdr.called("agent.start"), "an unknown profile must not spawn anything");
}
@Test
void spawnWorkerReusesExistingWorkerSpace() throws Exception {
// A space labelled "bridged-workers" already exists → no second workspace.create.
@@ -211,6 +302,137 @@ class BridgedAppTest {
assertEquals("w9:t2", params(herdr, "tab.close").get("tab_id"));
}
@Test
void messageReturnsTheWorkersStructuredReply() throws Exception {
// CB-104 (option C): the blocking send resolves on the worker's bridge_reply, not a scrape.
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // poller delivers the injection
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
var send = java.util.concurrent.CompletableFuture.supplyAsync(() -> {
try { return postMessage(port, "{\"content\":\"review this\",\"timeoutMs\":4000}"); }
catch (Exception e) { throw new RuntimeException(e); }
});
// Give the background send thread time to open its rendezvous waiter (CB-307: reply now
// queues in the inbox if no waiter is open, which would break the round-trip).
Thread.sleep(200);
HttpResponse<String> reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"LGTM ship it\"}");
assertEquals(200, reply.statusCode());
HttpResponse<String> res = send.get(6, java.util.concurrent.TimeUnit.SECONDS);
assertEquals(200, res.statusCode());
assertEquals("LGTM ship it", mapper.readTree(res.body()).get("reply").asText());
// (injection via agent.send is covered deterministically by the timeout-working test)
}
@Test
void asyncSendReturnsATicketThenPollReportsTheReply() throws Exception {
// CB-107 fire-and-poll: wait:false returns a ticket immediately; the result is polled.
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // poller delivers the injection
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> accepted = postMessage(port, "{\"content\":\"do it\",\"wait\":false}");
assertEquals(202, accepted.statusCode());
String ticket = mapper.readTree(accepted.body()).get("ticket").asText();
assertFalse(ticket.isBlank(), "an async send must return a ticket");
// Give the background async send thread time to open its rendezvous waiter.
Thread.sleep(200);
HttpResponse<String> reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"async LGTM\"}");
assertEquals(200, reply.statusCode());
// Polling the ticket now reports the finished delegation and its reply.
JsonNode task;
long deadline = System.currentTimeMillis() + 3000;
do {
task = mapper.readTree(req(port, "GET", "/tasks/" + ticket).body());
if ("done".equals(task.path("phase").asText())) break;
//noinspection BusyWait
Thread.sleep(10);
} while (System.currentTimeMillis() < deadline);
assertEquals("done", task.get("phase").asText());
assertEquals("async LGTM", task.get("reply").asText());
assertEquals("reply", task.get("replySource").asText());
}
@Test
void pollUnknownTicketIs404() throws Exception {
int port = startHealthy();
HttpResponse<String> res = req(port, "GET", "/tasks/task-999");
assertEquals(404, res.statusCode());
assertEquals("unknown_ticket", mapper.readTree(res.body()).get("error").asText());
}
@Test
void replyWithNoPendingSendQueuesInsteadOfConflict() throws Exception {
// CB-307: a reply with no open send now queues in the inbox, not a 409 conflict.
int port = startHealthy();
HttpResponse<String> res = postJson(port, "/sessions/term_a/reply", "{\"content\":\"orphan\"}");
assertEquals(200, res.statusCode());
// The queued reply is drainable.
HttpResponse<String> drain = req(port, "GET", "/sessions/term_a/replies");
assertEquals(200, drain.statusCode());
JsonNode body = mapper.readTree(drain.body());
assertEquals(1, body.get("replies").size());
assertEquals("orphan", body.get("replies").get(0).get("content").asText());
}
@Test
void drainRepliesReturnsEmptyForNoReplies() throws Exception {
int port = startHealthy();
HttpResponse<String> res = req(port, "GET", "/sessions/term_a/replies");
assertEquals(200, res.statusCode());
assertEquals(0, mapper.readTree(res.body()).get("replies").size());
}
@Test
void messageTimesOutQueuedWhenWorkerNeverInjectable() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("working"); // never injectable → never delivered
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":150}");
assertEquals(202, res.statusCode());
assertEquals("queued", mapper.readTree(res.body()).get("status").asText());
assertFalse(herdr.called("agent.send"), "no injection while the worker is mid-turn");
}
@Test
void messageTimesOutWorkingWhenDeliveredButNoReply() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // delivered, but nobody replies
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":250}");
assertEquals(202, res.statusCode());
assertEquals("working", mapper.readTree(res.body()).get("status").asText());
assertTrue(herdr.called("agent.send"), "message was injected");
}
@Test
void messageRejectsBlankContent() throws Exception {
int port = startHealthy();
assertEquals(400, postMessage(port, "{}").statusCode());
}
@Test
void sessionStatusReportsLiveAgentStatus() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("blocked");
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "GET", "/sessions/term_a/status");
assertEquals(200, res.statusCode());
assertEquals("blocked", mapper.readTree(res.body()).get("status").asText());
}
@Test
void sessionStatusReportsReadinessFromMcpPresence() throws Exception {
int port = start(new FakeHerdr(), "http://gx00.gw:8000", Set.of("gx00.gw"));
// Not yet seen on the bridge MCP → not ready.
assertFalse(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
// Worker connects its MCP client → available.
presence.markPresent("term_a");
assertTrue(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
}
@Test
void stopWorkerInPanePlacementClosesOnlyThePane() throws Exception {
FakeHerdr herdr = new FakeHerdr();
@@ -0,0 +1,130 @@
package dev.ltms.bridged.session;
import java.util.Collections;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.CopyOnWriteArrayList;
/** Recording fake {@link Worktrees} for CB-301-ext acceptance tests (no live git). */
public final class FakeWorktrees implements Worktrees {
public record AddCall(String repoRoot, String branch, String baseRef) {
}
public record RemoveCall(String repoRoot, String worktreePath) {
}
public record OverlayCall(String repoRoot, String worktreePath,
List<String> requested, List<String> copied, List<String> skipWorktree) {
}
public record RepoRootCall(String cwd) {
}
private final List<AddCall> addCalls = new CopyOnWriteArrayList<>();
private final List<RemoveCall> removeCalls = new CopyOnWriteArrayList<>();
private final List<OverlayCall> overlayCalls = new CopyOnWriteArrayList<>();
private final List<RepoRootCall> repoRootCalls = new CopyOnWriteArrayList<>();
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
private volatile RuntimeException addFailure;
private volatile String repoRoot = "/repo";
private volatile String prefix = "/worktrees";
public FakeWorktrees withRepoRoot(String root) {
this.repoRoot = root;
return this;
}
public FakeWorktrees withPrefix(String prefix) {
this.prefix = prefix;
return this;
}
/** Paths that exist in the primary repo and will be copied to the worktree. */
public FakeWorktrees exists(String... paths) {
Collections.addAll(existingPaths, paths);
return this;
}
/** Paths that exist AND are tracked, so overlayParity should --skip-worktree them. */
public FakeWorktrees track(String... paths) {
exists(paths);
Collections.addAll(trackedPaths, paths);
return this;
}
/** Make subsequent {@link #add} calls throw (simulates git worktree add failure). */
public FakeWorktrees failAdd(String message) {
this.addFailure = new WorktreeException(message);
return this;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
addCalls.add(new AddCall(repoRoot, branch, baseRef));
if (addFailure != null) {
throw addFailure;
}
// The branch already carries a unique nonce, so the derived path is distinct per acquire
// without an extra counter — keep it a pure function of the branch the test can predict.
return prefix + "/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
removeCalls.add(new RemoveCall(repoRoot, worktreePath));
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
List<String> copied = new java.util.ArrayList<>();
List<String> skipped = new java.util.ArrayList<>();
for (String rel : overlay) {
if (!existingPaths.contains(rel)) {
continue; // missing source is silently skipped
}
copied.add(rel);
if (trackedPaths.contains(rel)) {
skipped.add(rel);
}
}
overlayCalls.add(new OverlayCall(repoRoot, worktreePath, List.copyOf(overlay),
List.copyOf(copied), List.copyOf(skipped)));
}
@Override
public String repoRoot(String cwd) {
repoRootCalls.add(new RepoRootCall(cwd));
return repoRoot;
}
public List<AddCall> addCalls() {
return List.copyOf(addCalls);
}
public List<RemoveCall> removeCalls() {
return List.copyOf(removeCalls);
}
public List<OverlayCall> overlayCalls() {
return List.copyOf(overlayCalls);
}
public List<RepoRootCall> repoRootCalls() {
return List.copyOf(repoRootCalls);
}
public AddCall lastAdd() {
return addCalls.isEmpty() ? null : addCalls.getLast();
}
public RemoveCall lastRemove() {
return removeCalls.isEmpty() ? null : removeCalls.getLast();
}
public OverlayCall lastOverlay() {
return overlayCalls.isEmpty() ? null : overlayCalls.getLast();
}
}
@@ -0,0 +1,407 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-301 / CB-303 acceptance tests for the authoritative session registry, one-shot lifecycle FSM,
* and configurable lifecycle limits (idle TTL, context cap, drain).
* No live herdr — everything runs against the same {@link FakeHerdr} the rest of the project uses.
*/
class SessionManagerTest {
private SessionManager sessionManager(FakeHerdr herdr) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
return new SessionManager(workers);
}
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock) {
return sessionManager(herdr, clock, 0);
}
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
return new SessionManager(workers, new GitWorktrees(), clock, contextCap);
}
@Test
void acquireRegistersSpawningSessionWithDistinctPaneId() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession a = sessions.acquire("ltms-local", "/work/a", "/caller/a", "term_primary");
WorkerSession b = sessions.acquire("ltms-local", "/work/b", "/caller/b", "term_primary");
assertEquals(WorkerSession.State.SPAWNING, a.state(), "fresh session starts spawning");
assertEquals("ltms-local", a.profile());
assertEquals("/work/a", a.cwd(), "explicit requested cwd is recorded");
assertEquals("term_primary", a.ownerTerminal());
assertTrue(a.spawnedAtNanos() > 0);
assertNotNull(a.paneId());
assertNotNull(a.terminalId());
assertNotEquals(a.paneId(), b.paneId(), "no pane reuse");
assertNotEquals(a.terminalId(), b.terminalId(), "no terminal reuse");
assertEquals(2, sessions.roster().size(), "both sessions are registered");
}
@Test
void presenceMovesSpawningToReadyAndDeliveredTurnMovesToDone() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
assertEquals(WorkerSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"MCP presence moves SPAWNING → READY");
assertTrue(sessions.asPresence().isPresent(terminal), "presence is also recorded");
sessions.onDelivered(terminal);
assertEquals(WorkerSession.State.BUSY, sessions.get(session.paneId()).orElseThrow().state(),
"delivery moves READY → BUSY");
sessions.onTurnComplete(terminal);
assertEquals(WorkerSession.State.DONE, sessions.get(session.paneId()).orElseThrow().state(),
"turn completion moves BUSY → DONE");
}
@Test
void releaseTearsDownWorkerAndRemovesFromRosterAndIsIdempotent() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", null);
String paneId = session.paneId();
sessions.release(paneId);
assertTrue(herdr.called("pane.close"), "release tears the worker pane down");
assertTrue(sessions.get(paneId).isEmpty(), "released session is no longer retrievable");
assertTrue(sessions.roster().isEmpty(), "released session is no longer in the roster");
assertDoesNotThrow(() -> sessions.release(paneId), "a second release is harmless");
}
@Test
void onTurnFailedMovesSessionToFailed() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnFailed(terminal);
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(WorkerSession.State.FAILED, updated.state(), "turn failure moves to FAILED");
assertTrue(sessions.roster().contains(updated), "FAILED is still in acquired-minus-released roster");
}
@Test
void recycleProducesNewPaneIdAndOldOneIsGone() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession oldSession = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String oldPane = oldSession.paneId();
String oldTerminal = oldSession.terminalId();
WorkerSession fresh = sessions.recycle(oldPane);
assertNotEquals(oldPane, fresh.paneId(), "recycle yields a new pane id");
assertNotEquals(oldTerminal, fresh.terminalId(), "recycle yields a new terminal id");
assertEquals(oldSession.profile(), fresh.profile(), "profile is preserved");
assertEquals(oldSession.cwd(), fresh.cwd(), "cwd is preserved");
assertEquals(oldSession.ownerTerminal(), fresh.ownerTerminal(), "owner is preserved");
assertTrue(sessions.get(oldPane).isEmpty(), "old pane is deregistered");
assertEquals(1, sessions.roster().size(), "only the fresh session remains");
assertEquals(fresh.paneId(), sessions.roster().getFirst().paneId());
long paneCloseCount = herdr.calls.stream()
.filter(c -> "pane.close".equals(c.method()))
.filter(c -> oldPane.equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
assertEquals(1, paneCloseCount, "the old worker was torn down");
}
@Test
void rosterReflectsAcquiredMinusReleased() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession a = sessions.acquire("ltms-local", "/a", "/caller", "ownerA");
WorkerSession b = sessions.acquire("ltms-local", "/b", "/caller", "ownerB");
assertEquals(2, sessions.roster().size());
assertTrue(sessions.roster().stream().anyMatch(s -> s.paneId().equals(a.paneId())));
assertTrue(sessions.roster().stream().anyMatch(s -> s.paneId().equals(b.paneId())));
sessions.release(a.paneId());
assertEquals(1, sessions.roster().size());
assertEquals(b.paneId(), sessions.roster().getFirst().paneId());
}
// --- CB-303 lifecycle limits ----------------------------------------------------
@Test
void reapIdleDoesNothingWhenNoSessions() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L);
assertEquals(0, sessions.reapIdle(10));
assertTrue(sessions.roster().isEmpty());
}
@Test
void readySessionPastIdleTtlIsReaped() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
clock[0] = 11;
assertEquals(1, sessions.reapIdle(10), "READY session past TTL is reaped");
assertTrue(sessions.get(session.paneId()).isEmpty(), "reaped session is removed from registry");
assertTrue(herdr.called("pane.close"), "reaped session tears the pane down");
}
@Test
void readySessionWithinIdleTtlSurvives() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
clock[0] = 5;
assertEquals(0, sessions.reapIdle(10), "READY session within TTL is not reaped");
assertEquals(WorkerSession.State.READY,
sessions.get(session.paneId()).orElseThrow().state(),
"READY session survives");
}
@Test
void busySessionPastIdleTtlIsNotReaped() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
clock[0] = 100;
assertEquals(0, sessions.reapIdle(10), "BUSY session past TTL is never reaped");
assertEquals(WorkerSession.State.BUSY,
sessions.get(session.paneId()).orElseThrow().state(),
"BUSY session remains");
}
@Test
void doneSessionPastIdleTtlIsReaped() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
clock[0] = 21;
assertEquals(1, sessions.reapIdle(20), "DONE session past TTL is reaped");
assertTrue(sessions.get(session.paneId()).isEmpty(), "DONE session is removed");
}
@Test
void reapIdleReturnsCorrectCountAndSkipsBusy() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "owner1");
WorkerSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "owner2");
sessions.asPresence().markPresent(ready.terminalId());
sessions.asPresence().markPresent(busy.terminalId());
sessions.onDelivered(busy.terminalId());
clock[0] = 50;
assertEquals(1, sessions.reapIdle(30), "only READY past TTL is reaped");
assertTrue(sessions.get(ready.paneId()).isEmpty(), "READY session is gone");
assertEquals(WorkerSession.State.BUSY,
sessions.get(busy.paneId()).orElseThrow().state(),
"BUSY session is still registered");
}
@Test
void contextCapDisabledSessionSurvivesMultipleTurns() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 0);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(WorkerSession.State.DONE, updated.state(), "session finishes second turn");
assertEquals(2, updated.turnCount(), "turn count tracks both deliveries");
long releaseCloseCount = paneCloseCallsFor(herdr, session.paneId());
assertEquals(0, releaseCloseCount, "cap disabled — no forced release of the worker pane");
}
@Test
void contextCapTwoReleasesAfterSecondComplete() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 2);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
assertEquals(WorkerSession.State.DONE,
sessions.get(session.paneId()).orElseThrow().state(),
"first turn completes without release");
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
assertTrue(sessions.get(session.paneId()).isEmpty(), "session released after cap reached");
assertTrue(sessions.roster().isEmpty(), "released session leaves roster");
assertEquals(1, paneCloseCallsFor(herdr, session.paneId()),
"forced release tears the worker pane down exactly once");
}
@Test
void drainAllReleasesBusyAndReadySessionsAndWaitsForBusy() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "ownerR");
WorkerSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "ownerB");
sessions.asPresence().markPresent(ready.terminalId());
sessions.asPresence().markPresent(busy.terminalId());
sessions.onDelivered(busy.terminalId());
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
assertTrue(sessions.roster().isEmpty(), "drain clears the roster");
assertTrue(sessions.get(ready.paneId()).isEmpty(), "ready session is released");
assertTrue(sessions.get(busy.paneId()).isEmpty(), "busy session is released after timeout");
assertEquals(1, paneCloseCallsFor(herdr, ready.paneId()),
"ready worker pane is torn down");
assertEquals(1, paneCloseCallsFor(herdr, busy.paneId()),
"busy worker pane is torn down");
}
private static long paneCloseCallsFor(FakeHerdr herdr, String paneId) {
return herdr.calls.stream()
.filter(c -> "pane.close".equals(c.method()))
.filter(c -> paneId.equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
}
// --- CB-306 spawn-readiness gate: no half-registered session on timeout ----------------
@Test
void acquireThrowsPeerUnreachableWhenGateTimesOutAndRegistersNoSession() {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // never becomes injectable
long[] clock = {0};
// Gate-enabled launcher (1 ms timeout + no-op sleeper that advances clock past deadline)
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
1, () -> clock[0], () -> clock[0] += 10);
SessionManager sessions = new SessionManager(workers, new GitWorktrees(), () -> 0L, 0);
assertThrows(PeerUnreachableException.class,
() -> sessions.acquire("ltms-local", null, "/caller", "term_primary"),
"acquire must throw PeerUnreachableException when spawn times out");
// No half-registered session — the error happened inside spawn, before
// SessionManager could put() anything into the registry.
assertTrue(sessions.roster().isEmpty(),
"no session is registered when spawn times out (roster empty)");
}
// --- CB-516: release must notify, so a blocked send can be failed --------------------------
@Test
void releaseNotifiesTheListenerWithTheReleasedTerminal() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller", null);
sessions.release(s.paneId());
assertEquals(java.util.List.of(s.terminalId()), released,
"every teardown path funnels through release, so one hook must see the terminal");
}
@Test
void releasingAnUnknownPaneNotifiesNobody() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
sessions.release("w9:p404"); // idempotent teardown of something already gone
assertTrue(released.isEmpty(), "no session removed ⇒ no send was waiting on it");
}
@Test
void aThrowingReleaseListenerDoesNotBlockTheTeardown() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
sessions.onRelease(_ -> {
throw new IllegalStateException("listener blew up");
});
WorkerSession s = sessions.acquire("ltms-local", null, "/caller", null);
assertDoesNotThrow(() -> sessions.release(s.paneId()),
"a listener failure must never prevent the teardown it is reacting to");
assertTrue(sessions.get(s.paneId()).isEmpty(), "and the session is still deregistered");
}
}
@@ -0,0 +1,131 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Wrapper-behaviour tests for {@link SessionReaper} (the thread lifecycle). The TTL policy itself
* (SessionManager.reapIdle) is covered by SessionManagerTest and is deliberately not retested here.
* A real SessionManager is used, built the same way the rest of this package's tests do.
*/
class SessionReaperTest {
private static final long IDLE_TTL_SECONDS = 60;
private static final long SHORT_INTERVAL_MILLIS = 20;
private static ClaudeCodeLauncher launcher() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
}
/** A manager on the fake worktree seam — these tests never touch a real git checkout. */
private static SessionManager sessionManager() {
return new SessionManager(launcher(), new FakeWorktrees());
}
private static SessionReaper reaper() {
return new SessionReaper(sessionManager(), IDLE_TTL_SECONDS, SHORT_INTERVAL_MILLIS);
}
/**
* A double {@code start()} must leave exactly one live loop, so a single {@code stop()} still
* silences it. Asserting only "no throw" would pass against a reaper that never started at
* all — and against one that started twice — which is the entire point of the guard.
*/
@Test
void startIsIdempotent() throws InterruptedException {
AtomicLong ticks = new AtomicLong();
SessionReaper reaper = new SessionReaper(countingManager(ticks),
IDLE_TTL_SECONDS, SHORT_INTERVAL_MILLIS);
assertDoesNotThrow(() -> {
reaper.start();
reaper.start();
}, "a second start() must not throw");
assertTrue(awaitTicks(ticks, 2), "the loop is running after a double start()");
// One stop() for two start() calls: if the second start had spawned its own loop, a
// surviving thread would keep the counter climbing past this point.
reaper.stop();
Thread.sleep(SHORT_INTERVAL_MILLIS * 4);
long settled = ticks.get();
Thread.sleep(SHORT_INTERVAL_MILLIS * 4);
assertEquals(settled, ticks.get(),
"a single stop() must silence the reaper even after two start() calls");
}
/** A manager whose clock counts reads — every {@code reapIdle} reads it exactly once. */
private static SessionManager countingManager(AtomicLong ticks) {
return new SessionManager(launcher(), new FakeWorktrees(), () -> {
ticks.incrementAndGet();
return System.nanoTime();
});
}
/** Bounded wait for the loop to tick at least {@code n} times; avoids fixed-sleep flakiness. */
private static boolean awaitTicks(AtomicLong ticks, long n) throws InterruptedException {
long deadline = System.currentTimeMillis() + 2000;
while (ticks.get() < n && System.currentTimeMillis() < deadline) {
Thread.sleep(10);
}
return ticks.get() >= n;
}
@Test
void stopIsIdempotentAndSafeBeforeStart() {
SessionReaper reaper = reaper();
assertDoesNotThrow(reaper::stop, "stop() before start() must not throw");
assertDoesNotThrow(reaper::stop, "a second stop() must not throw");
}
/**
* The loop must actually iterate, and {@code stop()} must actually end it.
*
* <p>Observed through an injected clock rather than by sleeping and hoping: every
* {@code reapIdle} call reads {@code nowNanos} exactly once, so the tick count <em>is</em> the
* iteration count. Asserting merely "nothing threw" would pass even if {@code start()} were a
* no-op, which is the whole behaviour under test.
*/
@Test
void theLoopRunsRepeatedlyAndStopEndsIt() throws InterruptedException {
AtomicLong ticks = new AtomicLong();
SessionReaper reaper = new SessionReaper(countingManager(ticks),
IDLE_TTL_SECONDS, SHORT_INTERVAL_MILLIS);
reaper.start();
// Bounded wait rather than a fixed sleep + exact count: proves repetition without pinning
// a timing-derived number that would flake on a loaded machine.
boolean iterated = awaitTicks(ticks, 2);
long whileRunning = ticks.get();
reaper.stop();
assertTrue(iterated,
"the reaper loop must iterate repeatedly; observed " + whileRunning + " tick(s)");
// After stop() the loop must go quiet. Allow one in-flight iteration to finish, then
// confirm the count has stopped advancing.
Thread.sleep(SHORT_INTERVAL_MILLIS * 4);
long settled = ticks.get();
Thread.sleep(SHORT_INTERVAL_MILLIS * 4);
assertEquals(settled, ticks.get(), "stop() must end the loop, not just flag it");
}
}
@@ -0,0 +1,223 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-301-ext acceptance tests for worktree provisioning and config-parity overlay.
* No live git — every Worktrees call is handled by {@link FakeWorktrees} and every herdr
* call by {@link FakeHerdr}, matching the project's fake-based test style.
*/
class WorktreeSessionManagerTest {
private static ClaudeCodeLauncher workerService(FakeHerdr herdr) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
}
private static String startCwd(FakeHerdr herdr) {
@SuppressWarnings("unchecked")
Map<String, Object> start = (Map<String, Object>) herdr.lastCall("agent.start").params();
Object cwd = start.get("cwd");
return cwd == null ? null : cwd.toString();
}
@Test
void sharedTreeAcquireMakesNoWorktreesCallsAndRecordsNullWorktree() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary");
assertTrue(worktrees.addCalls().isEmpty(), "shared-tree acquire never adds a worktree");
assertTrue(worktrees.repoRootCalls().isEmpty(), "shared-tree acquire never resolves a repo root");
assertTrue(worktrees.overlayCalls().isEmpty(), "shared-tree acquire never overlays parity");
assertNull(s.worktree(), "shared-tree session has no worktree");
assertNull(s.branch(), "shared-tree session has no branch");
assertEquals("/caller/proj", s.cwd(), "shared-tree cwd is the caller's cwd");
assertEquals("/caller/proj", startCwd(herdr), "spawn receives the caller's cwd");
}
@Test
void worktreeAcquireProvisionsAndRecordsPathAndBranch() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-999", null));
assertEquals(1, worktrees.addCalls().size(), "one worktree was added");
FakeWorktrees.AddCall add = worktrees.lastAdd();
assertNotNull(add);
assertEquals("/repo", add.repoRoot());
assertTrue(add.branch().startsWith("worker/cb-999-"), "branch is worker/<slug>-<nonce>: " + add.branch());
assertNull(add.baseRef(), "null baseRef is passed through (HEAD default)");
String expectedPath = "/wt/" + add.branch().replace('/', '_');
assertEquals(expectedPath, s.worktree(), "session records the returned worktree path");
assertEquals(add.branch(), s.branch(), "session records the branch");
assertEquals(expectedPath, startCwd(herdr), "spawn receives the worktree path as cwd");
assertEquals(expectedPath, s.cwd(), "session cwd is the worktree path");
}
@Test
void worktreeAcquireRunsParityOverlayWithProfileDefaults() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.track(".mcp.json")
.exists(".claude/settings.local.json");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-888", null));
assertEquals(1, worktrees.overlayCalls().size());
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
assertNotNull(overlay);
assertEquals("/repo", overlay.repoRoot());
assertEquals(List.of(".mcp.json", ".claude/settings.local.json", ".env", ".envrc"),
overlay.requested(), "default parity overlay is used when unset");
assertEquals(List.of(".mcp.json", ".claude/settings.local.json"), overlay.copied(),
"existing paths are copied; missing paths are skipped");
assertEquals(List.of(".mcp.json"), overlay.skipWorktree(),
"tracked copied paths are --skip-worktree'd");
}
@Test
void releaseRemovesWorktreeButDoesNotDeleteBranch() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-666", null));
String paneId = s.paneId();
sessions.release(paneId);
assertTrue(herdr.called("pane.close"), "release still tears the worker pane down");
assertEquals(1, worktrees.removeCalls().size(), "worktree session triggers one remove");
FakeWorktrees.RemoveCall remove = worktrees.lastRemove();
assertNotNull(remove);
assertEquals("/repo", remove.repoRoot());
assertEquals(s.worktree(), remove.worktreePath());
// The fake records no branch-delete calls because Worktrees.remove only removes the checkout.
assertTrue(sessions.get(paneId).isEmpty(), "released session is no longer retrievable");
}
@Test
void sharedTreeReleaseMakesNoWorktreesCalls() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", null);
sessions.release(s.paneId());
assertTrue(herdr.called("pane.close"), "release tears the worker pane down");
assertTrue(worktrees.removeCalls().isEmpty(), "shared-tree release never removes a worktree");
}
@Test
void failedWorktreeAddUnwindsWithoutRegisteringSessionOrSpawning() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().failAdd("worktree add failed");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
assertThrows(WorktreeException.class, () ->
sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-555", null)));
assertEquals(0, sessions.size(), "failed acquire leaves no registry entry");
assertFalse(herdr.called("agent.start"), "spawn is never reached when add fails");
assertTrue(worktrees.removeCalls().isEmpty(), "no worktree was added, so none is removed");
}
@Test
void twoWorktreeAcquiresYieldDistinctBranchesAndPaths() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession a = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-444", null));
WorkerSession b = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-444", null));
assertNotEquals(a.branch(), b.branch(), "branches are distinct");
assertNotEquals(a.worktree(), b.worktree(), "paths are distinct");
assertEquals(2, worktrees.addCalls().size());
assertEquals(2, sessions.roster().size());
}
/**
* CB-507 regression. A plain REST spawn supplies neither a requested nor a caller cwd
* ({@code BridgedApp} hardcodes {@code callerCwd = null}), and the worktree branch used to
* resolve the repo root from just those two — yielding {@code null}, which the real
* {@code GitWorktrees} turns into {@code git -C null} and an NPE out of {@code ProcessBuilder}
* (HTTP 500).
*
* <p>Note this asserts on the <em>recorded</em> cwd rather than expecting a throw:
* {@link FakeWorktrees#repoRoot} only records its argument and returns a canned root, so a
* null flows through the fake harmlessly. That permissiveness is precisely why the whole
* suite stayed green while the feature was broken in production — so the assertion has to be
* "a usable cwd was passed down", not "an exception was raised".
*/
@Test
void worktreeAcquireWithNoRequestedOrCallerCwdStillResolvesANonNullRepoRoot() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
sessions.acquire("ltms-local", null, null, null, new WorktreeRequest("cb-507", null));
assertFalse(worktrees.repoRootCalls().isEmpty(),
"repoRoot should have been called to resolve the repo root");
String cwd = worktrees.repoRootCalls().getFirst().cwd();
assertNotNull(cwd, "a null cwd here becomes `git -C null` and NPEs in the real GitWorktrees");
assertFalse(cwd.isBlank(), "a blank cwd is as unusable as a null one");
}
/**
* The same line carried a second, quieter bug: it never consulted the profile's configured
* {@code cwd:}, so a worktree spawn silently ignored a pinned per-profile working directory.
* Routing through {@code effectiveCwd} honours it.
*/
@Test
void worktreeAcquireHonoursTheProfileConfiguredCwd() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
// Argument order matters: configDir is the 4th parameter, cwd the 11th (after mcpUrl).
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, "/pinned/dir", null);
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
SessionManager sessions = new SessionManager(launcher, worktrees);
sessions.acquire("ltms-local", null, null, null, new WorktreeRequest("cb-507b", null));
assertEquals(1, worktrees.repoRootCalls().size());
assertEquals("/pinned/dir", worktrees.repoRootCalls().getFirst().cwd(),
"the profile's configured cwd must reach repoRoot, not be ignored");
}
}
@@ -0,0 +1,472 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.*;
/** The step-4 launch-flag injection: the bridge MCP + reply charter are appended to the argv. */
class ClaudeCodeLauncherTest {
private ClaudeCodeLauncher service(FakeHerdr herdr, List<String> argv, String mcpUrl) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
argv, "tab", "bridged-workers", "worker: {profile} #{n}", mcpUrl, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
}
@SuppressWarnings("unchecked")
private List<String> spawnedArgv(FakeHerdr herdr) {
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("argv");
}
@Test
void appendsBridgeMcpAndReplyCharterWhenMcpUrlSet() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("ccs", "ltms-local"), "http://127.0.0.1:8765/mcp").spawn();
List<String> argv = spawnedArgv(herdr);
assertEquals(List.of("ccs", "ltms-local"), argv.subList(0, 2), "base command preserved first");
assertTrue(argv.contains("--mcp-config"));
assertTrue(argv.stream().anyMatch(a -> a.contains("\"bridge\"") && a.contains("http://127.0.0.1:8765/mcp")),
"inline bridge MCP config present");
assertTrue(argv.contains("--append-system-prompt"));
assertTrue(argv.stream().anyMatch(a -> a.contains("bridge_reply")), "reply charter present");
}
@Test
void noBridgeFlagsWhenMcpUrlAbsent() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("bash", "-c", "sleep 1"), null).spawn();
assertEquals(List.of("bash", "-c", "sleep 1"), spawnedArgv(herdr), "argv untouched without mcpUrl");
}
private ClaudeCodeLauncher multiProfile(FakeHerdr herdr) {
BridgedConfig.Worker gx10 = new BridgedConfig.Worker("gx10", "http://gx10.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "gx10"), "tab", "bridged-workers", "w #{n}", null, null, null);
BridgedConfig.Worker ollama = new BridgedConfig.Worker("ollama", "http://ollama.ltms.dev", null,
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "ollama"), "tab", "bridged-workers", "w #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx10.gw", "ollama.ltms.dev")),
Map.of("gx10", gx10, "ollama", ollama), "gx10", _ -> "tok");
}
@Test
@SuppressWarnings("unchecked")
void spawnPicksTheNamedProfilesBaseUrlAndArgv() {
FakeHerdr herdr = new FakeHerdr();
multiProfile(herdr).spawn("ollama");
Map<String, Object> start = (Map<String, Object>) herdr.lastCall("agent.start").params();
Map<String, String> env = (Map<String, String>) start.get("env");
assertEquals("http://ollama.ltms.dev", env.get("ANTHROPIC_BASE_URL"), "the named profile's base_url");
assertEquals(List.of("ccs", "ollama"), start.get("argv"), "the named profile's launch command");
}
@Test
void spawnRejectsAnUnknownProfile() {
try (FakeHerdr herdr = new FakeHerdr()) {
assertThrows(IllegalArgumentException.class, () -> multiProfile(herdr).spawn("nope"));
}
}
@SuppressWarnings("unchecked")
private static String startCwd(FakeHerdr herdr) {
// The worker's cwd is set on agent.start (an agent pane does not inherit the tab's cwd).
Object v = ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("cwd");
return v == null ? null : v.toString();
}
@Test
void requestedCwdRootsTheWorker() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("ccs", "ltms-local"), null).spawn("ltms-local", "/work/proj", "/caller/home");
assertEquals("/work/proj", startCwd(herdr), "an explicit spawn cwd wins over everything");
}
@Test
void profileConfigCwdBeatsTheCallerCwd() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker("ltms-local", "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"w #{n}", null, "/pinned/dir", null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("ltms-local", cfg), "ltms-local", _ -> null);
svc.spawn("ltms-local", null, "/caller/home");
assertEquals("/pinned/dir", startCwd(herdr), "a profile-pinned cwd overrides the caller's");
}
@Test
void inheritsTheCallerCwdWhenNothingElseIsSet() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("ccs", "ltms-local"), null).spawn("ltms-local", null, "/primary/project");
assertEquals("/primary/project", startCwd(herdr), "no explicit/config cwd → inherit the primary's");
}
// --- CB-302 git-forge token injection (worker checkpoint grant) ------------
@SuppressWarnings("unchecked")
private static Map<String, String> startEnv(FakeHerdr herdr) {
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("env");
}
@Test
void injectsForgeTokenAndHostWhenProfileGrantsIt() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"impl", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "impl"), "tab", "bridged-workers", "w #{n}", null, null, null,
"GITEA_ACCESS_TOKEN", null); // parityOverlay null; gitHostEnv null → defaults to GITEA_HOST
Function<String, String> host = name -> switch (name) {
case "GITEA_ACCESS_TOKEN" -> "gt-secret";
case "GITEA_HOST" -> "git.ltms.dev";
default -> null;
};
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("impl", cfg), "impl", host).spawn();
Map<String, String> env = startEnv(herdr);
assertEquals("gt-secret", env.get("GITEA_TOKEN"), "the forge token is injected for a granting profile");
assertEquals("git.ltms.dev", env.get("GITEA_HOST"), "the paired forge host rides along with the token");
}
@Test
void noForgeTokenWhenProfileDoesNotGrantIt() {
FakeHerdr herdr = new FakeHerdr();
// gitTokenEnv unset (12-arg ctor); the env would resolve a token if asked, proving the gate
// is the profile config, not a missing env var.
BridgedConfig.Worker cfg = new BridgedConfig.Worker("ltms-local", "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("ccs"), "tab", "bridged-workers", "w #{n}",
null, null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("ltms-local", cfg), "ltms-local",
_ -> "would-be-secret").spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("GITEA_TOKEN"), "no forge token when the profile does not opt in");
assertNull(env.get("GITEA_HOST"), "no forge host without a granted token");
}
// --- CB-117 orphan reap: the pure predicate --------------------------------
@Test
void isForeignWorkerMatchesOurSchemeWithANonSelfNonce() {
assertTrue(ClaudeCodeLauncher.isForeignWorker("claude-ollama-be09c2-2", "aaaaaa"),
"a bridge worker name with a different nonce is a prior daemon's orphan");
assertTrue(ClaudeCodeLauncher.isForeignWorker("claude-gx10-4127af-11", "aaaaaa"),
"profile and multi-digit seq are still parsed; foreign nonce ⇒ reap");
}
@Test
void isForeignWorkerSparesOurOwnLiveWorkersAndNonWorkers() {
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude-ollama-abcdef-3", "abcdef"),
"a worker with THIS process's nonce is ours and live — never reap it");
assertFalse(ClaudeCodeLauncher.isForeignWorker(null, "abcdef"), "an unnamed agent is not a worker");
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude", "abcdef"), "a bare kind name is not a worker");
assertFalse(ClaudeCodeLauncher.isForeignWorker("my-repl", "abcdef"), "a user's own label is not a worker");
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude-ollama-XYZ123-2", "abcdef"),
"a non-hex nonce does not match our scheme");
}
// --- CB-117 orphan reap: the wiring through stop() -------------------------
private static long paneCloseCount(FakeHerdr herdr, String paneId) {
return herdr.calls.stream()
.filter(c -> c.method().equals("pane.close"))
.filter(c -> paneId.equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
}
@Test
void reapsAForeignOrphanButSparesOurOwnWorkerAndUserSessions() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = multiProfile(herdr);
herdr.withAgent("claude-ollama-be09c2-2", "term_orphan", "wQ:pF", "wQ:t8") // prior daemon's leak
.withAgent("claude-gx10-" + svc.nameNonce() + "-1", "term_mine", "wQ:pMine", "wQ:tMine"); // ours, live
// (the fake's default unnamed term_a stands in for a user's own Claude session)
int reaped = svc.reapOrphanWorkers();
assertEquals(1, reaped, "exactly the one foreign-nonce orphan is reaped");
assertEquals(1, paneCloseCount(herdr, "wQ:pF"), "the orphan's pane is closed");
assertEquals(0, paneCloseCount(herdr, "wQ:pMine"), "our own live worker's pane is left running");
assertEquals(0, paneCloseCount(herdr, "w2:p7"), "a user's own session is never touched");
assertTrue(herdr.called("tab.close"), "the orphan's now-empty dedicated tab is closed too");
}
@Test
void reapCountsAnAlreadyGoneOrphanAsReaped() {
FakeHerdr herdr = new FakeHerdr().paneCloseFailsWith("pane_not_found");
ClaudeCodeLauncher svc = multiProfile(herdr);
herdr.withAgent("claude-ollama-0d856d-3", "term_gone", "wQ:pS", "wQ:tD");
assertEquals(1, svc.reapOrphanWorkers(),
"a pane that vanished between list and close is a successful reap, not a failure");
}
@Test
void reapIsSkippedWhenHerdrCannotBeListed() {
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.list throws
assertEquals(0, multiProfile(herdr).reapOrphanWorkers(), "a listing failure reaps nothing and does not throw");
}
// --- PeerHandle indirection ----------------------------------------------------------------
@Test
void spawnReturnsPeerHandleWithIdEqualToPaneId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn must return a non-null handle");
assertEquals("w9:pW_1", handle.id(), "handle.id() must equal the agent's paneId");
}
@Test
void spawnReturnsPeerHandleWithCorrectTerminalId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, "/caller"));
assertEquals("term_new_1", handle.terminalId(), "handle.terminalId() must equal the agent's terminalId");
}
@Test
void capabilitiesIncludeMidTurnAskWorktreeOrphanReap() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
Set<Capability> caps = svc.capabilities();
assertTrue(caps.contains(Capability.MID_TURN_ASK), "every Claude Code peer supports mid-turn ask");
assertTrue(caps.contains(Capability.WORKTREE), "every CLI peer supports worktree cwd");
assertTrue(caps.contains(Capability.ORPHAN_REAP), "every herdr launcher supports orphan reap");
}
@Test
void capabilitiesIncludeSelfPrWhenProfileHasGitToken() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"impl", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "impl"), "tab", "bridged-workers", "w #{n}", null, null, null,
"GITEA_ACCESS_TOKEN", null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("impl", cfg), "impl",
_ -> "tok");
assertTrue(svc.capabilities().contains(Capability.SELF_PR),
"a profile with a git token grants SELF_PR");
}
@Test
void capabilitiesExcludeSelfPrWhenNoGitToken() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
assertFalse(svc.capabilities().contains(Capability.SELF_PR),
"no git token profile → no SELF_PR capability");
}
@Test
void effectiveCwdViaSpawnRequestMatchesExistingResolution() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
String cwd = svc.effectiveCwd(new SpawnRequest("ltms-local", "/work/proj", "/caller/home"));
assertEquals("/work/proj", cwd, "effectiveCwd via SpawnRequest must match the three-arg resolution");
}
@Test
void profilesViaPeerLauncherMatchesExistingApi() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = multiProfile(herdr);
assertEquals(Set.of("gx10", "ollama"), svc.profiles(), "profiles() via PeerLauncher must match");
}
@Test
void defaultProfileViaPeerLauncherMatches() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = multiProfile(herdr);
assertEquals("gx10", svc.defaultProfile(), "defaultProfile() via PeerLauncher must match");
}
@Test
void stopViaPeerLauncherTearsDownByHandleId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
svc.stop(handle.id());
assertTrue(herdr.called("pane.close"), "stop via handle.id() must close the pane");
}
// --- CB-306 spawn-readiness gate -----------------------------------------------------------
private static Map<String, BridgedConfig.Worker> workerConfigMap(String profile, String mcpUrl) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
profile, "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", profile), "tab", "bridged-workers",
"worker: {profile} #{n}", mcpUrl, null, null);
return Map.of(cfg.profile(), cfg);
}
@Test
void spawnWaitsUntilInjectableThenReturnsHandle() {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // first status call sees UNKNOWN
long[] clock = {0};
boolean[] firstSleep = {true};
// The sleeper: advance the fake clock, and on the first call flip the
// agent status to IDLE so the next poll succeeds.
Runnable sleeper = () -> {
clock[0] += 300;
if (firstSleep[0]) {
herdr.agentStatus("idle");
firstSleep[0] = false;
}
};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
5000, () -> clock[0], sleeper);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn returns a handle when worker becomes injectable");
assertEquals("w9:pW_1", handle.id(), "handle id matches the started pane");
assertEquals(0, paneCloseCount(herdr, "w9:pW_1"),
"no pane.close when worker becomes injectable before timeout");
}
@Test
void spawnThrowsPeerUnreachableWhenNeverInjectableAndReapsPane() {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
1000, () -> clock[0], () -> clock[0] += 50);
PeerUnreachableException ex = assertThrows(
PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(ex.getMessage().contains("w9:pW_1"),
"exception message references the paneId: " + ex.getMessage());
assertTrue(ex.getMessage().contains("1000"),
"exception message references the timeout: " + ex.getMessage());
assertTrue(clock[0] >= 1000, "fake clock advanced past the timeout: " + clock[0]);
assertEquals(1, paneCloseCount(herdr, "w9:pW_1"),
"pane was closed on timeout (no orphan left behind)");
}
@Test
void spawnReturnsImmediatelyWhenGateIsDisabled() {
FakeHerdr herdr = new FakeHerdr();
// The default 6-arg constructor has spawnReadyTimeoutMs=0 (gate disabled).
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn returns a handle when the gate is disabled");
assertFalse(herdr.called("agent.get"),
"agent.get is never called when the gate is disabled (no polling)");
}
@Test
void spawnGateRespectsZeroTimeoutEvenWithFullConstructor() {
FakeHerdr herdr = new FakeHerdr();
long[] clock = {0};
// Explicit zero timeout with the full testability constructor — should
// skip polling entirely, just like the legacy default path.
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
0, () -> clock[0], () -> clock[0] += 1);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn still succeeds with zero timeout");
assertEquals(0, paneCloseCount(herdr, handle.id()),
"no orphan pane close from the gate path");
}
// --- CB-511: worker environment seeding -----------------------------------------------------
@Test
void workerInheritsTheDaemonPath() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
k -> "PATH".equals(k) ? "/opt/tools/bin:/usr/bin" : null).spawn();
assertEquals("/opt/tools/bin:/usr/bin", startEnv(herdr).get("PATH"),
"a worker with no PATH cannot run the build it is asked to run");
}
@Test
void profileEnvIsInjectedIntoTheWorker() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("JAVA_HOME", "/opt/jdk", "PATH", "/profile/bin"));
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
k -> "PATH".equals(k) ? "/daemon/bin" : null).spawn();
Map<String, String> env = startEnv(herdr);
assertEquals("/opt/jdk", env.get("JAVA_HOME"), "profile env: is passed through");
assertEquals("/profile/bin", env.get("PATH"), "an explicit profile PATH overrides the daemon's");
}
/**
* The security-relevant ordering. {@code SubscriptionGuard} is checked against the profile's
* {@code baseUrl} only, so if a profile's {@code env:} could overwrite ANTHROPIC_BASE_URL a
* worker could be pointed at an unguarded host while the guard passed on a benign one.
*/
@Test
void profileEnvCannotOverrideGuardCheckedAnthropicVars() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("ANTHROPIC_BASE_URL", "http://evil.example.com"));
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null).spawn();
assertEquals("http://gx00.gw:8000", startEnv(herdr).get("ANTHROPIC_BASE_URL"),
"the guard-checked baseUrl must win over any env: entry, or the boundary is bypassable");
}
}
@@ -0,0 +1,158 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
/**
* The composite router: profile → owning adapter for spawn/cwd/parity, pane id → owner for stop,
* and fleet-wide union/dedup for list/reap/caps/profiles. Exercised through two real adapters —
* claude-code + opencode — over one FakeHerdr, so each call is observed reaching the right adapter
* (the started herdr agent name carries that adapter's {@code claude-}/{@code opencode-} prefix).
*/
class CompositePeerLauncherTest {
private ClaudeCodeLauncher claudeAdapter(FakeHerdr herdr) {
// 12-arg back-compat Worker ctor → kind defaults to claude-code.
BridgedConfig.Worker claude = new BridgedConfig.Worker("claude", "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}",
null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("claude", claude), "claude", _ -> null);
}
private OpenCodeLauncher opencodeAdapter(FakeHerdr herdr) {
BridgedConfig.Worker gemini = new BridgedConfig.Worker("gemini", null, "google/gemini-2.5-pro",
null, "BRIDGED_WORKER_TOKEN", List.of("opencode"), "tab", "bridged-workers", "w #{n}",
null, null, null, "GITEA_ACCESS_TOKEN", null, BridgedConfig.Worker.KIND_OPENCODE);
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("gemini", gemini), "gemini", _ -> "tok");
}
private CompositePeerLauncher composite(FakeHerdr herdr) {
return new CompositePeerLauncher(
List.of(claudeAdapter(herdr), opencodeAdapter(herdr)), "claude");
}
@SuppressWarnings("unchecked")
private static String startedName(FakeHerdr herdr) {
return (String) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("name");
}
@Test
void spawnRoutesEachProfileToItsOwningAdapter() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
composite.spawn(new SpawnRequest("gemini", null, null));
assertTrue(startedName(herdr).startsWith("opencode-"),
"the gemini profile is spawned by the opencode adapter: " + startedName(herdr));
composite.spawn(new SpawnRequest("claude", null, null));
assertTrue(startedName(herdr).startsWith("claude-"),
"the claude profile is spawned by the claude-code adapter: " + startedName(herdr));
}
@Test
void nullProfileResolvesTheDefaultAndRoutesToItsOwner() {
FakeHerdr herdr = new FakeHerdr();
composite(herdr).spawn(new SpawnRequest(null, null, null));
assertTrue(startedName(herdr).startsWith("claude-"),
"a no-profile spawn resolves the default (claude) and routes to its adapter");
}
@Test
void unknownProfileIsRejected() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
assertThrows(IllegalArgumentException.class,
() -> composite.spawn(new SpawnRequest("nope", null, null)),
"a profile no adapter declares is an error");
}
@Test
void profilesAndDefaultAreExposedAcrossAdapters() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
assertEquals(Set.of("claude", "gemini"), composite.profiles(),
"profiles are the union of every adapter's profiles");
assertEquals("claude", composite.defaultProfile());
}
@Test
void capabilitiesAreTheUnionOfEveryAdapter() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher claude = claudeAdapter(herdr);
OpenCodeLauncher opencode = opencodeAdapter(herdr);
PeerLauncher composite = new CompositePeerLauncher(List.of(claude, opencode), "claude");
assertTrue(composite.capabilities().containsAll(claude.capabilities()),
"the fleet offers every claude-code capability");
assertTrue(composite.capabilities().containsAll(opencode.capabilities()),
"the fleet offers every opencode capability (incl. SELF_PR from its git-token profile)");
}
@Test
void listIsDeduplicatedByPaneIdAcrossAdaptersSharingHerdr() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
// Both adapters wrap the same herdr, so each list() returns the same global agent set;
// the composite must return each pane once, not once per adapter.
assertEquals(1, composite.list().size(),
"the single herdr-tracked pane appears once, not duplicated per adapter");
}
@Test
void reapSumsAcrossAdaptersAndEachAdapterReapsOnlyItsOwnPrefix() {
// One foreign opencode orphan + one foreign claude orphan, from a prior daemon (different nonce).
FakeHerdr herdr = new FakeHerdr()
.withAgent("opencode-gemini-ffffff-1", "term_o", "wQ:pO", "wQ:tO")
.withAgent("claude-claude-eeeeee-1", "term_c", "wQ:pC", "wQ:tC");
PeerLauncher composite = composite(herdr);
assertEquals(2, composite.reapOrphanWorkers(),
"both orphans are reaped — one by each adapter, summed by the composite");
}
@Test
void stopTearsDownAPaneSpawnedThroughTheComposite() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
PeerHandle handle = composite.spawn(new SpawnRequest("gemini", null, null));
composite.stop(handle.id());
assertTrue(herdr.calls.stream()
.anyMatch(c -> c.method().equals("pane.close")
&& handle.id().equals(((Map<?, ?>) c.params()).get("pane_id"))),
"stop routes to the spawning adapter and closes that worker's pane");
}
@Test
void constructorRejectsAProfileClaimedByTwoAdapters() {
FakeHerdr herdr = new FakeHerdr();
// Two opencode adapters both declaring "gemini" — a profile-name collision.
OpenCodeLauncher a = opencodeAdapter(herdr);
OpenCodeLauncher b = opencodeAdapter(herdr);
assertThrows(IllegalArgumentException.class,
() -> new CompositePeerLauncher(List.of(a, b), "gemini"),
"a profile two adapters both claim is a configuration error");
}
@Test
void constructorRejectsAnEmptyAdapterList() {
assertThrows(IllegalArgumentException.class,
() -> new CompositePeerLauncher(List.of(), "claude"),
"at least one adapter must be configured");
}
}
@@ -0,0 +1,269 @@
package dev.ltms.bridged.worker;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/**
* The opencode adapter's launch build: a file-based MCP mount + reply-charter instructions (no
* inline flags, no {@code ANTHROPIC_*}, no guard), the {@code -m} model flag, and the shared base
* transport (naming, reap, readiness gate) proving the {@link HerdrPeerLauncher} SPI is neutral.
*/
class OpenCodeLauncherTest {
private static BridgedConfig.Worker opencodeCfg(String model, String mcpUrl, String gitTokenEnv) {
return new BridgedConfig.Worker("gemini", null, model, null, "BRIDGED_WORKER_TOKEN",
List.of("opencode"), "tab", "bridged-workers", "opencode: {model} #{n}", mcpUrl,
null, null, gitTokenEnv, null, BridgedConfig.Worker.KIND_OPENCODE);
}
/** Gate-disabled launcher whose per-spawn config dirs land under an inspectable temp root. */
private OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, BridgedConfig.Worker cfg) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), k -> "GITEA_ACCESS_TOKEN".equals(k) ? "tok" : null,
0, System::currentTimeMillis, () -> { }, configRoot);
}
@SuppressWarnings("unchecked")
private static Map<String, Object> lastStart(FakeHerdr herdr) {
return (Map<String, Object>) herdr.lastCall("agent.start").params();
}
@SuppressWarnings("unchecked")
private static Map<String, String> startEnv(FakeHerdr herdr) {
return (Map<String, String>) lastStart(herdr).get("env");
}
@SuppressWarnings("unchecked")
private static List<String> startArgv(FakeHerdr herdr) {
return (List<String>) lastStart(herdr).get("argv");
}
@Test
void writesRemoteMcpConfigAndCharterInstructionsWhenMcpUrlSet(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null))
.spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"), "opencode carries no ANTHROPIC_* / subscription boundary");
String cfgPath = env.get("OPENCODE_CONFIG");
assertNotNull(cfgPath, "OPENCODE_CONFIG points the worker at the generated config file");
assertTrue(Path.of(cfgPath).startsWith(root), "config file is generated under the injected root");
// Assert on parsed structure, not substrings: the generated config is real JSON and its
// whitespace is the formatter's business, not the contract's.
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
JsonNode bridge = json.path("mcp").path("bridge");
assertEquals("remote", bridge.path("type").asText(), "bridge is mounted as a remote MCP server");
assertEquals("http://127.0.0.1:8765/mcp", bridge.path("url").asText(),
"the profile's bridge MCP url is present");
assertTrue(bridge.path("enabled").asBoolean(), "the bridge server is enabled");
assertTrue(json.path("instructions").isArray() && !json.path("instructions").isEmpty(),
"the reply charter is mounted via instructions");
// The instructions entry is a real file path holding the reply charter.
Path charter = Path.of(cfgPath).resolveSibling("reply-charter.md");
assertTrue(Files.exists(charter), "the charter file the config references was written");
assertTrue(Files.readString(charter).contains("bridge_reply"),
"the charter instructs the worker to answer via bridge_reply");
}
@Test
void noConfigFileWhenMcpUrlAbsent(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null)).spawn();
assertNull(startEnv(herdr).get("OPENCODE_CONFIG"),
"no bridge MCP url → no config file and no OPENCODE_CONFIG");
}
@Test
void passesTheModelAsDashMFlag(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null)).spawn();
List<String> argv = startArgv(herdr);
assertEquals("opencode", argv.getFirst(), "base opencode command preserved first");
int m = argv.indexOf("-m");
assertTrue(m >= 0, "model is selected with -m");
assertEquals("google/gemini-2.5-pro", argv.get(m + 1), "the provider/model selector follows -m");
}
@Test
void noModelFlagWhenModelBlank(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg(null, null, null)).spawn();
assertEquals(List.of("opencode"), startArgv(herdr), "no model → argv is the bare opencode command");
}
@Test
void injectsForgeTokenWhenProfileGrantsIt(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg(null, null, "GITEA_ACCESS_TOKEN")).spawn();
assertEquals("tok", startEnv(herdr).get("GITEA_TOKEN"),
"a git-token profile gets the peer-neutral GITEA_TOKEN grant, same as Claude");
}
@Test
void capabilitiesDeclareOrphanReapAndMcpAskAndConditionalSelfPr(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
assertEquals(java.util.Set.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP),
service(herdr, root, opencodeCfg(null, null, null)).capabilities(),
"no git token → no SELF_PR");
assertTrue(service(herdr, root, opencodeCfg(null, null, "GITEA_ACCESS_TOKEN"))
.capabilities().contains(Capability.SELF_PR),
"a git-token profile adds SELF_PR");
}
@Test
void foreignWorkerMatchesOpencodePrefixButNotClaude() {
String nonce = "abc123";
assertTrue(OpenCodeLauncher.isForeignWorker("opencode-gemini-def456-1", nonce),
"an opencode pane from another process is foreign");
assertFalse(OpenCodeLauncher.isForeignWorker("opencode-gemini-" + nonce + "-1", nonce),
"our own opencode pane (same nonce) is not foreign");
assertFalse(OpenCodeLauncher.isForeignWorker("claude-ltms-local-def456-1", nonce),
"a claude pane is never reaped by the opencode adapter");
}
@Test
void productionConstructorsWireThroughToTheBase() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = opencodeCfg(null, null, null);
// 5-arg (gate disabled) and 7-arg (gate enabled) production constructors both expose the profile.
OpenCodeLauncher disabled = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
OpenCodeLauncher gated = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null, 5000, 100);
assertEquals(java.util.Set.of("gemini"), disabled.profiles());
assertEquals("gemini", gated.defaultProfile());
}
@Test
void spawnGateThrowsPeerUnreachableWhenNeverInjectable(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // never injectable
long[] clock = {0};
OpenCodeLauncher svc = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("gemini", opencodeCfg(null, null, null)), "gemini", _ -> null,
1000, () -> clock[0], () -> clock[0] += 50, root);
PeerUnreachableException ex = assertThrows(PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(clock[0] >= 1000, "the fake clock advanced past the timeout: " + clock[0]);
long closes = herdr.calls.stream()
.filter(c -> c.method().equals("pane.close"))
.filter(c -> "w9:pW_1".equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
assertEquals(1, closes, "the worker pane was reaped on timeout (no orphan)");
assertNotNull(ex.getMessage());
}
@Test
void spawnReturnsHandleWhenGateDisabled(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
PeerHandle handle = service(herdr, root, opencodeCfg(null, null, null))
.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn returns a handle when the gate is disabled");
assertFalse(herdr.called("agent.get"), "no polling when the gate is disabled");
}
// --- CB-508: pinned OpenAI-compatible endpoint (e.g. a local vLLM) ---------------------------
/** A profile with a baseUrl but no model provider prefix cannot be resolved — fail loudly. */
private static BridgedConfig.Worker pinnedCfg(String model, String baseUrl, String mcpUrl) {
return new BridgedConfig.Worker("local", baseUrl, model, null, "BRIDGED_WORKER_TOKEN",
List.of("opencode"), "tab", "bridged-workers", "opencode: {model} #{n}", mcpUrl,
null, null, null, null, BridgedConfig.Worker.KIND_OPENCODE);
}
@Test
void baseUrlDeclaresACustomOpenAiCompatibleProvider(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, pinnedCfg("local-vllm/deepseek-v4-flash", "http://127.0.0.1:8000", null))
.spawn();
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(cfgPath, "a pinned endpoint needs a config file even with no bridge MCP url");
JsonNode provider = new ObjectMapper().readTree(Path.of(cfgPath).toFile())
.path("provider").path("local-vllm");
assertFalse(provider.isMissingNode(), "the provider id comes from the model selector");
assertEquals("@ai-sdk/openai-compatible", provider.path("npm").asText());
assertEquals("http://127.0.0.1:8000/v1", provider.path("options").path("baseURL").asText(),
"a bare host:port gets /v1 appended — that is where these servers mount the API");
assertFalse(provider.path("options").path("apiKey").asText().isBlank(),
"the AI SDK requires a non-empty key even when the server ignores it");
assertFalse(provider.path("models").path("deepseek-v4-flash").isMissingNode(),
"the model half of the selector is declared under the provider");
}
@Test
void aBaseUrlThatAlreadyCarriesAPathIsUsedVerbatim(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, pinnedCfg("local-vllm/m", "http://127.0.0.1:8000/openai/v1", null)).spawn();
JsonNode json = new ObjectMapper()
.readTree(Path.of(startEnv(herdr).get("OPENCODE_CONFIG")).toFile());
assertEquals("http://127.0.0.1:8000/openai/v1",
json.path("provider").path("local-vllm").path("options").path("baseURL").asText(),
"an endpoint mounted on a custom path must not have /v1 bolted on");
}
@Test
void aPinnedEndpointRejectsAModelWithNoProviderPrefix(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher =
service(herdr, root, pinnedCfg("deepseek-v4-flash", "http://127.0.0.1:8000", null));
// Silently falling back to the default gateway would point the worker at the wrong LLM
// while looking healthy — the one failure mode worth being loud about.
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, launcher::spawn);
assertTrue(e.getMessage().contains("<provider>/<model>"), "the error says how to fix it");
}
@Test
void aPinnedEndpointAndTheBridgeMcpCoexistInOneConfig(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, pinnedCfg("local-vllm/deepseek-v4-flash",
"http://127.0.0.1:8000", "http://127.0.0.1:8766/mcp")).spawn();
JsonNode json = new ObjectMapper()
.readTree(Path.of(startEnv(herdr).get("OPENCODE_CONFIG")).toFile());
assertEquals("remote", json.path("mcp").path("bridge").path("type").asText(),
"pinning an endpoint must not drop the bridge MCP mount");
assertFalse(json.path("provider").path("local-vllm").isMissingNode(),
"and the provider block is still declared alongside it");
assertTrue(json.path("instructions").isArray() && !json.path("instructions").isEmpty(),
"the reply charter survives too");
}
@Test
void noBaseUrlDeclaresNoProviderSoTheDefaultGatewayIsUsed(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("opencode/some-free-model", "http://127.0.0.1:8766/mcp", null))
.spawn();
JsonNode json = new ObjectMapper()
.readTree(Path.of(startEnv(herdr).get("OPENCODE_CONFIG")).toFile());
assertTrue(json.path("provider").isMissingNode(),
"without a baseUrl opencode resolves its own provider as before");
}
}
@@ -0,0 +1,37 @@
<configuration>
<!--
CB-506 — test-run logging. This file is NOT boilerplate; it exists to keep the test suite
out of the CB-505 audit trail.
main/resources/logback.xml routes the `audit` logger to a RollingFileAppender at
logs/audit.log. AuditLogTest and BridgedAppAuthTest exercise that same logger, so without
this file `mvn test` appends fabricated records — denied/forbidden SPAWN/STOP/SEND from
worker:term_a — to the production security log, byte-identical to real ones. An investigator
could not tell a test fixture from a genuine intrusion attempt. Logback prefers
logback-test.xml when it is on the test classpath, so this governs test runs only.
Two constraints if you edit this:
- NEVER add a FileAppender/RollingFileAppender here. That reintroduces the bug.
- Keep `audit` ENABLED (INFO, additivity=false). Setting it to OFF would silently break
AuditLogTest, which attaches its own ListAppender and asserts on emitted records.
-->
<appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
<encoder>
<pattern>%d{HH:mm:ss.SSS} [%thread] %-5level %logger{36} - %msg%n</pattern>
</encoder>
</appender>
<logger name="audit" level="INFO" additivity="false">
<appender-ref ref="STDOUT"/>
</logger>
<logger name="dev.ltms.bridged" level="WARN"/>
<logger name="org.eclipse.jetty" level="WARN"/>
<root level="INFO">
<appender-ref ref="STDOUT"/>
</root>
</configuration>
+61
View File
@@ -0,0 +1,61 @@
# CB-504 — systemd unit for bridged (Linux).
#
# The macOS launchd agent (deploy/dev.ltms.bridged.plist) is the supervision target for the
# current single-host deployment. This unit exists for the per-host gateways CB-308 introduces,
# which will run on Linux.
#
# Install (user service — bridged drives the user's herdr, not a system daemon):
# mkdir -p ~/.config/systemd/user
# cp deploy/bridged.service ~/.config/systemd/user/
# # edit ExecStart / WorkingDirectory / Environment below, then:
# systemctl --user daemon-reload
# systemctl --user enable --now bridged
# journalctl --user -u bridged -f
[Unit]
Description=bridged — claude-bridge message server
Documentation=https://git.ltms.dev/lms/claude-bridge/wiki
# Ordering only: herdr is a user process and its socket may appear after us. This is advisory —
# bridged retries the herdr socket rather than exiting, which is what actually makes a late
# socket survivable. Do NOT add Requires=: a herdr restart must not take bridged down with it.
After=herdr.service
Wants=herdr.service
[Service]
Type=simple
WorkingDirectory=%h/src/claude-bridge/bridged
ExecStart=/usr/lib/jvm/temurin-25-jdk/bin/java -jar target/bridged.jar bridged.yaml
Environment=HERDR_SOCKET_PATH=%h/.config/herdr/herdr.sock
# PATH matters more than it looks (CB-511): bridged propagates its own PATH to every worker it
# spawns, so this line decides whether the fleet can run a build at all. systemd does not source a
# login shell, so without it the daemon — and every worker — gets a bare default with no JDK/Maven.
Environment=PATH=/usr/lib/jvm/temurin-25-jdk/bin:/usr/share/maven/bin:/usr/local/bin:/usr/bin:/bin
# Secrets are NOT set here — this file is committed. Put the API/worker tokens in a private
# drop-in that systemd reads with restrictive permissions:
# systemctl --user edit bridged → [Service] / Environment=BRIDGED_API_TOKEN=...
# or point EnvironmentFile at a 0600 file:
# EnvironmentFile=%h/.config/bridged/env
Restart=on-failure
RestartSec=10s
# A bad config (e.g. a non-loopback bind without token auth) makes bridged fail fast by design.
# Give up rather than restart-loop on a permanent error.
StartLimitBurst=5
StartLimitIntervalSec=120
# The daemon reads the repo, writes worktrees, and talks to a Unix socket — it needs no more.
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=read-write
ProtectKernelTunables=true
ProtectControlGroups=true
RestrictSUIDSGID=true
StandardOutput=journal
StandardError=journal
SyslogIdentifier=bridged
[Install]
WantedBy=default.target
+80
View File
@@ -0,0 +1,80 @@
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<!--
CB-504 — launchd agent for bridged (macOS).
This is the real supervision target today: the dogfooded daemon runs on macOS, where there is
no systemd. A systemd unit ships alongside (deploy/bridged.service) for the Linux gateways
CB-308 introduces.
Install:
cp deploy/dev.ltms.bridged.plist ~/Library/LaunchAgents/
# edit the paths + JAVA_HOME below to match this host, then:
launchctl load -w ~/Library/LaunchAgents/dev.ltms.bridged.plist
launchctl list | grep bridged
Note on ordering: launchd has no "start after herdr" primitive for user agents, and neither
does systemd in a way that survives a socket appearing late. bridged retries the herdr socket
on startup instead, so an agent that comes up before herdr converges rather than dying — that
retry is the actual fix; KeepAlive below is the backstop.
-->
<plist version="1.0">
<dict>
<key>Label</key>
<string>dev.ltms.bridged</string>
<key>ProgramArguments</key>
<array>
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home/bin/java</string>
<string>-jar</string>
<string>/Users/CHANGEME/src/claude-bridge/bridged/target/bridged.jar</string>
<string>bridged.yaml</string>
</array>
<!-- Config path in ProgramArguments is relative, so the working directory must be the module. -->
<key>WorkingDirectory</key>
<string>/Users/CHANGEME/src/claude-bridge/bridged</string>
<key>EnvironmentVariables</key>
<dict>
<key>JAVA_HOME</key>
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home</string>
<key>HERDR_SOCKET_PATH</key>
<string>/Users/CHANGEME/.config/herdr/herdr.sock</string>
<!--
PATH matters more than it looks (CB-511): bridged propagates its own PATH to every worker
it spawns, so this line decides whether the fleet can run a build at all. launchd does NOT
source .zprofile/.zshrc, so without this the daemon (and therefore every worker) gets a
bare /usr/bin:/bin and no JDK or Maven. Keep the toolchain entries first.
-->
<key>PATH</key>
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home/bin:/Users/CHANGEME/Tool/apache-maven-3.9.16/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin</string>
<!--
Worker/API tokens are NOT set here: this file is committed. Export them from a private
launchd override or a wrapper script. bridged reads the API token from the env var named
by auth.tokenEnv (default BRIDGED_API_TOKEN) and only in auth.mode: token.
-->
</dict>
<key>RunAtLoad</key>
<true/>
<!-- Restart on crash, but not in a tight loop if the config is bad (bridged fails fast on a
non-loopback bind without token auth — that is a config error, not a transient one). -->
<key>KeepAlive</key>
<dict>
<key>SuccessfulExit</key>
<false/>
</dict>
<key>ThrottleInterval</key>
<integer>10</integer>
<key>StandardOutPath</key>
<string>/Users/CHANGEME/src/claude-bridge/bridged/logs/bridged.out.log</string>
<key>StandardErrorPath</key>
<string>/Users/CHANGEME/src/claude-bridge/bridged/logs/bridged.err.log</string>
<key>ProcessType</key>
<string>Background</string>
</dict>
</plist>
+60
View File
@@ -0,0 +1,60 @@
# LavinMQ — the AMQP broker behind bridged's durable ReplyInbox (CB-307 Stage 2).
#
# Why this file exists: the broker was previously run ad hoc and simply vanished from the host,
# which takes bridged down with it — AmqpReplyInbox.open throws on an unreachable broker and
# Bridged.java:187 does not guard it, so a missing broker is a hard startup failure, not a
# degraded mode. This pins the version, keeps the data, and brings itself back after a reboot.
#
# Usage:
# docker compose -f deploy/lavinmq/compose.yaml up -d
# docker compose -f deploy/lavinmq/compose.yaml ps
# docker compose -f deploy/lavinmq/compose.yaml logs -f
# docker compose -f deploy/lavinmq/compose.yaml down # keeps the volume
# docker compose -f deploy/lavinmq/compose.yaml down -v # DESTROYS held replies
#
# Management UI: http://127.0.0.1:15672 (guest / guest)
#
# This is bridged's OWN broker. Do not point bridged at any other AMQP server on this host —
# notably not the `local-rabbitmq` container, which belongs to a different project and would end
# up carrying this project's queues.
name: bridged-broker
services:
lavinmq:
# Pinned deliberately: :latest silently moves the broker under a running daemon.
image: cloudamqp/lavinmq:2.9.1
container_name: bridged-lavinmq
# The failure this deployment exists to prevent — survive reboots and Docker restarts, but
# stay down if it was stopped on purpose.
restart: unless-stopped
# Loopback-bound on purpose. LavinMQ ships a default guest/guest account, which is only
# acceptable because nothing off-host can reach it. bridged connects over 127.0.0.1, and
# binding 0.0.0.0 here would expose a broker with default credentials to the network.
ports:
- "127.0.0.1:5672:5672" # AMQP — bridged.yaml broker.uri points here
- "127.0.0.1:15672:15672" # HTTP management API + UI
# The whole point of Stage 2. Held-but-unacked replies live here; without a named volume a
# `docker compose down` would discard exactly what the durable inbox exists to protect.
volumes:
- lavinmq-data:/var/lib/lavinmq
healthcheck:
test: ["CMD", "lavinmqctl", "status"]
interval: 30s
timeout: 5s
retries: 3
start_period: 10s
logging:
driver: json-file
options:
max-size: "10m"
max-file: "3"
volumes:
lavinmq-data:
name: bridged-lavinmq-data
+131
View File
@@ -0,0 +1,131 @@
# CB-301 — Session Manager (one-shot, no reuse)
**Status:** design spec for review → delegate implementation.
**Grounded in:** `WorkerService`, `Injector`/`StatusPoller`/`TurnListener`, `MessageService`,
`BridgeMcp`, `BridgedApp` (see [wiki 9. Implementation](../wiki/9-Implementation.md)).
## Problem
`WorkerService` is **stateless about what it spawned**. Its own Javadoc says it plainly:
> "there is no registry; `list()` only asks herdr." — `WorkerService.reapOrphanWorkers` (line 288)
Consequences today:
- The daemon cannot answer "which workers did *I* spawn, in what lifecycle state, owned by whom,
since when?" without shelling to herdr for a raw agent list (no state, no ownership, no age).
- Cleanup of a worker that outlived its owning process depends entirely on the boot-time
name-nonce **reaper** (CB-117) — there is no live, authoritative roster during a run.
- `bridge_list` (CB-304) can only surface herdr's view, not a bridge-owned roster.
- There is no seam for per-session policy (checkpoint on teardown → CB-302; idle_ttl /
context_cap / drain → CB-303).
## Goal & non-goals
**Goal.** Introduce a `SessionManager` that owns an authoritative in-daemon registry of the worker
sessions this daemon process spawned, tracks each one's lifecycle state, and tears each down
deterministically. It becomes the single source of truth for the roster and the seam CB-302/303/304
build on.
**Non-goals (explicit — reuse policy chosen: one-shot, no reuse).**
- **No pooling / no reuse.** Every delegated task gets a fresh worker; a finished worker is torn
down, never handed to a later task. No "warm idle" pool, no `role@profile` keying.
- **No auto-teardown *timing*.** *When* a one-shot worker is released (immediately on turn
completion vs after an idle grace) is CB-303. CB-301 provides the **mechanism** (`release`) and
the registry; CB-303 sets the policy.
- **No checkpoint content.** Writing `STATE.md` + commit on teardown is CB-302; CB-301 only exposes
the release hook it will attach to.
"Recycle" under no-reuse is simply **release + fresh acquire** — a helper, not a pool operation.
## Design
`SessionManager` **wraps** `WorkerService` (does not replace it). `WorkerService` keeps doing the
subscription-guarded spawn/teardown mechanics; `SessionManager` adds the registry, lifecycle, and
ownership on top.
**Package:** new `dev.ltms.bridged.session` — keeps the registry/lifecycle concern separate from
the `worker` spawn mechanics. Holds `SessionManager` + `WorkerSession`.
**`recycle` is IN SCOPE for CB-301** (decided): implement `recycle(paneId, …)` = `release` the old
session then `acquire` a fresh one, asserting a new distinct paneId (the no-reuse invariant). It is
a thin convenience over the two primitives, shipped now so the no-reuse teardown+respawn path is
covered by a test from day one.
### `WorkerSession` (record or small mutable holder)
| Field | Source | Notes |
|---|---|---|
| `paneId` | `Agent.paneId()` | registry key |
| `terminalId` | `Agent.terminalId()` | for status/identity joins |
| `profile` | spawn arg | which profile spawned it |
| `cwd` | resolved cwd | the worker's working dir |
| `ownerTerminal` | caller identity (nullable) | the primary/turn that requested it; `null` = daemon/anon |
| `spawnedAtNanos` | `System.nanoTime()` | age basis for CB-303 (monotonic; no wall clock in tests) |
| `state` | lifecycle FSM | see below |
State is held in a `ConcurrentHashMap<String /*paneId*/, WorkerSession>`.
### Lifecycle state machine (one-shot)
```
SPAWNING --ready(MCP present)--> READY
READY --onDelivered--> BUSY
BUSY --onTurnComplete--> DONE
BUSY --onTurnFailed--> FAILED
READY|DONE|FAILED --release()--> RELEASED (deregistered)
SPAWNING|READY|BUSY|DONE --vanished/drop--> FAILED
```
- Transitions are driven by hooks the manager already has access to:
`WorkerPresence.markPresent` → `READY`; `TurnListener.onDelivered/onTurnComplete/onTurnFailed`
(the manager implements or decorates `TurnListener`) → `BUSY`/`DONE`/`FAILED`.
- `RELEASED` sessions are removed from the registry (teardown is terminal).
- Any state → `FAILED` on drop (worker vanished / injector `drop`), mirroring `Injector`.
### API
```java
final class SessionManager {
WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal);
void release(String paneId); // deterministic teardown + deregister
WorkerSession recycle(String paneId, ...); // release + acquire (no-reuse convenience)
Optional<WorkerSession> get(String paneId);
List<WorkerSession> roster(); // bridge-owned view (CB-304 consumes this)
// lifecycle hooks (package-private): onReady/onDelivered/onComplete/onFailed(target)
}
```
- `acquire` = `workerService.spawn(profile, requestedCwd, callerCwd)` → register `SPAWNING`.
- `release` = `workerService.stop(paneId)` → deregister. Idempotent (already-gone tolerated, matching
`WorkerService.stop`).
- `roster` joins the registry with live herdr status for a truthful "roster + live" (CB-304).
### Integration points
- **`Bridged.main`** — construct `SessionManager(workerService, ...)`; wire it as/decorating the
`TurnListener` alongside `CompletionResolver` so it sees turn boundaries, and give it the
`WorkerPresence` signal for `READY`.
- **`BridgeMcp.spawn` / `BridgedApp.spawnWorker`** — route spawn through `SessionManager.acquire`
(carry `callerTerminal` as `ownerTerminal`). **`bridge_stop` / `DELETE /workers/{paneId}`** →
`SessionManager.release`.
- **`bridge_list` / `GET /sessions` (CB-304 later)** — read `SessionManager.roster()`.
- **`MessageService`** — no change required for one-shot; a later CB-303 auto-release hook can call
`release` from `onTurnComplete` under policy.
## Acceptance (tests, no live herdr — fakes as elsewhere)
1. `acquire` registers a `SPAWNING` session with the right owner/profile/cwd; a second `acquire`
yields a **distinct** paneId and a **distinct** session (no reuse).
2. Presence signal moves `SPAWNING → READY`; a delivered turn moves `READY → BUSY → DONE`.
3. `release` tears the worker down via `WorkerService.stop` and removes it from `roster()`;
a second `release` on the same paneId is a harmless no-op.
4. `onTurnFailed` / drop moves the session to `FAILED` and it is absent from the live roster.
5. `recycle` produces a new paneId and the old one is gone (no-reuse invariant).
6. `roster()` reflects exactly the sessions acquired-minus-released, joined with live status.
## Seams left open (deliberately)
- **CB-302** — attach a checkpoint step (`STATE.md` + commit) to the `release` path.
- **CB-303** — a policy loop over `roster()` using `spawnedAtNanos`/state to auto-`release` on
`idle_ttl`, or drain on `context_cap`.
- **CB-304** — `bridge_list` reads `roster()` for a bridge-owned roster + live join.
+183
View File
@@ -0,0 +1,183 @@
# CB-301-ext — Worktree provisioning + config-parity overlay
**Status:** ✅ shipped — implemented at commit `97ecc71` (per-worker git worktree + config-parity
overlay). As-built: `session/GitWorktrees.java` behind the `Worktrees` port, wired in
`Bridged.main` and configurable via `worktreeRoot` / per-profile `parityOverlay`
(see `bridged.example.yaml`). Branch/worktree surface in `bridge_list` landed with CB-304
(`9fe04bf`); the worker-opened-PR checkpoint landed as CB-302 (`64e70ef`).
**Extends:** [CB-301 Session Manager](CB-301-Session-Manager.md) (shipped, commit `54d907c`).
**Realizes:** the config-parity requirement in [Worker Git Workflow](Worker-Git-Workflow.md).
**Grounded in:** `SessionManager`, `WorkerService.spawn/effectiveCwd`, `BridgedConfig.Worker`,
`inject/…LsofPeerPidLookup` (the `ProcessBuilder` exec pattern).
## Problem
CB-301 gives each worker a session record but every worker still runs in the **primary's own
working tree** (`cwd = callerCwd`). One worker at a time is safe; two **parallel implementers** would
stomp each other. We need each implementer session to get an **isolated git worktree on its own
branch** — *without* degrading the worker: a bare worktree checks out **tracked files only**, so it
silently drops the untracked/local config (`.claude/settings.local.json`, the locally-modified
`.mcp.json`, `.env`) that makes a session a full peer of the primary. **CB-301-ext provisions the
worktree AND hydrates it to config parity**, so a worker differs from the primary only in the LLM
provider.
## Decisions (locked)
1. **Opt-in, not default.** A worktree is provisioned **only** when the caller requests one. Absent a
request, `acquire` behaves exactly as it does today (shared primary tree) — auditors, smoke tests,
and conversational workers are unaffected. **Backward compatibility is a hard requirement.**
2. **Copy-overlay + `--skip-worktree`, not symlink.** Each parity file is **copied** primary→worktree
(isolation-friendly, no symlink type-change noise on tracked files). For a *tracked* overlay file
(`.mcp.json`) the worktree copy is then marked `git update-index --skip-worktree`, so the worker's
commits can **never** include the parity overlay. Ignored files (`settings.local.json`) stay
ignored in the worktree (shared `info/exclude`), so no marking is needed.
3. **Branch persists; worktree is disposable.** `release` runs `git worktree remove --force` (the
working dir is throwaway) but **never deletes the branch** — the branch holds the worker's commits
and its PR (CB-302). Teardown of the checkout ≠ teardown of the work.
4. **Git behind a seam.** SessionManager depends on a `Worktrees` interface (production impl shells
`git` via `ProcessBuilder`; tests use a fake). No live `git` in unit tests — mirrors the
`WorkerService`/`FakeHerdr` seam.
## Design
### `WorktreeRequest` (new, nullable = "no worktree")
```java
package dev.ltms.bridged.session;
/** Ask acquire() to provision an isolated worktree. null ⇒ run in the shared primary tree. */
public record WorktreeRequest(String ticketSlug, String baseRef) {
// ticketSlug seeds the branch name; baseRef null/blank ⇒ current HEAD of the repo.
}
```
### `WorkerSession` — two nullable fields added
| Field | Notes |
|---|---|
| `worktree` | absolute path of the provisioned worktree; `null` ⇒ shared tree |
| `branch` | the worker's branch (`worker/<slug>-<nonce>`); `null` ⇒ shared tree |
Add to the record + `withState`. A `null` worktree keeps every existing test and the shared-tree path
untouched.
### `Worktrees` seam (new)
```java
package dev.ltms.bridged.session;
public interface Worktrees {
/** git -C <repoRoot> worktree add <path> -b <branch> <baseRef|HEAD>. Returns the worktree path. */
String add(String repoRoot, String branch, String baseRef);
/** git -C <repoRoot> worktree remove --force <path>. Idempotent (already-gone tolerated). */
void remove(String repoRoot, String worktreePath);
/** Copy each existing overlay path repoRoot→worktree; mark tracked ones --skip-worktree. */
void overlayParity(String repoRoot, String worktreePath, List<String> overlay);
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
String repoRoot(String cwd);
}
```
- **Production impl** `GitWorktrees implements Worktrees` — `ProcessBuilder` per path, `redirectErrorStream(true)`, non-zero exit → a `WorktreeException`. Worktree location = `<worktreeRoot>/<nonce>` where `worktreeRoot` is a daemon setting (default: sibling `../.bridged-worktrees` of the repo root — **outside** the repo, never nested).
- `overlayParity` per file: skip if absent in `repoRoot`; else copy into the worktree; if `git -C <wt> ls-files --error-unmatch <path>` succeeds (tracked), run `git -C <wt> update-index --skip-worktree <path>`.
### `acquire` — extended, old signature preserved
```java
// existing (unchanged): shared tree
WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal);
// new overload: provision a worktree when wt != null
WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal,
WorktreeRequest wt);
```
When `wt != null`:
1. `repoRoot = worktrees.repoRoot(firstNonBlank(requestedCwd, callerCwd))`.
2. `branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce`.
3. `path = worktrees.add(repoRoot, branch, wt.baseRef())`.
4. `worktrees.overlayParity(repoRoot, path, cfg.parityOverlay())`.
5. `spawn(profile, path, callerCwd)` — **the worktree path becomes the worker's cwd** (highest
precedence in `WorkerService.resolveCwd`).
6. register the session with `worktree=path, branch=branch`.
7. **On any failure in 1–5, unwind**: if the worktree was added, `remove` it; do not leave a dangling
registry entry. (Guard/spawn already throw before herdr on a bad base_url — unchanged.)
### `release` — remove the worktree, keep the branch
```java
public void release(String paneId) {
WorkerSession s = registry.remove(paneId);
workerService.stop(paneId); // existing
if (s != null && s.worktree() != null) {
worktrees.remove(worktrees.repoRoot(s.cwd()), s.worktree()); // branch is NOT deleted
}
}
```
### Config — `BridgedConfig.Worker.parityOverlay` + a `worktreeRoot`
- Add `List<String> parityOverlay` to the `Worker` record (12th field). Compact-constructor default
when null/empty: `[".mcp.json", ".claude/settings.local.json", ".env", ".envrc"]` (missing paths are
silently skipped, so the default is safe across repos). Update `withProfile`.
- Add a top-level daemon setting `worktreeRoot` (String, nullable → `<repoParent>/.bridged-worktrees`).
- `@JsonIgnoreProperties(ignoreUnknown = true)` already set → additive, no parser breakage.
### Surface: MCP + REST
- `bridge_spawn` gains an optional `worktree` arg: `true`, or a ticket slug string. Truthy ⇒ build a
`WorktreeRequest(slug, null)` and call the 5-arg `acquire`.
- `POST /workers` gains `worktree` (+ optional `ticket`) in the body/query, same mapping.
- `workerView`/`view(WorkerSession)` include `worktree` and `branch` **when non-null** (omit for
shared-tree sessions, so existing response assertions for shared-tree spawns are unchanged).
## Flow
```mermaid
flowchart TD
A["acquire(profile, ..., WorktreeRequest?)"] --> B{"worktree<br/>requested?"}
B -->|"no (default)"| C["spawn(cwd = callerCwd)<br/>— shared tree, unchanged"]
B -->|yes| D["repoRoot = rev-parse --show-toplevel"]
D --> E["git worktree add path -b branch base"]
E --> F["overlayParity: copy local config in;<br/>--skip-worktree the tracked ones"]
F --> G["spawn(cwd = worktree path)"]
G --> H["register worktree + branch on the session"]
E -.->|"add/overlay/spawn fails"| X["unwind: remove worktree,<br/>no dangling registry entry"]:::warn
C --> R["worker is a full peer of the primary"]:::goal
H --> R
classDef warn fill:#b7791f,stroke:#7b341e,color:#ffffff;
classDef goal fill:#2b6cb0,stroke:#2a4365,color:#ffffff;
```
*Green = the invariant: whether shared-tree (parity for free) or worktree (parity via overlay), the
worker matches the primary. Amber = the failure-unwind path.*
## Acceptance (fake `Worktrees`, no live git)
1. **Backward-compat:** `acquire` with **no** `WorktreeRequest` makes **zero** `Worktrees` calls,
spawns with `cwd = callerCwd`, and records `worktree == null` / `branch == null`. (Every CB-301
test still passes.)
2. **Provision:** `acquire(..., new WorktreeRequest("cb-999", null))` calls `add(repoRoot,
"worker/cb-999-<nonce>", null)`, then `spawn` receives the returned worktree path as
`requestedCwd`; the session records that path + branch.
3. **Overlay:** `overlayParity` is invoked with the profile's `parityOverlay` (default list when
unset); the fake asserts tracked paths were `--skip-worktree`'d and missing paths skipped.
4. **Release removes worktree, keeps branch:** releasing a worktree session calls
`Worktrees.remove(repoRoot, path)` and performs **no** branch-delete; a shared-tree session's
release makes no `Worktrees` calls.
5. **Failure unwind:** a fake `add` that throws ⇒ `acquire` throws, the session is **not** registered,
and no worker is left running (spawn not reached / torn down).
6. **Distinct worktrees:** two worktree acquires yield **distinct** branches and paths (no collision).
## Constraints & exclusions (standing, non-negotiable)
- **Only a worker sets `ANTHROPIC_BASE_URL`.** Worktrees touch cwd + files only; env path is unchanged
— the guard still runs before any herdr call.
- **Worker cwd stays inside the primary repo** (the worktree is a checkout of it) — never `$HOME`.
- **`.mcp.json` and `wiki/` never enter a worker commit.** `.mcp.json` is overlaid for *reference* but
`--skip-worktree`'d so it can't be staged; `wiki/` is a submodule the worker must not touch. The
CB-302 commit step (and the implementer skill) exclude both.
- **The overlay list stays explicit + minimal** (trust: local secrets flow to an off-subscription
worker). No blanket tree copy.
## Seams left for later
- **CB-302** — the worker commit → push → PR checkpoint runs *inside* the worktree on its branch.
- **CB-304** — `roster()` rows surface `worktree`/`branch` for the fleet view.
+143
View File
@@ -0,0 +1,143 @@
# CB-306 — Spawn-Readiness Gate (launcher-owned terminal readiness)
**Status:** design note / delegation spec (branch `worker/cb-306-readiness`)
**Issue:** gitea `lms/claude-bridge` #4
**Owner of the behaviour:** `ClaudeCodeLauncher` (the `PeerLauncher` adapter) — NOT core.
## 1. Problem
`bridge_spawn` today returns a session the instant the herdr pane is started. The pane is not
yet a usable Claude REPL — it may still be sitting at the folder-trust prompt, or the CLI may
never come up at all. Nothing blocks or times out on that. Consequences:
- A `bridge_send` to a not-yet-ready worker surfaces as a **~60 s MCP-client timeout** (the send
blocks waiting for a turn that can't start) instead of a fast, explicit spawn failure.
- A worker stuck at the folder-trust prompt lingers in `SPAWNING` forever; nothing fails it.
We want **fail-fast spawn**: `spawn()` returns only once the peer is genuinely up and usable in
its terminal, or throws a clean error (and leaves no orphan pane) within a bounded timeout.
## 2. What already exists (do NOT rebuild)
`SessionManager` + `PresenceBridge.markPresent()` already flip a session `SPAWNING → READY` on
**any MCP contact from the worker** (`onReady(terminal)` → `transitionByTerminal(SPAWNING, READY)`,
SessionManager ~line 249, "worker became available on the bridge MCP"). That is the **delivery
lifecycle** and it stays exactly as-is. CB-306 does **not** touch it and does **not** replace it.
CB-306 adds a *complementary, launcher-side* gate: the launcher guarantees the **terminal** is a
live, interactive REPL before it hands a handle back. The two signals are layered:
| Signal | Owner | Means | CB-306 |
|---|---|---|---|
| terminal reaches interactive REPL (herdr `IDLE`) | `ClaudeCodeLauncher` (this ticket) | pane is past folder-trust, CLI is up | **NEW — the spawn gate** |
| first MCP contact → `SPAWNING→READY` | core (`SessionManager`/`PresenceBridge`) | worker spoke to the bridge | unchanged |
## 3. The readiness predicate (herdr status)
`AgentStatus.fromWire` maps herdr's wire strings to `IDLE | WORKING | BLOCKED | DONE | UNKNOWN`.
A freshly started pane that has **not** reached an interactive Claude — including one stalled at
the folder-trust prompt — reports **`UNKNOWN`** (herdr has not detected a Claude REPL yet). Once
the CLI is up and settled at its prompt it reports **`IDLE`**.
**Predicate:** the pane is *ready* when `AgentControl.status(target)` first returns an
**injectable** state (`IDLE`, `BLOCKED`, or `DONE` — reuse `AgentStatus.injectable()`). `UNKNOWN`
= not ready. `WORKING` alone is ambiguous this early and should not by itself satisfy readiness;
wait for an injectable state. (We do not need to know *why* a pane isn't ready — a trust stall,
a crash, and a slow start all present as "never becomes injectable" and all correctly time out.)
## 4. Contract change on `spawn(SpawnRequest)`
`ClaudeCodeLauncher.spawn(SpawnRequest)` becomes **block-until-ready-or-throw**:
1. Start the pane exactly as today (`spawn(profile, cwd, callerCwd) → Agent`, build env + guard +
argv, `spawnInTab`/`spawnAsPane`, unique-named).
2. **Poll** `agentControl.status(paneId)` every `pollIntervalMs` (~300 ms) until it is `injectable()`
or `spawnReadyTimeoutMs` elapses.
3. **Ready** → return the `WorkerHandle(paneId, terminalId)` as today.
4. **Timeout** → the launcher **closes the pane it started** (and its tab, via the same path
`release`/`stop` uses) and throws **`PeerUnreachableException`** (new, in `dev.ltms.bridged.peer`).
No orphan pane is left behind — the launcher cleans up its own failed birth.
`spawnReadyTimeoutMs == 0` (or unset) **disables** the gate = legacy non-blocking behaviour, so the
change is opt-in per deployment and existing tests that don't configure it keep their old semantics.
### Testability seam (required)
Do **not** call `Thread.sleep` directly in the poll loop against a real clock — unit tests must not
real-sleep. Introduce a small injectable seam, mirroring the existing `StatusPoller` style:
- a `LongSupplier nowMillis` (monotonic clock) **and** a sleep/wait hook (e.g. a
`Sleeper`/`Waiter` functional interface, or reuse whatever `StatusPoller` already uses), both
defaulting to the real implementations in the production constructor and overridable in tests.
Unit tests (add to the existing `ClaudeCodeLauncher` test):
- fake `AgentControl` returns `UNKNOWN` a few times then `IDLE` → `spawn` returns the handle; assert
no `close` was called.
- fake `AgentControl` always `UNKNOWN` → `spawn` throws `PeerUnreachableException`; assert the pane
**was closed** (verify `close(paneId)` invoked) and the fake clock advanced past the timeout.
- `spawnReadyTimeoutMs == 0` → `spawn` returns immediately without polling (legacy path).
## 5. Config
Add to the launcher-level config (a bridged-level knob, not per-profile) in `bridged.yaml` +
`BridgedConfig`:
```yaml
spawn_ready_timeout_ms: 20000 # 0 disables the gate (legacy non-blocking spawn)
spawn_ready_poll_ms: 300
```
Jackson ignores unknown keys, so omitting them in existing YAML is safe; pick sane defaults in code
(`20000` / `300`). Keep the names consistent with existing config field style in `BridgedConfig`.
## 6. Core / MCP propagation
`SessionManager.acquire(...)` already calls `launcher.spawn(req)`. A thrown
`PeerUnreachableException` must propagate out as a **clean spawn failure**:
- The **worktree** acquire path already has a try/catch that cleans up a provisioned worktree when
`spawn` throws — verify the new exception flows through it (worktree removed, nothing registered).
- The **non-worktree** path registers the session only *after* `spawn` returns, so a throw means no
half-live `SPAWNING` session is ever registered — confirm this and add a test.
- `bridge_spawn` (MCP verb) must return an **error result** carrying the exception message, not a
success with a dead session. Trace `BridgeMcp`/`BridgedApp` spawn handlers and make sure the
exception becomes a clean tool error, not an uncaught 500 with a stack trace.
**Out of scope (do NOT do here):** gating `bridge_send` on session `READY` (existing status-gate +
this spawn gate already close the window), MCP-handshake-as-readiness signal, the CB-307 broker,
any config `kind:` discriminator, any second adapter.
## 7. Definition of done
- `ClaudeCodeLauncher.spawn` blocks until injectable or throws `PeerUnreachableException` +
self-reaps the pane; gate disabled when timeout is 0.
- New `PeerUnreachableException` in `dev.ltms.bridged.peer`.
- Config knobs wired (`spawn_ready_timeout_ms`, `spawn_ready_poll_ms`) with safe defaults.
- Existing `SPAWNING→READY` MCP-contact transition untouched.
- New unit tests (ready / timeout+reap / disabled) green; **all existing tests still pass unchanged**.
- Build clean via the worker's own `mvn` (primary re-runs the authoritative IDE + `mvn clean install`
gate — self-reports are not verified facts).
## 8. Sequence
```mermaid
sequenceDiagram
participant SM as SessionManager.acquire
participant L as ClaudeCodeLauncher.spawn
participant T as herdr (AgentControl)
SM->>L: spawn(SpawnRequest)
L->>T: start pane (env+guard+argv)
loop until injectable or timeout
L->>T: status(paneId)
T-->>L: UNKNOWN / IDLE
end
alt reached injectable
L-->>SM: PeerHandle(id, terminalId)
else timed out
L->>T: close(paneId) + tab
L-->>SM: throw PeerUnreachableException
SM-->>SM: no session registered / worktree cleaned
end
```
*Figure — the launcher blocks in `spawn` until the pane is a usable REPL, else self-reaps and throws.*

Some files were not shown because too many files have changed in this diff Show More