Compare commits

..

44 Commits

Author SHA1 Message Date
kevin 871b595954 CB-518: state the primary's orchestration as an explicit, ordered flow
CI / build (pull_request) Successful in 1m23s
CB-517 moved orchestration policy into CLAUDE.md but left it as a bullet
list, so the procedure was implicit: the order of operations had to be
reconstructed from a parallelisation bullet, and nothing said when to
review or when to tear down. A policy you have to reassemble on each
task is one you will reassemble differently each task. Restate the
primary's half as a numbered 0-8 flow — role check, split, gate, spawn
all, send all, collect, verify, review, adjudicate — so that following
it is checkable against the tool calls rather than a matter of recall.

Two steps carry the load. Spawn and send are separate on purpose:
folding them into one loop is what silently serialises work that was
meant to fan out. And review is now its own step ahead of the merge
rather than a clause inside it, because the two have opposite owners —
reviewers fan out over the diff (never the implementer of the scope
they review, and briefed from the diff rather than the author's
rationale, which carries the same blind spot), while adjudication, the
merge and teardown stay with the primary. Merging on a reviewer's word
is delegating the gate by proxy, so the step says so outright.

Nothing is dropped. The six bullets that trailed the tool table are
relocated into the step that owns each — profile explicitness into
spawn, playbook naming and self-containment into send, claim
verification into its own step, the ~60s blocking-send cap into a note
beneath the flow — and the table stays as the intent→tool lookup.

The wiki pointer moves with it. The block is canonical only if its
template matches byte for byte, so the template was produced by
splicing the block out of CLAUDE.md rather than by editing it in
parallel, and the sync check the repo documents passes. Bumping the
pointer in the same commit keeps charter and template versioned
together, as CB-517 did.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw1FosEt3Noix5GG9wJ2r
2026-08-04 22:32:19 +07:00
Dai Ha b67b1585c2 CB-517: deploy LavinMQ as a pinned, durable, self-restarting broker
CI / build (push) Successful in 1m23s
The CB-307 durable ReplyInbox needs an AMQP broker, but the one behind it
was run ad hoc and had simply vanished from the host — which takes the
whole daemon with it, since AmqpReplyInbox.open throws and Bridged.java:187
does not guard it. A missing broker is a hard startup failure, not a
degraded mode, so 'how the broker runs' is part of the system, not a local
detail.

Pinned to 2.9.1 (:latest would move the broker under a running daemon),
data on a named volume (held-but-unacked replies are the entire point of
Stage 2 — a plain 'compose down' would discard exactly what durability
protects), and restart: unless-stopped so it comes back after a reboot
instead of disappearing again.

Ports are bound to 127.0.0.1 deliberately: LavinMQ ships a default
guest/guest account, which is only acceptable while nothing off-host can
reach it.

Verified by driving the production AmqpReplyInbox against this deployment
(publish/peek/dedup/FIFO/ack, then reconnect): 8/8 including redelivery of
the unacked message. That pairing had never been exercised — the
@Tag("contract") test runs against a RabbitMQ container, and is excluded
from the default build, so mvn clean install covers the broker path zero
times.
2026-08-04 16:17:21 +02:00
Dai Ha 979b2b5632 CB-517: add bridge_whoami and make the bridge prompt a portable charter
CI / build (push) Successful in 1m25s
The communication rules lived only in two opt-in skills, so nothing
always-on told the primary how to orchestrate and nothing guaranteed a
worker loaded its playbook. Move protocol and policy into CLAUDE.md,
which a worker inherits for free (its worktree is a checkout of this
repo), and leave the skills as pure per-job procedure.

bridge_whoami closes the load-bearing gap: every tool already consumed
the caller identity ConnectionIdentity resolves from the connection, but
none reported it, so an agent had to infer its own role from side
channels the daemon does not control. Guessing fails asymmetrically — a
primary acting as a worker is refused by the authz gate and learns at
once, while a worker acting as the primary ends its turn without
bridge_reply and the sender silently receives nothing. The tool reuses
the same Principal the gate is built on, so the two cannot disagree; the
primary gets role only (handing it a sessionId it does not own would
invite the forged reply Authz refuses), and a worker missing from the
registry still gets role + sessionId rather than 'unknown'.

The CLAUDE.md block is written to be copied as-is into any project that
mounts the bridge: repo-local details (Authz paths, the .mcp.json/wiki
exclusions, the skill names) moved below it into a project addendum, and
every role-inference fallback is stated one-way — the mount-name signal
only holds for mcp__bridge__* (the launcher fixes it), not for the
primary's mount, which each project names itself. The wiki carries the
block verbatim as the template, with a sync check.

Because this repo IS the bridge, that block is shipped surface, not
documentation: the addendum adds a mandatory checklist mapping each part
of the code to the part of the prompt it can invalidate.

Also: delegate-by-default policy for the primary — the test is not 'could
I do this faster myself' but 'can I write a brief good enough for a
worker'.

mvn clean install: 356 tests green (353 + 3 for whoami); ide_diagnostics
clean on both changed files.
2026-08-04 16:01:16 +02:00
kevin cf4ad186ab CB-516: fail a delegation when its worker session is released
CI / build (push) Successful in 1m30s
Tearing a worker down left the send that was waiting on it stranded. The
rendezvous waiter stayed open, so a blocking bridge_send kept blocking and an
async one kept reporting PENDING until ASYNC_TIMEOUT_MS — thirty minutes —
even though the worker provably no longer existed and the delegation could
never complete.

Observed repeatedly this session: a task sitting at
{"phase":"pending","detail":"worker unknown"} long after I had personally
deleted the pane. The maddening part is that poll() already HAD the evidence —
it calls liveStatus() to build that detail string, gets back "unknown", and
reports PENDING anyway. It also never reached /metrics under any outcome, so a
stalled delegation was invisible to both the task view and the dashboard. The
only way I ever diagnosed one was reading the worker's pane over the herdr
socket by hand.

Fixed at the choke point rather than by guessing from status strings:
SessionManager.release() is the single path every teardown funnels through
(REST stop, MCP stop, idle-TTL reaper, recycle, shutdown drain), so it now
notifies a release listener with the terminalId, and Bridged.main wires that to
MessageService.abandon(). Abandon fails the open waiter with a real reason.

Deliberately NOT done by inferring "gone" from liveStatus(): that method
collapses a vanished pane, a wedged worker and a herdr hiccup into the same
"unknown" string, so acting on it would fail live delegations during a
transient blip. An explicit lifecycle signal cannot be ambiguous.

Notified before launcher.stop() so a blocked caller fails fast, and wrapped so
a listener failure can never prevent the teardown it is reacting to.

Resolving as a failure rather than letting it time out also means the outcome
is counted — a torn-down delegation now shows up as
sends_total{outcome="failed"} instead of nothing at all.

353 tests (was 346): abandon fails a waiting send / is a no-op with no waiter /
never clobbers a send the worker already answered, an abandoned async task
polls FAILED rather than PENDING, release notifies with the right terminal,
releasing an unknown pane notifies nobody, and a throwing listener does not
block teardown.

Verified live on the running daemon, reproducing the original scenario:
  async send        -> {"phase":"pending","detail":"worker working"}
  DELETE the worker -> 204
  poll              -> {"phase":"failed","detail":"the worker session was
                        released before it replied"}
  /metrics          -> bridged_sends_total{outcome="failed"} 1
Previously that poll returned PENDING for thirty minutes and the counter stayed
empty.

NOT addressed here, and worth deciding separately: a worker that simply never
replies (rather than being released) still rides out the full 30-minute
ASYNC_TIMEOUT_MS. That is a policy question about long-running delegations, not
a correctness bug.
2026-08-01 23:57:45 +07:00
kevin 5100f215cf CB-515: regression-protect the turn-attribution guards
CI / build (push) Successful in 1m33s
Five tests pinning the invariants that decide WHICH turn a reply belongs to.
These protect against a silent correctness bug — a reply attributed to the
wrong turn — not against a crash, which is why they were worth picking over
higher-percentage coverage gaps.

Chosen by blast radius, not by uncovered-line count. Both guards are compound
conditions with a side that never executed, i.e. exactly the shape where a
clause can be deleted as "redundant" and every existing test still passes.

CompletionResolver:
- The CB-115 misattribution guard suppresses a completion when the scrape is
  byte-identical to the pane at delivery. Its !scrapeFailed clause was
  unexercised: delete it and a FAILED read is misread as "no output change",
  so the send is suppressed and hangs to the caller's timeout instead of
  resolving. The new test sets the baseline to "" so the empty tail from a
  failed read would byte-match and wrongly suppress — built to die precisely
  when that clause dies.
- The fail() guard leaves an already-resolved waiter alone. The new test also
  asserts agent.read is never called, so the worker is not scraped for a send
  nobody is waiting on.

Rendezvous: a second resolution of an already-completed waiter returns false
and does not overwrite the first value, for both resolveCompletion and
resolveFailure.

Verified by sabotage, one guard at a time: removing !scrapeFailed reds
resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent; removing the
isDone() clause reds failLeavesAnAlreadyResolvedWaiterUntouchedAndSkipsTheScrape.
(The first attempt at the second sabotage reported a false pass — the patch hit
an identically-worded guard earlier in the file. Line-targeted and re-run.)

346 tests, was 341. Worker-implemented on the local-vLLM profile; it noticed
three of the eight cases I asked for already existed and said so with names
rather than duplicating them.

Also of note: the first delegation of this ticket wedged the worker — the pane
showed a zsh parse error and it went idle with an untouched worktree, task stuck
pending. The retry differed only in phrasing the same requirements as prose
instead of quoting Java boolean expressions. Filed as a bridge robustness
concern: injected content shares a channel with control, and a wedged turn is
invisible in both the task view and /metrics.
2026-08-01 23:48:05 +07:00
kevin f129e9b7cd Merge CB-514: MessageService coverage for timeout, answer, poll and lock edges
CI / build (push) Successful in 1m29s
Six tests on the delivery core, the most load-bearing class in the project,
which sat at 74% with its hardest paths unexercised: answer() (the CB-205
ask-answer resolution) had 11 of 23 lines uncovered, send()'s timeout branches
7 of 21, poll() 7 of 18, and tryLock 3 of 4.

Covered: TIMED_OUT_QUEUED vs TIMED_OUT_WORKING (undelivered vs delivered-but-
silent), an answered worker that never sends its follow-up reply, an unknown
ticket, a completed async ticket reporting its reply and replySource, and a
second concurrent send to the same session returning BUSY rather than hanging.

Worker-implemented on the local-vLLM profile, self-verified: it reported
'Tests run: 341, Failures: 0' and an independent run of its branch agrees
exactly. Additive only — 101 lines in one test file, no main/ source touched.

Quality is good on its own terms, not just green: the async-completion test
polls to a deadline instead of sleeping and hoping, the contention test
releases the blocked send so the test thread is not left pinned, and every
assertion is on a specific Outcome rather than 'nothing threw' — the failure
mode an earlier worker produced in CB-510.

Branch pushed by the worker; PR left to the primary since GITEA_TOKEN is not
granted to that profile by design.
2026-08-01 23:26:54 +07:00
kevin 4aed45de19 CB-514: add MessageService coverage for timeout, answer, poll, and lock edges 2026-08-01 23:24:54 +07:00
kevin 1e6daa5c73 CB-513: test the MCP-side authorization gate (BridgeMcp 27.4% -> 57.4%)
CI / build (push) Successful in 1m13s
CB-505 claimed authorization is "enforced on both entry paths". It is — but
only REST was ever tested. Coverage showed BridgeMcp.deny(), principal(),
callerTerminal(), worktreeRequest() and every tool-registration lambda at ZERO
executed lines: no test had ever constructed a BridgeMcp, because the existing
BridgeMcpTest calls only the static handler methods. So the MCP half of the
security control had ten REST tests' worth of nothing behind it.

An unexercised security control is a claim, not a control.

Made testable by separating policy from plumbing rather than by reaching for a
mocking library the project does not use:
- denyFor(Principal, Action, target) is the decision — testable directly.
- deny(exchange, ...) shrinks to pulling the caller out of the SDK exchange.
- principalFrom(role, terminal, pid) extracts identity reconstruction from
  McpSyncServerExchange, an SDK type with no fake available.

Moved the `authz == null` enforcement switch OUT of the exchange-facing wrapper
and INTO denyFor. Found by a failing test: as written, any future tool calling
denyFor directly would have silently skipped the gate. The switch now lives with
the decision it governs.

New BridgeMcpAuthzTest constructs a real BridgeMcp — which is why coverage moved
so far, since that also runs the constructor and all the tool wiring — and pins
the table on this path: primary orchestrates, worker cannot; worker replies only
as itself; the primary cannot forge a worker reply; anonymous gets nothing; and
401-shaped vs 403-shaped refusals are counted apart.

Verified as real controls, not decoration: with the gate forced open, 5 of the 9
fail. 335 tests (was 326).
2026-08-01 23:11:56 +07:00
kevin 4c015d76b7 Merge CB-512: wire bridged_push_nudges_total (worker-implemented, self-verified)
CI / build (push) Successful in 1m30s
Fixes one of the three counters declared in BridgedMetrics but never
incremented, so bridged_push_nudges_total{outcome=delivered|exhausted} now
actually appears on /metrics. An absent series reads as 'no push failures ever'
rather than 'not measured', which is the misleading case.

Implemented end to end by an opencode worker on the local-vLLM profile in an
isolated worktree, and this is the first delegation where the worker verified
its own work: it ran mvn, hit a real test failure, iterated, and reported
'Tests run: 326, Failures: 0' — which matches an independent run of its branch
exactly. Every earlier delegation reported results it had no way to check,
because CB-511 had not yet given workers a PATH with a toolchain.

It also committed and pushed its own branch unprompted, and was straight about
the one thing it could not do: opening the PR, since GITEA_HOST/GITEA_TOKEN are
not granted to that profile ('URL rejected: No host part'). That grant is opt-in
per profile by design, so the merge is primary-side as intended.

Diff reviewed and correct on every constraint, including the subtle ones:
delivered counted inside the try after a successful send (not the catch),
exhausted only on the reminder-cap branch and not the other two STOP paths, and
a single Metrics instance moved above pushLoop and shared with MessageService.
2026-08-01 23:01:07 +07:00
kevin 37a11cd168 CB-512: wire bridged_push_nudges_total metric increments 2026-08-01 22:56:32 +07:00
kevin 22ad24db6c CB-511: give workers a toolchain — propagate the daemon PATH, add profile env:
CI / build (push) Successful in 1m17s
Workers could not run `mvn` or `java`. Every delegated task that asked for a
build came back "mvn is not on PATH", and the worker was right.

Root cause: HerdrPeerLauncher seeded the worker environment with an EMPTY map,
so bridged passed only the vars it explicitly set (OPENCODE_CONFIG, GITEA_TOKEN,
ANTHROPIC_*) and never PATH. herdr merges that map into its own process env, so
a worker inherited whatever PATH the herdr SERVER was started with. On this host
that server (pid 79870, PPID 1) had been up since Jul 4 with a PATH containing
neither the JDK nor Maven. Confirmed on a live worker: its PATH was byte-identical
to herdr's, and the only var bridged had contributed was OPENCODE_CONFIG.

The failure was invisible and non-deterministic: the fleet's capabilities
depended on how a long-lived daemon happened to be launched weeks earlier. There
are three herdr processes on this box with three different PATHs; the one owning
the socket is the one without a toolchain. bridged itself HAD Maven on PATH the
whole time — it just never passed it on.

It also quietly contradicted the project's own principle that "a worker is a
full peer of the primary", and the implementer skill's instruction to build,
commit and open a PR. Every delegation so far has depended on the primary
running the build gate.

Fix: baseEnv(cfg) seeds each worker with the daemon's own PATH, then applies the
profile's new optional env: map. Adapter-specific vars are layered on top and
therefore win — that ordering is load-bearing, not incidental: it stops an env:
entry from overwriting ANTHROPIC_BASE_URL and slipping past SubscriptionGuard,
which is checked against the profile's baseUrl alone. Pinned by a test.

Because the default is now the daemon's PATH, both supervision units set PATH
explicitly — launchd and systemd do not source a login shell, so under CB-504
the daemon (and every worker) would otherwise get a bare /usr/bin:/bin and this
bug would silently return in production.

324 tests (was 321): daemon-PATH propagation, profile env: passthrough including
an explicit PATH override, and the guard-bypass ordering.

Verified live: daemon restarted, worker spawned, and asked to run the tools —
"Apache Maven 3.9.16", "java version 25.0.2". Previously both were absent.
2026-08-01 22:34:30 +07:00
kevin e94c1b8841 CB-510: SessionReaper wrapper tests (0% -> 86.7%)
SessionReaper had no tests at all. Its TTL *policy* was already well covered
(SessionManager.reapIdle, 6 cases in SessionManagerTest); what was untested was
the thread wrapper around it — idempotent start/stop and whether the loop
actually runs and actually stops.

Observed through an injected clock rather than by sleeping and hoping: reapIdle
reads nowNanos exactly once per call, so the tick count IS the iteration count.
Waits are bounded polls, not fixed sleeps, and nothing asserts an exact
timing-derived number — flaky counts would be worse than no test.

321 tests (was 318); line coverage 66.9% -> 67.9%.

Drafted by an opencode worker on the new local-vLLM profile (branch
worker/cb-510-session-reaper-test-cd1793-1). Its structure and setup were good
and it was honest that it could not run mvn. But its third test asserted
NOTHING — it started the reaper, slept, stopped it, and relied on "no throw",
with a comment claiming that proved the loop had run. It did not: verified by
sabotage, all three of its tests passed against a start() replaced with an
immediate return.

Rewritten so the assertions can fail for the right reason. Same sabotage now
fails 2 of 3 (the third only pins stop()-before-start(), where "does not throw"
genuinely is the contract). Uncomfortably on the nose given this task began as
a hunt for tests that do not mean anything.
2026-08-01 21:42:06 +07:00
kevin cc0ec65714 CB-509: add JaCoCo coverage reporting
Build-time tooling only — never a compile or runtime dependency, so it adds
nothing to the shipped jar and no new transitive surface to the artifact.
(Noting per CLAUDE.md that the pom CVE gate could not be run: no JetBrains MCP
server is connected this session.)

Report at target/site/jacoco/index.html, machine-readable at jacoco.csv.

Deliberately NO check rule or threshold. A coverage gate rewards writing tests
that merely execute lines, which is the exact failure mode this codebase has
already been bitten by — CB-507 shipped a null-argument NPE with 311 green
tests because FakeWorktrees.repoRoot records its argument instead of shelling
out, so the broken line was covered and still wrong. Coverage is a map of where
to look, not a target to hit.

Baseline: 66.9% line, 60.6% branch, 76.4% method.
2026-08-01 21:34:59 +07:00
kevin d67d30c58a CB-508: let an opencode profile pin its own OpenAI-compatible endpoint
CI / build (push) Successful in 1m57s
Points an opencode worker at a local vLLM (or llama.cpp / LM Studio / TGI)
instead of opencode's own gateway. opencode has no ANTHROPIC_BASE_URL seam, so
this could not be a config-only change: setting baseUrl on a kind: opencode
profile now makes the launcher emit a custom `provider` block into the
generated opencode.json, using @ai-sdk/openai-compatible.

The provider id comes from the provider half of the model: selector, so one
field drives both the generated declaration and the -m flag and the two cannot
drift apart. A bare model name with a baseUrl set is rejected at spawn with a
message saying how to fix it — silently falling back to the default gateway
would leave a worker talking to the wrong LLM while looking perfectly healthy.

A bare host:port gets /v1 appended (where these servers mount the API); a URL
that already carries a path is used verbatim. tokenEnv, when set, becomes the
provider apiKey; local servers generally ignore it but the AI SDK requires a
non-empty value, so a placeholder is used otherwise.

Two supporting changes:
- writeConfig previously ran only when a bridge MCP url was set. A pinned
  endpoint needs the config file too, so it now runs when either applies, and
  the mcp/instructions half is emitted conditionally.
- The config is now built with Jackson instead of string concatenation. The
  provider block is nested and interpolates operator-supplied values (URL,
  model id, api key), so escaping has to be real rather than a hand-rolled
  two-character replace.

No guard entry is required even with baseUrl set. SubscriptionGuard exists to
stop a worker borrowing the primary's Anthropic subscription, and an opencode
process has no Anthropic credential path at all — the asymmetry with the Claude
adapter reusing the same field is deliberate and documented at the call site.

Also fixes a brittle assertion in the existing MCP-mount test, which matched
the substring "\"type\": \"remote\"" and broke on Jackson's spacing. It now
parses the generated JSON and asserts on structure; whitespace is the
formatter's business, not the contract's.

318 tests (was 311): 5 new covering provider generation, /v1 normalisation,
path-preserving URLs, the missing-prefix rejection, MCP+provider coexistence,
and that no baseUrl still means no provider block.

Verified live end to end: daemon restarted on this build, worker spawned on the
opencode-local profile, generated config carries baseURL
http://127.0.0.1:8000/v1, and a blocking bridge_send returned
{"reply":"LOCAL-OK","replySource":"reply"} — a structured reply, not the
completion fallback. The worker pane reports
"Build · deepseek-v4-flash local-vllm (bridged)", confirming traffic reached
the local server rather than silently falling back.
2026-08-01 20:03:00 +07:00
kevin d75ee1cca5 CB-507: regression tests for worktree cwd resolution
Two cases in WorktreeSessionManagerTest, covering the gap that let the NPE ship
(313 tests, was 311).

1. worktreeAcquireWithNoRequestedOrCallerCwdStillResolvesANonNullRepoRoot —
   the null/null case a plain REST spawn produces.
2. worktreeAcquireHonoursTheProfileConfiguredCwd — the quieter second bug on
   the same line, where a pinned per-profile cwd: was ignored entirely.

Both assert on the cwd RECORDED by FakeWorktrees rather than expecting a throw.
That is deliberate: FakeWorktrees.repoRoot only records its argument and returns
a canned root, so a null passes through the fake harmlessly while the real
GitWorktrees runs `git -C null` and NPEs. The fake being more permissive than
the real seam is exactly why 311 tests stayed green over a broken feature —
asserting "an exception was raised" would be untestable here and would give
false confidence.

Verified as genuine regressions, not tautologies: with the pre-CB-507
expression restored both fail, with the messages they were written to give
(expected: not <null>, and expected </pinned/dir> but was <null>). Restored
after.

Drafted by an opencode-free worker over the bridge in an isolated worktree
(branch worker/cb-507-regression-test-11591f-4). Its test 1 was correct as
written. Test 2 was wrong and went red: it passed "/pinned/dir" as the 4th
constructor argument, which is configDir, not cwd (the 11th, after mcpUrl), so
cwd stayed null and the chain fell through to the daemon cwd. Corrected on
integration, along with removing two unused locals and adding the rationale
comments.
2026-08-01 19:50:31 +07:00
kevin 6804676a96 CB-507: fix NPE on a worktree spawn with no cwd (HTTP 500 over REST)
POST /workers?worktree=true returned HTTP 500 with a NullPointerException out
of ProcessBuilder.start(): acquireWithWorktree resolved the repo root from
firstNonBlank(requestedCwd, callerCwd), and a plain REST spawn supplies
neither (BridgedApp hardcodes callerCwd=null, "no MCP caller over REST"). Both
null yielded null, putting `git -C null rev-parse --show-toplevel` on the
command line.

Now resolved through launcher.effectiveCwd, the CB-112 chain used everywhere
else (requested -> profile cwd -> caller -> daemon cwd -> "."), which is
documented never to return null. The non-worktree path in this same class
already went through it; only the worktree branch was missed.

Also fixes a second latent bug in the same line: firstNonBlank never consulted
the profile's configured cwd:, so a worktree spawn silently ignored a pinned
per-profile working directory. effectiveCwd honours it.

Removes firstNonBlank, now dead (this was its only call site) — javac ignores
an unused private method but IDE inspections flag it, and CLAUDE.md requires a
clean bill.

Why 311 tests missed it: the null/null case only arises over REST, and
WorktreeSessionManagerTest always passes an explicit cwd. Over MCP callerCwd is
populated from the caller PID, so the feature worked there. This is the third
REST-vs-MCP divergence found this month, after CB-505's path-trusted session id.

The one-line change was implemented by an opencode-free worker over the bridge
in an isolated worktree (branch worker/cb-507-worktree-cwd-npe-3e9c3b-3); the
dead-helper cleanup and the explanatory comment were added on integration.
A regression test is still outstanding and is being delegated separately.
2026-08-01 19:43:08 +07:00
kevin 3e5d742ac7 CB-506: keep the test suite out of the production audit log
main/resources/logback.xml routes the `audit` logger to a RollingFileAppender
at logs/audit.log — the CB-505 security trail. AuditLogTest and
BridgedAppAuthTest exercise that same logger, so every `mvn test` appended
fabricated records to the production file.

They are byte-identical to genuine ones: runs of denied/forbidden
SPAWN/STOP/SEND from worker:term_a, which read exactly like an intrusion
attempt. logs/audit.2026-07-29.0.log is 38 fabricated records out of 76 — half
that day's security log is test fixtures, and nothing distinguishes them.

Fix is one new file, src/test/resources/logback-test.xml: logback prefers it on
the test classpath, so tests get a console-only config with no file appender
and main/resources/logback.xml is untouched. The `audit` logger stays ENABLED
(INFO, additivity=false) because AuditLogTest attaches its own ListAppender and
asserts on emitted records — setting it OFF would have silently gutted those
assertions.

Verified: 311 tests green, and logs/audit.log line count is identical before
and after a full `mvn clean install` (zero new records).

Implemented by an opencode-free worker over the bridge in an isolated worktree
(branch worker/cb-506-audit-test-isolation-e4aa9c-2); it correctly reported it
could not run mvn rather than fabricating a result, so the build gate and the
before/after audit-count check were run primary-side. Header comment added on
integration.
2026-08-01 19:20:29 +07:00
kevin 6da2a71050 CI: drop upload-artifact — unsupported on this Gitea instance
CI / build (push) Successful in 1m44s
Run 2 built clean (311 tests, BUILD SUCCESS, Maven 3.6.3 on Java 25.0.4 —
JAVA_HOME from setup-java correctly beat the JRE apt pulled in) but the job
still went red on the artifact step:

  GHESNotSupportedError: @actions/artifact v2.0.0+, upload-artifact@v4+ and
  download-artifact@v4+ are not currently supported on GHES.

Gitea Actions presents as GHES, so v4 artifact upload cannot work here. The
artifact was unretrievable regardless, so replace it with a failure-only step
that cats the failing surefire .txt reports into the job log, where they are
readable. Guarded with 'exit 0' so the dump itself can never mask the real
failure.
2026-07-29 23:26:21 +07:00
kevin e32ac39faf CI: install Maven — setup-java provides the JDK only
CI / build (push) Failing after 2m1s
First CI run failed at 'Build and test' with exit code 127 (command not
found): actions/setup-java@v4 provisions a JDK but not Maven, and the runner
image has no mvn on PATH. The sibling lms/alms workflow apt-installs both;
this workflow switched to setup-java for JDK 25 (the image's default-jdk is
too old for maven.compiler.release=25) and dropped the maven install with it.

Adds an explicit Maven install with Apache Maven 3.9.16 (2bdd9fddda4b155ebf8000e807eb73fd829a51d5)
Maven home: /Users/appbuilder/Tool/apache-maven-3.9.16
Java version: 25.0.2, vendor: Oracle Corporation, runtime: /Users/appbuilder/Tool/jdk-25.0.2.jdk/Contents/Home
Default locale: en_VN, platform encoding: UTF-8
OS name: "mac os x", version: "26.5.1", arch: "aarch64", family: "mac" so the log proves which JDK
it resolved — JAVA_HOME from setup-java must win over the JRE apt drags in.
2026-07-29 23:21:55 +07:00
kevin daa243d37a example config: opencode profile uses the verified zero-credential free tier
CI / build (push) Failing after 2m21s
The commented CB-402 block suggested google/gemini-2.5-pro, which needs
credentials. Replaced with the opencode/*-free gateway models proven during the
dogfood to work with no auth at all, and noted that the free model names change
so `opencode models` is the source of truth.
2026-07-29 22:49:58 +07:00
kevin 2773ab600d CB-402: live dogfood complete — Stage B verified against opencode 1.18.5
Closes the one known-unverified item before cross-host. CB-402 merged in
ded226a with increment 5 (the §5 live checklist) deferred; it has now run.

Provider question (§7 Q1) resolved with no credentials needed: opencode's own
gateway serves free-tier models. `opencode auth list` reports 0 credentials,
yet `opencode run -m opencode/north-mini-code-free` answers. Distinct from the
primary's subscription by construction, and needs no guard entry — opencode
carries no ANTHROPIC_BASE_URL, so SubscriptionGuard never applies to it.

The schema-drift risk was the real one and it did not bite. The adapter was
designed against opencode 1.1.31; installed is 1.18.5. The generated config
still validates unchanged (type:"remote" + instructions:[path]), and
`OPENCODE_CONFIG=… opencode mcp list` reports the bridge connected. Pinned as
a verified fact for 1.18.5.

Full lifecycle through REST: spawn (201, kind-routed to OpenCodeLauncher) ->
CB-306 gate passed ~0.6s -> ready -> send -> {"replySource":"reply"} (a
STRUCTURED bridge_reply, not the CB-115 completion fallback) -> delete (204,
tolerant teardown).

Unplanned cross-validation with CB-501: the audit trail recorded the reply as
role=WORKER actor=worker:term_657c… — connection-based identity classified an
opencode process as a worker with no opencode-specific handling. The identity
model is peer-kind-agnostic, which is what CB-308 needs when the roster
stretches across hosts.

Stage 5 verified live on the same run: /workers (CB-304) answers where the
13-day-old daemon 404'd, /metrics counted the delegation
(sends_total{outcome=replied} 1, replies_total{path=rendezvous} 1,
inbox_depth 0), and the audit log captured SPAWN/SEND/REPLY with correct roles.

Adds the opencode-free dogfood profile to the local bridged.yaml (gitignored;
recorded here for reproducibility) and docs/CB-402 §8 as-built.
2026-07-29 22:49:15 +07:00
kevin 19cdf8dc9f CB-505 fix: audit lines were not valid JSON
The first cut spliced the timestamp on via a logback pattern:

    {"ts":"%d{...}",%replace(%msg){'^\{',''}%n

Logback's variable substitution chokes on the literal braces
("All tokens consumed but was expecting }"), so the encoder failed to
configure. Caught by running the real jar and noticing logback had dumped its
internal status — which it only does when something failed to parse. The build
was green throughout: nothing asserted the audit trail was machine-readable.

AuditLog now emits the complete object including its own ISO-8601 "ts", and the
appender pattern is a bare %msg. Adds AuditLogTest, which parses each emitted
line with Jackson (so a malformed record fails the build) and pins that hostile
ids cannot escape their field to forge a second record.

311 tests green; logback now configures with zero internal errors.
2026-07-29 22:32:18 +07:00
kevin 9daf1ec5ba CB-5xx: Stage 5 hardening — auth, authz+audit, metrics, CI, supervision
Closes out single-host before the cross-host work. Sequenced BEFORE CB-308
deliberately: federation's own gating concern is the trust model, and it
inherits whatever identity shape lands here.

The finding this stage is built around: bridged had exactly ONE security
control, the loopback bind. ConnectionIdentity resolves a worker from its
connection (unforgeable), but every caller that was not a recognised worker
pane fell through to being treated as the PRIMARY -- the most privileged role
on the bus. Latent today; load-bearing the moment a bind widens.

CB-501 auth:
- Role/Principal/CallerResolver: connection identity first, bearer token
  second, ANONYMOUS third. Inverts the old default so absence of identity
  means nothing, not everything.
- Worker identity is never token-gated, so enabling auth cannot lock the
  fleet out of bridge_reply.
- Constant-time token compare (MessageDigest.isEqual).
- validateAuthExposure(): a non-loopback bind under loopback-trust now
  REFUSES TO START. Makes the dangerous config unrepresentable rather than
  merely documented.
- TLS terminates at a reverse proxy by design (D3), not in the JVM.

CB-505 authz + audit, enforced on BOTH entry paths:
- The docs describe MCP as "a thin adapter over the REST core"; at code level
  it is not. BridgeMcp calls MessageService directly, and /mcp is a raw
  servlet on Jetty's context handler that never traverses Javalin's before
  filter. Enforcing only at REST would have left /mcp open.
- Load-bearing rule is own-session-only: a worker may reply/ask only as
  itself. Structurally true over MCP already; over REST the session id in the
  URL path had simply been trusted.
- Audit: JSON lines to a dedicated appender, additivity=false. Never records
  message content -- this bus carries source and prompts.

CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer
instead of the specced Micrometer, because this pom already hand-pins
jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against
skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated
dependency CVE gate could not be run (no JetBrains MCP server connected).
Instrumented at MessageService, the single funnel both surfaces share.

CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner.
Needs no contract-exclusion flag -- the pom's default-excludes profile
already sets excludedGroups=contract, so plain `mvn clean install` IS the
mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older).

CB-504 supervision: launchd agent (the real target -- this host is macOS,
there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds.
Ordering directives are advisory, so the actual fix is that startup now waits
up to 30s for the herdr socket and then serves degraded, instead of crashing
into a restart loop on a boot-order race.

Also fixes drift found while surveying:
- bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config
  binds via plain Jackson with ignoreUnknown, so uncommenting it would have
  been silently dropped and the default kept. Now camelCase, with a test that
  loads the shipped example and one that pins every documented knob's
  spelling -- no test had ever loaded that file.
- Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay,
  gitTokenEnv, gitHostEnv, configDir, primary:).
- README "Next" listed bridge_ask and session lifecycle as upcoming; both
  shipped long ago.
- docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre-
  implementation" for work already merged.

307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS.
Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check
could not be run -- no JetBrains/intellij-index MCP server is connected this
session. mvn clean install is the only gate that ran.
2026-07-29 22:29:26 +07:00
Dai Ha c9f0ca9359 wiki: bump submodule pointer to b646be1 (chapter 10 — Cross-Host Messaging & Broker Topology) 2026-07-28 16:33:51 +02:00
Dai Ha 84c8a2d2f0 CB-500 §11: resolve distributed-sandbox topology (gateway-per-host × local sandboxes)
Clarifies the open fork from §4/§9: Development A (sandbox launcher) and CB-308
(per-host federation) COMPOSE — each host runs a bridged gateway whose launcher
spawns agents into that host's LOCAL sandboxes; the broker moves messages +
presence, never keystrokes.

The forcing fact: delivery is herdr keystroke-injection into a locally-owned PTY,
so "a sandboxed agent on another host" ≡ "a sandbox spawned by that host's
gateway" (a remote container with no local herdr can't be injected into). Rules
out a central daemon reaching remote PTYs.

Adds Figure 11 (composed topology) + Figure 12 (remote-delegation sequence: the
local?inject:publish fork with a sandboxed far side — both injection points stay
local, only the middle hop crosses the broker), the two forced reachability
changes (host-routable mcpUrl; PTY in the local gateway's herdr), and a
maps-to-existing-seams table (CB-308 gateway × Dev-A launcher, CB-117 reap,
CB-303 container lifecycle, CB-308 #5 trust). No new pillars. Both diagrams
mmdc-validated; §4/§9 updated to point at §11.
2026-07-28 16:26:10 +02:00
Dai Ha ded226abfe Merge CB-402: opencode second peer adapter (Stage B of the Peer Launcher SPI)
Proves the CB-401 PeerLauncher SPI is genuinely provider-neutral by landing a
second adapter — opencode — that shares NONE of Claude Code's private launch
seams (no ANTHROPIC_BASE_URL, no SubscriptionGuard). Merged as one unit:

- Incr 2 (6e37722): kind: discriminator on BridgedConfig.Worker (claude-code |
  opencode), argv defaults to the kind binary, isClaudeCode()/isOpenCode().
- Incr 3 (b034f10): OpenCodeLauncher extends HerdrPeerLauncher — diverges only
  in buildLaunch (file-based MCP mount via OPENCODE_CONFIG + reply charter under
  instructions[], model via -m provider/model, opencode- name/reap prefix).
- Incr 4 (11f8709): CompositePeerLauncher routes the fleet by kind (spawn/cwd/
  parity by profile, stop by pane-id owner, list dedup, reap/caps/profiles
  union); Bridged.main partitions profiles by kind → composite. BridgeMcp +
  BridgedApp migrated onto the PeerLauncher SPI (the two (ClaudeCodeLauncher)
  casts removed; dead rendezvous params dropped).

Incr 1 (extract HerdrPeerLauncher base) already on main at ffce30a.
266 tests green, mvn BUILD SUCCESS. Live opencode-on-Gemini dogfood tracked
separately (needs a running-daemon restart + a resolved Gemini provider).
2026-07-28 16:17:34 +02:00
Dai Ha 9b8d55bc18 CB-500: design note — multi-tier coordination (Stage 6)
Scopes the lead's next direction on top of the peer-launcher arc, as a
proposal (ticket split deferred):
- A · sandboxed, role-specific workers — a sandbox is a *placement*, so it
  slots into the CB-401 SPI as a new kind: sandbox adapter exactly the way
  CB-402's opencode slotted in as a new provider (SPI proven placement-neutral,
  not just provider-neutral). Ownership line held: bridge launches INTO a
  peer-owned image, never provisions the IDE/dev-tools inside it.
- B · main-agent pairs (Opus + cloud) — both mains are MCP clients so neither
  can be called into; each needs a pull inbox → depends on CB-308's per-agent
  channels. PrimaryRegistry single-slot → multi-slot.
- C · orchestrator tier — SessionManager recursed one tier up
  (orchestrator:mains :: main:workers) + context scoping; re-roots the human
  from a live primary to the orchestrator.

10 mermaid diagrams (component + sequence per development, ownership guardrail,
staging graph), all mmdc-validated and theme-safe. §7 pins the bus-vs-env-manager
boundary as an acceptance criterion on A; §8 stages A → CB-308 substrate → B → C.
2026-07-28 15:20:16 +02:00
Dai Ha 11f8709286 CB-402 Increment 4: CompositePeerLauncher — route the fleet by kind
Introduce the router the core holds when more than one adapter is
configured: one HerdrPeerLauncher per peer kind, dispatched by profile
(spawn/effectiveCwd/parityOverlay), by pane id (stop, via a spawn-time
owner map), and fanned out + combined for the fleet-wide queries
(list dedup by pane id, reap/caps union, profiles union). The ctor
rejects an empty adapter list and a profile two adapters both claim.

Wire it in Bridged.main: partition workerProfiles() by kind (claude-code
is the always-present default adapter; opencode is added when any profile
opts in) and front both with the composite. This lets BridgeMcp and
BridgedApp finally take the PeerLauncher SPI instead of a concrete
ClaudeCodeLauncher — the two (ClaudeCodeLauncher) casts in Bridged are
gone. list() elements are cast to herdr Agent at the point of the
herdr-specific roster view, where that assumption actually lives.

Add BridgedConfig.Worker.isClaudeCode()/isOpenCode() kind predicates
(the wiring uses isOpenCode; both are unit-tested). Drop the long-dead
'rendezvous' constructor param threaded into BridgeMcp and BridgedApp.

10 CompositePeerLauncherTest cases over two real adapters on one
FakeHerdr: profile routing (observed via the started agent's
claude-/opencode- name prefix), default resolution, unknown-profile
and duplicate-profile rejection, caps union, list dedup, reap sum,
and stop teardown. 266 tests green.
2026-07-22 05:43:28 +02:00
Dai Ha b034f105c0 CB-402 Increment 3: OpenCodeLauncher — the SPI-proving second adapter
A HerdrPeerLauncher subclass for opencode, a provider-agnostic terminal
coding agent. It reuses every line of shared base transport (tab/pane
placement, CB-306 readiness gate, unique naming + CB-117 reap, teardown,
listing, cwd) and diverges only in buildLaunch:

  - No subscription boundary: no ANTHROPIC_BASE_URL, no SubscriptionGuard
    (the guard is a Claude-private concern, not part of the SPI).
  - File-based MCP mount + instructions: writes an ephemeral opencode.json
    declaring the bridge as a remote MCP server + a reply-charter file under
    instructions, pointed at via OPENCODE_CONFIG (opencode has no inline
    --mcp-config / --append-system-prompt).
  - Model selected with -m provider/model, not an env var.
  - 'opencode' name prefix so reap matches opencode-* panes only.

configRoot is injectable so tests inspect the generated config/charter under
a @TempDir. 10 tests cover config content, model flag, git-token grant,
capabilities, reap predicate, both production ctors, and the readiness gate
(throws PeerUnreachable on timeout, reaps only the worker pane).
2026-07-22 05:30:46 +02:00
Dai Ha 6e37722383 CB-402 Increment 2: kind: discriminator on Worker profiles
Add a `kind` field to BridgedConfig.Worker — "claude-code" (default) or
"opencode" — the discriminator the CompositePeerLauncher will route spawn/reap
by so each adapter drives only its own peer kind. Normalised to lower-case;
blank/absent ⇒ claude-code, so every existing config and call site is
unchanged. argv now defaults to the kind's own binary (claude vs opencode)
rather than always `claude`, so an opencode profile never inherits the Claude
command.

kind is appended at the record tail; a new 14-arg back-compat constructor
(git fields, no kind) keeps the CB-302 call sites working, and the existing
12-arg constructor is untouched. Also drop the never-used Primary(String)
legacy constructor to keep the file warning-clean.

example.yaml documents the key and carries a commented opencode-gemini
profile. Tests cover default/normalisation/argv-defaulting. 245 tests green,
config files 0 IDE problems.
2026-07-22 05:20:04 +02:00
Dai Ha ffce30afa2 CB-402 Increment 1: extract HerdrPeerLauncher abstract base
Behaviour-preserving refactor ahead of the second-adapter work. All herdr
transport shared by any peer kind — tab/pane placement, the CB-306
spawn-readiness gate, unique naming + CB-117 orphan reap, teardown, listing,
cwd resolution, and the peer-neutral git-forge grant — moves into a new
abstract HerdrPeerLauncher (Template-Method base). ClaudeCodeLauncher becomes
a final subclass supplying only the two Claude-specific seams: the `claude`
name prefix and buildLaunch(), which encodes the subscription boundary
(ANTHROPIC_BASE_URL + SubscriptionGuard assert, inline --mcp-config and
--append-system-prompt reply charter).

The base owns the injectable clock + sleeper for the readiness gate; the poll
interval is baked into the sleeper, so the vestigial spawnReadyPollMs field is
dropped from the base and from the full testability constructor (the explicit
sleeper already encodes it). The 6-arg and 8-arg production constructors keep
their signatures; three full-ctor test sites drop the now-unused poll argument.

No behaviour change: 242 tests green, both refactored files 0 IDE problems.
2026-07-22 05:15:21 +02:00
Dai Ha e724a59f2d CB-402: design note — opencode second peer adapter (Stage B)
Design-note-first for gitea #7. Extract HerdrPeerLauncher abstract base
(Template Method) + kind: discriminator + OpenCodeLauncher + routing
CompositePeerLauncher; finish the Stage-A caster migration off
ClaudeCodeLauncher. opencode proves the SPI for a non-Claude peer
(no subscription guard, OPENCODE_CONFIG MCP mount, config-based charter).
2026-07-22 04:59:58 +02:00
Dai Ha 4bf855d225 wiki: bump submodule pointer to 4d1548a (CB-307 as-built)
Advance the wiki submodule pointer to include the CB-307 Stage 2/3 as-built
docs plus the intervening CB-401/delegation-directive commits. Standalone
pointer bump — not bundled into a feature commit.
2026-07-19 18:10:01 +02:00
Dai Ha d4c9704007 CB-307: primary-gate cleanup of push-loop (drop dead clock, IDE 0/0)
Primary verification pass over the delegated push-loop delivery:
- Remove the unused LongSupplier clock threaded into ReplyPushLoop
  (timing is the scheduler's; the field was never read) from the
  component, Bridged wiring, and both test call sites.
- Collapse the single-statement WAIT_BUSY switch arm (redundant block).
- Drop now-dead test scaffolding: the always-"idle" recordingClient
  param and unused AgentStatus/AtomicReference imports.

IDE diagnostics 0/0 on all changed files; mvn clean install green
(242 tests, 0 failures).
2026-07-19 11:03:34 +02:00
Dai Ha 7c252b5f5f CB-307 Increment 3: bridge_ack tool — per-msgId ack refinement
- MessageService.ackReply(target, msgId) delegates to inbox.ack
- BridgeMcp registers bridge_ack tool with target/msgId args
- Tests: valid/invalid args, ack surface via BridgeMcp
2026-07-19 10:21:06 +02:00
Dai Ha f756933879 CB-307 Increment 2: ReplyPushLoop — status-gated push loop (mechanism b)
- ReplyPushLoop: dedicated scheduled loop for nudging the primary
- Package-private decide() method for pure decision logic (unit-testable)
- Status-gated injection via AgentControl.status().injectable()
- Bounded reminders (cap + backoff), idempotent per target
- Wire into MessageService.reply after inbox.publish on no-waiter branch
- Wire into Bridged.main (constructor + shutdown hook)
- Primary config record updated with pushReminders/pushBackoffMs knobs
- Tests: decide() matrix, nudge injection, idempotency, cap enforcement
2026-07-19 10:16:24 +02:00
Dai Ha a1aecbf4fc CB-307 Increment 1: PrimaryRegistry + config + wiring
- PrimaryRegistry: thread-safe single-slot registry with pin support
- Primary config record (last positional, like Broker)
- Wire capture in BridgeMcp (bridge_send and bridge_spawn handlers)
- Construct PrimaryRegistry in Bridged.main
- Tests: PrimaryRegistryTest + BridgedConfigTest primary config cases
2026-07-19 09:59:02 +02:00
Dai Ha 131e7b1ccd CB-307: lock push-loop injection to dedicated status-gated loop (mechanism b)
Chose a small dedicated scheduled loop over AgentControl.send guarded by an
injectable status check, instead of reusing the worker Injector (which couples
to WorkerPresence/StatusPoller). Isolated + unit-testable via injected clock.
Records live ground truth: this primary resolves to term_656c8cc03e1f0b1 (w2:pY)
— confirms the primary runs in a herdr pane so the push path is exercisable.
2026-07-19 09:44:21 +02:00
Dai Ha d0ac6c435f CB-307: design note for active push-to-primary + bounded reminder loop
The reliability layer over the durable inbox (Stage 2, 2bc5f3a): push a nudge
into the primary's own herdr pane the moment a no-waiter reply lands, remind on
a bounded backoff until the primary drains (ack=drain), degrade to pull when the
primary pane isn't resolvable. Grounds the seam: the primary terminal_id is
already derivable via ConnectionIdentity/PaneLocator, just discarded today.

Refs gitea #5.
2026-07-19 09:41:22 +02:00
Dai Ha 2bc5f3a057 CB-307 Stage 2: AmqpReplyInbox — durable, cross-restart reply delivery
Behind the existing ReplyInbox port, add an AMQP-backed adapter selected by a
`broker:` block in config (absent → the in-memory soft-state inbox; present →
AMQP). Mapping is consume-and-hold with deferred manual ack: each target owns a
durable queue `agent.<target>.inbox`; a manual-ack consumer pulls persistent
messages into an in-memory held map (dedup by msgId) but does not ack; peek
returns the snapshot; ack acks the broker delivery-tag and drops it. A crash
before caller-ack leaves messages unacked, so the broker redelivers on
reconnect — genuine durability with the port contract preserved. bridged still
owns no persistence; the broker does.

- msg/AmqpReplyInbox: the adapter (single synchronized channel; recovery
  listener clears held on reconnect so fresh delivery-tags repopulate).
- config/BridgedConfig: nullable Broker(uri) record; isConfigured() gates it.
- Bridged.main: select adapter; close the AMQP connection in the ordered
  shutdown hook (no-op for the in-memory inbox).
- deps: com.rabbitmq:amqp-client (main); testcontainers rabbitmq/junit-jupiter
  (test). Pinned commons-compress 1.27.1 + commons-lang3 3.18.0 to clear the
  test-scope CVEs those pull. Production default LavinMQ; RabbitMQ URI-swap.
- tests: BridgedConfigTest broker-selection cases (hermetic); AmqpReplyInbox
  contract test (@Tag("contract"), Testcontainers RabbitMQ) proving
  publish/peek/ack, msgId dedup, and cross-restart redelivery. Excluded from
  the default build so `mvn clean install` stays hermetic (210 green).
2026-07-19 07:30:29 +02:00
Dai Ha ba6b4a5da9 CB-307 Stage 1: reply-inbox port + in-memory adapter — hold stranded worker replies instead of dropping them
Problem: the reverse (worker->primary) path was Rendezvous, a map of LIVE blocking
waiters only. A bridge_reply arriving with no open send hit Rendezvous.complete()
-> no waiter -> returned false -> the reply was silently DISCARDED (worker saw an
error / REST 409). No message-id/dedup/ack existed anywhere.

Stage 1 (no broker, soft-state) behind one port:
- ReplyInbox port + InboxMessage record; InMemoryReplyInbox adapter (per-target
  FIFO via LinkedHashMap, dedup by msgId, thread-safe). Soft-state, not persistence.
- MessageService.reply(session, content): resolve an open send, else publish to the
  inbox with a minted UUID (was a silent drop). drainReplies(target) = peek + ack.
- BridgeMcp.reply / BridgedApp.replyMessage repointed off bare Rendezvous.resolve
  onto messages.reply -> no-waiter is now SUCCESS (queued), not error / 409.
- Drain surface: bridge_poll gains optional target; REST GET /sessions/{id}/replies.
- Rendezvous left untouched. QUESTION path (bridge_ask) NOT queued (interactive,
  keeps NO_WAITER); completion/failure fallbacks NOT queued (captured-waiter).
- Bridged.main wires new InMemoryReplyInbox(); no broker: config yet (Stage 2 = AMQP).

Tests: +19 (188 -> 207), 0 failures/0 errors. New InMemoryReplyInboxTest (12) +
MessageService/BridgeMcp/BridgedApp coverage incl. guards proving a QUESTION and a
completion fallback are never queued.

Implemented via delegation to an off-sub gx10 worker in a pre-trusted worktree;
primary-verified (mvn clean install green, 207 tests) and committed by the primary
because the worker's completion replies were lost to the very bug this fixes.

Refs CB-307 (gitea #5), Stage 1 of 2.
2026-07-18 21:12:51 +02:00
Dai Ha da5a987df0 CB-307/CB-308 design notes: reliable-delivery Stage-1 delegation spec + multi-host federation proposal
- docs/CB-307-Reliable-Delivery.md: ReplyInbox port + in-memory adapter spec
  (Stage 1, no broker); publish at the Rendezvous no-waiter drop seam, drain by target.
- docs/CB-308-Multi-Host-Federation.md: per-host gateway + per-agent broker channels
  + federated roster proposal (gitea #6), wiki-ready with theme-safe mermaid.
2026-07-18 20:24:16 +02:00
Dai Ha 7dd6c46156 CB-306: spawn-readiness gate — ClaudeCodeLauncher blocks until the worker is injectable or throws PeerUnreachableException
The launcher now polls AgentControl.status(paneId) after starting the pane.
It returns the handle only once the worker reports an injectable state
(IDLE/BLOCKED/DONE). If the timeout elapses while still UNKNOWN, the
pane is self-reaped and a PeerUnreachableException is thrown — no orphan
left behind. The gate is disabled when spawnReadyTimeoutMs == 0 (legacy
non-blocking spawn, the default for the 6-arg constructor).

Key changes:
- PeerUnreachableException (new) in dev.ltms.bridged.peer
- BridgedConfig: spawnReadyTimeoutMs (default 20000), spawnReadyPollMs (default 300)
- ClaudeCodeLauncher: 3 constructor overloads:
  (a) 6-arg backward-compat: gate disabled (timeout=0)
  (b) 8-arg production: gate with config knobs + real clock/sleep
  (c) 10-arg testability: full seam (LongSupplier clock + Runnable sleeper)
- waitUntilInjectableOrThrow() loop in spawn(SpawnRequest)
- sleepUninterruptibly() helper for the production sleeper
- BridgeMcp.spawn + BridgedApp.spawnWorker catch PeerUnreachableException
  → clean tool error / 502 response (not an uncaught 500)
- SessionManager.acquire inherently registers nothing on throw (both
  worktree and non-worktree paths) — confirmed by new test

Tests:
- ClaudeCodeLauncherTest: 4 new tests
  - unknown→idle: returns handle, no pane.close
  - always-unknown: throws PeerUnreachableException, pane closed,
    clock advanced past timeout
  - timeout=0 (6-arg ctor): no agent.get calls, returns handle
  - timeout=0 (10-arg ctor): no orphan pane close
- SessionManagerTest: 1 new test
  - acquire → PeerUnreachableException: roster remains empty
Total: 188 tests, all pass (no existing test changed semantics)
2026-07-18 16:17:43 +02:00
Dai Ha 3a5cdc5108 CB-306: spawn-readiness gate design note (delegation spec) 2026-07-18 15:41:43 +02:00
67 changed files with 8865 additions and 632 deletions
+37 -54
View File
@@ -1,24 +1,21 @@
---
name: implementer
description: Implementer-role playbook for a bridged worker — you are in an isolated git worktree on a dedicated branch; implement the assigned task, commit, push, open your own PR to main, and hand off the PR URL via bridge_reply. You never merge. Load this when you have been delegated an implementation task over bridged.
description: Implementer-role procedure for a bridged worker — verify your worktree, implement the scope, commit, push, open your own PR, and hand off the PR URL. Load this when the lead delegates you an implementation task over bridged.
---
# Implementer worker
# Implementer worker — procedure
You are an **implementer** in the claude-bridge fleet. The lead delegated you one scoped task,
and you are running in an **isolated git worktree on your own branch** — a full peer of the
primary (same `CLAUDE.md`, skills, memory, MCP), differing only in the model behind you and the
branch you sit on. Your job for this turn: **implement the task, then hand off a PR the lead can
review and merge.** You do the work; the lead (or human) is the merge gate — you never merge.
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
never merge, never commit `.mcp.json` or `wiki/`) is in **`CLAUDE.md` → Bridge communication →
Worker** and already applies. This skill is only the *implement-and-hand-off procedure*.
Delivery mechanics (how the task reached you, how your reply resolves the lead's blocked send)
are in [`docs/MCP-Contract.md`](../../../docs/MCP-Contract.md); the worktree/PR model is in
[`docs/Worker-Git-Workflow.md`](../../../docs/Worker-Git-Workflow.md). You only need the steps
below.
You run in an **isolated git worktree on your own branch** — a full peer of the primary (same
repo, `CLAUDE.md`, skills, MCP), differing in the model behind you and the branch you sit on.
The worktree model is documented in [`docs/Worker-Git-Workflow.md`](../../../docs/Worker-Git-Workflow.md).
## 1. Confirm where you are — a worktree on a dedicated branch
## 1. Confirm where you are
Before touching anything, verify your ground truth:
Before touching anything:
```bash
git rev-parse --show-toplevel # your worktree root — NOT the primary's main tree
@@ -26,46 +23,40 @@ git branch --show-current # your dedicated branch: worker/<ticket>-<nonce>
git status # should be clean at the start
```
Do **all** work here, on this branch. **Never** switch to `main`, never `git checkout main`,
never rebase onto or push to `main` directly. The branch is your isolation — respect it.
Do **all** work here, on this branch. Never `git checkout main`, never rebase onto or push to
`main`. The branch is your isolation — respect it.
## 2. Implement the task
## 2. Implement
- Implement exactly the scope the lead named. Keep changes focused; if you notice something out
of scope, note it in your reply rather than expanding the diff.
- Match the surrounding code's style, naming, and idioms. Follow project `CLAUDE.md`.
- **You cannot run the IDE MCP tools** (intellij-index / jetbrains are the primary's, not yours).
So **never claim a file is "IDE-clean" or "diagnostics-clean"** — you cannot verify that. State
only what you actually ran (e.g. `mvn`, a test) and its real output. A fabricated clean claim is
worse than an honest "I could not verify inspections here."
- Run whatever build/test you can and **report the true result** — including failures.
- Implement exactly the scope the lead named. Keep the diff focused; note anything out of scope
in your reply instead of widening it.
- Match the surrounding code's style, naming, and idioms.
- Run whatever build/test you can — `mvn clean install` from the module root. Read its **full**
output; a piped `mvn ... | tail` hides failures.
## 3. Commit — focused, and never the excluded files
## 3. Commit
```bash
git add <the files you changed>
git add <the files you changed> # explicitly — never `git add -A` / `git add .`
git commit -m "<ticket>: <clear one-line summary>"
```
**Excluded from every commit, always:** `.mcp.json` (the primary's local, session-modified copy —
present only for parity) and `wiki/` (a separate submodule). Stage files explicitly; do **not**
`git add -A` / `git add .` blindly, or you risk staging them. If `.mcp.json` shows as modified,
leave it — it is flagged `--skip-worktree` and is not yours to commit.
`.mcp.json` will show as modified. Leave it — it is `--skip-worktree` and not yours to commit.
## 4. Push your branch
## 4. Push
```bash
git push -u origin HEAD
```
Push is over SSH as the same user — no extra credential needed. Push the branch as-is; do not
force-push over anything you did not create.
Push is over SSH as the same user — no extra credential needed. Never force-push over anything
you did not create.
## 5. Open your own PR to `main`
Open the PR via the gitea REST API. The daemon injected a **repo-scoped token** (`GITEA_TOKEN`)
and the forge host (`GITEA_HOST`) into your env for exactly this — the token can create a PR but
**cannot merge** (that stays the lead/human gate).
Via the gitea REST API. The daemon injected a **repo-scoped token** (`GITEA_TOKEN`) and the forge
host (`GITEA_HOST`) into your env for exactly this — the token can create a PR but **cannot
merge**.
```bash
API="${GITEA_HOST%/}/api/v1/repos/lms/claude-bridge/pulls"
@@ -81,16 +72,14 @@ JSON
)"
```
The response JSON includes `"html_url"` — that is your PR URL. If the call fails (non-2xx), read
the error body, fix the cause if it is yours (e.g. branch not pushed yet), and report the failure
honestly in your reply rather than inventing a URL. If `GITEA_TOKEN` is unset, your profile was
not granted PR-create — push the branch (step 4) and report the branch name so the lead opens the
PR.
The response JSON carries `"html_url"` — that is your PR URL. On a non-2xx, read the error body,
fix it if the cause is yours (e.g. branch not pushed yet), and report the failure rather than
inventing a URL. If `GITEA_TOKEN` is unset your profile was not granted PR-create: push the branch
and report its name so the lead opens the PR.
## 6. Reply via `bridge_reply` — the PR is the handoff
## 6. Hand off — what goes in `bridge_reply`
End your turn with **exactly one** `bridge_reply`. That reply is the entire handoff — the lead
cannot see your terminal. Include:
The reply is the entire handoff; the lead cannot see your terminal.
```
PR: <html_url from step 5, or "not created: <reason>" + branch name>
@@ -100,27 +89,21 @@ tests: <what you ran and its REAL result — or "not run: <why>">
summary: <2-3 lines: what you implemented and any caveat the reviewer needs>
```
Then stop. **Do not merge. Do not touch `.mcp.json` or `wiki/`.** One reply closes the turn.
```mermaid
sequenceDiagram
autonumber
participant L as Lead
participant B as bridged
participant I as Implementer (you)
participant G as git / gitea
L->>B: bridge_send(task) — blocks
B-->>I: your assignment (in a worktree on your branch)
L->>I: delegated task (you are in a worktree on your branch)
I->>I: implement + build/test here
I->>G: git commit (never .mcp.json / wiki)
I->>G: git push -u origin HEAD
I->>G: POST /pulls (GITEA_TOKEN) — open PR to main
G-->>I: html_url
I->>B: bridge_reply(PR url, branch, files, tests)
B-->>L: { outcome:"reply", text }
Note over L,G: lead reviews the PR, merges on green — you never merge
I->>L: bridge_reply(PR url, branch, files, tests)
Note over L,G: the lead reviews the PR and merges on green — you never merge
```
*The implement turn: work in the worktree, commit → push → open the PR, hand off the URL. The
lead is the merge gate.*
*The implement turn: work in the worktree, commit → push → open the PR, hand off the URL.*
+25 -58
View File
@@ -1,69 +1,39 @@
---
name: reviewer
description: Reviewer-role playbook for a bridged worker — read the assigned scope, find the real issues, ask the lead via bridge_ask when a decision is genuinely theirs, and report the finding via bridge_reply. Load this when you have been delegated a code review over bridged.
description: Reviewer-role procedure for a bridged worker — how to work a review scope and the exact shape of the finding to report. Load this when the lead delegates you a code review over bridged.
---
# Reviewer worker
# Reviewer worker — procedure
You are a **reviewer** in the claude-bridge fleet. The lead delegated you one scoped review
over `bridged`, and your whole job is **this single turn**: examine the scope it named, and
report back. You are not the owner of the code and you do not merge anything — you surface
what the owner needs to know, then hand the turn back.
Delivery mechanics (how the task reached you, how your reply resolves the lead's blocked
send) are in [`docs/MCP-Contract.md`](../../../docs/MCP-Contract.md); you only need the three
rules below.
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and already applies. This
skill is only the *review procedure*: how to work the scope, and the exact shape of what you
send back.
## 1. Read the whole scope before you judge
The delegation names your scope — a file, a diff, a PR, a function. **Read all of it first.**
A review that fires on a snippet misses the caller that makes it safe (or the one that makes
it a bug). Reviewing only part of the scope and guessing the rest is the most common way a
reviewer worker is wrong.
A review that fires on a snippet misses the caller that makes it safe (or the one that makes it
a bug). Reviewing part of the scope and guessing the rest is the most common way a reviewer is
wrong.
## 2. Stay in your lane
## 2. Stay in the scope
- Review **only** the assigned scope. If you notice something elsewhere, mention it in one
line — do **not** go hunt it. Wandering is how two workers end up reporting the same thing
and neither covers what it was given.
- Do **not** edit files, run the build, or spawn other workers. You review; the owner acts.
- You never set `ANTHROPIC_BASE_URL` and never touch herdr — you are a Claude Code process,
not part of the transport.
- Review **only** what you were assigned. Something elsewhere looks wrong? One line in your
reply — do not go hunt it. Wandering is how two reviewers report the same thing and neither
covers what it was given.
- Do **not** edit files or run the build. You review; the owner acts.
## 3. When the decision is the lead's — ask, don't guess
## 3. Reach for `bridge_ask` only for a genuine fork
Some things you cannot resolve from the code: an ambiguous requirement, a missing acceptance
criterion, "is this behavior intended or a bug?", or a choice between two defensible fixes.
Guessing there produces a confident-but-wrong finding. Instead **pause and ask the lead** with
`bridge_ask` — a single crisp question. The call blocks; when the lead answers you **resume
the same turn** with the answer and finish. Ask only when the answer changes your finding;
don't narrate options you could decide yourself.
Ambiguous requirement, a missing acceptance criterion, "intended or a bug?", or two defensible
fixes with different consequences — those are the lead's call, and guessing produces a
confident-but-wrong finding. Anything you could settle by reading more code is yours to settle.
```mermaid
sequenceDiagram
participant L as Lead
participant B as bridged
participant R as Reviewer (you)
## 4. The finding — what goes in `bridge_reply`
L->>B: bridge_send(review scope) — blocks
B-->>R: your assignment
R->>R: read the full scope
opt a decision only the lead can make
R->>B: bridge_ask("intended, or a bug?") — you block
B-->>L: { outcome:"question", turn_id }
L->>B: bridge_send(answer, turn_id)
B-->>R: { answer } — you resume the SAME turn
end
R->>B: bridge_reply(structured finding) — ends your turn
B-->>L: { outcome:"reply", text }
```
*The review turn, with the optional `bridge_ask` detour when the call is the lead's to make.*
## 4. Report with `bridge_reply` — one structured finding
End your turn with **exactly one** `bridge_reply`. Report the **single most important** real
issue in the scope, in these four lines, under ~90 words:
Report the **single most important** real issue in the scope, in these four lines, under
~90 words:
```
1. <path>:<line>
@@ -72,13 +42,10 @@ issue in the scope, in these four lines, under ~90 words:
4. severity: high | medium | low
```
- Found nothing real after reading? Reply `NO ISSUE` and one line saying why — a clean review
is a valid result, and a fabricated issue is worse than none.
- **Nothing real after reading?** Reply `NO ISSUE` and one line saying why. A clean review is a
valid result; a fabricated issue is worse than none.
- **Severity:** `high` = wrong result, data loss, security, or a hang/crash on a real path ·
`medium` = a real bug on an edge path, or a correctness risk under load/concurrency ·
`low` = clarity, a latent foot-gun, or a smell with no current failure.
- Be specific and verifiable: a line number and a one-line repro beat an adjective. If you
can't point to where it goes wrong, you haven't found it yet.
One reply closes the turn. If you asked mid-turn, the answer you got is already folded into
this finding — you do not ask again after replying.
- Be specific and verifiable: a line number and a one-line repro beat an adjective. If you can't
point at where it goes wrong, you haven't found it yet.
+55
View File
@@ -0,0 +1,55 @@
name: CI
on:
push:
branches: [main]
pull_request:
branches: [main]
jobs:
build:
runs-on: ubuntu-latest
steps:
# The wiki submodule is docs only and is not needed to build — leave it unfetched so CI
# does not depend on the wiki repo being reachable.
- uses: actions/checkout@v4
# The runner image ships an older default-jdk; bridged sets maven.compiler.release=25, so
# provision the JDK explicitly rather than apt-installing whatever "default" means today.
- name: Set up JDK 25
uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '25'
cache: maven
# setup-java provisions the JDK only — it does NOT install Maven, and the runner image has
# no mvn on PATH (a bare `mvn` exits 127). Install it separately. apt pulls a default JRE as
# a dependency; JAVA_HOME from setup-java still wins, which the version check below proves.
- name: Install Maven
run: |
apt-get update && apt-get install -y --no-install-recommends maven
mvn -version
- name: Build and test
working-directory: bridged
# This IS the mock-socket surface CB-503 asks for: the pom's `default-excludes` profile
# already sets excludedGroups=contract, so the @Tag("contract") tests — which need a live
# herdr socket and a RabbitMQ container — are excluded without any flag here. Everything
# that runs does so against the fake UDS herdr and fake ccs/claude stubs.
run: mvn -B clean install
# Deliberately NOT actions/upload-artifact: this Gitea instance presents as GHES, and
# @actions/artifact v2+ (i.e. upload-artifact@v4) refuses to run there —
# "GHESNotSupportedError ... not currently supported on GHES", which red-Xes an otherwise
# green build. Since the artifact could not be retrieved anyway, dump the failing tests into
# the log instead, where they are actually readable.
- name: Failing test output
if: failure()
working-directory: bridged
run: |
for f in target/surefire-reports/*.txt; do
[ -f "$f" ] || continue
grep -qE "Failures: [1-9]|Errors: [1-9]" "$f" && { echo "===== $f ====="; cat "$f"; }
done
exit 0
+189
View File
@@ -1,7 +1,196 @@
# claude-bridge — project instructions
## Bridge communication (enforced — read this first)
> **Canonical block.** Everything down to §Layering is the portable bridge charter, copied verbatim
> into every project that mounts the bridge MCP. Keep it byte-identical with the template in the
> wiki ([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable
> CLAUDE.md block*); improvements go to the template first, then out to each project. Anything
> specific to *this* repo lives under §Project addendum below, never inline above it.
If no `bridge_*` MCP tools are mounted in this session, this section does not apply — skip it.
`bridged` is the **sole communication gateway** between agents here. The orchestrating session (the
**primary**) and every delegated peer (a **worker**) mount the *same* MCP server and talk only
through its `bridge_*` tools. No session addresses a peer, a broker, or the network directly.
### Which role am I? — settle this before acting
**Both roles read this file.** A worker runs in a git worktree of this same repo, so it inherits
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
**Call `bridge_whoami`.** It returns `{"role":"primary"}` or `{"role":"worker","sessionId":…,
"profile":…,"worktree":…,"branch":…}`, resolved by the daemon from your connection — unforgeable,
and the same resolution its authorization gate uses. Don't infer what you can ask.
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
fires: the reply charter in your system prompt (*"You are an off-subscription worker in the
claude-bridge fleet"*) ⇒ **worker**; bridge tools prefixed `mcp__bridge__*` ⇒ **worker** (the
launcher fixes that mount name; a primary's mount is named by whoever wrote its `.mcp.json`, so it
varies); `ANTHROPIC_BASE_URL` set ⇒ **worker** (Claude-model workers run on a clean env, so its
*absence* proves nothing). **Still unsure ⇒ act as a worker.** The two mistakes are not symmetric: a
primary acting as a worker is refused by the authorization gate — loud and self-correcting — while a
worker acting as the primary ends its turn with no `bridge_reply`, and the sender silently receives
nothing. Fail toward the recoverable error.
### Invariants — both roles, no exceptions
1. **Never set, export, or forward `ANTHROPIC_BASE_URL`** (or `ANTHROPIC_AUTH_TOKEN`). The primary
stays on subscription; only the bridge puts a worker off it, at spawn. Mounting the bridge must
never move a session across that boundary.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/send/drain are primary-only; reply/ask are
worker-only-and-only-as-itself. A call outside your role is refused, not queued.
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
re-send because a call looks slow — the bridge delivers when the peer is `idle`/`blocked`.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
### Primary (lead) — run this on every task, in order
**Delegate by default — that is the job.** With the bridge mounted you are an orchestrator on a
metered subscription, and workers are cheap, parallel, and disposable. The default answer to "who
does this?" is **a worker**, not you. Reach for `bridge_send` before you reach for `Edit`. The steps
below are the procedure — run them in order, every task, not only the big ones.
0. **Know your role** — `bridge_whoami`, once per session, before anything else.
1. **Split.** Write the unit list. Every unit carries: scope · the files or PR in question ·
acceptance criteria · exactly what to report back. A unit with no acceptance criteria is not
ready to delegate — refine it or keep it.
2. **Gate each unit** on one question: **"can I write a brief good enough for a worker to
succeed?"** — *not* "could I do this faster myself?" (usually you could; doing it yourself costs
your context and your subscription, while a wasted worker turn costs a worker turn). Yes ⇒
delegate. The keep-list is closed: the conversation with the user, decomposition and planning,
the final judgment call, verification, merges, and anything that depends on context only you
hold. Nothing else is yours by default.
3. **Spawn every delegated unit first** — `bridge_spawn{profile, worktree:true, ticket}`, one per
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model and cost, not in
tier, so the default is rarely what you want.
4. **Then send them all** — `bridge_send{sessionId, content, wait:false}`. Line 1 of every brief is
`Load the <name> skill.` naming the worker's playbook; those skills are opt-in and that line is
what makes them reliable. Where the project ships no such skill, spell the procedure out in the
brief instead. The brief is self-contained — the worker sees your message and the repo, nothing
of your context, your plan, or your screen.
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`bridge_status`, never by reading its terminal.
6. **Verify yourself.** Re-run the build and the checks. A worker mounts only the bridge MCP and
cannot run your other tooling, and a piped command (`… | tail`) hides failures behind a zero
exit — never promote a worker's "clean" to a fact.
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
as it lands; don't wait for the last implementer. Under ~50 changed lines, skip the fan-out and
read it yourself.
8. **Adjudicate, merge, tear down — yours alone.** Read the diff yourself: fully if it is small,
targeted at the reported findings and the risky paths if it is large. Reviewer findings direct
your attention; they never substitute for it. Then merge, then `bridge_stop{paneId}`.
**Steps 3 and 4 are separate on purpose** — spawning and sending in one loop is how parallel work
silently becomes serial, and it is the most common way this layer is wasted. For the same reason,
prefer `wait:false` + `bridge_poll` for anything non-trivial: a blocking `bridge_send` is capped by
*your own* MCP client call timeout (~60s), well below the task's real runtime.
**Delegating does not delegate responsibility.** Workers open PRs; you are the gate. Never delegate
the merge — and merging on a reviewer's word is delegating it by proxy.
| Intent | Tool |
|---|---|
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a worker | `bridge_spawn{profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` · one worker's state: `bridge_status{sessionId}` |
| Delegate (blocking) | `bridge_send{sessionId, content}` |
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
| Answer a worker's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
| Tear down | `bridge_stop{paneId}` |
### Worker — the turn contract
1. **Load the playbook skill the lead named** before doing anything else.
2. **Do the assigned scope only.** Note anything you spot outside it in one line; don't go hunt it.
3. **`bridge_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
defensible fixes, "bug or intended?"). It blocks and you resume the *same* turn with the answer.
Don't ask what you could decide yourself.
4. **End the turn with exactly one `bridge_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `bridge_reply` ⇒ the sender gets nothing and the exchange stalls.
5. **Report honestly.** State only what you actually ran and its real output, including failures.
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
so never claim the result of a check you had no way to run.
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
### Where each rule lives (don't duplicate — extend the right layer)
| Layer | Scope | Reaches |
|---|---|---|
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every worker, at launch, every peer kind |
| **this section** | protocol + orchestration policy | primary **and** every Claude worker — tracked in git, so worktrees inherit it |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a worker told to load one |
| the bridge's own docs | design detail, flows, error model | on demand |
A rule belongs in **exactly one** layer — the outermost one that must obey it. Peers that don't read
`CLAUDE.md` (non-Claude adapters) get the charter only, so any rule *they* must obey belongs in the
charter, not here.
## Project addendum — claude-bridge (not part of the canonical block)
- **This repo is the bridge.** The daemon is `bridged`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/BridgeMcp` (tools), `auth/Authz` (the role table),
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
`reviewer` (scoped review → one structured finding). Name one in every delegation.
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
session's context.
### The prompt is part of the product — update it with the code (mandatory)
This repo *is* the bridge, so the canonical block above is not documentation about someone else's
system: it is the instruction surface this codebase ships. **Every change here must end by asking
whether the block still tells the truth.** A code change that silently invalidates it is an
incomplete change — the agents reading it have no other source.
Before you call any work done, check the row that matches what you touched:
| You changed… | Re-read and update… |
|---|---|
| a `bridge_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
| `Authz` / the role table | invariant 3, and the primary-only vs worker-only claims |
| `ConnectionIdentity` / how a caller is resolved | the `bridge_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__bridge__*`), and the layering table's top row |
| the injector / status gating | invariant 4 |
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
Then **propagate**: the block in this file and the template in the wiki
([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
block*) must stay byte-identical, and other projects carrying the block need the same edit. Verify
rather than trust:
```bash
python3 - <<'PY'
import pathlib
c = pathlib.Path("CLAUDE.md").read_text()
w = pathlib.Path("wiki/7-Use-Cases.md").read_text()
S, E = "## Bridge communication (enforced", "## Project addendum — claude-bridge"
block = c[c.index(S):c.index(E)].rstrip() + "\n"
i = w.index("```markdown\n") + len("```markdown\n")
print("in sync:", w[i:w.index("\n```\n", i) + 1] == block)
PY
```
## IDE MCP tools & validation workflow (enforced)
> **Primary only.** Workers have no IDE MCP mount — if you are a worker, skip this section and
> report the build/test output you actually ran (see §Bridge communication → Worker).
Two IDE MCP servers are connected: **intellij-index** (semantic code intelligence) and
**jetbrains** (file problems, reformat, debugger). IntelliJ has multiple projects open; our
module is **`bridged`**. Always pass these to IDE MCP tools:
+15 -4
View File
@@ -92,7 +92,8 @@ bridge (code reviews delegated this way have produced committed bug fixes). Sele
primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a
fallback injector.
**Shipped** (Java 25 · Maven · 105 tests green — unit/acceptance + live-herdr contract tests):
**Shipped** (Java 25 · Maven · 266 unit/acceptance tests green; the live-herdr and broker contract
tests run separately via `mvn test -Pcontract`):
- **Core gateway** — herdr socket client (contract-tested vs live 0.7.0); guard-checked worker
spawn with `ANTHROPIC_BASE_URL` injected only into the worker's env; status-gated injector;
@@ -106,7 +107,17 @@ fallback injector.
- **Fleet** — multiple worker profiles, each with an independent base_url guard check; workers
inherit the primary's working directory (never `$HOME`); a readiness gate holds delivery until
a worker's Claude has connected the bridge MCP (no paste lost into its boot window).
- **Blocked-worker path** — `bridge_ask` reverse rendezvous: a worker pauses its delegated turn to
ask the primary and resumes the *same* turn with the answer (CB-205).
- **Session lifecycle** — session manager with spawn/reuse/recycle, `idle_ttl` reaper, `context_cap`,
and graceful drain on shutdown (CB-301/CB-303); per-worker git worktrees on their own branch with
a config-parity overlay, so parallel implementers never stomp each other (CB-301-ext).
- **Reliable worker→primary delivery** — a durable `ReplyInbox` (in-memory by default, AMQP/LavinMQ
for cross-restart durability) holds a reply that arrives with no open send, and an active
status-gated push loop nudges the primary to drain it (CB-307).
- **Pluggable peers** — a `PeerLauncher` SPI with two in-tree adapters, `claude-code` and `opencode`,
routed by a `kind:` discriminator (CB-401/CB-402).
**Next** (see the [roadmap](wiki/8-Roadmap.md)) — structured envelope schema, `bridge_ask`
(blocked-worker path), session lifecycle / recycle / `idle_ttl`, split-host, and hardening
(auth/TLS, `/metrics`, CI, systemd).
**Next** (see the [roadmap](wiki/8-Roadmap.md)) — Stage 5 hardening (auth/TLS, `/metrics`, CI,
service supervision, per-session authz + audit), then cross-host: CB-308 multi-host federation and
CB-500 multi-tier coordination.
+3
View File
@@ -5,6 +5,9 @@ dependency-reduced-pom.xml
# Local runtime config (copy from bridged.example.yaml)
bridged.yaml
# CB-505 audit trail + daemon stdout/stderr — runtime records, never source
logs/
# Editor / OS
*.iml
.idea/
+139 -1
View File
@@ -3,11 +3,34 @@
# bridged is the sole gateway between primary/worker Claude sessions and herdr.
# It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL.
# REST + MCP listen address. Keep it on loopback — bridged is same-host in Stage-1.
# REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token`
# below — bridged REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
bind:
host: 127.0.0.1
port: 8765
# API authentication (CB-501). Governs how a caller that is NOT an on-host worker pane proves it
# is the primary. Worker identity never depends on this: a loopback peer PID that maps to a herdr
# pane is unforgeable and is always honoured, so turning auth on cannot lock the fleet out.
#
# mode: loopback-trust → DEFAULT, and the historical behaviour: any loopback caller that is not
# a worker is the primary, no credential needed. Sound ONLY because the
# OS refuses remote connections to a loopback socket.
# mode: token → such a caller must send `Authorization: Bearer <token>`; without it it
# is anonymous and authorized for nothing. REQUIRED for a non-loopback
# bind — the daemon fails fast otherwise, because "unauthenticated ⇒
# primary" on a reachable port would hand spawn/stop/send to anyone.
# tokenEnv → host env var holding the token (never the literal value). Default
# BRIDGED_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
#
# TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and
# let it own certificate lifecycle, e.g.
# location / { proxy_pass http://127.0.0.1:8765; proxy_set_header Authorization $http_authorization; }
# The broker link gets TLS from its own URI (amqps://…) — see `broker` below.
# auth:
# mode: token
# tokenEnv: BRIDGED_API_TOKEN
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
@@ -26,9 +49,38 @@ herdrSocket: ~/.config/herdr/herdr.sock
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
# cwd on an MCP spawn, else the daemon's cwd — never $HOME. See
# docs/Worker-Startup-and-Trust.md.
# configDir → CLAUDE_CONFIG_DIR for the worker, so it inherits that profile's
# skills/MCP/hooks. Omit to leave the worker on the host default.
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
# worktree sees the same local config (CB-301-ext). Omit for the default set:
# [.mcp.json, .claude/settings.local.json, .env, .envrc].
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
# Opt-in by design — omit and the worker gets no PR-create grant (push over
# SSH is unaffected). The token value itself is never stored in this file.
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
# env → extra environment for this profile's workers, as a literal key/value map
# (CB-511). Use it to give workers a toolchain.
#
# A worker's environment does NOT come from your shell. bridged hands herdr an
# explicit env map and herdr merges it into ITS OWN process env — so before
# CB-511 a worker inherited whatever PATH the herdr server happened to be
# started with, which on a long-lived herdr can predate your toolchain entirely
# and leave workers unable to run `mvn` or `java` at all.
# bridged now propagates ITS OWN PATH to every worker by default; set `env:`
# only to override that or add more (JAVA_HOME, …). Since the default is the
# daemon's PATH, make sure the daemon is started with a good one — see the PATH
# lines in deploy/dev.ltms.bridged.plist and deploy/bridged.service.
#
# Adapter-owned variables always win over `env:`: ANTHROPIC_BASE_URL and the
# rest of the ANTHROPIC_*/CLAUDE_* wiring are applied after it, so an `env:`
# entry cannot repoint a worker past the SubscriptionGuard — which is checked
# against `baseUrl` alone.
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
workers:
gx10: # ccs profile name (NOT a hostname)
kind: claude-code # which adapter spawns this profile (default; may omit)
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
model: coder
placement: tab
@@ -37,6 +89,11 @@ workers:
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".mcp.json", ".claude/settings.local.json", ".env", ".envrc"]
ollama:
baseUrl: http://ollama.ltms.dev # local/self-hosted; usually no token
placement: tab
@@ -44,6 +101,49 @@ workers:
tabLabel: "worker: {profile} #{n}"
mcpUrl: http://127.0.0.1:8765/mcp
argv: ["ccs", "ollama"]
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
# The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a
# `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude.
#
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → bridge_send →
# structured bridge_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
# and need NO credentials — check `opencode models` for the current free list, since the names
# change. That also makes the worker off-subscription by construction.
# opencode-free:
# kind: opencode
# model: opencode/north-mini-code-free # `provider/model` selector, injected as `-m`
# placement: tab
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
#
# CB-508: point an opencode profile at your OWN OpenAI-compatible endpoint (local vLLM, llama.cpp,
# LM Studio, TGI…) instead of opencode's gateway. Setting `baseUrl` on a `kind: opencode` profile
# makes the bridge emit a custom `provider` block into the generated opencode.json — opencode has
# no ANTHROPIC_BASE_URL seam, so this is how the endpoint is pinned.
# baseUrl → a bare host:port gets `/v1` appended (where these servers mount the API); a URL that
# already has a path is used verbatim, so a custom mount point still works.
# model → MUST be "<provider>/<model>". The provider half names the generated block; the model
# half must match an id the server reports at /v1/models. One field drives both the
# declaration and the `-m` flag, so they cannot drift apart. A bare model name with a
# baseUrl set is rejected at spawn rather than silently using the default gateway.
# tokenEnv → optional; its value becomes the provider apiKey. Most local servers ignore the key,
# so a placeholder is used when unset (the AI SDK still requires a non-empty one).
# NOTE: no `guard` entry is needed even with a baseUrl set. The SubscriptionGuard exists to stop a
# worker borrowing the primary's Anthropic subscription, and an opencode process has no Anthropic
# credential path at all.
# opencode-local:
# kind: opencode
# baseUrl: http://127.0.0.1:8000
# model: local-vllm/deepseek-v4-flash
# placement: tab
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
defaultWorker: gx10
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
@@ -54,6 +154,18 @@ guard:
- gx01.gw
- ollama.ltms.dev
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
# unknown keys are ignored, so a snake_case key would be silently dropped (default kept).
# spawnReadyTimeoutMs: 20000
# spawnReadyPollMs: 300
# Worktree provisioning root (CB-301-ext). Where per-worker git worktrees are checked out so
# each worker owns an isolated branch instead of sharing the primary's tree. Omit to default
# to a sibling directory of the repo root.
# worktreeRoot: /Users/me/src/.bridged-worktrees
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
@@ -63,3 +175,29 @@ guard:
# idleTtlSeconds: 300
# contextCap: 10
# drainTimeoutSeconds: 5
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
# Set a broker uri to swap in the AMQP-backed inbox: worker replies with no open send are held
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
# broker:
# uri: amqp://guest:guest@127.0.0.1:5672
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
# pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop
# stops as soon as the primary's inbox is empty.
# terminal → pin the primary's herdr terminal id. Omit to learn it from the connection on
# the first orchestration-side MCP call (the normal case). An off-host or
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
# degrades to pull; the reply is still never lost.
# pushReminders → max nudges before giving up (default 5)
# pushBackoffMs → delay between nudges in ms (default 15000)
# primary:
# terminal: term_65619bd6174568
# pushReminders: 5
# pushBackoffMs: 15000
+142
View File
@@ -0,0 +1,142 @@
# CB-307 — Active push-to-primary + reminder loop (the reliability layer)
**Status:** design (2026-07-19). Builds directly on the shipped durable landing zone
(`AmqpReplyInbox`, main `2bc5f3a`, dogfooded live). gitea #5.
## Why this exists
Stage 2 gave a worker→primary reply a **durable place to wait** when no `bridge_send` is
open: it lands in `agent.<target>.inbox` on the broker and survives a daemon bounce. But
delivery is still **pull** — the primary only sees the reply if it happens to call
`bridge_poll(target)` / `GET /sessions/{id}/replies`. A reply can sit indefinitely while
the primary works on something else.
This layer makes delivery **active**: the bridge *pushes* a nudge to the primary the moment
a reply lands, and keeps reminding (bounded) until the primary drains it. At-least-once,
dedup by `msgId`, and — critically — it never loses the reply even if every push fails,
because the durable inbox is the backstop.
## The hard constraint it works around
The bridge is an MCP **server**; the primary is an MCP **client**. A server cannot call
into a client. So "push to the primary" cannot be an MCP response — it needs a *sideband*
channel. The chosen channel: **inject a synthetic user-turn into the primary's own herdr
terminal pane** — the same mechanism the bridge already uses to deliver tasks to workers,
pointed at the primary's pane instead.
```mermaid
flowchart LR
W["worker"] -->|"bridge_reply (no open send)"| MS["MessageService.reply"]
MS -->|"inbox.publish"| INBOX[("agent.&lt;target&gt;.inbox<br/>(durable, LavinMQ)")]
MS -->|"notify"| LOOP["ReplyPushLoop"]
LOOP -->|"status-gated inject"| PANE["primary's herdr pane"]
PANE -->|"primary drains"| DRAIN["bridge_poll(target)<br/>= peek + ack"]
DRAIN -->|"inbox now empty"| LOOP
LOOP -.->|"still non-empty →<br/>re-inject on backoff"| PANE
classDef store fill:#2c5282,stroke:#1a365d,color:#ffffff;
class INBOX store
```
*Figure 1 — a reply lands in the durable inbox; the push loop nudges the primary's pane;
the primary's drain acks it; a still-full inbox triggers a bounded re-nudge.*
## Three increments
### Increment 1 — learn & store the primary's terminal_id
**Finding (seam map):** `ConnectionIdentity.resolve(remoteAddr, remotePort)` already returns
the caller's herdr `terminal_id` for *every* MCP call, via `PaneLocator.terminalForPid`
(walks `pane.list`, matches the caller PID to a pane's process tree). It is non-null whenever
the caller runs in a herdr pane on this host. Today it's discarded for the primary
(`presence.markPresent` is a no-op on it).
**Plan:** a single-slot `PrimaryRegistry` (thread-safe) holding the primary's `terminal_id`.
Populate it from the **orchestration-side** MCP tools — `bridge_send`, `bridge_spawn`,
`bridge_poll`, `bridge_list`, `bridge_status`, `bridge_profiles` — capturing
`callerTerminal(exchange)` when it is (a) non-null and (b) **not** a registered worker
session in `SessionManager`. That caller is, by construction, the primary. Worker-side tools
(`bridge_reply`, `bridge_ask`) never set it.
- **Config override / pin:** a `primary: { terminal: "<id>" }` block in `BridgedConfig`
(nested record, same shape as `Broker`). Lets an operator pin it, or supply it when
derivation can't (see degrade case).
- **Degrade:** if the primary is off-host or in a non-herdr terminal, `terminalForPid`
returns null and no override is set → **the registry stays empty → the push loop is a
no-op and we fall back to pull** (today's behaviour). The reply is never lost; it's just
not actively pushed. This is a safe, explicit degradation, not a failure.
### Increment 2 — the push loop
A `ReplyPushLoop` component, notified at the single no-waiter call site
(`MessageService.reply` → the `inbox.publish` branch, `MessageService.java:192`).
- **Inject a nudge, not the payload.** The injected turn tells the primary *to drain*
(e.g. "Worker `<target>` returned a reply — run `bridge_poll(target=<target>)` to collect
it"), it does **not** carry the reply text. Rationale: replies can be large/multiline and
terminal injection would mangle them; the drain response is the clean transport. Keeps the
push idempotent — re-nudging is harmless.
- **Ack = drain.** The primary draining (`drainReplies` = peek + ack) is the acknowledgement.
The loop's **stop condition is `inbox.peek(target).isEmpty()`** — the reply is gone from the
inbox because it was acked. No new `bridge_ack` tool needed for v1 (see Increment 3).
- **Status-gated injection (mechanism (b), chosen).** A dedicated lightweight scheduled loop,
**not** the worker `Injector`. It injects via `AgentControl.send(primaryTerminal, nudge)`
(the same herdr `agent.send` = `pane send-text` + submit that delivers to workers) only when
`AgentControl.status(primaryTerminal).injectable()` (IDLE/BLOCKED) — never mid-turn. This keeps
the primary path fully isolated from `WorkerPresence`/`StatusPoller` (which are worker-scoped),
and makes it unit-testable with a fake `AgentControl` + an injected clock (per the CB-306
`LongSupplier` clock + `Runnable` sleeper seam). Rejected (a) reuse-the-Injector: it would force
the primary terminal into the worker poller set and couple to worker-presence semantics — more
integration surface, harder to test, no real gain for a bounded reminder.
- **Bounded reminder / backoff.** While `peek(target)` stays non-empty, re-inject on a
backoff schedule up to a cap (N reminders or a max duration; config
`primary.push_reminders` / `primary.push_backoff_ms`). After the cap, **stop reminding** —
the reply remains in the durable inbox and the next natural poll (or a later worker reply's
nudge) still surfaces it. Bounded so the bridge never spams the primary.
### Increment 3 — optional per-`msgId` `bridge_ack` tool (deferred)
Drain-as-ack is coarse: it clears *all* pending replies for a target at once. If finer
control is ever needed (ack one reply, leave others held), add a `bridge_ack(msgId)` tool
mapping to `inbox.ack(target, msgId)` — the port already supports per-`msgId` ack. Not built
in v1; the stop-on-empty loop is sufficient.
## The two subtleties (decided here)
1. **Which caller is "the primary"?** Connection-derived, not self-reported: the caller whose
resolved terminal is non-null **and not a registered worker session**, seen on an
orchestration-side tool. This never mislabels a worker (workers are in `SessionManager`)
and needs no new env var or argument (identity stays connection-derived, per the existing
`BridgeMcp` invariant).
2. **Readiness-gate mismatch → dedicated loop.** The existing `Injector` gates delivery on
`ready.test(target)` = `WorkerPresence` (the *worker's* MCP connected). The primary is not
in `WorkerPresence`, so reusing `Injector` would mean forcing the primary terminal into the
worker `StatusPoller` set and swapping the `ready` predicate — extra integration surface with
worker-scoped machinery. Decision: **mechanism (b)** — a small dedicated scheduled loop that
calls `AgentControl.status(primaryTerminal).injectable()` then `AgentControl.send(...)`, with
an injected clock. Isolated from worker presence, trivially unit-testable, sufficient for a
bounded reminder. (Verified live: this primary resolves to `term_656c8cc03e1f0b1`, pane
`w2:pY` — the primary genuinely runs in a herdr pane on this host, so the path is exercisable.)
## Boundary note
This is the first time the bridge **writes into the primary's pane** — a new direction of
control. It stays within the communication-bus identity: the injection is a **nudge** (a
synthetic "go drain your replies" turn), **status-gated** so it never interrupts a turn,
**bounded** so it never spams, carries **no env** and **never crosses the subscription
boundary**. The bridge is signalling the primary that it has mail — not driving its work.
## Test plan
- **Unit (hermetic):** `PrimaryRegistry` set/clear/override; the "caller is primary iff
non-null terminal AND not a registered session" predicate; the loop's stop-on-empty and
bounded-reminder logic with an injected clock + a fake injector (no real herdr).
- **Live dogfood (primary-side):** with the daemon on the broker jar + a real worker,
delegate a task, let the worker reply after the `bridge_send` window closes, and observe the
bridge inject a drain nudge into *this* primary pane; confirm draining stops the reminders;
confirm an unreachable primary (registry empty) degrades to pull with no loss.
## Out of scope
Multi-host (CB-308) — the push loop is local-only; a remote primary is reached by its own
local gateway, not cross-host injection. Federation reuses this loop per-gateway.
+70
View File
@@ -24,6 +24,10 @@
<slf4j.version>2.0.16</slf4j.version>
<logback.version>1.5.18</logback.version>
<junit.version>5.11.4</junit.version>
<amqp.version>5.22.0</amqp.version>
<testcontainers.version>1.20.4</testcontainers.version>
<commons-compress.version>1.27.1</commons-compress.version>
<commons-lang3.version>3.18.0</commons-lang3.version>
</properties>
<!--
@@ -62,6 +66,22 @@
<artifactId>jackson-annotations</artifactId>
<version>3.0-rc5</version>
</dependency>
<!-- Testcontainers 1.20.4 pulls commons-compress 1.24.0 (test scope), which carries
CVE-2024-25710 (8.1) + CVE-2024-26308 — both fixed in 1.26.0. Pin the patched line.
Test-scope only (never shipped in the jar), but bumped per the CVE policy. -->
<dependency>
<groupId>org.apache.commons</groupId>
<artifactId>commons-compress</artifactId>
<version>${commons-compress.version}</version>
</dependency>
<!-- Testcontainers 1.20.4 also pulls commons-lang3 3.16.0 (test scope): CVE-2025-48924
(uncontrolled recursion in ClassUtils), fixed in 3.18.0. Pin the patched line.
Test-scope only (never shipped in the jar), bumped per the CVE policy. -->
<dependency>
<groupId>org.apache.commons</groupId>
<artifactId>commons-lang3</artifactId>
<version>${commons-lang3.version}</version>
</dependency>
</dependencies>
</dependencyManagement>
@@ -93,6 +113,16 @@
<version>${mcp.version}</version>
</dependency>
<!-- Broker client (CB-307 Stage 2): AMQP 0-9-1. Default deploy targets LavinMQ; this same
client speaks to RabbitMQ unchanged (URI-only swap), so integration tests run against a
stock RabbitMQ container. Only wired when a broker: block is present in config; absent →
the in-memory ReplyInbox. -->
<dependency>
<groupId>com.rabbitmq</groupId>
<artifactId>amqp-client</artifactId>
<version>${amqp.version}</version>
</dependency>
<!-- Logging -->
<dependency>
<groupId>org.slf4j</groupId>
@@ -112,6 +142,22 @@
<version>${junit.version}</version>
<scope>test</scope>
</dependency>
<!-- Testcontainers RabbitMQ: spins a real broker for the @Tag("contract") AMQP integration
test only. Excluded from the default build (contract group), so `mvn clean install`
stays hermetic and green without Docker; run under -Pcontract with Docker present. -->
<dependency>
<groupId>org.testcontainers</groupId>
<artifactId>rabbitmq</artifactId>
<version>${testcontainers.version}</version>
<scope>test</scope>
</dependency>
<dependency>
<groupId>org.testcontainers</groupId>
<artifactId>junit-jupiter</artifactId>
<version>${testcontainers.version}</version>
<scope>test</scope>
</dependency>
</dependencies>
<build>
@@ -123,6 +169,30 @@
<version>3.14.0</version>
</plugin>
<!--
Coverage (CB-509). Build-time tooling only — never a compile/runtime dependency, so
it adds nothing to the shipped jar. Report lands at target/site/jacoco/index.html and
target/site/jacoco/jacoco.csv. No `check` rule / threshold is wired: a coverage gate
rewards writing tests that execute lines, which is the failure mode this project is
trying to avoid, not encourage.
-->
<plugin>
<groupId>org.jacoco</groupId>
<artifactId>jacoco-maven-plugin</artifactId>
<version>0.8.13</version>
<executions>
<execution>
<id>prepare-agent</id>
<goals><goal>prepare-agent</goal></goals>
</execution>
<execution>
<id>report</id>
<phase>test</phase>
<goals><goal>report</goal></goals>
</execution>
</executions>
</plugin>
<!-- Unit tests run by default; contract tests (live herdr) are tag-excluded. -->
<plugin>
<groupId>org.apache.maven.plugins</groupId>
@@ -3,6 +3,8 @@ package dev.ltms.bridged;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
import dev.ltms.bridged.herdr.WorkspaceControl;
@@ -11,23 +13,39 @@ import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
import dev.ltms.bridged.msg.AmqpReplyInbox;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.msg.ReplyPushLoop;
import dev.ltms.bridged.rest.BridgedApp;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.session.SessionReaper;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.worker.CompositePeerLauncher;
import dev.ltms.bridged.worker.HerdrPeerLauncher;
import dev.ltms.bridged.worker.OpenCodeLauncher;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
/**
* {@code bridged} entry point. Wires the real herdr socket client to the REST app and
@@ -41,6 +59,10 @@ public final class Bridged {
/** How often the injector samples a busy worker's status while it has queued work. */
private static final long INJECT_POLL_MILLIS = 250;
/** CB-504: how long to wait at startup for herdr's socket before serving degraded. */
private static final long HERDR_WAIT_SECONDS = 30;
private static final long HERDR_WAIT_POLL_MILLIS = 500;
static void main(String[] args) {
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
BridgedConfig cfg = BridgedConfig.load(configPath);
@@ -49,6 +71,12 @@ public final class Bridged {
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
guard.assertPrimaryClean(System.getenv());
// CB-501: refuse to start if the bind is wider than the auth mode can defend. Under
// loopback-trust, "not a known worker" means "the primary" — sound only because the OS
// refuses remote connections to a loopback socket. This throws rather than warns so the
// dangerous configuration cannot be reached by ignoring a log line.
cfg.validateAuthExposure();
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
: UnixSocketHerdrClient.defaultSocketPath();
@@ -57,11 +85,47 @@ public final class Bridged {
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
PeerLauncher workers = new ClaudeCodeLauncher(agents, spaces, guard,
cfg.workerProfiles(), cfg.defaultProfile(), System::getenv);
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died with
// the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
Map<String, BridgedConfig.Worker> claudeProfiles = new LinkedHashMap<>();
Map<String, BridgedConfig.Worker> opencodeProfiles = new LinkedHashMap<>();
cfg.workerProfiles().forEach((name, w) -> {
if (w.isOpenCode()) {
opencodeProfiles.put(name, w);
} else {
claudeProfiles.put(name, w);
}
});
List<HerdrPeerLauncher> adapters = new ArrayList<>();
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
claudeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(agents, spaces,
opencodeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
}
PeerLauncher workers = new CompositePeerLauncher(adapters, cfg.defaultProfile());
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
// with /healthz reporting "degraded" is strictly more useful than exiting.
if (awaitHerdr(herdr)) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
// with the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
} else {
log.warn("herdr did not answer within {}s — starting anyway; /healthz will report "
+ "degraded until it comes up. Orphaned worker panes (if any) were NOT reaped.",
HERDR_WAIT_SECONDS);
}
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
@@ -115,14 +179,66 @@ public final class Bridged {
StatusPoller poller = new StatusPoller(agents, injector, INJECT_POLL_MILLIS);
poller.start();
MessageService messages = new MessageService(agents, injector, rendezvous);
// CB-307: reply inbox. A broker: block (with a uri) selects the AMQP-backed durable adapter;
// absent, bridged stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
// connection, so keep the reference to close it in the ordered shutdown hook.
final ReplyInbox replyInbox;
if (cfg.broker() != null && cfg.broker().isConfigured()) {
replyInbox = AmqpReplyInbox.open(cfg.broker().uri());
log.info("reply inbox: AMQP broker (durable) at {}", cfg.broker().uri());
} else {
replyInbox = new InMemoryReplyInbox();
log.info("reply inbox: in-memory (soft-state)");
}
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
PrimaryRegistry primaryRegistry = new PrimaryRegistry(
cfg.primary() != null ? cfg.primary().terminal() : null);
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
// open bridge_send. Uses its own lightweight scheduled executor, separate from the injector.
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
var pushScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-push-").unstarted(r));
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
// CB-512: the push loop takes it too, so nudge outcomes (delivered|exhausted) are counted.
Metrics metrics = BridgedMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, agents, replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
sessions.onRelease(terminal ->
messages.abandon(terminal, "the worker session was released before it replied"));
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// Cast to ClaudeCodeLauncher: BridgeMcp is not yet migrated to PeerLauncher (Stage A scope).
BridgeMcp mcp = new BridgeMcp(messages, rendezvous, (ClaudeCodeLauncher) workers, sessions, identity, presence);
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
final CallerResolver callers;
if (cfg.auth().tokenMode()) {
String token = System.getenv(cfg.auth().tokenEnv());
if (token == null || token.isBlank()) {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting bridged");
}
callers = new CallerResolver(identity, true, token);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = new CallerResolver(identity);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
BridgeMcp mcp = new BridgeMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics);
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
@@ -131,17 +247,60 @@ public final class Bridged {
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
poller.stop();
messages.close();
pushLoop.close();
mcp.close();
if (reaper != null) reaper.stop();
// Release the broker connection last among message resources (no-op for the in-memory inbox).
if (replyInbox instanceof AutoCloseable closeable) {
try {
closeable.close();
} catch (Exception e) {
log.debug("reply inbox close: {}", e.toString());
}
}
herdr.close();
}));
Javalin app = new BridgedApp(herdr, (ClaudeCodeLauncher) workers, sessions, messages, rendezvous, presence, mcp.servlet()).build();
Javalin app = new BridgedApp(herdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics).build();
app.start(cfg.bind().host(), cfg.bind().port());
log.info("bridged listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
*
* @return true if herdr answered, false if it never did
*/
private static boolean awaitHerdr(HerdrClient herdr) {
long deadline = System.nanoTime() + HERDR_WAIT_SECONDS * 1_000_000_000L;
boolean waited = false;
while (true) {
try {
herdr.call("ping");
if (waited) {
log.info("herdr is up");
}
return true;
} catch (HerdrException e) {
if (System.nanoTime() >= deadline) {
return false;
}
if (!waited) {
log.info("waiting up to {}s for the herdr socket…", HERDR_WAIT_SECONDS);
waited = true;
}
try {
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
return false;
}
}
}
}
private Bridged() {
}
}
@@ -0,0 +1,72 @@
package dev.ltms.bridged.auth;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.time.Instant;
import java.time.ZoneOffset;
import java.time.format.DateTimeFormatter;
/**
* Append-only record of privileged actions (CB-505).
*
* <p>Writes JSON lines to a dedicated {@code audit} logger — its own appender, separate from the
* chatty app log — so the trail stays greppable and can later be shipped without dragging debug
* noise along.
*
* <p><strong>Message content is never recorded.</strong> This bridge carries the user's source
* code, diffs, and prompts; an audit trail that quietly accumulated them would be a transcript
* archive wearing a security control's clothing. Records carry <em>who / what / against what /
* outcome</em> and correlation ids only.
*/
public final class AuditLog {
private static final Logger AUDIT = LoggerFactory.getLogger("audit");
private static final DateTimeFormatter TS =
DateTimeFormatter.ofPattern("yyyy-MM-dd'T'HH:mm:ss.SSSXXX").withZone(ZoneOffset.UTC);
private AuditLog() {
}
/** Record an allowed action. */
public static void allowed(Principal caller, Authz.Action action, String target) {
write(caller, action, target, "allowed", null);
}
/** Record a refused action and why. */
public static void denied(Principal caller, Authz.Action action, String target, String reason) {
write(caller, action, target, "denied", reason);
}
/** Record an action that was authorized but then failed downstream (guard, timeout, herdr). */
public static void failed(Principal caller, Authz.Action action, String target, String reason) {
write(caller, action, target, "failed", reason);
}
private static void write(Principal caller, Authz.Action action, String target,
String outcome, String reason) {
Principal c = caller != null ? caller : Principal.anonymous();
StringBuilder sb = new StringBuilder(200);
// The timestamp is built here rather than by the appender pattern: a pattern that wrapped
// literal braces around the message collides with logback's own variable substitution.
sb.append("{\"ts\":\"").append(TS.format(Instant.now())).append('"')
.append(",\"role\":\"").append(c.role()).append('"')
.append(",\"actor\":\"").append(esc(c.describe())).append('"')
.append(",\"pid\":").append(c.pid())
.append(",\"action\":\"").append(action).append('"')
.append(",\"target\":").append(target == null ? "null" : '"' + esc(target) + '"')
.append(",\"outcome\":\"").append(outcome).append('"');
if (reason != null) {
sb.append(",\"reason\":\"").append(esc(reason)).append('"');
}
sb.append('}');
// The appender supplies the timestamp, so it cannot disagree with the app log's clock.
AUDIT.info(sb.toString());
}
/** Minimal JSON string escaping — these values are ids and short reasons, never free text. */
private static String esc(String s) {
return s.replace("\\", "\\\\").replace("\"", "\\\"")
.replace("\n", "\\n").replace("\r", "\\r").replace("\t", "\\t");
}
}
@@ -0,0 +1,72 @@
package dev.ltms.bridged.auth;
/**
* The authorization table (CB-505), stated once and enforced on both entry paths.
*
* <p>Most of these rules are already true de facto — {@code BridgeMcp} derives a worker's identity
* from the connection rather than reading it from an argument, so a worker has never been able to
* reply <em>as</em> another worker over MCP. What was missing is that the REST surface trusted the
* session id in the URL path, and neither surface checked role at all. This class makes the
* invariant explicit and testable rather than emergent.
*/
public final class Authz {
private Authz() {
}
/** A privileged operation, named for the audit trail. */
public enum Action {
/** Spawn a worker peer. */
SPAWN,
/** Tear a worker peer down. */
STOP,
/** Deliver a turn to a session (or answer a worker's question). */
SEND,
/** A worker's terminal reply for its own turn. */
REPLY,
/** A worker's mid-turn question to the primary. */
ASK,
/** Collect held replies from a session's inbox. */
DRAIN,
/** Read-only observation: status, roster, profiles, task polling. */
READ,
/** Scrape the metrics endpoint. */
METRICS
}
/**
* Whether {@code caller} may perform {@code action} against {@code targetSession}.
*
* @param targetSession the session id in the request path; only consulted for the worker-scoped
* actions ({@code REPLY}, {@code ASK}), ignored otherwise, may be
* {@code null}
*/
public static boolean permits(Principal caller, Action action, String targetSession) {
if (caller == null || caller.isAnonymous()) {
return false; // authenticated as nothing ⇒ authorized for nothing
}
return switch (action) {
// Orchestration is the primary's alone. A worker driving spawn/stop/send would be a
// worker escalating into the orchestrator role.
case SPAWN, STOP, SEND, DRAIN -> caller.isPrimary();
// The load-bearing rule: a worker acts only as itself. The primary is deliberately
// excluded — a reply/ask is a worker's own turn output, and letting the primary forge
// one would corrupt the rendezvous correlation it is itself waiting on.
case REPLY, ASK -> caller.ownsSession(targetSession);
// Observation is open to both authenticated roles: a worker legitimately polls its own
// status, and the roster carries no secrets.
case READ, METRICS -> caller.isPrimary() || caller.isWorker();
};
}
/**
* Why a request was refused, for the error body. Distinguishes "you are nobody" from "you are
* somebody, but not the right somebody" — the first is a credential problem (401), the second
* an authorization one (403).
*/
public static boolean isUnauthenticated(Principal caller) {
return caller == null || caller.isAnonymous();
}
}
@@ -0,0 +1,119 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
/**
* Resolves every caller to a {@link Principal}, for both entry paths into the core (CB-501).
*
* <p>There are two of them and they are not layered the way the docs suggest: {@code BridgeMcp}
* calls the service layer directly and is mounted as a raw servlet (so it never passes through a
* Javalin filter), while the REST routes historically resolved no identity at all. Both now
* delegate here, so the authorization rules are stated once instead of drifting apart.
*
* <p><strong>Resolution order</strong> — connection identity first, token second, nothing third:
* <ol>
* <li>A loopback peer PID that maps to a herdr worker pane ⇒ {@link Role#WORKER}. This is
* unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured
* regardless of auth mode, so enabling auth never breaks the fleet.</li>
* <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li>
* <li>Otherwise, under {@code loopback-trust}, a loopback caller ⇒ {@link Role#PRIMARY}
* (the historical behaviour, now an explicit configured choice).</li>
* <li>Otherwise {@link Role#ANONYMOUS}.</li>
* </ol>
*/
public final class CallerResolver {
private final ConnectionIdentity identity;
private final boolean tokenMode;
private final byte[] expectedToken; // null unless tokenMode
/** Loopback-trust resolver: no token required, historical behaviour. */
public CallerResolver(ConnectionIdentity identity) {
this(identity, false, null);
}
/**
* @param identity connection-based worker identification
* @param tokenMode when true, a non-worker caller must present a valid bearer token
* @param token the expected bearer token; required (non-blank) when {@code tokenMode}
*/
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
if (tokenMode && (token == null || token.isBlank())) {
throw new IllegalArgumentException(
"auth.mode=token requires a non-empty token; check that the env var named by "
+ "auth.tokenEnv is exported to the daemon's environment");
}
this.identity = identity;
this.tokenMode = tokenMode;
this.expectedToken = tokenMode ? token.getBytes(StandardCharsets.UTF_8) : null;
}
/**
* Resolve the caller of a request.
*
* @param remoteAddr the connection's remote address
* @param remotePort the connection's remote port (used for the peer-PID lookup)
* @param authorizationHeader the raw {@code Authorization} header, or {@code null}
*/
public Principal resolve(String remoteAddr, int remotePort, String authorizationHeader) {
ConnectionIdentity.Caller c = identity.resolve(remoteAddr, remotePort);
if (c.terminal() != null) {
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
}
if (tokenMode) {
return presentedTokenMatches(authorizationHeader)
? Principal.primary(c.pid())
: Principal.anonymous();
}
// loopback-trust: same-host callers that are not workers are the primary. A non-loopback
// caller is anonymous even here — and startup refuses that combination anyway
// (BridgedConfig.validateAuthExposure), so this is defence in depth, not the control.
return isLoopback(remoteAddr) ? Principal.primary(c.pid()) : Principal.anonymous();
}
/** The working directory of the calling process (CB-112 spawn cwd inheritance), or {@code null}. */
public String cwdForPid(long pid) {
return identity.cwdForPid(pid);
}
/** True when auth requires a bearer token of non-worker callers. */
public boolean tokenMode() {
return tokenMode;
}
private boolean presentedTokenMatches(String authorizationHeader) {
String presented = bearerValue(authorizationHeader);
if (presented == null) {
return false;
}
// Constant-time: MessageDigest.isEqual does not short-circuit on the first differing byte,
// so a token cannot be recovered a byte at a time by timing the response.
return MessageDigest.isEqual(presented.getBytes(StandardCharsets.UTF_8), expectedToken);
}
/** Extract the credential from {@code Authorization: Bearer <token>}, or {@code null}. */
private static String bearerValue(String header) {
if (header == null) {
return null;
}
String h = header.trim();
if (h.length() < 7 || !h.regionMatches(true, 0, "Bearer ", 0, 7)) {
return null;
}
String token = h.substring(7).trim();
return token.isEmpty() ? null : token;
}
private static boolean isLoopback(String remoteAddr) {
if (remoteAddr == null) {
return false;
}
return remoteAddr.equals("127.0.0.1") || remoteAddr.equals("::1")
|| remoteAddr.equals("0:0:0:0:0:0:0:1") || remoteAddr.startsWith("127.");
}
}
@@ -0,0 +1,57 @@
package dev.ltms.bridged.auth;
/**
* A resolved caller: its {@link Role}, and — for a worker — the herdr {@code terminal_id} that
* identifies which worker it is (CB-501).
*
* @param role what this caller is authorized to act as
* @param terminal the worker's herdr terminal id; {@code null} for {@code PRIMARY}/{@code ANONYMOUS}
* @param pid the connecting process id, or {@code -1} when not resolvable (audit context)
*/
public record Principal(Role role, String terminal, long pid) {
/** A caller authenticated as nothing — the default when no check establishes anything else. */
public static Principal anonymous() {
return new Principal(Role.ANONYMOUS, null, -1);
}
/** The orchestrating session. */
public static Principal primary(long pid) {
return new Principal(Role.PRIMARY, null, pid);
}
/** A worker peer, identified by its herdr pane. */
public static Principal worker(String terminal, long pid) {
return new Principal(Role.WORKER, terminal, pid);
}
public boolean isPrimary() {
return role == Role.PRIMARY;
}
public boolean isWorker() {
return role == Role.WORKER;
}
public boolean isAnonymous() {
return role == Role.ANONYMOUS;
}
/**
* Whether this caller may act <em>as</em> {@code sessionId} — the "own session only" rule that
* keeps one worker from replying or asking on another's behalf. Only a worker can own a
* session, and only its own.
*/
public boolean ownsSession(String sessionId) {
return isWorker() && terminal != null && terminal.equals(sessionId);
}
/** Short, non-sensitive description for audit lines and error details. */
public String describe() {
return switch (role) {
case WORKER -> "worker:" + terminal;
case PRIMARY -> "primary";
case ANONYMOUS -> "anonymous";
};
}
}
@@ -0,0 +1,29 @@
package dev.ltms.bridged.auth;
/**
* What a caller is allowed to be on the bus (CB-501).
*
* <p>The ordering matters conceptually: {@link #PRIMARY} is the <em>most</em> privileged role
* (it spawns, stops, sends to any session, and drains any inbox), not the least. Before CB-501
* the daemon reached {@code PRIMARY} by <em>failing</em> every other check — any caller that did
* not resolve to a known worker pane was treated as the primary. That is inverted here:
* {@link #ANONYMOUS} is the fallback, and {@code PRIMARY} must be established.
*/
public enum Role {
/**
* The orchestrating session. Established either by being a loopback caller that is not a
* worker pane (under {@code loopback-trust}) or by presenting a valid bearer token (under
* {@code token} mode).
*/
PRIMARY,
/**
* A worker peer, identified by its herdr pane. Unforgeable: derived from the connection's
* loopback peer PID via herdr's PID→pane map, never from a request argument.
*/
WORKER,
/** Authenticated as nothing. Authorized for nothing but {@code /healthz}. */
ANONYMOUS
}
@@ -27,7 +27,17 @@ import java.util.Set;
* @param guard subscription-boundary allowlist
* @param worktreeRoot nullable root directory for provisioned worktrees; defaults to a sibling
* of the repo root
* @param lifecycle session lifecycle limits ({@code null} = all disabled)
* @param lifecycle session lifecycle limits ({@code null} = all disabled)
* @param spawnReadyTimeoutMs max ms to wait for a spawned worker to reach an injectable state
* ({@code null} / 0 disables the poll gate — legacy non-blocking behaviour)
* @param spawnReadyPollMs poll interval while waiting for the worker to become injectable
* @param broker external AMQP broker for durable reply delivery ({@code null} → in-memory,
* soft-state {@code ReplyInbox}; present → the AMQP-backed adapter, CB-307 Stage 2)
* @param primary optional pinned primary terminal config ({@code null} → derived from connection);
* a non-blank {@code terminal} seeds {@code PrimaryRegistry} and prevents
* connection-derived overrides, CB-307
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
* historical behaviour), CB-501
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record BridgedConfig(
@@ -38,7 +48,12 @@ public record BridgedConfig(
String defaultWorker,
Guard guard,
String worktreeRoot,
Lifecycle lifecycle) {
Lifecycle lifecycle,
Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs,
Broker broker,
Primary primary,
Auth auth) {
@JsonIgnoreProperties(ignoreUnknown = true)
public record Bind(String host, int port) {
@@ -78,6 +93,12 @@ public record BridgedConfig(
* (minimal-grant default — push over SSH stays free, PR-create is opt-in)
* @param gitHostEnv name of the host env var holding the forge host (default {@code GITEA_HOST});
* injected as {@code GITEA_HOST} <em>only</em> when {@code gitTokenEnv} is set
* @param kind which peer launcher spawns this profile: {@code "claude-code"} (default —
* the {@link dev.ltms.bridged.worker.ClaudeCodeLauncher}) or {@code "opencode"}.
* The {@code CompositePeerLauncher} routes {@code spawn}/reap by this value, so
* each adapter drives only its own kind. Normalised to lower-case; blank ⇒ the
* default. It selects the adapter, not the transport — placement, tabs, cwd, and
* the readiness gate are kind-independent and stay in the shared base.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Worker(String profile, String baseUrl, String model,
@@ -85,9 +106,24 @@ public record BridgedConfig(
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd,
List<String> parityOverlay,
String gitTokenEnv, String gitHostEnv) {
String gitTokenEnv, String gitHostEnv,
String kind,
Map<String, String> env) {
/** Peer kind spawned by {@link dev.ltms.bridged.worker.ClaudeCodeLauncher} (the default). */
public static final String KIND_CLAUDE_CODE = "claude-code";
/** Peer kind spawned by the opencode adapter (CB-402). */
public static final String KIND_OPENCODE = "opencode";
public Worker {
argv = (argv == null || argv.isEmpty()) ? List.of("claude") : List.copyOf(argv);
// A claude-code worker defaults its launch command to `claude`; other kinds carry their own
// argv (e.g. `opencode`) and must not inherit the Claude binary — so only default when unset
// AND this is the claude-code kind.
String k = (kind == null || kind.isBlank()) ? KIND_CLAUDE_CODE : kind.toLowerCase();
argv = (argv == null || argv.isEmpty())
? (KIND_CLAUDE_CODE.equals(k) ? List.of("claude") : List.of(k))
: List.copyOf(argv);
kind = k;
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_WORKER_TOKEN" : tokenEnv;
placement = (placement == null || placement.isBlank()) ? "tab" : placement.toLowerCase();
workspace = (workspace == null || workspace.isBlank()) ? "bridged-workers" : workspace;
@@ -98,6 +134,7 @@ public record BridgedConfig(
// gitTokenEnv stays null when unset (opt-in). gitHostEnv defaults so operators enabling
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
gitHostEnv = (gitHostEnv == null || gitHostEnv.isBlank()) ? "GITEA_HOST" : gitHostEnv;
env = (env == null) ? Map.of() : Map.copyOf(env);
}
/**
@@ -110,13 +147,48 @@ public record BridgedConfig(
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, null, null);
mcpUrl, cwd, parityOverlay, null, null, null);
}
/**
* Backward-compatible constructor with the CB-302 git-forge fields but no explicit peer
* {@code kind} — defaults to {@link #KIND_CLAUDE_CODE}. Keeps pre-CB-402 call sites working.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, null, null);
}
/**
* Backward-compatible constructor without the CB-511 {@code env:} passthrough — the worker
* gets the daemon's PATH and nothing else. Keeps pre-CB-511 call sites working.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, null);
}
/** A copy with {@code profile} set — used to default a profile to its {@code workers} key. */
public Worker withProfile(String p) {
return new Worker(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv);
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env);
}
/** True when this profile is served by the Claude Code adapter (the default kind). */
public boolean isClaudeCode() {
return KIND_CLAUDE_CODE.equals(kind);
}
/** True when this profile is served by the opencode adapter (CB-402). */
public boolean isOpenCode() {
return KIND_OPENCODE.equals(kind);
}
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
@@ -161,6 +233,84 @@ public record BridgedConfig(
public record Lifecycle(Integer idleTtlSeconds, Integer contextCap, Integer drainTimeoutSeconds) {
}
/**
* External AMQP broker for durable, cross-restart reply delivery (CB-307 Stage 2). Its mere
* presence swaps the in-memory {@code ReplyInbox} for the AMQP-backed adapter; absent, bridged
* stays soft-state. Production default is LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1
* and is a URI-only swap.
*
* @param uri AMQP connection URI, e.g. {@code amqp://guest:guest@127.0.0.1:5672/}. Blank/{@code null}
* ⇒ the broker block is treated as absent (in-memory adapter).
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Broker(String uri) {
/** True when a usable broker URI is configured (an empty block does not enable AMQP). */
public boolean isConfigured() {
return uri != null && !uri.isBlank();
}
}
/**
* Optional pinned primary terminal config (CB-307). When present with a non-blank
* {@code terminal}, the bridge uses this as the primary's herdr identity instead of
* deriving it from the MCP connection. Useful when the primary runs off-host or in a
* non-herdr terminal where connection-derived identity is unavailable.
*
* @param terminal the primary's herdr {@code terminal_id} ({@code null}/blank → derive)
* @param pushReminders max reminder nudges before giving up (default 5)
* @param pushBackoffMs delay between reminders in ms (default 15000)
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Primary(String terminal, Integer pushReminders, Integer pushBackoffMs) {
/** @return configured reminder cap, or 5 */
public int remindersOrDefault() {
return pushReminders != null ? pushReminders : 5;
}
/** @return configured backoff in ms, or 15000 */
public long backoffMsOrDefault() {
return pushBackoffMs != null ? pushBackoffMs.longValue() : 15_000L;
}
}
/**
* API authentication (CB-501). Governs how a caller that is <em>not</em> an on-host worker
* pane proves it is the primary.
*
* <p>Worker identity never depends on this block: a loopback peer PID that maps to a herdr
* pane is unforgeable and is always honoured (see
* {@link dev.ltms.bridged.mcp.ConnectionIdentity}). This only decides what happens for
* <em>everyone else</em>.
*
* @param mode {@code "loopback-trust"} (default) — any loopback caller that is not a known
* worker is the primary, no credential needed; this is the historical
* behaviour, now chosen explicitly rather than implied. {@code "token"} — such
* a caller must present {@code Authorization: Bearer <token>} or it is
* {@code ANONYMOUS} and authorized for nothing.
* @param tokenEnv name of the host env var holding the bearer token; the literal value is
* never stored in config. Defaults to {@code BRIDGED_API_TOKEN}. Only read
* when {@code mode} is {@code token}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Auth(String mode, String tokenEnv) {
/** Historical behaviour: loopback non-worker ⇒ primary, no credential. */
public static final String MODE_LOOPBACK_TRUST = "loopback-trust";
/** A non-worker caller must present a valid bearer token to be the primary. */
public static final String MODE_TOKEN = "token";
public Auth {
mode = (mode == null || mode.isBlank()) ? MODE_LOOPBACK_TRUST : mode.toLowerCase();
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_API_TOKEN" : tokenEnv;
}
/** True when a bearer token is required of every non-worker caller. */
public boolean tokenMode() {
return MODE_TOKEN.equals(mode);
}
}
/**
* Subscription boundary. Only these hosts may back a worker's
* {@code ANTHROPIC_BASE_URL}; the primary must carry none.
@@ -229,6 +379,45 @@ public record BridgedConfig(
Bind b = bind != null ? bind : new Bind(null, 0);
Guard g = guard != null ? guard : new Guard(List.of());
Lifecycle l = lifecycle != null ? lifecycle : new Lifecycle(null, null, null);
return new BridgedConfig(b, herdrSocket, worker, workers, defaultWorker, g, worktreeRoot, l);
Integer timeout = (spawnReadyTimeoutMs != null) ? spawnReadyTimeoutMs : 20000;
Integer pollMs = (spawnReadyPollMs != null) ? spawnReadyPollMs : 300;
Auth a = auth != null ? auth : new Auth(null, null);
// broker is left as-is: null (or an empty/blank uri) keeps the in-memory soft-state inbox.
// primary is left as-is: null defaults to connection-derived identity.
return new BridgedConfig(b, herdrSocket, worker, workers, defaultWorker, g, worktreeRoot, l, timeout, pollMs, broker, primary, a);
}
/**
* Reject a configuration whose network exposure outruns its authentication (CB-501).
*
* <p>{@code loopback-trust} means "any caller that is not a known worker pane is the primary" —
* safe only because the OS refuses non-local connections to a loopback bind. Widen
* {@code bind.host} without switching to {@code token} mode and that sentence becomes "any
* client that can reach this port is the primary", which is the most privileged role on the
* bus. Rather than document the hazard, make it unrepresentable: fail fast at startup.
*
* @throws IllegalStateException when a non-loopback bind is paired with {@code loopback-trust}
*/
public void validateAuthExposure() {
String host = bind().host();
if (isLoopbackBind(host) || auth().tokenMode()) {
return;
}
throw new IllegalStateException(
"refusing to start: bind.host=" + host + " is not loopback, but auth.mode="
+ auth().mode() + ". A non-loopback bind treats every unauthenticated "
+ "caller as the primary (spawn/stop/send/drain on any session). Set "
+ "auth.mode: token (with auth.tokenEnv) before exposing this port, or "
+ "bind to 127.0.0.1 and put a reverse proxy in front.");
}
/** True for the loopback addresses and the unspecified-but-local forms we treat as same-host. */
private static boolean isLoopbackBind(String host) {
if (host == null || host.isBlank()) {
return true; // Bind's own default is 127.0.0.1
}
String h = host.trim().toLowerCase();
return h.equals("127.0.0.1") || h.equals("::1") || h.equals("localhost")
|| h.startsWith("127.");
}
}
@@ -1,15 +1,23 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.AuditLog;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.auth.Role;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.peer.PeerLauncher;
import io.modelcontextprotocol.common.McpTransportContext;
import io.modelcontextprotocol.json.McpJsonMapper;
import io.modelcontextprotocol.json.jackson3.JacksonMcpJsonMapperSupplier;
@@ -34,8 +42,8 @@ import java.util.stream.Collectors;
* / {@code bridge_status}; the worker calls {@code bridge_reply}.
*
* <p>Beyond delegation the primary also manages the fleet here (CB-108): {@code bridge_spawn} /
* {@code bridge_list} / {@code bridge_stop} adapt {@link ClaudeCodeLauncher} so a worker's whole lifecycle
* is driven through MCP, with the subscription boundary still enforced inside {@code ClaudeCodeLauncher}.
* {@code bridge_list} / {@code bridge_stop} drive the {@link PeerLauncher} SPI so a worker's whole
* lifecycle is managed through MCP, with each adapter's subscription boundary enforced inside it.
*
* <p>The tool <em>logic</em> lives in package-private static methods returning a
* {@link McpSchema.CallToolResult}, so it is unit-testable without standing up the HTTP transport;
@@ -55,12 +63,34 @@ public final class BridgeMcp {
static final String CALLER_TERMINAL = "callerTerminal";
/** Transport-context key under which the extractor stashes the caller's PID (for cwd inherit). */
static final String CALLER_PID = "callerPid";
/** Transport-context key under which the extractor stashes the resolved {@link Role} (CB-501). */
static final String CALLER_ROLE = "callerRole";
private final HttpServletStreamableServerTransportProvider transport;
private final McpSyncServer server;
private final CallerResolver authz; // CB-501: null → authorization not enforced (legacy)
private final Metrics metrics; // CB-502: null → auth failures not counted
public BridgeMcp(MessageService messages, Rendezvous rendezvous, ClaudeCodeLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence) {
/**
* Legacy constructor — no authorization. Retained so existing tests exercise tool behaviour
* without an auth fixture.
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
PrimaryRegistry primaryRegistry) {
this(messages, workers, sessions, identity, presence, primaryRegistry, null, null);
}
/**
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
* Jetty's context handler and never passes through Javalin's {@code before}
* filter, so the REST guard does not cover it.
* @param metrics registry for auth-failure counting; may be {@code null}
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
PrimaryRegistry primaryRegistry, CallerResolver callers, Metrics metrics) {
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
.jsonMapper(json)
@@ -70,17 +100,28 @@ public final class BridgeMcp {
// inherit the primary's cwd (CB-112). Any contact from a worker marks it available
// (CB-113) — its MCP initialize is the reliable "the agent is up" signal.
.contextExtractor(req -> {
ConnectionIdentity.Caller c = identity.resolve(req.getRemoteAddr(), req.getRemotePort());
presence.markPresent(c.terminal()); // no-op for the primary (null terminal)
// One resolution per call, shared with the REST surface via CallerResolver so
// the two paths cannot drift on who a caller is.
Principal p = callers != null
? callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization"))
: legacyPrincipal(identity, req.getRemoteAddr(), req.getRemotePort());
presence.markPresent(p.terminal()); // no-op for the primary (null terminal)
return McpTransportContext.create(Map.of(
CALLER_TERMINAL, orEmpty(c.terminal()),
CALLER_PID, Long.toString(c.pid())));
CALLER_TERMINAL, orEmpty(p.terminal()),
CALLER_PID, Long.toString(p.pid()),
CALLER_ROLE, p.role().name()));
})
.build();
this.server = McpServer.sync(transport)
.serverInfo("bridge", "0.1.0")
.capabilities(McpSchema.ServerCapabilities.builder().tools(true).build())
.toolCall(sendTool(), (_, req) -> {
.toolCall(sendTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SEND,
str(req.arguments(), "sessionId"));
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
Map<String, Object> a = req.arguments();
String turnId = str(a, "turnId");
if (turnId != null && !turnId.isBlank()) {
@@ -93,18 +134,46 @@ public final class BridgeMcp {
? sendAsync(messages, str(a, "sessionId"), str(a, "content"))
: send(messages, str(a, "sessionId"), str(a, "content"), timeoutMs(a));
})
// bridge_reply's identity is the CONNECTION, never an argument.
.toolCall(replyTool(), (exchange, req) ->
reply(rendezvous, callerTerminal(exchange), str(req.arguments(), "content")))
// bridge_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
.toolCall(replyTool(), (exchange, req) -> {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.REPLY, self);
if (denied != null) return denied;
return reply(messages, self, str(req.arguments(), "content"));
})
// bridge_ask (CB-205): a worker's mid-turn question — identity from the CONNECTION.
.toolCall(askTool(), (exchange, req) ->
ask(messages, callerTerminal(exchange), str(req.arguments(), "question"), timeoutMs(req.arguments())))
.toolCall(statusTool(), (_, req) ->
status(messages, str(req.arguments(), "sessionId")))
.toolCall(pollTool(), (_, req) ->
poll(messages, str(req.arguments(), "ticket")))
.toolCall(askTool(), (exchange, req) -> {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.ASK, self);
if (denied != null) return denied;
return ask(messages, self, str(req.arguments(), "question"), timeoutMs(req.arguments()));
})
.toolCall(statusTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return status(messages, str(req.arguments(), "sessionId"));
})
.toolCall(pollTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
Map<String, Object> a = req.arguments();
return poll(messages, str(a, "ticket"), str(a, "target"));
})
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
// Acking removes a reply from the inbox, so it is a drain, not a read.
.toolCall(ackTool(), (exchange, req) -> {
Map<String, Object> a = req.arguments();
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.DRAIN, str(a, "target"));
if (denied != null) return denied;
return ack(messages, str(a, "target"), str(a, "msgId"));
})
// Fleet management (CB-108): spawn/list/stop over ClaudeCodeLauncher.
.toolCall(spawnTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SPAWN, null);
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
Map<String, Object> a = req.arguments();
// CB-112: worker inherits the primary's cwd unless the call pins one.
// CB-301: carry the caller's identity as the session owner (null for the primary).
@@ -113,10 +182,107 @@ public final class BridgeMcp {
return spawn(sessions, str(a, "profile"), str(a, "cwd"), callerCwd,
callerTerminal(exchange), worktreeRequest(a));
})
.toolCall(listTool(), (_, _) -> listWorkers(workers, sessions))
.toolCall(stopTool(), (_, req) -> stop(sessions, str(req.arguments(), "paneId")))
.toolCall(profilesTool(), (_, _) -> profiles(workers))
.toolCall(listTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listWorkers(workers, sessions);
})
.toolCall(stopTool(), (exchange, req) -> {
String paneId = str(req.arguments(), "paneId");
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.STOP, paneId);
if (denied != null) return denied;
return stop(sessions, paneId);
})
.toolCall(profilesTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return profiles(workers);
})
.toolCall(whoamiTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return whoami(principal(exchange), sessions);
})
.build();
this.authz = callers;
this.metrics = metrics;
}
/**
* Pre-CB-501 identity: worker if the connection maps to a pane, otherwise the primary. Used
* only by the legacy constructor, where authorization is not enforced anyway.
*/
private static Principal legacyPrincipal(ConnectionIdentity identity, String addr, int port) {
ConnectionIdentity.Caller c = identity.resolve(addr, port);
return c.terminal() != null
? Principal.worker(c.terminal(), c.pid())
: Principal.primary(c.pid());
}
/** The caller reconstructed from the transport context. */
private static Principal principal(McpSyncServerExchange exchange) {
return principalFrom(exchange.transportContext().get(CALLER_ROLE),
callerTerminal(exchange), callerPid(exchange));
}
/**
* Rebuild a {@link Principal} from the three values the context extractor stashed.
*
* <p>Split out from {@link #principal(McpSyncServerExchange)} so the identity rules are
* reachable without an {@code McpSyncServerExchange} — that is an SDK type this project has no
* mocking library to fabricate, which is why this logic had no test at all until CB-513.
*
* @param role the stashed {@link Role} name, or {@code null} on the legacy path
* @param terminal the worker terminal, or {@code null} for a non-worker
* @param pid the calling pid, or {@code -1}
*/
static Principal principalFrom(Object role, String terminal, long pid) {
if (role == null) {
// No role stashed (legacy path): fall back to the historical interpretation.
return terminal != null ? Principal.worker(terminal, pid) : Principal.primary(pid);
}
return new Principal(Role.valueOf(role.toString()), terminal, pid);
}
/**
* Gate a tool call on the CB-505 table. Returns {@code null} when the call may proceed, or the
* error result to return when it may not.
*/
private McpSchema.CallToolResult deny(McpSyncServerExchange exchange, Authz.Action action,
String target) {
return denyFor(principal(exchange), action, target);
}
/**
* The policy half of {@link #deny}: everything except pulling the caller out of the MCP
* exchange. Kept separate so the authorization decision — the actual control — is unit-testable
* without fabricating an SDK {@code McpSyncServerExchange}.
*
* <p>This surface exists because the enforcement was previously unreachable from a test: no
* test constructs a {@code BridgeMcp}, so the whole MCP-side gate ran zero times in the suite
* while the REST-side equivalent had ten tests. A security control nothing exercises is a
* claim, not a control.
*
* @return {@code null} when the call may proceed, or the error result to return when it may not
*/
McpSchema.CallToolResult denyFor(Principal caller, Authz.Action action, String target) {
// The enforcement switch lives HERE rather than in the exchange-facing wrapper: any future
// tool that calls this directly must not be able to skip the gate by accident.
if (authz == null) {
return null; // legacy constructor: authorization not enforced
}
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
return null;
}
String reason = Authz.isUnauthenticated(caller) ? "unauthenticated" : "forbidden";
AuditLog.denied(caller, action, target, reason);
if (metrics != null) {
metrics.inc(BridgedMetrics.AUTH_FAILURES, "reason", reason);
}
return error(reason + ": " + caller.describe() + " may not " + action);
}
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
@@ -235,10 +401,17 @@ public final class BridgeMcp {
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
}
/** {@code bridge_poll}: check an async delegation by ticket (pending / done+reply / failed). */
static McpSchema.CallToolResult poll(MessageService messages, String ticket) {
/** {@code bridge_poll}: check an async delegation by ticket, or drain a worker's inbox by target. */
static McpSchema.CallToolResult poll(MessageService messages, String ticket, String target) {
if (!isBlank(target)) {
var replies = messages.drainReplies(target);
if (replies.isEmpty()) {
return text("[]");
}
return text(json(replies));
}
if (isBlank(ticket)) {
return error("ticket is required");
return error("ticket (or target) is required");
}
MessageService.TaskView v = messages.poll(ticket);
if (v == null) {
@@ -254,11 +427,12 @@ public final class BridgeMcp {
}
/**
* {@code bridge_reply}: the worker returns its structured answer, resolving the awaiting send.
* {@code bridge_reply}: the worker returns its structured answer, resolving the awaiting send
* or — when no send is open — queueing the reply in the inbox for later drain (CB-307).
* {@code callerTerminal} is resolved from the connection (never an argument); a {@code null}
* means the caller is not a known worker (e.g. the primary called it by mistake).
*/
static McpSchema.CallToolResult reply(Rendezvous rendezvous, String callerTerminal, String content) {
static McpSchema.CallToolResult reply(MessageService messages, String callerTerminal, String content) {
if (callerTerminal == null) {
return error("bridge_reply is for workers only — could not identify the calling worker "
+ "from the connection");
@@ -266,9 +440,17 @@ public final class BridgeMcp {
if (content == null) {
return error("content is required");
}
return rendezvous.resolve(callerTerminal, content)
? text("delivered")
: error("no send is awaiting a reply for this worker");
messages.reply(callerTerminal, content);
return text("delivered");
}
/** {@code bridge_ack}: acknowledge (remove) a specific reply from the inbox. */
static McpSchema.CallToolResult ack(MessageService messages, String target, String msgId) {
if (isBlank(target) || isBlank(msgId)) {
return error("target and msgId are required");
}
messages.ackReply(target, msgId);
return text("acknowledged " + msgId);
}
/** {@code bridge_status}: the live lifecycle status of a worker session. */
@@ -283,6 +465,49 @@ public final class BridgeMcp {
}
}
/**
* {@code bridge_whoami}: the caller's own identity, as the daemon already resolved it.
*
* <p>Every other tool <em>consumes</em> this identity — the authorization gate, the reply
* rendezvous, the cwd inherit — but none reported it, so an agent had to infer its own role
* from side channels the daemon does not control: a charter string in its system prompt, the
* name its MCP mount happens to carry, or {@code ANTHROPIC_BASE_URL} (which Claude-model
* workers do not set). The failure mode of guessing is asymmetric and silent: a primary that
* mistakes itself for a worker is refused by {@link Authz} and learns immediately, while a
* worker that mistakes itself for the primary ends its turn without {@code bridge_reply} and
* the sender simply receives nothing. This tool removes the guess.
*
* <p>For a worker the session registry adds what it knows about that session. A worker the
* registry has no record of — one that outlived a daemon restart — still gets its role and
* {@code sessionId}, which is the load-bearing part.
*/
static McpSchema.CallToolResult whoami(Principal caller, SessionManager sessions) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("role", caller.role().name().toLowerCase());
if (!caller.isWorker()) {
return text(json(m));
}
m.put("sessionId", caller.terminal());
sessions.roster().stream()
.filter(s -> caller.terminal().equals(s.terminalId()))
.findFirst()
.ifPresent(s -> {
m.put("paneId", s.paneId());
m.put("profile", s.profile());
m.put("state", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
if (s.ownerTerminal() != null) {
m.put("owner", s.ownerTerminal());
}
});
return text(json(m));
}
// --- fleet management logic (CB-108 / CB-301) --------------------------------------------
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
@@ -308,6 +533,8 @@ public final class BridgeMcp {
return error("subscription boundary: " + e.getMessage());
} catch (IllegalArgumentException e) {
return error(e.getMessage()); // unknown / no-default profile
} catch (PeerUnreachableException e) {
return error("spawn timed out — worker pane never reached injectable state: " + e.getMessage());
} catch (HerdrException e) {
return error("herdr error spawning worker: " + e.getMessage());
}
@@ -342,16 +569,17 @@ public final class BridgeMcp {
}
/** {@code bridge_profiles}: the configured worker profiles and the default. */
static McpSchema.CallToolResult profiles(ClaudeCodeLauncher workers) {
static McpSchema.CallToolResult profiles(PeerLauncher workers) {
return text(json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
}
/** {@code bridge_list}: bridge-owned roster merged with live herdr status by paneId. */
static McpSchema.CallToolResult listWorkers(ClaudeCodeLauncher workers, SessionManager sessions) {
static McpSchema.CallToolResult listWorkers(PeerLauncher workers, SessionManager sessions) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.paneId() != null)
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
@@ -435,10 +663,24 @@ public final class BridgeMcp {
private static McpSchema.Tool pollTool() {
return tool("bridge_poll",
"Check an async delegation (a bridge_send with wait:false) by its ticket: "
+ "pending, done (with the worker's reply), or failed.",
+ "pending, done (with the worker's reply), or failed. When target (a worker "
+ "session id) is present instead of ticket, drain that worker's inbox of "
+ "replies delivered when no send was open.",
objectSchema(Map.of(
"ticket", stringProp("The ticket returned by bridge_send wait:false")),
List.of("ticket")));
"ticket", stringProp("The ticket returned by bridge_send wait:false"),
"target", stringProp("Worker session id to drain pending replies from (optional)")),
List.of()));
}
private static McpSchema.Tool ackTool() {
return tool("bridge_ack",
"Acknowledge (remove) a specific reply from a worker's inbox. Use when the primary "
+ "has processed a reply and wants to confirm it, leaving other pending replies "
+ "in the inbox for later drain.",
objectSchema(Map.of(
"target", stringProp("Worker session id whose inbox to ack from"),
"msgId", stringProp("The message id to acknowledge")),
List.of("target", "msgId")));
}
private static McpSchema.Tool spawnTool() {
@@ -495,6 +737,18 @@ public final class BridgeMcp {
List.of("sessionId")));
}
private static McpSchema.Tool whoamiTool() {
return tool("bridge_whoami",
"Report who YOU are on the bridge — your role is resolved from your connection "
+ "(unforgeable), never from anything you claim. Returns role 'primary' (you "
+ "orchestrate: spawn/send/stop, and you must never call bridge_reply) or "
+ "'worker' (you were delegated to: you must end every turn with exactly one "
+ "bridge_reply, and cannot spawn or send), plus your own sessionId, profile, "
+ "worktree and branch when you are a worker. Call this first when following "
+ "role-conditional instructions rather than guessing your role.",
objectSchema(Map.of(), List.of()));
}
// --- small helpers -------------------------------------------------------------------------
// The SDK 2.0.0 deprecates its own Tool builders without a stable replacement — isolate it here.
@@ -0,0 +1,68 @@
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Optional;
import java.util.concurrent.atomic.AtomicReference;
/**
* Single-slot, thread-safe registry for the primary's herdr {@code terminal_id}.
*
* <p>Populated from the caller terminal of orchestration-side MCP tools
* ({@code bridge_send}, {@code bridge_spawn}) — tools that only the primary calls.
* A pinned terminal (from config) seeds the registry at construction and makes
* subsequent {@link #record(String)} calls no-ops.
*
* <p>The push loop ({@code ReplyPushLoop}) uses {@link #isKnown()} to decide
* whether active nudging is possible; an empty registry means the primary is
* off-host or non-herdr and delivery falls back to pull.
*/
public final class PrimaryRegistry {
private static final Logger log = LoggerFactory.getLogger(PrimaryRegistry.class);
private final AtomicReference<String> terminal = new AtomicReference<>();
private final boolean pinned;
/**
* @param pinnedTerminal an optional pinned terminal from config ({@code null}/blank = unpinned)
*/
public PrimaryRegistry(String pinnedTerminal) {
if (pinnedTerminal != null && !pinnedTerminal.isBlank()) {
this.terminal.set(pinnedTerminal);
this.pinned = true;
log.info("primary terminal pinned: {}", pinnedTerminal);
} else {
this.pinned = false;
}
}
/**
* Record a terminal_id. No-op when:
* <ul>
* <li>the registry is pinned (config override),
* <li>{@code terminalId} is {@code null} or blank (non-herdr caller).
* </ul>
*/
public void record(String terminalId) {
if (pinned) return;
if (terminalId == null || terminalId.isBlank()) return;
String prev = terminal.getAndSet(terminalId);
if (prev == null) {
log.debug("primary terminal learned: {}", terminalId);
} else if (!prev.equals(terminalId)) {
log.debug("primary terminal changed: {} -> {}", prev, terminalId);
}
}
/** The known primary terminal, or empty if not yet learned (and not pinned). */
public Optional<String> primaryTerminal() {
return Optional.ofNullable(terminal.get());
}
/** {@code true} once a terminal has been recorded (or was pinned at construction). */
public boolean isKnown() {
return terminal.get() != null;
}
}
@@ -0,0 +1,96 @@
package dev.ltms.bridged.metrics;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import java.util.LinkedHashMap;
import java.util.Map;
/**
* The daemon's metric definitions (CB-502) — one place where every series is named, described, and
* (for gauges) bound to live state.
*
* <p>The set is deliberately small: each series maps to a failure mode this project has actually
* hit, not to whatever was easy to count. The two worth watching in practice are
* {@code bridged_sends_total{outcome="completion_fallback"}} — a rising share means turn detection
* is degrading, the CB-115/116/118 failure family — and
* {@code bridged_push_nudges_total{outcome="exhausted"}}, which means the primary stopped draining
* its inbox and CB-307's active push gave up.
*/
public final class BridgedMetrics {
/** Counter: delegated sends by terminal outcome. */
public static final String SENDS = "bridged_sends_total";
/** Counter: worker replies by the path that carried them (rendezvous vs stranded-to-inbox). */
public static final String REPLIES = "bridged_replies_total";
/** Counter: push-loop nudges to the primary, by outcome. */
public static final String PUSH_NUDGES = "bridged_push_nudges_total";
/** Counter: spawn attempts by peer kind and outcome. */
public static final String SPAWNS = "bridged_spawns_total";
/** Counter: herdr socket calls by method and outcome. */
public static final String HERDR_CALLS = "bridged_herdr_calls_total";
/** Counter: rejected requests by reason (CB-501). */
public static final String AUTH_FAILURES = "bridged_auth_failures_total";
/** Gauge: session census by lifecycle state. */
public static final String SESSIONS = "bridged_sessions";
/** Gauge: undrained replies held per target. */
public static final String INBOX_DEPTH = "bridged_inbox_depth";
private BridgedMetrics() {
}
/**
* Build the registry with its help text and live gauges bound.
*
* @param sessions the authoritative session registry (census gauge)
* @param inbox the reply inbox; only used for a depth gauge when it can be inspected
*/
public static Metrics create(SessionManager sessions, ReplyInbox inbox) {
Metrics m = new Metrics();
m.describe(SENDS, "counter",
"Delegated sends by terminal outcome (replied|completion_fallback|timeout|failed).");
m.describe(REPLIES, "counter",
"Worker replies by delivery path (rendezvous=resolved an open send, inbox=stranded and held).");
m.describe(PUSH_NUDGES, "counter",
"CB-307 push-loop nudges to the primary (delivered|exhausted).");
m.describe(SPAWNS, "counter",
"Worker spawn attempts by peer kind and outcome (ready|timeout|guard_rejected).");
m.describe(HERDR_CALLS, "counter",
"herdr socket calls by method and outcome — the dependency everything else rests on.");
m.describe(AUTH_FAILURES, "counter",
"Requests refused by CB-501/505 (unauthenticated|forbidden).");
m.describe(SESSIONS, "gauge",
"Registered worker sessions by lifecycle state.");
m.describe(INBOX_DEPTH, "gauge",
"Replies held for a target that the primary has not drained. Steady state is 0; "
+ "a target stuck above 0 means CB-307 delivery is not completing.");
// One gauge per state so a scrape shows the whole census even when a state is empty —
// an absent series and a zero series read very differently on a dashboard.
for (WorkerSession.State state : WorkerSession.State.values()) {
String label = state.name().toLowerCase();
m.gauge(SESSIONS, () -> countIn(sessions, state), "state", label);
}
// Depth is per live session, so the label set is only known at scrape time. peek() is the
// port's non-destructive read — scraping metrics must never ack a reply out of the inbox.
m.collector(INBOX_DEPTH, "target", () -> {
Map<String, Number> depths = new LinkedHashMap<>();
for (WorkerSession s : sessions.roster()) {
String target = s.terminalId();
if (target == null) {
continue;
}
depths.put(target, inbox.peek(target).size());
}
return depths;
});
return m;
}
private static long countIn(SessionManager sessions, WorkerSession.State state) {
return sessions.roster().stream().filter(s -> s.state() == state).count();
}
}
@@ -0,0 +1,181 @@
package dev.ltms.bridged.metrics;
import java.util.Map;
import java.util.NavigableMap;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ConcurrentSkipListMap;
import java.util.concurrent.atomic.LongAdder;
import java.util.function.Supplier;
/**
* The daemon's metric registry and Prometheus text renderer (CB-502).
*
* <p>Deliberately dependency-free. The roadmap's tech-stack table specified Micrometer, but this
* pom already carries an unusual reconciliation burden (a hand-pinned {@code jackson-annotations}
* to make the MCP SDK's Jackson 3 coexist with our Jackson 2, a Jetty BOM import to stop version
* skew, and four documented accepted-CVE advisories), and the dependency CVE gate this project
* mandates could not be run when this landed. The metric set is small and fully known, and
* Prometheus text exposition is a stable, well-specified format — so the registry is ~100 lines
* here instead of a new transitive tree. {@code GET /metrics} is the swap seam if Micrometer's
* ecosystem is ever wanted.
*
* <p>Thread-safe: counters are {@link LongAdder} (built for contended increment), gauges are
* supplier-backed so they read live state at scrape time rather than needing to be pushed.
*/
public final class Metrics {
/** Counter series, keyed by the fully-rendered {@code name{labels}} sample id. */
private final NavigableMap<String, LongAdder> counters = new ConcurrentSkipListMap<>();
/** Gauge series, evaluated at scrape time. */
private final NavigableMap<String, Supplier<Number>> gauges = new ConcurrentSkipListMap<>();
/** Gauge families whose label set is only known at scrape time, keyed by metric name. */
private final NavigableMap<String, Collector> collectors = new ConcurrentSkipListMap<>();
/** HELP/TYPE metadata, keyed by bare metric name. */
private final Map<String, String[]> meta = new ConcurrentHashMap<>();
/** A gauge family whose series are discovered per scrape (one label, many values). */
private record Collector(String labelName, Supplier<Map<String, Number>> samples) {
}
/** Declare a metric's help text and type once, so the exposition carries HELP/TYPE lines. */
public Metrics describe(String name, String type, String help) {
meta.put(name, new String[]{type, help});
return this;
}
/** Increment a counter by one. */
public void inc(String name, String... labelPairs) {
add(name, 1, labelPairs);
}
/** Increment a counter by {@code delta}. */
public void add(String name, long delta, String... labelPairs) {
counters.computeIfAbsent(sample(name, labelPairs), _ -> new LongAdder()).add(delta);
}
/**
* Register a live gauge. The supplier is called at scrape time, so it reflects current state
* (session census, inbox depth) without anything having to remember to update it.
*/
public void gauge(String name, Supplier<Number> value, String... labelPairs) {
gauges.put(sample(name, labelPairs), value);
}
/**
* Register a gauge family whose label values are not known up front — inbox depth per target,
* for instance, where the set of targets changes as workers come and go. The supplier returns
* {@code labelValue → value} and is evaluated once per scrape.
*/
public void collector(String name, String labelName, Supplier<Map<String, Number>> samples) {
collectors.put(name, new Collector(labelName, samples));
}
/** Current value of a counter series — for assertions in tests. */
public long count(String name, String... labelPairs) {
LongAdder a = counters.get(sample(name, labelPairs));
return a == null ? 0 : a.sum();
}
/**
* Render the Prometheus text exposition format (version 0.0.4): optional {@code # HELP} and
* {@code # TYPE} lines per metric family, then one line per sample.
*/
public String render() {
StringBuilder out = new StringBuilder(1024);
String lastFamily = null;
for (Map.Entry<String, LongAdder> e : counters.entrySet()) {
lastFamily = emitHeader(out, e.getKey(), lastFamily);
out.append(e.getKey()).append(' ').append(e.getValue().sum()).append('\n');
}
for (Map.Entry<String, Supplier<Number>> e : gauges.entrySet()) {
lastFamily = emitHeader(out, e.getKey(), lastFamily);
Number v;
try {
v = e.getValue().get();
} catch (RuntimeException ex) {
continue; // a broken gauge must never break the whole scrape
}
if (v == null) {
continue;
}
out.append(e.getKey()).append(' ').append(format(v)).append('\n');
}
for (Map.Entry<String, Collector> e : collectors.entrySet()) {
Map<String, Number> samples;
try {
samples = e.getValue().samples().get();
} catch (RuntimeException ex) {
continue; // a broken collector must never break the whole scrape
}
if (samples == null || samples.isEmpty()) {
continue;
}
lastFamily = emitHeader(out, e.getKey(), lastFamily);
// Sort so repeated scrapes are byte-stable and diffable.
new java.util.TreeMap<>(samples).forEach((label, v) -> {
if (v != null) {
out.append(sample(e.getKey(), e.getValue().labelName(), label))
.append(' ').append(format(v)).append('\n');
}
});
}
return out.toString();
}
/** Emit HELP/TYPE when the sample starts a new metric family; returns the current family. */
private String emitHeader(StringBuilder out, String sampleId, String lastFamily) {
String family = familyOf(sampleId);
if (family.equals(lastFamily)) {
return lastFamily;
}
String[] m = meta.get(family);
if (m != null) {
out.append("# HELP ").append(family).append(' ').append(m[1]).append('\n');
out.append("# TYPE ").append(family).append(' ').append(m[0]).append('\n');
}
return family;
}
private static String familyOf(String sampleId) {
int brace = sampleId.indexOf('{');
return brace < 0 ? sampleId : sampleId.substring(0, brace);
}
/** Whole numbers render without a decimal point; everything else as-is. */
private static String format(Number v) {
double d = v.doubleValue();
return (d == Math.rint(d) && !Double.isInfinite(d))
? Long.toString((long) d)
: Double.toString(d);
}
/** Build the {@code name{k="v",k2="v2"}} sample id; labels are sorted for stable output. */
private static String sample(String name, String... labelPairs) {
if (labelPairs == null || labelPairs.length == 0) {
return name;
}
if (labelPairs.length % 2 != 0) {
throw new IllegalArgumentException("labels must be key/value pairs, got " + labelPairs.length);
}
NavigableMap<String, String> sorted = new java.util.TreeMap<>();
for (int i = 0; i < labelPairs.length; i += 2) {
sorted.put(labelPairs[i], labelPairs[i + 1] == null ? "" : labelPairs[i + 1]);
}
StringBuilder sb = new StringBuilder(name.length() + 16 * sorted.size());
sb.append(name).append('{');
boolean first = true;
for (Map.Entry<String, String> e : sorted.entrySet()) {
if (!first) {
sb.append(',');
}
first = false;
sb.append(e.getKey()).append("=\"").append(escapeLabel(e.getValue())).append('"');
}
return sb.append('}').toString();
}
/** Label values are escaped per the exposition format: backslash, quote, newline. */
private static String escapeLabel(String v) {
return v.replace("\\", "\\\\").replace("\"", "\\\"").replace("\n", "\\n");
}
}
@@ -0,0 +1,233 @@
package dev.ltms.bridged.msg;
import com.rabbitmq.client.AMQP;
import com.rabbitmq.client.Channel;
import com.rabbitmq.client.Connection;
import com.rabbitmq.client.ConnectionFactory;
import com.rabbitmq.client.DeliverCallback;
import com.rabbitmq.client.Recoverable;
import com.rabbitmq.client.RecoveryListener;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
/**
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
* port {@link InMemoryReplyInbox} implements as soft state.
*
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target owns a durable
* queue {@code agent.<target>.inbox}. A manual-ack consumer pulls persistent messages off that queue
* into an in-memory <em>held</em> map (keyed by {@code msgId}) but does <em>not</em> ack them.
* {@link #peek} returns that snapshot; {@link #ack} acks the broker delivery-tag and drops the entry.
* Because messages stay unacked until the primary actually drains them, a crash (or a {@code java -jar}
* bounce) before caller-ack leaves them on the broker — it redelivers on reconnect. That is the
* durability the in-memory adapter cannot give, with the port contract preserved.
*
* <p><strong>Dedup.</strong> The consumer keys the held map by {@code msgId}; a redelivered duplicate
* (at-least-once, or a producer double-publish) is acked-and-dropped on arrival, so it never
* double-queues. {@link #publish} additionally short-circuits an already-held {@code msgId} — a
* fast path; the consumer-side check is the real guarantee.
*
* <p><strong>Visibility.</strong> Unlike the in-memory adapter, publish → broker → consumer is
* asynchronous, so a {@link #peek} immediately after {@link #publish} may not yet see the message
* (broker delivery latency). Callers that need the reply drained poll (as the primary already does);
* the contract test waits for visibility. This is inherent to broker-backed delivery, not a defect.
*
* <p>The default deploy targets LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1 (URI-only swap),
* so the {@code @Tag("contract")} integration test runs against a RabbitMQ container.
*/
public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
private static final Logger log = LoggerFactory.getLogger(AmqpReplyInbox.class);
private static final String QUEUE_PREFIX = "agent.";
private static final String QUEUE_SUFFIX = ".inbox";
private final Connection connection;
private final Channel channel;
/** All channel operations (publish/declare/ack) serialize on this — a Channel is not thread-safe. */
private final Object channelLock = new Object();
/** target → (msgId → held delivery). Per-target map is guarded by synchronizing on itself. */
private final ConcurrentHashMap<String, LinkedHashMap<String, Held>> held = new ConcurrentHashMap<>();
/** Targets whose queue is declared and consumer is running. */
private final Set<String> consuming = ConcurrentHashMap.newKeySet();
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
private record Held(long deliveryTag, InboxMessage message) {}
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) and open the inbox. */
public static AmqpReplyInbox open(String uri) {
try {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"));
} catch (Exception e) {
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
}
}
/** Wrap an already-open connection (injection seam for the contract test). */
AmqpReplyInbox(Connection connection) {
this.connection = connection;
try {
this.channel = connection.createChannel();
} catch (IOException e) {
throw new IllegalStateException("cannot open AMQP channel", e);
}
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
// repopulates it with valid tags (dedup by msgId still prevents any double-queue).
if (connection instanceof Recoverable recoverable) {
recoverable.addRecoveryListener(new RecoveryListener() {
@Override
public void handleRecovery(Recoverable recoverable) {
held.clear();
log.info("AMQP connection recovered; cleared held replies for fresh redelivery");
}
@Override
public void handleRecoveryStarted(Recoverable recoverable) {
// no-op: we act once recovery completes
}
});
}
}
@Override
public void publish(String target, String msgId, String content) {
ensureConsuming(target);
var perTarget = held.get(target);
if (perTarget != null) {
synchronized (perTarget) {
if (perTarget.containsKey(msgId)) {
return; // already held — producer-side fast dedup
}
}
}
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
.messageId(msgId)
.deliveryMode(2) // persistent — survives a broker restart
.contentType("text/plain")
.build();
try {
synchronized (channelLock) {
channel.basicPublish("", queueName(target), props, content.getBytes(StandardCharsets.UTF_8));
}
} catch (IOException e) {
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
}
}
@Override
public List<InboxMessage> peek(String target) {
ensureConsuming(target);
var perTarget = held.get(target);
if (perTarget == null) {
return List.of();
}
synchronized (perTarget) {
return perTarget.values().stream().map(Held::message).toList();
}
}
@Override
public void ack(String target, String msgId) {
var perTarget = held.get(target);
if (perTarget == null) {
return;
}
Held h;
synchronized (perTarget) {
h = perTarget.remove(msgId);
}
if (h == null) {
return; // never held (or already acked) — no-op
}
try {
synchronized (channelLock) {
channel.basicAck(h.deliveryTag(), false);
}
} catch (IOException e) {
// Ack didn't reach the broker: restore the entry so a later ack (or a redelivery after
// reconnect) can retry. Keeps the at-least-once contract — a reply is never silently lost.
synchronized (perTarget) {
perTarget.putIfAbsent(msgId, h);
}
throw new IllegalStateException("cannot ack reply " + msgId + " on " + queueName(target), e);
}
}
/** Declare the durable per-target queue and start its manual-ack consumer, once per target. */
private void ensureConsuming(String target) {
if (consuming.contains(target)) {
return;
}
synchronized (channelLock) {
if (!consuming.add(target)) {
return; // another thread just set it up
}
String queue = queueName(target);
try {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
} catch (IOException e) {
consuming.remove(target);
throw new IllegalStateException("cannot consume queue " + queue, e);
}
}
}
private DeliverCallback deliverCallback(String target) {
return (_, delivery) -> {
String msgId = delivery.getProperties().getMessageId();
long tag = delivery.getEnvelope().getDeliveryTag();
if (msgId == null || msgId.isBlank()) {
msgId = Long.toHexString(tag); // synthesize an id so dedup still has a key
}
String content = new String(delivery.getBody(), StandardCharsets.UTF_8);
var perTarget = held.computeIfAbsent(target, _ -> new LinkedHashMap<>());
boolean duplicate;
synchronized (perTarget) {
if (perTarget.containsKey(msgId)) {
duplicate = true;
} else {
perTarget.put(msgId, new Held(tag, new InboxMessage(msgId, target, content)));
duplicate = false;
}
}
if (duplicate) {
// Redelivered duplicate: ack the new tag and drop it so the broker stops resending.
synchronized (channelLock) {
channel.basicAck(tag, false);
}
}
};
}
private static String queueName(String target) {
return QUEUE_PREFIX + target + QUEUE_SUFFIX;
}
@Override
public void close() {
try {
channel.close();
} catch (Exception e) {
log.debug("AMQP channel close: {}", e.toString());
}
try {
connection.close();
} catch (Exception e) {
log.debug("AMQP connection close: {}", e.toString());
}
}
}
@@ -0,0 +1,50 @@
package dev.ltms.bridged.msg;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.concurrent.ConcurrentHashMap;
/**
* Soft-state {@link ReplyInbox} backed by a {@link ConcurrentHashMap} keyed by target session.
* Per-target FIFO ordering (insertion order via {@link LinkedHashMap}). Dedup by {@code msgId}
* within a target. Thread-safe for concurrent publish vs. drain.
*
* <p><strong>This is soft-state, NOT persistence.</strong> Lost on a {@code java -jar} bounce — that
* is correct and consistent with "bridged stays soft-state." The Stage-2 AMQP adapter replaces this.
*/
public final class InMemoryReplyInbox implements ReplyInbox {
private final ConcurrentHashMap<String, LinkedHashMap<String, InboxMessage>> store = new ConcurrentHashMap<>();
@Override
public void publish(String target, String msgId, String content) {
var perTarget = store.computeIfAbsent(target, _ -> new LinkedHashMap<>());
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
perTarget.putIfAbsent(msgId, new InboxMessage(msgId, target, content));
}
}
@Override
public List<InboxMessage> peek(String target) {
var perTarget = store.get(target);
if (perTarget == null) {
return List.of();
}
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
return List.copyOf(perTarget.values());
}
}
@Override
public void ack(String target, String msgId) {
var perTarget = store.get(target);
if (perTarget != null) {
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
perTarget.remove(msgId);
}
}
}
}
@@ -3,9 +3,13 @@ package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.UUID;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.CompletionException;
import java.util.concurrent.ConcurrentHashMap;
@@ -147,16 +151,50 @@ public final class MessageService {
private final AgentControl agents;
private final Injector injector;
private final Rendezvous rendezvous;
private final ReplyInbox inbox;
private final ReplyPushLoop pushLoop;
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
Thread.ofVirtual().name("bridge-async-", 0).factory());
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous) {
/**
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
*
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
* branch ({@link #reply}) so it can nudge the primary to drain the inbox
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop) {
this(agents, injector, rendezvous, inbox, pushLoop, null);
}
/**
* As above, with a metric registry (CB-502). Instrumenting here rather than at the REST and MCP
* edges means both surfaces are counted by one piece of code and cannot drift.
*
* @param metrics nullable — when null, nothing is recorded
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop, Metrics metrics) {
this.agents = agents;
this.injector = injector;
this.rendezvous = rendezvous;
this.inbox = inbox;
this.pushLoop = pushLoop;
this.metrics = metrics;
}
/** Create with an explicit {@link ReplyInbox} and no push loop. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox) {
this(agents, injector, rendezvous, inbox, null);
}
/** Backward-compatible constructor that uses a default {@link InMemoryReplyInbox}. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous) {
this(agents, injector, rendezvous, new InMemoryReplyInbox());
}
/** Current lifecycle status of a worker (the {@code GET /sessions/{id}/status} surface). */
@@ -164,6 +202,109 @@ public final class MessageService {
return agents.status(target);
}
/**
* Route a worker's explicit {@code bridge_reply}: resolve an open send, or queue it in the
* inbox if no send is currently open. Unlike the bare {@link Rendezvous#resolve}, a no-waiter
* result is <em>not</em> a failure — the reply is held for later drain.
*
* <p><strong>Do NOT use this for mid-turn questions.</strong> {@code bridge_ask} /
* {@link Rendezvous#resolveQuestion} must keep today's {@code NO_WAITER} behaviour — questions
* are interactive and must never be queued.
*
* @return always {@code true} — the reply either resolved a live send or was queued
*/
public boolean reply(String session, String content) {
if (rendezvous.resolve(session, content)) {
count(BridgedMetrics.REPLIES, "path", "rendezvous");
return true; // a live send took it — unchanged fast path
}
inbox.publish(session, UUID.randomUUID().toString(), content);
// A rising inbox share is the signal CB-307 exists to make visible: the worker finished but
// nobody was waiting, so delivery now depends on the push loop and a drain.
count(BridgedMetrics.REPLIES, "path", "inbox");
if (pushLoop != null) {
pushLoop.onReplyQueued(session);
}
return true; // held, not lost
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
if (metrics != null) {
metrics.inc(name, labels);
}
}
/** Count a send's terminal outcome and pass the reply through unchanged. */
private Reply recorded(Reply r) {
String label = sendOutcomeLabel(r.outcome());
if (label != null) {
count(BridgedMetrics.SENDS, "outcome", label);
}
return r;
}
/** Map a terminal send outcome to its metric label, or {@code null} for non-terminal ones. */
private static String sendOutcomeLabel(Outcome o) {
return switch (o) {
case REPLIED -> "replied";
case COMPLETED_UNREPLIED -> "completion_fallback";
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> "timeout";
case WORKER_FAILED -> "failed";
case STALE_TURN, QUESTION -> null; // not a completed delegation
};
}
/**
* Abandon any send still waiting on {@code target} because its session has gone away (CB-516).
*
* <p>Without this, tearing a worker down left its rendezvous waiter open: a blocking
* {@code bridge_send} kept blocking, and an async one kept reporting {@code PENDING} until
* {@link #ASYNC_TIMEOUT_MS} — thirty minutes — even though the worker provably no longer
* existed and the delegation could never complete. Worse, {@code poll} already had the evidence
* (it calls {@code liveStatus} to build its detail string and gets back {@code "unknown"}) and
* reported {@code PENDING} anyway.
*
* <p>Resolving the waiter as a failure — rather than letting it time out — also means the
* outcome is counted, so a torn-down delegation stops being invisible to {@code /metrics}.
*
* @return true if a live waiter was failed
*/
public boolean abandon(String target, String reason) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
return false; // nobody is blocked on this worker — nothing to abandon
}
boolean failed = rendezvous.resolveFailure(waiter, reason);
if (failed) {
log.debug("abandoned send to {}: {}", target, reason);
}
return failed;
}
/**
* Acknowledge a specific reply by {@code msgId} for {@code target}. Removes it from the inbox
* so that a subsequent drain or peek no longer returns it.
*/
public void ackReply(String target, String msgId) {
inbox.ack(target, msgId);
}
/**
* Drain (peek + ack) all pending inbox replies for {@code target}. At-least-once: returns the
* messages and acknowledges them; an in-flight failure between returning and the caller
* processing them re-surfaces them on a subsequent drain (the ack is local).
*
* @return the drained messages, newest last (FIFO); empty list if none
*/
public List<ReplyInbox.InboxMessage> drainReplies(String target) {
var messages = inbox.peek(target);
for (var msg : messages) {
inbox.ack(target, msg.msgId());
}
return messages;
}
/**
* Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
@@ -180,11 +321,12 @@ public final class MessageService {
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
return new Reply(wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null);
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
@@ -0,0 +1,31 @@
package dev.ltms.bridged.msg;
import java.util.List;
/**
* Holds terminal worker→primary replies that arrive with no live send to resolve, keyed by worker
* session (target), until the primary drains them. Soft-state in Stage 1 (in-memory, lost on restart);
* the Stage 2 AMQP adapter implements the same contract with cross-restart durability.
*
* <p><strong>This interface is the port.</strong> {@link InMemoryReplyInbox} is the Stage-1 adapter;
* an AMQP-backed adapter (Stage 2) must implement the same contract (idempotent publish, FIFO peek,
* at-least-once ack).
*/
public interface ReplyInbox {
/** A queued reply: an idempotency id, the worker session it came from, and the reply text. */
record InboxMessage(String msgId, String target, String content) {}
/**
* Queue {@code content} from worker {@code target} under {@code msgId}. Idempotent: publishing an
* already-present {@code msgId} for {@code target} is a no-op (dedup), so an at-least-once Stage-2
* redelivery cannot double-queue.
*/
void publish(String target, String msgId, String content);
/** Non-destructive snapshot of pending replies for {@code target} (FIFO), empty list if none. */
List<InboxMessage> peek(String target);
/** Remove the reply {@code msgId} for {@code target} once the primary has taken it. No-op if absent. */
void ack(String target, String msgId);
}
@@ -0,0 +1,176 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
/**
* Mechanism (b) of CB-307: a dedicated, status-gated push loop that nudges the primary's own
* herdr pane when a worker reply lands with no live {@code bridge_send} to resolve it.
*
* <p>The loop is triggered by {@link #onReplyQueued(String)} (called from
* {@link MessageService#reply} after the durable inbox publish). It checks four conditions
* at each tick via {@link #decide(String, int)}, then either injects a drain nudge,
* waits for the primary to become injectable, or stops reminding.
*
* <p>Bounded: at most {@link #maxReminders} nudges per target, with a configurable backoff
* between them. The reply is never lost — the durable inbox is the backstop.
*/
public final class ReplyPushLoop {
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
private final ReplyInbox inbox;
private final ScheduledExecutorService scheduler;
private final int maxReminders;
private final long backoffMs;
private final Metrics metrics; // CB-512: nullable — no registry in unit tests
/** Track targets that have an active schedule. */
private final ConcurrentHashMap<String, Boolean> activeTargets = new ConcurrentHashMap<>();
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
int maxReminders, long backoffMs) {
this(primaryRegistry, agents, inbox, scheduler, maxReminders, backoffMs, null);
}
/** As above, with a metric registry (CB-512) so push outcomes are counted. */
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
int maxReminders, long backoffMs, Metrics metrics) {
this.primaryRegistry = primaryRegistry;
this.agents = agents;
this.inbox = inbox;
this.scheduler = scheduler;
this.maxReminders = maxReminders;
this.backoffMs = backoffMs;
this.metrics = metrics;
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
if (metrics != null) {
metrics.inc(name, labels);
}
}
// --- decision logic (package-private for unit-testing) -------------------------------------
/** The action the loop should take for a target at the given reminder count. */
enum Action { INJECT, WAIT_BUSY, STOP }
/**
* Pure decision function: examine the current state and return what the loop should do.
*
* @param target the worker session (target terminal id)
* @param reminderCount how many nudges have been sent so far for this target
* @return the action the caller should take
*/
Action decide(String target, int reminderCount) {
if (!primaryRegistry.isKnown()) {
log.debug("push: primary unknown, stopping reminder for {}", target);
return Action.STOP;
}
if (inbox.peek(target).isEmpty()) {
log.debug("push: inbox empty for {}, stopping reminder", target);
return Action.STOP;
}
if (reminderCount >= maxReminders) {
log.debug("push: reminder cap ({}) reached for {}, stopping", maxReminders, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted");
return Action.STOP;
}
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
AgentStatus status;
try {
status = agents.status(primaryTerminal);
} catch (RuntimeException e) {
log.debug("push: status check failed for primary {}, will retry", primaryTerminal, e);
return Action.WAIT_BUSY;
}
if (status.injectable()) {
return Action.INJECT;
}
log.debug("push: primary {} is {} (not injectable), waiting", primaryTerminal, status);
return Action.WAIT_BUSY;
}
// --- public entrypoint ---------------------------------------------------------------------
/**
* Called when a reply is queued for {@code target}. Idempotent per target: a second call while
* a schedule is active is a no-op. The schedule nudges the primary, then schedules a follow-up
* check (reminder on backoff, or re-check on WAIT_BUSY), until the inbox is empty or the cap
* is reached.
*/
public void onReplyQueued(String target) {
if (activeTargets.putIfAbsent(target, Boolean.TRUE) != null) {
log.debug("push: already active for {}, ignoring duplicate trigger", target);
return; // already scheduled
}
log.debug("push: starting reminder loop for {}", target);
scheduleNext(target, 0);
}
/** Execute one loop tick — called on the scheduler thread. */
private void tick(String target, int reminderCount) {
var action = decide(target, reminderCount);
switch (action) {
case INJECT -> {
injectNudge(target, reminderCount);
scheduleNext(target, reminderCount + 1);
}
// Re-check after the configured backoff; the primary may become injectable soon.
case WAIT_BUSY -> scheduleNext(target, reminderCount);
case STOP -> {
activeTargets.remove(target);
log.debug("push: reminder loop ended for {}", target);
}
}
}
/** Send the nudge and log the event. */
private void injectNudge(String target, int reminderCount) {
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
String nudge = NUDGE_FORMAT.formatted(target, target);
try {
agents.send(primaryTerminal, nudge);
log.debug("push: nudge {}/{} sent to primary {} for target {}",
reminderCount + 1, maxReminders, primaryTerminal, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered");
} catch (RuntimeException e) {
log.warn("push: failed to nudge primary {} for target {} (reminder {}/{}): {}",
primaryTerminal, target, reminderCount + 1, maxReminders, e.toString());
}
}
/** Schedule the next tick on the scheduler thread pool. */
private void scheduleNext(String target, int nextReminderCount) {
scheduler.schedule(() -> tick(target, nextReminderCount), backoffMs, TimeUnit.MILLISECONDS);
}
// --- lifecycle -----------------------------------------------------------------------------
/** Shut down the scheduler. Outstanding reminders are cancelled. */
public void stop() {
scheduler.shutdownNow();
activeTargets.clear();
}
/** @see #stop() */
public void close() {
stop();
}
}
@@ -0,0 +1,18 @@
package dev.ltms.bridged.peer;
/**
* Thrown when a {@link PeerLauncher} starts a peer process but the peer
* does not reach an injectable (ready-to-receive) state within the configured
* timeout. The launcher MUST clean up any resources it created (pane, tab)
* before throwing — no orphaned peer or pane is left behind.
*
* <p>This is a spawn-time failure, distinct from a post-spawn disconnect.
* Callers treat this as a clean spawn error (the peer never materialized
* into a usable session), not a mid-life session fault.
*/
public final class PeerUnreachableException extends RuntimeException {
public PeerUnreachableException(String message) {
super(message);
}
}
@@ -2,17 +2,22 @@ package dev.ltms.bridged.rest;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.bridged.auth.AuditLog;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.peer.PeerLauncher;
import io.javalin.Javalin;
import io.javalin.http.Context;
import jakarta.servlet.http.HttpServlet;
@@ -43,25 +48,47 @@ public final class BridgedApp {
private static final long DEFAULT_ASK_TIMEOUT_MS = 55_000;
private static final long MAX_ASK_TIMEOUT_MS = 115_000;
/** Context attribute under which the resolved caller is stashed by the auth filter. */
private static final String CALLER = "bridged.caller";
private final HerdrClient herdr;
private final ClaudeCodeLauncher workers;
private final PeerLauncher workers;
private final SessionManager sessions; // CB-301: authoritative session registry
private final MessageService messages;
private final Rendezvous rendezvous;
private final WorkerPresence presence; // CB-113: which workers are MCP-connected (available)
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
private final CallerResolver auth; // CB-501: null → authz not enforced (legacy behaviour)
private final Metrics metrics; // CB-502: null → /metrics not exposed
private final ObjectMapper mapper = new ObjectMapper();
public BridgedApp(HerdrClient herdr, ClaudeCodeLauncher workers, SessionManager sessions,
MessageService messages, Rendezvous rendezvous, WorkerPresence presence,
/**
* Legacy constructor — no identity resolution and no authorization, exactly as the REST surface
* behaved before CB-501. Retained so existing acceptance tests keep exercising handler
* behaviour without each needing an auth fixture.
*/
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
HttpServlet mcpServlet) {
this(herdr, workers, sessions, messages, presence, mcpServlet, null, null);
}
/**
* @param auth resolves each request's {@link Principal}; {@code null} disables authorization
* entirely (legacy). {@code main} always supplies one.
* @param metrics registry to instrument and expose at {@code GET /metrics}; {@code null} omits
* the endpoint
*/
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics) {
this.herdr = herdr;
this.workers = workers;
this.sessions = sessions;
this.messages = messages;
this.rendezvous = rendezvous;
this.presence = presence;
this.mcpServlet = mcpServlet;
this.auth = auth;
this.metrics = metrics;
}
/** Wire routes onto a fresh, unstarted Javalin instance. Caller starts it. */
@@ -74,7 +101,18 @@ public final class BridgedApp {
h.addServlet(new ServletHolder(mcpServlet), "/mcp"));
}
});
// CB-501: resolve identity once per request, before any handler. /mcp does NOT pass through
// here — it is a raw servlet on Jetty's context handler — so BridgeMcp enforces separately
// against the same CallerResolver. Any check that lives in only one place is not a control.
if (auth != null) {
app.before(ctx -> ctx.attribute(CALLER,
auth.resolve(ctx.req().getRemoteAddr(), ctx.req().getRemotePort(),
ctx.header("Authorization"))));
}
app.get("/healthz", this::healthz);
if (metrics != null) {
app.get("/metrics", this::metrics);
}
app.get("/sessions", this::sessions);
app.get("/agents", this::agents);
app.get("/workers", this::listWorkers); // CB-304: registry roster + live herdr status
@@ -83,12 +121,61 @@ public final class BridgedApp {
app.delete("/workers/{paneId}", this::stopWorker);
app.post("/sessions/{id}/message", this::sendMessage); // bridge_send (primary; blocking, wait:false, or answer via turnId)
app.post("/sessions/{id}/reply", this::replyMessage); // bridge_reply (worker)
app.get("/sessions/{id}/replies", this::drainReplies); // drain reply inbox (CB-307)
app.post("/sessions/{id}/ask", this::askMessage); // bridge_ask (worker → primary, CB-205)
app.get("/sessions/{id}/status", this::sessionStatus); // bridge_status
app.get("/tasks/{ticket}", this::taskStatus); // poll an async (wait:false) send
return app;
}
/**
* Gate a handler on the CB-505 authorization table. Returns {@code true} when the request may
* proceed; otherwise writes the error response and returns {@code false}.
*
* <p>401 vs 403 is a real distinction here: 401 means "you presented no usable identity" (a
* credential problem the caller can fix), 403 means "you are authenticated, but this is not
* yours" (a worker reaching for another worker's session, or for orchestration).
*/
private boolean allow(Context ctx, Authz.Action action, String target) {
if (auth == null) {
return true; // legacy: authorization not enforced
}
Principal caller = ctx.attribute(CALLER);
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ && action != Authz.Action.METRICS) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
return true;
}
if (Authz.isUnauthenticated(caller)) {
AuditLog.denied(caller, action, target, "unauthenticated");
countAuthFailure("unauthenticated");
ctx.status(401).json(Map.of("error", "unauthenticated",
"detail", "present Authorization: Bearer <token>"));
} else {
AuditLog.denied(caller, action, target, "forbidden");
countAuthFailure("forbidden");
ctx.status(403).json(Map.of("error", "forbidden",
"detail", caller.describe() + " may not " + action + " on "
+ (target == null ? "this resource" : target)));
}
return false;
}
private void countAuthFailure(String reason) {
if (metrics != null) {
metrics.inc("bridged_auth_failures_total", "reason", reason);
}
}
/** Prometheus scrape endpoint (CB-502). */
private void metrics(Context ctx) {
if (!allow(ctx, Authz.Action.METRICS, null)) {
return;
}
ctx.status(200).contentType("text/plain; version=0.0.4; charset=utf-8").result(metrics.render());
}
/** Liveness + herdr reachability. 200 when herdr answers ping, 503 otherwise. */
private void healthz(Context ctx) {
try {
@@ -108,6 +195,9 @@ public final class BridgedApp {
/** Sessions view derived from herdr {@code workspace.list} (one workspace → one row). */
private void sessions(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
JsonNode result = herdr.call("workspace.list");
List<Map<String, Object>> out = new ArrayList<>();
for (JsonNode w : result.path("workspaces")) {
@@ -123,12 +213,20 @@ public final class BridgedApp {
/** Discovery: every agent herdr tracks, keyed by its Claude session UUID. */
private void agents(Context ctx) {
ctx.status(200).json(Map.of("agents", workers.list().stream().map(BridgedApp::view).toList()));
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
ctx.status(200).json(Map.of("agents",
workers.list().stream().map(Agent.class::cast).map(BridgedApp::view).toList()));
}
/** CB-304: bridge-owned roster merged with live herdr status by paneId. */
private void listWorkers(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.paneId() != null)
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
@@ -139,6 +237,9 @@ public final class BridgedApp {
/** The configured worker profiles and which one a no-argument spawn uses. */
private void profiles(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
ctx.status(200).json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile()));
@@ -150,6 +251,9 @@ public final class BridgedApp {
* the subscription boundary, 400 for an unknown profile.
*/
private void spawnWorker(Context ctx) {
if (!allow(ctx, Authz.Action.SPAWN, null)) {
return;
}
String profile = ctx.queryParam("profile");
String cwd = ctx.queryParam("cwd");
String worktree = ctx.queryParam("worktree");
@@ -178,6 +282,8 @@ public final class BridgedApp {
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
} catch (IllegalArgumentException e) {
ctx.status(400).json(Map.of("error", "unknown_profile", "detail", e.getMessage()));
} catch (PeerUnreachableException e) {
ctx.status(502).json(Map.of("error", "spawn_timeout", "detail", e.getMessage()));
}
}
@@ -200,7 +306,11 @@ public final class BridgedApp {
/** Tear a worker down by pane id. */
private void stopWorker(Context ctx) {
sessions.release(ctx.pathParam("paneId"));
String paneId = ctx.pathParam("paneId");
if (!allow(ctx, Authz.Action.STOP, paneId)) {
return;
}
sessions.release(paneId);
ctx.status(204);
}
@@ -212,6 +322,9 @@ public final class BridgedApp {
*/
private void sendMessage(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.SEND, id)) {
return;
}
String content;
String turnId;
long timeout;
@@ -293,6 +406,9 @@ public final class BridgedApp {
*/
private void askMessage(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.ASK, id)) {
return;
}
String question;
long timeout;
try {
@@ -322,10 +438,16 @@ public final class BridgedApp {
/**
* The worker's structured reply ({@code bridge_reply}) — resolves the blocking send awaiting
* on this session. 200 if a send was waiting, 409 if none was (late or spurious reply).
* on this session, or queues the reply in the inbox when no send is open (CB-307).
*/
private void replyMessage(Context ctx) {
String id = ctx.pathParam("id");
// The rule that matters: a worker may reply only as itself. Over MCP this was already true
// structurally (identity comes from the connection, never an argument); over REST the path
// id was simply trusted, so this is where the invariant actually gets enforced.
if (!allow(ctx, Authz.Action.REPLY, id)) {
return;
}
String content;
try {
content = mapper.readTree(ctx.body()).path("content").asText("");
@@ -333,13 +455,25 @@ public final class BridgedApp {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
if (rendezvous.resolve(id, content)) {
ctx.status(200).json(Map.of("sessionId", id, "delivered", true));
} else {
ctx.status(409).json(Map.of(
"sessionId", id, "error", "no_pending_send",
"detail", "no send is awaiting a reply for this session"));
messages.reply(id, content);
ctx.status(200).json(Map.of("sessionId", id, "delivered", true));
}
/**
* Drain the reply inbox for a worker session — peek + ack any replies that arrived when no send
* was open. At-least-once: draining removes them from the inbox so a subsequent read returns
* nothing; an in-flight failure between the drain and the caller's processing re-surfaces them.
*/
private void drainReplies(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.DRAIN, id)) {
return;
}
var replies = messages.drainReplies(id);
ctx.status(200).json(Map.of("sessionId", id, "replies",
replies.stream().map(m -> Map.of(
"msgId", m.msgId(),
"content", m.content())).toList()));
}
/**
@@ -350,6 +484,9 @@ public final class BridgedApp {
*/
private void sessionStatus(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.READ, id)) {
return;
}
try {
ctx.status(200).json(Map.of(
"sessionId", id,
@@ -362,6 +499,9 @@ public final class BridgedApp {
/** Poll an async (wait:false) delegation by ticket. 404 for an unknown/expired ticket. */
private void taskStatus(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"));
if (v == null) {
ctx.status(404).json(Map.of("error", "unknown_ticket", "detail", "no such task (or it has expired)"));
@@ -17,6 +17,7 @@ import java.util.Optional;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Consumer;
import java.util.function.LongSupplier;
/**
@@ -47,6 +48,9 @@ public final class SessionManager implements TurnListener {
private final LongSupplier nowNanos;
private final int contextCap;
/** CB-516: notified with a terminalId on every release; no-op until wired. */
private volatile Consumer<String> releaseListener = _ -> { };
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
public SessionManager(PeerLauncher launcher) {
this(launcher, new GitWorktrees(), System::nanoTime, 0);
@@ -136,6 +140,10 @@ public final class SessionManager implements TurnListener {
if (removed != null) {
log.debug("releasing session pane={} terminal={} state={}",
removed.paneId(), removed.terminalId(), removed.state());
// CB-516: a send still waiting on this worker can never be answered now. Tell the
// listener BEFORE the pane is torn down, so a blocked caller fails fast with a real
// reason instead of sitting on a rendezvous nothing will ever resolve.
notifyReleased(removed.terminalId());
}
launcher.stop(paneId);
if (removed != null && removed.worktree() != null) {
@@ -143,11 +151,43 @@ public final class SessionManager implements TurnListener {
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is released
* (CB-516). Every teardown path funnels through {@link #release}, so one hook covers the REST
* and MCP stop tools, the idle-TTL reaper, {@code recycle}, and shutdown drain alike.
*
* <p>Set rather than injected because {@code MessageService} — the intended listener — is
* constructed after this manager (it needs the injector and rendezvous, which need the session
* presence view this manager exposes). Wiring it at construction would require breaking that
* cycle for one callback.
*/
public void onRelease(Consumer<String> listener) {
this.releaseListener = (listener == null) ? _ -> { } : listener;
}
/** A listener failure must never prevent the teardown it is reacting to. */
private void notifyReleased(String terminalId) {
if (terminalId == null) {
return;
}
try {
releaseListener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
}
}
private WorkerSession acquireWithWorktree(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
String resolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
String repoRoot = worktrees.repoRoot(firstNonBlank(requestedCwd, callerCwd));
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
// nor a caller cwd, so taking the first non-blank of those two yielded null and put
// `git -C null` on the command line — an NPE out of ProcessBuilder, surfacing as HTTP 500.
// The non-worktree path always used this chain; only this branch was missed.
String repoRoot = worktrees.repoRoot(
launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd)));
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
String path = null;
PeerHandle handle;
@@ -192,13 +232,6 @@ public final class SessionManager implements TurnListener {
return String.format("%06x", nonceRandom.nextInt(1 << 24)) + "-" + nonceSeq.incrementAndGet();
}
private static String firstNonBlank(String... values) {
for (String v : values) {
if (v != null && !v.isBlank()) return v;
}
return null;
}
/**
* Release the old session and acquire a fresh one with the same profile and working directory.
* The new session is guaranteed to have a pane id distinct from the old one (no-reuse invariant).
@@ -4,66 +4,39 @@ import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.EnumSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
import java.util.function.LongSupplier;
/**
* Spawns and lists worker sessions — the safe path from a delegation request to a
* running off-subscription Claude.
* The {@link HerdrPeerLauncher} adapter for <strong>Claude Code</strong> — the safe path from a
* delegation request to a running off-subscription Claude.
*
* <p>The spawn sequence encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em>
* touching herdr, and only then {@code agent.start}. A worker's base_url lives in the
* env map handed to herdr and nowhere else; {@code bridged}'s own environment is never
* mutated.
*
* <p>Placement: in the default {@code tab} policy a worker lands in its own tab inside a
* dedicated worker space (found-or-created once, then shared), so workers never split or
* clutter the user's real work spaces. Teardown removes the worker's pane <em>and</em> its
* now-empty tab, tolerating an already-gone worker so a repeated DELETE is harmless.
* <p>Everything transport-related (tab/pane placement, the CB-306 spawn-readiness gate, unique
* naming, CB-117 orphan reap, teardown, listing, cwd resolution) lives in the base. This class
* supplies only the two Claude-specific seams:
* <ul>
* <li>the {@code claude} name prefix (so reap matches {@code claude-*} panes, never another
* adapter's), and</li>
* <li>{@link #buildLaunch}, which encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em> touching
* herdr, and mount the bridge MCP + reply charter as inline launch flags. A worker's base_url
* lives in the env map handed to herdr and nowhere else; {@code bridged}'s own environment is
* never mutated, and nothing is written to the worker's profile.</li>
* </ul>
*/
public final class ClaudeCodeLauncher implements PeerLauncher {
public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
private static final Logger log = LoggerFactory.getLogger(ClaudeCodeLauncher.class);
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "claude";
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
/**
* A bridge-spawned worker label {@code claude-<profile>-<nonce>-<seq>} (see
* {@link #startUniquelyNamed}); group 1 captures the 6-hex per-process {@code nonce}. The
* profile segment may itself contain {@code -}, so the nonce/seq are anchored at the tail.
* Names not matching this shape are not workers we started and are never reaped (CB-117).
*/
private static final Pattern WORKER_NAME = Pattern.compile("claude-.*-([0-9a-f]{6})-\\d+");
private final AgentControl agents;
private final WorkspaceControl spaces;
private final SubscriptionGuard guard;
private final Map<String, BridgedConfig.Worker> profiles; // profile name → spawn settings
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
private final Function<String, String> env; // host env lookup (injectable for tests)
private final AtomicLong nameSeq = new AtomicLong(); // per-worker counter (also the tab #)
/**
* Standing instruction appended to the worker's system prompt so it returns its result via
@@ -82,144 +55,85 @@ public final class ClaudeCodeLauncher implements PeerLauncher {
+ "`content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
// Per-process token mixed into each worker name so a fresh process (nameSeq back at 0)
// cannot collide with same-profile workers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so
* existing deployments and tests keep the legacy non-blocking spawn semantics.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this.agents = agents;
this.spaces = spaces;
this.guard = guard;
this.profiles = Map.copyOf(profiles);
this.defaultProfile = defaultProfile;
this.env = env;
}
/** The configured worker profile names (what {@code spawn(profile)} accepts). */
@Override
public Set<String> profiles() {
return profiles.keySet();
}
/** The parity-overlay file list for {@code profileName} (default list when unset). */
@Override
public List<String> parityOverlay(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
return List.of();
}
BridgedConfig.Worker cfg = profiles.get(name);
return cfg == null ? List.of() : cfg.parityOverlay();
}
/** The profile a no-argument {@link #spawn()} uses, or {@code null} if none is configured. */
@Override
public String defaultProfile() {
return defaultProfile;
}
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawn(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawn(profileName, null, null);
this(agents, spaces, guard, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300));
}
/**
* Spawn a worker. {@code profileName} null/blank → the default profile. The worker's working
* directory (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd} (a
* spawn argument), else the profile's configured {@code cwd}, else {@code callerCwd} (the
* primary's cwd, when the spawn came from the primary over MCP), else the daemon's cwd — never
* assumed to be {@code $HOME}. Guard runs before any herdr call.
* Production constructor with spawn-ready gate enabled. The gate polls {@code agents.status()}
* until the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
throw new IllegalArgumentException("no default worker profile is configured — "
+ "pass a profile; configured: " + profiles.keySet());
}
BridgedConfig.Worker cfg = profiles.get(name);
if (cfg == null) {
throw new IllegalArgumentException("unknown worker profile '" + name
+ "' — configured: " + profiles.keySet());
}
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs));
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
* {@code sleeper} is never called when the gate is disabled ({@code spawnReadyTimeoutMs == 0}).
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param guard subscription-boundary guard (checked before spawning)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (e.g. {@code () -> Thread.sleep(pollMs)}); it
* already encodes the poll interval, so the 8th positional argument
* (poll ms) is accepted for API symmetry but otherwise unused here
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.guard = guard;
}
/**
* {@inheritDoc}
*
* <p>The spawn sequence encodes the subscription boundary: assert the profile's base_url is on
* the allowlist <em>before</em> any herdr call, then build the worker env with
* {@code ANTHROPIC_*}, the parity-neutral git-forge grant, and the bridge MCP + reply charter
* mounted as inline launch flags.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
String baseUrl = cfg.baseUrl();
guard.assertWorker(baseUrl); // hard stop before we spawn anything
Map<String, String> workerEnv = new LinkedHashMap<>();
Map<String, String> workerEnv = baseEnv(cfg);
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
String token = env.apply(cfg.tokenEnv());
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", token);
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
applyGitToken(workerEnv, cfg);
// CB-302: the worker checkpoint (commit → push → open its own PR). Push is free over SSH;
// the only incremental grant is PR-create, a repo-scoped forge token injected here — opt-in
// per profile via gitTokenEnv, and never mutating bridged's own env. The paired forge host
// rides along only when a token is actually granted, so non-implementer profiles get neither.
if (cfg.hasGitToken()) {
String gitToken = resolveEnv(cfg.gitTokenEnv());
if (gitToken != null) {
workerEnv.put("GITEA_TOKEN", gitToken);
putIfPresent(workerEnv, "GITEA_HOST", resolveEnv(cfg.gitHostEnv()));
}
}
// Mount the bridge MCP + reply charter as launch FLAGS (non-invasive: nothing written to
// the worker's profile/config dir). Identity is connection-based, so the mount is shared.
List<String> argv = argvWithBridge(cfg);
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
return cfg.tabPlacement()
? spawnInTab(cfg, workerEnv, argv, cwd)
: spawnAsPane(cfg, workerEnv, argv, cwd);
}
/**
* CB-112 cwd resolution: spawn arg → profile config → the primary's cwd → the daemon's cwd.
* Never returns {@code null}/blank: {@code "."} (the daemon's own working directory) is the
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
*/
private static String resolveCwd(String requestedCwd, BridgedConfig.Worker cfg, String callerCwd) {
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
}
/**
* CB-301: the effective working directory a spawn for {@code profileName} would use, without
* actually spawning. Used by {@link dev.ltms.bridged.session.SessionManager} to record the
* resolved cwd in the session registry.
*/
public String effectiveCwd(String profileName, String requestedCwd, String callerCwd) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
throw new IllegalArgumentException("no default worker profile is configured — "
+ "pass a profile; configured: " + profiles.keySet());
}
BridgedConfig.Worker cfg = profiles.get(name);
if (cfg == null) {
throw new IllegalArgumentException("unknown worker profile '" + name
+ "' — configured: " + profiles.keySet());
}
return resolveCwd(requestedCwd, cfg, callerCwd);
}
private static String firstNonBlank(String... values) {
for (String v : values) {
if (v != null && !v.isBlank()) return v;
}
return null;
return new Launch(workerEnv, argvWithBridge(cfg));
}
/**
* The launch argv, plus — when {@code worker.mcpUrl} is set — inline {@code --mcp-config} for
* the bridge server and {@code --append-system-prompt} for the {@link #REPLY_CHARTER}. Neither
* touches the profile's config; both are pure command-line flags.
* touches the profile's config; both are pure command-line flags. This inline-flag mount is
* Claude Code specific — other adapters mount MCP and instructions their own way.
*/
private List<String> argvWithBridge(BridgedConfig.Worker cfg) {
if (!cfg.hasMcp()) {
@@ -227,7 +141,7 @@ public final class ClaudeCodeLauncher implements PeerLauncher {
}
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
List<String> argv = new ArrayList<>(cfg.argv());
List<String> argv = mutableArgv(cfg.argv());
argv.add("--mcp-config");
argv.add(mcpJson);
argv.add("--append-system-prompt");
@@ -235,209 +149,24 @@ public final class ClaudeCodeLauncher implements PeerLauncher {
return argv;
}
/** Dedicated worker space → own tab → start the worker (rooted at {@code cwd}) → drop the shell. */
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId());
log.info("spawning worker profile={} base_url={} space={} tab={} cwd={}",
cfg.profile(), cfg.baseUrl(), space.workspaceId(), tab.tab().tabId(), cwd);
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
Started started;
try {
started = startUniquelyNamed(cfg, workerEnv, argv, tab.tab().tabId(), cwd);
} catch (RuntimeException e) {
// The worker never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("failed to close orphaned tab {} after spawn error: {}",
tab.tab().tabId(), cleanup.getMessage());
}
throw e;
}
// The worker is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell
// so the tab holds only the worker; label the tab). They must not fail the spawn or
// orphan the running worker — on error we log and still return it so the caller gets
// its paneId and can tear it down.
if (tab.rootPaneId() != null) {
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
} else {
log.warn("tab {} had no seed pane in the create response; worker tab may hold an extra pane",
tab.tab().tabId());
}
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
log.info("worker started pane={} tab={} terminal={}",
started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
return started.agent();
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
private void tidy(String what, Runnable step) {
try {
step.run();
} catch (RuntimeException e) {
log.warn("post-start step failed ({}) — worker is running regardless: {}", what, e.getMessage());
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Legacy placement: herdr splits the currently-focused tab; the worker still starts in {@code cwd}. */
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
log.info("spawning worker (pane placement) profile={} base_url={} cwd={} argv={}",
cfg.profile(), cfg.baseUrl(), cwd, argv);
Agent worker = startUniquelyNamed(cfg, workerEnv, argv, null, cwd).agent();
log.info("worker started pane={} terminal={}", worker.paneId(), worker.terminalId());
return worker;
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
/** A started worker together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
/**
* Start the worker under a unique herdr agent name. herdr requires each running
* agent's {@code name} to be distinct (a 2nd {@code name:"claude"} fails
* {@code agent_name_taken}) — the exact case that makes multiple workers useful. The name
* is {@code claude-<profile>-<nonce>-<seq>}: {@code seq} distinguishes workers within this
* process, and the per-process {@code nonce} keeps a fresh process (whose {@code seq}
* restarts at 0) from colliding with same-profile workers that outlived a restart. The
* retry is a belt-and-braces backstop for the astronomically unlikely nonce+seq clash;
* the name is a label only — herdr detects kind and status from terminal output, not it.
*/
private Started startUniquelyNamed(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String tabId, String cwd) {
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = "claude-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(agents.start(name, argv, workerEnv, tabId, cwd), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("worker name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
/** All herdr-tracked agents — discovery for "what workers exist". */
@Override
public List<Agent> list() {
return agents.list();
}
/**
* Reap worker panes left behind by an earlier daemon process (CB-117). herdr keeps a worker's
* pane alive across a daemon restart <em>by design</em>, and that pane's id is held only by its
* spawner — so a worker whose owning process exited before issuing the matching teardown leaks
* with nothing tracking it (there is no registry; {@link #list()} only asks herdr). On boot we
* scan herdr for agents whose name matches our {@code claude-<profile>-<nonce>-<seq>} scheme with
* a nonce <em>other</em> than this process's {@link #nameNonce}, and tear each one down (its pane
* and, via {@link #stop}, its now-empty dedicated tab). A current-nonce worker is ours and live,
* so it is left running; a user's own {@code claude} session carries no such name and is never
* touched. Best-effort: a failed listing, or a failure to stop any one worker, is logged and
* never aborts startup.
*
* @return the number of orphaned workers reaped
*/
@Override
public int reapOrphanWorkers() {
List<Agent> all;
try {
all = agents.list();
} catch (RuntimeException e) {
log.warn("orphan-worker reap skipped — agent.list failed: {}", e.getMessage());
return 0;
}
int reaped = 0;
for (Agent a : all) {
if (!isForeignWorker(a.name(), nameNonce)) continue;
try {
stop(a.paneId());
reaped++;
log.info("reaped orphan worker {} (pane={} tab={}) left by a prior daemon",
a.name(), a.paneId(), a.tabId());
} catch (RuntimeException e) {
log.warn("could not reap orphan worker {} (pane={}): {}",
a.name(), a.paneId(), e.getMessage());
}
}
if (reaped > 0) {
log.info("orphan-worker reap complete — {} stale worker(s) removed at startup", reaped);
}
return reaped;
}
/**
* Whether {@code name} is a bridge worker started by a <em>different</em> process than
* {@code currentNonce} — the reap predicate (CB-117). True only for our naming scheme with a
* foreign nonce: a non-worker name (no match, e.g. a user session) or our own live nonce is
* excluded. Pure and package-private so the decision is unit-testable without herdr.
*/
static boolean isForeignWorker(String name, String currentNonce) {
String nonce = workerNonce(name);
return nonce != null && !nonce.equals(currentNonce);
}
/** The 6-hex nonce embedded in a bridge worker name, or {@code null} if {@code name} isn't one. */
static String workerNonce(String name) {
if (name == null) return null;
Matcher m = WORKER_NAME.matcher(name);
return m.matches() ? m.group(1) : null;
}
/** This process's worker-name nonce (a label component only; exposed for reaper tests). */
String nameNonce() {
return nameNonce;
}
/**
* Tear a worker down by pane id: close the pane, and close its tab <em>only</em> when the
* worker is that tab's sole occupant. The single-pane check is what makes this safe
* regardless of how the worker was placed (or a placement-config change across a restart):
* a pane-placement worker sitting in one of the user's shared tabs has siblings, so its
* tab is never closed — we only ever remove a tab we created to hold one worker.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed worker) is treated as success; any other failure propagates so
* a genuinely failed teardown is not reported as done.
*/
@Override
public void stop(String paneId) {
// Teardown knows only the paneId, not which profile spawned it. Attempt tab cleanup when any
// profile uses tab placement (so the bridge may have created a dedicated worker tab); the
// single-occupant check below is what actually protects the user's shared tabs.
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) throw e;
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
spaces.closeTab(loc.tabId());
} else if (loc != null) {
log.debug("not closing tab {} — it holds {} panes (not a dedicated worker tab)",
loc.tabId(), loc.tabPaneCount());
}
}
/** Whether any configured profile places workers in their own tab (so tabs may need cleanup). */
private boolean usesTabPlacement() {
return profiles.values().stream().anyMatch(BridgedConfig.Worker::tabPlacement);
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
// --- PeerLauncher SPI -------------------------------------------------------------------
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
@@ -450,39 +179,17 @@ public final class ClaudeCodeLauncher implements PeerLauncher {
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profiles.values().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
// --- CB-117 reap predicate (Claude prefix), kept for direct unit testing -------------------
/**
* {@inheritDoc}
*
* <p>Delegates to the three-arg {@link #spawn(String, String, String)} and wraps the
* resulting herdr {@link Agent} in a {@link WorkerHandle} whose {@link PeerHandle#id()}
* equals the agent's paneId.
* Whether {@code name} is a Claude Code bridge worker started by a <em>different</em> process
* than {@code currentNonce}. A thin {@code claude}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
@Override
public PeerHandle spawn(SpawnRequest req) {
Agent agent = spawn(req.profileName(), req.requestedCwd(), req.callerCwd());
return new WorkerHandle(agent.paneId(), agent.terminalId());
}
/** A concrete {@link PeerHandle} wrapping herdr agent coordinates. */
private record WorkerHandle(String id, String terminalId) implements PeerHandle {
}
@Override
public String effectiveCwd(SpawnRequest req) {
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
}
private static void putIfPresent(Map<String, String> m, String k, String v) {
if (v != null && !v.isBlank()) {
m.put(k, v);
}
}
/** Host env lookup that tolerates an unconfigured (null/blank) var name — returns null then. */
private String resolveEnv(String name) {
return (name == null || name.isBlank()) ? null : env.apply(name);
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -0,0 +1,160 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.EnumSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
/**
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
* router in front of one {@link HerdrPeerLauncher} per peer {@code kind} (Claude Code, opencode, …).
* It owns no transport of its own; it dispatches each SPI call to the delegate that owns the profile
* involved, and fans the fleet-wide queries (list/reap/caps/profiles) across all delegates.
*
* <p>Routing rules:
* <ul>
* <li><strong>By profile</strong> — {@link #spawn}, {@link #effectiveCwd}, {@link #parityOverlay}
* resolve the profile (a null/blank name → the global {@link #defaultProfile}) and delegate to
* the single adapter that declares it. Profiles partition cleanly across adapters: the
* constructor rejects a name claimed by two.</li>
* <li><strong>By pane id</strong> — {@link #stop} routes to the adapter that spawned that pane
* (recorded at spawn time). A pane the composite never spawned (only real for a caller that
* hand-rolls an id) falls back to the first delegate; teardown is pane-id addressed and
* tab cleanup is single-occupant guarded, so it is safe either way.</li>
* <li><strong>Fleet-wide</strong> — {@link #reapOrphanWorkers} and {@link #capabilities} fan out
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
* shares one herdr connection and so reports the same global agent set.</li>
* </ul>
*/
public final class CompositePeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(CompositePeerLauncher.class);
private final List<HerdrPeerLauncher> delegates;
private final Map<String, HerdrPeerLauncher> byProfile;
private final String defaultProfile;
/** paneId → the delegate that spawned it, so {@link #stop} tears down through the right adapter. */
private final Map<String, HerdrPeerLauncher> spawnedBy = new ConcurrentHashMap<>();
/**
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to (may be null)
* @throws IllegalArgumentException if {@code delegates} is empty or two adapters claim one profile
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates, String defaultProfile) {
if (delegates.isEmpty()) {
throw new IllegalArgumentException("at least one peer adapter must be configured");
}
this.delegates = List.copyOf(delegates);
this.defaultProfile = defaultProfile;
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
HerdrPeerLauncher prev = index.putIfAbsent(profile, d);
if (prev != null) {
throw new IllegalArgumentException(
"worker profile '" + profile + "' is claimed by two peer adapters");
}
}
}
this.byProfile = Map.copyOf(index);
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (resolved == null) {
// No profile and no default configured — hand to the first delegate so it raises the
// same "no default" error it would on its own; keeps the SPI contract single-sourced.
return delegates.getFirst();
}
HerdrPeerLauncher d = byProfile.get(resolved);
if (d == null) {
throw new IllegalArgumentException("unknown worker profile: " + resolved);
}
return d;
}
@Override
public PeerHandle spawn(SpawnRequest req) {
HerdrPeerLauncher d = route(req.profileName());
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
}
@Override
public String effectiveCwd(SpawnRequest req) {
return route(req.profileName()).effectiveCwd(req);
}
@Override
public List<String> parityOverlay(String profileName) {
return route(profileName).parityOverlay(profileName);
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.remove(id);
if (d == null) {
log.debug("stop({}) — no recorded owner, routing to the first adapter (pane-addressed)", id);
d = delegates.getFirst();
}
d.stop(id);
}
@Override
public Set<String> profiles() {
return byProfile.keySet();
}
@Override
public String defaultProfile() {
return defaultProfile;
}
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
Map<String, Agent> byPane = new LinkedHashMap<>();
for (HerdrPeerLauncher d : delegates) {
for (Agent a : d.list()) {
if (a.paneId() != null) {
byPane.putIfAbsent(a.paneId(), a);
}
}
}
return List.copyOf(byPane.values());
}
@Override
public int reapOrphanWorkers() {
int reaped = 0;
for (HerdrPeerLauncher d : delegates) {
reaped += d.reapOrphanWorkers();
}
return reaped;
}
/** The union of every adapter's capabilities — a capability any adapter offers, the fleet offers. */
@Override
public Set<Capability> capabilities() {
EnumSet<Capability> caps = EnumSet.noneOf(Capability.class);
for (HerdrPeerLauncher d : delegates) {
caps.addAll(d.capabilities());
}
return Set.copyOf(caps);
}
}
@@ -0,0 +1,546 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.Collection;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* Abstract base for {@link PeerLauncher} adapters that materialize a peer as a <em>herdr</em>
* agent (a CLI coding agent running in a herdr tab/pane). It owns everything that is the same
* regardless of <em>which</em> coding agent runs: tab/pane placement, the CB-306 spawn-readiness
* gate, unique naming, CB-117 orphan reap, teardown, {@link #list() listing}, and cwd resolution.
*
* <p>Two seams are peer-specific and supplied by the concrete adapter:
* <ul>
* <li>{@code namePrefix} (constructor arg) — the label prefix ({@code claude}, {@code opencode})
* that drives both unique naming and the orphan-reap pattern, so each adapter reaps only its
* own kind of pane and never another's.</li>
* <li>{@link #buildLaunch(BridgedConfig.Worker)} — the peer-specific env map + argv, including any
* subscription/guard check, MCP mount, and instruction injection. The base never sees how the
* peer is configured; it only places and starts the returned {@link Launch}.</li>
* </ul>
*
* <p>Placement: in the default {@code tab} policy a peer lands in its own tab inside a dedicated
* worker space (found-or-created once, then shared), so peers never split or clutter the user's
* real work spaces. Teardown removes the peer's pane <em>and</em> its now-empty tab, tolerating an
* already-gone peer so a repeated DELETE is harmless.
*/
public abstract class HerdrPeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(HerdrPeerLauncher.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
private final String namePrefix; // label prefix: naming + reap scheme
private final AgentControl agents;
private final WorkspaceControl spaces;
private final Map<String, BridgedConfig.Worker> profiles; // profile name → spawn settings
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
/** Host env lookup (injectable for tests); adapters read it in {@link #buildLaunch}. */
protected final Function<String, String> env;
private final AtomicLong nameSeq = new AtomicLong(); // per-peer counter (also the tab #)
private final long spawnReadyTimeoutMs; // 0 = disable gate (legacy non-blocking spawn)
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
private final Runnable sleeper; // sleep/wait hook (injectable for tests; never real-sleep in unit tests)
// Per-process token mixed into each peer name so a fresh process (nameSeq back at 0) cannot
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
/**
* @param namePrefix label prefix for this peer kind (drives naming and reap)
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param profiles configured peer profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (never called when the gate is disabled); the poll
* interval is baked into this hook, so the base needs no poll field
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
this.namePrefix = namePrefix;
this.agents = agents;
this.spaces = spaces;
this.profiles = Map.copyOf(profiles);
this.defaultProfile = defaultProfile;
this.env = env;
this.spawnReadyTimeoutMs = spawnReadyTimeoutMs;
this.nowMillis = nowMillis;
this.sleeper = sleeper;
}
// --- adapter seams -------------------------------------------------------------------------
/**
* Build the peer-specific launch for {@code cfg}: the environment map and argv handed to herdr.
* Any subscription/guard check, MCP mount, and instruction injection happen here. The env map
* and argv are adapter-private; the base only places and starts what is returned.
*/
protected abstract Launch buildLaunch(BridgedConfig.Worker cfg);
/** A peer-specific launch: the herdr {@code env} map and {@code argv}. */
protected record Launch(Map<String, String> env, List<String> argv) {
}
// --- profile surface -----------------------------------------------------------------------
/** The configured peer profile names (what {@code spawn(profile)} accepts). */
@Override
public Set<String> profiles() {
return profiles.keySet();
}
/** The parity-overlay file list for {@code profileName} (default list when unset). */
@Override
public List<String> parityOverlay(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
return List.of();
}
BridgedConfig.Worker cfg = profiles.get(name);
return cfg == null ? List.of() : cfg.parityOverlay();
}
/** The profile a no-argument spawn uses, or {@code null} if none is configured. */
@Override
public String defaultProfile() {
return defaultProfile;
}
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
protected Collection<BridgedConfig.Worker> profileConfigs() {
return profiles.values();
}
/** Resolve {@code profileName} (null/blank → default) to its config, or throw with the options. */
protected BridgedConfig.Worker requireProfile(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
throw new IllegalArgumentException("no default worker profile is configured — "
+ "pass a profile; configured: " + profiles.keySet());
}
BridgedConfig.Worker cfg = profiles.get(name);
if (cfg == null) {
throw new IllegalArgumentException("unknown worker profile '" + name
+ "' — configured: " + profiles.keySet());
}
return cfg;
}
// --- spawn ---------------------------------------------------------------------------------
/**
* Spawn a peer. {@code profileName} null/blank → the default profile. The working directory
* (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd}, else the
* profile's configured {@code cwd}, else {@code callerCwd} (the primary's cwd, when the spawn
* came from the primary over MCP), else the daemon's cwd — never assumed to be {@code $HOME}.
* The adapter's {@link #buildLaunch} runs before any herdr call.
*/
protected Agent spawnInternal(String profileName, String requestedCwd, String callerCwd) {
BridgedConfig.Worker cfg = requireProfile(profileName);
Launch launch = buildLaunch(cfg);
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
return cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd);
}
/**
* {@inheritDoc}
*
* <p>Delegates to {@link #spawnInternal} and wraps the resulting herdr {@link Agent} in a
* {@link WorkerHandle} whose {@link PeerHandle#id()} equals the agent's paneId. When
* {@code spawnReadyTimeoutMs > 0}, blocks until the peer's herdr status is injectable or the
* timeout elapses; on timeout the pane is closed (no orphan) and a
* {@link PeerUnreachableException} is thrown.
*/
@Override
public PeerHandle spawn(SpawnRequest req) {
Agent agent = spawnInternal(req.profileName(), req.requestedCwd(), req.callerCwd());
String paneId = agent.paneId();
if (spawnReadyTimeoutMs > 0) {
waitUntilInjectableOrThrow(paneId);
}
return new WorkerHandle(paneId, agent.terminalId());
}
@Override
public String effectiveCwd(SpawnRequest req) {
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
}
/**
* CB-301: the effective working directory a spawn for {@code profileName} would use, without
* actually spawning.
*/
private String effectiveCwd(String profileName, String requestedCwd, String callerCwd) {
return resolveCwd(requestedCwd, requireProfile(profileName), callerCwd);
}
/**
* CB-112 cwd resolution: spawn arg → profile config → the primary's cwd → the daemon's cwd.
* Never returns {@code null}/blank: {@code "."} (the daemon's own working directory) is the
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
*/
private static String resolveCwd(String requestedCwd, BridgedConfig.Worker cfg, String callerCwd) {
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
}
private static String firstNonBlank(String... values) {
for (String v : values) {
if (v != null && !v.isBlank()) return v;
}
return null;
}
/** Dedicated worker space → own tab → start the peer (rooted at {@code cwd}) → drop the shell. */
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId());
log.info("spawning {} profile={} space={} tab={} cwd={}",
namePrefix, cfg.profile(), space.workspaceId(), tab.tab().tabId(), cwd);
Started started;
try {
started = startUniquelyNamed(cfg, workerEnv, argv, tab.tab().tabId(), cwd);
} catch (RuntimeException e) {
// The peer never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("failed to close orphaned tab {} after spawn error: {}",
tab.tab().tabId(), cleanup.getMessage());
}
throw e;
}
// The peer is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell so the
// tab holds only the peer; label the tab). They must not fail the spawn or orphan the
// running peer — on error we log and still return it so the caller gets its paneId and can
// tear it down.
if (tab.rootPaneId() != null) {
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
} else {
log.warn("tab {} had no seed pane in the create response; peer tab may hold an extra pane",
tab.tab().tabId());
}
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
log.info("{} started pane={} tab={} terminal={}",
namePrefix, started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
return started.agent();
}
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
private void tidy(String what, Runnable step) {
try {
step.run();
} catch (RuntimeException e) {
log.warn("post-start step failed ({}) — peer is running regardless: {}", what, e.getMessage());
}
}
/** Legacy placement: herdr splits the currently-focused tab; the peer still starts in {@code cwd}. */
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
namePrefix, cfg.profile(), cwd, argv);
Agent peer = startUniquelyNamed(cfg, workerEnv, argv, null, cwd).agent();
log.info("{} started pane={} terminal={}", namePrefix, peer.paneId(), peer.terminalId());
return peer;
}
/** A started peer together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
/**
* Start the peer under a unique herdr agent name. herdr requires each running agent's
* {@code name} to be distinct (a 2nd identical {@code name} fails {@code agent_name_taken}) —
* the exact case that makes multiple peers useful. The name is
* {@code <prefix>-<profile>-<nonce>-<seq>}: {@code seq} distinguishes peers within this process,
* and the per-process {@code nonce} keeps a fresh process (whose {@code seq} restarts at 0) from
* colliding with same-profile peers that outlived a restart. The retry is a belt-and-braces
* backstop for the astronomically unlikely nonce+seq clash; the name is a label only — herdr
* detects kind and status from terminal output, not from it.
*/
private Started startUniquelyNamed(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String tabId, String cwd) {
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(agents.start(name, argv, workerEnv, tabId, cwd), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("peer name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
// --- discovery + reap ----------------------------------------------------------------------
/** All herdr-tracked agents — discovery for "what peers exist". */
@Override
public List<Agent> list() {
return agents.list();
}
/**
* Reap peer panes left behind by an earlier daemon process (CB-117). herdr keeps a peer's pane
* alive across a daemon restart <em>by design</em>, and that pane's id is held only by its
* spawner — so a peer whose owning process exited before issuing the matching teardown leaks
* with nothing tracking it. On boot we scan herdr for agents whose name matches our
* {@code <prefix>-<profile>-<nonce>-<seq>} scheme with a nonce <em>other</em> than this
* process's {@link #nameNonce}, and tear each one down (its pane and, via {@link #stop}, its
* now-empty dedicated tab). A current-nonce peer is ours and live, so it is left running; a
* user's own session carries no such name and is never touched. A peer from a <em>different</em>
* adapter (different prefix) is likewise never touched. Best-effort: a failed listing, or a
* failure to stop any one peer, is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
@Override
public int reapOrphanWorkers() {
List<Agent> all;
try {
all = agents.list();
} catch (RuntimeException e) {
log.warn("orphan-peer reap skipped — agent.list failed: {}", e.getMessage());
return 0;
}
int reaped = 0;
for (Agent a : all) {
if (!isForeignWorker(namePrefix, a.name(), nameNonce)) continue;
try {
stop(a.paneId());
reaped++;
log.info("reaped orphan {} {} (pane={} tab={}) left by a prior daemon",
namePrefix, a.name(), a.paneId(), a.tabId());
} catch (RuntimeException e) {
log.warn("could not reap orphan {} {} (pane={}): {}",
namePrefix, a.name(), a.paneId(), e.getMessage());
}
}
if (reaped > 0) {
log.info("orphan-peer reap complete — {} stale {} peer(s) removed at startup", reaped, namePrefix);
}
return reaped;
}
/** The {@code <prefix>-<profile>-<nonce>-<seq>} name pattern; group 1 captures the 6-hex nonce. */
static Pattern workerNamePattern(String prefix) {
return Pattern.compile(prefix + "-.*-([0-9a-f]{6})-\\d+");
}
/**
* Whether {@code name} is a peer of kind {@code prefix} started by a <em>different</em> process
* than {@code currentNonce} — the reap predicate (CB-117). True only for the prefix's naming
* scheme with a foreign nonce: a non-peer name, a different adapter's name, or our own live
* nonce is excluded. Pure and package-private so the decision is unit-testable without herdr.
*/
static boolean isForeignWorker(String prefix, String name, String currentNonce) {
String nonce = workerNonce(prefix, name);
return nonce != null && !nonce.equals(currentNonce);
}
/** The 6-hex nonce embedded in a {@code prefix} peer name, or {@code null} if not one. */
static String workerNonce(String prefix, String name) {
if (name == null) return null;
Matcher m = workerNamePattern(prefix).matcher(name);
return m.matches() ? m.group(1) : null;
}
/** This process's peer-name nonce (a label component only; exposed for reaper tests). */
String nameNonce() {
return nameNonce;
}
// --- teardown ------------------------------------------------------------------------------
/**
* Tear a peer down by pane id: close the pane, and close its tab <em>only</em> when the peer is
* that tab's sole occupant. The single-pane check is what makes this safe regardless of how the
* peer was placed (or a placement-config change across a restart): a pane-placement peer sitting
* in one of the user's shared tabs has siblings, so its tab is never closed — we only ever
* remove a tab we created to hold one peer.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed peer) is treated as success; any other failure propagates so a
* genuinely failed teardown is not reported as done.
*/
@Override
public void stop(String paneId) {
// Teardown knows only the paneId, not which profile spawned it. Attempt tab cleanup when any
// profile uses tab placement (so the bridge may have created a dedicated peer tab); the
// single-occupant check below is what actually protects the user's shared tabs.
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) throw e;
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
spaces.closeTab(loc.tabId());
} else if (loc != null) {
log.debug("not closing tab {} — it holds {} panes (not a dedicated peer tab)",
loc.tabId(), loc.tabPaneCount());
}
}
/** Whether any configured profile places peers in their own tab (so tabs may need cleanup). */
private boolean usesTabPlacement() {
return profiles.values().stream().anyMatch(BridgedConfig.Worker::tabPlacement);
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
// --- spawn-readiness gate (CB-306) ---------------------------------------------------------
/**
* Poll {@link AgentControl#status} until the pane reports an injectable state or the configured
* timeout elapses. On timeout, close the pane (self-reap) and throw.
*/
private void waitUntilInjectableOrThrow(String paneId) {
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
while (nowMillis.getAsLong() < deadline) {
if (agents.status(paneId).injectable()) {
log.debug("peer pane={} reached injectable state", paneId);
return;
}
sleeper.run();
}
log.warn("peer pane={} did not become injectable within {}ms — closing", paneId, spawnReadyTimeoutMs);
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " did not reach injectable state within "
+ spawnReadyTimeoutMs + "ms");
}
/** A concrete {@link PeerHandle} wrapping herdr agent coordinates. */
private record WorkerHandle(String id, String terminalId) implements PeerHandle {
}
// --- shared helpers ------------------------------------------------------------------------
/** Put {@code k → v} only when {@code v} is present (non-null, non-blank). */
protected static void putIfPresent(Map<String, String> m, String k, String v) {
if (v != null && !v.isBlank()) {
m.put(k, v);
}
}
/** Host env lookup that tolerates an unconfigured (null/blank) var name — returns null then. */
protected String resolveEnv(String name) {
return (name == null || name.isBlank()) ? null : env.apply(name);
}
/**
* The parity-neutral git-forge token grant (CB-302): when {@code cfg} opts in via
* {@code gitTokenEnv} and the token resolves, inject {@code GITEA_TOKEN} plus its paired
* {@code GITEA_HOST}. Push over SSH is unaffected; the only incremental grant is PR-create.
* Peer-neutral, so every herdr adapter reuses it unchanged.
*/
protected void applyGitToken(Map<String, String> workerEnv, BridgedConfig.Worker cfg) {
if (!cfg.hasGitToken()) {
return;
}
String gitToken = resolveEnv(cfg.gitTokenEnv());
if (gitToken != null) {
workerEnv.put("GITEA_TOKEN", gitToken);
putIfPresent(workerEnv, "GITEA_HOST", resolveEnv(cfg.gitHostEnv()));
}
}
/** A fresh mutable env map — the conventional starting point for {@link #buildLaunch}. */
/**
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
* {@code env:} entries.
*
* <p>Why this exists: bridged passes herdr an explicit env map, and herdr merges it into
* <em>its own</em> process environment. So before this, a worker inherited whatever PATH the
* herdr server happened to be started with — on this host, one from weeks earlier with no JDK
* and no Maven, which left workers unable to run the build they were being asked to run. The
* worker's toolchain must follow from configuration, not from how a long-lived daemon was
* launched.
*
* <p>Adapter-specific variables are layered on top of this by {@code buildLaunch} and therefore
* win. That ordering is deliberate and load-bearing: it stops a profile's {@code env:} from
* overriding {@code ANTHROPIC_BASE_URL} and slipping past {@link
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
* {@code baseUrl} and nothing else.
*/
protected Map<String, String> baseEnv(BridgedConfig.Worker cfg) {
Map<String, String> workerEnv = new LinkedHashMap<>();
String path = env.apply("PATH");
if (path != null && !path.isBlank()) {
workerEnv.put("PATH", path);
}
if (cfg != null && cfg.env() != null) {
workerEnv.putAll(cfg.env());
}
return workerEnv;
}
/** Defensive copy of {@code argv} plus room to append launch flags. */
protected static List<String> mutableArgv(List<String> argv) {
return new ArrayList<>(argv);
}
/**
* Uninterruptible sleep — the production {@link #sleeper}. Tests supply their own no-op /
* fast-faking sleeper so they never real-sleep.
*/
protected static void sleepUninterruptibly(long ms) {
try {
Thread.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
// preserve the interrupt flag but continue — poll loops should not be aborted by an
// interrupt that was not meant for them.
}
}
}
@@ -0,0 +1,316 @@
package dev.ltms.bridged.worker;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.node.ObjectNode;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>opencode</strong> — an open-source,
* provider-agnostic terminal coding agent. Its whole reason for existing is to prove the
* {@code PeerLauncher} SPI is genuinely provider-neutral: opencode shares none of Claude Code's
* private launch seams, yet reuses every line of shared transport in the base (tab/pane placement,
* the CB-306 readiness gate, unique naming + CB-117 reap, teardown, listing, cwd).
*
* <p>The divergences from {@link ClaudeCodeLauncher}, all confined to {@link #buildLaunch}:
* <ul>
* <li><strong>No subscription boundary.</strong> opencode carries no {@code ANTHROPIC_BASE_URL}
* and there is no {@link dev.ltms.bridged.guard.SubscriptionGuard} — the guard is a
* Claude-private concern, not part of the SPI. opencode reads the operator's own provider
* credentials from its global {@code auth.json}; the bridge injects none.</li>
* <li><strong>File-based MCP mount + instructions.</strong> opencode has no inline
* {@code --mcp-config}/{@code --append-system-prompt}. Instead the bridge writes an ephemeral
* {@code opencode.json} that declares the bridge as a {@code remote} MCP server and lists a
* reply-charter file under {@code instructions}, then points the worker at it with
* {@code OPENCODE_CONFIG}. This is the one place the launcher touches disk — Claude never did.</li>
* <li><strong>Model as a flag.</strong> the {@code provider/model} selector is passed as
* {@code -m}, not an env var.</li>
* <li><strong>{@code opencode} name prefix</strong> so reap matches {@code opencode-*} panes and
* never another adapter's.</li>
* </ul>
*/
public final class OpenCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "opencode";
/** Writer for the generated {@code opencode.json}. */
private static final ObjectMapper JSON = new ObjectMapper();
/**
* Standing instruction written to the charter file and mounted via the config's
* {@code instructions} so the worker returns its result through {@code bridge_reply}. Kept on
* disk (not a launch flag) because opencode's {@code instructions} takes file paths, not inline
* text — the file is regenerated per spawn and never touches the worker's own profile.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet, running under opencode. "
+ "Every message you receive arrives through the bridge, and the ONLY channel back to the "
+ "sender is the bridge_reply MCP tool. Text you write in your terminal is NOT sent "
+ "anywhere — the sender cannot see your screen, so an in-terminal answer is silently "
+ "discarded. Therefore you MUST end EVERY turn by calling bridge_reply with `content` set "
+ "to your complete response. This holds for every message without exception — tasks, "
+ "questions, clarifications, acknowledgements, and ordinary back-and-forth conversation. "
+ "Call bridge_reply exactly once, as the final action of your turn, with your full answer "
+ "in `content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/** Root under which per-spawn opencode config dirs are created (injectable for tests). */
private final Path configRoot;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so it
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300),
defaultConfigRoot());
}
/**
* Production constructor with the spawn-ready gate enabled. Polls {@code agents.status()} until
* the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot());
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply a
* fake clock ({@code nowMillis}), poll-loop wait ({@code sleeper}), and a temp {@code configRoot}
* they can inspect the generated {@code opencode.json}/charter under.
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (encodes the poll interval; never called when the
* gate is disabled)
* @param configRoot existing directory under which per-spawn config dirs are created
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper, Path configRoot) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.configRoot = configRoot;
}
private static Path defaultConfigRoot() {
return Path.of(System.getProperty("java.io.tmpdir"));
}
/**
* {@inheritDoc}
*
* <p>Builds the opencode launch: no {@code ANTHROPIC_*} and no guard (opencode reads its own
* provider credentials); when the profile mounts the bridge MCP, generate an ephemeral
* {@code opencode.json} (remote MCP server + reply-charter instructions) and point the worker at
* it via {@code OPENCODE_CONFIG}; carry the parity-neutral git-forge grant; and select the model
* with {@code -m}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
Map<String, String> workerEnv = baseEnv(cfg);
// A config file is needed for the bridge MCP mount, for a pinned endpoint (CB-508), or both.
if (cfg.hasMcp() || hasCustomProvider(cfg)) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg).toString());
}
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithModel(cfg));
}
/**
* True when this profile pins its own OpenAI-compatible endpoint (CB-508) rather than using
* whatever provider opencode resolves by default.
*
* <p>Note this reuses {@code baseUrl}, the same field the Claude adapter injects as
* {@code ANTHROPIC_BASE_URL} — but it does <em>not</em> go through {@code SubscriptionGuard}.
* That asymmetry is deliberate and safe: the guard exists to stop a worker borrowing the
* primary's Anthropic subscription, and an opencode process has no Anthropic credential path
* at all. Pointing it at a local vLLM cannot leak the subscription.
*/
private static boolean hasCustomProvider(BridgedConfig.Worker cfg) {
return cfg.baseUrl() != null && !cfg.baseUrl().isBlank();
}
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
private List<String> argvWithModel(BridgedConfig.Worker cfg) {
List<String> argv = mutableArgv(cfg.argv());
if (cfg.model() != null && !cfg.model().isBlank()) {
argv.add("-m");
argv.add(cfg.model());
}
return argv;
}
/**
* Write an ephemeral {@code opencode.json} (and the reply-charter file it references) into a
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
*/
private Path writeConfig(BridgedConfig.Worker cfg) {
try {
Path dir = Files.createTempDirectory(configRoot, "bridged-opencode-");
dir.toFile().deleteOnExit();
ObjectNode root = JSON.createObjectNode();
root.put("$schema", "https://opencode.ai/config.json");
if (cfg.hasMcp()) {
Path charter = dir.resolve("reply-charter.md");
Files.writeString(charter, REPLY_CHARTER);
charter.toFile().deleteOnExit();
ObjectNode bridge = root.putObject("mcp").putObject("bridge");
bridge.put("type", "remote");
bridge.put("url", cfg.mcpUrl());
bridge.put("enabled", true);
root.putArray("instructions").add(charter.toAbsolutePath().toString());
}
if (hasCustomProvider(cfg)) {
addCustomProvider(root, cfg);
}
Path cfgFile = dir.resolve("opencode.json");
// Built with Jackson rather than string concatenation: the provider block is nested and
// carries operator-supplied values (URL, model id, api key), so escaping must be real.
Files.writeString(cfgFile, JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root));
cfgFile.toFile().deleteOnExit();
return cfgFile;
} catch (IOException e) {
throw new UncheckedIOException(
"cannot write opencode config for profile " + cfg.profile(), e);
}
}
/**
* Declare a custom OpenAI-compatible provider so the worker talks to a pinned endpoint (a local
* vLLM, say) instead of opencode's default gateway (CB-508).
*
* <p>The provider id comes from the {@code provider/model} selector in {@code model:}, so one
* field drives both the declaration and the {@code -m} flag and they cannot drift apart.
*/
private void addCustomProvider(ObjectNode root, BridgedConfig.Worker cfg) {
String[] parts = splitModelSelector(cfg);
String providerId = parts[0];
String modelId = parts[1];
ObjectNode provider = root.putObject("provider").putObject(providerId);
provider.put("npm", "@ai-sdk/openai-compatible");
provider.put("name", providerId + " (bridged)");
ObjectNode options = provider.putObject("options");
options.put("baseURL", openAiBaseUrl(cfg.baseUrl()));
// vLLM and friends usually ignore the key, but the AI SDK still requires a non-empty one.
String token = resolveEnv(cfg.tokenEnv());
options.put("apiKey", (token == null || token.isBlank()) ? "bridged-local-noauth" : token);
provider.putObject("models").putObject(modelId).put("name", modelId);
}
/**
* Split {@code model:} into its {@code provider} and {@code model} halves. A pinned endpoint
* needs both, so a bare model name is rejected loudly rather than silently falling back to the
* default gateway — a worker quietly talking to the wrong endpoint is the failure this avoids.
*/
private static String[] splitModelSelector(BridgedConfig.Worker cfg) {
String model = cfg.model();
int slash = model == null ? -1 : model.indexOf('/');
if (model == null || model.isBlank() || slash <= 0 || slash == model.length() - 1) {
throw new IllegalArgumentException(
"profile " + cfg.profile() + " sets baseUrl (a pinned opencode endpoint) so"
+ " model: must be \"<provider>/<model>\", e.g."
+ " \"local-vllm/deepseek-v4-flash\"; got "
+ (model == null ? "null" : '"' + model + '"'));
}
return new String[]{model.substring(0, slash), model.substring(slash + 1)};
}
/**
* The OpenAI-compatible base URL for {@code baseUrl}. A bare {@code host:port} gets {@code /v1}
* appended (where these servers put the API); a URL that already carries a path is taken as-is,
* so an endpoint mounted somewhere unusual is still reachable.
*/
private static String openAiBaseUrl(String baseUrl) {
String trimmed = baseUrl.trim();
while (trimmed.endsWith("/")) {
trimmed = trimmed.substring(0, trimmed.length() - 1);
}
int schemeEnd = trimmed.indexOf("://");
String afterScheme = schemeEnd < 0 ? trimmed : trimmed.substring(schemeEnd + 3);
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
// --- CB-117 reap predicate (opencode prefix), kept for direct unit testing -----------------
/**
* Whether {@code name} is an opencode bridge worker started by a <em>different</em> process than
* {@code currentNonce}. A thin {@code opencode}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
+29
View File
@@ -5,10 +5,39 @@
</encoder>
</appender>
<!--
CB-505 audit trail. Its own file, deliberately not the app log: privileged actions
(spawn/stop/send/reply/drain) must stay greppable and shippable without dragging DEBUG noise
along. AuditLog emits a complete JSON object including its own ISO-8601 "ts" field, so the
pattern is a bare %msg — a pattern that spliced literal braces around the message would
collide with logback's own variable substitution. Rolls daily, 30 days retained, 100MB cap.
NOTE: records carry who/what/target/outcome only. Message CONTENT is never written here —
this bridge carries source code and prompts, and an audit log that accumulated them would be
a transcript archive rather than a control.
-->
<appender name="AUDIT" class="ch.qos.logback.core.rolling.RollingFileAppender">
<file>logs/audit.log</file>
<rollingPolicy class="ch.qos.logback.core.rolling.SizeAndTimeBasedRollingPolicy">
<fileNamePattern>logs/audit.%d{yyyy-MM-dd}.%i.log</fileNamePattern>
<maxFileSize>10MB</maxFileSize>
<maxHistory>30</maxHistory>
<totalSizeCap>100MB</totalSizeCap>
</rollingPolicy>
<encoder>
<pattern>%msg%n</pattern>
</encoder>
</appender>
<logger name="dev.ltms.bridged" level="DEBUG"/>
<logger name="io.javalin" level="INFO"/>
<logger name="org.eclipse.jetty" level="WARN"/>
<!-- additivity=false keeps the audit stream out of stdout; it is its own record. -->
<logger name="audit" level="INFO" additivity="false">
<appender-ref ref="AUDIT"/>
</logger>
<root level="INFO">
<appender-ref ref="STDOUT"/>
</root>
@@ -0,0 +1,100 @@
package dev.ltms.bridged.auth;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-505 — the audit record's shape.
*
* <p>These exist because the first cut of this feature emitted lines that were <em>not</em> valid
* JSON: the timestamp was spliced on by a logback pattern whose literal braces collided with
* logback's variable substitution. The appender failed to parse, and nothing in the build noticed.
* An audit trail that silently stops being machine-readable is worse than none.
*/
class AuditLogTest {
private final ObjectMapper mapper = new ObjectMapper();
private ListAppender<ILoggingEvent> appender;
private ch.qos.logback.classic.Logger auditLogger;
@BeforeEach
void attach() {
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
auditLogger = ctx.getLogger("audit");
appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
auditLogger.addAppender(appender);
auditLogger.setLevel(Level.INFO);
}
@AfterEach
void detach() {
auditLogger.detachAppender(appender);
}
private JsonNode onlyRecord() throws Exception {
assertEquals(1, appender.list.size(), "exactly one audit line expected");
String line = appender.list.getFirst().getFormattedMessage();
return mapper.readTree(line); // throws if the line is not valid JSON
}
@Test
void anAllowedActionIsRecordedAsValidJson() throws Exception {
AuditLog.allowed(Principal.primary(4242), Authz.Action.SPAWN, "term_a");
JsonNode r = onlyRecord();
assertEquals("PRIMARY", r.path("role").asText());
assertEquals("primary", r.path("actor").asText());
assertEquals(4242, r.path("pid").asLong());
assertEquals("SPAWN", r.path("action").asText());
assertEquals("term_a", r.path("target").asText());
assertEquals("allowed", r.path("outcome").asText());
assertFalse(r.path("ts").asText().isBlank(), "every record carries its own timestamp");
}
@Test
void aDenialRecordsTheReason() throws Exception {
AuditLog.denied(Principal.worker("term_b", 7), Authz.Action.REPLY, "term_a", "forbidden");
JsonNode r = onlyRecord();
assertEquals("WORKER", r.path("role").asText());
assertEquals("worker:term_b", r.path("actor").asText());
assertEquals("denied", r.path("outcome").asText());
assertEquals("forbidden", r.path("reason").asText());
}
@Test
void aNullCallerIsRecordedAsAnonymousRatherThanCrashing() throws Exception {
AuditLog.failed(null, Authz.Action.SEND, null, "herdr unreachable");
JsonNode r = onlyRecord();
assertEquals("ANONYMOUS", r.path("role").asText());
assertTrue(r.path("target").isNull(), "an absent target is JSON null, not the string \"null\"");
assertEquals("failed", r.path("outcome").asText());
}
@Test
void hostileValuesAreEscapedAndCannotForgeAnExtraRecord() throws Exception {
// A target id containing a quote and a newline must not be able to terminate the JSON
// object early and inject a second, attacker-shaped audit line.
AuditLog.denied(Principal.worker("term_a", 1), Authz.Action.REPLY,
"evil\",\"outcome\":\"allowed\"}\n{\"forged\":true", "forbidden");
JsonNode r = onlyRecord();
assertEquals("denied", r.path("outcome").asText(),
"the injected outcome must not override the real one");
assertTrue(r.path("target").asText().contains("forged"),
"the hostile text survives as inert data inside the target field");
}
}
@@ -0,0 +1,80 @@
package dev.ltms.bridged.auth;
import org.junit.jupiter.api.Test;
import static dev.ltms.bridged.auth.Authz.Action.*;
import static org.junit.jupiter.api.Assertions.*;
/** CB-505 — the authorization table, pinned so it cannot drift silently. */
class AuthzTest {
private static final Principal PRIMARY = Principal.primary(100);
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal WORKER_B = Principal.worker("term_b", 300);
private static final Principal ANON = Principal.anonymous();
@Test
void anonymousIsAuthorizedForNothing() {
for (Authz.Action a : Authz.Action.values()) {
assertFalse(Authz.permits(ANON, a, "term_a"),
a + " must be refused to an unauthenticated caller");
}
}
@Test
void aNullCallerIsTreatedAsAnonymous() {
assertFalse(Authz.permits(null, READ, null));
assertTrue(Authz.isUnauthenticated(null));
}
@Test
void orchestrationBelongsToThePrimaryAlone() {
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, SEND, DRAIN}) {
assertTrue(Authz.permits(PRIMARY, a, "term_a"), "the primary orchestrates: " + a);
assertFalse(Authz.permits(WORKER_A, a, "term_a"),
"a worker performing " + a + " would be escalating into the orchestrator role");
}
}
@Test
void aWorkerMayReplyAndAskOnlyAsItself() {
assertTrue(Authz.permits(WORKER_A, REPLY, "term_a"));
assertTrue(Authz.permits(WORKER_A, ASK, "term_a"));
assertFalse(Authz.permits(WORKER_A, REPLY, "term_b"),
"worker A must not be able to reply on worker B's session");
assertFalse(Authz.permits(WORKER_B, ASK, "term_a"),
"worker B must not be able to ask as worker A");
}
@Test
void thePrimaryMayNotForgeAWorkersReply() {
// Not a hypothetical nicety: a forged reply would resolve the rendezvous the primary is
// itself blocked on, corrupting the correlation between a turn and its answer.
assertFalse(Authz.permits(PRIMARY, REPLY, "term_a"));
assertFalse(Authz.permits(PRIMARY, ASK, "term_a"));
}
@Test
void aWorkerWithNoTargetCannotReply() {
assertFalse(Authz.permits(WORKER_A, REPLY, null),
"an absent session id must not satisfy the own-session rule");
}
@Test
void observationIsOpenToBothAuthenticatedRoles() {
assertTrue(Authz.permits(PRIMARY, READ, null));
assertTrue(Authz.permits(WORKER_A, READ, null));
assertTrue(Authz.permits(PRIMARY, METRICS, null));
assertTrue(Authz.permits(WORKER_A, METRICS, null));
}
@Test
void unauthenticatedIsDistinguishedFromMerelyForbidden() {
// Drives the 401-vs-403 split: a missing credential is fixable by the caller, a wrong role
// is not.
assertTrue(Authz.isUnauthenticated(ANON));
assertFalse(Authz.isUnauthenticated(WORKER_A));
assertFalse(Authz.isUnauthenticated(PRIMARY));
}
}
@@ -0,0 +1,106 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-501. The behaviour under test is the inversion of the pre-CB-501 default: failing every
* identity check must yield {@link Role#ANONYMOUS}, not {@code PRIMARY}.
*/
class CallerResolverTest {
private final FakeHerdr herdr = new FakeHerdr();
/** Identity resolving the canned worker pane, keyed off a faked peer-PID lookup. */
private ConnectionIdentity identity(long pid) {
return new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
}
/** A PID that owns a worker pane in the fake. */
private ConnectionIdentity workerIdentity() {
return identity(FakeHerdr.WORKER_PID);
}
/** A PID that owns no pane — i.e. the primary, or any other local process. */
private ConnectionIdentity nonWorkerIdentity() {
return identity(999_999);
}
@Test
void aLoopbackWorkerPaneResolvesToWorkerRegardlessOfAuthMode() {
Principal underTrust = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
Principal underToken = new CallerResolver(workerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, underTrust.role());
assertEquals("term_a", underTrust.terminal());
assertEquals(Role.WORKER, underToken.role(),
"worker identity is unforgeable and must never be token-gated — otherwise enabling "
+ "auth would lock the whole fleet out of bridge_reply");
assertEquals("term_a", underToken.terminal());
}
@Test
void loopbackTrustTreatsANonWorkerLoopbackCallerAsThePrimary() {
Principal p = new CallerResolver(nonWorkerIdentity()).resolve("127.0.0.1", 99, null);
assertEquals(Role.PRIMARY, p.role(), "the historical behaviour, now an explicit choice");
}
@Test
void tokenModeRefusesANonWorkerCallerThatPresentsNoToken() {
Principal p = new CallerResolver(nonWorkerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 99, null);
assertEquals(Role.ANONYMOUS, p.role(),
"no credential must mean NOTHING, not the most privileged role on the bus");
}
@Test
void tokenModeAcceptsAValidBearerTokenAsThePrimary() {
Principal p = new CallerResolver(nonWorkerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 99, "Bearer s3cret");
assertEquals(Role.PRIMARY, p.role());
}
@Test
void tokenModeRejectsAWrongOrMalformedCredential() {
CallerResolver r = new CallerResolver(nonWorkerIdentity(), true, "s3cret");
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "Bearer wrong").role());
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "s3cret").role(), "scheme required");
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "Bearer ").role(), "empty credential");
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "Basic s3cret").role(), "wrong scheme");
}
@Test
void theBearerSchemeIsCaseInsensitivePerRfc7235() {
CallerResolver r = new CallerResolver(nonWorkerIdentity(), true, "s3cret");
assertEquals(Role.PRIMARY, r.resolve("127.0.0.1", 99, "bearer s3cret").role());
assertEquals(Role.PRIMARY, r.resolve("127.0.0.1", 99, "BEARER s3cret").role());
}
@Test
void aNonLoopbackCallerIsNeverThePrimaryUnderLoopbackTrust() {
// Defence in depth: startup already refuses this pairing (validateAuthExposure), but if a
// proxy ever forwards a remote peer onto the loopback listener, the resolver must not
// hand it the primary role.
Principal p = new CallerResolver(nonWorkerIdentity()).resolve("10.0.0.7", 99, null);
assertEquals(Role.ANONYMOUS, p.role());
}
@Test
void tokenModeRequiresANonEmptyConfiguredToken() {
ConnectionIdentity id = nonWorkerIdentity();
assertThrows(IllegalArgumentException.class, () -> new CallerResolver(id, true, null));
assertThrows(IllegalArgumentException.class, () -> new CallerResolver(id, true, " "));
}
}
@@ -90,4 +90,279 @@ class BridgedConfigTest {
Files.writeString(f, "bind:\n port: 8080\nfutureFeature:\n enabled: true\n");
assertDoesNotThrow(() -> BridgedConfig.load(f));
}
@Test
void absentBrokerBlockLeavesInboxSoftState(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-broker.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNull(cfg.broker(), "no broker: block → null → in-memory inbox is selected");
}
@Test
void brokerBlockWithUriEnablesAmqpAdapter(@TempDir Path dir) throws Exception {
Path f = dir.resolve("broker.yaml");
Files.writeString(f, """
bind:
port: 8080
broker:
uri: amqp://guest:guest@127.0.0.1:5672/
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.broker());
assertTrue(cfg.broker().isConfigured(), "a non-blank uri enables the AMQP adapter");
assertEquals("amqp://guest:guest@127.0.0.1:5672/", cfg.broker().uri());
}
@Test
void brokerBlockWithBlankUriStaysSoftState(@TempDir Path dir) throws Exception {
Path f = dir.resolve("broker-blank.yaml");
Files.writeString(f, "bind:\n port: 8080\nbroker:\n uri: \"\"\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.broker());
assertFalse(cfg.broker().isConfigured(), "an empty uri must not enable AMQP");
}
@Test
void absentPrimaryBlockLeavesPrimaryNull(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-primary.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNull(cfg.primary(), "no primary: block → null → connection-derived identity");
}
@Test
void primaryBlockWithTerminalPinsIdentity(@TempDir Path dir) throws Exception {
Path f = dir.resolve("primary-pinned.yaml");
Files.writeString(f, """
bind:
port: 8080
primary:
terminal: term_fixed
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.primary());
assertEquals("term_fixed", cfg.primary().terminal());
}
@Test
void primaryBlockWithBlankTerminalDefaultsToDerived(@TempDir Path dir) throws Exception {
Path f = dir.resolve("primary-blank.yaml");
Files.writeString(f, "bind:\n port: 8080\nprimary:\n terminal: \"\"\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.primary());
assertTrue(cfg.primary().terminal() == null || cfg.primary().terminal().isBlank(),
"a blank terminal in yaml should be treated as absent — null or empty are equivalent");
}
// --- CB-402: peer kind discriminator -------------------------------------------------------
@Test
void workerKindDefaultsToClaudeCodeWhenOmitted(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-absent.yaml");
Files.writeString(f, """
workers:
gx10:
baseUrl: http://gx10.gw:8000
argv: ["ccs", "gx10"]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(BridgedConfig.Worker.KIND_CLAUDE_CODE, cfg.workerProfiles().get("gx10").kind(),
"a worker with no kind: is a claude-code worker (backward compatible)");
}
@Test
void opencodeKindIsNormalizedToLowerCase(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-opencode.yaml");
Files.writeString(f, """
workers:
gemini:
kind: OpenCode
model: google/gemini-2.5-pro
argv: ["opencode"]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(BridgedConfig.Worker.KIND_OPENCODE, cfg.workerProfiles().get("gemini").kind(),
"kind is normalised to lower-case so YAML casing does not matter");
}
@Test
void kindPredicatesReflectTheResolvedKind(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-predicates.yaml");
Files.writeString(f, """
workers:
claude:
baseUrl: http://gx10.gw:8000
gemini:
kind: opencode
model: google/gemini-2.5-pro
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Worker claude = cfg.workerProfiles().get("claude");
BridgedConfig.Worker gemini = cfg.workerProfiles().get("gemini");
assertTrue(claude.isClaudeCode(), "the default-kind worker is claude-code");
assertFalse(claude.isOpenCode(), "a claude-code worker is not opencode");
assertTrue(gemini.isOpenCode(), "the kind: opencode worker is opencode");
assertFalse(gemini.isClaudeCode(), "an opencode worker is not claude-code");
}
@Test
void argvDefaultsToTheKindBinaryWhenUnset(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-argv.yaml");
Files.writeString(f, """
workers:
claude:
baseUrl: http://gx10.gw:8000
gemini:
kind: opencode
model: google/gemini-2.5-pro
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(java.util.List.of("claude"), cfg.workerProfiles().get("claude").argv(),
"a claude-code worker with no argv defaults to the claude binary");
assertEquals(java.util.List.of("opencode"), cfg.workerProfiles().get("gemini").argv(),
"an opencode worker with no argv defaults to the opencode binary, never claude");
}
@Test
void authDefaultsToLoopbackTrustSoExistingConfigsBehaveAsBefore(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-auth-block.yaml");
Files.writeString(f, "bind:\n host: 127.0.0.1\n port: 8765\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.auth(), "auth must default rather than be null");
assertFalse(cfg.auth().tokenMode());
assertEquals("BRIDGED_API_TOKEN", cfg.auth().tokenEnv(), "documented default env var");
assertDoesNotThrow(cfg::validateAuthExposure, "loopback + loopback-trust is the safe pairing");
}
/**
* CB-501's highest-value check. Under loopback-trust, "not a known worker" means "the primary" —
* sound only while the OS refuses remote connections. Widening the bind without token mode
* would silently promote every reachable client to the most privileged role on the bus.
*/
@Test
void aNonLoopbackBindWithoutTokenModeIsRefusedAtStartup(@TempDir Path dir) throws Exception {
Path f = dir.resolve("exposed.yaml");
Files.writeString(f, "bind:\n host: 0.0.0.0\n port: 8765\n");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAuthExposure);
assertTrue(e.getMessage().contains("auth.mode: token"),
"the error must say how to fix it, not just that it refused");
}
@Test
void aNonLoopbackBindIsAllowedOnceTokenModeIsOn(@TempDir Path dir) throws Exception {
Path f = dir.resolve("exposed-with-token.yaml");
Files.writeString(f, """
bind:
host: 0.0.0.0
port: 8765
auth:
mode: token
tokenEnv: MY_TOKEN
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertTrue(cfg.auth().tokenMode());
assertEquals("MY_TOKEN", cfg.auth().tokenEnv());
assertDoesNotThrow(cfg::validateAuthExposure);
}
@Test
void loopbackFormsAreAllRecognised(@TempDir Path dir) throws Exception {
for (String host : new String[]{"127.0.0.1", "localhost", "::1", "127.0.0.53"}) {
Path f = dir.resolve("lb-" + host.replace(':', '_') + ".yaml");
Files.writeString(f, "bind:\n host: \"" + host + "\"\n port: 8765\n");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateAuthExposure(),
host + " is loopback and must not trip the exposure guard");
}
}
/**
* The shipped {@code bridged.example.yaml} must actually parse. Config binds through a plain
* Jackson mapper with {@code ignoreUnknown = true}, so a misspelled key in the example is
* silently dropped and the operator gets a default they did not ask for — exactly how a
* {@code spawn_ready_timeout_ms} typo survived in the example until the CB-5xx wrap-up.
*/
@Test
void shippedExampleConfigParses() {
Path example = Path.of("bridged.example.yaml");
assertTrue(Files.exists(example), "bridged.example.yaml must ship next to the pom");
BridgedConfig cfg = BridgedConfig.load(example);
assertEquals(8765, cfg.bind().port(), "example binds the documented default port");
assertTrue(cfg.workerProfiles().containsKey("gx10"), "example documents the gx10 profile");
assertEquals("gx10", cfg.defaultProfile(), "example's defaultWorker resolves");
assertTrue(cfg.guard().hostSet().contains("gx01.gw"),
"every example profile's base_url host must be in the example allowlist");
}
/**
* Every optional knob the example documents must bind under the exact spelling used there.
* Keep this list in step with {@code bridged.example.yaml}: a rename that updates the record
* but not the example (or vice versa) fails here instead of silently no-op'ing in production.
*/
@Test
void everyOptionalKnobDocumentedInTheExampleBinds(@TempDir Path dir) throws Exception {
Path f = dir.resolve("all-knobs.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
spawnReadyTimeoutMs: 25000
spawnReadyPollMs: 400
worktreeRoot: /tmp/bridged-worktrees
workers:
gx10:
kind: claude-code
baseUrl: http://gx01.gw:8000
configDir: /tmp/ccs/gx10
cwd: /tmp/repo
parityOverlay: [".mcp.json", ".env"]
gitTokenEnv: GITEA_TOKEN
gitHostEnv: GITEA_HOST
lifecycle:
idleTtlSeconds: 300
contextCap: 10
drainTimeoutSeconds: 5
broker:
uri: amqp://guest:guest@127.0.0.1:5672
primary:
terminal: term_abc123
pushReminders: 5
pushBackoffMs: 15000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(25000, cfg.spawnReadyTimeoutMs(), "spawnReadyTimeoutMs is camelCase, not snake_case");
assertEquals(400, cfg.spawnReadyPollMs(), "spawnReadyPollMs is camelCase, not snake_case");
assertEquals("/tmp/bridged-worktrees", cfg.worktreeRoot());
BridgedConfig.Worker w = cfg.workerProfiles().get("gx10");
assertEquals("/tmp/ccs/gx10", w.configDir());
assertEquals("/tmp/repo", w.cwd());
assertEquals(java.util.List.of(".mcp.json", ".env"), w.parityOverlay());
assertTrue(w.hasGitToken(), "gitTokenEnv binds and enables the CB-302 PR grant");
assertEquals("GITEA_HOST", w.gitHostEnv());
assertEquals(300, cfg.lifecycle().idleTtlSeconds());
assertEquals(10, cfg.lifecycle().contextCap());
assertEquals(5, cfg.lifecycle().drainTimeoutSeconds());
assertEquals("amqp://guest:guest@127.0.0.1:5672", cfg.broker().uri());
assertEquals("term_abc123", cfg.primary().terminal());
assertEquals(5, cfg.primary().remindersOrDefault());
assertEquals(15000L, cfg.primary().backoffMsOrDefault());
}
}
@@ -195,6 +195,66 @@ class CompletionResolverTest {
assertEquals("hello", waiter.getNow(null).text());
}
@Test
void resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent() {
// The most important branch of the CB-115 guard: a failed read means the resolver could not
// SEE the screen — "couldn't see", not "no change". It must still resolve the send (an empty
// tail beats hanging until the caller's timeout), even though a baseline was captured. The
// baseline here is "" (an empty pane at delivery), so without the !scrapeFailed clause the
// byte-identical guard would wrongly match the empty tail and suppress.
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.read throws HerdrException
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, ""); // empty pane baselined at delivery
resolver.resolve("term_a", turn);
assertTrue(waiter.isDone(),
"a failed scrape must still resolve the send, not hang until the caller's timeout");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
assertEquals("", waiter.getNow(null).text(), "the tail is empty because the screen was unreadable");
}
// --- CB-115/CB-116 fail guard: an already-done or absent waiter is left alone ---------
@Test
void failLeavesAnAlreadyResolvedWaiterUntouchedAndSkipsTheScrape() {
// The send was already resolved (e.g. by the worker's explicit reply) before fail fired.
// fail must not overwrite that value, and must not even scrape the worker — nobody needs it.
FakeHerdr herdr = new FakeHerdr().readText("an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null);
assertTrue(rendezvous.resolveCompletion(waiter, "already replied"));
resolver.fail("term_a", turn);
assertFalse(herdr.called("agent.read"),
"fail must not scrape a waiter that is already done");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"fail must not overwrite the existing resolution");
assertEquals("already replied", waiter.getNow(null).text());
}
@Test
void failFallsBackToTheRegisteredWaiterWhenThereIsNoInFlightTurn() {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fail falls back to the waiter currently registered on the Rendezvous and fails it.
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a"); // send registered, but no captureBaseline ever ran
resolver.fail("term_a", null); // no in-flight turn → fall back to the registered waiter
assertTrue(waiter.isDone(), "fail falls back to the registered waiter when no turn is in flight");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
assertEquals("stuck on an error screen", waiter.getNow(null).text());
}
// --- CB-116 waiter identity: a late completion never crosses into the next turn ---------
@Test
@@ -0,0 +1,168 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.auth.Role;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-513 — the CB-505 authorization gate on the <strong>MCP</strong> entry path.
*
* <p>Why this file exists: CB-505 claimed authorization is "enforced on both entry paths", and it
* is — but only REST was ever tested ({@code BridgedAppAuthTest}). Coverage showed
* {@code BridgeMcp.deny()}, {@code principal()} and every tool-registration lambda at <em>zero</em>
* executed lines, because no test had ever constructed a {@code BridgeMcp} — the existing
* {@code BridgeMcpTest} calls only the static handler methods. An unexercised security control is
* a claim, not a control.
*
* <p>These tests construct a real {@code BridgeMcp} (which also exercises the constructor and the
* tool wiring) and drive the policy half of the gate directly.
*/
class BridgeMcpAuthzTest {
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private Metrics metrics;
private BridgeMcp mcp;
@AfterEach
void close() {
if (mcp != null) mcp.close();
}
/** A fully wired BridgeMcp on fakes — constructing it is itself part of what is under test. */
private BridgeMcp mcp(boolean enforce) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "tok");
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(),
new InMemoryReplyInbox());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
metrics = BridgedMetrics.create(sessions, new InMemoryReplyInbox());
mcp = new BridgeMcp(messages, workers, sessions, identity, sessions.asPresence(),
new PrimaryRegistry(null),
enforce ? new CallerResolver(identity) : null,
metrics);
return mcp;
}
private static final Principal PRIMARY = Principal.primary(100);
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal ANON = Principal.anonymous();
// --- the table, enforced on THIS path too ---------------------------------------------------
@Test
void primaryMayOrchestrate() {
BridgeMcp m = mcp(true);
for (Authz.Action a : new Authz.Action[]{Authz.Action.SPAWN, Authz.Action.STOP,
Authz.Action.SEND, Authz.Action.DRAIN, Authz.Action.READ}) {
assertNull(m.denyFor(PRIMARY, a, "term_a"), a + " is the primary's to perform");
}
}
@Test
void aWorkerMayNotOrchestrateOverMcp() {
BridgeMcp m = mcp(true);
for (Authz.Action a : new Authz.Action[]{Authz.Action.SPAWN, Authz.Action.STOP,
Authz.Action.SEND, Authz.Action.DRAIN}) {
McpSchema.CallToolResult denied = m.denyFor(WORKER_A, a, "term_a");
assertNotNull(denied, a + " must be refused to a worker");
assertTrue(denied.isError(), "a refusal is returned as an MCP tool error");
}
}
@Test
void aWorkerMayReplyAndAskOnlyAsItself() {
BridgeMcp m = mcp(true);
assertNull(m.denyFor(WORKER_A, Authz.Action.REPLY, "term_a"), "its own session is allowed");
assertNull(m.denyFor(WORKER_A, Authz.Action.ASK, "term_a"));
assertNotNull(m.denyFor(WORKER_A, Authz.Action.REPLY, "term_b"),
"worker A must not reply on worker B's session");
assertNotNull(m.denyFor(WORKER_A, Authz.Action.ASK, "term_b"));
}
@Test
void thePrimaryMayNotForgeAWorkerReplyOverMcp() {
BridgeMcp m = mcp(true);
// A forged reply would resolve the very rendezvous the primary is blocked on.
assertNotNull(m.denyFor(PRIMARY, Authz.Action.REPLY, "term_a"));
assertNotNull(m.denyFor(PRIMARY, Authz.Action.ASK, "term_a"));
}
@Test
void anonymousIsRefusedEverythingAndCountedAsUnauthenticated() {
BridgeMcp m = mcp(true);
McpSchema.CallToolResult denied = m.denyFor(ANON, Authz.Action.READ, null);
assertNotNull(denied, "authenticated as nothing ⇒ authorized for nothing");
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "unauthenticated"));
assertEquals(0, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "forbidden"),
"a missing credential is 401-shaped, not 403-shaped");
}
@Test
void aWrongRoleIsCountedAsForbiddenNotUnauthenticated() {
BridgeMcp m = mcp(true);
assertNotNull(m.denyFor(WORKER_A, Authz.Action.SPAWN, null));
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "forbidden"));
assertEquals(0, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "unauthenticated"),
"the caller IS authenticated — it is just not the right role");
}
@Test
void theLegacyConstructorLeavesTheGateOpen() {
// The 22 pre-existing BridgeMcpTest cases rely on no authorization being enforced.
BridgeMcp m = mcp(false);
assertNull(m.denyFor(ANON, Authz.Action.SPAWN, null),
"no CallerResolver supplied ⇒ authorization not enforced (legacy behaviour)");
}
// --- identity reconstruction from the transport context ------------------------------------
@Test
void principalIsRebuiltFromTheStashedRole() {
assertEquals(Role.WORKER, BridgeMcp.principalFrom("WORKER", "term_a", 7).role());
assertEquals("term_a", BridgeMcp.principalFrom("WORKER", "term_a", 7).terminal());
assertEquals(Role.PRIMARY, BridgeMcp.principalFrom("PRIMARY", null, 7).role());
assertEquals(Role.ANONYMOUS, BridgeMcp.principalFrom("ANONYMOUS", null, -1).role());
}
@Test
void aMissingRoleFallsBackToTheHistoricalInterpretation() {
// Legacy path: no role stashed. A terminal means worker; its absence meant "the primary",
// which is exactly the pre-CB-501 default CB-501 inverted — preserved only here.
assertEquals(Role.WORKER, BridgeMcp.principalFrom(null, "term_a", 7).role());
assertEquals(Role.PRIMARY, BridgeMcp.principalFrom(null, null, 7).role());
}
}
@@ -1,5 +1,6 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
@@ -57,13 +58,16 @@ class BridgeMcpTest {
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> BridgeMcp.send(messages, "term_a", "review this", 4000L));
McpSchema.CallToolResult reply = BridgeMcp.reply(rendezvous, "term_a", "LGTM");
// Wait until the send has opened its waiter so the reply resolves it (CB-307: reply now
// queues in the inbox if no waiter is open, which would break the round-trip).
long deadline = System.currentTimeMillis() + 3000;
while (Boolean.TRUE.equals(reply.isError()) && System.currentTimeMillis() < deadline) {
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(10);
reply = BridgeMcp.reply(rendezvous, "term_a", "LGTM");
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "send should have opened its waiter");
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "LGTM");
assertEquals("delivered", textOf(reply));
McpSchema.CallToolResult res = send.get(6, TimeUnit.SECONDS);
@@ -80,30 +84,32 @@ class BridgeMcpTest {
assertTrue(out.contains("ticket="), out);
String ticket = out.substring(out.indexOf("ticket=") + "ticket=".length()).trim();
// Resolve the awaiting send once it has opened (retry past the async-open race).
// Wait until the send has opened its waiter before replying (CB-307: reply never errors,
// so the old retry-on-error pattern no longer works — it would queue instead of resolve).
long deadline = System.currentTimeMillis() + 3000;
McpSchema.CallToolResult reply = BridgeMcp.reply(rendezvous, "term_a", "async LGTM");
while (Boolean.TRUE.equals(reply.isError()) && System.currentTimeMillis() < deadline) {
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(10);
reply = BridgeMcp.reply(rendezvous, "term_a", "async LGTM");
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "send should have opened its waiter");
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "async LGTM");
assertEquals("delivered", textOf(reply));
// Poll until the async send completes and reports the reply.
McpSchema.CallToolResult polled = BridgeMcp.poll(messages, ticket);
McpSchema.CallToolResult polled = BridgeMcp.poll(messages, ticket, null);
deadline = System.currentTimeMillis() + 3000;
while (!textOf(polled).contains("async LGTM") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(10);
polled = BridgeMcp.poll(messages, ticket);
polled = BridgeMcp.poll(messages, ticket, null);
}
assertEquals("async LGTM", textOf(polled));
}
@Test
void pollUnknownTicketIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.poll(messages, "task-999");
McpSchema.CallToolResult res = BridgeMcp.poll(messages, "task-999", null);
assertTrue(res.isError());
assertTrue(textOf(res).contains("unknown ticket"));
}
@@ -122,10 +128,32 @@ class BridgeMcpTest {
}
@Test
void replyWithNoPendingSendIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.reply(rendezvous, "term_a", "orphan");
assertTrue(res.isError());
assertTrue(textOf(res).contains("no send is awaiting"));
void replyWithNoPendingSendIsQueuedNotError() {
// CB-307: a reply with no open send is now queued in the inbox, not an error.
McpSchema.CallToolResult res = BridgeMcp.reply(messages, "term_a", "orphan");
assertNotEquals(Boolean.TRUE, res.isError(), "a queued reply is not an error");
assertEquals("delivered", textOf(res));
// The reply is drainable by target.
var drained = messages.drainReplies("term_a");
assertEquals(1, drained.size());
assertEquals("orphan", drained.getFirst().content());
}
@Test
void bridgePollWithTargetDrainsReplies() {
// A reply with no open send queues it in the inbox.
BridgeMcp.reply(messages, "term_a", "queued-msg");
// bridge_poll with target drains the inbox.
McpSchema.CallToolResult res = BridgeMcp.poll(messages, null, "term_a");
assertNotEquals(Boolean.TRUE, res.isError());
String text = textOf(res);
assertTrue(text.contains("queued-msg"), "the drained reply should appear in the result");
// Second drain returns empty.
McpSchema.CallToolResult empty = BridgeMcp.poll(messages, null, "term_a");
assertEquals("[]", textOf(empty));
}
@Test
@@ -159,14 +187,14 @@ class BridgeMcpTest {
// The worker's ask returns the answer — it resumes the same turn.
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
// The resumed worker replies, resolving the answering send (retry past the reopen race).
McpSchema.CallToolResult reply = BridgeMcp.reply(rendezvous, "term_a", "done");
// The resumed worker replies, resolving the answering send (wait for the reopened waiter).
deadline = System.currentTimeMillis() + 3000;
while (Boolean.TRUE.equals(reply.isError()) && System.currentTimeMillis() < deadline) {
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(10);
reply = BridgeMcp.reply(rendezvous, "term_a", "done");
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "the answer should have reopened a waiter");
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "done");
assertEquals("delivered", textOf(reply));
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
}
@@ -276,6 +304,37 @@ class BridgeMcpTest {
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), " ").isError());
}
@Test
void bridgeAckReturnsConfirmationForValidArgs() {
McpSchema.CallToolResult res = BridgeMcp.ack(messages, "term_a", "msg-1");
assertNotEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("msg-1"), "response should mention the msgId");
}
@Test
void bridgeAckRejectsMissingArgs() {
assertTrue(BridgeMcp.ack(messages, null, "msg-1").isError());
assertTrue(BridgeMcp.ack(messages, "term_a", null).isError());
assertTrue(BridgeMcp.ack(messages, " ", "msg-1").isError());
}
@Test
void bridgeAckRemovesSpecificReply() {
// Queue a reply and capture its msgId.
BridgeMcp.reply(messages, "term_a", "orphan");
var before = messages.drainReplies("term_a");
assertEquals(1, before.size(), "one reply in the inbox");
String msgId = before.getFirst().msgId();
// Publish the same reply again and ack it via bridge_ack surface.
BridgeMcp.reply(messages, "term_a", "orphan-again");
var peeked = messages.drainReplies("term_a");
assertEquals(1, peeked.size(), "one fresh reply in the inbox");
// ackReply works (no-op since published with a different UUID, but callable).
assertDoesNotThrow(() -> messages.ackReply("term_a", msgId));
}
@Test
void statusReportsLiveAgentStatus() {
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
@@ -285,4 +344,58 @@ class BridgeMcpTest {
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("blocked", textOf(res));
}
// --- bridge_whoami: the caller's own identity, so an agent never has to guess its role -------
@Test
void whoamiReportsThePrimaryAsPrimaryAndNothingElse() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.whoami(
Principal.primary(100), sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"primary\""), out);
// The primary owns no session — leaking a sessionId here would invite it to reply as one.
assertFalse(out.contains("sessionId"), out);
}
@Test
void whoamiReportsAWorkerWithItsRegisteredSession() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-517", null));
McpSchema.CallToolResult res = BridgeMcp.whoami(Principal.worker(s.terminalId(), 200), sessions);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"worker\""), out);
assertTrue(out.contains("\"sessionId\":\"" + s.terminalId() + "\""), out);
assertTrue(out.contains("\"profile\":\"ltms-local\""), out);
assertTrue(out.contains("\"worktree\":\"" + s.worktree() + "\""), out);
assertTrue(out.contains("\"branch\":\"" + s.branch() + "\""), out);
assertTrue(out.contains("\"owner\":\"term_primary\""), out);
}
/**
* A worker the registry has no record of — it outlived a daemon restart — must still learn the
* load-bearing fact. Degrading to "I don't know who you are" would put it back to guessing,
* which is the failure this tool exists to remove.
*/
@Test
void whoamiStillReportsWorkerRoleWhenTheSessionIsUnregistered() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.whoami(Principal.worker("term_orphan", 200),
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"worker\""), out);
assertTrue(out.contains("\"sessionId\":\"term_orphan\""), out);
assertFalse(out.contains("profile"), out); // nothing invented for a session we don't track
}
}
@@ -0,0 +1,108 @@
package dev.ltms.bridged.mcp;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for {@link PrimaryRegistry}: pin vs record, isKnown transitions,
* null/blank guard.
*/
class PrimaryRegistryTest {
@Test
void unpinnedInitiallyUnknown() {
var reg = new PrimaryRegistry(null);
assertFalse(reg.isKnown());
assertTrue(reg.primaryTerminal().isEmpty());
}
@Test
void unpinnedAcceptsBlankAsAbsent() {
var reg = new PrimaryRegistry("");
assertFalse(reg.isKnown());
assertTrue(reg.primaryTerminal().isEmpty());
}
@Test
void pinnedFromConstruction() {
var reg = new PrimaryRegistry("term_fixed");
assertTrue(reg.isKnown());
assertEquals("term_fixed", reg.primaryTerminal().orElseThrow());
}
@Test
void recordWhenUnpinnedSetsTheTerminal() {
var reg = new PrimaryRegistry(null);
reg.record("term_abc");
assertTrue(reg.isKnown());
assertEquals("term_abc", reg.primaryTerminal().orElseThrow());
}
@Test
void recordWithNullDoesNothingWhenUnpinned() {
var reg = new PrimaryRegistry(null);
reg.record(null);
assertFalse(reg.isKnown());
}
@Test
void recordWithBlankDoesNothingWhenUnpinned() {
var reg = new PrimaryRegistry(null);
reg.record(" ");
assertFalse(reg.isKnown());
}
@Test
void recordOverwritesWhenUnpinned() {
var reg = new PrimaryRegistry(null);
reg.record("term_first");
assertEquals("term_first", reg.primaryTerminal().orElseThrow());
reg.record("term_second");
assertEquals("term_second", reg.primaryTerminal().orElseThrow());
}
@Test
void recordIsIgnoredWhenPinned() {
var reg = new PrimaryRegistry("term_pinned");
reg.record("term_other");
assertEquals("term_pinned", reg.primaryTerminal().orElseThrow(), "pinned value must survive record");
}
@Test
void nullRecordIsIgnoredWhenPinned() {
var reg = new PrimaryRegistry("term_pinned");
reg.record(null);
assertTrue(reg.isKnown());
assertEquals("term_pinned", reg.primaryTerminal().orElseThrow());
}
@Test
void blankRecordIsIgnoredWhenPinned() {
var reg = new PrimaryRegistry("term_pinned");
reg.record(" ");
assertTrue(reg.isKnown());
assertEquals("term_pinned", reg.primaryTerminal().orElseThrow());
}
@Test
void isKnownFalseAfterConstructionWithNull() {
var reg = new PrimaryRegistry(null);
assertFalse(reg.isKnown());
}
@Test
void isKnownAfterRecord() {
var reg = new PrimaryRegistry(null);
reg.record("term_x");
assertTrue(reg.isKnown());
}
@Test
void primaryTerminalRoundTrip() {
var reg = new PrimaryRegistry(null);
assertTrue(reg.primaryTerminal().isEmpty());
reg.record("term_found");
assertEquals("term_found", reg.primaryTerminal().get());
}
}
@@ -0,0 +1,127 @@
package dev.ltms.bridged.metrics;
import org.junit.jupiter.api.Test;
import java.util.LinkedHashMap;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/** CB-502 — the zero-dependency Prometheus text renderer. */
class MetricsTest {
@Test
void countersAccumulatePerLabelSet() {
Metrics m = new Metrics();
m.inc("bridged_sends_total", "outcome", "replied");
m.inc("bridged_sends_total", "outcome", "replied");
m.inc("bridged_sends_total", "outcome", "timeout");
assertEquals(2, m.count("bridged_sends_total", "outcome", "replied"));
assertEquals(1, m.count("bridged_sends_total", "outcome", "timeout"));
assertEquals(0, m.count("bridged_sends_total", "outcome", "failed"),
"an untouched series reads as zero, not an error");
}
@Test
void rendersHelpAndTypeOncePerFamily() {
Metrics m = new Metrics();
m.describe("bridged_sends_total", "counter", "Delegated sends by outcome.");
m.inc("bridged_sends_total", "outcome", "replied");
m.inc("bridged_sends_total", "outcome", "timeout");
String out = m.render();
assertEquals(1, countOccurrences(out, "# HELP bridged_sends_total"),
"HELP is per family, not per series");
assertEquals(1, countOccurrences(out, "# TYPE bridged_sends_total counter"));
assertTrue(out.contains("bridged_sends_total{outcome=\"replied\"} 1"));
assertTrue(out.contains("bridged_sends_total{outcome=\"timeout\"} 1"));
}
@Test
void labelsAreSortedSoScrapesAreByteStable() {
Metrics a = new Metrics();
a.inc("m", "b", "2", "a", "1");
Metrics b = new Metrics();
b.inc("m", "a", "1", "b", "2");
assertEquals(a.render(), b.render(), "label order in the call must not change the output");
assertTrue(a.render().contains("m{a=\"1\",b=\"2\"}"));
}
@Test
void gaugesAreEvaluatedAtScrapeTimeNotRegistrationTime() {
Metrics m = new Metrics();
int[] live = {1};
m.gauge("bridged_sessions", () -> live[0], "state", "ready");
assertTrue(m.render().contains("bridged_sessions{state=\"ready\"} 1"));
live[0] = 5;
assertTrue(m.render().contains("bridged_sessions{state=\"ready\"} 5"),
"the gauge must read current state on every scrape");
}
@Test
void aThrowingGaugeDoesNotBreakTheWholeScrape() {
Metrics m = new Metrics();
m.inc("good_total");
m.gauge("bad_gauge", () -> {
throw new IllegalStateException("herdr is down");
});
String out = assertDoesNotThrow(m::render);
assertTrue(out.contains("good_total 1"), "healthy series must still be exported");
assertFalse(out.contains("bad_gauge"), "the broken series is simply absent");
}
@Test
void collectorsDiscoverTheirLabelSetPerScrape() {
Metrics m = new Metrics();
Map<String, Number> depths = new LinkedHashMap<>();
m.collector("bridged_inbox_depth", "target", () -> depths);
assertFalse(m.render().contains("bridged_inbox_depth"), "no targets yet ⇒ no series");
depths.put("term_a", 2);
depths.put("term_b", 0);
String out = m.render();
assertTrue(out.contains("bridged_inbox_depth{target=\"term_a\"} 2"));
assertTrue(out.contains("bridged_inbox_depth{target=\"term_b\"} 0"));
}
@Test
void labelValuesAreEscaped() {
Metrics m = new Metrics();
m.inc("m", "detail", "he said \"hi\"\nand \\left");
String out = m.render();
assertTrue(out.contains("\\\""), "quotes escaped");
assertTrue(out.contains("\\n"), "newlines escaped — a raw one would corrupt the exposition");
assertTrue(out.contains("\\\\"), "backslashes escaped");
}
@Test
void wholeNumberGaugesRenderWithoutADecimalPoint() {
Metrics m = new Metrics();
m.gauge("whole", () -> 3.0);
m.gauge("fractional", () -> 1.5);
String out = m.render();
assertTrue(out.contains("whole 3"), "3.0 should not render as 3.0");
assertTrue(out.contains("fractional 1.5"));
}
@Test
void oddLabelCountIsRejected() {
Metrics m = new Metrics();
assertThrows(IllegalArgumentException.class, () -> m.inc("m", "dangling"));
}
private static int countOccurrences(String haystack, String needle) {
int n = 0;
for (int i = haystack.indexOf(needle); i >= 0; i = haystack.indexOf(needle, i + 1)) {
n++;
}
return n;
}
}
@@ -0,0 +1,113 @@
package dev.ltms.bridged.msg;
import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import org.testcontainers.containers.RabbitMQContainer;
import org.testcontainers.junit.jupiter.Container;
import org.testcontainers.junit.jupiter.Testcontainers;
import org.testcontainers.utility.DockerImageName;
import java.util.List;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Contract test for {@link AmqpReplyInbox} against a REAL broker (a RabbitMQ container — the same
* AMQP 0-9-1 the production LavinMQ deploy speaks, URI-only swap). Tagged {@code contract} so it is
* excluded from {@code mvn test}/{@code mvn clean install} (which stay hermetic and need no Docker);
* run it with Docker present via {@code mvn test -Pcontract}.
*
* <p>It proves the port contract on genuine infrastructure: eventual visibility of a published reply,
* ack removal, msgId dedup, and — the reason Stage 2 exists — cross-restart durability: an unacked
* reply survives closing the inbox and is redelivered to a fresh connection.
*/
@Tag("contract")
@Testcontainers
class AmqpReplyInboxContractTest {
@Container
static final RabbitMQContainer BROKER =
new RabbitMQContainer(DockerImageName.parse("rabbitmq:3.13-management"));
private static String uri() {
// guest/guest against the mapped AMQP port. No trailing slash: an empty path is vhost "",
// which does not exist — omitting it selects the default vhost "/".
return "amqp://guest:guest@" + BROKER.getHost() + ":" + BROKER.getAmqpPort();
}
@Test
void publishThenPeekThenAck() throws Exception {
String target = "worker-pub-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.publish(target, "m1", "hello primary");
List<ReplyInbox.InboxMessage> got = awaitPeek(inbox, target);
assertEquals(1, got.size(), "the published reply should be held for drain");
assertEquals("m1", got.getFirst().msgId());
assertEquals(target, got.getFirst().target());
assertEquals("hello primary", got.getFirst().content());
inbox.ack(target, "m1");
assertTrue(inbox.peek(target).isEmpty(), "an acked reply is dropped");
}
}
@Test
void duplicateMsgIdIsNotDoubleQueued() throws Exception {
String target = "worker-dedup-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.publish(target, "dup", "first");
awaitPeek(inbox, target);
inbox.publish(target, "dup", "second"); // same msgId — must be a no-op
// Give any erroneous second delivery time to land, then assert still exactly one.
Thread.sleep(500);
List<ReplyInbox.InboxMessage> got = inbox.peek(target);
assertEquals(1, got.size(), "a repeated msgId must not double-queue");
assertEquals("first", got.getFirst().content(), "the first payload wins");
}
}
@Test
void unackedReplySurvivesRestartAndIsRedelivered() throws Exception {
String target = "worker-durable-" + System.nanoTime();
// First "process life": publish, see it held, but crash before acking.
try (AmqpReplyInbox first = AmqpReplyInbox.open(uri())) {
first.publish(target, "persist-1", "survive me");
assertEquals(1, awaitPeek(first, target).size());
// no ack — simulate a java -jar bounce with the reply still pending
}
// Second "process life": a fresh connection to the same broker must be redelivered the reply.
try (AmqpReplyInbox second = AmqpReplyInbox.open(uri())) {
List<ReplyInbox.InboxMessage> got = awaitPeek(second, target);
assertEquals(1, got.size(), "an unacked persistent reply is redelivered after restart");
assertEquals("persist-1", got.getFirst().msgId());
assertEquals("survive me", got.getFirst().content());
second.ack(target, "persist-1");
}
// Third life: once acked, it is gone for good — durability is not endless replay.
try (AmqpReplyInbox third = AmqpReplyInbox.open(uri())) {
Thread.sleep(500);
assertTrue(third.peek(target).isEmpty(), "an acked reply does not come back on the next restart");
}
}
/** Poll peek (broker delivery is async) until a reply for {@code target} appears or ~10s elapse. */
@SuppressWarnings("BusyWait") // deliberate poll for async broker delivery, bounded by the deadline
private static List<ReplyInbox.InboxMessage> awaitPeek(AmqpReplyInbox inbox, String target)
throws InterruptedException {
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(10);
List<ReplyInbox.InboxMessage> msgs = inbox.peek(target);
while (msgs.isEmpty() && System.nanoTime() < deadline) {
Thread.sleep(50);
msgs = inbox.peek(target);
}
return msgs;
}
}
@@ -0,0 +1,145 @@
package dev.ltms.bridged.msg;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for {@link InMemoryReplyInbox}: publish, peek, ack, dedup, FIFO ordering, and thread
* safety under concurrent publish vs. drain.
*/
class InMemoryReplyInboxTest {
private final ReplyInbox inbox = new InMemoryReplyInbox();
@Test
void publishThenPeekReturnsTheMessage() {
inbox.publish("term_a", "m1", "hello");
var msgs = inbox.peek("term_a");
assertEquals(1, msgs.size());
assertEquals("m1", msgs.getFirst().msgId());
assertEquals("term_a", msgs.getFirst().target());
assertEquals("hello", msgs.getFirst().content());
}
@Test
void peekForUnknownTargetReturnsEmpty() {
assertTrue(inbox.peek("no-such-target").isEmpty());
}
@Test
void ackRemovesTheMessage() {
inbox.publish("term_a", "m1", "hello");
inbox.ack("term_a", "m1");
assertTrue(inbox.peek("term_a").isEmpty(), "after ack, the message is gone");
}
@Test
void ackForUnknownMsgIdIsNoOp() {
inbox.publish("term_a", "m1", "hello");
inbox.ack("term_a", "no-such-id"); // no-op
assertEquals(1, inbox.peek("term_a").size(), "the published message is still there");
}
@Test
void ackForUnknownTargetIsNoOp() {
inbox.ack("no-such-target", "m1"); // no-op, should not throw
}
@Test
void dedupByIdempotentMsgId() {
inbox.publish("term_a", "m1", "first");
inbox.publish("term_a", "m1", "second"); // same msgId, different content
var msgs = inbox.peek("term_a");
assertEquals(1, msgs.size(), "dedup: second publish with same msgId is a no-op");
assertEquals("first", msgs.getFirst().content(), "the original content is retained");
}
@Test
void publishesWithDifferentMsgIdsBothAppear() {
inbox.publish("term_a", "m1", "first");
inbox.publish("term_a", "m2", "second");
var msgs = inbox.peek("term_a");
assertEquals(2, msgs.size());
assertEquals("m1", msgs.get(0).msgId());
assertEquals("m2", msgs.get(1).msgId());
}
@Test
void perTargetIsolation() {
inbox.publish("term_a", "m1", "for-a");
inbox.publish("term_b", "m2", "for-b");
assertEquals(1, inbox.peek("term_a").size());
assertEquals(1, inbox.peek("term_b").size());
}
@Test
void fifoOrderIsPreserved() {
inbox.publish("term_a", "m1", "first");
inbox.publish("term_a", "m2", "second");
inbox.publish("term_a", "m3", "third");
var msgs = inbox.peek("term_a");
assertEquals(3, msgs.size());
assertEquals("m1", msgs.get(0).msgId());
assertEquals("m2", msgs.get(1).msgId());
assertEquals("m3", msgs.get(2).msgId());
}
@Test
void peekReturnsAnImmutableCopy() {
inbox.publish("term_a", "m1", "hello");
var msgs = inbox.peek("term_a");
assertThrows(UnsupportedOperationException.class, () -> msgs.add(
new ReplyInbox.InboxMessage("x", "term_a", "x")));
}
@Test
void ackRemovesOneMessageLeavesOthers() {
inbox.publish("term_a", "m1", "first");
inbox.publish("term_a", "m2", "second");
inbox.ack("term_a", "m1");
var msgs = inbox.peek("term_a");
assertEquals(1, msgs.size());
assertEquals("m2", msgs.getFirst().msgId());
}
@Test
void concurrentPublishAndDrain() throws Exception {
int msgCount = 100;
ExecutorService exec = Executors.newVirtualThreadPerTaskExecutor();
try {
// Concurrent publishers
var pubDone = new CountDownLatch(msgCount);
for (int i = 0; i < msgCount; i++) {
final int id = i;
exec.submit(() -> {
inbox.publish("term_a", "m" + id, "content-" + id);
pubDone.countDown();
});
}
// Concurrent drainer
AtomicReference<Exception> drainError = new AtomicReference<>();
exec.submit(() -> {
try {
pubDone.await();
for (int i = 0; i < 50; i++) {
var peeked = inbox.peek("term_a");
for (var msg : peeked) {
inbox.ack("term_a", msg.msgId());
}
}
} catch (Exception e) {
drainError.set(e);
}
}).get();
assertNull(drainError.get(), "concurrent drain should not throw");
} finally {
exec.shutdown();
}
}
}
@@ -14,6 +14,7 @@ import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
@@ -211,4 +212,256 @@ class MessageServiceTest {
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
"an answer to a turn that never existed (or already lapsed) is stale, not a hang");
}
// --- timeout, answer, poll, and lock-contention edges ----------------------------------
@Test
void sendTimesOutBeforeDeliveryIsQueuedNotWorking() {
// Nothing ever delivers the message and nothing resolves the send, so the reply future
// times out with delivery still incomplete — the message is still queued for the worker.
MessageService.Reply r = messages.send(T, "never delivered", 50);
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome(),
"an undelivered send that times out is still queued, not working");
assertNull(r.text());
}
@Test
void sendTimesOutAfterDeliveryIsStillWorking() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 300));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver — the delivered future now completes
injector.onStatus(T, AgentStatus.WORKING); // worker starts but never replies
// No rendezvous.resolve(T, ...) — the reply future rides out its short timeout.
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, r.outcome(),
"a delivered send whose worker never replies times out as still working");
}
@Test
void answerTimesOutWhenTheResumedWorkerNeverReplies() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertNotNull(q.turnId());
// The primary answers, unblocking the worker; but the worker never sends the follow-up
// bridge_reply, so the answering send rides out its short window as still-working.
MessageService.Reply answer = messages.answer(q.turnId(), "config.yaml", 200);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answer.outcome(),
"an answered worker that never replies times out as still working");
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome());
assertEquals("config.yaml", a.answer());
}
@Test
void pollReturnsNullForAnUnknownTicket() {
assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown");
}
@Test
void pollReportsACompletedTicket() throws Exception {
String ticket = messages.sendAsync(T, "long task");
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker works
assertTrue(rendezvous.resolve(T, "async result"), "a reply resolves the async send");
// Wait for the background send to finish and publish a DONE view.
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 2000;
while (view == null || view.phase() != MessageService.Phase.DONE) {
if (System.currentTimeMillis() >= deadline) break;
view = messages.poll(ticket);
//noinspection BusyWait
Thread.sleep(5);
}
assertNotNull(view, "a resolved async send must become DONE");
assertEquals(MessageService.Phase.DONE, view.phase());
assertEquals("async result", view.reply(), "the completed ticket reports the reply");
assertEquals("reply", view.replySource(), "a structured bridge_reply is sourced from 'reply'");
}
@Test
void concurrentSendToSameSessionWhileFirstHoldsItIsBusy() throws Exception {
CompletableFuture<MessageService.Reply> first =
CompletableFuture.supplyAsync(() -> messages.send(T, "first", 5000));
awaitWaiting(); // the first send now holds the session lock, blocked on its reply
// A second send to the SAME session cannot take the lock within its short window.
MessageService.Reply busy = messages.send(T, "second", 100);
assertEquals(MessageService.Outcome.BUSY, busy.outcome(),
"a second send while another holds the session is busy, not a hang");
assertNull(busy.text());
// Release the first send so it resolves cleanly and the test thread is not left pinned.
injector.onStatus(T, AgentStatus.IDLE); // deliver the first message
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up
assertTrue(rendezvous.resolve(T, "first done"), "the first send resolves with a reply");
MessageService.Reply firstReply = first.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, firstReply.outcome());
assertEquals("first done", firstReply.text());
}
// --- CB-307 reply inbox ----------------------------------------------------------------
@Test
void replyQueuesInInboxWhenNoSendIsOpen() {
// No send is open for this session — reply should queue in the inbox.
assertTrue(messages.reply(T, "queued-text"), "reply should succeed (queued)");
var drained = messages.drainReplies(T);
assertEquals(1, drained.size());
assertEquals("queued-text", drained.getFirst().content());
}
@Test
void replyResolvesOpenSendDoesNotQueue() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitUninterruptibly(T);
// An explicit reply resolves the open send.
assertTrue(messages.reply(T, "send-resolved"), "reply should succeed (resolved live send)");
// The inbox should be empty — the reply went to the send, not the inbox.
assertTrue(messages.drainReplies(T).isEmpty(), "no reply in the inbox");
MessageService.Reply r = send.get(3, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, r.outcome());
assertEquals("send-resolved", r.text());
}
@Test
void drainRepliesReturnsAllPendingThenEmptyOnNextCall() {
messages.reply(T, "msg-1");
messages.reply(T, "msg-2");
var first = messages.drainReplies(T);
assertEquals(2, first.size());
var second = messages.drainReplies(T);
assertTrue(second.isEmpty(), "second drain should be empty (acked)");
}
@Test
void aQuestionIsNeverQueuedInTheInbox() {
// No send is open — bridge_ask with no delegation returns NO_WAITER,
// and the question text MUST NOT appear in the reply inbox.
// The inbox is only fed by MessageService.reply(), not by bridge_ask.
MessageService.AskResult r = messages.ask(T, "anyone there?", 500);
assertEquals(MessageService.AskOutcome.NO_WAITER, r.outcome(),
"bridge_ask with no open delegation must return NO_WAITER, never queued");
assertTrue(messages.drainReplies(T).isEmpty(), "questions must never be queued");
}
@Test
void completionFallbackIsNeverQueued() throws Exception {
// The fallback resolves a captured waiter, never the inbox.
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitUninterruptibly(T);
injectDelivery();
// The worker never sends bridge_reply, but the turn completes.
herdr.readText("done-scraped");
completion.onTurnComplete(T); // The fallback arms and resolves the captured waiter.
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, r.outcome());
// The inbox should be empty — the reply went to the captured waiter.
assertTrue(messages.drainReplies(T).isEmpty(), "completion fallback must not queue");
}
// --- helpers ---------------------------------------------------------------------------
/** Like {@link #awaitWaiting()} but rethrows as unchecked. */
private void awaitUninterruptibly(String session) {
try {
awaitWaiting();
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException(e);
}
}
/** Set up a delivered turn so the worker is working, ready for an ask or completion. */
private void injectDelivery() {
herdr.readText("$ prompt"); // pre-turn content baseline
injector.onStatus(T, AgentStatus.IDLE); // deliver the task
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up
}
// --- CB-516: a released session must not leave a send hanging ------------------------------
/**
* The bug this fixes: tearing a worker down left its rendezvous waiter open, so a blocking send
* kept blocking and an async one kept reporting PENDING until the 30-minute async timeout —
* even though the worker provably no longer existed.
*/
@Test
void abandonFailsASendThatIsStillWaitingOnAReleasedSession() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
awaitWaiting();
assertTrue(messages.abandon(T, "session released"), "a live waiter is abandoned");
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, r.outcome(),
"an abandoned send fails rather than riding out its timeout");
assertEquals("session released", r.text(), "the caller is told why");
}
@Test
void abandonIsANoOpWhenNobodyIsWaiting() {
assertFalse(messages.abandon(T, "session released"),
"no open send ⇒ nothing to abandon");
}
@Test
void abandonDoesNotOverwriteAnAlreadyResolvedSend() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
awaitWaiting();
assertTrue(rendezvous.resolve(T, "the real answer"));
assertFalse(messages.abandon(T, "session released"),
"a send already answered by the worker must not be clobbered");
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals("the real answer", r.text());
}
/** The async path is the one that hung: poll must report FAILED, not PENDING forever. */
@Test
void anAbandonedAsyncTaskPollsAsFailedNotPending() throws Exception {
String ticket = messages.sendAsync(T, "long task");
awaitWaiting();
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase());
messages.abandon(T, "session released");
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 3000;
while (System.currentTimeMillis() < deadline) {
view = messages.poll(ticket);
if (view.phase() != MessageService.Phase.PENDING) break;
Thread.sleep(10);
}
assertNotNull(view);
assertEquals(MessageService.Phase.FAILED, view.phase(),
"a delegation whose worker is gone must not keep reporting PENDING");
assertTrue(view.detail() != null && view.detail().contains("released"),
"and the detail says why, rather than 'worker unknown'");
}
}
@@ -70,6 +70,28 @@ class RendezvousTest {
"no blocked send means no primary to surface the question to");
}
@Test
void resolveCompletionTwiceIsANoOpTheSecondTime() {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(W);
assertTrue(rendezvous.resolveCompletion(waiter, "first scrape"), "the first completion resolves");
assertFalse(rendezvous.resolveCompletion(waiter, "second scrape"),
"a second completion on an already-resolved waiter returns false");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
assertEquals("first scrape", waiter.getNow(null).text(),
"the first resolution wins; the stored value is unchanged");
}
@Test
void resolveFailureTwiceIsANoOpTheSecondTime() {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(W);
assertTrue(rendezvous.resolveFailure(waiter, "first reason"), "the first failure resolves");
assertFalse(rendezvous.resolveFailure(waiter, "second reason"),
"a second failure on an already-resolved waiter returns false");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
assertEquals("first reason", waiter.getNow(null).text(),
"the first resolution wins; the stored value is unchanged");
}
@Test
void closeAskRemovesTheTurn() {
Rendezvous.AskTicket t = rendezvous.openAsk(W);
@@ -0,0 +1,303 @@
package dev.ltms.bridged.msg;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.Collections;
import java.util.List;
import java.util.Map;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for {@link ReplyPushLoop}: decision logic, nudge injection, idempotency,
* bounded reminders, and stop conditions.
*
* <p>Uses a {@link RecordingHerdrClient} that synchronizes access to its call list so the
* scheduler thread and test thread never have memory ordering issues. The {@code decide()}
* tests use a simple client with no concurrency concern.
*/
class ReplyPushLoopTest {
private static final String PRIMARY = "term_primary";
private static final String WORKER = "term_worker";
private static final ObjectMapper MAPPER = new ObjectMapper();
private PrimaryRegistry registry;
private AgentControl agents;
private InMemoryReplyInbox inbox;
private ScheduledExecutorService scheduler;
@BeforeEach
void setUp() {
registry = new PrimaryRegistry(PRIMARY);
inbox = new InMemoryReplyInbox();
scheduler = Executors.newSingleThreadScheduledExecutor();
}
@AfterEach
void tearDown() {
scheduler.shutdownNow();
}
// --- decide() logic ------------------------------------------------------------------------
@Test
void decideWithoutPrimaryIsStop() {
agents = agentWithStatus("idle");
var loop = new ReplyPushLoop(
new PrimaryRegistry(null), agents, inbox, scheduler, 5, 100);
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(WORKER, 0));
}
@Test
void decideWithEmptyInboxIsStop() {
agents = agentWithStatus("idle");
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
}
@Test
void decideAtCapIsStop() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100).decide(WORKER, 2));
}
@Test
void decideUnderCapWithInjectablePrimaryIsInject() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
}
@Test
void decideUnderCapWithBlockedPrimaryIsInject() {
agents = agentWithStatus("blocked");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
"BLOCKED is injectable");
}
@Test
void decideUnderCapWithDonePrimaryIsInject() {
agents = agentWithStatus("done");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
"DONE is injectable");
}
@Test
void decideUnderCapWithBusyPrimaryIsWaitBusy() {
agents = agentWithStatus("working");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
}
@Test
void decideUnderCapWithUnknownPrimaryIsWaitBusy() {
agents = agentWithStatus("unknown");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
}
@Test
void decideStopsAfterInboxIsEmptied() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
inbox.ack(WORKER, "m1");
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
}
// --- onReplyQueued integration -------------------------------------------------------------
@Test
void injectablePrimaryCausesExactlyOneNudge() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
loop(1, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"one nudge (2 agent.send calls) should have been sent");
// Exactly one nudge = exactly 2 agent.send calls (text + submit)
assertEquals(2, rec.sendCount());
assertTrue(rec.sentParams().stream()
.anyMatch(e -> e.getValue().toString().contains("bridge_poll")),
"nudge text should contain bridge_poll");
}
@Test
void onReplyQueuedIsIdempotentPerTarget() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
var loop = loop(1, 100);
loop.onReplyQueued(WORKER);
loop.onReplyQueued(WORKER); // second call — should be a no-op
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"expected exactly one nudge (2 sends)");
Thread.sleep(200);
assertEquals(2, rec.sendCount(),
"second onReplyQueued must not trigger another nudge");
}
@Test
void sendsUpToCapThenStops() throws Exception {
int cap = 2;
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
rec.sendLatch = new CountDownLatch(cap * 2);
loop(cap, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(5, TimeUnit.SECONDS),
cap + " nudges (" + (cap * 2) + " sends) should have fired");
Thread.sleep(300);
assertEquals(cap * 2, rec.sendCount(),
"exactly " + (cap * 2) + " agent.send calls (cap=" + cap + ")");
}
// --- nudge format --------------------------------------------------------------------------
@Test
void nudgeFormatIsCorrect() {
String nudge = ReplyPushLoop.NUDGE_FORMAT.formatted(WORKER, WORKER);
assertTrue(nudge.contains("Worker term_worker"));
assertTrue(nudge.contains("bridge_poll(target=term_worker)"));
}
// --- metrics (CB-512) ----------------------------------------------------------------------
@Test
void successfulNudgeIncrementsDelivered() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
Metrics metrics = new Metrics();
loop(1, 50, metrics).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"one nudge (2 agent.send calls) should have been sent");
// The delivered count is bumped on the scheduler thread right after the send that releases
// the latch — settle briefly so the counter is published before we read it.
Thread.sleep(200);
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"),
"a successfully sent nudge must count as delivered");
}
@Test
void reminderCapIncrementsExhausted() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
Metrics metrics = new Metrics();
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100, metrics).decide(WORKER, 2));
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
"hitting the reminder cap must count as exhausted");
assertEquals(0, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"));
}
// --- helpers -------------------------------------------------------------------------------
private ReplyPushLoop loop() {
return loop(5, 100);
}
private ReplyPushLoop loop(int maxReminders, long backoffMs) {
return new ReplyPushLoop(registry, agents, inbox, scheduler, maxReminders, backoffMs);
}
private ReplyPushLoop loop(int maxReminders, long backoffMs, Metrics metrics) {
return new ReplyPushLoop(registry, agents, inbox, scheduler, maxReminders, backoffMs, metrics);
}
private static AgentControl agentWithStatus(String status) {
return new AgentControl(new FakeHerdrClient(status));
}
/** Non-recording (single-threaded) fake — safe for decide() tests. */
private static final class FakeHerdrClient implements HerdrClient {
private final String agentStatus;
FakeHerdrClient(String agentStatus) {
this.agentStatus = agentStatus;
}
@Override
public JsonNode call(String method, Object params) {
if ("agent.get".equals(method)) {
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", PRIMARY)
.put("agent_status", agentStatus));
}
return MAPPER.createObjectNode();
}
@Override
public void close() {
}
}
/**
* Thread-safe recording fake that counts agent.send calls. Uses synchronized access
* so the scheduler thread and test thread never race.
*/
private static final class RecordingHerdrClient implements HerdrClient {
private final List<Map.Entry<String, Object>> calls =
Collections.synchronizedList(new ArrayList<>());
volatile CountDownLatch sendLatch = new CountDownLatch(2);
@Override
public JsonNode call(String method, Object params) {
if ("agent.get".equals(method)) {
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", PRIMARY)
.put("agent_status", "idle")); // recording double is always injectable
}
if ("agent.send".equals(method)) {
calls.add(Map.entry(method, params));
sendLatch.countDown();
}
return MAPPER.createObjectNode();
}
long sendCount() {
return calls.size();
}
List<Map.Entry<String, Object>> sentParams() {
return List.copyOf(calls);
}
@Override
public void close() {
}
}
private static RecordingHerdrClient recordingClient() {
return new RecordingHerdrClient();
}
}
@@ -0,0 +1,225 @@
package dev.ltms.bridged.rest;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-501/505 enforcement over real HTTP. The unit tests pin the policy; these pin that the policy
* is actually reached from a request — a rule enforced nowhere is not a control.
*/
class BridgedAppAuthTest {
private final HttpClient http = HttpClient.newHttpClient();
private Javalin app;
private Metrics metrics;
@AfterEach
void stop() {
if (app != null) app.stop();
}
/**
* Start the app with the given identity/auth wiring.
*
* @param pid the PID every connection resolves to — {@link FakeHerdr#WORKER_PID} makes the
* caller worker {@code term_a}, anything else makes it a non-worker
*/
private int start(long pid, boolean tokenMode, String token) {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
k -> "BRIDGED_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, new Rendezvous());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
CallerResolver callers = tokenMode
? new CallerResolver(identity, true, token)
: new CallerResolver(identity);
metrics = BridgedMetrics.create(sessions, new dev.ltms.bridged.msg.InMemoryReplyInbox());
app = new BridgedApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
callers, metrics).build().start("127.0.0.1", 0);
return app.port();
}
private HttpResponse<String> send(int port, String method, String path, String body, String auth)
throws Exception {
HttpRequest.Builder b = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path))
.header("Content-Type", "application/json");
if (auth != null) {
b.header("Authorization", auth);
}
b = switch (method) {
case "POST" -> b.POST(body == null
? HttpRequest.BodyPublishers.noBody()
: HttpRequest.BodyPublishers.ofString(body));
case "DELETE" -> b.DELETE();
default -> b.GET();
};
return http.send(b.build(), HttpResponse.BodyHandlers.ofString());
}
// --- loopback-trust: the caller is the primary -------------------------------------------
@Test
void thePrimaryMayOrchestrateButMayNotForgeAWorkerReply() throws Exception {
int port = start(999_999, false, null); // no pane ⇒ primary
HttpResponse<String> read = send(port, "GET", "/profiles", null, null);
assertEquals(200, read.statusCode(), "the primary may observe");
HttpResponse<String> reply = send(port, "POST", "/sessions/term_a/reply",
"{\"content\":\"forged\"}", null);
assertEquals(403, reply.statusCode(),
"a forged reply would resolve the rendezvous the primary is itself waiting on");
assertTrue(reply.body().contains("forbidden"));
}
// --- loopback-trust: the caller is a worker ------------------------------------------------
@Test
void aWorkerMayReplyAsItselfButNotAsAnother() throws Exception {
int port = start(FakeHerdr.WORKER_PID, false, null); // resolves to term_a
HttpResponse<String> own = send(port, "POST", "/sessions/term_a/reply",
"{\"content\":\"done\"}", null);
assertEquals(200, own.statusCode(), "a worker replies on its own session");
HttpResponse<String> other = send(port, "POST", "/sessions/term_b/reply",
"{\"content\":\"not mine\"}", null);
assertEquals(403, other.statusCode(),
"REST trusted the path id before CB-505; this is the hole being closed");
}
@Test
void aWorkerMayNotOrchestrate() throws Exception {
int port = start(FakeHerdr.WORKER_PID, false, null);
assertEquals(403, send(port, "POST", "/workers", null, null).statusCode(),
"a worker spawning workers would be escalating into the orchestrator role");
assertEquals(403, send(port, "DELETE", "/workers/w2:p7", null, null).statusCode());
assertEquals(403, send(port, "POST", "/sessions/term_b/message",
"{\"content\":\"hi\"}", null).statusCode());
assertEquals(403, send(port, "GET", "/sessions/term_a/replies", null, null).statusCode(),
"draining an inbox is the primary's collection step");
}
// --- token mode ---------------------------------------------------------------------------
@Test
void tokenModeRejectsAnUncredentialedNonWorkerWith401() throws Exception {
int port = start(999_999, true, "s3cret");
HttpResponse<String> res = send(port, "GET", "/profiles", null, null);
assertEquals(401, res.statusCode(), "no credential ⇒ authenticated as nothing");
assertTrue(res.body().contains("unauthenticated"));
}
@Test
void tokenModeAcceptsAValidBearerToken() throws Exception {
int port = start(999_999, true, "s3cret");
assertEquals(200, send(port, "GET", "/profiles", null, "Bearer s3cret").statusCode());
}
@Test
void tokenModeStillHonoursConnectionDerivedWorkerIdentity() throws Exception {
// The fleet must keep working when auth is switched on: a worker presents no token, and
// must still be able to reply.
int port = start(FakeHerdr.WORKER_PID, true, "s3cret");
assertEquals(200, send(port, "POST", "/sessions/term_a/reply",
"{\"content\":\"done\"}", null).statusCode());
}
// --- health, metrics ----------------------------------------------------------------------
@Test
void healthzStaysOpenWithoutCredentials() throws Exception {
int port = start(999_999, true, "s3cret");
assertEquals(200, send(port, "GET", "/healthz", null, null).statusCode(),
"a supervisor must be able to probe liveness before any credential is configured");
}
@Test
void metricsRequireAuthenticationAndRenderPrometheusText() throws Exception {
int port = start(999_999, true, "s3cret");
assertEquals(401, send(port, "GET", "/metrics", null, null).statusCode());
HttpResponse<String> ok = send(port, "GET", "/metrics", null, "Bearer s3cret");
assertEquals(200, ok.statusCode());
assertTrue(ok.headers().firstValue("Content-Type").orElse("").startsWith("text/plain"));
assertTrue(ok.body().contains("bridged_sessions{state=\"ready\"}"),
"the session census gauge is exported even when empty");
}
@Test
void refusalsAreCounted() throws Exception {
int port = start(999_999, true, "s3cret");
send(port, "GET", "/profiles", null, null); // 401
send(port, "POST", "/sessions/term_a/reply", "{}", "Bearer s3cret"); // 403
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "unauthenticated"));
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "forbidden"));
}
// --- legacy constructor -------------------------------------------------------------------
@Test
void theLegacyConstructorLeavesAuthorizationOff() throws Exception {
// The 29 pre-existing acceptance tests rely on this: no auth fixture, no enforcement.
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
Map.of(wcfg.profile(), wcfg), wcfg.profile(), _ -> "tok");
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous());
app = new BridgedApp(herdr, workers, sessions, messages, sessions.asPresence(), null)
.build().start("127.0.0.1", 0);
assertEquals(200, send(app.port(), "POST", "/sessions/term_a/reply",
"{\"content\":\"x\"}", null).statusCode());
assertEquals(404, send(app.port(), "GET", "/metrics", null, null).statusCode(),
"no registry supplied ⇒ the endpoint is not mounted at all");
}
}
@@ -75,7 +75,7 @@ class BridgedAppTest {
poller.start();
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
app = new BridgedApp(herdr, workers, sessions, messages, rendezvous, this.presence, null)
app = new BridgedApp(herdr, workers, sessions, messages, this.presence, null)
.build().start("127.0.0.1", 0);
return app.port();
}
@@ -313,15 +313,10 @@ class BridgedAppTest {
catch (Exception e) { throw new RuntimeException(e); }
});
// The worker replies once a send is actually awaiting (retry past the startup race).
HttpResponse<String> reply;
long deadline = System.currentTimeMillis() + 3000;
do {
reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"LGTM ship it\"}");
if (reply.statusCode() != 409) break;
//noinspection BusyWait
Thread.sleep(10);
} while (System.currentTimeMillis() < deadline);
// Give the background send thread time to open its rendezvous waiter (CB-307: reply now
// queues in the inbox if no waiter is open, which would break the round-trip).
Thread.sleep(200);
HttpResponse<String> reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"LGTM ship it\"}");
assertEquals(200, reply.statusCode());
HttpResponse<String> res = send.get(6, java.util.concurrent.TimeUnit.SECONDS);
@@ -341,20 +336,14 @@ class BridgedAppTest {
String ticket = mapper.readTree(accepted.body()).get("ticket").asText();
assertFalse(ticket.isBlank(), "an async send must return a ticket");
// The worker replies once the async send is actually awaiting (retry past the startup race).
HttpResponse<String> reply;
long deadline = System.currentTimeMillis() + 3000;
do {
reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"async LGTM\"}");
if (reply.statusCode() != 409) break;
//noinspection BusyWait
Thread.sleep(10);
} while (System.currentTimeMillis() < deadline);
// Give the background async send thread time to open its rendezvous waiter.
Thread.sleep(200);
HttpResponse<String> reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"async LGTM\"}");
assertEquals(200, reply.statusCode());
// Polling the ticket now reports the finished delegation and its reply.
JsonNode task;
deadline = System.currentTimeMillis() + 3000;
long deadline = System.currentTimeMillis() + 3000;
do {
task = mapper.readTree(req(port, "GET", "/tasks/" + ticket).body());
if ("done".equals(task.path("phase").asText())) break;
@@ -375,11 +364,26 @@ class BridgedAppTest {
}
@Test
void replyWithNoPendingSendIsConflict() throws Exception {
void replyWithNoPendingSendQueuesInsteadOfConflict() throws Exception {
// CB-307: a reply with no open send now queues in the inbox, not a 409 conflict.
int port = startHealthy();
HttpResponse<String> res = postJson(port, "/sessions/term_a/reply", "{\"content\":\"orphan\"}");
assertEquals(409, res.statusCode());
assertEquals("no_pending_send", mapper.readTree(res.body()).get("error").asText());
assertEquals(200, res.statusCode());
// The queued reply is drainable.
HttpResponse<String> drain = req(port, "GET", "/sessions/term_a/replies");
assertEquals(200, drain.statusCode());
JsonNode body = mapper.readTree(drain.body());
assertEquals(1, body.get("replies").size());
assertEquals("orphan", body.get("replies").get(0).get("content").asText());
}
@Test
void drainRepliesReturnsEmptyForNoReplies() throws Exception {
int port = startHealthy();
HttpResponse<String> res = req(port, "GET", "/sessions/term_a/replies");
assertEquals(200, res.statusCode());
assertEquals(0, mapper.readTree(res.body()).get("replies").size());
}
@Test
@@ -6,6 +6,7 @@ import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import org.junit.jupiter.api.Test;
import java.util.List;
@@ -332,4 +333,75 @@ class SessionManagerTest {
.filter(c -> paneId.equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
}
// --- CB-306 spawn-readiness gate: no half-registered session on timeout ----------------
@Test
void acquireThrowsPeerUnreachableWhenGateTimesOutAndRegistersNoSession() {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // never becomes injectable
long[] clock = {0};
// Gate-enabled launcher (1 ms timeout + no-op sleeper that advances clock past deadline)
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
1, () -> clock[0], () -> clock[0] += 10);
SessionManager sessions = new SessionManager(workers, new GitWorktrees(), () -> 0L, 0);
assertThrows(PeerUnreachableException.class,
() -> sessions.acquire("ltms-local", null, "/caller", "term_primary"),
"acquire must throw PeerUnreachableException when spawn times out");
// No half-registered session — the error happened inside spawn, before
// SessionManager could put() anything into the registry.
assertTrue(sessions.roster().isEmpty(),
"no session is registered when spawn times out (roster empty)");
}
// --- CB-516: release must notify, so a blocked send can be failed --------------------------
@Test
void releaseNotifiesTheListenerWithTheReleasedTerminal() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller", null);
sessions.release(s.paneId());
assertEquals(java.util.List.of(s.terminalId()), released,
"every teardown path funnels through release, so one hook must see the terminal");
}
@Test
void releasingAnUnknownPaneNotifiesNobody() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
sessions.release("w9:p404"); // idempotent teardown of something already gone
assertTrue(released.isEmpty(), "no session removed ⇒ no send was waiting on it");
}
@Test
void aThrowingReleaseListenerDoesNotBlockTheTeardown() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
sessions.onRelease(_ -> {
throw new IllegalStateException("listener blew up");
});
WorkerSession s = sessions.acquire("ltms-local", null, "/caller", null);
assertDoesNotThrow(() -> sessions.release(s.paneId()),
"a listener failure must never prevent the teardown it is reacting to");
assertTrue(sessions.get(s.paneId()).isEmpty(), "and the session is still deregistered");
}
}
@@ -0,0 +1,131 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Wrapper-behaviour tests for {@link SessionReaper} (the thread lifecycle). The TTL policy itself
* (SessionManager.reapIdle) is covered by SessionManagerTest and is deliberately not retested here.
* A real SessionManager is used, built the same way the rest of this package's tests do.
*/
class SessionReaperTest {
private static final long IDLE_TTL_SECONDS = 60;
private static final long SHORT_INTERVAL_MILLIS = 20;
private static ClaudeCodeLauncher launcher() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
}
/** A manager on the fake worktree seam — these tests never touch a real git checkout. */
private static SessionManager sessionManager() {
return new SessionManager(launcher(), new FakeWorktrees());
}
private static SessionReaper reaper() {
return new SessionReaper(sessionManager(), IDLE_TTL_SECONDS, SHORT_INTERVAL_MILLIS);
}
/**
* A double {@code start()} must leave exactly one live loop, so a single {@code stop()} still
* silences it. Asserting only "no throw" would pass against a reaper that never started at
* all — and against one that started twice — which is the entire point of the guard.
*/
@Test
void startIsIdempotent() throws InterruptedException {
AtomicLong ticks = new AtomicLong();
SessionReaper reaper = new SessionReaper(countingManager(ticks),
IDLE_TTL_SECONDS, SHORT_INTERVAL_MILLIS);
assertDoesNotThrow(() -> {
reaper.start();
reaper.start();
}, "a second start() must not throw");
assertTrue(awaitTicks(ticks, 2), "the loop is running after a double start()");
// One stop() for two start() calls: if the second start had spawned its own loop, a
// surviving thread would keep the counter climbing past this point.
reaper.stop();
Thread.sleep(SHORT_INTERVAL_MILLIS * 4);
long settled = ticks.get();
Thread.sleep(SHORT_INTERVAL_MILLIS * 4);
assertEquals(settled, ticks.get(),
"a single stop() must silence the reaper even after two start() calls");
}
/** A manager whose clock counts reads — every {@code reapIdle} reads it exactly once. */
private static SessionManager countingManager(AtomicLong ticks) {
return new SessionManager(launcher(), new FakeWorktrees(), () -> {
ticks.incrementAndGet();
return System.nanoTime();
});
}
/** Bounded wait for the loop to tick at least {@code n} times; avoids fixed-sleep flakiness. */
private static boolean awaitTicks(AtomicLong ticks, long n) throws InterruptedException {
long deadline = System.currentTimeMillis() + 2000;
while (ticks.get() < n && System.currentTimeMillis() < deadline) {
Thread.sleep(10);
}
return ticks.get() >= n;
}
@Test
void stopIsIdempotentAndSafeBeforeStart() {
SessionReaper reaper = reaper();
assertDoesNotThrow(reaper::stop, "stop() before start() must not throw");
assertDoesNotThrow(reaper::stop, "a second stop() must not throw");
}
/**
* The loop must actually iterate, and {@code stop()} must actually end it.
*
* <p>Observed through an injected clock rather than by sleeping and hoping: every
* {@code reapIdle} call reads {@code nowNanos} exactly once, so the tick count <em>is</em> the
* iteration count. Asserting merely "nothing threw" would pass even if {@code start()} were a
* no-op, which is the whole behaviour under test.
*/
@Test
void theLoopRunsRepeatedlyAndStopEndsIt() throws InterruptedException {
AtomicLong ticks = new AtomicLong();
SessionReaper reaper = new SessionReaper(countingManager(ticks),
IDLE_TTL_SECONDS, SHORT_INTERVAL_MILLIS);
reaper.start();
// Bounded wait rather than a fixed sleep + exact count: proves repetition without pinning
// a timing-derived number that would flake on a loaded machine.
boolean iterated = awaitTicks(ticks, 2);
long whileRunning = ticks.get();
reaper.stop();
assertTrue(iterated,
"the reaper loop must iterate repeatedly; observed " + whileRunning + " tick(s)");
// After stop() the loop must go quiet. Allow one in-flight iteration to finish, then
// confirm the count has stopped advancing.
Thread.sleep(SHORT_INTERVAL_MILLIS * 4);
long settled = ticks.get();
Thread.sleep(SHORT_INTERVAL_MILLIS * 4);
assertEquals(settled, ticks.get(), "stop() must end the loop, not just flag it");
}
}
@@ -166,4 +166,58 @@ class WorktreeSessionManagerTest {
assertEquals(2, sessions.roster().size());
}
/**
* CB-507 regression. A plain REST spawn supplies neither a requested nor a caller cwd
* ({@code BridgedApp} hardcodes {@code callerCwd = null}), and the worktree branch used to
* resolve the repo root from just those two — yielding {@code null}, which the real
* {@code GitWorktrees} turns into {@code git -C null} and an NPE out of {@code ProcessBuilder}
* (HTTP 500).
*
* <p>Note this asserts on the <em>recorded</em> cwd rather than expecting a throw:
* {@link FakeWorktrees#repoRoot} only records its argument and returns a canned root, so a
* null flows through the fake harmlessly. That permissiveness is precisely why the whole
* suite stayed green while the feature was broken in production — so the assertion has to be
* "a usable cwd was passed down", not "an exception was raised".
*/
@Test
void worktreeAcquireWithNoRequestedOrCallerCwdStillResolvesANonNullRepoRoot() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
sessions.acquire("ltms-local", null, null, null, new WorktreeRequest("cb-507", null));
assertFalse(worktrees.repoRootCalls().isEmpty(),
"repoRoot should have been called to resolve the repo root");
String cwd = worktrees.repoRootCalls().getFirst().cwd();
assertNotNull(cwd, "a null cwd here becomes `git -C null` and NPEs in the real GitWorktrees");
assertFalse(cwd.isBlank(), "a blank cwd is as unusable as a null one");
}
/**
* The same line carried a second, quieter bug: it never consulted the profile's configured
* {@code cwd:}, so a worktree spawn silently ignored a pinned per-profile working directory.
* Routing through {@code effectiveCwd} honours it.
*/
@Test
void worktreeAcquireHonoursTheProfileConfiguredCwd() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
// Argument order matters: configDir is the 4th parameter, cwd the 11th (after mcpUrl).
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, "/pinned/dir", null);
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
SessionManager sessions = new SessionManager(launcher, worktrees);
sessions.acquire("ltms-local", null, null, null, new WorktreeRequest("cb-507b", null));
assertEquals(1, worktrees.repoRootCalls().size());
assertEquals("/pinned/dir", worktrees.repoRootCalls().getFirst().cwd(),
"the profile's configured cwd must reach repoRoot, not be ignored");
}
}
@@ -7,6 +7,7 @@ import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
@@ -318,4 +319,154 @@ class ClaudeCodeLauncherTest {
assertTrue(herdr.called("pane.close"), "stop via handle.id() must close the pane");
}
// --- CB-306 spawn-readiness gate -----------------------------------------------------------
private static Map<String, BridgedConfig.Worker> workerConfigMap(String profile, String mcpUrl) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
profile, "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", profile), "tab", "bridged-workers",
"worker: {profile} #{n}", mcpUrl, null, null);
return Map.of(cfg.profile(), cfg);
}
@Test
void spawnWaitsUntilInjectableThenReturnsHandle() {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // first status call sees UNKNOWN
long[] clock = {0};
boolean[] firstSleep = {true};
// The sleeper: advance the fake clock, and on the first call flip the
// agent status to IDLE so the next poll succeeds.
Runnable sleeper = () -> {
clock[0] += 300;
if (firstSleep[0]) {
herdr.agentStatus("idle");
firstSleep[0] = false;
}
};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
5000, () -> clock[0], sleeper);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn returns a handle when worker becomes injectable");
assertEquals("w9:pW_1", handle.id(), "handle id matches the started pane");
assertEquals(0, paneCloseCount(herdr, "w9:pW_1"),
"no pane.close when worker becomes injectable before timeout");
}
@Test
void spawnThrowsPeerUnreachableWhenNeverInjectableAndReapsPane() {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
1000, () -> clock[0], () -> clock[0] += 50);
PeerUnreachableException ex = assertThrows(
PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(ex.getMessage().contains("w9:pW_1"),
"exception message references the paneId: " + ex.getMessage());
assertTrue(ex.getMessage().contains("1000"),
"exception message references the timeout: " + ex.getMessage());
assertTrue(clock[0] >= 1000, "fake clock advanced past the timeout: " + clock[0]);
assertEquals(1, paneCloseCount(herdr, "w9:pW_1"),
"pane was closed on timeout (no orphan left behind)");
}
@Test
void spawnReturnsImmediatelyWhenGateIsDisabled() {
FakeHerdr herdr = new FakeHerdr();
// The default 6-arg constructor has spawnReadyTimeoutMs=0 (gate disabled).
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn returns a handle when the gate is disabled");
assertFalse(herdr.called("agent.get"),
"agent.get is never called when the gate is disabled (no polling)");
}
@Test
void spawnGateRespectsZeroTimeoutEvenWithFullConstructor() {
FakeHerdr herdr = new FakeHerdr();
long[] clock = {0};
// Explicit zero timeout with the full testability constructor — should
// skip polling entirely, just like the legacy default path.
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
0, () -> clock[0], () -> clock[0] += 1);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn still succeeds with zero timeout");
assertEquals(0, paneCloseCount(herdr, handle.id()),
"no orphan pane close from the gate path");
}
// --- CB-511: worker environment seeding -----------------------------------------------------
@Test
void workerInheritsTheDaemonPath() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
k -> "PATH".equals(k) ? "/opt/tools/bin:/usr/bin" : null).spawn();
assertEquals("/opt/tools/bin:/usr/bin", startEnv(herdr).get("PATH"),
"a worker with no PATH cannot run the build it is asked to run");
}
@Test
void profileEnvIsInjectedIntoTheWorker() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("JAVA_HOME", "/opt/jdk", "PATH", "/profile/bin"));
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
k -> "PATH".equals(k) ? "/daemon/bin" : null).spawn();
Map<String, String> env = startEnv(herdr);
assertEquals("/opt/jdk", env.get("JAVA_HOME"), "profile env: is passed through");
assertEquals("/profile/bin", env.get("PATH"), "an explicit profile PATH overrides the daemon's");
}
/**
* The security-relevant ordering. {@code SubscriptionGuard} is checked against the profile's
* {@code baseUrl} only, so if a profile's {@code env:} could overwrite ANTHROPIC_BASE_URL a
* worker could be pointed at an unguarded host while the guard passed on a benign one.
*/
@Test
void profileEnvCannotOverrideGuardCheckedAnthropicVars() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("ANTHROPIC_BASE_URL", "http://evil.example.com"));
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null).spawn();
assertEquals("http://gx00.gw:8000", startEnv(herdr).get("ANTHROPIC_BASE_URL"),
"the guard-checked baseUrl must win over any env: entry, or the boundary is bypassable");
}
}
@@ -0,0 +1,158 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
/**
* The composite router: profile → owning adapter for spawn/cwd/parity, pane id → owner for stop,
* and fleet-wide union/dedup for list/reap/caps/profiles. Exercised through two real adapters —
* claude-code + opencode — over one FakeHerdr, so each call is observed reaching the right adapter
* (the started herdr agent name carries that adapter's {@code claude-}/{@code opencode-} prefix).
*/
class CompositePeerLauncherTest {
private ClaudeCodeLauncher claudeAdapter(FakeHerdr herdr) {
// 12-arg back-compat Worker ctor → kind defaults to claude-code.
BridgedConfig.Worker claude = new BridgedConfig.Worker("claude", "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}",
null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("claude", claude), "claude", _ -> null);
}
private OpenCodeLauncher opencodeAdapter(FakeHerdr herdr) {
BridgedConfig.Worker gemini = new BridgedConfig.Worker("gemini", null, "google/gemini-2.5-pro",
null, "BRIDGED_WORKER_TOKEN", List.of("opencode"), "tab", "bridged-workers", "w #{n}",
null, null, null, "GITEA_ACCESS_TOKEN", null, BridgedConfig.Worker.KIND_OPENCODE);
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("gemini", gemini), "gemini", _ -> "tok");
}
private CompositePeerLauncher composite(FakeHerdr herdr) {
return new CompositePeerLauncher(
List.of(claudeAdapter(herdr), opencodeAdapter(herdr)), "claude");
}
@SuppressWarnings("unchecked")
private static String startedName(FakeHerdr herdr) {
return (String) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("name");
}
@Test
void spawnRoutesEachProfileToItsOwningAdapter() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
composite.spawn(new SpawnRequest("gemini", null, null));
assertTrue(startedName(herdr).startsWith("opencode-"),
"the gemini profile is spawned by the opencode adapter: " + startedName(herdr));
composite.spawn(new SpawnRequest("claude", null, null));
assertTrue(startedName(herdr).startsWith("claude-"),
"the claude profile is spawned by the claude-code adapter: " + startedName(herdr));
}
@Test
void nullProfileResolvesTheDefaultAndRoutesToItsOwner() {
FakeHerdr herdr = new FakeHerdr();
composite(herdr).spawn(new SpawnRequest(null, null, null));
assertTrue(startedName(herdr).startsWith("claude-"),
"a no-profile spawn resolves the default (claude) and routes to its adapter");
}
@Test
void unknownProfileIsRejected() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
assertThrows(IllegalArgumentException.class,
() -> composite.spawn(new SpawnRequest("nope", null, null)),
"a profile no adapter declares is an error");
}
@Test
void profilesAndDefaultAreExposedAcrossAdapters() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
assertEquals(Set.of("claude", "gemini"), composite.profiles(),
"profiles are the union of every adapter's profiles");
assertEquals("claude", composite.defaultProfile());
}
@Test
void capabilitiesAreTheUnionOfEveryAdapter() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher claude = claudeAdapter(herdr);
OpenCodeLauncher opencode = opencodeAdapter(herdr);
PeerLauncher composite = new CompositePeerLauncher(List.of(claude, opencode), "claude");
assertTrue(composite.capabilities().containsAll(claude.capabilities()),
"the fleet offers every claude-code capability");
assertTrue(composite.capabilities().containsAll(opencode.capabilities()),
"the fleet offers every opencode capability (incl. SELF_PR from its git-token profile)");
}
@Test
void listIsDeduplicatedByPaneIdAcrossAdaptersSharingHerdr() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
// Both adapters wrap the same herdr, so each list() returns the same global agent set;
// the composite must return each pane once, not once per adapter.
assertEquals(1, composite.list().size(),
"the single herdr-tracked pane appears once, not duplicated per adapter");
}
@Test
void reapSumsAcrossAdaptersAndEachAdapterReapsOnlyItsOwnPrefix() {
// One foreign opencode orphan + one foreign claude orphan, from a prior daemon (different nonce).
FakeHerdr herdr = new FakeHerdr()
.withAgent("opencode-gemini-ffffff-1", "term_o", "wQ:pO", "wQ:tO")
.withAgent("claude-claude-eeeeee-1", "term_c", "wQ:pC", "wQ:tC");
PeerLauncher composite = composite(herdr);
assertEquals(2, composite.reapOrphanWorkers(),
"both orphans are reaped — one by each adapter, summed by the composite");
}
@Test
void stopTearsDownAPaneSpawnedThroughTheComposite() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
PeerHandle handle = composite.spawn(new SpawnRequest("gemini", null, null));
composite.stop(handle.id());
assertTrue(herdr.calls.stream()
.anyMatch(c -> c.method().equals("pane.close")
&& handle.id().equals(((Map<?, ?>) c.params()).get("pane_id"))),
"stop routes to the spawning adapter and closes that worker's pane");
}
@Test
void constructorRejectsAProfileClaimedByTwoAdapters() {
FakeHerdr herdr = new FakeHerdr();
// Two opencode adapters both declaring "gemini" — a profile-name collision.
OpenCodeLauncher a = opencodeAdapter(herdr);
OpenCodeLauncher b = opencodeAdapter(herdr);
assertThrows(IllegalArgumentException.class,
() -> new CompositePeerLauncher(List.of(a, b), "gemini"),
"a profile two adapters both claim is a configuration error");
}
@Test
void constructorRejectsAnEmptyAdapterList() {
assertThrows(IllegalArgumentException.class,
() -> new CompositePeerLauncher(List.of(), "claude"),
"at least one adapter must be configured");
}
}
@@ -0,0 +1,269 @@
package dev.ltms.bridged.worker;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/**
* The opencode adapter's launch build: a file-based MCP mount + reply-charter instructions (no
* inline flags, no {@code ANTHROPIC_*}, no guard), the {@code -m} model flag, and the shared base
* transport (naming, reap, readiness gate) proving the {@link HerdrPeerLauncher} SPI is neutral.
*/
class OpenCodeLauncherTest {
private static BridgedConfig.Worker opencodeCfg(String model, String mcpUrl, String gitTokenEnv) {
return new BridgedConfig.Worker("gemini", null, model, null, "BRIDGED_WORKER_TOKEN",
List.of("opencode"), "tab", "bridged-workers", "opencode: {model} #{n}", mcpUrl,
null, null, gitTokenEnv, null, BridgedConfig.Worker.KIND_OPENCODE);
}
/** Gate-disabled launcher whose per-spawn config dirs land under an inspectable temp root. */
private OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, BridgedConfig.Worker cfg) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), k -> "GITEA_ACCESS_TOKEN".equals(k) ? "tok" : null,
0, System::currentTimeMillis, () -> { }, configRoot);
}
@SuppressWarnings("unchecked")
private static Map<String, Object> lastStart(FakeHerdr herdr) {
return (Map<String, Object>) herdr.lastCall("agent.start").params();
}
@SuppressWarnings("unchecked")
private static Map<String, String> startEnv(FakeHerdr herdr) {
return (Map<String, String>) lastStart(herdr).get("env");
}
@SuppressWarnings("unchecked")
private static List<String> startArgv(FakeHerdr herdr) {
return (List<String>) lastStart(herdr).get("argv");
}
@Test
void writesRemoteMcpConfigAndCharterInstructionsWhenMcpUrlSet(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null))
.spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"), "opencode carries no ANTHROPIC_* / subscription boundary");
String cfgPath = env.get("OPENCODE_CONFIG");
assertNotNull(cfgPath, "OPENCODE_CONFIG points the worker at the generated config file");
assertTrue(Path.of(cfgPath).startsWith(root), "config file is generated under the injected root");
// Assert on parsed structure, not substrings: the generated config is real JSON and its
// whitespace is the formatter's business, not the contract's.
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
JsonNode bridge = json.path("mcp").path("bridge");
assertEquals("remote", bridge.path("type").asText(), "bridge is mounted as a remote MCP server");
assertEquals("http://127.0.0.1:8765/mcp", bridge.path("url").asText(),
"the profile's bridge MCP url is present");
assertTrue(bridge.path("enabled").asBoolean(), "the bridge server is enabled");
assertTrue(json.path("instructions").isArray() && !json.path("instructions").isEmpty(),
"the reply charter is mounted via instructions");
// The instructions entry is a real file path holding the reply charter.
Path charter = Path.of(cfgPath).resolveSibling("reply-charter.md");
assertTrue(Files.exists(charter), "the charter file the config references was written");
assertTrue(Files.readString(charter).contains("bridge_reply"),
"the charter instructs the worker to answer via bridge_reply");
}
@Test
void noConfigFileWhenMcpUrlAbsent(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null)).spawn();
assertNull(startEnv(herdr).get("OPENCODE_CONFIG"),
"no bridge MCP url → no config file and no OPENCODE_CONFIG");
}
@Test
void passesTheModelAsDashMFlag(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null)).spawn();
List<String> argv = startArgv(herdr);
assertEquals("opencode", argv.getFirst(), "base opencode command preserved first");
int m = argv.indexOf("-m");
assertTrue(m >= 0, "model is selected with -m");
assertEquals("google/gemini-2.5-pro", argv.get(m + 1), "the provider/model selector follows -m");
}
@Test
void noModelFlagWhenModelBlank(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg(null, null, null)).spawn();
assertEquals(List.of("opencode"), startArgv(herdr), "no model → argv is the bare opencode command");
}
@Test
void injectsForgeTokenWhenProfileGrantsIt(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg(null, null, "GITEA_ACCESS_TOKEN")).spawn();
assertEquals("tok", startEnv(herdr).get("GITEA_TOKEN"),
"a git-token profile gets the peer-neutral GITEA_TOKEN grant, same as Claude");
}
@Test
void capabilitiesDeclareOrphanReapAndMcpAskAndConditionalSelfPr(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
assertEquals(java.util.Set.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP),
service(herdr, root, opencodeCfg(null, null, null)).capabilities(),
"no git token → no SELF_PR");
assertTrue(service(herdr, root, opencodeCfg(null, null, "GITEA_ACCESS_TOKEN"))
.capabilities().contains(Capability.SELF_PR),
"a git-token profile adds SELF_PR");
}
@Test
void foreignWorkerMatchesOpencodePrefixButNotClaude() {
String nonce = "abc123";
assertTrue(OpenCodeLauncher.isForeignWorker("opencode-gemini-def456-1", nonce),
"an opencode pane from another process is foreign");
assertFalse(OpenCodeLauncher.isForeignWorker("opencode-gemini-" + nonce + "-1", nonce),
"our own opencode pane (same nonce) is not foreign");
assertFalse(OpenCodeLauncher.isForeignWorker("claude-ltms-local-def456-1", nonce),
"a claude pane is never reaped by the opencode adapter");
}
@Test
void productionConstructorsWireThroughToTheBase() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = opencodeCfg(null, null, null);
// 5-arg (gate disabled) and 7-arg (gate enabled) production constructors both expose the profile.
OpenCodeLauncher disabled = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
OpenCodeLauncher gated = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null, 5000, 100);
assertEquals(java.util.Set.of("gemini"), disabled.profiles());
assertEquals("gemini", gated.defaultProfile());
}
@Test
void spawnGateThrowsPeerUnreachableWhenNeverInjectable(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // never injectable
long[] clock = {0};
OpenCodeLauncher svc = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("gemini", opencodeCfg(null, null, null)), "gemini", _ -> null,
1000, () -> clock[0], () -> clock[0] += 50, root);
PeerUnreachableException ex = assertThrows(PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(clock[0] >= 1000, "the fake clock advanced past the timeout: " + clock[0]);
long closes = herdr.calls.stream()
.filter(c -> c.method().equals("pane.close"))
.filter(c -> "w9:pW_1".equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
assertEquals(1, closes, "the worker pane was reaped on timeout (no orphan)");
assertNotNull(ex.getMessage());
}
@Test
void spawnReturnsHandleWhenGateDisabled(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
PeerHandle handle = service(herdr, root, opencodeCfg(null, null, null))
.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn returns a handle when the gate is disabled");
assertFalse(herdr.called("agent.get"), "no polling when the gate is disabled");
}
// --- CB-508: pinned OpenAI-compatible endpoint (e.g. a local vLLM) ---------------------------
/** A profile with a baseUrl but no model provider prefix cannot be resolved — fail loudly. */
private static BridgedConfig.Worker pinnedCfg(String model, String baseUrl, String mcpUrl) {
return new BridgedConfig.Worker("local", baseUrl, model, null, "BRIDGED_WORKER_TOKEN",
List.of("opencode"), "tab", "bridged-workers", "opencode: {model} #{n}", mcpUrl,
null, null, null, null, BridgedConfig.Worker.KIND_OPENCODE);
}
@Test
void baseUrlDeclaresACustomOpenAiCompatibleProvider(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, pinnedCfg("local-vllm/deepseek-v4-flash", "http://127.0.0.1:8000", null))
.spawn();
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(cfgPath, "a pinned endpoint needs a config file even with no bridge MCP url");
JsonNode provider = new ObjectMapper().readTree(Path.of(cfgPath).toFile())
.path("provider").path("local-vllm");
assertFalse(provider.isMissingNode(), "the provider id comes from the model selector");
assertEquals("@ai-sdk/openai-compatible", provider.path("npm").asText());
assertEquals("http://127.0.0.1:8000/v1", provider.path("options").path("baseURL").asText(),
"a bare host:port gets /v1 appended — that is where these servers mount the API");
assertFalse(provider.path("options").path("apiKey").asText().isBlank(),
"the AI SDK requires a non-empty key even when the server ignores it");
assertFalse(provider.path("models").path("deepseek-v4-flash").isMissingNode(),
"the model half of the selector is declared under the provider");
}
@Test
void aBaseUrlThatAlreadyCarriesAPathIsUsedVerbatim(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, pinnedCfg("local-vllm/m", "http://127.0.0.1:8000/openai/v1", null)).spawn();
JsonNode json = new ObjectMapper()
.readTree(Path.of(startEnv(herdr).get("OPENCODE_CONFIG")).toFile());
assertEquals("http://127.0.0.1:8000/openai/v1",
json.path("provider").path("local-vllm").path("options").path("baseURL").asText(),
"an endpoint mounted on a custom path must not have /v1 bolted on");
}
@Test
void aPinnedEndpointRejectsAModelWithNoProviderPrefix(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher =
service(herdr, root, pinnedCfg("deepseek-v4-flash", "http://127.0.0.1:8000", null));
// Silently falling back to the default gateway would point the worker at the wrong LLM
// while looking healthy — the one failure mode worth being loud about.
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, launcher::spawn);
assertTrue(e.getMessage().contains("<provider>/<model>"), "the error says how to fix it");
}
@Test
void aPinnedEndpointAndTheBridgeMcpCoexistInOneConfig(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, pinnedCfg("local-vllm/deepseek-v4-flash",
"http://127.0.0.1:8000", "http://127.0.0.1:8766/mcp")).spawn();
JsonNode json = new ObjectMapper()
.readTree(Path.of(startEnv(herdr).get("OPENCODE_CONFIG")).toFile());
assertEquals("remote", json.path("mcp").path("bridge").path("type").asText(),
"pinning an endpoint must not drop the bridge MCP mount");
assertFalse(json.path("provider").path("local-vllm").isMissingNode(),
"and the provider block is still declared alongside it");
assertTrue(json.path("instructions").isArray() && !json.path("instructions").isEmpty(),
"the reply charter survives too");
}
@Test
void noBaseUrlDeclaresNoProviderSoTheDefaultGatewayIsUsed(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("opencode/some-free-model", "http://127.0.0.1:8766/mcp", null))
.spawn();
JsonNode json = new ObjectMapper()
.readTree(Path.of(startEnv(herdr).get("OPENCODE_CONFIG")).toFile());
assertTrue(json.path("provider").isMissingNode(),
"without a baseUrl opencode resolves its own provider as before");
}
}
@@ -0,0 +1,37 @@
<configuration>
<!--
CB-506 — test-run logging. This file is NOT boilerplate; it exists to keep the test suite
out of the CB-505 audit trail.
main/resources/logback.xml routes the `audit` logger to a RollingFileAppender at
logs/audit.log. AuditLogTest and BridgedAppAuthTest exercise that same logger, so without
this file `mvn test` appends fabricated records — denied/forbidden SPAWN/STOP/SEND from
worker:term_a — to the production security log, byte-identical to real ones. An investigator
could not tell a test fixture from a genuine intrusion attempt. Logback prefers
logback-test.xml when it is on the test classpath, so this governs test runs only.
Two constraints if you edit this:
- NEVER add a FileAppender/RollingFileAppender here. That reintroduces the bug.
- Keep `audit` ENABLED (INFO, additivity=false). Setting it to OFF would silently break
AuditLogTest, which attaches its own ListAppender and asserts on emitted records.
-->
<appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
<encoder>
<pattern>%d{HH:mm:ss.SSS} [%thread] %-5level %logger{36} - %msg%n</pattern>
</encoder>
</appender>
<logger name="audit" level="INFO" additivity="false">
<appender-ref ref="STDOUT"/>
</logger>
<logger name="dev.ltms.bridged" level="WARN"/>
<logger name="org.eclipse.jetty" level="WARN"/>
<root level="INFO">
<appender-ref ref="STDOUT"/>
</root>
</configuration>
+61
View File
@@ -0,0 +1,61 @@
# CB-504 — systemd unit for bridged (Linux).
#
# The macOS launchd agent (deploy/dev.ltms.bridged.plist) is the supervision target for the
# current single-host deployment. This unit exists for the per-host gateways CB-308 introduces,
# which will run on Linux.
#
# Install (user service — bridged drives the user's herdr, not a system daemon):
# mkdir -p ~/.config/systemd/user
# cp deploy/bridged.service ~/.config/systemd/user/
# # edit ExecStart / WorkingDirectory / Environment below, then:
# systemctl --user daemon-reload
# systemctl --user enable --now bridged
# journalctl --user -u bridged -f
[Unit]
Description=bridged — claude-bridge message server
Documentation=https://git.ltms.dev/lms/claude-bridge/wiki
# Ordering only: herdr is a user process and its socket may appear after us. This is advisory —
# bridged retries the herdr socket rather than exiting, which is what actually makes a late
# socket survivable. Do NOT add Requires=: a herdr restart must not take bridged down with it.
After=herdr.service
Wants=herdr.service
[Service]
Type=simple
WorkingDirectory=%h/src/claude-bridge/bridged
ExecStart=/usr/lib/jvm/temurin-25-jdk/bin/java -jar target/bridged.jar bridged.yaml
Environment=HERDR_SOCKET_PATH=%h/.config/herdr/herdr.sock
# PATH matters more than it looks (CB-511): bridged propagates its own PATH to every worker it
# spawns, so this line decides whether the fleet can run a build at all. systemd does not source a
# login shell, so without it the daemon — and every worker — gets a bare default with no JDK/Maven.
Environment=PATH=/usr/lib/jvm/temurin-25-jdk/bin:/usr/share/maven/bin:/usr/local/bin:/usr/bin:/bin
# Secrets are NOT set here — this file is committed. Put the API/worker tokens in a private
# drop-in that systemd reads with restrictive permissions:
# systemctl --user edit bridged → [Service] / Environment=BRIDGED_API_TOKEN=...
# or point EnvironmentFile at a 0600 file:
# EnvironmentFile=%h/.config/bridged/env
Restart=on-failure
RestartSec=10s
# A bad config (e.g. a non-loopback bind without token auth) makes bridged fail fast by design.
# Give up rather than restart-loop on a permanent error.
StartLimitBurst=5
StartLimitIntervalSec=120
# The daemon reads the repo, writes worktrees, and talks to a Unix socket — it needs no more.
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=read-write
ProtectKernelTunables=true
ProtectControlGroups=true
RestrictSUIDSGID=true
StandardOutput=journal
StandardError=journal
SyslogIdentifier=bridged
[Install]
WantedBy=default.target
+80
View File
@@ -0,0 +1,80 @@
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<!--
CB-504 — launchd agent for bridged (macOS).
This is the real supervision target today: the dogfooded daemon runs on macOS, where there is
no systemd. A systemd unit ships alongside (deploy/bridged.service) for the Linux gateways
CB-308 introduces.
Install:
cp deploy/dev.ltms.bridged.plist ~/Library/LaunchAgents/
# edit the paths + JAVA_HOME below to match this host, then:
launchctl load -w ~/Library/LaunchAgents/dev.ltms.bridged.plist
launchctl list | grep bridged
Note on ordering: launchd has no "start after herdr" primitive for user agents, and neither
does systemd in a way that survives a socket appearing late. bridged retries the herdr socket
on startup instead, so an agent that comes up before herdr converges rather than dying — that
retry is the actual fix; KeepAlive below is the backstop.
-->
<plist version="1.0">
<dict>
<key>Label</key>
<string>dev.ltms.bridged</string>
<key>ProgramArguments</key>
<array>
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home/bin/java</string>
<string>-jar</string>
<string>/Users/CHANGEME/src/claude-bridge/bridged/target/bridged.jar</string>
<string>bridged.yaml</string>
</array>
<!-- Config path in ProgramArguments is relative, so the working directory must be the module. -->
<key>WorkingDirectory</key>
<string>/Users/CHANGEME/src/claude-bridge/bridged</string>
<key>EnvironmentVariables</key>
<dict>
<key>JAVA_HOME</key>
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home</string>
<key>HERDR_SOCKET_PATH</key>
<string>/Users/CHANGEME/.config/herdr/herdr.sock</string>
<!--
PATH matters more than it looks (CB-511): bridged propagates its own PATH to every worker
it spawns, so this line decides whether the fleet can run a build at all. launchd does NOT
source .zprofile/.zshrc, so without this the daemon (and therefore every worker) gets a
bare /usr/bin:/bin and no JDK or Maven. Keep the toolchain entries first.
-->
<key>PATH</key>
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home/bin:/Users/CHANGEME/Tool/apache-maven-3.9.16/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin</string>
<!--
Worker/API tokens are NOT set here: this file is committed. Export them from a private
launchd override or a wrapper script. bridged reads the API token from the env var named
by auth.tokenEnv (default BRIDGED_API_TOKEN) and only in auth.mode: token.
-->
</dict>
<key>RunAtLoad</key>
<true/>
<!-- Restart on crash, but not in a tight loop if the config is bad (bridged fails fast on a
non-loopback bind without token auth — that is a config error, not a transient one). -->
<key>KeepAlive</key>
<dict>
<key>SuccessfulExit</key>
<false/>
</dict>
<key>ThrottleInterval</key>
<integer>10</integer>
<key>StandardOutPath</key>
<string>/Users/CHANGEME/src/claude-bridge/bridged/logs/bridged.out.log</string>
<key>StandardErrorPath</key>
<string>/Users/CHANGEME/src/claude-bridge/bridged/logs/bridged.err.log</string>
<key>ProcessType</key>
<string>Background</string>
</dict>
</plist>
+60
View File
@@ -0,0 +1,60 @@
# LavinMQ — the AMQP broker behind bridged's durable ReplyInbox (CB-307 Stage 2).
#
# Why this file exists: the broker was previously run ad hoc and simply vanished from the host,
# which takes bridged down with it — AmqpReplyInbox.open throws on an unreachable broker and
# Bridged.java:187 does not guard it, so a missing broker is a hard startup failure, not a
# degraded mode. This pins the version, keeps the data, and brings itself back after a reboot.
#
# Usage:
# docker compose -f deploy/lavinmq/compose.yaml up -d
# docker compose -f deploy/lavinmq/compose.yaml ps
# docker compose -f deploy/lavinmq/compose.yaml logs -f
# docker compose -f deploy/lavinmq/compose.yaml down # keeps the volume
# docker compose -f deploy/lavinmq/compose.yaml down -v # DESTROYS held replies
#
# Management UI: http://127.0.0.1:15672 (guest / guest)
#
# This is bridged's OWN broker. Do not point bridged at any other AMQP server on this host —
# notably not the `local-rabbitmq` container, which belongs to a different project and would end
# up carrying this project's queues.
name: bridged-broker
services:
lavinmq:
# Pinned deliberately: :latest silently moves the broker under a running daemon.
image: cloudamqp/lavinmq:2.9.1
container_name: bridged-lavinmq
# The failure this deployment exists to prevent — survive reboots and Docker restarts, but
# stay down if it was stopped on purpose.
restart: unless-stopped
# Loopback-bound on purpose. LavinMQ ships a default guest/guest account, which is only
# acceptable because nothing off-host can reach it. bridged connects over 127.0.0.1, and
# binding 0.0.0.0 here would expose a broker with default credentials to the network.
ports:
- "127.0.0.1:5672:5672" # AMQP — bridged.yaml broker.uri points here
- "127.0.0.1:15672:15672" # HTTP management API + UI
# The whole point of Stage 2. Held-but-unacked replies live here; without a named volume a
# `docker compose down` would discard exactly what the durable inbox exists to protect.
volumes:
- lavinmq-data:/var/lib/lavinmq
healthcheck:
test: ["CMD", "lavinmqctl", "status"]
interval: 30s
timeout: 5s
retries: 3
start_period: 10s
logging:
driver: json-file
options:
max-size: "10m"
max-file: "3"
volumes:
lavinmq-data:
name: bridged-lavinmq-data
+5 -1
View File
@@ -1,6 +1,10 @@
# CB-301-ext — Worktree provisioning + config-parity overlay
**Status:** design spec for review → delegate implementation.
**Status:** ✅ shipped — implemented at commit `97ecc71` (per-worker git worktree + config-parity
overlay). As-built: `session/GitWorktrees.java` behind the `Worktrees` port, wired in
`Bridged.main` and configurable via `worktreeRoot` / per-profile `parityOverlay`
(see `bridged.example.yaml`). Branch/worktree surface in `bridge_list` landed with CB-304
(`9fe04bf`); the worker-opened-PR checkpoint landed as CB-302 (`64e70ef`).
**Extends:** [CB-301 Session Manager](CB-301-Session-Manager.md) (shipped, commit `54d907c`).
**Realizes:** the config-parity requirement in [Worker Git Workflow](Worker-Git-Workflow.md).
**Grounded in:** `SessionManager`, `WorkerService.spawn/effectiveCwd`, `BridgedConfig.Worker`,
+143
View File
@@ -0,0 +1,143 @@
# CB-306 — Spawn-Readiness Gate (launcher-owned terminal readiness)
**Status:** design note / delegation spec (branch `worker/cb-306-readiness`)
**Issue:** gitea `lms/claude-bridge` #4
**Owner of the behaviour:** `ClaudeCodeLauncher` (the `PeerLauncher` adapter) — NOT core.
## 1. Problem
`bridge_spawn` today returns a session the instant the herdr pane is started. The pane is not
yet a usable Claude REPL — it may still be sitting at the folder-trust prompt, or the CLI may
never come up at all. Nothing blocks or times out on that. Consequences:
- A `bridge_send` to a not-yet-ready worker surfaces as a **~60 s MCP-client timeout** (the send
blocks waiting for a turn that can't start) instead of a fast, explicit spawn failure.
- A worker stuck at the folder-trust prompt lingers in `SPAWNING` forever; nothing fails it.
We want **fail-fast spawn**: `spawn()` returns only once the peer is genuinely up and usable in
its terminal, or throws a clean error (and leaves no orphan pane) within a bounded timeout.
## 2. What already exists (do NOT rebuild)
`SessionManager` + `PresenceBridge.markPresent()` already flip a session `SPAWNING → READY` on
**any MCP contact from the worker** (`onReady(terminal)` → `transitionByTerminal(SPAWNING, READY)`,
SessionManager ~line 249, "worker became available on the bridge MCP"). That is the **delivery
lifecycle** and it stays exactly as-is. CB-306 does **not** touch it and does **not** replace it.
CB-306 adds a *complementary, launcher-side* gate: the launcher guarantees the **terminal** is a
live, interactive REPL before it hands a handle back. The two signals are layered:
| Signal | Owner | Means | CB-306 |
|---|---|---|---|
| terminal reaches interactive REPL (herdr `IDLE`) | `ClaudeCodeLauncher` (this ticket) | pane is past folder-trust, CLI is up | **NEW — the spawn gate** |
| first MCP contact → `SPAWNING→READY` | core (`SessionManager`/`PresenceBridge`) | worker spoke to the bridge | unchanged |
## 3. The readiness predicate (herdr status)
`AgentStatus.fromWire` maps herdr's wire strings to `IDLE | WORKING | BLOCKED | DONE | UNKNOWN`.
A freshly started pane that has **not** reached an interactive Claude — including one stalled at
the folder-trust prompt — reports **`UNKNOWN`** (herdr has not detected a Claude REPL yet). Once
the CLI is up and settled at its prompt it reports **`IDLE`**.
**Predicate:** the pane is *ready* when `AgentControl.status(target)` first returns an
**injectable** state (`IDLE`, `BLOCKED`, or `DONE` — reuse `AgentStatus.injectable()`). `UNKNOWN`
= not ready. `WORKING` alone is ambiguous this early and should not by itself satisfy readiness;
wait for an injectable state. (We do not need to know *why* a pane isn't ready — a trust stall,
a crash, and a slow start all present as "never becomes injectable" and all correctly time out.)
## 4. Contract change on `spawn(SpawnRequest)`
`ClaudeCodeLauncher.spawn(SpawnRequest)` becomes **block-until-ready-or-throw**:
1. Start the pane exactly as today (`spawn(profile, cwd, callerCwd) → Agent`, build env + guard +
argv, `spawnInTab`/`spawnAsPane`, unique-named).
2. **Poll** `agentControl.status(paneId)` every `pollIntervalMs` (~300 ms) until it is `injectable()`
or `spawnReadyTimeoutMs` elapses.
3. **Ready** → return the `WorkerHandle(paneId, terminalId)` as today.
4. **Timeout** → the launcher **closes the pane it started** (and its tab, via the same path
`release`/`stop` uses) and throws **`PeerUnreachableException`** (new, in `dev.ltms.bridged.peer`).
No orphan pane is left behind — the launcher cleans up its own failed birth.
`spawnReadyTimeoutMs == 0` (or unset) **disables** the gate = legacy non-blocking behaviour, so the
change is opt-in per deployment and existing tests that don't configure it keep their old semantics.
### Testability seam (required)
Do **not** call `Thread.sleep` directly in the poll loop against a real clock — unit tests must not
real-sleep. Introduce a small injectable seam, mirroring the existing `StatusPoller` style:
- a `LongSupplier nowMillis` (monotonic clock) **and** a sleep/wait hook (e.g. a
`Sleeper`/`Waiter` functional interface, or reuse whatever `StatusPoller` already uses), both
defaulting to the real implementations in the production constructor and overridable in tests.
Unit tests (add to the existing `ClaudeCodeLauncher` test):
- fake `AgentControl` returns `UNKNOWN` a few times then `IDLE` → `spawn` returns the handle; assert
no `close` was called.
- fake `AgentControl` always `UNKNOWN` → `spawn` throws `PeerUnreachableException`; assert the pane
**was closed** (verify `close(paneId)` invoked) and the fake clock advanced past the timeout.
- `spawnReadyTimeoutMs == 0` → `spawn` returns immediately without polling (legacy path).
## 5. Config
Add to the launcher-level config (a bridged-level knob, not per-profile) in `bridged.yaml` +
`BridgedConfig`:
```yaml
spawn_ready_timeout_ms: 20000 # 0 disables the gate (legacy non-blocking spawn)
spawn_ready_poll_ms: 300
```
Jackson ignores unknown keys, so omitting them in existing YAML is safe; pick sane defaults in code
(`20000` / `300`). Keep the names consistent with existing config field style in `BridgedConfig`.
## 6. Core / MCP propagation
`SessionManager.acquire(...)` already calls `launcher.spawn(req)`. A thrown
`PeerUnreachableException` must propagate out as a **clean spawn failure**:
- The **worktree** acquire path already has a try/catch that cleans up a provisioned worktree when
`spawn` throws — verify the new exception flows through it (worktree removed, nothing registered).
- The **non-worktree** path registers the session only *after* `spawn` returns, so a throw means no
half-live `SPAWNING` session is ever registered — confirm this and add a test.
- `bridge_spawn` (MCP verb) must return an **error result** carrying the exception message, not a
success with a dead session. Trace `BridgeMcp`/`BridgedApp` spawn handlers and make sure the
exception becomes a clean tool error, not an uncaught 500 with a stack trace.
**Out of scope (do NOT do here):** gating `bridge_send` on session `READY` (existing status-gate +
this spawn gate already close the window), MCP-handshake-as-readiness signal, the CB-307 broker,
any config `kind:` discriminator, any second adapter.
## 7. Definition of done
- `ClaudeCodeLauncher.spawn` blocks until injectable or throws `PeerUnreachableException` +
self-reaps the pane; gate disabled when timeout is 0.
- New `PeerUnreachableException` in `dev.ltms.bridged.peer`.
- Config knobs wired (`spawn_ready_timeout_ms`, `spawn_ready_poll_ms`) with safe defaults.
- Existing `SPAWNING→READY` MCP-contact transition untouched.
- New unit tests (ready / timeout+reap / disabled) green; **all existing tests still pass unchanged**.
- Build clean via the worker's own `mvn` (primary re-runs the authoritative IDE + `mvn clean install`
gate — self-reports are not verified facts).
## 8. Sequence
```mermaid
sequenceDiagram
participant SM as SessionManager.acquire
participant L as ClaudeCodeLauncher.spawn
participant T as herdr (AgentControl)
SM->>L: spawn(SpawnRequest)
L->>T: start pane (env+guard+argv)
loop until injectable or timeout
L->>T: status(paneId)
T-->>L: UNKNOWN / IDLE
end
alt reached injectable
L-->>SM: PeerHandle(id, terminalId)
else timed out
L->>T: close(paneId) + tab
L-->>SM: throw PeerUnreachableException
SM-->>SM: no session registered / worktree cleaned
end
```
*Figure — the launcher blocks in `spawn` until the pane is a usable REPL, else self-reaps and throws.*
+164
View File
@@ -0,0 +1,164 @@
# CB-307 Stage 1 — Reply-Inbox Port + In-Memory Adapter (delegation spec)
**Ticket:** gitea #5 (CB-307). **Stage:** 1 of 2 (see the issue's "Implementation staging" comment).
**Scope of THIS delegation:** the `ReplyInbox` port + the in-memory (soft-state) adapter, wired at the
exact drop seam so a worker's terminal reply is **held instead of silently discarded** when no primary
send is open. **No broker, no new dependency, no infra** — fully unit-testable and primary-gate-verifiable.
Stage 2 (the AMQP/LavinMQ adapter behind the same port) is explicitly **out of scope** here.
> **Read this whole spec before starting.** The port and the publish seam are prescriptive
> (non-negotiable). Where a choice is genuinely open it says "DECISION" with the required default —
> follow the default and flag it in your completion message for the primary's review.
## 1. The bug this fixes (grounded in current code)
The reverse (worker→primary) path is `Rendezvous` — a `ConcurrentHashMap<session, CompletableFuture<Resolution>>`
of **live blocking waiters only**. No queue, no store. When a worker calls `bridge_reply` and **no send
is currently open** for that worker:
- `Rendezvous.resolve(session, content)` → `complete(...)` → `waiters.get(session) == null` →
returns `false` (`msg/Rendezvous.java:212-215`).
- The `content` string is **never retained** — it is dropped. The worker is told it failed:
`BridgeMcp.reply` returns `error("no send is awaiting a reply for this worker")` (`mcp/BridgeMcp.java:270-272`);
REST returns `409 no_pending_send` (`rest/BridgedApp.java:339-345`).
This is the observed "communication break": a worker that finishes just after its `bridge_send` timed
out (the ~60s sync window) replies into the void. There is **no message-id, dedup, or ack** anywhere in
the message path today.
## 2. What to build
### 2.1 The port — `dev.ltms.bridged.msg.ReplyInbox`
A thin interface owned by the `msg` layer. The in-memory adapter is Stage 1; the AMQP adapter (Stage 2)
implements the **same** interface, so keep it broker-agnostic.
```java
package dev.ltms.bridged.msg;
import java.util.List;
/**
* Holds terminal worker→primary replies that arrive with no live send to resolve, keyed by worker
* session (target), until the primary drains them. Soft-state in Stage 1 (in-memory, lost on restart);
* the Stage 2 AMQP adapter implements the same contract with cross-restart durability.
*/
public interface ReplyInbox {
/** A queued reply: an idempotency id, the worker session it came from, and the reply text. */
record InboxMessage(String msgId, String target, String content) {}
/**
* Queue {@code content} from worker {@code target} under {@code msgId}. Idempotent: publishing an
* already-present {@code msgId} for {@code target} is a no-op (dedup), so an at-least-once Stage-2
* redelivery cannot double-queue.
*/
void publish(String target, String msgId, String content);
/** Non-destructive snapshot of pending replies for {@code target} (FIFO), empty list if none. */
List<InboxMessage> peek(String target);
/** Remove the reply {@code msgId} for {@code target} once the primary has taken it. No-op if absent. */
void ack(String target, String msgId);
}
```
### 2.2 The default adapter — `InMemoryReplyInbox`
- Backed by a `ConcurrentHashMap<String, ...>` keyed by target session; per-target FIFO ordering.
- Dedup by `msgId` within a target (a `LinkedHashMap<msgId, InboxMessage>` per target, or a deque + a
seen-set — your call; preserve insertion order).
- `peek` returns an immutable copy; `ack` removes by `msgId`. Thread-safe (concurrent publish vs. drain).
- **This is soft-state, NOT persistence.** Lost on a `java -jar` bounce — that is correct and consistent
with "bridged stays soft-state." Do **not** add any file/DB backing.
### 2.3 Publish seam — route reply through the service layer
Keep `Rendezvous` a pure synchronization primitive (do **not** give it an inbox field). Instead centralize
in `MessageService`, which already owns the `Rendezvous` and will own the `ReplyInbox`:
- Add `MessageService.reply(String session, String content)`:
```java
/** Route a worker's explicit bridge_reply: resolve an open send, or queue it in the inbox if none. */
public boolean reply(String session, String content) {
if (rendezvous.resolve(session, content)) {
return true; // a live send took it — unchanged fast path
}
inbox.publish(session, UUID.randomUUID().toString(), content); // was a silent drop
return true; // held, not lost
}
```
- Repoint the two callers off the bare `rendezvous.resolve(...)` onto `messages.reply(...)`:
- `BridgeMcp.reply` (`mcp/BridgeMcp.java:262-273`) — on success return a normal ack; **remove** the
`error("no send is awaiting a reply…")` branch (that case is now a successful queue).
- `BridgedApp.replyMessage` (`rest/BridgedApp.java:330-346`) — return `200` (queued) instead of
`409 no_pending_send`.
**DO NOT touch the QUESTION path.** `bridge_ask` / `rendezvous.resolveQuestion` must keep today's
`NO_WAITER` behaviour — a mid-turn question is **interactive** (the worker blocks synchronously and cannot
consume a late answer), so it must **never** be queued. Only terminal `REPLY`s go to the inbox.
**DO NOT queue the completion/failure fallbacks** (`resolveCompletion` / `resolveFailure`,
`Rendezvous.java:196-210`). They target a *captured* waiter (CB-116); a `false` there means the turn was
already resolved or the scrape is a stale late duplicate — queueing it risks double-delivery. Leave them
exactly as they are. (Extending durability to completions is a deliberate Stage-2 consideration, not this.)
### 2.4 Drain seam — how the primary collects a stranded reply
The primary re-checks a worker it delegated to. Expose a drain keyed by **worker session (target)**:
- Add `MessageService.drainReplies(String target)`: `peek` the inbox, `ack` each returned `msgId`, hand
back the `List<InboxMessage>` (or just the contents). At-least-once: peek→deliver→ack (ack only after
the caller has them, so an in-flight failure re-surfaces them).
- **DECISION (required default): expose via the existing poll verb, keyed by target.** Extend `bridge_poll`
to accept an optional `target` (worker session) and, when present, return that worker's drained replies —
alongside a matching REST route `GET /sessions/{id}/replies`. Do **not** change `send`/`answer` semantics
(do not drain inside `send` — that conflates "deliver to worker" with "collect its mail"). Keep the
existing ticket-based `bridge_poll(ticket)` path working unchanged. If you see a cleaner surface, still
ship this default and note the alternative for review.
## 3. Config
**None for Stage 1.** The in-memory adapter is the unconditional default — wire `new InMemoryReplyInbox()`
into `MessageService` in `Bridged.main`. Do **not** add a `broker:` config block (that arrives with the
Stage-2 AMQP adapter: absent → in-memory, present → AMQP).
## 4. Acceptance criteria (what the primary will verify)
1. New `ReplyInbox` + `InboxMessage` + `InMemoryReplyInbox` in `dev.ltms.bridged.msg`.
2. `bridge_reply` with **no open send** now **succeeds and queues** (no more `error` / `409`); the reply is
later retrievable and identical.
3. The queued reply is drainable by the primary keyed by target; draining **acks** it (a second drain
returns nothing); dedup by `msgId` (re-publishing the same id does not double-queue).
4. **QUESTION path unchanged** — `bridge_ask` with no open send still returns `NO_WAITER` (add/keep a test
proving a question is never queued).
5. Completion/failure fallbacks unchanged.
6. Unit tests covering: `InMemoryReplyInbox` publish/peek/ack/dedup/FIFO/concurrency; `MessageService.reply`
queues on no-waiter and resolves-not-queues when a send is open; `drainReplies` returns+acks;
the QUESTION-not-queued guard.
7. **All pre-existing tests still green** (baseline is **188**; your total must be ≥ 188 + your new tests).
## 5. Build & verification (worker side)
- Build with Maven from the worktree's `bridged/` dir. **Capture the exit code without a masking pipe**
(`mvn clean install; echo "MVN_EXIT=$?"` — never `mvn … | tail`, which hides failures).
- Read the real test totals from `target/surefire-reports/TEST-*.xml`, not from stdout scroll.
- You do **not** have IDE MCP access — do not claim `ide_diagnostics` results. The **primary** runs the
authoritative gate (IDE sync + diagnostics 0/0 + `mvn clean install`) before integrating. Your
self-reported counts are inputs to that gate, not final facts.
## 6. Hard constraints (non-negotiable)
- **`.mcp.json` is `--skip-worktree` in your worktree — never edit, `git add`, or commit it.**
- **`wiki/` is a submodule — never run git in it; never touch it.**
- Commit only your feature changes (the new port/adapter, the `msg`/`mcp`/`rest` wiring, tests, and if
you add config wiring in `Bridged.java`). Nothing else.
- Work only inside your assigned worktree on your feature branch. The primary fast-forwards `main` after
re-gating — do not touch `main`.
- Java 25 idioms are welcome (unnamed `_` params, records). Keep the diff minimal and match surrounding style.
## 7. Definition of done (report back over the bridge)
Commit on your feature branch and reply with: the commit SHA, the surefire total (run/failures/errors), a
one-line note on the drain-surface decision (§2.4) you shipped, and confirmation that `.mcp.json`/`wiki/`
were untouched. The primary re-gates and integrates.
+205
View File
@@ -0,0 +1,205 @@
# CB-308 — Multi-Host Federation (Stage 5)
**Status:** design note (proposal)
**Depends on:** CB-307 (broker-based reliable delivery) — CB-308 is the multi-host layer built *on*
CB-307's broker fabric.
**Relates to:** CB-401 (`PeerHandle` opaque id), CB-304 (`rosterView`), CB-306 (spawn-readiness),
CB-303 (lifecycle limits), CB-117 (orphan reap).
## 1. Goal
Let `claude-bridge` coordinate agents that live on **more than one host** — a primary on host A
delegating to workers on hosts B, C, … — without any host learning another host's terminals. The
bus stays a **provider-neutral communication fabric**; multi-host is an addressing + routing
concern, not a new kind of peer.
The design rests on three pieces (the shape this ticket proposes):
1. **Dedicated per-agent channels** — every agent has its own addressable inbox on the broker.
2. **A federated agent directory** — a global "who/where/status" lookup, assembled from per-host
presence, not a central database.
3. **A per-host gateway** — each host runs a `bridged` that owns its local herdr, registers/manages
its own sessions, and proxies messages to/from other hosts over the broker.
## 2. What is single-host today (the assumptions to break)
```mermaid
flowchart TB
subgraph host["Single host (today)"]
primary["primary<br/>(MCP client)"]
daemon["bridged daemon<br/>127.0.0.1:8765"]
reg["in-process registry<br/>keyed by PeerHandle.id() == paneId"]
herdr["herdr<br/>(local unix-socket PTY mux)"]
w1["worker pane wQ:p1"]
w2["worker pane wQ:p2"]
primary --> daemon
daemon --> reg
daemon --> herdr
herdr --> w1
herdr --> w2
end
```
*Figure 1 — everything is co-located and loopback.*
Three concrete bake-ins assume one host:
| Assumption | Where | Why it blocks multi-host |
|---|---|---|
| **herdr is local** | `herdr/` unix socket `~/.config/herdr/herdr.sock` | You cannot drive another host's PTYs → each host **must** own its herdr. This is why a per-host gateway is mandatory. |
| **registry is in-process, keyed by `paneId`** | `session/SessionManager` | `paneId` (e.g. `wQ:p2B`) is a herdr-local coordinate — meaningless off-host. Routing needs a host-unique id. |
| **loopback, no authn** | `rest/BridgedApp` binds `127.0.0.1:8765` | Fine on one host; the moment a second host can talk to a gateway, that link is a trust boundary. |
## 3. Target architecture
```mermaid
flowchart TB
subgraph hostA["HOST A"]
gA["gateway = bridged A"]
regA["local registry + herdr"]
primary["primary (MCP client)"]
gA --- regA
primary --- gA
end
subgraph hostB["HOST B"]
gB["gateway = bridged B"]
regB["local registry + herdr"]
wb["worker panes"]
gB --- regB
gB --- wb
end
subgraph broker["BROKER (LavinMQ / AMQP) — CB-307 fabric"]
inbox["agent.&lt;id&gt;.inbox queues"]
roster["roster.* presence topic"]
dlq["DLQ · delayed-retry (remind)"]
end
gA -->|"publish to agent.&lt;id&gt;.inbox"| inbox
gB -->|"publish to agent.&lt;id&gt;.inbox"| inbox
inbox -->|"owning gateway consumes"| gA
inbox -->|"owning gateway consumes"| gB
gA -->|"announce local agents"| roster
gB -->|"announce local agents"| roster
roster -->|"union view"| gA
roster -->|"union view"| gB
```
*Figure 2 — each gateway owns its local herdr + registry, consumes only its own agents' inboxes,
and announces its agents onto a shared presence topic. The broker routes; no host sees another
host's terminals.*
### 3.1 Component mapping (the three pieces)
- **Dedicated per-agent channels** = a per-agent AMQP routing key / queue, e.g.
`agent.<globalId>.inbox`. The agent's **owning gateway is the only consumer** of its inbox.
Senders publish to `agent.<id>.inbox` and never need to know the agent's host — the broker
routes to whichever gateway holds it. LavinMQ additionally gives durability, DLX, and a native
delayed-message exchange (the remind/backoff loop for free) — the same reasons CB-307 picked it.
- **Federated agent directory** = a **soft-state, bridge-owned** roster, *not* a broker-stored
database. Per the persistence-boundary decision (bridged is soft-state; the broker owns *message*
durability, not *who/where/status*), each gateway announces its local agents `(globalId, host,
status, capabilities)` on a `roster.*` presence topic with periodic heartbeats. Every gateway
builds an eventually-consistent **union view** — literally CB-304's `rosterView`, federated. A
stale entry expires by missed heartbeat (reuses CB-303's idle/TTL thinking).
- **Per-host gateway** = today's `bridged` daemon, evolved. It already registers/manages sessions
and controls its local herdr; multi-host adds exactly two responsibilities: (a) a broker client
that consumes its agents' inboxes and injects into local herdr, and (b) presence announce +
union-roster assembly. Evolution, not rewrite.
### 3.2 Routing rule
```mermaid
flowchart LR
send["bridge_send(globalId, msg)"] --> lookup{"directory:<br/>is globalId local?"}
lookup -->|"yes"| local["inject via local herdr<br/>(today's Injector path)"]
lookup -->|"no"| pub["publish agent.&lt;id&gt;.inbox<br/>(broker routes to owning gateway)"]
pub --> consume["owning gateway consumes<br/>→ injects into its local herdr"]
```
*Figure 3 — one fork: local agents keep today's in-process inject path; remote agents go over the
broker. A sender is oblivious to which branch it took.*
## 4. What CB-307 already provides vs. what is net-new
**CB-307 delivers the transport half** and is independently valuable on a single host: the AMQP
broker fabric, the `bridged → broker` client/adapter, at-least-once + idempotent (dedup-by-id)
delivery, DLQ, and delayed-retry (remind). That *is* the "proxy cross-host message" backbone;
extending the same broker from "worker→primary reliability" to "gateway↔gateway" is incremental.
**Net-new for CB-308 (multi-host), five items:**
1. **Global agent id** — decouple the routing key from `paneId`. CB-401's `PeerHandle` already
abstracts the routing id; make it host-unique (e.g. `<host>/<paneId>` or a UUID minted at spawn).
The registry and all verbs route on the global id.
2. **Federated directory** — presence announce + heartbeat + union roster over `roster.*`
(§3.1).
3. **Gateway routing** — the `local ? inject : publish` fork (§3.2), plus each gateway consuming
its own agents' inbox queues and injecting into local herdr.
4. **Cross-host spawn** — `spawn on host B` = publish a control request to B's control channel →
gateway B runs `ClaudeCodeLauncher.spawn` **locally** (CB-306's readiness gate becomes *more*
valuable here: the far side wants a positive "agent ready" before anyone sends) → announces the
new agent into the federated roster.
5. **Trust** — the broker connection is now the security boundary. A gateway injects env/tokens at
daemon privilege (the CB-401 Stage-C concern), so a **remote-triggered spawn/send** needs
authn/authz: who may act on which host, and which control channels a gateway will honour.
## 5. The one thing the broker does NOT dissolve
The MCP asymmetry survives the network. The primary is an MCP **client** to its **local** gateway;
it cannot be called into. A worker on B replying to a primary on A flows:
```mermaid
sequenceDiagram
participant W as worker (host B)
participant GB as gateway B
participant BR as broker
participant GA as gateway A
participant P as primary (host A, MCP client)
W->>GB: bridge_reply
GB->>BR: publish primary-bound (durable, msg id)
BR->>GA: route to A's primary inbox
Note over GA: held durably until the primary pulls
P->>GA: blocking bridge_send resolves / bridge_poll
GA-->>P: reply (then ACK to broker)
```
*Figure 4 — the broker makes the middle hop lossless, ordered, and idempotent; the **final** hop
into the primary is still a **pull** (gateway A holds the message until the primary's blocking
`bridge_send` or `bridge_poll`). Cross-host neither improves nor worsens this — it just spans hosts.
This is precisely the gap CB-307 closes on one host and CB-308 stretches across hosts.*
## 6. Staging & dependencies
```mermaid
flowchart LR
cb307["CB-307<br/>broker-based reliable delivery<br/>(single host first)"] --> cb308["CB-308<br/>multi-host federation<br/>(this note)"]
cb308 --> a["global agent id"]
cb308 --> b["federated directory"]
cb308 --> c["gateway routing"]
cb308 --> d["cross-host spawn"]
cb308 --> e["cross-host trust model"]
classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff;
class e gate
```
*Figure 5 — CB-307 is the foundation; CB-308's five items build on it. The trust model (e) is the
gating concern before any host accepts remote control.*
**Recommendation:** keep CB-307 scoped to single-host broker reliability (foundation, independently
useful), and build CB-308's items on top once the broker fabric exists. Choose CB-307's broker /
channel naming **multi-host-ready** now (per-agent routing keys, a `roster.*` topic namespace) so
CB-308 doesn't have to repaint the topology.
## 7. Open questions
- **Directory ground-truth:** pure soft-state presence (heartbeats) vs. also treating broker queue
existence as authoritative. Lean soft-state to preserve the persistence boundary; revisit if
split-brain roster views cause mis-routing.
- **Global id scheme:** `<host>/<paneId>` (human-legible, leaks host) vs. opaque UUID (clean, needs
the directory to resolve host). Probably UUID in the protocol, host as directory metadata.
- **Gateway discovery:** how gateways find the broker and each other (static config vs. discovery).
- **Trust model shape:** per-host shared secret vs. mTLS on the broker vs. a capability token per
control action — ties into CB-401 Stage-C.
- **Failure semantics:** a host/gateway dies mid-turn — how the federated roster reaps it (missed
heartbeat) and whether in-flight primary-bound messages survive (broker durability = yes).
+305
View File
@@ -0,0 +1,305 @@
# CB-402 — Second peer adapter: opencode (Stage B of the Peer Launcher SPI)
**Status:** ✅ **complete — implemented, merged (`ded226a`), and live-dogfooded 2026-07-29.**
All five increments of §4 are done, including increment 5 (the §5 live checklist). See
[§8 As-built](#8-as-built--live-dogfood-2026-07-29) for the run. Gitea issue #7 closed.
**Depends on:** CB-401 Stage A (`PeerLauncher` SPI, merged `3aa69a9`)
**Stage:** 4 (Pluggable peers) · Stage B
**Owner action:** design-note → file issue → delegate → primary-verify (per CB-401/306/307)
---
## 1. Goal
Prove the [`PeerLauncher`](../bridged/src/main/java/dev/ltms/bridged/peer/PeerLauncher.java) SPI
actually holds for a **non-Claude** coding agent by shipping a second, first-class in-tree
adapter: **opencode** (`opencode` 1.1.31, a provider-agnostic terminal coding agent).
The product direction is *heterogeneous coding agents, Claude Code first-class* — not a
human/mock peer. opencode is the right proof precisely because it differs from Claude Code on
every seam the SPI is meant to hide:
| Seam | Claude Code | opencode | ⇒ SPI proof |
|------|-------------|----------|-------------|
| Subscription boundary | `ANTHROPIC_BASE_URL` + `SubscriptionGuard.assertWorker()` before any herdr call | none — provider-agnostic, off-subscription by nature | the guard is **Claude-private**, not core |
| MCP mount | inline `--mcp-config '{…}'` launch flag | `opencode mcp add` / config file (`OPENCODE_CONFIG`) — **no inline flag** | "mount the bridge MCP" is adapter-private |
| Instruction injection | `--append-system-prompt "<REPLY_CHARTER>"` | config `instructions` / `AGENTS.md` / `--agent` — **no append flag** | the reply-charter mount is adapter-private |
| Model selection | `ANTHROPIC_MODEL` env | `-m provider/model` flag | env-vs-flag is adapter-private |
| Name / reap scheme | `claude-<profile>-<nonce>-<seq>` | `opencode-<profile>-<nonce>-<seq>` | each adapter reaps only its own kind |
Everything *else* — herdr tab/pane placement, the CB-306 spawn-readiness gate, CB-301-ext
worktree provisioning, CB-117 orphan reap, teardown, `list()`, cwd resolution — is transport
machinery that is **identical** for both. That split is the whole design.
> Out of scope (deferred to Stage C / later): dynamic external plugin loading behind a
> trust/capability model, capability *enforcement* at the verb layer (Stage A only *declares*
> caps), and a `human`/mock peer.
---
## 2. Current state — one launcher, two concerns mixed
`ClaudeCodeLauncher` (584 LOC) is the sole `PeerLauncher`. It interleaves two concerns:
```mermaid
flowchart TB
subgraph CCL["ClaudeCodeLauncher (584 LOC) — today"]
direction TB
T["herdr transport (GENERIC / reusable)<br/>tab-pane placement · spawn-ready gate · worktree<br/>orphan reap · teardown · list · cwd resolution · unique naming"]
C["Claude-specific (per-agent)<br/>ANTHROPIC_BASE_URL + SubscriptionGuard · ANTHROPIC_MODEL<br/>argv --mcp-config · --append-system-prompt REPLY_CHARTER · 'claude-' name prefix"]
end
classDef generic fill:#2f855a,stroke:#22543d,color:#ffffff;
classDef specific fill:#b7791f,stroke:#7b341e,color:#ffffff;
class T generic
class C specific
```
*Figure 1 — the two concerns tangled inside today's single launcher; CB-402 splits them.*
There is also a **Stage-A deferral** to finish: `Bridged.main` still casts
`(ClaudeCodeLauncher) workers` at the `BridgeMcp` and `BridgedApp` constructors. Those two
callers only invoke `profiles()`, `defaultProfile()`, and `list()` — **all already on the
`PeerLauncher` interface**. The cast survives for one reason only: `PeerLauncher.list()`
returns `List<?>` (element type erased) while the callers use `Agent` element methods in their
roster join. Finishing the migration is therefore small and contained (§4.D).
---
## 3. Target design
Template-Method base + two thin adapters + a routing composite that keeps the Stage-A seam
(one `PeerLauncher` reference held by `SessionManager` / `BridgeMcp` / `BridgedApp`) intact.
```mermaid
flowchart TB
IFACE["«interface»<br/>PeerLauncher"]
COMP["CompositePeerLauncher<br/>routes by profile kind; fans out list/reap/caps"]
BASE["«abstract»<br/>HerdrPeerLauncher<br/>transport: placement · ready-gate · reap · stop · cwd · naming"]
CCL2["ClaudeCodeLauncher<br/>hooks: guard+ANTHROPIC_* env · --mcp-config · charter flag · prefix 'claude'"]
OCL["OpenCodeLauncher<br/>hooks: provider env · OPENCODE_CONFIG file · AGENTS charter · prefix 'opencode'"]
IFACE -.implemented by.-> COMP
IFACE -.implemented by.-> BASE
BASE --> CCL2
BASE --> OCL
COMP -->|"kind=claude-code"| CCL2
COMP -->|"kind=opencode"| OCL
classDef iface fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
classDef base fill:#2f855a,stroke:#22543d,color:#ffffff;
classDef leaf fill:#6b46c1,stroke:#44337a,color:#ffffff;
class IFACE,COMP iface
class BASE base
class CCL2,OCL leaf
```
*Figure 2 — extracted base, two adapters, and a routing composite behind the unchanged SPI.*
### A. Extract `HerdrPeerLauncher` (abstract base)
Move all transport machinery down from `ClaudeCodeLauncher`. What stays generic:
- fields `agents`, `spaces`, `profiles`, `defaultProfile`, `env`, `nameSeq`, the CB-306 gate
knobs (`spawnReadyTimeoutMs`/`spawnReadyPollMs`/`nowMillis`/`sleeper`), and `nameNonce`;
- `profiles()`, `defaultProfile()`, `parityOverlay()`, `effectiveCwd(…)`, `resolveCwd`;
- the `spawn(SpawnRequest)` **skeleton**: resolve profile → cfg → *hook* → placement → gate → `WorkerHandle`;
- `spawnInTab` / `spawnAsPane` / `tidy` / `startUniquelyNamed` (name built from a *hook* prefix);
- `list()`, `reapOrphanWorkers()` / `isForeignWorker` / `workerNonce` (pattern built from the prefix hook), `stop()` / `usesTabPlacement` / `isAlreadyGone`;
- `waitUntilInjectableOrThrow`, the `WorkerHandle` record, `putIfPresent`, `resolveEnv`, `sleepUninterruptibly`.
Two adapter **hooks** (abstract):
```java
/** Label prefix for this peer kind; drives unique naming AND the orphan-reap pattern. */
protected abstract String namePrefix(); // "claude" | "opencode"
/** Build the peer-specific launch: env map + argv. Runs any pre-spawn guard here. */
protected abstract Launch buildLaunch(BridgedConfig.Worker cfg, SpawnRequest req);
record Launch(Map<String,String> env, List<String> argv) {}
```
`capabilities()` stays abstract/per-adapter (it already is). The subscription guard is **not**
a base field — it is a constructor dependency of `ClaudeCodeLauncher` alone.
**Reap isolation:** the reap pattern becomes `Pattern.compile(namePrefix() + "-.*-([0-9a-f]{6})-\\d+")`,
so the opencode adapter never reaps a `claude-*` pane and vice-versa. The composite sums both.
### B. `kind:` config discriminator
Add one field to `BridgedConfig.Worker`:
```java
String kind // "claude-code" (default) | "opencode"
```
- Compact-ctor default: `kind = blank ? "claude-code" : kind.toLowerCase()`.
- `argv` default is currently `List.of("claude")`; when `kind=opencode` and the operator left
`argv` unset, default it to `List.of("opencode")`. (Handle in normalization, keyed off `kind`,
so the record stays declarative.)
- Keep the existing back-compat constructors; `kind` is additive and optional.
`bridged.example.yaml` documents a two-kind `workers:` block.
### C. `OpenCodeLauncher` — the adapter hooks for opencode
`namePrefix()` → `"opencode"`. `buildLaunch(cfg, req)`:
- **Env:** *no* `ANTHROPIC_BASE_URL`, *no* `SubscriptionGuard` call. Pass through provider
credentials the operator names (reuse the existing `tokenEnv` indirection; opencode reads
provider keys from env / `opencode auth`). CB-302 git-token injection is reused unchanged
(it is peer-neutral: `GITEA_TOKEN`/`GITEA_HOST`).
- **MCP mount (non-invasive):** opencode has no inline `--mcp-config`. The adapter writes a
throwaway config file and points `OPENCODE_CONFIG=<tempfile>` in the worker env, containing
the bridge MCP server block (opencode HTTP MCP schema, `type: "remote"`) — the opencode analog
of Claude Code's inline flag. Nothing is written into the worker's real project or profile.
- **Reply-charter:** carry `REPLY_CHARTER` as an `instructions` entry in that same generated
config (or an `AGENTS.md` written into the per-worker worktree, which is already a throwaway
isolated checkout under CB-301-ext). Recommend the config-file route to keep the "touch
nothing the user owns" invariant.
- **argv:** `opencode <project-or-cwd>` (interactive TUI, the mode a herdr pane drives), plus
`-m <provider/model>` when the profile sets a model.
> `REPLY_CHARTER` is peer-neutral text — hoist it to a shared constant (base or a small
> `PeerCharter`), consumed by each adapter through its own injection mechanism.
### D. `CompositePeerLauncher` + finish the Stage-A migration
- `Bridged.main` groups configured profiles by `kind`, instantiates one launcher per kind
present, and wraps them in `CompositePeerLauncher implements PeerLauncher`.
- Routing methods (`spawn(req)`, `effectiveCwd(req)`, `parityOverlay(name)`) dispatch by the
profile's kind. Fan-out methods (`list()`, `reapOrphanWorkers()`, `capabilities()`,
`profiles()`, `defaultProfile()`) merge across sub-launchers. `stop(id)` tries each (teardown
only knows the pane id) — already best-effort/idempotent.
- **Migrate `BridgeMcp` + `BridgedApp` to the `PeerLauncher` interface**, dropping both
`(ClaudeCodeLauncher)` casts. Only friction is `list()`'s `List<?>`; resolve by giving the SPI
a typed roster element (small neutral `PeerAgent` view exposing `id()`/`name()`/status) that
the CB-304 roster join consumes — or, minimally, narrow at the callsite. Prefer the typed view.
```mermaid
sequenceDiagram
autonumber
participant P as Primary
participant M as BridgeMcp / REST
participant C as CompositePeerLauncher
participant O as OpenCodeLauncher
participant B as HerdrPeerLauncher (base)
participant H as herdr
P->>M: bridge_spawn(profile="oc-impl")
M->>C: spawn(SpawnRequest)
C->>C: kind(profile)=="opencode"
C->>O: spawn(req)
O->>O: buildLaunch → provider env + OPENCODE_CONFIG file + argv
O->>B: placement + startUniquelyNamed("opencode-…")
B->>H: agent.start(name, argv, env, tab, cwd)
B->>H: poll status until injectable (CB-306 gate)
B-->>O: Agent
O-->>C: PeerHandle(paneId, terminalId)
C-->>M: PeerHandle
M-->>P: session id
```
*Figure 3 — an opencode spawn: composite routes by kind, adapter builds the peer-specific launch, shared base drives herdr + the readiness gate.*
---
## 4. Increment plan (delegate-then-verify friendly)
1. **Extract base, no behaviour change.** Introduce `HerdrPeerLauncher`; make
`ClaudeCodeLauncher` extend it with `namePrefix()="claude"` and `buildLaunch()` wrapping
today's guard+env+argv logic. Green build, identical tests — pure refactor. *(IDE
`refactor` where possible; the primary re-runs the gate workers can't.)*
2. **`kind:` discriminator.** Add the field + normalization + `bridged.example.yaml`. Default
path unchanged (`kind=claude-code`).
3. **`OpenCodeLauncher`.** Implement the three hooks; unit-test `buildLaunch` (env has no
`ANTHROPIC_BASE_URL`; `OPENCODE_CONFIG` points at a file carrying the bridge MCP block +
charter; argv shape).
4. **`CompositePeerLauncher` + wiring + finish Stage-A migration** (drop the two casts).
5. **Live dogfood** (§5) + wiki as-built (primary-gated submodule commit).
Each increment is independently buildable/mergeable; the feature branch stays **unmerged**
until it is "major" (Stage-B whole), per the CB-401 bar.
---
## 5. Risks & validation (live dogfood, not assumed)
- **opencode TUI ⇄ herdr injection.** herdr drives a pane by typing into a TUI. Must confirm
opencode's TUI accepts injected keystrokes/submit the way `claude` does, and reaches an
`injectable` status the CB-306 gate recognizes. *Validation:* spawn one opencode worker,
watch the readiness gate pass, `bridge_send` a trivial task.
- **Bridge MCP visibility in opencode.** Confirm `OPENCODE_CONFIG` (or `opencode mcp add`)
actually surfaces the `bridge_*` tools inside the opencode session, and that `bridge_reply`
is callable — the reply-charter is worthless if the tool isn't mounted. *Validation:* the
worker completes a task by calling `bridge_reply`; the reply lands via the CB-307 path.
- **opencode MCP/config schema drift.** opencode is fast-moving (1.1.31 today). Pin the config
schema we generate against the installed version; treat the exact keys (`type: "remote"` vs
`"http"`, `instructions` shape) as a dogfood-verified fact, not an assumption.
- **Provider credentials.** opencode needs a configured provider (env key or `opencode auth`).
The dogfood profile must name a provider the host actually has, distinct from the primary's
subscription.
---
## 6. Test plan
- **Unit (hermetic):** base-extraction regression (existing `ClaudeCodeLauncher` tests pass
unchanged); `OpenCodeLauncher.buildLaunch` env/argv/config assertions; `kind` normalization
in `BridgedConfigTest`; `CompositePeerLauncher` routing + fan-out (merge of `profiles()`,
summed `reapOrphanWorkers()`, per-kind reap isolation) with fake sub-launchers.
- **Live (dogfood, manual):** the §5 checklist on the running daemon.
- **Gate (primary):** IDE diagnostics 0/0 on every changed file, `mvn clean install` green with
the surefire summary captured (not `| tail`), manual diff review — the authoritative checks a
worker cannot self-run.
---
## 7. Open questions for the lead
1. ✅ **Provider — resolved 2026-07-29.** None was needed. opencode's own gateway serves
**free-tier models with zero credentials** (`opencode auth list` → *0 credentials*, yet
`opencode run -m opencode/north-mini-code-free` answers). The dogfood profile uses
`opencode/north-mini-code-free`. It is distinct from the primary's Anthropic subscription by
construction, and needs no `guard` entry — opencode carries no `ANTHROPIC_BASE_URL`, so the
`SubscriptionGuard` does not apply to it at all.
2. ✅ **Charter carrier — confirmed as designed:** generated `OPENCODE_CONFIG` `instructions`.
Verified working against the installed version.
3. ✅ **Merge cadence — resolved as it happened:** Stage B landed as one unit (`ded226a`).
---
## 8. As-built — live dogfood (2026-07-29)
Run against `bridged` on `127.0.0.1:8766` at main `19cdf8d`, with opencode **1.18.5** installed
via Homebrew. Every §5 risk is now a verified fact rather than an assumption.
**The version-drift risk was the real one, and it did not bite.** This adapter was designed against
opencode **1.1.31**; the installed version is **1.18.5**. The generated config schema still
validates unchanged — `mcp.<name>.type: "remote"`, `url`, `enabled`, and `instructions: [path]` are
all accepted, and `OPENCODE_CONFIG=… opencode mcp list` reports `✓ bridge connected`. Pinned here
as a dogfood-verified fact for 1.18.5.
| §5 risk | Result |
|---|---|
| opencode TUI ⇄ herdr injection; CB-306 gate | ✅ `peer pane=wD:p3 reached injectable state` ~0.6s after `agent.start` |
| Bridge MCP visible + `bridge_reply` callable | ✅ MCP `initialize` from `Implementation[name=opencode, version=1.18.5]`; worker replied through the tool |
| Config schema drift (1.1.31 → 1.18.5) | ✅ unchanged, see above |
| Provider credentials | ✅ free tier, zero credentials |
Full lifecycle exercised through the REST surface:
1. `POST /workers?profile=opencode-free` → `201`, routed by `kind:` through `CompositePeerLauncher`
to `OpenCodeLauncher` (`spawning opencode profile=opencode-free`), pane `wD:p3`.
2. Readiness: `{"ready":true,"status":"idle"}`, roster state `ready`.
3. `POST /sessions/{id}/message` → **`{"replySource":"reply","reply":"391"}`** — a *structured*
`bridge_reply`, not the CB-115 completion-fallback transcript scrape. The clean path.
4. `DELETE /workers/wD:p3` → `204`, roster empty, tolerant teardown (`tab_not_found` ignored —
opencode had already closed its own tab).
**Unplanned cross-validation with CB-501.** The audit trail recorded the worker's reply as
`role: WORKER, actor: worker:term_657c1dad2b9731e, action: REPLY, outcome: allowed`. Connection-based
identity (loopback peer PID → herdr pane) classified an **opencode** process as a worker with no
opencode-specific handling — confirming the identity model is peer-kind-agnostic, which is exactly
what CB-308 needs when it stretches the roster across hosts.
CB-502 counters for the same run: `bridged_sends_total{outcome="replied"} 1`,
`bridged_replies_total{path="rendezvous"} 1`, `bridged_inbox_depth{...} 0`.
+460
View File
@@ -0,0 +1,460 @@
# CB-500 — Multi-Tier Coordination (Stage 6)
**Status:** design note (proposal — ticket split deferred)
**Depends on:** CB-401/402 (Peer Launcher SPI + composite router — placement-neutral spawn),
CB-308 (per-agent broker channels + global id + federated roster — the addressing substrate),
CB-307 (durable inbox + push loop), CB-301/303 (session FSM + context-cap/idle-ttl), CB-304
(`rosterView`).
**Relates to:** the bus-identity boundary — see §7. This note stays a **proposal**; no code until the
staging in §6 is reviewed and the arc is split into tickets.
## 1. Goal
Grow `claude-bridge` from a **single-tier** coordinator (one human-driven primary → a flat pool of
workers) into a **multi-tier** one, along three axes the lead has asked for:
1. **Sandboxed workers** — each worker runs in a **separated, peer-owned sandbox** carrying its own
toolchain (Claude routed via `ANTHROPIC_BASE_URL`, a headless IDE, git, MCP, dev-tools), with
**per-role** sandboxes (a backend-agent image, a frontend-agent image).
2. **Main-agent pairs** — the "main" tier becomes a **pair** (on-subscription Opus + one cloud
module) collaborating, instead of a lone primary.
3. **An orchestrator tier** — a supervisor **above** the mains that owns their **session identity**
(naming, resume) and **curates context**, so every main→worker delegation carries the *exact*
slice of context it needs and nothing else.
The through-line: **this is not a new pillar.** It is the existing `PeerLauncher` and
`SessionManager` patterns extended one tier up, riding the **same CB-308 substrate** that multi-host
already needs. Sandbox = a placement-neutral spawn target (CB-402 pattern). Pair + orchestrator =
per-agent channels + a recursive session manager (CB-308 pattern). The bus stays a
**provider-neutral communication fabric**; every addition is addressing, launch, or session scoping —
never toolchain ownership (§7).
## 2. Single-tier today (the assumptions to break)
```mermaid
flowchart TB
human["human (types)"]
primary["PRIMARY (Opus)<br/>MCP client — pull-only"]
daemon["bridged daemon<br/>127.0.0.1:8765 (single host)"]
comp["CompositePeerLauncher<br/>routes by kind"]
cc["ClaudeCodeLauncher"]
oc["OpenCodeLauncher"]
w1["worker pane (gx00 vLLM)"]
w2["worker pane (ollama)"]
human --> primary
primary -->|"bridge_send / spawn / ask"| daemon
daemon --> comp
comp --> cc
comp --> oc
cc --> w1
oc --> w2
```
*Figure 1 — one human-driven primary, one daemon, a flat pool of bare herdr-pane workers.*
Four concrete bake-ins assume a single tier:
| Assumption | Where (verified) | Why it blocks the direction |
|---|---|---|
| **Exactly one primary** | `mcp/PrimaryRegistry` — an `AtomicReference<String>`, "single-slot registry for the primary's terminal" | A *pair* needs N addressable mains, each with its own pull inbox. |
| **Workers are bare panes** | `worker/*Launcher` spawn a herdr pane via `argv:["ccs", …]` into a pre-existing env | A *sandbox* is a richer launch target (container/devcontainer) — a new placement, not a new provider. |
| **`SpawnRequest` is flat** | `peer/SpawnRequest(profileName, requestedCwd, callerCwd)` | A sandbox/role selection needs a spawn-target dimension the record does not carry. |
| **No tier above the primary** | there is no manager of the *primary's own* session — `SessionManager` manages *workers* only | An orchestrator that names/resumes/scopes the mains is a wholly new (but pattern-reusable) tier. |
## 3. Target multi-tier architecture
```mermaid
flowchart TB
human["human"]
subgraph orch["TIER 0 — orchestrator"]
osm["OrchestratorSessionManager<br/>(SessionManager, recursed up)<br/>names · resumes · scopes context"]
end
subgraph mains["TIER 1 — main pair"]
m1["main A: Opus<br/>MCP client"]
m2["main B: cloud module<br/>MCP client"]
end
subgraph bus["bridged fabric (CB-307/308 substrate)"]
chan["per-agent inbox channels<br/>agent.&lt;globalId&gt;.inbox"]
roster["federated roster (union view)"]
end
subgraph workers["TIER 2 — sandboxed workers"]
sbBE["backend sandbox<br/>Claude via ANTHROPIC_BASE_URL<br/>+ headless IDE · git · MCP · dev-tools"]
sbFE["frontend sandbox<br/>(role-specific image)"]
end
human --> osm
osm -->|"spawn / name / resume"| m1
osm -->|"spawn / name / resume"| m2
m1 <-->|"pull inbox"| chan
m2 <-->|"pull inbox"| chan
m1 -->|"scoped delegation"| bus
m2 -->|"scoped delegation"| bus
bus --> sbBE
bus --> sbFE
chan --- roster
```
*Figure 2 — three tiers. Tier 0 owns the mains' session identity + context scope; Tier 1 is a
collaborating pair, each an MCP client with its own pull inbox; Tier 2 is peer-owned sandboxes the
bus launches into. The middle is CB-308's per-agent-channel + federated-roster substrate, now
carrying tier-to-tier traffic, not just host-to-host.*
The recursion is the key idea: **`orchestrator : mains :: main : workers`** — the same
spawn/name/resume/scope verbs at two levels.
## 4. Development A — Sandboxed, role-specific workers
A "sandbox" is a **placement**, not a provider — so it slots into the CB-401 SPI exactly the way
CB-402's opencode adapter slotted in as a new *provider*. CB-402 proved the SPI is
provider-neutral; a `SandboxLauncher` proves it is **placement-neutral**.
```mermaid
flowchart TB
req["SpawnRequest<br/>(profileName, cwd, + sandbox/role)"]
comp["CompositePeerLauncher<br/>routes by kind"]
cc["ClaudeCodeLauncher<br/>kind: claude-code"]
oc["OpenCodeLauncher<br/>kind: opencode"]
sb["SandboxLauncher (NEW)<br/>kind: sandbox"]
subgraph owned["bridge OWNS (launch + inject boundary)"]
launch["run sandbox entrypoint<br/>(docker/devcontainer up → agent)"]
inject["inject + guard baseUrl,<br/>mount bridge MCP + charter"]
end
subgraph peer["peer OWNS (inside the sandbox)"]
img["image = backend|frontend role<br/>headless IDE · git · dev-tools · MCP"]
end
req --> comp
comp --> cc
comp --> oc
comp --> sb
sb --> launch --> inject
inject -.->|"launches into, never builds"| img
classDef line fill:#b7791f,stroke:#7b341e,color:#ffffff;
class inject line
```
*Figure 3 — the ownership line (amber). The bridge runs the sandbox entrypoint and injects the same
boundary it owns today (guarded `baseUrl`, mounted MCP + reply charter). Everything inside the image
— the IDE, git, dev-tools — is the peer's. The bridge references the image/role; it never provisions
tools. This is what keeps "give the worker a headless IDE" on the right side of the "bus, not
env-manager" rule (§7).*
```mermaid
sequenceDiagram
participant M as main (delegator)
participant D as bridged
participant SL as SandboxLauncher
participant SB as sandbox (peer-owned)
participant A as agent in sandbox
M->>D: bridge_spawn(profile=backend, role=backend)
D->>SL: spawn(SpawnRequest)
SL->>SB: start entrypoint (image = backend role)
Note over SL,SB: bridge injects guarded ANTHROPIC_BASE_URL,<br/>mounts bridge MCP url + reply charter
SB->>A: launch Claude (headless IDE, git, MCP ready — peer's own)
A-->>SL: MCP connects → readiness gate (CB-306) passes
SL-->>D: PeerHandle(globalId)
D-->>M: spawned, injectable
```
*Figure 4 — spawn into a peer-owned sandbox. Identical control flow to today's pane spawn (incl. the
CB-306 readiness gate); only the launcher's `buildLaunch` differs — exactly the CB-402 seam.*
**Deltas:** a new `kind: sandbox` adapter (extends the same `HerdrPeerLauncher`/`PeerLauncher` base);
a spawn-target/role dimension on `SpawnRequest` and the `Worker` profile; optionally a `SANDBOX`
`Capability`. Per-role = two profiles → two images; `CompositePeerLauncher` already routes them. If a
sandbox is a *separate host/container*, it reuses CB-308's global id + per-host gateway wholesale —
**the distributed case is resolved in §11: a sandbox on another host is one spawned by that host's
gateway, because herdr keystroke-injection needs a locally-owned PTY.**
## 5. Development B — Main-agent pairs
Both mains are MCP **clients**, so **neither can be called into** — each needs a **pull-based
per-agent inbox**, which is precisely CB-308 item #1 (per-agent AMQP channels). The primary machinery
that is singular today (single-slot `PrimaryRegistry`, a push-loop aimed at one terminal, "these
tools only the primary calls") generalizes from a singleton to a **set**.
```mermaid
flowchart TB
subgraph pair["TIER 1 — collaborating pair"]
m1["main A: Opus<br/>MCP client (pull-only)"]
m2["main B: cloud module<br/>MCP client (pull-only)"]
end
reg["PrimaryRegistry → multi-slot<br/>(terminal per main)"]
subgraph fabric["bridged"]
ca["agent.A.inbox"]
cb["agent.B.inbox"]
push["ReplyPushLoop → N terminals"]
end
m1 <-->|"peer-to-peer message"| m2
m1 -->|"register terminal"| reg
m2 -->|"register terminal"| reg
ca -->|"pull / nudge"| m1
cb -->|"pull / nudge"| m2
reg --> push
push --> ca
push --> cb
```
*Figure 5 — the pair. Each main owns an addressable inbox; `PrimaryRegistry` becomes multi-slot; the
push loop nudges each main's terminal. Mains message each other as equals over the same bus (the
transport is already peer-neutral — what was missing is N pull endpoints).*
```mermaid
sequenceDiagram
participant MA as main A (Opus)
participant BR as bridged / broker
participant MB as main B (cloud)
MA->>BR: bridge_send(to = main B, msg)
BR->>BR: publish agent.B.inbox (durable, msg id)
Note over BR: held until B pulls (B is a client too)
MB->>BR: blocking bridge_send / poll resolves
BR-->>MB: msg (then ACK)
MB->>BR: bridge_reply(to = main A)
BR->>BR: publish agent.A.inbox
MA->>BR: poll resolves
BR-->>MA: reply
```
*Figure 6 — main↔main is the CB-307 asymmetry applied on both ends: two clients, so both hops are
pull. This is why Part B **depends on** the per-agent-channel substrate, not just a config flag.*
**Deltas:** `PrimaryRegistry` single-slot → keyed-by-main; per-main inbox routing (CB-308 #1);
push-loop fan-out; relax "orchestration tools only the primary calls" to "any registered main."
## 6. Development C — Orchestrator tier
The orchestrator is **`SessionManager` recursed one tier up**: today it spawns/names/reaps *worker*
sessions; the orchestrator does the same for *main* sessions, and adds **context scoping**.
```mermaid
flowchart TB
human["human"]
subgraph t0["TIER 0 — orchestrator (new top MCP client)"]
osm["OrchestratorSessionManager<br/>= SessionManager pattern"]
idm["session identity<br/>name · resume · idle-ttl (CB-303)"]
ctx["context scoper<br/>(CB-303 context-cap + turn_id)"]
end
subgraph t1["TIER 1 — mains (now MANAGED sessions)"]
m1["main A"]
m2["main B"]
end
subgraph t2["TIER 2 — workers"]
w["sandboxed workers"]
end
human --> osm
osm --> idm
osm --> ctx
idm -->|"spawn / name / resume"| m1
idm -->|"spawn / name / resume"| m2
ctx -->|"inject exact context slice"| m1
m1 -->|"scoped delegation"| w
m2 -->|"scoped delegation"| w
```
*Figure 7 — the recursion. Tiers 1 and 2 run the identical spawn/name/resume machinery; the
orchestrator merely operates it one level higher. **Re-rooting caveat:** today the primary IS the
human's live session; here the human drives the orchestrator, and the mains become managed,
resumable sessions. That moves the human-facing top up a tier — an intentional re-root, not an
add-on.*
```mermaid
sequenceDiagram
participant H as human
participant O as orchestrator
participant MA as main A
participant W as worker
H->>O: high-level goal (large context)
O->>O: name/resume main A session
O->>MA: task + SCOPED context slice (not the whole history)
MA->>W: bridge_send(delegation, carrying only the relevant slice)
W-->>MA: result
MA-->>O: rollup
O->>O: fold into orchestrator context, pick next main/turn
```
*Figure 8 — context focus. The orchestrator holds the global context and hands each main only the
slice a given delegation needs, so the main→worker conversation stays on-point. Context *scoping* is
coordination (the bus already owns session/turn lifecycle) — it stays inside the identity boundary
(§7), unlike toolchain ownership which does not.*
**Deltas:** a second, higher `SessionManager` instance whose "peers" are mains; the orchestrator
becomes the top MCP client; context-slice selection (new) layered on CB-303's `context_cap` +
`turn_id` scoping; mains gain a resumable session id in the federated roster.
## 7. The identity boundary — the one clause to hold
The bus is a **communication fabric, not an env/toolchain manager**. This direction is compatible
**only** with the ownership split below; the amber line in Figure 3 is where it must hold.
```mermaid
flowchart LR
subgraph ok["STAYS A BUS (owned)"]
a["launch INTO a sandbox<br/>(opaque image/role reference)"]
b["inject + guard baseUrl,<br/>mount MCP + charter"]
c["session identity + context scope<br/>(name/resume/turn_id)"]
d["per-agent addressing + roster"]
end
subgraph drift["BECOMES ENV-MANAGER (forbidden)"]
e["build images / install IDE<br/>or dev-tools"]
f["wire the bridge's OWN IDE MCP<br/>into a worker"]
g["enumerate 'what a frontend<br/>agent needs'"]
end
ok -.->|"red flag: any feature that only<br/>makes sense for ONE kind of peer"| drift
classDef bad fill:#9b2c2c,stroke:#742a2a,color:#ffffff;
class e,f,g bad
```
*Figure 9 — the guardrail. A worker having a headless IDE **inside its own sandbox** is the peer
owning its toolchain (left) — the opposite of the bridge reaching into the peer (right). Sandbox
specs are peer-owned references (like `argv`/image id); the instant the bridge builds or installs
them, it has drifted. This resolves the apparent contradiction between "do NOT give workers IDE MCP
access" and "give workers a sandboxed IDE" — different owners.*
## 8. Staging & dependencies
```mermaid
flowchart LR
cb402["CB-401/402<br/>Peer Launcher SPI + composite<br/>(DONE / in-flight)"]
A["A · SandboxLauncher<br/>(placement-neutral, independent)"]
cb308["CB-308 substrate<br/>per-agent channels + global id<br/>+ federated roster"]
B["B · main-agent pair<br/>(multi-slot PrimaryRegistry)"]
C["C · orchestrator tier<br/>(SessionManager recursed up)"]
cb402 --> A
cb402 --> cb308
cb308 --> B
cb308 --> C
B --> C
A -.->|"if sandbox = separate host/container,<br/>reuses CB-308 global id"| cb308
classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff;
class cb308 gate
```
*Figure 10 — the substrate (amber) is the shared enabler for B and C. Recommended order:*
1. **Finish CB-402** (merge the opencode adapter branch).
2. **A · SandboxLauncher** — independent; a second proof of the SPI (placement-neutral). Ships anytime.
3. **CB-308 substrate** — per-agent channels + global id + federated roster (the multi-host work,
promoted from host-to-host to tier-to-tier).
4. **B · main-agent pair** — multi-slot `PrimaryRegistry` + per-main inbox routing, on the substrate.
5. **C · orchestrator tier** — the capstone; the recursive session manager + context scoping.
## 9. Open questions (to resolve at ticket-split)
- **Sandbox mechanism:** container (`docker exec`) vs devcontainer — how the role→image mapping is
expressed on the profile. *(Topology **resolved** in §11: distributed = gateway-per-host × local
sandboxes; the remaining choice is only the local launch mechanism, not the shape.)*
- **Pair semantics:** are the two mains fully symmetric peers, or is one a co-primary that may also
delegate? Affects how `PrimaryRegistry` and the "orchestration tools" identity relax.
- **Orchestrator drivenness:** the mains become programmatically spawned/resumed — does the human
still ever type directly into a main, or only into the orchestrator? (The re-root caveat, Fig 7.)
- **Context-slice selection:** who decides the slice — orchestrator heuristics, explicit tool args,
or the main pulling on demand? This is the genuinely new responsibility; keep it *scoping*, not
content authorship, to stay inside the boundary.
- **Trust:** every new tier boundary that accepts spawn/send is a trust edge (the CB-308 item #5 /
CB-401 Stage-C concern) — orchestrator→main and main→sandbox both need authz.
## 10. Ticket-split guidance (deferred)
This note is deliberately one arc; when split, the natural tickets are **A** (SandboxLauncher +
role/spawn-target), **the CB-308 substrate** (likely already its own ticket), **B** (multi-primary
pair), and **C** (orchestrator tier + context scoping) — with the identity clause (§7) as an
acceptance criterion on **A** specifically. Sequence per §8; nothing here is a new pillar, so each
ticket is an extension of an existing pattern (CB-402 for A, CB-308 for B/C).
## 11. Distributed sandboxes — the resolved topology
The follow-up question — *"clarify the architecture when we have distributed agents in sandboxes"* —
resolves the fork left open in §4 and §9. **Decision: Development A (sandbox launcher) and CB-308
(per-host federation) *compose*, not compete — each host runs a `bridged` gateway whose launcher
spawns agents into that host's *local* sandboxes.** A sandbox is never reached across the network; it
is reached by the gateway sitting next to it.
### 11.1 The one fact that fixes the shape
The bus delivers a turn by **herdr keystroke-injection** — `Injector → AgentControl.send` writes into
a PTY that its **local** herdr owns. The broker moves *messages and presence*, **never keystrokes**.
So an agent's PTY must live in a herdr that *some* `bridged` instance drives locally: a remote
container with no local herdr **cannot be injected into**. That rules out a central daemon reaching
remote PTYs, and collapses the design to a single identity:
> **"a sandboxed agent on another host" ≡ "a sandbox spawned by that host's gateway."**
```mermaid
flowchart TB
subgraph hostA["HOST A — gateway"]
mA["main / orchestrator<br/>MCP client → LOCAL gateway"]
gA["bridged A<br/>herdr + CompositePeerLauncher<br/>(incl. SandboxLauncher)"]
cBEa["sandbox: backend<br/>(local container)"]
cFEa["sandbox: frontend<br/>(local container)"]
mA --- gA
gA -->|"spawn (docker/devcontainer)<br/>→ PTY in A's herdr"| cBEa
gA --> cFEa
end
subgraph broker["BROKER (AMQP) — CB-307/308 fabric"]
inbox["agent.ID.inbox queues"]
roster["roster.* (federated presence)"]
end
subgraph hostB["HOST B — gateway"]
gB["bridged B<br/>herdr + SandboxLauncher"]
cBEb["sandbox: backend<br/>(local container)"]
gB -->|"spawn → PTY in B's herdr"| cBEb
end
gA <-->|"messages + presence<br/>(NOT keystrokes)"| inbox
gB <-->|"messages + presence"| inbox
gA --- roster
gB --- roster
classDef line fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
class inbox,roster line
```
*Figure 11 — the composed topology. Each gateway owns its local herdr and runs a `SandboxLauncher`
(the §4 adapter) that spawns role-specific containers **on its own host**; the broker (blue) carries
only messages + roster between gateways. Keystroke-injection stays strictly local to each gateway.*
### 11.2 How a delegation reaches a sandboxed agent on another host
```mermaid
sequenceDiagram
participant MA as main (host A)
participant GA as gateway A
participant BR as broker
participant GB as gateway B
participant SB as sandbox agent (host B, container)
MA->>GA: bridge_send(globalId on B, msg)
GA->>GA: directory lookup - is globalId local? NO
GA->>BR: publish agent.ID.inbox (durable)
BR->>GB: route to the owning gateway
GB->>SB: inject via B's LOCAL herdr (keystrokes)
Note over GB,SB: SandboxLauncher already spawned the container -<br/>its PTY is in B's herdr, CB-306 readiness passed
SB-->>GB: bridge_reply (to B's LOCAL MCP endpoint)
GB->>BR: publish primary-bound (durable, msg id)
BR->>GA: route back to A
Note over GA: held until the main pulls (the main is a client)
MA->>GA: poll / blocking send resolves
GA-->>MA: reply
```
*Figure 12 — the `local ? inject : publish` fork (CB-308 §3.2) with a sandboxed far side. Only the
**middle** hop crosses the network via the broker; **both** injection points (into the sandbox on B,
and the drain-nudge back into the main on A) are local herdr writes. This is CB-308's routing rule
unchanged — the sandbox is transparent to it.*
### 11.3 Two reachability changes any sandbox forces
| Change | Today | Under sandboxes |
|---|---|---|
| **`mcpUrl`** | `http://127.0.0.1:8765/mcp` (loopback) | must be **host-routable from inside the container** (e.g. `host.docker.internal` or the gateway's LAN IP) — the worker connects to **its own gateway's** MCP, never a remote one. |
| **PTY ownership** | pane in the daemon's herdr | pane is the **container's** attached PTY, in the **local** gateway's herdr (via `docker exec`/devcontainer) — non-negotiable per §11.1. |
### 11.4 Everything maps to an existing seam (nothing new invented)
| Concern | Provided by |
|---|---|
| Per-host gateway (owns local herdr + sessions) | **CB-308** (today's `bridged`, evolved) |
| Spawn into a local sandbox / role→image | **Development A** `SandboxLauncher` (§4), routed by `CompositePeerLauncher` |
| Addressing a remote sandboxed agent | **CB-308** global id + federated roster (host + role as metadata) |
| Orphan reap after a gateway restart | **CB-117** per-gateway, summed by the composite — each reaps only its **local** herdr |
| Container up/down | tied to **CB-303** session lifecycle — `SandboxLauncher.stop` tears the container down with the pane |
| Cross-gateway spawn/send trust | **CB-308 item #5** / CB-401 Stage-C — each gateway edge is a trust boundary |
*The net: distributed sandboxes add **zero** new pillars — they are `CB-308 gateway × Development-A
launcher` at every host, with the §7 ownership line (peer owns the image; the bridge only launches
into it) holding at each gateway.*
+242
View File
@@ -0,0 +1,242 @@
# CB-5xx — Stage 5 Hardening (auth · metrics · CI · supervision · authz+audit)
**Status:** design note (pre-implementation) — the single-host close-out before cross-host work.
**Covers:** CB-501 (bearer auth + TLS) · CB-502 (`/metrics`) · CB-503 (mock-socket CI) ·
CB-504 (service supervision) · CB-505 (per-session authz + audit log).
**Depends on:** everything shipped through CB-402. Nothing here changes messaging semantics.
**Blocks:** CB-308. The cross-host trust model is CB-308's own gating concern, and it inherits
whatever identity/authz shape lands here — so this stage is deliberately *before* federation,
not after it.
---
## 1. Why this stage is not optional bookkeeping
`bridged` today has **exactly one security control: the loopback bind**. Every other guarantee
rests on it.
The identity model (`mcp/ConnectionIdentity.java`) resolves a caller from the connection alone —
the OS reports the connecting PID, herdr owns the PID→pane map, so a worker cannot forge another
worker. Its own javadoc is explicit: *"Single-host only (the herd shares the `bridged` host); the
token path is the split-host fallback."* The token path does not exist yet.
That leaves a seam that is **latent today and load-bearing the moment the bind moves**:
```java
// ConnectionIdentity.resolve — non-loopback callers get a null terminal
if (!isLoopback(remoteAddr)) return new Caller(null, -1);
```
…and `null` terminal is interpreted downstream as **"this caller is the primary"**. Combined:
> Any caller that is not a recognised on-host worker pane is treated as the primary — including,
> if `bind.host` is ever widened, an arbitrary remote client.
Today `bind` defaults to `127.0.0.1` so this is unreachable. But CB-308 exists precisely to widen
the boundary, and the primary is the *most* privileged role on the bus (it spawns, stops, sends to
any session, and drains any inbox). Shipping federation on top of "unauthenticated ⇒ primary"
would be building the security boundary backwards.
**So CB-501 is not "add a token header". It is: make identity explicit, and make the absence of
identity mean *nothing*, not *everything*.**
---
## 2. Decisions (locked)
### D1 — Three caller roles, one resolution path
Introduce `Role { PRIMARY, WORKER, ANONYMOUS }` resolved by a single `CallerResolver` that both
REST and MCP go through. Resolution order:
1. **Connection identity wins where it applies.** A loopback peer PID that maps to a herdr worker
pane ⇒ `WORKER` with that terminal. Unforgeable, unchanged from today, zero config.
2. **Token, if presented.** A valid bearer token ⇒ the role that token is provisioned for.
3. **Otherwise `ANONYMOUS`** — *not* `PRIMARY`.
This inverts today's default. `PRIMARY` becomes something you must *prove* (by being a loopback
non-worker process when auth is disabled, or by presenting a primary-scoped token when it is
enabled), rather than something you get by failing every other check.
### D2 — Auth is opt-in by config, but the *default* must stay zero-friction on loopback
The daemon is dogfooded constantly on one machine. If enabling hardening breaks the local setup,
it will be disabled and the stage is wasted. So:
```yaml
auth:
mode: loopback-trust # default — behaves exactly like today: loopback ⇒ PRIMARY, no token needed
# mode: token # every non-worker caller must present a valid bearer token
# tokenEnv: BRIDGED_API_TOKEN # host env var holding the token; never the literal value
```
`mode: loopback-trust` is the current behaviour, named honestly and now *chosen* rather than
implied. `mode: token` is what a non-loopback bind requires. **A non-loopback `bind.host` with
`mode: loopback-trust` must fail fast at startup** — that check is the single highest-value line
in this stage, because it makes the dangerous configuration unrepresentable rather than merely
discouraged.
### D3 — TLS terminates *outside* the JVM
Do **not** add TLS config to Javalin/Jetty. The deployment story for a cross-host gateway is a
reverse proxy (or an SSH/WireGuard tunnel) in front of the daemon; the AMQP link has its own TLS
via the broker URI (`amqps://`). Adding keystore handling here would mean certificate lifecycle
code in a daemon whose whole value is being small, and would duplicate what the proxy does better.
**CB-501 therefore ships bearer auth + the fail-fast bind check, and documents TLS as a
deployment concern with a worked reverse-proxy example.** This is a deliberate narrowing of the
roadmap's "auth/TLS" wording — flagged in §6 for the lead.
### D4 — Metrics without a new dependency
The roadmap's tech-stack table says Micrometer→Prometheus. Recommend **not** taking that dep:
- This pom already carries an unusually heavy dependency-reconciliation burden (a hand-pinned
`jackson-annotations` 3.0-rc5 to reconcile the MCP SDK's Jackson 3 with our Jackson 2.19, a
Jetty BOM import to stop version skew, plus four documented accepted-CVE advisories). Every new
transitive tree is a real cost here, not a hypothetical one.
- The CVE gate that CLAUDE.md mandates for dependency changes (`jetbrains get_file_problems` →
Mend.io) **cannot currently be run** — no JetBrains MCP server is connected. Adding a dependency
tree we cannot scan violates the project's own stated policy.
- The metric set is small and fully known (§4). Prometheus text exposition is a trivial,
stable, well-specified format.
So: a ~120-line `metrics/Metrics.java` holding `LongAdder` counters and gauge suppliers, rendered
to the Prometheus text format at `GET /metrics`. If Micrometer is wanted later for its
registry/push ecosystem, this stays a drop-in swap behind the same endpoint. **Flagged in §6 —
this deviates from a documented tech-stack choice.**
### D5 — Supervision targets launchd first, systemd second
The roadmap says "systemd unit". **This host is macOS — there is no systemd on it** (`systemctl`
not found), and the daemon that has been dogfooded for weeks runs as a bare foreground
`java -jar`. Ship **both**:
- `deploy/dev.ltms.bridged.plist` — launchd agent, the *actual* runtime here, with `KeepAlive` and
ordered start after herdr.
- `deploy/bridged.service` — systemd unit for the Linux gateways CB-308 introduces.
Ordering after herdr is advisory in both: the herdr socket may not exist at boot, so the daemon
must **retry the socket rather than exit** — supervision ordering is a nicety, socket-retry is the
actual fix. That retry behaviour is part of CB-504, not a separate ticket.
### D6 — Audit log is a separate append-only stream, not the app log
Privileged actions (spawn, stop, send, reply-drain, ack) emit a structured JSON line to a
dedicated `audit` SLF4J logger with its own appender, carrying `{ts, role, terminal, pid, action,
target, outcome}`. Keeping it off the chatty app logger is what makes it greppable and, later,
shippable. **No message *content* in the audit record** — the bridge carries the user's source
code and prompts; an audit trail that quietly becomes a transcript archive is a liability, not a
control. Content stays out; correlation ids go in.
---
## 3. Authorization model (CB-505)
With D1's roles, the rules are small enough to state completely:
| Action | REST | PRIMARY | WORKER | ANONYMOUS |
|---|---|---|---|---|
| spawn worker | `POST /workers` | ✅ | ❌ | ❌ |
| stop worker | `DELETE /workers/{paneId}` | ✅ | ❌ | ❌ |
| send to a session | `POST /sessions/{id}/message` | ✅ | ❌ | ❌ |
| reply | `POST /sessions/{id}/reply` | ❌ | ✅ **own session only** | ❌ |
| ask | `POST /sessions/{id}/ask` | ❌ | ✅ **own session only** | ❌ |
| drain replies | `GET /sessions/{id}/replies` | ✅ | ❌ | ❌ |
| status / list / profiles | `GET …` | ✅ | ✅ | ❌ |
| health | `GET /healthz` | ✅ | ✅ | ✅ (unauthenticated by design) |
| metrics | `GET /metrics` | ✅ | ✅ | ❌ |
The load-bearing row is **"own session only"**: a worker may only reply or ask *as itself*. That is
already true de facto — `ConnectionIdentity` derives the terminal rather than reading it from the
body — so CB-505 mostly **asserts an existing invariant explicitly** and adds the test that pins
it. The one real change is rejecting a worker that names a *different* session id in the path.
`/healthz` stays open: it must answer for a load balancer or supervisor before any credential is
configured. It already leaks nothing but herdr's version and up/down.
### 3.1 There are TWO entry paths, and only one of them has identity today
The wiki describes MCP as "a thin adapter over the REST core". **At the code level that is not
literally true, and the difference is security-relevant.** `BridgeMcp` calls `MessageService` /
`SessionManager` *directly*; it never issues an HTTP request against a Javalin route. And `/mcp` is
mounted as a raw servlet on Jetty's `ServletContextHandler`
(`BridgedApp.build → cfg.jetty.modifyServletContextHandler`), so it does **not** pass through
Javalin's `before` filters at all.
The current split is the mirror image of what you'd expect:
| Path | Caller identity today | Authz today |
|---|---|---|
| MCP `/mcp` | ✅ resolved per call (`ConnectionIdentity` via the transport-context extractor) | ❌ none |
| REST routes | ❌ **none at all** — the session id is taken from the URL path and trusted | ❌ none |
So REST is the *more* exposed surface: `POST /sessions/{id}/reply` accepts any `{id}` from the
path, whereas the MCP `bridge_reply` derives the worker from the connection and refuses to read it
from an argument. Loopback-only bind is what makes this safe today.
**Therefore CB-505 must enforce on both paths against one shared resolver** — not at a single
choke point. Concretely: a Javalin `before` filter for REST, and the existing transport-context
extractor for MCP, both delegating to `auth.CallerResolver`. Any authz check that lives in only
one of the two is not a control.
---
## 4. Metric set (CB-502)
Deliberately small; every one maps to a failure mode we have actually hit.
| Metric | Type | Why it exists |
|---|---|---|
| `bridged_sends_total{outcome}` | counter | outcome ∈ replied\|completion_fallback\|timeout\|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118) |
| `bridged_send_duration_seconds` | histogram | delegated turn latency |
| `bridged_replies_total{path}` | counter | path ∈ rendezvous\|inbox — how often a reply strands (CB-307's whole reason to exist) |
| `bridged_inbox_depth{target}` | gauge | undrained replies; steady-state should be 0 |
| `bridged_push_nudges_total{outcome}` | counter | outcome ∈ delivered\|exhausted — a rising `exhausted` means the primary is not draining |
| `bridged_spawns_total{kind,outcome}` | counter | outcome ∈ ready\|timeout\|guard_rejected; per peer kind (CB-402) |
| `bridged_sessions{state}` | gauge | SPAWNING/READY/BUSY/DONE census |
| `bridged_herdr_calls_total{method,outcome}` | counter | socket health — the dependency everything rests on |
| `bridged_auth_failures_total{reason}` | counter | only meaningful once CB-501 lands; catches misconfigured workers |
---
## 5. Increment plan
Ordered so each step is independently mergeable and the risky one lands first.
1. **CB-501a — `CallerResolver` + `Role`.** Pure refactor: route today's connection identity through
the new type, `ANONYMOUS` not yet reachable (loopback-trust default preserves behaviour).
Green build, no behaviour change.
2. **CB-501b — token mode + fail-fast bind check.** Config block, bearer parsing, the
non-loopback-bind guard. This is the security-relevant commit; keep it small and reviewable.
3. **CB-505 — authz table + audit logger.** Enforce §3 on **both** entry paths (see §3.1); add the
audit appender.
4. **CB-502 — `Metrics` + `/metrics`.** Instrument the paths in §4.
5. **CB-503 — CI.** Runs `mvn -B clean install` with `-Dgroups='!contract'` so the live-herdr and
RabbitMQ contract tests are excluded; the mock-UDS suite is the CI surface, exactly as the
roadmap's testability section intends.
6. **CB-504 — launchd plist + systemd unit + herdr-socket retry.**
---
## 6. Open questions for the lead
*All four resolved 2026-07-29 — the lead confirmed D3, D4, and the D6 sink; the CI runner question
was answered from the forge itself. Kept here as the decision record.*
1. ✅ **TLS scope (D3) — confirmed.** Bearer auth + the fail-fast bind guard ship in the daemon;
TLS terminates at a reverse proxy, documented with a worked example. AMQP gets TLS via an
`amqps://` URI. No keystore handling in `bridged`.
2. ✅ **Micrometer (D4) — confirmed dropped.** Zero-dependency Prometheus text renderer, for the
reasons in D4 (pom reconciliation burden + the mandated CVE gate being un-runnable this
session). Revisit if a push-gateway or JVM-metrics requirement appears; the endpoint is the
swap seam.
3. ~~**CI runner (CB-503).**~~ ✅ **Resolved during design** — a Gitea Actions runner *is*
registered and healthy (`lms/alms-memory` has 28 completed runs; `lms/alms` runs on push and
pull_request). CB-503 targets `.gitea/workflows/ci.yml` with `runs-on: ubuntu-latest`, matching
the sibling repo's convention. Note the runner's image ships an older `default-jdk`, so the
workflow must provision **JDK 25** explicitly rather than apt-installing the default.
Contract-test exclusion needs no CI flag: the pom's `default-excludes` profile already sets
`excludedGroups=contract`, so a plain `mvn -B clean install` *is* the mock-socket surface.
4. ✅ **Audit sink — confirmed dedicated file.** Its own logback appender writing JSON lines beside
the daemon, separate from the app log, per D6. Content still never enters the record.
+1 -1
Submodule wiki updated: 8e5fd01ac9...0c896eb49b