Compare commits

...

113 Commits

Author SHA1 Message Date
Dai Ha 6f5b3a1fd6 CB-528b: provision an isolated CODEX_HOME for a Codex peer 2026-08-10 21:25:12 +02:00
Dai Ha ef49835c4f CB-528: the CodexHome seam, ahead of the adapter that uses it
Codex reads everything from CODEX_HOME — config, credentials, sessions, skills,
plugins, state. Pointing a peer at the operator's own ~/.codex would hand it the
operator's tool surface and let it write into the operator's session history:
the same failure CB-525 exists to prevent on the Claude side, in a runtime where
there is no --mcp-config to neutralize.

Extracted as an interface rather than a launcher method because provisioning is
filesystem work with its own failure modes. The common one is a missing
credential, which Codex surfaces as an opaque 401 mid-turn instead of a spawn
error — so the contract says provision() must fail loudly there. Splitting it
also lets the launcher be tested without touching a real home directory.

Lands before the adapter so the launcher and the provisioner can be built
against a fixed seam instead of against each other.
2026-08-10 20:47:26 +02:00
Dai Ha ef1e014b41 CB-527: ship the bridge as an installable Claude Code plugin
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m21s
The orchestration contract had no distributable form. Every consuming project
hand-copied a block of CLAUDE.md and hand-wrote an .mcp.json, and we keep a
script whose only job is to notice those copies drifting apart. A plugin is
versioned, installed once, and updates in place.

Ships no credentials, deliberately: every secret is referenced by environment
variable NAME and the value never enters a file, which is what makes the
artifact safe to publish. The setup skill states the two rules that are easy to
get wrong for the right-sounding reasons — the PR token must not be able to
merge (a worker opens, the primary gates), and ANTHROPIC_BASE_URL must never be
set by setup, because mounting the bridge must not move a session off
subscription.

The plugin root is plugin/, not the repo root. An installed plugin's .mcp.json
is a committed file, while this repo's root .mcp.json is local-only and
--skip-worktree; rooting the plugin at the repo would commit the primary's IDE
servers and hand them to every worker — the exact failure CB-525 exists to
prevent.

Scope is client-side setup only. herdr and bridged stay separate services with
their own lifecycles, and the skill refuses to install them rather than guess.
It also refuses to accept /healthz as proof: health reports only that the daemon
can reach herdr, and CB-521 showed it staying green while every spawn failed, so
verification ends with a real spawn.

Both manifests pass `claude plugin validate --strict`.
2026-08-10 19:55:17 +02:00
Dai Ha e3c8393d1b CB-527: retire the ollama profile from the worker choice
The ollama backend is decommissioned, so the example config stops pointing
readers at a dead host and the guard allowlist stops carrying an entry with
no profile behind it — a stale entry there is dead permission, and that list
is the only thing keeping a worker off the primary's subscription.

The second illustrative profile survives as gx11: the example exists to show
`placement: weighted` having something to choose between, and a one-profile
example would quietly stop demonstrating that.

It also moves the CB-523 auto-compact override onto the surviving profile.
That guard had been attached to `ollama` alone, so retiring the profile would
have removed the fleet's only protection against the failure it was written
for — a worker whose prompt is rejected before auto-compact ever fires. The
window belongs on every profile, not on whichever one happened to hit it.
2026-08-10 19:55:17 +02:00
kevin 379e03f9d0 CB-308: fold in the adversarial review; bump wiki to 4320c1c
CI / build (push) Successful in 59s
CI / contract (push) Successful in 1m16s
Four new resolved decisions (§7.7-7.10): turn state split by where the signals
are, with ABANDONED explicitly belt-and-braces over the waiter timeout; the
dual ack model with spawn idempotence by construction (gid stored IN the herdr
pane — the load-bearing detail of the no-ledger position — plus an in-flight
reservation for redelivery during a slow spawn); enforced publish semantics
(confirms + mandatory on a separate channel, return-before-confirm caveat);
queue lifecycle = session lifecycle with .v2 names for the redeclare hazard.

§8 reworked: global id scheme resolved and moved up; control authorization
sharpened into the hard gate on U4 (per-host allowlist beside the peer keys);
key distribution/rotation added. New §9: implementation order, each step
verifiable single-host, U4 gated, U8 last.

Wiki pointer bumped to 4320c1c (chapter 10 same-pass changes).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 22:32:00 +07:00
kevin e13921aa8a CB-308: resolve the cross-host design review; bump wiki to 36bb865
CI / build (push) Successful in 54s
CI / contract (push) Successful in 1m5s
Six decisions recorded in §7, replacing the matching open questions: signed
messages (identity from the key, extending identity-from-connection across the
broker), target-host-owned profiles advertised via presence, repo provisioning
by pinned forge clone, live-only asks with TTL + TOO_LATE notice, spawn-id
dedup on the target, and broker-outage semantics (local unaffected, remote
fails fast, gateway stays soft-state). §8 keeps what is genuinely still open,
with control *authorization* now separated from the resolved *authenticity*.

The broker-level half (U8 broadcast, exclusive consumers, inbox caps, TLS,
schema version, trace id) lands in wiki chapter 10 §10 — pointer bumped
(also picks up 1710a77, chapter 11 Features).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 21:54:30 +07:00
kevin 8a549d8610 Release 1.0.0 — one leader, one host, complete
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m36s
Bump bridged to 1.0.0 and add the release notes: the single-leader,
single-host scope is closed — gateway, lifecycle, two-way delivery,
pluggable peers, auth/authz/audit, supervision, CI. Cross-host
federation (CB-308) is the next major line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 20:58:06 +07:00
Dai Ha 84081b2bd8 CB-526: make a shipped capability undocumentable-by-accident
CI / contract (push) Successful in 49s
CI / build (push) Successful in 1m23s
CLAUDE.md already carries a mandatory before-done checklist ("the prompt is part
of the product"). It covered the instruction surface but not the operator-facing
one, and the result was measurable: CB-506 through CB-525 shipped without a
single wiki mention, while the Roadmap went on claiming Stage 5 was finished.

Discipline is what already failed, so this rides the existing gate rather than
adding a new habit to remember: one more row, firing when a change touches
anything an operator can use, configure, or observe. The row names where the
other two kinds of change go too (contracts to Implementation, coverage to the
Roadmap), so "nothing to document" is a decision the table makes rather than a
default you fall into.

Addendum-only — the canonical block is untouched and still byte-identical to the
wiki template (verified).
2026-08-09 19:38:34 +02:00
Dai Ha defe3365c4 CB-519: re-point pane assertions at the protocol-19 coordinate
CI / build (push) Successful in 1m22s
CI / contract (push) Successful in 2m32s
Rebase integration only, no behaviour change. CB-519's tests named the pane the
pre-protocol-19 fake produced (w9:pW_n); upstream's herdr 0.8.0 port creates the
pane through tab.create and starts the agent into it, so the fake now reports
w9:pRoot_n. Five assertions were therefore counting closes of a pane that never
existed and reading 0.

mvn clean install: Tests run: 399, Failures: 0, Errors: 0 — BUILD SUCCESS
2026-08-09 05:59:14 +02:00
Dai Ha 4e6201ecd1 CB-525: isolate a worker's tool surface to what its launcher mounts
A worker in a provisioned worktree was inheriting the primary's MCP servers by
two independent routes: the repo commits a .mcp.json declaring the IDE servers,
so a fresh checkout mounts them, and the default parity overlay then copied the
primary's own copy over the top.

Those servers are bound to the primary's IntelliJ project, so every path they
hand back points into the primary's checkout. A CB-523 worker made all 59 of its
edits there while running `mvn -f bridged/pom.xml` against its worktree — every
build it ran was of code that did not contain its changes, and it passed. The
worker's own `ls` of the file it had "edited" returned "No such file".

GitWorktrees now neutralizes .mcp.json at provisioning: an explicitly empty
server map, --skip-worktree'd when tracked so it never reads as pending work a
worker might commit. Unconditional, because the overlay was only half the leak.
The bridge itself is unaffected — it reaches a worker through the launcher's
--mcp-config flag, not the project file, so bridge_reply still works.

- BridgedConfig: .mcp.json out of the default parity overlay
- GitWorktrees: isolateToolSurface() on add(), with the rationale in javadoc
- GitWorktreesTest: 4 real-git acceptance tests (2 fail if the call is removed)
- implementer skill: work from $PWD, and quote a green unpiped `mvn clean
  install` from the worktree as the acceptance criterion

mvn clean install: Tests run: 392, Failures: 0, Errors: 0 — BUILD SUCCESS
2026-08-09 05:57:30 +02:00
Dai Ha ccf50f950e CB-524: make worker placement order deterministic across JVM runs
The weighted policy breaks an exact-weight tie on candidate list order
(WeightedRoundRobinPolicy picks the first candidate with a strictly greater
score), and that list comes from CompositePeerLauncher.candidates(), which
iterates profileConfigs. Both that map and BridgedConfig.workerProfiles() were
built with Map.copyOf, whose iteration order is salted per JVM run — so the
"in definition order" contract candidates() documents was not held.

Two consequences. In production, a config with equal weights (ollama 0.5 /
gx10 0.5) placed its first worker on a profile chosen at random on every daemon
restart. In the suite, CompositePeerLauncherTest.failoverRetriesNextCandidate-
WhenProfileIsUnreachable failed roughly one run in four, because whether
profile "a" was tried first depended on the salt.

Preserve definition order at every layer: unmodifiable LinkedHashMap for
workerProfiles(), profileConfigs, and byProfile (which also feeds the
user-visible bridge_profiles listing). The tests build profile maps with an
ordered helper rather than Map.of, which is salted for the same reason.

Guarded by a pair of tests declaring the same two profiles in opposite order
and asserting opposite first attempts, so any order-scrambling implementation
must fail one of them. Verified by mutation: reverting profileConfigs to
Map.copyOf fails 8/8 runs (6 caught by the original test, 2 only by the new
reversed-order one); with the fix, 10/10 fresh JVMs pass, 388 tests green.
2026-08-09 05:57:30 +02:00
Dai Ha a7f0211e2f CB-518: weighted placement policy for worker spawns 2026-08-09 05:57:30 +02:00
Dai Ha 049ce4828c CB-520: split ReplyInbox into explicit own/release and publish halves 2026-08-09 05:57:12 +02:00
Dai Ha c1173346ef CB-521: make the AMQP contract test runnable locally and in CI 2026-08-09 05:56:45 +02:00
Dai Ha 7ace184fe6 CB-519: make PeerHandle.id() a host-unique opaque UUID, decoupled from the herdr pane id 2026-08-09 05:56:45 +02:00
Dai Ha 54b314ace5 CB-522: let the primary run inside a herdr pane
Caller identity resolved any loopback PID that mapped to a herdr pane as a
WORKER, and PaneLocator scans every pane -- not just bridged-spawned ones. A
primary running inside a herdr pane therefore classified itself as a worker and
was refused SPAWN/SEND/STOP, i.e. every orchestration verb it exists to call.

The failure is self-locking: PrimaryRegistry only learns the primary's terminal
from bridge_send/bridge_spawn, the exact calls being refused, so the learned
value can never bootstrap. Only an operator-set pin breaks the cycle.

CallerResolver now consults primary.terminal from config *before* the pane
lookup. Deliberately the pinned value only, never the learned one -- the learned
terminal is populated by the callers this method is itself classifying, so
trusting it would be circular. Config is operator input, never network input, so
this widens no attack surface; bridge_whoami and the authz gate still share one
resolution.

Fixing that exposed a second, older bug. BridgeMcp's context extractor forwards
the caller's terminal into markPresent on every MCP call, documented as "no-op
for the primary (null terminal)". WorkerPresence.markPresent honours that, but
PresenceBridge overrides it and forwards the same null into SessionManager.
onReady -> transitionByTerminal -> findByTerminal, which called
terminalId.equals(...) unguarded. It only reached the scan once the registry was
non-empty, so the primary's first spawn succeeded and every later call NPE'd
with an HTTP 500 -- and it would have fired for ANY primary not living in a
herdr pane, pinned or not.

findByTerminal is now total. That covers onReady, onDelivered, onTurnComplete
and onTurnFailed at once; a null id could never match a registered session
anyway, so "no match" is the honest answer rather than taking down an unrelated
tool call.

Also drops two dead pass-throughs on CallerResolver (cwdForPid, tokenMode) that
IDE inspections flagged -- callers use ConnectionIdentity and BridgedConfig.Auth
directly.

The example config now states that primary.terminal is REQUIRED, not just a
push-loop optimisation, when the primary shares a herdr pane.

mvn clean install: 360 tests, 0 failures. Verified live: daemon restarted on
this jar, bridge_whoami reports primary, and four concurrent worktree spawns --
the exact shape that NPE'd -- now all succeed.
2026-08-09 05:55:12 +02:00
kevin 224b344445 CB-522: resolve the pinned primary.terminal pane as the primary, not a worker
CI / build (push) Successful in 2m54s
A primary running INSIDE a herdr pane was resolved as a worker by the
pane-match rule and refused every orchestration tool — the exact lockout
bridge_whoami surfaced on this deployment. The CB-307 primary.terminal pin
always claimed to replace connection-derived identity but only fed the push
loop; it now short-circuits CallerResolver ahead of the pane→worker rule
(the pane mapping is as unforgeable as a worker's, so no credential needed,
even in token mode). bridged.example.yaml documents the block.

Also guard the presence bridge against the primary's null terminal: the MCP
context extractor marks presence on every request, and the first genuine
primary contact NPEd into the SPAWNING→READY transition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
2026-08-08 21:52:38 +07:00
kevin 0b28b4cb0f CB-521: port the herdr adapter to protocol 19 (herdr 0.8.0)
herdr 0.8.0 redesigned the agent API out from under the daemon: agent.start
now launches a supported kind INTO an existing pane, env/cwd move to pane
creation (tab.create / pane.split — the subscription-boundary seam now),
agent.send is replaced by agent.prompt (self-submitting) plus agent.send_keys
for the Enter nudge, and terminal ids are no longer valid agent.* targets.

- AgentControl: start(name, kind, args, paneId); prompt/send_keys delivery;
  cached terminal→pane target translation (invalidated on agent_not_found).
- WorkspaceControl: tab.create carries cwd+env; pane.split for legacy placement.
- HerdrPeerLauncher: the seed pane IS the worker pane (no drop step); retry
  agent.start while the seed shell boots (agent_pane_busy).
- FakeHerdr and the test suite model protocol 19 (unique seed panes, required
  kind/pane_id, prompt-based delivery); contract tests probe the seed shell
  instead of arbitrary-command agents, which protocol 19 removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
2026-08-08 21:52:25 +07:00
kevin 83129e165c Merge CB-518: state the primary's orchestration as an explicit, ordered flow
CI / build (push) Successful in 1m19s
Turns the primary's half of the bridge charter from a bullet list of
policies into a numbered 0-8 procedure, and splits delegated review out
of the merge step it used to sit beside. Wiki template kept byte-identical
by splicing; pointer bumped to 0c896eb in the same commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw1FosEt3Noix5GG9wJ2r
2026-08-04 22:37:17 +07:00
kevin 871b595954 CB-518: state the primary's orchestration as an explicit, ordered flow
CI / build (pull_request) Successful in 1m23s
CB-517 moved orchestration policy into CLAUDE.md but left it as a bullet
list, so the procedure was implicit: the order of operations had to be
reconstructed from a parallelisation bullet, and nothing said when to
review or when to tear down. A policy you have to reassemble on each
task is one you will reassemble differently each task. Restate the
primary's half as a numbered 0-8 flow — role check, split, gate, spawn
all, send all, collect, verify, review, adjudicate — so that following
it is checkable against the tool calls rather than a matter of recall.

Two steps carry the load. Spawn and send are separate on purpose:
folding them into one loop is what silently serialises work that was
meant to fan out. And review is now its own step ahead of the merge
rather than a clause inside it, because the two have opposite owners —
reviewers fan out over the diff (never the implementer of the scope
they review, and briefed from the diff rather than the author's
rationale, which carries the same blind spot), while adjudication, the
merge and teardown stay with the primary. Merging on a reviewer's word
is delegating the gate by proxy, so the step says so outright.

Nothing is dropped. The six bullets that trailed the tool table are
relocated into the step that owns each — profile explicitness into
spawn, playbook naming and self-containment into send, claim
verification into its own step, the ~60s blocking-send cap into a note
beneath the flow — and the table stays as the intent→tool lookup.

The wiki pointer moves with it. The block is canonical only if its
template matches byte for byte, so the template was produced by
splicing the block out of CLAUDE.md rather than by editing it in
parallel, and the sync check the repo documents passes. Bumping the
pointer in the same commit keeps charter and template versioned
together, as CB-517 did.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw1FosEt3Noix5GG9wJ2r
2026-08-04 22:32:19 +07:00
Dai Ha b67b1585c2 CB-517: deploy LavinMQ as a pinned, durable, self-restarting broker
CI / build (push) Successful in 1m23s
The CB-307 durable ReplyInbox needs an AMQP broker, but the one behind it
was run ad hoc and had simply vanished from the host — which takes the
whole daemon with it, since AmqpReplyInbox.open throws and Bridged.java:187
does not guard it. A missing broker is a hard startup failure, not a
degraded mode, so 'how the broker runs' is part of the system, not a local
detail.

Pinned to 2.9.1 (:latest would move the broker under a running daemon),
data on a named volume (held-but-unacked replies are the entire point of
Stage 2 — a plain 'compose down' would discard exactly what durability
protects), and restart: unless-stopped so it comes back after a reboot
instead of disappearing again.

Ports are bound to 127.0.0.1 deliberately: LavinMQ ships a default
guest/guest account, which is only acceptable while nothing off-host can
reach it.

Verified by driving the production AmqpReplyInbox against this deployment
(publish/peek/dedup/FIFO/ack, then reconnect): 8/8 including redelivery of
the unacked message. That pairing had never been exercised — the
@Tag("contract") test runs against a RabbitMQ container, and is excluded
from the default build, so mvn clean install covers the broker path zero
times.
2026-08-04 16:17:21 +02:00
Dai Ha 979b2b5632 CB-517: add bridge_whoami and make the bridge prompt a portable charter
CI / build (push) Successful in 1m25s
The communication rules lived only in two opt-in skills, so nothing
always-on told the primary how to orchestrate and nothing guaranteed a
worker loaded its playbook. Move protocol and policy into CLAUDE.md,
which a worker inherits for free (its worktree is a checkout of this
repo), and leave the skills as pure per-job procedure.

bridge_whoami closes the load-bearing gap: every tool already consumed
the caller identity ConnectionIdentity resolves from the connection, but
none reported it, so an agent had to infer its own role from side
channels the daemon does not control. Guessing fails asymmetrically — a
primary acting as a worker is refused by the authz gate and learns at
once, while a worker acting as the primary ends its turn without
bridge_reply and the sender silently receives nothing. The tool reuses
the same Principal the gate is built on, so the two cannot disagree; the
primary gets role only (handing it a sessionId it does not own would
invite the forged reply Authz refuses), and a worker missing from the
registry still gets role + sessionId rather than 'unknown'.

The CLAUDE.md block is written to be copied as-is into any project that
mounts the bridge: repo-local details (Authz paths, the .mcp.json/wiki
exclusions, the skill names) moved below it into a project addendum, and
every role-inference fallback is stated one-way — the mount-name signal
only holds for mcp__bridge__* (the launcher fixes it), not for the
primary's mount, which each project names itself. The wiki carries the
block verbatim as the template, with a sync check.

Because this repo IS the bridge, that block is shipped surface, not
documentation: the addendum adds a mandatory checklist mapping each part
of the code to the part of the prompt it can invalidate.

Also: delegate-by-default policy for the primary — the test is not 'could
I do this faster myself' but 'can I write a brief good enough for a
worker'.

mvn clean install: 356 tests green (353 + 3 for whoami); ide_diagnostics
clean on both changed files.
2026-08-04 16:01:16 +02:00
kevin cf4ad186ab CB-516: fail a delegation when its worker session is released
CI / build (push) Successful in 1m30s
Tearing a worker down left the send that was waiting on it stranded. The
rendezvous waiter stayed open, so a blocking bridge_send kept blocking and an
async one kept reporting PENDING until ASYNC_TIMEOUT_MS — thirty minutes —
even though the worker provably no longer existed and the delegation could
never complete.

Observed repeatedly this session: a task sitting at
{"phase":"pending","detail":"worker unknown"} long after I had personally
deleted the pane. The maddening part is that poll() already HAD the evidence —
it calls liveStatus() to build that detail string, gets back "unknown", and
reports PENDING anyway. It also never reached /metrics under any outcome, so a
stalled delegation was invisible to both the task view and the dashboard. The
only way I ever diagnosed one was reading the worker's pane over the herdr
socket by hand.

Fixed at the choke point rather than by guessing from status strings:
SessionManager.release() is the single path every teardown funnels through
(REST stop, MCP stop, idle-TTL reaper, recycle, shutdown drain), so it now
notifies a release listener with the terminalId, and Bridged.main wires that to
MessageService.abandon(). Abandon fails the open waiter with a real reason.

Deliberately NOT done by inferring "gone" from liveStatus(): that method
collapses a vanished pane, a wedged worker and a herdr hiccup into the same
"unknown" string, so acting on it would fail live delegations during a
transient blip. An explicit lifecycle signal cannot be ambiguous.

Notified before launcher.stop() so a blocked caller fails fast, and wrapped so
a listener failure can never prevent the teardown it is reacting to.

Resolving as a failure rather than letting it time out also means the outcome
is counted — a torn-down delegation now shows up as
sends_total{outcome="failed"} instead of nothing at all.

353 tests (was 346): abandon fails a waiting send / is a no-op with no waiter /
never clobbers a send the worker already answered, an abandoned async task
polls FAILED rather than PENDING, release notifies with the right terminal,
releasing an unknown pane notifies nobody, and a throwing listener does not
block teardown.

Verified live on the running daemon, reproducing the original scenario:
  async send        -> {"phase":"pending","detail":"worker working"}
  DELETE the worker -> 204
  poll              -> {"phase":"failed","detail":"the worker session was
                        released before it replied"}
  /metrics          -> bridged_sends_total{outcome="failed"} 1
Previously that poll returned PENDING for thirty minutes and the counter stayed
empty.

NOT addressed here, and worth deciding separately: a worker that simply never
replies (rather than being released) still rides out the full 30-minute
ASYNC_TIMEOUT_MS. That is a policy question about long-running delegations, not
a correctness bug.
2026-08-01 23:57:45 +07:00
kevin 5100f215cf CB-515: regression-protect the turn-attribution guards
CI / build (push) Successful in 1m33s
Five tests pinning the invariants that decide WHICH turn a reply belongs to.
These protect against a silent correctness bug — a reply attributed to the
wrong turn — not against a crash, which is why they were worth picking over
higher-percentage coverage gaps.

Chosen by blast radius, not by uncovered-line count. Both guards are compound
conditions with a side that never executed, i.e. exactly the shape where a
clause can be deleted as "redundant" and every existing test still passes.

CompletionResolver:
- The CB-115 misattribution guard suppresses a completion when the scrape is
  byte-identical to the pane at delivery. Its !scrapeFailed clause was
  unexercised: delete it and a FAILED read is misread as "no output change",
  so the send is suppressed and hangs to the caller's timeout instead of
  resolving. The new test sets the baseline to "" so the empty tail from a
  failed read would byte-match and wrongly suppress — built to die precisely
  when that clause dies.
- The fail() guard leaves an already-resolved waiter alone. The new test also
  asserts agent.read is never called, so the worker is not scraped for a send
  nobody is waiting on.

Rendezvous: a second resolution of an already-completed waiter returns false
and does not overwrite the first value, for both resolveCompletion and
resolveFailure.

Verified by sabotage, one guard at a time: removing !scrapeFailed reds
resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent; removing the
isDone() clause reds failLeavesAnAlreadyResolvedWaiterUntouchedAndSkipsTheScrape.
(The first attempt at the second sabotage reported a false pass — the patch hit
an identically-worded guard earlier in the file. Line-targeted and re-run.)

346 tests, was 341. Worker-implemented on the local-vLLM profile; it noticed
three of the eight cases I asked for already existed and said so with names
rather than duplicating them.

Also of note: the first delegation of this ticket wedged the worker — the pane
showed a zsh parse error and it went idle with an untouched worktree, task stuck
pending. The retry differed only in phrasing the same requirements as prose
instead of quoting Java boolean expressions. Filed as a bridge robustness
concern: injected content shares a channel with control, and a wedged turn is
invisible in both the task view and /metrics.
2026-08-01 23:48:05 +07:00
kevin f129e9b7cd Merge CB-514: MessageService coverage for timeout, answer, poll and lock edges
CI / build (push) Successful in 1m29s
Six tests on the delivery core, the most load-bearing class in the project,
which sat at 74% with its hardest paths unexercised: answer() (the CB-205
ask-answer resolution) had 11 of 23 lines uncovered, send()'s timeout branches
7 of 21, poll() 7 of 18, and tryLock 3 of 4.

Covered: TIMED_OUT_QUEUED vs TIMED_OUT_WORKING (undelivered vs delivered-but-
silent), an answered worker that never sends its follow-up reply, an unknown
ticket, a completed async ticket reporting its reply and replySource, and a
second concurrent send to the same session returning BUSY rather than hanging.

Worker-implemented on the local-vLLM profile, self-verified: it reported
'Tests run: 341, Failures: 0' and an independent run of its branch agrees
exactly. Additive only — 101 lines in one test file, no main/ source touched.

Quality is good on its own terms, not just green: the async-completion test
polls to a deadline instead of sleeping and hoping, the contention test
releases the blocked send so the test thread is not left pinned, and every
assertion is on a specific Outcome rather than 'nothing threw' — the failure
mode an earlier worker produced in CB-510.

Branch pushed by the worker; PR left to the primary since GITEA_TOKEN is not
granted to that profile by design.
2026-08-01 23:26:54 +07:00
kevin 4aed45de19 CB-514: add MessageService coverage for timeout, answer, poll, and lock edges 2026-08-01 23:24:54 +07:00
kevin 1e6daa5c73 CB-513: test the MCP-side authorization gate (BridgeMcp 27.4% -> 57.4%)
CI / build (push) Successful in 1m13s
CB-505 claimed authorization is "enforced on both entry paths". It is — but
only REST was ever tested. Coverage showed BridgeMcp.deny(), principal(),
callerTerminal(), worktreeRequest() and every tool-registration lambda at ZERO
executed lines: no test had ever constructed a BridgeMcp, because the existing
BridgeMcpTest calls only the static handler methods. So the MCP half of the
security control had ten REST tests' worth of nothing behind it.

An unexercised security control is a claim, not a control.

Made testable by separating policy from plumbing rather than by reaching for a
mocking library the project does not use:
- denyFor(Principal, Action, target) is the decision — testable directly.
- deny(exchange, ...) shrinks to pulling the caller out of the SDK exchange.
- principalFrom(role, terminal, pid) extracts identity reconstruction from
  McpSyncServerExchange, an SDK type with no fake available.

Moved the `authz == null` enforcement switch OUT of the exchange-facing wrapper
and INTO denyFor. Found by a failing test: as written, any future tool calling
denyFor directly would have silently skipped the gate. The switch now lives with
the decision it governs.

New BridgeMcpAuthzTest constructs a real BridgeMcp — which is why coverage moved
so far, since that also runs the constructor and all the tool wiring — and pins
the table on this path: primary orchestrates, worker cannot; worker replies only
as itself; the primary cannot forge a worker reply; anonymous gets nothing; and
401-shaped vs 403-shaped refusals are counted apart.

Verified as real controls, not decoration: with the gate forced open, 5 of the 9
fail. 335 tests (was 326).
2026-08-01 23:11:56 +07:00
kevin 4c015d76b7 Merge CB-512: wire bridged_push_nudges_total (worker-implemented, self-verified)
CI / build (push) Successful in 1m30s
Fixes one of the three counters declared in BridgedMetrics but never
incremented, so bridged_push_nudges_total{outcome=delivered|exhausted} now
actually appears on /metrics. An absent series reads as 'no push failures ever'
rather than 'not measured', which is the misleading case.

Implemented end to end by an opencode worker on the local-vLLM profile in an
isolated worktree, and this is the first delegation where the worker verified
its own work: it ran mvn, hit a real test failure, iterated, and reported
'Tests run: 326, Failures: 0' — which matches an independent run of its branch
exactly. Every earlier delegation reported results it had no way to check,
because CB-511 had not yet given workers a PATH with a toolchain.

It also committed and pushed its own branch unprompted, and was straight about
the one thing it could not do: opening the PR, since GITEA_HOST/GITEA_TOKEN are
not granted to that profile ('URL rejected: No host part'). That grant is opt-in
per profile by design, so the merge is primary-side as intended.

Diff reviewed and correct on every constraint, including the subtle ones:
delivered counted inside the try after a successful send (not the catch),
exhausted only on the reminder-cap branch and not the other two STOP paths, and
a single Metrics instance moved above pushLoop and shared with MessageService.
2026-08-01 23:01:07 +07:00
kevin 37a11cd168 CB-512: wire bridged_push_nudges_total metric increments 2026-08-01 22:56:32 +07:00
kevin 22ad24db6c CB-511: give workers a toolchain — propagate the daemon PATH, add profile env:
CI / build (push) Successful in 1m17s
Workers could not run `mvn` or `java`. Every delegated task that asked for a
build came back "mvn is not on PATH", and the worker was right.

Root cause: HerdrPeerLauncher seeded the worker environment with an EMPTY map,
so bridged passed only the vars it explicitly set (OPENCODE_CONFIG, GITEA_TOKEN,
ANTHROPIC_*) and never PATH. herdr merges that map into its own process env, so
a worker inherited whatever PATH the herdr SERVER was started with. On this host
that server (pid 79870, PPID 1) had been up since Jul 4 with a PATH containing
neither the JDK nor Maven. Confirmed on a live worker: its PATH was byte-identical
to herdr's, and the only var bridged had contributed was OPENCODE_CONFIG.

The failure was invisible and non-deterministic: the fleet's capabilities
depended on how a long-lived daemon happened to be launched weeks earlier. There
are three herdr processes on this box with three different PATHs; the one owning
the socket is the one without a toolchain. bridged itself HAD Maven on PATH the
whole time — it just never passed it on.

It also quietly contradicted the project's own principle that "a worker is a
full peer of the primary", and the implementer skill's instruction to build,
commit and open a PR. Every delegation so far has depended on the primary
running the build gate.

Fix: baseEnv(cfg) seeds each worker with the daemon's own PATH, then applies the
profile's new optional env: map. Adapter-specific vars are layered on top and
therefore win — that ordering is load-bearing, not incidental: it stops an env:
entry from overwriting ANTHROPIC_BASE_URL and slipping past SubscriptionGuard,
which is checked against the profile's baseUrl alone. Pinned by a test.

Because the default is now the daemon's PATH, both supervision units set PATH
explicitly — launchd and systemd do not source a login shell, so under CB-504
the daemon (and every worker) would otherwise get a bare /usr/bin:/bin and this
bug would silently return in production.

324 tests (was 321): daemon-PATH propagation, profile env: passthrough including
an explicit PATH override, and the guard-bypass ordering.

Verified live: daemon restarted, worker spawned, and asked to run the tools —
"Apache Maven 3.9.16", "java version 25.0.2". Previously both were absent.
2026-08-01 22:34:30 +07:00
kevin e94c1b8841 CB-510: SessionReaper wrapper tests (0% -> 86.7%)
SessionReaper had no tests at all. Its TTL *policy* was already well covered
(SessionManager.reapIdle, 6 cases in SessionManagerTest); what was untested was
the thread wrapper around it — idempotent start/stop and whether the loop
actually runs and actually stops.

Observed through an injected clock rather than by sleeping and hoping: reapIdle
reads nowNanos exactly once per call, so the tick count IS the iteration count.
Waits are bounded polls, not fixed sleeps, and nothing asserts an exact
timing-derived number — flaky counts would be worse than no test.

321 tests (was 318); line coverage 66.9% -> 67.9%.

Drafted by an opencode worker on the new local-vLLM profile (branch
worker/cb-510-session-reaper-test-cd1793-1). Its structure and setup were good
and it was honest that it could not run mvn. But its third test asserted
NOTHING — it started the reaper, slept, stopped it, and relied on "no throw",
with a comment claiming that proved the loop had run. It did not: verified by
sabotage, all three of its tests passed against a start() replaced with an
immediate return.

Rewritten so the assertions can fail for the right reason. Same sabotage now
fails 2 of 3 (the third only pins stop()-before-start(), where "does not throw"
genuinely is the contract). Uncomfortably on the nose given this task began as
a hunt for tests that do not mean anything.
2026-08-01 21:42:06 +07:00
kevin cc0ec65714 CB-509: add JaCoCo coverage reporting
Build-time tooling only — never a compile or runtime dependency, so it adds
nothing to the shipped jar and no new transitive surface to the artifact.
(Noting per CLAUDE.md that the pom CVE gate could not be run: no JetBrains MCP
server is connected this session.)

Report at target/site/jacoco/index.html, machine-readable at jacoco.csv.

Deliberately NO check rule or threshold. A coverage gate rewards writing tests
that merely execute lines, which is the exact failure mode this codebase has
already been bitten by — CB-507 shipped a null-argument NPE with 311 green
tests because FakeWorktrees.repoRoot records its argument instead of shelling
out, so the broken line was covered and still wrong. Coverage is a map of where
to look, not a target to hit.

Baseline: 66.9% line, 60.6% branch, 76.4% method.
2026-08-01 21:34:59 +07:00
kevin d67d30c58a CB-508: let an opencode profile pin its own OpenAI-compatible endpoint
CI / build (push) Successful in 1m57s
Points an opencode worker at a local vLLM (or llama.cpp / LM Studio / TGI)
instead of opencode's own gateway. opencode has no ANTHROPIC_BASE_URL seam, so
this could not be a config-only change: setting baseUrl on a kind: opencode
profile now makes the launcher emit a custom `provider` block into the
generated opencode.json, using @ai-sdk/openai-compatible.

The provider id comes from the provider half of the model: selector, so one
field drives both the generated declaration and the -m flag and the two cannot
drift apart. A bare model name with a baseUrl set is rejected at spawn with a
message saying how to fix it — silently falling back to the default gateway
would leave a worker talking to the wrong LLM while looking perfectly healthy.

A bare host:port gets /v1 appended (where these servers mount the API); a URL
that already carries a path is used verbatim. tokenEnv, when set, becomes the
provider apiKey; local servers generally ignore it but the AI SDK requires a
non-empty value, so a placeholder is used otherwise.

Two supporting changes:
- writeConfig previously ran only when a bridge MCP url was set. A pinned
  endpoint needs the config file too, so it now runs when either applies, and
  the mcp/instructions half is emitted conditionally.
- The config is now built with Jackson instead of string concatenation. The
  provider block is nested and interpolates operator-supplied values (URL,
  model id, api key), so escaping has to be real rather than a hand-rolled
  two-character replace.

No guard entry is required even with baseUrl set. SubscriptionGuard exists to
stop a worker borrowing the primary's Anthropic subscription, and an opencode
process has no Anthropic credential path at all — the asymmetry with the Claude
adapter reusing the same field is deliberate and documented at the call site.

Also fixes a brittle assertion in the existing MCP-mount test, which matched
the substring "\"type\": \"remote\"" and broke on Jackson's spacing. It now
parses the generated JSON and asserts on structure; whitespace is the
formatter's business, not the contract's.

318 tests (was 311): 5 new covering provider generation, /v1 normalisation,
path-preserving URLs, the missing-prefix rejection, MCP+provider coexistence,
and that no baseUrl still means no provider block.

Verified live end to end: daemon restarted on this build, worker spawned on the
opencode-local profile, generated config carries baseURL
http://127.0.0.1:8000/v1, and a blocking bridge_send returned
{"reply":"LOCAL-OK","replySource":"reply"} — a structured reply, not the
completion fallback. The worker pane reports
"Build · deepseek-v4-flash local-vllm (bridged)", confirming traffic reached
the local server rather than silently falling back.
2026-08-01 20:03:00 +07:00
kevin d75ee1cca5 CB-507: regression tests for worktree cwd resolution
Two cases in WorktreeSessionManagerTest, covering the gap that let the NPE ship
(313 tests, was 311).

1. worktreeAcquireWithNoRequestedOrCallerCwdStillResolvesANonNullRepoRoot —
   the null/null case a plain REST spawn produces.
2. worktreeAcquireHonoursTheProfileConfiguredCwd — the quieter second bug on
   the same line, where a pinned per-profile cwd: was ignored entirely.

Both assert on the cwd RECORDED by FakeWorktrees rather than expecting a throw.
That is deliberate: FakeWorktrees.repoRoot only records its argument and returns
a canned root, so a null passes through the fake harmlessly while the real
GitWorktrees runs `git -C null` and NPEs. The fake being more permissive than
the real seam is exactly why 311 tests stayed green over a broken feature —
asserting "an exception was raised" would be untestable here and would give
false confidence.

Verified as genuine regressions, not tautologies: with the pre-CB-507
expression restored both fail, with the messages they were written to give
(expected: not <null>, and expected </pinned/dir> but was <null>). Restored
after.

Drafted by an opencode-free worker over the bridge in an isolated worktree
(branch worker/cb-507-regression-test-11591f-4). Its test 1 was correct as
written. Test 2 was wrong and went red: it passed "/pinned/dir" as the 4th
constructor argument, which is configDir, not cwd (the 11th, after mcpUrl), so
cwd stayed null and the chain fell through to the daemon cwd. Corrected on
integration, along with removing two unused locals and adding the rationale
comments.
2026-08-01 19:50:31 +07:00
kevin 6804676a96 CB-507: fix NPE on a worktree spawn with no cwd (HTTP 500 over REST)
POST /workers?worktree=true returned HTTP 500 with a NullPointerException out
of ProcessBuilder.start(): acquireWithWorktree resolved the repo root from
firstNonBlank(requestedCwd, callerCwd), and a plain REST spawn supplies
neither (BridgedApp hardcodes callerCwd=null, "no MCP caller over REST"). Both
null yielded null, putting `git -C null rev-parse --show-toplevel` on the
command line.

Now resolved through launcher.effectiveCwd, the CB-112 chain used everywhere
else (requested -> profile cwd -> caller -> daemon cwd -> "."), which is
documented never to return null. The non-worktree path in this same class
already went through it; only the worktree branch was missed.

Also fixes a second latent bug in the same line: firstNonBlank never consulted
the profile's configured cwd:, so a worktree spawn silently ignored a pinned
per-profile working directory. effectiveCwd honours it.

Removes firstNonBlank, now dead (this was its only call site) — javac ignores
an unused private method but IDE inspections flag it, and CLAUDE.md requires a
clean bill.

Why 311 tests missed it: the null/null case only arises over REST, and
WorktreeSessionManagerTest always passes an explicit cwd. Over MCP callerCwd is
populated from the caller PID, so the feature worked there. This is the third
REST-vs-MCP divergence found this month, after CB-505's path-trusted session id.

The one-line change was implemented by an opencode-free worker over the bridge
in an isolated worktree (branch worker/cb-507-worktree-cwd-npe-3e9c3b-3); the
dead-helper cleanup and the explanatory comment were added on integration.
A regression test is still outstanding and is being delegated separately.
2026-08-01 19:43:08 +07:00
kevin 3e5d742ac7 CB-506: keep the test suite out of the production audit log
main/resources/logback.xml routes the `audit` logger to a RollingFileAppender
at logs/audit.log — the CB-505 security trail. AuditLogTest and
BridgedAppAuthTest exercise that same logger, so every `mvn test` appended
fabricated records to the production file.

They are byte-identical to genuine ones: runs of denied/forbidden
SPAWN/STOP/SEND from worker:term_a, which read exactly like an intrusion
attempt. logs/audit.2026-07-29.0.log is 38 fabricated records out of 76 — half
that day's security log is test fixtures, and nothing distinguishes them.

Fix is one new file, src/test/resources/logback-test.xml: logback prefers it on
the test classpath, so tests get a console-only config with no file appender
and main/resources/logback.xml is untouched. The `audit` logger stays ENABLED
(INFO, additivity=false) because AuditLogTest attaches its own ListAppender and
asserts on emitted records — setting it OFF would have silently gutted those
assertions.

Verified: 311 tests green, and logs/audit.log line count is identical before
and after a full `mvn clean install` (zero new records).

Implemented by an opencode-free worker over the bridge in an isolated worktree
(branch worker/cb-506-audit-test-isolation-e4aa9c-2); it correctly reported it
could not run mvn rather than fabricating a result, so the build gate and the
before/after audit-count check were run primary-side. Header comment added on
integration.
2026-08-01 19:20:29 +07:00
kevin 6da2a71050 CI: drop upload-artifact — unsupported on this Gitea instance
CI / build (push) Successful in 1m44s
Run 2 built clean (311 tests, BUILD SUCCESS, Maven 3.6.3 on Java 25.0.4 —
JAVA_HOME from setup-java correctly beat the JRE apt pulled in) but the job
still went red on the artifact step:

  GHESNotSupportedError: @actions/artifact v2.0.0+, upload-artifact@v4+ and
  download-artifact@v4+ are not currently supported on GHES.

Gitea Actions presents as GHES, so v4 artifact upload cannot work here. The
artifact was unretrievable regardless, so replace it with a failure-only step
that cats the failing surefire .txt reports into the job log, where they are
readable. Guarded with 'exit 0' so the dump itself can never mask the real
failure.
2026-07-29 23:26:21 +07:00
kevin e32ac39faf CI: install Maven — setup-java provides the JDK only
CI / build (push) Failing after 2m1s
First CI run failed at 'Build and test' with exit code 127 (command not
found): actions/setup-java@v4 provisions a JDK but not Maven, and the runner
image has no mvn on PATH. The sibling lms/alms workflow apt-installs both;
this workflow switched to setup-java for JDK 25 (the image's default-jdk is
too old for maven.compiler.release=25) and dropped the maven install with it.

Adds an explicit Maven install with Apache Maven 3.9.16 (2bdd9fddda4b155ebf8000e807eb73fd829a51d5)
Maven home: /Users/appbuilder/Tool/apache-maven-3.9.16
Java version: 25.0.2, vendor: Oracle Corporation, runtime: /Users/appbuilder/Tool/jdk-25.0.2.jdk/Contents/Home
Default locale: en_VN, platform encoding: UTF-8
OS name: "mac os x", version: "26.5.1", arch: "aarch64", family: "mac" so the log proves which JDK
it resolved — JAVA_HOME from setup-java must win over the JRE apt drags in.
2026-07-29 23:21:55 +07:00
kevin daa243d37a example config: opencode profile uses the verified zero-credential free tier
CI / build (push) Failing after 2m21s
The commented CB-402 block suggested google/gemini-2.5-pro, which needs
credentials. Replaced with the opencode/*-free gateway models proven during the
dogfood to work with no auth at all, and noted that the free model names change
so `opencode models` is the source of truth.
2026-07-29 22:49:58 +07:00
kevin 2773ab600d CB-402: live dogfood complete — Stage B verified against opencode 1.18.5
Closes the one known-unverified item before cross-host. CB-402 merged in
ded226a with increment 5 (the §5 live checklist) deferred; it has now run.

Provider question (§7 Q1) resolved with no credentials needed: opencode's own
gateway serves free-tier models. `opencode auth list` reports 0 credentials,
yet `opencode run -m opencode/north-mini-code-free` answers. Distinct from the
primary's subscription by construction, and needs no guard entry — opencode
carries no ANTHROPIC_BASE_URL, so SubscriptionGuard never applies to it.

The schema-drift risk was the real one and it did not bite. The adapter was
designed against opencode 1.1.31; installed is 1.18.5. The generated config
still validates unchanged (type:"remote" + instructions:[path]), and
`OPENCODE_CONFIG=… opencode mcp list` reports the bridge connected. Pinned as
a verified fact for 1.18.5.

Full lifecycle through REST: spawn (201, kind-routed to OpenCodeLauncher) ->
CB-306 gate passed ~0.6s -> ready -> send -> {"replySource":"reply"} (a
STRUCTURED bridge_reply, not the CB-115 completion fallback) -> delete (204,
tolerant teardown).

Unplanned cross-validation with CB-501: the audit trail recorded the reply as
role=WORKER actor=worker:term_657c… — connection-based identity classified an
opencode process as a worker with no opencode-specific handling. The identity
model is peer-kind-agnostic, which is what CB-308 needs when the roster
stretches across hosts.

Stage 5 verified live on the same run: /workers (CB-304) answers where the
13-day-old daemon 404'd, /metrics counted the delegation
(sends_total{outcome=replied} 1, replies_total{path=rendezvous} 1,
inbox_depth 0), and the audit log captured SPAWN/SEND/REPLY with correct roles.

Adds the opencode-free dogfood profile to the local bridged.yaml (gitignored;
recorded here for reproducibility) and docs/CB-402 §8 as-built.
2026-07-29 22:49:15 +07:00
kevin 19cdf8dc9f CB-505 fix: audit lines were not valid JSON
The first cut spliced the timestamp on via a logback pattern:

    {"ts":"%d{...}",%replace(%msg){'^\{',''}%n

Logback's variable substitution chokes on the literal braces
("All tokens consumed but was expecting }"), so the encoder failed to
configure. Caught by running the real jar and noticing logback had dumped its
internal status — which it only does when something failed to parse. The build
was green throughout: nothing asserted the audit trail was machine-readable.

AuditLog now emits the complete object including its own ISO-8601 "ts", and the
appender pattern is a bare %msg. Adds AuditLogTest, which parses each emitted
line with Jackson (so a malformed record fails the build) and pins that hostile
ids cannot escape their field to forge a second record.

311 tests green; logback now configures with zero internal errors.
2026-07-29 22:32:18 +07:00
kevin 9daf1ec5ba CB-5xx: Stage 5 hardening — auth, authz+audit, metrics, CI, supervision
Closes out single-host before the cross-host work. Sequenced BEFORE CB-308
deliberately: federation's own gating concern is the trust model, and it
inherits whatever identity shape lands here.

The finding this stage is built around: bridged had exactly ONE security
control, the loopback bind. ConnectionIdentity resolves a worker from its
connection (unforgeable), but every caller that was not a recognised worker
pane fell through to being treated as the PRIMARY -- the most privileged role
on the bus. Latent today; load-bearing the moment a bind widens.

CB-501 auth:
- Role/Principal/CallerResolver: connection identity first, bearer token
  second, ANONYMOUS third. Inverts the old default so absence of identity
  means nothing, not everything.
- Worker identity is never token-gated, so enabling auth cannot lock the
  fleet out of bridge_reply.
- Constant-time token compare (MessageDigest.isEqual).
- validateAuthExposure(): a non-loopback bind under loopback-trust now
  REFUSES TO START. Makes the dangerous config unrepresentable rather than
  merely documented.
- TLS terminates at a reverse proxy by design (D3), not in the JVM.

CB-505 authz + audit, enforced on BOTH entry paths:
- The docs describe MCP as "a thin adapter over the REST core"; at code level
  it is not. BridgeMcp calls MessageService directly, and /mcp is a raw
  servlet on Jetty's context handler that never traverses Javalin's before
  filter. Enforcing only at REST would have left /mcp open.
- Load-bearing rule is own-session-only: a worker may reply/ask only as
  itself. Structurally true over MCP already; over REST the session id in the
  URL path had simply been trusted.
- Audit: JSON lines to a dedicated appender, additivity=false. Never records
  message content -- this bus carries source and prompts.

CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer
instead of the specced Micrometer, because this pom already hand-pins
jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against
skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated
dependency CVE gate could not be run (no JetBrains MCP server connected).
Instrumented at MessageService, the single funnel both surfaces share.

CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner.
Needs no contract-exclusion flag -- the pom's default-excludes profile
already sets excludedGroups=contract, so plain `mvn clean install` IS the
mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older).

CB-504 supervision: launchd agent (the real target -- this host is macOS,
there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds.
Ordering directives are advisory, so the actual fix is that startup now waits
up to 30s for the herdr socket and then serves degraded, instead of crashing
into a restart loop on a boot-order race.

Also fixes drift found while surveying:
- bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config
  binds via plain Jackson with ignoreUnknown, so uncommenting it would have
  been silently dropped and the default kept. Now camelCase, with a test that
  loads the shipped example and one that pins every documented knob's
  spelling -- no test had ever loaded that file.
- Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay,
  gitTokenEnv, gitHostEnv, configDir, primary:).
- README "Next" listed bridge_ask and session lifecycle as upcoming; both
  shipped long ago.
- docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre-
  implementation" for work already merged.

307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS.
Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check
could not be run -- no JetBrains/intellij-index MCP server is connected this
session. mvn clean install is the only gate that ran.
2026-07-29 22:29:26 +07:00
Dai Ha c9f0ca9359 wiki: bump submodule pointer to b646be1 (chapter 10 — Cross-Host Messaging & Broker Topology) 2026-07-28 16:33:51 +02:00
Dai Ha 84c8a2d2f0 CB-500 §11: resolve distributed-sandbox topology (gateway-per-host × local sandboxes)
Clarifies the open fork from §4/§9: Development A (sandbox launcher) and CB-308
(per-host federation) COMPOSE — each host runs a bridged gateway whose launcher
spawns agents into that host's LOCAL sandboxes; the broker moves messages +
presence, never keystrokes.

The forcing fact: delivery is herdr keystroke-injection into a locally-owned PTY,
so "a sandboxed agent on another host" ≡ "a sandbox spawned by that host's
gateway" (a remote container with no local herdr can't be injected into). Rules
out a central daemon reaching remote PTYs.

Adds Figure 11 (composed topology) + Figure 12 (remote-delegation sequence: the
local?inject:publish fork with a sandboxed far side — both injection points stay
local, only the middle hop crosses the broker), the two forced reachability
changes (host-routable mcpUrl; PTY in the local gateway's herdr), and a
maps-to-existing-seams table (CB-308 gateway × Dev-A launcher, CB-117 reap,
CB-303 container lifecycle, CB-308 #5 trust). No new pillars. Both diagrams
mmdc-validated; §4/§9 updated to point at §11.
2026-07-28 16:26:10 +02:00
Dai Ha ded226abfe Merge CB-402: opencode second peer adapter (Stage B of the Peer Launcher SPI)
Proves the CB-401 PeerLauncher SPI is genuinely provider-neutral by landing a
second adapter — opencode — that shares NONE of Claude Code's private launch
seams (no ANTHROPIC_BASE_URL, no SubscriptionGuard). Merged as one unit:

- Incr 2 (6e37722): kind: discriminator on BridgedConfig.Worker (claude-code |
  opencode), argv defaults to the kind binary, isClaudeCode()/isOpenCode().
- Incr 3 (b034f10): OpenCodeLauncher extends HerdrPeerLauncher — diverges only
  in buildLaunch (file-based MCP mount via OPENCODE_CONFIG + reply charter under
  instructions[], model via -m provider/model, opencode- name/reap prefix).
- Incr 4 (11f8709): CompositePeerLauncher routes the fleet by kind (spawn/cwd/
  parity by profile, stop by pane-id owner, list dedup, reap/caps/profiles
  union); Bridged.main partitions profiles by kind → composite. BridgeMcp +
  BridgedApp migrated onto the PeerLauncher SPI (the two (ClaudeCodeLauncher)
  casts removed; dead rendezvous params dropped).

Incr 1 (extract HerdrPeerLauncher base) already on main at ffce30a.
266 tests green, mvn BUILD SUCCESS. Live opencode-on-Gemini dogfood tracked
separately (needs a running-daemon restart + a resolved Gemini provider).
2026-07-28 16:17:34 +02:00
Dai Ha 9b8d55bc18 CB-500: design note — multi-tier coordination (Stage 6)
Scopes the lead's next direction on top of the peer-launcher arc, as a
proposal (ticket split deferred):
- A · sandboxed, role-specific workers — a sandbox is a *placement*, so it
  slots into the CB-401 SPI as a new kind: sandbox adapter exactly the way
  CB-402's opencode slotted in as a new provider (SPI proven placement-neutral,
  not just provider-neutral). Ownership line held: bridge launches INTO a
  peer-owned image, never provisions the IDE/dev-tools inside it.
- B · main-agent pairs (Opus + cloud) — both mains are MCP clients so neither
  can be called into; each needs a pull inbox → depends on CB-308's per-agent
  channels. PrimaryRegistry single-slot → multi-slot.
- C · orchestrator tier — SessionManager recursed one tier up
  (orchestrator:mains :: main:workers) + context scoping; re-roots the human
  from a live primary to the orchestrator.

10 mermaid diagrams (component + sequence per development, ownership guardrail,
staging graph), all mmdc-validated and theme-safe. §7 pins the bus-vs-env-manager
boundary as an acceptance criterion on A; §8 stages A → CB-308 substrate → B → C.
2026-07-28 15:20:16 +02:00
Dai Ha 11f8709286 CB-402 Increment 4: CompositePeerLauncher — route the fleet by kind
Introduce the router the core holds when more than one adapter is
configured: one HerdrPeerLauncher per peer kind, dispatched by profile
(spawn/effectiveCwd/parityOverlay), by pane id (stop, via a spawn-time
owner map), and fanned out + combined for the fleet-wide queries
(list dedup by pane id, reap/caps union, profiles union). The ctor
rejects an empty adapter list and a profile two adapters both claim.

Wire it in Bridged.main: partition workerProfiles() by kind (claude-code
is the always-present default adapter; opencode is added when any profile
opts in) and front both with the composite. This lets BridgeMcp and
BridgedApp finally take the PeerLauncher SPI instead of a concrete
ClaudeCodeLauncher — the two (ClaudeCodeLauncher) casts in Bridged are
gone. list() elements are cast to herdr Agent at the point of the
herdr-specific roster view, where that assumption actually lives.

Add BridgedConfig.Worker.isClaudeCode()/isOpenCode() kind predicates
(the wiring uses isOpenCode; both are unit-tested). Drop the long-dead
'rendezvous' constructor param threaded into BridgeMcp and BridgedApp.

10 CompositePeerLauncherTest cases over two real adapters on one
FakeHerdr: profile routing (observed via the started agent's
claude-/opencode- name prefix), default resolution, unknown-profile
and duplicate-profile rejection, caps union, list dedup, reap sum,
and stop teardown. 266 tests green.
2026-07-22 05:43:28 +02:00
Dai Ha b034f105c0 CB-402 Increment 3: OpenCodeLauncher — the SPI-proving second adapter
A HerdrPeerLauncher subclass for opencode, a provider-agnostic terminal
coding agent. It reuses every line of shared base transport (tab/pane
placement, CB-306 readiness gate, unique naming + CB-117 reap, teardown,
listing, cwd) and diverges only in buildLaunch:

  - No subscription boundary: no ANTHROPIC_BASE_URL, no SubscriptionGuard
    (the guard is a Claude-private concern, not part of the SPI).
  - File-based MCP mount + instructions: writes an ephemeral opencode.json
    declaring the bridge as a remote MCP server + a reply-charter file under
    instructions, pointed at via OPENCODE_CONFIG (opencode has no inline
    --mcp-config / --append-system-prompt).
  - Model selected with -m provider/model, not an env var.
  - 'opencode' name prefix so reap matches opencode-* panes only.

configRoot is injectable so tests inspect the generated config/charter under
a @TempDir. 10 tests cover config content, model flag, git-token grant,
capabilities, reap predicate, both production ctors, and the readiness gate
(throws PeerUnreachable on timeout, reaps only the worker pane).
2026-07-22 05:30:46 +02:00
Dai Ha 6e37722383 CB-402 Increment 2: kind: discriminator on Worker profiles
Add a `kind` field to BridgedConfig.Worker — "claude-code" (default) or
"opencode" — the discriminator the CompositePeerLauncher will route spawn/reap
by so each adapter drives only its own peer kind. Normalised to lower-case;
blank/absent ⇒ claude-code, so every existing config and call site is
unchanged. argv now defaults to the kind's own binary (claude vs opencode)
rather than always `claude`, so an opencode profile never inherits the Claude
command.

kind is appended at the record tail; a new 14-arg back-compat constructor
(git fields, no kind) keeps the CB-302 call sites working, and the existing
12-arg constructor is untouched. Also drop the never-used Primary(String)
legacy constructor to keep the file warning-clean.

example.yaml documents the key and carries a commented opencode-gemini
profile. Tests cover default/normalisation/argv-defaulting. 245 tests green,
config files 0 IDE problems.
2026-07-22 05:20:04 +02:00
Dai Ha ffce30afa2 CB-402 Increment 1: extract HerdrPeerLauncher abstract base
Behaviour-preserving refactor ahead of the second-adapter work. All herdr
transport shared by any peer kind — tab/pane placement, the CB-306
spawn-readiness gate, unique naming + CB-117 orphan reap, teardown, listing,
cwd resolution, and the peer-neutral git-forge grant — moves into a new
abstract HerdrPeerLauncher (Template-Method base). ClaudeCodeLauncher becomes
a final subclass supplying only the two Claude-specific seams: the `claude`
name prefix and buildLaunch(), which encodes the subscription boundary
(ANTHROPIC_BASE_URL + SubscriptionGuard assert, inline --mcp-config and
--append-system-prompt reply charter).

The base owns the injectable clock + sleeper for the readiness gate; the poll
interval is baked into the sleeper, so the vestigial spawnReadyPollMs field is
dropped from the base and from the full testability constructor (the explicit
sleeper already encodes it). The 6-arg and 8-arg production constructors keep
their signatures; three full-ctor test sites drop the now-unused poll argument.

No behaviour change: 242 tests green, both refactored files 0 IDE problems.
2026-07-22 05:15:21 +02:00
Dai Ha e724a59f2d CB-402: design note — opencode second peer adapter (Stage B)
Design-note-first for gitea #7. Extract HerdrPeerLauncher abstract base
(Template Method) + kind: discriminator + OpenCodeLauncher + routing
CompositePeerLauncher; finish the Stage-A caster migration off
ClaudeCodeLauncher. opencode proves the SPI for a non-Claude peer
(no subscription guard, OPENCODE_CONFIG MCP mount, config-based charter).
2026-07-22 04:59:58 +02:00
Dai Ha 4bf855d225 wiki: bump submodule pointer to 4d1548a (CB-307 as-built)
Advance the wiki submodule pointer to include the CB-307 Stage 2/3 as-built
docs plus the intervening CB-401/delegation-directive commits. Standalone
pointer bump — not bundled into a feature commit.
2026-07-19 18:10:01 +02:00
Dai Ha d4c9704007 CB-307: primary-gate cleanup of push-loop (drop dead clock, IDE 0/0)
Primary verification pass over the delegated push-loop delivery:
- Remove the unused LongSupplier clock threaded into ReplyPushLoop
  (timing is the scheduler's; the field was never read) from the
  component, Bridged wiring, and both test call sites.
- Collapse the single-statement WAIT_BUSY switch arm (redundant block).
- Drop now-dead test scaffolding: the always-"idle" recordingClient
  param and unused AgentStatus/AtomicReference imports.

IDE diagnostics 0/0 on all changed files; mvn clean install green
(242 tests, 0 failures).
2026-07-19 11:03:34 +02:00
Dai Ha 7c252b5f5f CB-307 Increment 3: bridge_ack tool — per-msgId ack refinement
- MessageService.ackReply(target, msgId) delegates to inbox.ack
- BridgeMcp registers bridge_ack tool with target/msgId args
- Tests: valid/invalid args, ack surface via BridgeMcp
2026-07-19 10:21:06 +02:00
Dai Ha f756933879 CB-307 Increment 2: ReplyPushLoop — status-gated push loop (mechanism b)
- ReplyPushLoop: dedicated scheduled loop for nudging the primary
- Package-private decide() method for pure decision logic (unit-testable)
- Status-gated injection via AgentControl.status().injectable()
- Bounded reminders (cap + backoff), idempotent per target
- Wire into MessageService.reply after inbox.publish on no-waiter branch
- Wire into Bridged.main (constructor + shutdown hook)
- Primary config record updated with pushReminders/pushBackoffMs knobs
- Tests: decide() matrix, nudge injection, idempotency, cap enforcement
2026-07-19 10:16:24 +02:00
Dai Ha a1aecbf4fc CB-307 Increment 1: PrimaryRegistry + config + wiring
- PrimaryRegistry: thread-safe single-slot registry with pin support
- Primary config record (last positional, like Broker)
- Wire capture in BridgeMcp (bridge_send and bridge_spawn handlers)
- Construct PrimaryRegistry in Bridged.main
- Tests: PrimaryRegistryTest + BridgedConfigTest primary config cases
2026-07-19 09:59:02 +02:00
Dai Ha 131e7b1ccd CB-307: lock push-loop injection to dedicated status-gated loop (mechanism b)
Chose a small dedicated scheduled loop over AgentControl.send guarded by an
injectable status check, instead of reusing the worker Injector (which couples
to WorkerPresence/StatusPoller). Isolated + unit-testable via injected clock.
Records live ground truth: this primary resolves to term_656c8cc03e1f0b1 (w2:pY)
— confirms the primary runs in a herdr pane so the push path is exercisable.
2026-07-19 09:44:21 +02:00
Dai Ha d0ac6c435f CB-307: design note for active push-to-primary + bounded reminder loop
The reliability layer over the durable inbox (Stage 2, 2bc5f3a): push a nudge
into the primary's own herdr pane the moment a no-waiter reply lands, remind on
a bounded backoff until the primary drains (ack=drain), degrade to pull when the
primary pane isn't resolvable. Grounds the seam: the primary terminal_id is
already derivable via ConnectionIdentity/PaneLocator, just discarded today.

Refs gitea #5.
2026-07-19 09:41:22 +02:00
Dai Ha 2bc5f3a057 CB-307 Stage 2: AmqpReplyInbox — durable, cross-restart reply delivery
Behind the existing ReplyInbox port, add an AMQP-backed adapter selected by a
`broker:` block in config (absent → the in-memory soft-state inbox; present →
AMQP). Mapping is consume-and-hold with deferred manual ack: each target owns a
durable queue `agent.<target>.inbox`; a manual-ack consumer pulls persistent
messages into an in-memory held map (dedup by msgId) but does not ack; peek
returns the snapshot; ack acks the broker delivery-tag and drops it. A crash
before caller-ack leaves messages unacked, so the broker redelivers on
reconnect — genuine durability with the port contract preserved. bridged still
owns no persistence; the broker does.

- msg/AmqpReplyInbox: the adapter (single synchronized channel; recovery
  listener clears held on reconnect so fresh delivery-tags repopulate).
- config/BridgedConfig: nullable Broker(uri) record; isConfigured() gates it.
- Bridged.main: select adapter; close the AMQP connection in the ordered
  shutdown hook (no-op for the in-memory inbox).
- deps: com.rabbitmq:amqp-client (main); testcontainers rabbitmq/junit-jupiter
  (test). Pinned commons-compress 1.27.1 + commons-lang3 3.18.0 to clear the
  test-scope CVEs those pull. Production default LavinMQ; RabbitMQ URI-swap.
- tests: BridgedConfigTest broker-selection cases (hermetic); AmqpReplyInbox
  contract test (@Tag("contract"), Testcontainers RabbitMQ) proving
  publish/peek/ack, msgId dedup, and cross-restart redelivery. Excluded from
  the default build so `mvn clean install` stays hermetic (210 green).
2026-07-19 07:30:29 +02:00
Dai Ha ba6b4a5da9 CB-307 Stage 1: reply-inbox port + in-memory adapter — hold stranded worker replies instead of dropping them
Problem: the reverse (worker->primary) path was Rendezvous, a map of LIVE blocking
waiters only. A bridge_reply arriving with no open send hit Rendezvous.complete()
-> no waiter -> returned false -> the reply was silently DISCARDED (worker saw an
error / REST 409). No message-id/dedup/ack existed anywhere.

Stage 1 (no broker, soft-state) behind one port:
- ReplyInbox port + InboxMessage record; InMemoryReplyInbox adapter (per-target
  FIFO via LinkedHashMap, dedup by msgId, thread-safe). Soft-state, not persistence.
- MessageService.reply(session, content): resolve an open send, else publish to the
  inbox with a minted UUID (was a silent drop). drainReplies(target) = peek + ack.
- BridgeMcp.reply / BridgedApp.replyMessage repointed off bare Rendezvous.resolve
  onto messages.reply -> no-waiter is now SUCCESS (queued), not error / 409.
- Drain surface: bridge_poll gains optional target; REST GET /sessions/{id}/replies.
- Rendezvous left untouched. QUESTION path (bridge_ask) NOT queued (interactive,
  keeps NO_WAITER); completion/failure fallbacks NOT queued (captured-waiter).
- Bridged.main wires new InMemoryReplyInbox(); no broker: config yet (Stage 2 = AMQP).

Tests: +19 (188 -> 207), 0 failures/0 errors. New InMemoryReplyInboxTest (12) +
MessageService/BridgeMcp/BridgedApp coverage incl. guards proving a QUESTION and a
completion fallback are never queued.

Implemented via delegation to an off-sub gx10 worker in a pre-trusted worktree;
primary-verified (mvn clean install green, 207 tests) and committed by the primary
because the worker's completion replies were lost to the very bug this fixes.

Refs CB-307 (gitea #5), Stage 1 of 2.
2026-07-18 21:12:51 +02:00
Dai Ha da5a987df0 CB-307/CB-308 design notes: reliable-delivery Stage-1 delegation spec + multi-host federation proposal
- docs/CB-307-Reliable-Delivery.md: ReplyInbox port + in-memory adapter spec
  (Stage 1, no broker); publish at the Rendezvous no-waiter drop seam, drain by target.
- docs/CB-308-Multi-Host-Federation.md: per-host gateway + per-agent broker channels
  + federated roster proposal (gitea #6), wiki-ready with theme-safe mermaid.
2026-07-18 20:24:16 +02:00
Dai Ha 7dd6c46156 CB-306: spawn-readiness gate — ClaudeCodeLauncher blocks until the worker is injectable or throws PeerUnreachableException
The launcher now polls AgentControl.status(paneId) after starting the pane.
It returns the handle only once the worker reports an injectable state
(IDLE/BLOCKED/DONE). If the timeout elapses while still UNKNOWN, the
pane is self-reaped and a PeerUnreachableException is thrown — no orphan
left behind. The gate is disabled when spawnReadyTimeoutMs == 0 (legacy
non-blocking spawn, the default for the 6-arg constructor).

Key changes:
- PeerUnreachableException (new) in dev.ltms.bridged.peer
- BridgedConfig: spawnReadyTimeoutMs (default 20000), spawnReadyPollMs (default 300)
- ClaudeCodeLauncher: 3 constructor overloads:
  (a) 6-arg backward-compat: gate disabled (timeout=0)
  (b) 8-arg production: gate with config knobs + real clock/sleep
  (c) 10-arg testability: full seam (LongSupplier clock + Runnable sleeper)
- waitUntilInjectableOrThrow() loop in spawn(SpawnRequest)
- sleepUninterruptibly() helper for the production sleeper
- BridgeMcp.spawn + BridgedApp.spawnWorker catch PeerUnreachableException
  → clean tool error / 502 response (not an uncaught 500)
- SessionManager.acquire inherently registers nothing on throw (both
  worktree and non-worktree paths) — confirmed by new test

Tests:
- ClaudeCodeLauncherTest: 4 new tests
  - unknown→idle: returns handle, no pane.close
  - always-unknown: throws PeerUnreachableException, pane closed,
    clock advanced past timeout
  - timeout=0 (6-arg ctor): no agent.get calls, returns handle
  - timeout=0 (10-arg ctor): no orphan pane close
- SessionManagerTest: 1 new test
  - acquire → PeerUnreachableException: roster remains empty
Total: 188 tests, all pass (no existing test changed semantics)
2026-07-18 16:17:43 +02:00
Dai Ha 3a5cdc5108 CB-306: spawn-readiness gate design note (delegation spec) 2026-07-18 15:41:43 +02:00
Dai Ha 3aa69a9e32 CB-401 Stage A follow-up: rename WorkerService -> ClaudeCodeLauncher
Name the first-class Claude Code adapter explicitly, per the Peer Launcher SPI:
WorkerService was the de-facto Claude-Code launcher; as an in-tree PeerLauncher impl
it should say so. Pure IDE rename (class + file + WorkerServiceTest) plus stale
Javadoc/comment mentions swept to the new name. No behaviour change.

Gate: IDE diagnostics 0/0 on touched files; mvn clean install BUILD SUCCESS,
MVN_EXIT=0, 183 tests pass. Deferral #1 from issue #3 cleared.
2026-07-18 14:39:56 +02:00
Dai Ha e056c7e1fa CB-401: Stage A - extract PeerLauncher SPI in-tree 2026-07-18 07:48:55 +02:00
Dai Ha d63273d082 CB-401: Peer Launcher SPI design note (Stage 4)
Design-only. Audits the as-built Claude/herdr coupling (concentrated in
WorkerService), defines a PeerLauncher SPI + opaque PeerHandle + capability
model so the bus delegates peer materialization to a config-selected adapter.
ClaudeCodeLauncher = adapted WorkerService. Stages A/B/C with the Stage-C
plugin-loading security gate called out. No production code touched.
2026-07-17 17:49:21 +02:00
Dai Ha 0efb65ca0a CB-303: session lifecycle limits — idle_ttl reaper, context_cap, graceful drain
Verified on primary: ide_diagnostics clean (incl. weak warnings), mvn clean install
BUILD SUCCESS, 173 tests. Delegated impl (worker/cb-303-80ec1a-3, 3 parts), primary-gated.
2026-07-17 10:08:55 +02:00
Dai Ha 9fe04bfb08 CB-304: bridge_list roster + live herdr join (surface worktree/branch); add GET /workers
Verified on primary: ide_diagnostics clean (incl. weak warnings), mvn clean install
BUILD SUCCESS, 165 tests. Delegated impl (worker/cb-304-bd1e4f-2), primary-gated.
2026-07-17 10:08:46 +02:00
Dai Ha 09d3948acf CB-303 part 3: graceful drain on shutdown 2026-07-17 09:59:57 +02:00
Dai Ha 954351a80b CB-303 part 2: context_cap turn budget 2026-07-17 09:56:46 +02:00
Dai Ha 8d51066ddd CB-303 part 1: idle_ttl session reaper (injectable clock + SessionReaper) 2026-07-17 09:52:40 +02:00
Dai Ha 84102baab4 CB-304: bridge_list roster + live herdr join (worktree/branch); add GET /workers 2026-07-17 09:50:12 +02:00
Dai Ha 64e70efdf1 CB-302: worker checkpoint — repo-scoped forge token injection + implementer skill
The worker "checkpoint" is commit → push → open its own PR. Push is free over SSH
(same user, same keys); the only incremental grant is PR-create, so the daemon injects
a repo-scoped gitea token into the worker env — opt-in per profile, never mutating
bridged's own environment.

- BridgedConfig.Worker: gitTokenEnv/gitHostEnv fields (opt-in; gitHostEnv defaults to
  GITEA_HOST). Backward-compat 12-arg constructor keeps pre-CB-302 call sites + YAML
  working. hasGitToken() gates injection.
- WorkerService.spawn: inject GITEA_TOKEN (and paired GITEA_HOST) only when the profile
  grants a token AND the host env resolves one. resolveEnv() tolerates unset var names.
- WorkerServiceTest: injection present for a granting profile; absent when not (proving
  the gate is config, not a missing env var).
- .claude/skills/implementer/SKILL.md: worktree-aware playbook — confirm the worktree/
  branch, implement, commit (never .mcp.json/wiki), push, open PR via gitea REST with
  GITEA_TOKEN, hand off the PR URL via bridge_reply. Never merge; workers can't run IDE
  diagnostics so never claim IDE-clean.

Whole-project gate: mvn clean install green, 164 tests, 0 failures.
2026-07-17 08:51:27 +02:00
Dai Ha 97ecc7136e CB-301-ext: per-worker git worktree + config-parity overlay
Opt-in isolated worktree so parallel implementers don't stomp the shared
tree, hydrated to config parity so a worker differs from the primary only
in LLM provider.

- Worktrees seam (interface) behind SessionManager; GitWorktrees shells git
  via ProcessBuilder (non-zero exit -> WorktreeException), FakeWorktrees for
  tests. No live git in unit tests.
- acquire() 5-arg overload provisions add -> overlayParity -> spawn(cwd=wt)
  -> register, unwinding the worktree on any failure before registration.
  4-arg overload and shared-tree behavior unchanged (backward compatible).
- release() removes the checkout but never deletes the branch (it holds the
  worker's commits + PR, CB-302).
- overlayParity copies local config (.mcp.json, settings.local.json, .env/
  .envrc) into the worktree; tracked ones get --skip-worktree so a worker
  can never stage the parity overlay.
- WorkerSession gains nullable worktree/branch; BridgedConfig.Worker gains
  parityOverlay (default list) + top-level worktreeRoot.
- bridge_spawn / POST /workers gain an optional worktree(+ticket) arg; the
  worker view includes worktree/branch only when non-null.

Verify fixes on the delegated impl: strip trailing dashes in slug()
(^-+|-+$, was ^-+|^-+$); make FakeWorktrees.add a pure fn of the branch
(nonce already unique); MCP worktreeRequest treats blank/"false" string as
no-worktree, matching the REST builder.

162 tests, 0 failures.
2026-07-17 06:46:20 +02:00
Dai Ha f9073e2320 docs: CB-301-ext spec — worktree provisioning + config-parity overlay
Opt-in per-acquire worktree (shared-tree default preserved). git behind a
Worktrees seam (ProcessBuilder impl, fake in tests). acquire provisions
worktree+branch, overlays local config (copy + --skip-worktree on tracked
files so the worker can't commit .mcp.json), spawns with cwd=worktree.
release removes the worktree but keeps the branch (holds commits/PR).
WorkerSession gains nullable worktree/branch; BridgedConfig.Worker gains
parityOverlay. 6 fake-based acceptance tests incl. backward-compat + unwind.
2026-07-17 06:32:09 +02:00
Dai Ha 54d907c314 CB-301: SessionManager — authoritative worker session registry + one-shot FSM
Adds dev.ltms.bridged.session with WorkerSession (immutable record) and
SessionManager wrapping WorkerService: a ConcurrentHashMap registry keyed by
paneId, the one-shot lifecycle FSM (SPAWNING->READY->BUSY->DONE, ->FAILED on
drop/turn-failure, ->RELEASED on teardown), ownership (ownerTerminal), and
recycle = release + fresh acquire (no-reuse invariant). Driven by TurnListener
(BUSY/DONE/FAILED) and a WorkerPresence bridge (READY).

Wiring: Bridged.main constructs it and composes it into the TurnListener
alongside CompletionResolver; bridge_spawn / POST /workers route through
acquire (carrying caller identity as owner); bridge_stop / DELETE /workers
route through release. WorkerService gains effectiveCwd(); WorkerPresence
de-finalized so the manager can present a READY-driving view.

asPresence() returns a single cached bridge (a fresh one per call would
fragment the shared present set). roster() is the registry snapshot; the live
herdr join is left for CB-304. 6 fake-based acceptance tests; full suite green
(155/155).

Delegated to an off-subscription worker against docs/CB-301-Session-Manager.md;
primary verified (ide diagnostics clean, mvn clean install green) + fixed the
asPresence caching bug.
2026-07-16 19:30:56 +02:00
Dai Ha 82c7d6553a docs: worker git workflow — daemon worktree + config parity + worker-opened PR
Worktree is code-only isolation; SessionManager hydrates it to full config
parity (overlay untracked local settings/.mcp.json/.env) so a worker is a
full peer of the primary, differing only in the LLM provider. Worker commits,
pushes over SSH, and opens its own PR (gitea REST + repo-scoped token).
2026-07-16 19:25:41 +02:00
Dai Ha 19b10e3216 docs: CB-301 session-manager design spec (one-shot, no reuse; recycle in scope)
The as-built audit surfaced that WorkerService keeps no registry of what it
spawned (its own Javadoc: 'there is no registry; list() only asks herdr').
CB-301 adds a SessionManager wrapping WorkerService: an authoritative in-daemon
roster with a per-session lifecycle FSM (SPAWNING/READY/BUSY/DONE/RELEASED/
FAILED), deterministic release, and recycle (= release + fresh acquire, no
reuse). Leaves clean seams for CB-302 (checkpoint on release), CB-303 (idle_ttl/
context_cap/drain policy over roster), CB-304 (bridge_list reads roster).
2026-07-16 19:10:13 +02:00
Dai Ha 0f79e6bed5 wiki: bump submodule to page-9 as-built implementation architecture (8e5fd01) 2026-07-16 16:45:31 +02:00
Dai Ha aa0cf814ff e2e: capture bridge_ask transcript confirming single turnId after coalescing
Post-4aa9d03 live run: the duplicate-ask retry now coalesces onto one turn
(turnId=term_...#1, was #2 before the fix). Clean round-trip, RESULT OK.
2026-07-16 16:32:30 +02:00
Dai Ha 4aa9d03da2 CB-205: coalesce duplicate bridge_ask calls onto one turn
A worker's tool execution is single-threaded, so two bridge_ask calls from the
same session can only be a transport retry — yet openAsk minted a fresh turnId
each time and both raced to resolve the primary's single forward waiter, the
loser returning NO_WAITER and leaving a dangling turn (observed as turnId #2 in
the live e2e). Now openAsk is idempotent per session: a second open ask coalesces
onto the existing turnId + answer future (fresh=false), and only the fresh owner
surfaces the question and tears the turn down. closeAsk clears the per-session
index (conditional by value) so a later ask reopens fresh.

Implemented by an off-subscription worker delegated over the bridge; reviewed and
validated on the primary (IDE-clean, mvn clean install 149 green, +2 tests:
concurrent double-ask coalescing + post-close reopen).
2026-07-16 16:27:10 +02:00
Dai Ha 37abfdfd6e wiki: bump submodule to Stage-2-complete roadmap update (0c81b51) 2026-07-16 16:15:15 +02:00
Dai Ha d5c3ede215 CB-202: reviewer-role skill for bridged workers
The playbook a reviewer worker loads when the lead delegates a scoped review:
read the whole scope before judging, stay in the assigned lane, ask the lead
via bridge_ask when the call is genuinely theirs (resuming the same turn with
the answer), and report exactly one structured finding via bridge_reply. Pairs
the already-shipped bridge_reply/bridge_ask tools with the role guidance that
tells a worker how to use them. Mermaid validated with mmdc.
2026-07-16 16:10:36 +02:00
Dai Ha 426855e378 e2e: live bridge_ask reverse-rendezvous harness (CB-205)
Drives the reverse path end to end over REST loopback: a worker is delegated
a task it cannot finish without asking, calls bridge_ask mid-turn, and the
primary answers on the surfaced turnId so the worker resumes the SAME turn.
Two blocking sends, no polling. Verified live: worker asked in ~9s, resumed
and replied CHOSEN=BLUE via clean bridge_reply after the primary answered.

Subscription-safe by construction (REST face only; never sets ANTHROPIC_BASE_URL).
2026-07-16 16:08:21 +02:00
Dai Ha 358c6970b5 CB-205/CB-201: bridge_ask reverse rendezvous + lightweight question/turn kind
A worker can now pause its delegated turn to ask the primary a question and
resume the same turn with the answer — the reverse of bridge_send.

- Rendezvous: QUESTION kind carrying a turnId, plus a reverse-ask registry
  (openAsk/resolveQuestion/askSession/answerAsk/closeAsk).
- MessageService.ask(): surface a worker's question to the primary's open send,
  block for the answer. answer(): resolve the worker's ask by turnId, then block
  for its eventual bridge_reply (session derived from turnId, not an argument).
- BridgeMcp: bridge_ask tool (worker-only, identity from the connection);
  bridge_send routes a turnId to the answer path. Shared formatReply().
- BridgedApp REST parity: POST /sessions/{id}/ask, turnId on /message.
- Lightweight CB-201: the kind vocabulary is the QUESTION outcome + turn_id
  correlation, not a rigid from/to/corr envelope (the connection-identity
  mechanism already covers addressing more robustly).

Tests: MessageServiceTest ask/answer round-trip + NO_WAITER/timeout/stale;
new RendezvousTest for the reverse registry; BridgeMcp ask/answer parity.
mvn: 147 green.
2026-07-16 15:32:56 +02:00
Dai Ha 55ebd5b949 e2e: sustained back-and-forth conversation harness (5-min stateful continuity) 2026-07-16 15:19:21 +02:00
Dai Ha 2a61fe69f1 CB-118: clip the completion baseline so the CB-115 guard survives >cap blocks
captureBaseline stored the raw, unclipped last-assistant block while resolve()
compares against clip(...) capped at MAX_SCRAPE_CHARS. For a block longer than
4000 chars the two capped representations never match even when the pane is
unchanged, defeating the CB-115 misattribution guard and letting a stale
completion resolve a rapid back-to-back send. Clip the baseline identically.

Regression test: an unchanged >cap block stays suppressed.

Surfaced by the fan-out issue-hunt E2E (1 primary -> 3 concurrent workers,
e2e/issue_hunt_test.py, added here). The same hunt's WorkerService.stop() and
Rendezvous.complete() findings were verified as false positives (locatePane is
already guarded; the sender's finally-close already removes the waiter).

Closes #2
2026-07-16 09:11:20 +02:00
Dai Ha ff6aacdc78 CB-117: reap orphaned worker panes on startup
herdr keeps worker panes alive across a daemon restart by design, and a
worker's paneId is held only by its spawner — so a worker whose owning
process exited before its DELETE leaks with nothing tracking it (there is
no registry; list() only asks herdr). Observed as three idle claude-ollama
panes left in the worker space from earlier runs.

On boot, WorkerService.reapOrphanWorkers() scans herdr for agents whose
name matches our claude-<profile>-<nonce>-<seq> scheme with a nonce other
than this process's nameNonce, and tears each down (pane + its now-empty
dedicated tab). A current-nonce worker is ours and live (spared); a user's
own claude session carries no such name (untouched). Keyed on the nonce so
it survives kill -9 and reaps a *previous* daemon's leaks — the actual case
shutdown-hook reaping and an in-memory registry both miss.

- Agent now projects herdr's 'name' (was dropped) so the reaper can key on it.
- isForeignWorker/workerNonce are pure + package-private for unit testing.
- FakeHerdr.withAgent seeds named agents into agent.list.
- Wired best-effort into Bridged startup before serving.

Closes lms/claude-bridge#1
2026-07-16 08:39:37 +02:00
Dai Ha 5f0ec034d9 e2e: standard bridge conversation test harness
A repeatable multi-turn primary↔worker conversation driven entirely through the
bridge's loopback REST face (async fire-and-poll) — never sets ANTHROPIC_BASE_URL
and never touches herdr, so it is subscription-safe by construction. Records every
turn to a transcript, grades each (OK / DEGRADED / EMPTY / FAILED / WEDGE), and
exits non-zero if any turn fails to deliver-and-reply, so it is CI-usable.

This is the repeatable form of the ad-hoc channel test that surfaced the CB-115 and
CB-116 gaps.
2026-07-16 08:27:50 +02:00
Dai Ha 31e34d177b CB-115/CB-116: reliable turn completion — status refinement, clean scrape, waiter identity
A 5-turn primary↔worker conversation test (see the e2e harness) surfaced three
delegation-channel gaps; this closes them.

CB-115 — status + scrape correctness:
- AgentStatus gains DONE (herdr's explicit turn-complete marker) so a finished
  turn is no longer misread as UNKNOWN and left to wedge or false-fail.
- StatusRefiner reclassifies content-bearing UNKNOWN samples (StatusPoller wired
  to it), and the Injector baselines pane content on delivery (TurnListener gains
  onDelivered) to guard completion against previous-turn misattribution.
- CompletionResolver.lastAssistantBlock stops at the first hard TUI boundary, so a
  scrape returns only the assistant answer — no input box, prompt echo, spinner,
  tips or warnings.

CB-116 — waiter identity (the cross-turn stale reply):
- The completion/failure fallback ran on a virtual thread and resolved whichever
  waiter was currently registered for the session. Since the rendezvous holds one
  waiter per session and sends serialize, turn N's late completion could land on
  turn N+1's waiter and deliver turn N's stale scrape as turn N+1's answer. The
  baseline guard missed it because turn N was resolved by bridge_reply, which never
  updates the completion baseline.
- Fix: capture the exact waiter (and pre-turn baseline) when a turn is delivered,
  on the poller thread before any next-turn delivery can overwrite it, and resolve
  THAT waiter — a no-op if it was already resolved. Rendezvous.resolveCompletion/
  resolveFailure now take the captured CompletableFuture; currentWaiter exposes the
  registered one for capture. A late completion for turn N can no longer touch turn
  N+1's send.

Verified: 128 unit tests green (incl. a CB-116 regression asserting a late
completion never resolves the next turn's waiter); a re-run of the conversation
test passes with turn 5 resolving to its own reply rather than turn 4's text.
2026-07-16 08:27:38 +02:00
Dai Ha b2d85af78b wiki: bump submodule to roadmap Stage-1-complete update (4304dc4) 2026-07-16 06:49:27 +02:00
Dai Ha 968a5c68b6 docs(README): update Status to reflect the shipped implementation
The Status section still read 'Design selected' — but bridged is built and
dogfooded (CB-101..114, 105 tests, live MCP tools, multi-profile, cwd inheritance,
readiness gate). Replace it with an accurate shipped/next breakdown, and drop the
bridge_ask overclaim from the gateway bullet (bridge_ask is roadmap, not built).
2026-07-16 06:46:40 +02:00
Dai Ha 3c05823f19 CB-114: resolveCwd never returns null (honor the 'daemon cwd, never $HOME' contract)
Final review-sweep finding: firstNonBlank(requestedCwd, cfg.cwd(), callerCwd,
user.dir) returns null if all are blank (pathological env with user.dir unset),
after which AgentControl drops the cwd and herdr defaults the pane to $HOME —
violating CB-112's documented contract. Append "." (the daemon's own cwd) as a
guaranteed non-blank last resort. Near-impossible trigger; makes the code honor its
own javadoc. 105 tests green.
2026-07-16 06:43:05 +02:00
Dai Ha 2fb46f670f CB-114: readiness-gate timeout + presence cleanup (delegated review findings)
An off-sub worker's review of CB-113 (delegated through the bridge) surfaced two
real gaps in the readiness gate:

- A worker that herdr reports idle but whose Claude never connects the bridge MCP
  (crashed during boot, or wedged on a startup prompt) left its message queued
  forever: ready.test() never passed, the target was polled indefinitely, and the
  caller's future never completed (async waiter hung for the full 30-min window).
  Injector now counts injectable-but-not-ready samples and, after a ~60s grace
  (READINESS_GRACE_POLLS, deliberately longer than the UNKNOWN stall grace since a
  first boot is slower than an in-turn blip), fails the queued messages, fires
  onTurnFailed so blocking/async waiters resolve WORKER_FAILED, and reclaims the
  target. Mirrors the CB-109 UNKNOWN-stall path.

- WorkerPresence.forget had no caller, so a worker's readiness lingered past its
  life. Injector now clears presence via a forget callback on drop() (pane crash)
  and on the readiness timeout.

+3 InjectorTest cases (never-ready failure, ready-within-grace delivery, drop clears
presence). 105 tests green.
2026-07-16 05:01:58 +02:00
Dai Ha 37f7ad4185 CB-113: reliable worker readiness gate + submission nudge
Delegating right after spawn failed: herdr reports 'idle' during the worker's
boot, so the injector delivered into a not-ready TUI (paste lost) and wedged the
worker. Two fixes:

- Readiness gate: a worker is 'available' only once its Claude connects the bridge
  MCP (the daemon observes it via peer-PID->terminal). WorkerPresence tracks it; the
  injector holds the first delivery until present, so it never pastes into the boot
  window. Exposed as 'ready' on GET /sessions/{id}/status.
- Submission nudge: the Enter accompanying a delivery can race the paste (esp. right
  as the TUI becomes ready), leaving text unsubmitted. While a delivered message
  stays idle (not picked up), the injector re-sends Enter each poll until the worker
  starts (WORKING) or the grace expires.

Validated live: spawn + immediate delegate now holds during boot (ready=false),
delivers on MCP-connect, re-nudges Enter, worker replies. 102 tests green.
2026-07-15 19:14:57 +02:00
Dai Ha 926724a279 CB-112: workers inherit the primary's working directory (not $HOME)
A worker now opens the same directory the primary is in, unless told otherwise.
Resolution: explicit spawn cwd → per-profile config cwd → the primary's cwd
(auto-detected from the bridge_spawn caller's PID via lsof -d cwd) → the daemon's
cwd. Never $HOME.

Mechanism (found by live probe, corrects the earlier assumption): an agent.start
pane does NOT inherit its tab's or workspace's cwd — it starts in $HOME. herdr's
agent.start honours an (undocumented) cwd param, so the resolved cwd is threaded
onto agent.start {cwd} (both tab and pane placement), not tab.create.

Surfaces: bridge_spawn {cwd?} + auto-detect via ConnectionIdentity.resolve (peer
PID) + ProcessCwdLookup (lsof); REST POST /workers ?cwd= / body cwd; per-profile
'cwd:' config. Validated live: explicit cwd → worker rooted there; no cwd over
REST → daemon cwd, not $HOME. Also clears the ccs folder-trust prompt when the
project dir is already trusted (see docs/Worker-Startup-and-Trust.md).
2026-07-15 16:33:46 +02:00
Dai Ha 959f04bc96 docs: worker startup — working directory & the folder-trust prompt
Documents (a) the rule that a worker inherits the PRIMARY's working directory,
never $HOME — resolved from an explicit cwd, else the bridge_spawn caller's PID
(lsof -d cwd), else the daemon cwd; herdr's seam is workspace.create {cwd}, since
agent.start has no cwd; (b) ccs (Claude Code) as the assumed launcher and its
folder-trust model (hasTrustDialogAccepted per project in each instance's
.claude.json), so trusting the project dir once per profile clears the prompt;
(c) that other CLIs have their own startup gates, documented per launcher.
Marks the cwd-inheritance as target design (follow-up), not yet wired.
2026-07-15 16:11:18 +02:00
Dai Ha 7f1b6b3a0d CB-111: multi-profile workers — named backends selectable at spawn
The bridge was hard-wired to one worker profile. Config now takes a 'workers'
map keyed by profile name plus 'defaultWorker'; WorkerService holds the map and
gains spawn(profile) (spawn() uses the default). Selection threads through the
surfaces: REST POST /workers ?profile= / {"profile":…} + GET /profiles; MCP
bridge_spawn {profile?} + new bridge_profiles. Each profile's base_url is
guard-checked independently, so gx10 and ollama can run side by side and you
address each worker by its returned sessionId. Backward-compatible: the legacy
singular 'worker:' block still loads as a one-entry profile map.
2026-07-15 15:54:31 +02:00
Dai Ha ab771ea24e CB-110: cover drop of a delivered-but-unpicked-up turn
Adds the drop test for the delivered/awaitingPickup/!turnObserved state (worker
vanished after delivery but before a WORKING sample) — a gap the other two drop
tests missed. Surfaced by an off-sub worker's code review of CB-110, delegated
through the bridge itself.
2026-07-15 15:30:20 +02:00
Dai Ha a629a7ee73 CB-110: fail an in-flight delegation when its worker vanishes
Companion to CB-109. When a worker disappears mid-turn (pane crash → herdr
*_not_found), the status poller drops the target, which failed only *queued*
messages — a message already DELIVERED is out of the queue, so its send's
rendezvous was left hanging until the 30-min async timeout. Injector.drop now
fires onTurnFailed for a delivered-but-unresolved turn (awaitingCompletion), so
the send resolves as WORKER_FAILED. Reuses the CB-109 resolver path; the failure
reason is neutral to cover both wedge (stuck) and vanish (gone).
2026-07-15 15:23:17 +02:00
Dai Ha 2052929768 CB-109: fail a delegation whose worker wedges in an unknown state
Dogfood found the gap: a turn that dies into an error screen herdr reports as
'unknown' (e.g. the worker hitting API ENOTFOUND) never produces a working->idle
boundary, so CB-106 never fires and the async send rides its full 30-min timeout.

The injector now counts consecutive 'unknown' samples while a delegation is
outstanding; any working/idle sample resets the streak, so only a genuine wedge
(~30s continuous unknown) trips it. It then fires TurnListener.onTurnFailed;
CompletionResolver scrapes the error screen and resolves the send via
Rendezvous.resolveFailure (Kind.FAILED -> Outcome.WORKER_FAILED), surfaced as
async phase=failed / REST status=failed / a [worker failed] MCP note, with the
error context as the reason. This also frees a delivery that wedged before pickup,
which the injectable-only pickup grace could never release.
2026-07-15 15:16:55 +02:00
Dai Ha a97c287aee CB-108: fleet-management MCP tools (bridge_spawn / bridge_list / bridge_stop)
The primary could delegate to a worker but not create or reap one over MCP —
spawning was a raw REST POST /workers. BridgeMcp now adapts WorkerService so a
worker's whole lifecycle runs through MCP: bridge_spawn returns the new worker's
sessionId (for bridge_send) and paneId (for bridge_stop); bridge_list projects the
tracked workers; bridge_stop tears one down. The subscription boundary stays
enforced inside WorkerService (bridge_spawn surfaces a guard breach as a tool error
without touching herdr). Tools are thin static adapters, unit-tested by parity.
2026-07-15 14:22:50 +02:00
Dai Ha 8ed2370fbe CB-107: async fire-and-poll delegation (wait:false + ticket poll)
A caller's MCP client caps a blocking bridge_send at ~60s, but a real delegated
task runs for minutes. sendAsync runs the same blocking send on a background
virtual thread and returns a ticket; poll(ticket) reports pending/done/failed.
Async reuses the blocking path (and its per-target serialization), so it inherits
reply + completion resolution for free. Surfaces: REST POST message wait:false ->
202 {ticket} + GET /tasks/{ticket}; MCP bridge_send wait flag + new bridge_poll.
Terminal tickets are pruned after a TTL so the registry stays bounded.
2026-07-15 14:19:14 +02:00
Dai Ha 5b26caca0c CB-106: completion fallback — resolve a send when the worker's turn ends without bridge_reply
The blocking send previously resolved only on an explicit bridge_reply; a real
delegated task (edit files, run a build) finishes and goes idle without ever
calling it, so the send always timed out. The injector now reports a confirmed
working -> idle turn boundary via a TurnListener; CompletionResolver scrapes the
worker's transcript tail and resolves the awaiting send (Rendezvous.resolveCompletion,
Kind.COMPLETION -> Outcome.COMPLETED_UNREPLIED), surfaced as replySource=transcript
at the REST/MCP edges. Completion is synthesized only from a confirmed turn (a
sampled 'working'), never from the pickup-grace path, so it can't race the explicit
reply or fire on a turn that never ran.
2026-07-15 14:14:04 +02:00
Dai Ha 6988bfe88f CB-103: injector submits the prompt (Enter as a separate keystroke)
A delegated message was delivered into the worker's input box but never
submitted, so no task was ever processed: bridge_send blocked until timeout
while the worker sat idle with the prompt typed but not entered.

herdr's agent.send delivers text as a bracketed paste; a trailing carriage
return in that same call is swallowed as literal newline content, not Enter.
AgentControl.send now emits two keystroke events — the payload, then a
standalone "\r" — so the worker actually submits and runs the task.

Verified live (Claude Code v2.1.210, off-sub worker): full hands-off
bridge_send -> auto-submit -> worker computes -> bridge_reply round trip
returns the reply in ~41s. 62 tests green.
2026-07-15 09:40:18 +02:00
Dai Ha 07722a6007 CB-1xx: worker launch via ccs + inline bridge MCP mount (step 4) + port 8765
Spawn a real worker with 'ccs ltms-local' (profile sets CLAUDE_CONFIG_DIR + off-sub base_url; auto mode preconfigured as defaultMode:auto). bridged appends the bridge MCP mount (--mcp-config, inline JSON) and the reply charter (--append-system-prompt) as launch FLAGS — non-invasive, nothing written to the worker's profile (safer than provisioning its config dir, which would clobber it). Identity is connection-based so the mount is shared. config: worker.mcpUrl. Default bind port 8080 -> 8765. The reply charter is guidance; the send timeout catches a non-cooperative worker. 60 tests green, IDE-clean.
2026-07-15 09:19:31 +02:00
Dai Ha fa570ab32a CB-105: connection-based MCP caller identity (peer PID → herdr pane)
Resolve who is calling an MCP tool from the connection, not a spoofable argument (per docs/MCP-Contract.md). ConnectionIdentity ties the loopback peer PID (LsofPeerPidLookup) to a herdr pane (PaneLocator via pane.list/pane.process_info) → the caller's terminal_id; a caller owning no pane is the primary. bridge_reply now takes only content and resolves the worker from the connection (no sessionId). Live-verified: real PID→pane (contract test), and a non-pane MCP caller correctly gets a workers-only error. Single-host; the token path stays the split-host fallback. 58 tests green, IDE-clean.
2026-07-15 09:19:31 +02:00
kevin 46f87b9cd3 wiki: sync docs with MCP contract (idle-edge turn-done, bridge_list/status naming)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 20:01:23 +07:00
Dai Ha 8ffbfd1f06 CB-105: MCP server (bridge_send/bridge_reply/bridge_status) over the REST core
Streamable-HTTP MCP server (io.modelcontextprotocol.sdk:mcp 2.0.0) mounted on the daemon's Jetty at /mcp, exposing three tools as thin adapters over MessageService/Rendezvous — the primary calls bridge_send/bridge_status, the worker calls bridge_reply. Tool logic in unit-testable static methods (parity tests); SDK owns the wire protocol. Resolves the Jackson 2/3 split by pinning jackson-annotations 3.0-rc5 (works for both our Jackson 2.19 and the SDK's Jackson 3). Live-verified: initialize handshake + tools/list return all three tools. Documents the residual Jackson-3 CVE (loopback, trusted clients). 52 tests green, IDE-clean.
2026-07-14 14:21:11 +02:00
Dai Ha 597ac2e562 CB-104: blocking bridge_send with rendezvous reply (POST /sessions/{id}/message + /reply)
Per-session-serialized blocking send that enqueues via the CB-103 injector (poller delivers) and blocks on a rendezvous resolved by the worker's structured bridge_reply, or a typed 200/202 outcome. No status-polling completion, no terminal scrape. GET /sessions/{id}/status. Reworked from an initial poll+scrape draft after a high-effort review found the polling completion unreliable; all findings fixed.
2026-07-14 14:21:11 +02:00
Dai Ha bc04637694 deps: bump to latest patched (jackson 2.19.0, javalin 6.7.0, jetty 11.0.25, logback 1.5.18)
Clears jetty CVE-2024-8184/CVE-2024-6763. Pins all Jetty modules via jetty-bom (no skew). Documents residual advisories with no upstream fix (jetty-http 11.x EOL, logback config-file CVEs, jackson WS-2026-0003) as accepted for this loopback daemon. CLAUDE.md: validate CVEs with the jetbrains analyzer (intellij-index is stale after pom edits).
2026-07-14 14:21:11 +02:00
kevin c7b58f2195 docs: add Team.md under docs/
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 19:14:25 +07:00
kevin 01cfda7965 Add MCP design doc 2026-07-14 12:30:20 +07:00
141 changed files with 20621 additions and 422 deletions
+18
View File
@@ -0,0 +1,18 @@
{
"name": "claude-bridge",
"description": "Tooling for orchestrating a fleet of delegated coding agents through the bridged MCP gateway.",
"owner": {
"name": "LTMS"
},
"plugins": [
{
"name": "claude-bridge",
"source": "./plugin",
"description": "Make a project bridge-ready: mount the bridged MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.1.0",
"author": {
"name": "LTMS"
}
}
]
}
+137
View File
@@ -0,0 +1,137 @@
---
name: implementer
description: Implementer-role procedure for a bridged worker — verify your worktree, implement the scope, commit, push, open your own PR, and hand off the PR URL. Load this when the lead delegates you an implementation task over bridged.
---
# Implementer worker — procedure
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
never merge, never commit `.mcp.json` or `wiki/`) is in **`CLAUDE.md` → Bridge communication →
Worker** and already applies. This skill is only the *implement-and-hand-off procedure*.
You run in an **isolated git worktree on your own branch** — a full peer of the primary (same repo,
`CLAUDE.md`, skills), differing in the model behind you and the branch you sit on. Your MCP surface
is **only what your launcher mounted** (the bridge): the primary's IDE and forge servers are not
yours, and the worktree's `.mcp.json` is deliberately emptied so you cannot inherit them.
The worktree model is documented in [`docs/Worker-Git-Workflow.md`](../../../docs/Worker-Git-Workflow.md).
## 1. Confirm where you are — then never leave
Before touching anything:
```bash
git rev-parse --show-toplevel # your worktree root — NOT the primary's main tree
git branch --show-current # your dedicated branch: worker/<ticket>-<nonce>
git status # should be clean at the start
```
Do **all** work here, on this branch. Never `git checkout main`, never rebase onto or push to
`main`. The branch is your isolation — respect it.
**Every path you read, edit, or build is relative to that root.** Work from `$PWD`; if a tool, a
brief, or your own memory hands you an absolute path, check it starts with your worktree root
before you touch it, and stop if it doesn't. An absolute path pointing anywhere else is the
primary's checkout — editing there while building here means **every build you run is of code that
does not contain your changes**, and it passes while your work goes nowhere. This has happened:
a worker made all 59 of its edits in the primary's tree and never noticed.
```bash
test "$(git rev-parse --show-toplevel)" = "$PWD" || cd "$(git rev-parse --show-toplevel)"
```
## 2. Implement
- Implement exactly the scope the lead named. Keep the diff focused; note anything out of scope
in your reply instead of widening it.
- Match the surrounding code's style, naming, and idioms.
**Acceptance criterion — a green build, quoted.** Your work is not done until this passes *inside
your worktree*:
```bash
cd "$(git rev-parse --show-toplevel)/bridged" && mvn clean install
echo "exit=$?"
```
Read its **full** output — never pipe it through `tail`/`head`/`grep`, which hide a failure behind
a zero exit. Then quote the real `Tests run: … Failures: … Errors: …` line and the
`BUILD SUCCESS`/`FAILURE` verbatim in your reply. If it does not go green, say so with the actual
error; a failing build honestly reported is a usable result, a claimed-green one is not. You have
no IDE MCP tools, so `mvn` is your only verification — never claim a check you had no way to run.
## 3. Commit
```bash
git add <the files you changed> # explicitly — never `git add -A` / `git add .`
git commit -m "<ticket>: <clear one-line summary>"
```
`.mcp.json` is neutralized and `--skip-worktree` in your worktree — never `git add` it, and never
"restore" it from the primary's copy. Same for `wiki/` (a submodule with its own remote).
## 4. Push
```bash
git push -u origin HEAD
```
Push is over SSH as the same user — no extra credential needed. Never force-push over anything
you did not create.
## 5. Open your own PR to `main`
Via the gitea REST API. The daemon injected a **repo-scoped token** (`GITEA_TOKEN`) and the forge
host (`GITEA_HOST`) into your env for exactly this — the token can create a PR but **cannot
merge**.
```bash
API="${GITEA_HOST%/}/api/v1/repos/lms/claude-bridge/pulls"
BRANCH="$(git branch --show-current)"
curl -sS -X POST "$API" \
-H "Authorization: token ${GITEA_TOKEN}" \
-H "Content-Type: application/json" \
-d "$(cat <<JSON
{"head": "${BRANCH}", "base": "main",
"title": "<ticket>: <concise change summary>",
"body": "<what changed and why; reference the ticket; note tests run and their result>"}
JSON
)"
```
The response JSON carries `"html_url"` — that is your PR URL. On a non-2xx, read the error body,
fix it if the cause is yours (e.g. branch not pushed yet), and report the failure rather than
inventing a URL. If `GITEA_TOKEN` is unset your profile was not granted PR-create: push the branch
and report its name so the lead opens the PR.
## 6. Hand off — what goes in `bridge_reply`
The reply is the entire handoff; the lead cannot see your terminal.
```
PR: <html_url from step 5, or "not created: <reason>" + branch name>
branch: <your branch>
root: <git rev-parse --show-toplevel — proves you worked in your own worktree>
files: <worktree-relative paths you changed>
build: <the verbatim "Tests run: …" and BUILD SUCCESS/FAILURE lines — or "not run: <why>">
summary: <2-3 lines: what you implemented and any caveat the reviewer needs>
```
```mermaid
sequenceDiagram
autonumber
participant L as Lead
participant I as Implementer (you)
participant G as git / gitea
L->>I: delegated task (you are in a worktree on your branch)
I->>I: "implement here — every path under $PWD"
I->>I: "mvn clean install in this worktree, unpiped, until green"
I->>G: git commit (never .mcp.json / wiki)
I->>G: git push -u origin HEAD
I->>G: POST /pulls (GITEA_TOKEN) — open PR to main
G-->>I: html_url
I->>L: bridge_reply(PR url, branch, files, tests)
Note over L,G: the lead reviews the PR and merges on green — you never merge
```
*The implement turn: work in the worktree, commit → push → open the PR, hand off the URL.*
+51
View File
@@ -0,0 +1,51 @@
---
name: reviewer
description: Reviewer-role procedure for a bridged worker — how to work a review scope and the exact shape of the finding to report. Load this when the lead delegates you a code review over bridged.
---
# Reviewer worker — procedure
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and already applies. This
skill is only the *review procedure*: how to work the scope, and the exact shape of what you
send back.
## 1. Read the whole scope before you judge
The delegation names your scope — a file, a diff, a PR, a function. **Read all of it first.**
A review that fires on a snippet misses the caller that makes it safe (or the one that makes it
a bug). Reviewing part of the scope and guessing the rest is the most common way a reviewer is
wrong.
## 2. Stay in the scope
- Review **only** what you were assigned. Something elsewhere looks wrong? One line in your
reply — do not go hunt it. Wandering is how two reviewers report the same thing and neither
covers what it was given.
- Do **not** edit files or run the build. You review; the owner acts.
## 3. Reach for `bridge_ask` only for a genuine fork
Ambiguous requirement, a missing acceptance criterion, "intended or a bug?", or two defensible
fixes with different consequences — those are the lead's call, and guessing produces a
confident-but-wrong finding. Anything you could settle by reading more code is yours to settle.
## 4. The finding — what goes in `bridge_reply`
Report the **single most important** real issue in the scope, in these four lines, under
~90 words:
```
1. <path>:<line>
2. issue: <one sentence — what is wrong and why it matters>
3. fix: <one line — the concrete change>
4. severity: high | medium | low
```
- **Nothing real after reading?** Reply `NO ISSUE` and one line saying why. A clean review is a
valid result; a fabricated issue is worse than none.
- **Severity:** `high` = wrong result, data loss, security, or a hang/crash on a real path ·
`medium` = a real bug on an edge path, or a correctness risk under load/concurrency ·
`low` = clarity, a latent foot-gun, or a smell with no current failure.
- Be specific and verifiable: a line number and a one-line repro beat an adjective. If you can't
point at where it goes wrong, you haven't found it yet.
+107
View File
@@ -0,0 +1,107 @@
name: CI
on:
push:
branches: [main]
pull_request:
branches: [main]
jobs:
build:
runs-on: ubuntu-latest
steps:
# The wiki submodule is docs only and is not needed to build — leave it unfetched so CI
# does not depend on the wiki repo being reachable.
- uses: actions/checkout@v4
# The runner image ships an older default-jdk; bridged sets maven.compiler.release=25, so
# provision the JDK explicitly rather than apt-installing whatever "default" means today.
- name: Set up JDK 25
uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '25'
cache: maven
# setup-java provisions the JDK only — it does NOT install Maven, and the runner image has
# no mvn on PATH (a bare `mvn` exits 127). Install it separately. apt pulls a default JRE as
# a dependency; JAVA_HOME from setup-java still wins, which the version check below proves.
- name: Install Maven
run: |
apt-get update && apt-get install -y --no-install-recommends maven
mvn -version
- name: Build and test
working-directory: bridged
# This IS the mock-socket surface CB-503 asks for: the pom's `default-excludes` profile
# already sets excludedGroups=contract, so the @Tag("contract") tests — which need a live
# herdr socket and a RabbitMQ container — are excluded without any flag here. Everything
# that runs does so against the fake UDS herdr and fake ccs/claude stubs.
run: mvn -B clean install
# Deliberately NOT actions/upload-artifact: this Gitea instance presents as GHES, and
# @actions/artifact v2+ (i.e. upload-artifact@v4) refuses to run there —
# "GHESNotSupportedError ... not currently supported on GHES", which red-Xes an otherwise
# green build. Since the artifact could not be retrieved anyway, dump the failing tests into
# the log instead, where they are actually readable.
- name: Failing test output
if: failure()
working-directory: bridged
run: |
for f in target/surefire-reports/*.txt; do
[ -f "$f" ] || continue
grep -qE "Failures: [1-9]|Errors: [1-9]" "$f" && { echo "===== $f ====="; cat "$f"; }
done
exit 0
# CB-521 — actually run the AMQP contract test in CI, against a REAL broker. The broker is a
# RabbitMQ SERVICE CONTAINER, not Testcontainers-with-Docker: the runner image has no Docker, so
# AmqpReplyInboxContractTest reads AMQP_URI (set below to the service's network alias) and binds
# straight to it — no Docker, no skipped tests. This separation (build job hermetic and
# Docker-free; contract job broker-provided) is deliberate — see the default-excludes/contract
# profiles in bridged/pom.xml. `setup-java` provides the JDK only; Maven is installed separately,
# exactly as in the build job above.
contract:
runs-on: ubuntu-latest
services:
rabbitmq:
image: rabbitmq:3.13 # same AMQP 0-9-1 engine the local Testcontainers fixture uses
env:
RABBITMQ_DEFAULT_USER: guest
RABBITMQ_DEFAULT_PASS: guest
ports:
- 5672:5672
env:
# Service containers are reachable from the job by their network alias on their internal port.
AMQP_URI: amqp://guest:guest@rabbitmq:5672
steps:
- uses: actions/checkout@v4
- name: Set up JDK 25
uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '25'
cache: maven
- name: Install Maven
run: |
apt-get update && apt-get install -y --no-install-recommends maven
mvn -version
# The `contract` profile clears the default-excludes group, so the @Tag("contract") AMQP test
# runs against the RabbitMQ service container (AMQP_URI). Pinned to the one contract test to
# avoid re-running the unit suite already covered by the `build` job.
- name: Contract tests
working-directory: bridged
run: mvn -B -Pcontract test -Dtest=AmqpReplyInboxContractTest
- name: Failing test output
if: failure()
working-directory: bridged
run: |
for f in target/surefire-reports/*.txt; do
[ -f "$f" ] || continue
grep -qE "Failures: [1-9]|Errors: [1-9]" "$f" && { echo "===== $f ====="; cat "$f"; }
done
exit 0
+207
View File
@@ -1,7 +1,205 @@
# claude-bridge — project instructions
## Bridge communication (enforced — read this first)
> **Canonical block.** Everything down to §Layering is the portable bridge charter, copied verbatim
> into every project that mounts the bridge MCP. Keep it byte-identical with the template in the
> wiki ([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable
> CLAUDE.md block*); improvements go to the template first, then out to each project. Anything
> specific to *this* repo lives under §Project addendum below, never inline above it.
If no `bridge_*` MCP tools are mounted in this session, this section does not apply — skip it.
`bridged` is the **sole communication gateway** between agents here. The orchestrating session (the
**primary**) and every delegated peer (a **worker**) mount the *same* MCP server and talk only
through its `bridge_*` tools. No session addresses a peer, a broker, or the network directly.
### Which role am I? — settle this before acting
**Both roles read this file.** A worker runs in a git worktree of this same repo, so it inherits
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
**Call `bridge_whoami`.** It returns `{"role":"primary"}` or `{"role":"worker","sessionId":…,
"profile":…,"worktree":…,"branch":…}`, resolved by the daemon from your connection — unforgeable,
and the same resolution its authorization gate uses. Don't infer what you can ask.
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
fires: the reply charter in your system prompt (*"You are an off-subscription worker in the
claude-bridge fleet"*) ⇒ **worker**; bridge tools prefixed `mcp__bridge__*` ⇒ **worker** (the
launcher fixes that mount name; a primary's mount is named by whoever wrote its `.mcp.json`, so it
varies); `ANTHROPIC_BASE_URL` set ⇒ **worker** (Claude-model workers run on a clean env, so its
*absence* proves nothing). **Still unsure ⇒ act as a worker.** The two mistakes are not symmetric: a
primary acting as a worker is refused by the authorization gate — loud and self-correcting — while a
worker acting as the primary ends its turn with no `bridge_reply`, and the sender silently receives
nothing. Fail toward the recoverable error.
### Invariants — both roles, no exceptions
1. **Never set, export, or forward `ANTHROPIC_BASE_URL`** (or `ANTHROPIC_AUTH_TOKEN`). The primary
stays on subscription; only the bridge puts a worker off it, at spawn. Mounting the bridge must
never move a session across that boundary.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/send/drain are primary-only; reply/ask are
worker-only-and-only-as-itself. A call outside your role is refused, not queued.
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
re-send because a call looks slow — the bridge delivers when the peer is `idle`/`blocked`.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
### Primary (lead) — run this on every task, in order
**Delegate by default — that is the job.** With the bridge mounted you are an orchestrator on a
metered subscription, and workers are cheap, parallel, and disposable. The default answer to "who
does this?" is **a worker**, not you. Reach for `bridge_send` before you reach for `Edit`. The steps
below are the procedure — run them in order, every task, not only the big ones.
0. **Know your role** — `bridge_whoami`, once per session, before anything else.
1. **Split.** Write the unit list. Every unit carries: scope · the files or PR in question ·
acceptance criteria · exactly what to report back. A unit with no acceptance criteria is not
ready to delegate — refine it or keep it.
2. **Gate each unit** on one question: **"can I write a brief good enough for a worker to
succeed?"** — *not* "could I do this faster myself?" (usually you could; doing it yourself costs
your context and your subscription, while a wasted worker turn costs a worker turn). Yes ⇒
delegate. The keep-list is closed: the conversation with the user, decomposition and planning,
the final judgment call, verification, merges, and anything that depends on context only you
hold. Nothing else is yours by default.
3. **Spawn every delegated unit first** — `bridge_spawn{profile, worktree:true, ticket}`, one per
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model and cost, not in
tier, so the default is rarely what you want.
4. **Then send them all** — `bridge_send{sessionId, content, wait:false}`. Line 1 of every brief is
`Load the <name> skill.` naming the worker's playbook; those skills are opt-in and that line is
what makes them reliable. Where the project ships no such skill, spell the procedure out in the
brief instead. The brief is self-contained — the worker sees your message and the repo, nothing
of your context, your plan, or your screen.
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`bridge_status`, never by reading its terminal.
6. **Verify yourself.** Re-run the build and the checks. A worker mounts only the bridge MCP and
cannot run your other tooling, and a piped command (`… | tail`) hides failures behind a zero
exit — never promote a worker's "clean" to a fact.
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
as it lands; don't wait for the last implementer. Under ~50 changed lines, skip the fan-out and
read it yourself.
8. **Adjudicate, merge, tear down — yours alone.** Read the diff yourself: fully if it is small,
targeted at the reported findings and the risky paths if it is large. Reviewer findings direct
your attention; they never substitute for it. Then merge, then `bridge_stop{paneId}`.
**Steps 3 and 4 are separate on purpose** — spawning and sending in one loop is how parallel work
silently becomes serial, and it is the most common way this layer is wasted. For the same reason,
prefer `wait:false` + `bridge_poll` for anything non-trivial: a blocking `bridge_send` is capped by
*your own* MCP client call timeout (~60s), well below the task's real runtime.
**Delegating does not delegate responsibility.** Workers open PRs; you are the gate. Never delegate
the merge — and merging on a reviewer's word is delegating it by proxy.
| Intent | Tool |
|---|---|
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a worker | `bridge_spawn{profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` · one worker's state: `bridge_status{sessionId}` |
| Delegate (blocking) | `bridge_send{sessionId, content}` |
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
| Answer a worker's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
| Tear down | `bridge_stop{paneId}` |
### Worker — the turn contract
1. **Load the playbook skill the lead named** before doing anything else.
2. **Do the assigned scope only.** Note anything you spot outside it in one line; don't go hunt it.
3. **`bridge_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
defensible fixes, "bug or intended?"). It blocks and you resume the *same* turn with the answer.
Don't ask what you could decide yourself.
4. **End the turn with exactly one `bridge_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `bridge_reply` ⇒ the sender gets nothing and the exchange stalls.
5. **Report honestly.** State only what you actually ran and its real output, including failures.
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
so never claim the result of a check you had no way to run.
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
### Where each rule lives (don't duplicate — extend the right layer)
| Layer | Scope | Reaches |
|---|---|---|
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every worker, at launch, every peer kind |
| **this section** | protocol + orchestration policy | primary **and** every Claude worker — tracked in git, so worktrees inherit it |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a worker told to load one |
| the bridge's own docs | design detail, flows, error model | on demand |
A rule belongs in **exactly one** layer — the outermost one that must obey it. Peers that don't read
`CLAUDE.md` (non-Claude adapters) get the charter only, so any rule *they* must obey belongs in the
charter, not here.
## Project addendum — claude-bridge (not part of the canonical block)
- **This repo is the bridge.** The daemon is `bridged`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/BridgeMcp` (tools), `auth/Authz` (the role table),
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
`reviewer` (scoped review → one structured finding). Name one in every delegation.
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
session's context.
### The prompt is part of the product — update it with the code (mandatory)
This repo *is* the bridge, so the canonical block above is not documentation about someone else's
system: it is the instruction surface this codebase ships. **Every change here must end by asking
whether the block still tells the truth.** A code change that silently invalidates it is an
incomplete change — the agents reading it have no other source.
Before you call any work done, check the row that matches what you touched:
| You changed… | Re-read and update… |
|---|---|
| a `bridge_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
| `Authz` / the role table | invariant 3, and the primary-only vs worker-only claims |
| `ConnectionIdentity` / how a caller is resolved | the `bridge_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__bridge__*`), and the layering table's top row |
| the injector / status gating | invariant 4 |
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
| **anything an operator can use, configure, or observe** — an MCP tool, a `bridged.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
That last row is not bookkeeping. Chapters 1–10 answer *how is this built* and *why this way*;
none of them has a home for *what can it do and how do I turn it on*, so for twenty tickets a
shipped capability landed nowhere and the Roadmap went on claiming the stage was finished. The
*why* line is the one that matters — without it a decision gets re-litigated from scratch a month
later. Internal contract changes go to `wiki/9-Implementation.md` instead; test and coverage work
is a Roadmap line. A change that touches none of the three earns no entry, and that is a normal
outcome rather than an omission.
Then **propagate**: the block in this file and the template in the wiki
([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
block*) must stay byte-identical, and other projects carrying the block need the same edit. Verify
rather than trust:
```bash
python3 - <<'PY'
import pathlib
c = pathlib.Path("CLAUDE.md").read_text()
w = pathlib.Path("wiki/7-Use-Cases.md").read_text()
S, E = "## Bridge communication (enforced", "## Project addendum — claude-bridge"
block = c[c.index(S):c.index(E)].rstrip() + "\n"
i = w.index("```markdown\n") + len("```markdown\n")
print("in sync:", w[i:w.index("\n```\n", i) + 1] == block)
PY
```
## IDE MCP tools & validation workflow (enforced)
> **Primary only.** Workers have no IDE MCP mount — if you are a worker, skip this section and
> report the build/test output you actually ran (see §Bridge communication → Worker).
Two IDE MCP servers are connected: **intellij-index** (semantic code intelligence) and
**jetbrains** (file problems, reformat, debugger). IntelliJ has multiple projects open; our
module is **`bridged`**. Always pass these to IDE MCP tools:
@@ -20,6 +218,15 @@ module is **`bridged`**. Always pass these to IDE MCP tools:
test run). A per-file-clean file can still break the build or another module. This is the
whole-project gate before declaring work done or committing.
**Whenever dependencies change (or a `pom.xml` edit), validate CVEs with
`jetbrains get_file_problems{filePath: "bridged/pom.xml"}`** — its Mend.io check reflects the
dependencies on disk. (Note: `ide_diagnostics` / intellij-index does NOT re-resolve dependencies
after a pom edit without a full Maven reimport, so it reports stale CVE results — don't trust it
for this.) Treat a CVE warning like any other: bump to a patched version and confirm
`mvn clean install` still passes. If the latest available version is still flagged (EOL line,
"insufficient information", or config-file-only advisories), document it as accepted in the pom
rather than chasing a fix that doesn't exist.
### Use IDE MCP tools for navigation, refactoring, and diagnostics only
- **Navigate (prefer over Grep/Read for symbols):** `ide_find_definition`, `ide_find_class`,
+38 -5
View File
@@ -50,8 +50,9 @@ flowchart LR
plain daemon (no Anthropic quota), so it may poll/subscribe freely.
- **One gateway (unified MCP setup):** `bridged` is the **sole communication path** for every
Claude session. Primary and workers each mount it as an MCP server (one `claude mcp add`
line, same on both) and talk over MCP tools — `bridge_send` / `bridge_reply` / `bridge_ask` /
`bridge_status`. **No Claude session ever addresses a broker, a peer, or the network
line, same on both) and talk over MCP tools — `bridge_send` / `bridge_reply` /
`bridge_status` (with `bridge_ask` planned for the blocked-worker path). **No Claude session
ever addresses a broker, a peer, or the network
directly**; any queue is `bridged`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so
mounting the bridge is subscription-safe by construction.
- **How the primary consumes a reply:** a single **blocking MCP call** (`bridge_send`);
@@ -85,6 +86,38 @@ Gitea wiki.
## Status
🟢 Design — herdr-centric **`bridged`** message server selected as the primary approach
(2026-07-11), superseding the AgentAPI plan (2026-07-08). AgentAPI retained as fallback
injector.
🟢 **Implemented & dogfooded** — the herdr-centric **`bridged`** message server is built and in
real use: an Opus primary delegates tasks to off-subscription workers that reply through the
bridge (code reviews delegated this way have produced committed bug fixes). Selected as the
primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a
fallback injector.
**Shipped** (Java 25 · Maven · 266 unit/acceptance tests green; the live-herdr and broker contract
tests run separately via `mvn test -Pcontract`):
- **Core gateway** — herdr socket client (contract-tested vs live 0.7.0); guard-checked worker
spawn with `ANTHROPIC_BASE_URL` injected only into the worker's env; status-gated injector;
blocking `bridge_send` with reply rendezvous; MCP server as a thin adapter over the REST core.
- **MCP tools** — `bridge_send` / `bridge_reply` / `bridge_status` (messaging) and `bridge_spawn`
/ `bridge_list` / `bridge_stop` / `bridge_profiles` / `bridge_poll` (fleet). Caller identity is
connection-based (loopback peer PID → herdr pane), so the same mount serves primary and workers.
- **Delivery reliability** — completion fallback (a confirmed `working→idle` turn resolves a
send); async fire-and-poll (beats the caller's MCP call timeout for long tasks); and failure
detection for wedged (`unknown`), vanished, and never-ready workers so a send never hangs.
- **Fleet** — multiple worker profiles, each with an independent base_url guard check; workers
inherit the primary's working directory (never `$HOME`); a readiness gate holds delivery until
a worker's Claude has connected the bridge MCP (no paste lost into its boot window).
- **Blocked-worker path** — `bridge_ask` reverse rendezvous: a worker pauses its delegated turn to
ask the primary and resumes the *same* turn with the answer (CB-205).
- **Session lifecycle** — session manager with spawn/reuse/recycle, `idle_ttl` reaper, `context_cap`,
and graceful drain on shutdown (CB-301/CB-303); per-worker git worktrees on their own branch with
a config-parity overlay, so parallel implementers never stomp each other (CB-301-ext).
- **Reliable worker→primary delivery** — a durable `ReplyInbox` (in-memory by default, AMQP/LavinMQ
for cross-restart durability) holds a reply that arrives with no open send, and an active
status-gated push loop nudges the primary to drain it (CB-307).
- **Pluggable peers** — a `PeerLauncher` SPI with two in-tree adapters, `claude-code` and `opencode`,
routed by a `kind:` discriminator (CB-401/CB-402).
**Next** (see the [roadmap](wiki/8-Roadmap.md)) — Stage 5 hardening (auth/TLS, `/metrics`, CI,
service supervision, per-session authz + audit), then cross-host: CB-308 multi-host federation and
CB-500 multi-tier coordination.
+3
View File
@@ -5,6 +5,9 @@ dependency-reduced-pom.xml
# Local runtime config (copy from bridged.example.yaml)
bridged.yaml
# CB-505 audit trail + daemon stdout/stderr — runtime records, never source
logs/
# Editor / OS
*.iml
.idea/
+227 -16
View File
@@ -3,31 +3,242 @@
# bridged is the sole gateway between primary/worker Claude sessions and herdr.
# It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL.
# REST + MCP listen address. Keep it on loopback — bridged is same-host in Stage-1.
# REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token`
# below — bridged REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
bind:
host: 127.0.0.1
port: 8080
port: 8765
# API authentication (CB-501). Governs how a caller that is NOT an on-host worker pane proves it
# is the primary. Worker identity never depends on this: a loopback peer PID that maps to a herdr
# pane is unforgeable and is always honoured, so turning auth on cannot lock the fleet out.
#
# mode: loopback-trust → DEFAULT, and the historical behaviour: any loopback caller that is not
# a worker is the primary, no credential needed. Sound ONLY because the
# OS refuses remote connections to a loopback socket.
# mode: token → such a caller must send `Authorization: Bearer <token>`; without it it
# is anonymous and authorized for nothing. REQUIRED for a non-loopback
# bind — the daemon fails fast otherwise, because "unauthenticated ⇒
# primary" on a reachable port would hand spawn/stop/send to anyone.
# tokenEnv → host env var holding the token (never the literal value). Default
# BRIDGED_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
#
# TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and
# let it own certificate lifecycle, e.g.
# location / { proxy_pass http://127.0.0.1:8765; proxy_set_header Authorization $http_authorization; }
# The broker link gets TLS from its own URI (amqps://…) — see `broker` below.
# auth:
# mode: token
# tokenEnv: BRIDGED_API_TOKEN
# Optional pinned primary terminal (CB-307). Names the herdr pane the PRIMARY itself runs in:
# a caller whose connection maps to this pane resolves as the primary (no credential needed —
# the pane mapping is as unforgeable as a worker's), and reply nudges are pushed to it.
# REQUIRED when the primary runs inside a herdr pane — without it the pane match reads the
# primary as a worker and refuses spawn/send/stop. Get the id from bridge_whoami; re-pin if
# the primary moves panes.
# primary:
# terminal: term_0123456789abcd
# pushReminders: 5 # max nudges before giving up (default 5)
# pushBackoffMs: 15000 # delay between nudges (default 15000)
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
# How a worker session is spawned. Stage-1 uses the existing ccs `ltms-local`
# profile, whose .claude.json routes to the gx00 vLLM below.
worker:
profile: ltms-local
baseUrl: http://gx00.gw:8000 # the gx00 vLLM (models: coder / deepseek-v4-flash)
model: coder
# Placement: each worker lands in its OWN tab inside a dedicated worker space, so it
# never splits or clutters your real work spaces. Use `pane` for the legacy behaviour
# (split the currently-focused tab).
placement: tab # tab | pane
workspace: bridged-workers # the dedicated worker space (found-or-created, shared)
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
# How worker sessions are spawned. Define one or more named profiles (backends) under
# `workers`; each key is the profile name (also the ccs profile). `defaultWorker` picks
# which one a no-argument spawn uses (bridge_spawn with no profile / POST /workers).
#
# Shared knobs (placement/workspace/tabLabel) can be repeated per profile; they usually match.
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
# Use `pane` for the legacy behaviour (split the focused tab).
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
# (--append-system-prompt) as launch flags; nothing is written to the profile.
# tokenEnv → host env var holding the worker's auth token (value never stored in config);
# omit for a backend that needs no token (e.g. a local ollama).
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
# cwd on an MCP spawn, else the daemon's cwd — never $HOME. See
# docs/Worker-Startup-and-Trust.md.
# configDir → CLAUDE_CONFIG_DIR for the worker, so it inherits that profile's
# skills/MCP/hooks. Omit to leave the worker on the host default.
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
# worktree sees the same local config (CB-301-ext). Omit for the default set:
# [.claude/settings.local.json, .env, .envrc].
#
# Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher
# mounts — the bridge, and nothing else. Replicating the primary's MCP config
# handed a worker the primary's IDE servers, which are bound to the primary's
# checkout, so its navigation returned paths OUTSIDE its own worktree: one
# worker made all 59 of its edits in the primary tree while compiling its
# worktree, and every build it ran was of code that did not contain them.
# bridged neutralizes a provisioned worktree's .mcp.json for this reason;
# listing it here would copy the primary's back over that.
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
# Opt-in by design — omit and the worker gets no PR-create grant (push over
# SSH is unaffected). The token value itself is never stored in this file.
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
# env → extra environment for this profile's workers, as a literal key/value map
# (CB-511). Use it to give workers a toolchain.
#
# A worker's environment does NOT come from your shell. bridged hands herdr an
# explicit env map and herdr merges it into ITS OWN process env — so before
# CB-511 a worker inherited whatever PATH the herdr server happened to be
# started with, which on a long-lived herdr can predate your toolchain entirely
# and leave workers unable to run `mvn` or `java` at all.
# bridged now propagates ITS OWN PATH to every worker by default; set `env:`
# only to override that or add more (JAVA_HOME, …). Since the default is the
# daemon's PATH, make sure the daemon is started with a good one — see the PATH
# lines in deploy/dev.ltms.bridged.plist and deploy/bridged.service.
#
# Adapter-owned variables always win over `env:`: ANTHROPIC_BASE_URL and the
# rest of the ANTHROPIC_*/CLAUDE_* wiring are applied after it, so an `env:`
# entry cannot repoint a worker past the SubscriptionGuard — which is checked
# against `baseUrl` alone.
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
workers:
gx10: # ccs profile name (NOT a hostname)
kind: claude-code # which adapter spawns this profile (default; may omit)
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
model: coder
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
weight: 0.5 # relative selection weight for placement: weighted
maxLoad: 2 # max live workers on this profile (omit for unlimited)
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
gx11: # a second backend, so `placement: weighted` has a choice
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}"
mcpUrl: http://127.0.0.1:8765/mcp
argv: ["ccs", "gx11"]
weight: 0.5
maxLoad: 2
# Pin an auto-compact window BELOW the served model's context ceiling. The global
# ~/.claude/settings.json value is shared by every ccs instance and the primary, so the
# per-profile override belongs here. Equal to the ceiling means auto-compact never fires
# before the server rejects the prompt, which kills a worker mid-turn (CB-523).
env:
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "280000"
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
# The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a
# `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude.
#
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → bridge_send →
# structured bridge_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
# and need NO credentials — check `opencode models` for the current free list, since the names
# change. That also makes the worker off-subscription by construction.
# opencode-free:
# kind: opencode
# model: opencode/north-mini-code-free # `provider/model` selector, injected as `-m`
# placement: tab
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
#
# CB-508: point an opencode profile at your OWN OpenAI-compatible endpoint (local vLLM, llama.cpp,
# LM Studio, TGI…) instead of opencode's gateway. Setting `baseUrl` on a `kind: opencode` profile
# makes the bridge emit a custom `provider` block into the generated opencode.json — opencode has
# no ANTHROPIC_BASE_URL seam, so this is how the endpoint is pinned.
# baseUrl → a bare host:port gets `/v1` appended (where these servers mount the API); a URL that
# already has a path is used verbatim, so a custom mount point still works.
# model → MUST be "<provider>/<model>". The provider half names the generated block; the model
# half must match an id the server reports at /v1/models. One field drives both the
# declaration and the `-m` flag, so they cannot drift apart. A bare model name with a
# baseUrl set is rejected at spawn rather than silently using the default gateway.
# tokenEnv → optional; its value becomes the provider apiKey. Most local servers ignore the key,
# so a placeholder is used when unset (the AI SDK still requires a non-empty one).
# NOTE: no `guard` entry is needed even with a baseUrl set. The SubscriptionGuard exists to stop a
# worker borrowing the primary's Anthropic subscription, and an opencode process has no Anthropic
# credential path at all.
# opencode-local:
# kind: opencode
# baseUrl: http://127.0.0.1:8000
# model: local-vllm/deepseek-v4-flash
# placement: tab
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
placement: weighted
defaultWorker: gx10
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
# must carry none. Grounded in ltms-local's real endpoints.
# must carry none. Every profile above must have its host listed here.
guard:
offSubscriptionHosts:
- gx00.gw
- ollama.ltms.dev
- gx01.gw
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
# unknown keys are ignored, so a snake_case key would be silently dropped (default kept).
# spawnReadyTimeoutMs: 20000
# spawnReadyPollMs: 300
# Worktree provisioning root (CB-301-ext). Where per-worker git worktrees are checked out so
# each worker owns an isolated branch instead of sharing the primary's tree. Omit to default
# to a sibling directory of the repo root.
# worktreeRoot: /Users/me/src/.bridged-worktrees
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
# contextCap → force-release a session after this many delegated turns
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
# lifecycle:
# idleTtlSeconds: 300
# contextCap: 10
# drainTimeoutSeconds: 5
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
# Set a broker uri to swap in the AMQP-backed inbox: worker replies with no open send are held
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
# broker:
# uri: amqp://guest:guest@127.0.0.1:5672
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
# pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop
# stops as soon as the primary's inbox is empty.
# terminal → pin the primary's herdr terminal id. Omit to learn it from the connection on
# the first orchestration-side MCP call (the normal case). An off-host or
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
# degrades to pull; the reply is still never lost.
#
# REQUIRED (CB-522) if the primary itself runs inside a herdr pane. Caller
# identity resolves a loopback PID to its herdr pane, and PaneLocator scans
# EVERY pane — not just bridged-spawned ones — so such a primary is otherwise
# classified as a WORKER and refused SPAWN/SEND/STOP. That failure is
# self-locking: the learned terminal is populated by the very orchestration
# calls being refused, so only this pinned value can break the cycle. Read the
# id off bridge_whoami (it reports the current terminal even while
# misclassified) and re-pin whenever the primary moves panes.
# pushReminders → max nudges before giving up (default 5)
# pushBackoffMs → delay between nudges in ms (default 15000)
# primary:
# terminal: term_65619bd6174568
# pushReminders: 5
# pushBackoffMs: 15000
+142
View File
@@ -0,0 +1,142 @@
# CB-307 — Active push-to-primary + reminder loop (the reliability layer)
**Status:** design (2026-07-19). Builds directly on the shipped durable landing zone
(`AmqpReplyInbox`, main `2bc5f3a`, dogfooded live). gitea #5.
## Why this exists
Stage 2 gave a worker→primary reply a **durable place to wait** when no `bridge_send` is
open: it lands in `agent.<target>.inbox` on the broker and survives a daemon bounce. But
delivery is still **pull** — the primary only sees the reply if it happens to call
`bridge_poll(target)` / `GET /sessions/{id}/replies`. A reply can sit indefinitely while
the primary works on something else.
This layer makes delivery **active**: the bridge *pushes* a nudge to the primary the moment
a reply lands, and keeps reminding (bounded) until the primary drains it. At-least-once,
dedup by `msgId`, and — critically — it never loses the reply even if every push fails,
because the durable inbox is the backstop.
## The hard constraint it works around
The bridge is an MCP **server**; the primary is an MCP **client**. A server cannot call
into a client. So "push to the primary" cannot be an MCP response — it needs a *sideband*
channel. The chosen channel: **inject a synthetic user-turn into the primary's own herdr
terminal pane** — the same mechanism the bridge already uses to deliver tasks to workers,
pointed at the primary's pane instead.
```mermaid
flowchart LR
W["worker"] -->|"bridge_reply (no open send)"| MS["MessageService.reply"]
MS -->|"inbox.publish"| INBOX[("agent.&lt;target&gt;.inbox<br/>(durable, LavinMQ)")]
MS -->|"notify"| LOOP["ReplyPushLoop"]
LOOP -->|"status-gated inject"| PANE["primary's herdr pane"]
PANE -->|"primary drains"| DRAIN["bridge_poll(target)<br/>= peek + ack"]
DRAIN -->|"inbox now empty"| LOOP
LOOP -.->|"still non-empty →<br/>re-inject on backoff"| PANE
classDef store fill:#2c5282,stroke:#1a365d,color:#ffffff;
class INBOX store
```
*Figure 1 — a reply lands in the durable inbox; the push loop nudges the primary's pane;
the primary's drain acks it; a still-full inbox triggers a bounded re-nudge.*
## Three increments
### Increment 1 — learn & store the primary's terminal_id
**Finding (seam map):** `ConnectionIdentity.resolve(remoteAddr, remotePort)` already returns
the caller's herdr `terminal_id` for *every* MCP call, via `PaneLocator.terminalForPid`
(walks `pane.list`, matches the caller PID to a pane's process tree). It is non-null whenever
the caller runs in a herdr pane on this host. Today it's discarded for the primary
(`presence.markPresent` is a no-op on it).
**Plan:** a single-slot `PrimaryRegistry` (thread-safe) holding the primary's `terminal_id`.
Populate it from the **orchestration-side** MCP tools — `bridge_send`, `bridge_spawn`,
`bridge_poll`, `bridge_list`, `bridge_status`, `bridge_profiles` — capturing
`callerTerminal(exchange)` when it is (a) non-null and (b) **not** a registered worker
session in `SessionManager`. That caller is, by construction, the primary. Worker-side tools
(`bridge_reply`, `bridge_ask`) never set it.
- **Config override / pin:** a `primary: { terminal: "<id>" }` block in `BridgedConfig`
(nested record, same shape as `Broker`). Lets an operator pin it, or supply it when
derivation can't (see degrade case).
- **Degrade:** if the primary is off-host or in a non-herdr terminal, `terminalForPid`
returns null and no override is set → **the registry stays empty → the push loop is a
no-op and we fall back to pull** (today's behaviour). The reply is never lost; it's just
not actively pushed. This is a safe, explicit degradation, not a failure.
### Increment 2 — the push loop
A `ReplyPushLoop` component, notified at the single no-waiter call site
(`MessageService.reply` → the `inbox.publish` branch, `MessageService.java:192`).
- **Inject a nudge, not the payload.** The injected turn tells the primary *to drain*
(e.g. "Worker `<target>` returned a reply — run `bridge_poll(target=<target>)` to collect
it"), it does **not** carry the reply text. Rationale: replies can be large/multiline and
terminal injection would mangle them; the drain response is the clean transport. Keeps the
push idempotent — re-nudging is harmless.
- **Ack = drain.** The primary draining (`drainReplies` = peek + ack) is the acknowledgement.
The loop's **stop condition is `inbox.peek(target).isEmpty()`** — the reply is gone from the
inbox because it was acked. No new `bridge_ack` tool needed for v1 (see Increment 3).
- **Status-gated injection (mechanism (b), chosen).** A dedicated lightweight scheduled loop,
**not** the worker `Injector`. It injects via `AgentControl.send(primaryTerminal, nudge)`
(the same herdr `agent.send` = `pane send-text` + submit that delivers to workers) only when
`AgentControl.status(primaryTerminal).injectable()` (IDLE/BLOCKED) — never mid-turn. This keeps
the primary path fully isolated from `WorkerPresence`/`StatusPoller` (which are worker-scoped),
and makes it unit-testable with a fake `AgentControl` + an injected clock (per the CB-306
`LongSupplier` clock + `Runnable` sleeper seam). Rejected (a) reuse-the-Injector: it would force
the primary terminal into the worker poller set and couple to worker-presence semantics — more
integration surface, harder to test, no real gain for a bounded reminder.
- **Bounded reminder / backoff.** While `peek(target)` stays non-empty, re-inject on a
backoff schedule up to a cap (N reminders or a max duration; config
`primary.push_reminders` / `primary.push_backoff_ms`). After the cap, **stop reminding** —
the reply remains in the durable inbox and the next natural poll (or a later worker reply's
nudge) still surfaces it. Bounded so the bridge never spams the primary.
### Increment 3 — optional per-`msgId` `bridge_ack` tool (deferred)
Drain-as-ack is coarse: it clears *all* pending replies for a target at once. If finer
control is ever needed (ack one reply, leave others held), add a `bridge_ack(msgId)` tool
mapping to `inbox.ack(target, msgId)` — the port already supports per-`msgId` ack. Not built
in v1; the stop-on-empty loop is sufficient.
## The two subtleties (decided here)
1. **Which caller is "the primary"?** Connection-derived, not self-reported: the caller whose
resolved terminal is non-null **and not a registered worker session**, seen on an
orchestration-side tool. This never mislabels a worker (workers are in `SessionManager`)
and needs no new env var or argument (identity stays connection-derived, per the existing
`BridgeMcp` invariant).
2. **Readiness-gate mismatch → dedicated loop.** The existing `Injector` gates delivery on
`ready.test(target)` = `WorkerPresence` (the *worker's* MCP connected). The primary is not
in `WorkerPresence`, so reusing `Injector` would mean forcing the primary terminal into the
worker `StatusPoller` set and swapping the `ready` predicate — extra integration surface with
worker-scoped machinery. Decision: **mechanism (b)** — a small dedicated scheduled loop that
calls `AgentControl.status(primaryTerminal).injectable()` then `AgentControl.send(...)`, with
an injected clock. Isolated from worker presence, trivially unit-testable, sufficient for a
bounded reminder. (Verified live: this primary resolves to `term_656c8cc03e1f0b1`, pane
`w2:pY` — the primary genuinely runs in a herdr pane on this host, so the path is exercisable.)
## Boundary note
This is the first time the bridge **writes into the primary's pane** — a new direction of
control. It stays within the communication-bus identity: the injection is a **nudge** (a
synthetic "go drain your replies" turn), **status-gated** so it never interrupts a turn,
**bounded** so it never spams, carries **no env** and **never crosses the subscription
boundary**. The bridge is signalling the primary that it has mail — not driving its work.
## Test plan
- **Unit (hermetic):** `PrimaryRegistry` set/clear/override; the "caller is primary iff
non-null terminal AND not a registered session" predicate; the loop's stop-on-empty and
bounded-reminder logic with an injected clock + a fake injector (no real herdr).
- **Live dogfood (primary-side):** with the daemon on the broker jar + a real worker,
delegate a task, let the worker reply after the `bridge_send` window closes, and observe the
bridge inject a drain nudge into *this* primary pane; confirm draining stops the reminders;
confirm an unreachable primary (registry empty) degrades to pull with no loss.
## Out of scope
Multi-host (CB-308) — the push loop is local-only; a remote primary is reached by its own
local gateway, not cross-host injection. Federation reuses this loop per-gateway.
+144 -5
View File
@@ -6,7 +6,7 @@
<groupId>dev.ltms</groupId>
<artifactId>bridged</artifactId>
<version>0.1.0-SNAPSHOT</version>
<version>1.0.0</version>
<packaging>jar</packaging>
<name>bridged</name>
@@ -17,13 +17,74 @@
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<mainClass>dev.ltms.bridged.Bridged</mainClass>
<jackson.version>2.18.2</jackson.version>
<javalin.version>6.3.0</javalin.version>
<jackson.version>2.19.0</jackson.version>
<javalin.version>6.7.0</javalin.version>
<jetty.version>11.0.25</jetty.version>
<mcp.version>2.0.0</mcp.version>
<slf4j.version>2.0.16</slf4j.version>
<logback.version>1.5.15</logback.version>
<logback.version>1.5.18</logback.version>
<junit.version>5.11.4</junit.version>
<amqp.version>5.22.0</amqp.version>
<testcontainers.version>1.20.4</testcontainers.version>
<commons-compress.version>1.27.1</commons-compress.version>
<commons-lang3.version>3.18.0</commons-lang3.version>
</properties>
<!--
Dependency security (validate with the JetBrains analyzer's Mend.io check on this pom).
Deps are pinned to the latest available versions. Residual advisories with NO upstream fix,
accepted for this loopback-bound daemon that processes no untrusted config:
- jetty-http 11.0.25 (via Javalin): CVE-2026-2332, CVE-2025-11143 — Jetty 11 is EOL;
fixed only in Jetty 12, which needs a Javalin major (6.x rides Jetty 11).
- logback-core 1.5.18: CVE-2025-11226, CVE-2026-1225 — both require a MALICIOUS
logback.xml (attacker with config write already has code execution); ours is trusted.
- jackson-core 2.19.0: WS-2026-0003 — "insufficient information", no fixed version published.
- tools.jackson.core (Jackson 3) 3.0.3 via the MCP SDK: CVE-2026-29062 (nesting-depth
resource exhaustion). The SDK 2.0.0 is pinned to Jackson 3.0.3 + jackson-annotations
3.0-rc5; bumping Jackson 3 to the patched 3.2.x breaks the SDK (annotation mismatch).
Only the loopback /mcp endpoint parses this JSON, from trusted local Claude clients.
The 11.0.23 -> 11.0.25 bump did clear jetty CVE-2024-8184 (5.9) and CVE-2024-6763.
-->
<!-- Force the latest patched Jetty 11.x across all Javalin-pulled Jetty modules (no version
skew). Javalin 6.x rides Jetty 11; a move to Jetty 12 needs a Javalin major. -->
<dependencyManagement>
<dependencies>
<dependency>
<groupId>org.eclipse.jetty</groupId>
<artifactId>jetty-bom</artifactId>
<version>${jetty.version}</version>
<type>pom</type>
<scope>import</scope>
</dependency>
<!-- The MCP SDK (Jackson 3) needs jackson-annotations with JsonFormat.Shape.POJO
(the 3.0 line); it shares the com.fasterxml.jackson.annotation package with our
Jackson 2.19 databind, so both must resolve to the same jar. 3.0 is built to work
with Jackson 2.19 databind too — pin it to reconcile the two. -->
<dependency>
<groupId>com.fasterxml.jackson.core</groupId>
<artifactId>jackson-annotations</artifactId>
<version>3.0-rc5</version>
</dependency>
<!-- Testcontainers 1.20.4 pulls commons-compress 1.24.0 (test scope), which carries
CVE-2024-25710 (8.1) + CVE-2024-26308 — both fixed in 1.26.0. Pin the patched line.
Test-scope only (never shipped in the jar), but bumped per the CVE policy. -->
<dependency>
<groupId>org.apache.commons</groupId>
<artifactId>commons-compress</artifactId>
<version>${commons-compress.version}</version>
</dependency>
<!-- Testcontainers 1.20.4 also pulls commons-lang3 3.16.0 (test scope): CVE-2025-48924
(uncontrolled recursion in ClassUtils), fixed in 3.18.0. Pin the patched line.
Test-scope only (never shipped in the jar), bumped per the CVE policy. -->
<dependency>
<groupId>org.apache.commons</groupId>
<artifactId>commons-lang3</artifactId>
<version>${commons-lang3.version}</version>
</dependency>
</dependencies>
</dependencyManagement>
<dependencies>
<!-- JSON + YAML (config, herdr wire format, REST bodies) -->
<dependency>
@@ -44,6 +105,24 @@
<version>${javalin.version}</version>
</dependency>
<!-- MCP server: the SERVER face. Streamable-HTTP servlet mounted on Javalin's Jetty at
/mcp, exposing bridge_send/bridge_reply/bridge_status as thin adapters over REST. -->
<dependency>
<groupId>io.modelcontextprotocol.sdk</groupId>
<artifactId>mcp</artifactId>
<version>${mcp.version}</version>
</dependency>
<!-- Broker client (CB-307 Stage 2): AMQP 0-9-1. Default deploy targets LavinMQ; this same
client speaks to RabbitMQ unchanged (URI-only swap), so integration tests run against a
stock RabbitMQ container. Only wired when a broker: block is present in config; absent →
the in-memory ReplyInbox. -->
<dependency>
<groupId>com.rabbitmq</groupId>
<artifactId>amqp-client</artifactId>
<version>${amqp.version}</version>
</dependency>
<!-- Logging -->
<dependency>
<groupId>org.slf4j</groupId>
@@ -63,6 +142,22 @@
<version>${junit.version}</version>
<scope>test</scope>
</dependency>
<!-- Testcontainers RabbitMQ: spins a real broker for the @Tag("contract") AMQP integration
test only. Excluded from the default build (contract group), so `mvn clean install`
stays hermetic and green without Docker; run under -Pcontract with Docker present. -->
<dependency>
<groupId>org.testcontainers</groupId>
<artifactId>rabbitmq</artifactId>
<version>${testcontainers.version}</version>
<scope>test</scope>
</dependency>
<dependency>
<groupId>org.testcontainers</groupId>
<artifactId>junit-jupiter</artifactId>
<version>${testcontainers.version}</version>
<scope>test</scope>
</dependency>
</dependencies>
<build>
@@ -74,6 +169,30 @@
<version>3.14.0</version>
</plugin>
<!--
Coverage (CB-509). Build-time tooling only — never a compile/runtime dependency, so
it adds nothing to the shipped jar. Report lands at target/site/jacoco/index.html and
target/site/jacoco/jacoco.csv. No `check` rule / threshold is wired: a coverage gate
rewards writing tests that execute lines, which is the failure mode this project is
trying to avoid, not encourage.
-->
<plugin>
<groupId>org.jacoco</groupId>
<artifactId>jacoco-maven-plugin</artifactId>
<version>0.8.13</version>
<executions>
<execution>
<id>prepare-agent</id>
<goals><goal>prepare-agent</goal></goals>
</execution>
<execution>
<id>report</id>
<phase>test</phase>
<goals><goal>report</goal></goals>
</execution>
</executions>
</plugin>
<!-- Unit tests run by default; contract tests (live herdr) are tag-excluded. -->
<plugin>
<groupId>org.apache.maven.plugins</groupId>
@@ -116,7 +235,27 @@
</profile>
<profile>
<id>contract</id>
<properties><excludedGroups></excludedGroups></properties>
<properties><excludedGroups/></properties>
<build>
<plugins>
<plugin>
<artifactId>maven-surefire-plugin</artifactId>
<configuration>
<!-- Docker-engine compat (see "Running the contract tests" in
docs/CB-307-Reliable-Delivery.md): Testcontainers 1.20.4's docker-java
client defaults to Docker API 1.32 when no version is set, but modern
engines (OrbStack on this dev host, min 1.40) reject that as too old —
which surfaces as "Could not find a valid Docker environment". Pinning
api.version=1.43 works on OrbStack and Docker 24+, and is overridable
per-host via -Dapi.version. Only active under -Pcontract, so the
default hermetic build never sets it. -->
<systemPropertyVariables>
<api.version>1.43</api.version>
</systemPropertyVariables>
</configuration>
</plugin>
</plugins>
</build>
</profile>
</profiles>
</project>
@@ -3,17 +3,52 @@ package dev.ltms.bridged;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
import dev.ltms.bridged.msg.AmqpReplyInbox;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.msg.ReplyPushLoop;
import dev.ltms.bridged.rest.BridgedApp;
import dev.ltms.bridged.worker.WorkerService;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.session.SessionReaper;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.placement.PlacementPolicies;
import dev.ltms.bridged.worker.CompositePeerLauncher;
import dev.ltms.bridged.worker.HerdrPeerLauncher;
import dev.ltms.bridged.worker.OpenCodeLauncher;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
/**
* {@code bridged} entry point. Wires the real herdr socket client to the REST app and
@@ -27,6 +62,10 @@ public final class Bridged {
/** How often the injector samples a busy worker's status while it has queued work. */
private static final long INJECT_POLL_MILLIS = 250;
/** CB-504: how long to wait at startup for herdr's socket before serving degraded. */
private static final long HERDR_WAIT_SECONDS = 30;
private static final long HERDR_WAIT_POLL_MILLIS = 500;
static void main(String[] args) {
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
BridgedConfig cfg = BridgedConfig.load(configPath);
@@ -35,30 +74,252 @@ public final class Bridged {
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
guard.assertPrimaryClean(System.getenv());
// CB-501: refuse to start if the bind is wider than the auth mode can defend. Under
// loopback-trust, "not a known worker" means "the primary" — sound only because the OS
// refuses remote connections to a loopback socket. This throws rather than warns so the
// dangerous configuration cannot be reached by ignoring a log line.
cfg.validateAuthExposure();
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
: UnixSocketHerdrClient.defaultSocketPath();
UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect(socket, new com.fasterxml.jackson.databind.ObjectMapper());
Runtime.getRuntime().addShutdownHook(new Thread(herdr::close));
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
WorkerService workers = new WorkerService(agents, spaces, guard, cfg.worker(), System::getenv);
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
Map<String, BridgedConfig.Worker> claudeProfiles = new LinkedHashMap<>();
Map<String, BridgedConfig.Worker> opencodeProfiles = new LinkedHashMap<>();
cfg.workerProfiles().forEach((name, w) -> {
if (w.isOpenCode()) {
opencodeProfiles.put(name, w);
} else {
claudeProfiles.put(name, w);
}
});
List<HerdrPeerLauncher> adapters = new ArrayList<>();
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
claudeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(agents, spaces,
opencodeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(name -> 0);
PeerLauncher workers = new CompositePeerLauncher(
adapters,
cfg.defaultProfile(),
cfg.workerProfiles(),
PlacementPolicies.fromName(cfg.placement()),
profileName -> liveCountRef.get().apply(profileName));
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
// with /healthz reporting "degraded" is strictly more useful than exiting.
if (awaitHerdr(herdr)) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
// with the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
} else {
log.warn("herdr did not answer within {}s — starting anyway; /healthz will report "
+ "degraded until it comes up. Orphaned worker panes (if any) were NOT reaped.",
HERDR_WAIT_SECONDS);
}
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
int contextCap = 0;
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
&& cfg.lifecycle().contextCap() > 0) {
contextCap = cfg.lifecycle().contextCap();
}
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()), contextCap);
liveCountRef.set(profileName -> (int) sessions.roster().stream()
.filter(s -> profileName.equals(s.profile()))
.count());
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
final SessionReaper reaper;
if (cfg.lifecycle() != null
&& cfg.lifecycle().idleTtlSeconds() != null
&& cfg.lifecycle().idleTtlSeconds() > 0) {
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
reaper.start();
} else {
reaper = null;
}
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
Injector injector = new Injector(agents);
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
Rendezvous rendezvous = new Rendezvous();
CompletionResolver completion = new CompletionResolver(agents, rendezvous);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
WorkerPresence presence = sessions.asPresence();
TurnListener turnListener = new TurnListener() {
@Override
public void onTurnComplete(String target) {
completion.onTurnComplete(target);
sessions.onTurnComplete(target);
}
@Override
public void onDelivered(String target) {
completion.onDelivered(target);
sessions.onDelivered(target);
}
@Override
public void onTurnFailed(String target) {
completion.onTurnFailed(target);
sessions.onTurnFailed(target);
}
};
Injector injector = new Injector(agents, turnListener, presence::isPresent, presence::forget);
StatusPoller poller = new StatusPoller(agents, injector, INJECT_POLL_MILLIS);
poller.start();
Runtime.getRuntime().addShutdownHook(new Thread(poller::stop));
Javalin app = new BridgedApp(herdr, workers).build();
// CB-307: reply inbox. A broker: block (with a uri) selects the AMQP-backed durable adapter;
// absent, bridged stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
// connection, so keep the reference to close it in the ordered shutdown hook.
final ReplyInbox replyInbox;
if (cfg.broker() != null && cfg.broker().isConfigured()) {
replyInbox = AmqpReplyInbox.open(cfg.broker().uri());
log.info("reply inbox: AMQP broker (durable) at {}", cfg.broker().uri());
} else {
replyInbox = new InMemoryReplyInbox();
log.info("reply inbox: in-memory (soft-state)");
}
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
// The pin also feeds CallerResolver below: a primary running inside a herdr pane would
// otherwise resolve as a worker and be refused every orchestration tool.
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
// open bridge_send. Uses its own lightweight scheduled executor, separate from the injector.
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
var pushScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-push-").unstarted(r));
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
// CB-512: the push loop takes it too, so nudge outcomes (delivered|exhausted) are counted.
Metrics metrics = BridgedMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, agents, replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
sessions.onAcquire(replyInbox::own);
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
sessions.onRelease(terminal -> {
messages.abandon(terminal, "the worker session was released before it replied");
replyInbox.release(terminal);
});
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
final CallerResolver callers;
if (cfg.auth().tokenMode()) {
String token = System.getenv(cfg.auth().tokenEnv());
if (token == null || token.isBlank()) {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting bridged");
}
callers = new CallerResolver(identity, true, token, pinnedPrimaryTerminal);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = new CallerResolver(identity, false, null, pinnedPrimaryTerminal);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
BridgeMcp mcp = new BridgeMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics);
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
// last. This replaces the earlier independent hooks that could race and close herdr early.
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
poller.stop();
messages.close();
pushLoop.close();
mcp.close();
if (reaper != null) reaper.stop();
// Release the broker connection last among message resources (no-op for the in-memory inbox).
if (replyInbox instanceof AutoCloseable closeable) {
try {
closeable.close();
} catch (Exception e) {
log.debug("reply inbox close: {}", e.toString());
}
}
herdr.close();
}));
Javalin app = new BridgedApp(herdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics).build();
app.start(cfg.bind().host(), cfg.bind().port());
log.info("bridged listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
*
* @return true if herdr answered, false if it never did
*/
private static boolean awaitHerdr(HerdrClient herdr) {
long deadline = System.nanoTime() + HERDR_WAIT_SECONDS * 1_000_000_000L;
boolean waited = false;
while (true) {
try {
herdr.call("ping");
if (waited) {
log.info("herdr is up");
}
return true;
} catch (HerdrException e) {
if (System.nanoTime() >= deadline) {
return false;
}
if (!waited) {
log.info("waiting up to {}s for the herdr socket…", HERDR_WAIT_SECONDS);
waited = true;
}
try {
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
return false;
}
}
}
}
private Bridged() {
}
}
@@ -0,0 +1,72 @@
package dev.ltms.bridged.auth;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.time.Instant;
import java.time.ZoneOffset;
import java.time.format.DateTimeFormatter;
/**
* Append-only record of privileged actions (CB-505).
*
* <p>Writes JSON lines to a dedicated {@code audit} logger — its own appender, separate from the
* chatty app log — so the trail stays greppable and can later be shipped without dragging debug
* noise along.
*
* <p><strong>Message content is never recorded.</strong> This bridge carries the user's source
* code, diffs, and prompts; an audit trail that quietly accumulated them would be a transcript
* archive wearing a security control's clothing. Records carry <em>who / what / against what /
* outcome</em> and correlation ids only.
*/
public final class AuditLog {
private static final Logger AUDIT = LoggerFactory.getLogger("audit");
private static final DateTimeFormatter TS =
DateTimeFormatter.ofPattern("yyyy-MM-dd'T'HH:mm:ss.SSSXXX").withZone(ZoneOffset.UTC);
private AuditLog() {
}
/** Record an allowed action. */
public static void allowed(Principal caller, Authz.Action action, String target) {
write(caller, action, target, "allowed", null);
}
/** Record a refused action and why. */
public static void denied(Principal caller, Authz.Action action, String target, String reason) {
write(caller, action, target, "denied", reason);
}
/** Record an action that was authorized but then failed downstream (guard, timeout, herdr). */
public static void failed(Principal caller, Authz.Action action, String target, String reason) {
write(caller, action, target, "failed", reason);
}
private static void write(Principal caller, Authz.Action action, String target,
String outcome, String reason) {
Principal c = caller != null ? caller : Principal.anonymous();
StringBuilder sb = new StringBuilder(200);
// The timestamp is built here rather than by the appender pattern: a pattern that wrapped
// literal braces around the message collides with logback's own variable substitution.
sb.append("{\"ts\":\"").append(TS.format(Instant.now())).append('"')
.append(",\"role\":\"").append(c.role()).append('"')
.append(",\"actor\":\"").append(esc(c.describe())).append('"')
.append(",\"pid\":").append(c.pid())
.append(",\"action\":\"").append(action).append('"')
.append(",\"target\":").append(target == null ? "null" : '"' + esc(target) + '"')
.append(",\"outcome\":\"").append(outcome).append('"');
if (reason != null) {
sb.append(",\"reason\":\"").append(esc(reason)).append('"');
}
sb.append('}');
// The appender supplies the timestamp, so it cannot disagree with the app log's clock.
AUDIT.info(sb.toString());
}
/** Minimal JSON string escaping — these values are ids and short reasons, never free text. */
private static String esc(String s) {
return s.replace("\\", "\\\\").replace("\"", "\\\"")
.replace("\n", "\\n").replace("\r", "\\r").replace("\t", "\\t");
}
}
@@ -0,0 +1,72 @@
package dev.ltms.bridged.auth;
/**
* The authorization table (CB-505), stated once and enforced on both entry paths.
*
* <p>Most of these rules are already true de facto — {@code BridgeMcp} derives a worker's identity
* from the connection rather than reading it from an argument, so a worker has never been able to
* reply <em>as</em> another worker over MCP. What was missing is that the REST surface trusted the
* session id in the URL path, and neither surface checked role at all. This class makes the
* invariant explicit and testable rather than emergent.
*/
public final class Authz {
private Authz() {
}
/** A privileged operation, named for the audit trail. */
public enum Action {
/** Spawn a worker peer. */
SPAWN,
/** Tear a worker peer down. */
STOP,
/** Deliver a turn to a session (or answer a worker's question). */
SEND,
/** A worker's terminal reply for its own turn. */
REPLY,
/** A worker's mid-turn question to the primary. */
ASK,
/** Collect held replies from a session's inbox. */
DRAIN,
/** Read-only observation: status, roster, profiles, task polling. */
READ,
/** Scrape the metrics endpoint. */
METRICS
}
/**
* Whether {@code caller} may perform {@code action} against {@code targetSession}.
*
* @param targetSession the session id in the request path; only consulted for the worker-scoped
* actions ({@code REPLY}, {@code ASK}), ignored otherwise, may be
* {@code null}
*/
public static boolean permits(Principal caller, Action action, String targetSession) {
if (caller == null || caller.isAnonymous()) {
return false; // authenticated as nothing ⇒ authorized for nothing
}
return switch (action) {
// Orchestration is the primary's alone. A worker driving spawn/stop/send would be a
// worker escalating into the orchestrator role.
case SPAWN, STOP, SEND, DRAIN -> caller.isPrimary();
// The load-bearing rule: a worker acts only as itself. The primary is deliberately
// excluded — a reply/ask is a worker's own turn output, and letting the primary forge
// one would corrupt the rendezvous correlation it is itself waiting on.
case REPLY, ASK -> caller.ownsSession(targetSession);
// Observation is open to both authenticated roles: a worker legitimately polls its own
// status, and the roster carries no secrets.
case READ, METRICS -> caller.isPrimary() || caller.isWorker();
};
}
/**
* Why a request was refused, for the error body. Distinguishes "you are nobody" from "you are
* somebody, but not the right somebody" — the first is a credential problem (401), the second
* an authorization one (403).
*/
public static boolean isUnauthenticated(Principal caller) {
return caller == null || caller.isAnonymous();
}
}
@@ -0,0 +1,142 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
/**
* Resolves every caller to a {@link Principal}, for both entry paths into the core (CB-501).
*
* <p>There are two of them and they are not layered the way the docs suggest: {@code BridgeMcp}
* calls the service layer directly and is mounted as a raw servlet (so it never passes through a
* Javalin filter), while the REST routes historically resolved no identity at all. Both now
* delegate here, so the authorization rules are stated once instead of drifting apart.
*
* <p><strong>Resolution order</strong> — connection identity first, token second, nothing third:
* <ol>
* <li>A loopback peer PID that maps to the pinned {@code primary.terminal} pane (CB-307) ⇒
* {@link Role#PRIMARY}. The pane mapping is as unforgeable as a worker's, and the config
* explicitly names that pane as the primary's own — without this rule a primary running
* <em>inside</em> a herdr pane is misread as a worker and locked out of orchestration.</li>
* <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#WORKER}. This is
* unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured
* regardless of auth mode, so enabling auth never breaks the fleet.</li>
* <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li>
* <li>Otherwise, under {@code loopback-trust}, a loopback caller ⇒ {@link Role#PRIMARY}
* (the historical behaviour, now an explicit configured choice).</li>
* <li>Otherwise {@link Role#ANONYMOUS}.</li>
* </ol>
*/
public final class CallerResolver {
private final ConnectionIdentity identity;
private final boolean tokenMode;
private final byte[] expectedToken; // null unless tokenMode
private final String pinnedPrimaryTerminal; // null unless primary.terminal is configured
/** Loopback-trust resolver: no token required, historical behaviour. */
public CallerResolver(ConnectionIdentity identity) {
this(identity, false, null, null);
}
/** As {@link #CallerResolver(ConnectionIdentity, boolean, String, String)} with no pin. */
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
this(identity, tokenMode, token, null);
}
/**
* @param identity connection-based worker identification
* @param tokenMode when true, a non-worker caller must present a valid bearer token
* @param token the expected bearer token; required (non-blank) when
* {@code tokenMode}
* @param pinnedPrimaryTerminal the primary's own herdr {@code terminal_id} from
* {@code primary.terminal} ({@code null}/blank = unpinned); a
* caller resolving to this pane is the primary, not a worker
*/
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
String pinnedPrimaryTerminal) {
if (tokenMode && (token == null || token.isBlank())) {
throw new IllegalArgumentException(
"auth.mode=token requires a non-empty token; check that the env var named by "
+ "auth.tokenEnv is exported to the daemon's environment");
}
this.identity = identity;
this.tokenMode = tokenMode;
this.expectedToken = tokenMode ? token.getBytes(StandardCharsets.UTF_8) : null;
this.pinnedPrimaryTerminal =
pinnedPrimaryTerminal == null || pinnedPrimaryTerminal.isBlank()
? null : pinnedPrimaryTerminal;
}
/**
* Resolve the caller of a request.
*
* @param remoteAddr the connection's remote address
* @param remotePort the connection's remote port (used for the peer-PID lookup)
* @param authorizationHeader the raw {@code Authorization} header, or {@code null}
*/
public Principal resolve(String remoteAddr, int remotePort, String authorizationHeader) {
ConnectionIdentity.Caller c = identity.resolve(remoteAddr, remotePort);
if (c.terminal() != null) {
if (c.terminal().equals(pinnedPrimaryTerminal)) {
// The config names this pane as the primary's own. The pane mapping is exactly as
// unforgeable as a worker's, so it outranks the token path — no credential needed.
return Principal.primary(c.pid());
}
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
}
if (tokenMode) {
return presentedTokenMatches(authorizationHeader)
? Principal.primary(c.pid())
: Principal.anonymous();
}
// loopback-trust: same-host callers that are not workers are the primary. A non-loopback
// caller is anonymous even here — and startup refuses that combination anyway
// (BridgedConfig.validateAuthExposure), so this is defence in depth, not the control.
return isLoopback(remoteAddr) ? Principal.primary(c.pid()) : Principal.anonymous();
}
/** The working directory of the calling process (CB-112 spawn cwd inheritance), or {@code null}. */
public String cwdForPid(long pid) {
return identity.cwdForPid(pid);
}
/** True when auth requires a bearer token of non-worker callers. */
public boolean tokenMode() {
return tokenMode;
}
private boolean presentedTokenMatches(String authorizationHeader) {
String presented = bearerValue(authorizationHeader);
if (presented == null) {
return false;
}
// Constant-time: MessageDigest.isEqual does not short-circuit on the first differing byte,
// so a token cannot be recovered a byte at a time by timing the response.
return MessageDigest.isEqual(presented.getBytes(StandardCharsets.UTF_8), expectedToken);
}
/** Extract the credential from {@code Authorization: Bearer <token>}, or {@code null}. */
private static String bearerValue(String header) {
if (header == null) {
return null;
}
String h = header.trim();
if (h.length() < 7 || !h.regionMatches(true, 0, "Bearer ", 0, 7)) {
return null;
}
String token = h.substring(7).trim();
return token.isEmpty() ? null : token;
}
private static boolean isLoopback(String remoteAddr) {
if (remoteAddr == null) {
return false;
}
return remoteAddr.equals("127.0.0.1") || remoteAddr.equals("::1")
|| remoteAddr.equals("0:0:0:0:0:0:0:1") || remoteAddr.startsWith("127.");
}
}
@@ -0,0 +1,57 @@
package dev.ltms.bridged.auth;
/**
* A resolved caller: its {@link Role}, and — for a worker — the herdr {@code terminal_id} that
* identifies which worker it is (CB-501).
*
* @param role what this caller is authorized to act as
* @param terminal the worker's herdr terminal id; {@code null} for {@code PRIMARY}/{@code ANONYMOUS}
* @param pid the connecting process id, or {@code -1} when not resolvable (audit context)
*/
public record Principal(Role role, String terminal, long pid) {
/** A caller authenticated as nothing — the default when no check establishes anything else. */
public static Principal anonymous() {
return new Principal(Role.ANONYMOUS, null, -1);
}
/** The orchestrating session. */
public static Principal primary(long pid) {
return new Principal(Role.PRIMARY, null, pid);
}
/** A worker peer, identified by its herdr pane. */
public static Principal worker(String terminal, long pid) {
return new Principal(Role.WORKER, terminal, pid);
}
public boolean isPrimary() {
return role == Role.PRIMARY;
}
public boolean isWorker() {
return role == Role.WORKER;
}
public boolean isAnonymous() {
return role == Role.ANONYMOUS;
}
/**
* Whether this caller may act <em>as</em> {@code sessionId} — the "own session only" rule that
* keeps one worker from replying or asking on another's behalf. Only a worker can own a
* session, and only its own.
*/
public boolean ownsSession(String sessionId) {
return isWorker() && terminal != null && terminal.equals(sessionId);
}
/** Short, non-sensitive description for audit lines and error details. */
public String describe() {
return switch (role) {
case WORKER -> "worker:" + terminal;
case PRIMARY -> "primary";
case ANONYMOUS -> "anonymous";
};
}
}
@@ -0,0 +1,29 @@
package dev.ltms.bridged.auth;
/**
* What a caller is allowed to be on the bus (CB-501).
*
* <p>The ordering matters conceptually: {@link #PRIMARY} is the <em>most</em> privileged role
* (it spawns, stops, sends to any session, and drains any inbox), not the least. Before CB-501
* the daemon reached {@code PRIMARY} by <em>failing</em> every other check — any caller that did
* not resolve to a known worker pane was treated as the primary. That is inverted here:
* {@link #ANONYMOUS} is the fallback, and {@code PRIMARY} must be established.
*/
public enum Role {
/**
* The orchestrating session. Established either by being a loopback caller that is not a
* worker pane (under {@code loopback-trust}) or by presenting a valid bearer token (under
* {@code token} mode).
*/
PRIMARY,
/**
* A worker peer, identified by its herdr pane. Unforgeable: derived from the connection's
* loopback peer PID via herdr's PID→pane map, never from a request argument.
*/
WORKER,
/** Authenticated as nothing. Authorized for nothing but {@code /healthz}. */
ANONYMOUS
}
@@ -8,7 +8,10 @@ import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Collections;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
/**
@@ -16,23 +19,51 @@ import java.util.Set;
* {@code bridged.example.yaml}). Unknown keys are ignored so config can grow ahead
* of the code.
*
* @param bind REST/MCP listen host:port
* @param herdrSocket path to herdr's Unix socket ({@code null} → client default)
* @param worker worker-spawn settings
* @param guard subscription-boundary allowlist
* @param bind REST/MCP listen host:port
* @param herdrSocket path to herdr's Unix socket ({@code null} → client default)
* @param worker single worker profile (legacy; superseded by {@code workers})
* @param workers named worker profiles, keyed by profile name (multi-backend fleet)
* @param defaultWorker which {@code workers} key a no-argument spawn uses ({@code null} → the
* single {@code worker}, or the sole/first profile)
* @param guard subscription-boundary allowlist
* @param worktreeRoot nullable root directory for provisioned worktrees; defaults to a sibling
* of the repo root
* @param lifecycle session lifecycle limits ({@code null} = all disabled)
* @param spawnReadyTimeoutMs max ms to wait for a spawned worker to reach an injectable state
* ({@code null} / 0 disables the poll gate — legacy non-blocking behaviour)
* @param spawnReadyPollMs poll interval while waiting for the worker to become injectable
* @param broker external AMQP broker for durable reply delivery ({@code null} → in-memory,
* soft-state {@code ReplyInbox}; present → the AMQP-backed adapter, CB-307 Stage 2)
* @param primary optional pinned primary terminal config ({@code null} → derived from connection);
* a non-blank {@code terminal} seeds {@code PrimaryRegistry} and prevents
* connection-derived overrides, CB-307
* @param placement how to choose a worker profile for an unqualified spawn:
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
* historical behaviour), CB-501
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record BridgedConfig(
Bind bind,
String herdrSocket,
Worker worker,
Guard guard) {
Map<String, Worker> workers,
String defaultWorker,
Guard guard,
String worktreeRoot,
Lifecycle lifecycle,
Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs,
Broker broker,
Primary primary,
String placement,
Auth auth) {
@JsonIgnoreProperties(ignoreUnknown = true)
public record Bind(String host, int port) {
public Bind {
if (host == null || host.isBlank()) host = "127.0.0.1";
if (port <= 0) port = 8080;
if (port <= 0) port = 8765;
}
}
@@ -53,17 +84,140 @@ public record BridgedConfig(
* @param tabLabel template for a worker tab's label; {@code {profile}}/{@code {model}}
* and {@code {n}} (per-worker number, to keep sibling tabs distinct)
* are substituted (default {@code "worker: {profile} #{n}"})
* @param mcpUrl bridge MCP URL to provision into the worker's {@code configDir} so it
* can call {@code bridge_reply} ({@code null}/blank → no provisioning; the
* worker won't reply, only the fallback/timeout resolves the send)
* @param cwd fixed working directory for this profile's workers (CB-112 "told otherwise");
* {@code null}/blank → inherit the primary's cwd, else the daemon's
* @param parityOverlay repo-relative paths copied primary→worktree for config parity; null/empty
* defaults to a sensible set of local config files
* @param gitTokenEnv name of the host env var holding the git-forge API token; when set, its
* value is injected as {@code GITEA_TOKEN} so the worker can open its own PR
* at checkpoint (CB-302). {@code null}/blank ⇒ no token is injected
* (minimal-grant default — push over SSH stays free, PR-create is opt-in)
* @param gitHostEnv name of the host env var holding the forge host (default {@code GITEA_HOST});
* injected as {@code GITEA_HOST} <em>only</em> when {@code gitTokenEnv} is set
* @param weight relative selection weight for {@code placement: weighted}. Absent or
* non-positive ⇒ 1.0. Weights are normalised by the policy, so they need
* not sum to 1.0.
* @param maxLoad max live workers allowed on this profile at one time; absent or
* non-positive ⇒ unlimited. Live means any session the registry still owns
* (acquired and not yet released), in any state.
* @param kind which peer launcher spawns this profile: {@code "claude-code"} (default —
* the {@link dev.ltms.bridged.worker.ClaudeCodeLauncher}) or {@code "opencode"}.
* The {@code CompositePeerLauncher} routes {@code spawn}/reap by this value, so
* each adapter drives only its own kind. Normalised to lower-case; blank ⇒ the
* default. It selects the adapter, not the transport — placement, tabs, cwd, and
* the readiness gate are kind-independent and stay in the shared base.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel) {
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd,
List<String> parityOverlay,
String gitTokenEnv, String gitHostEnv,
String kind,
Map<String, String> env,
Float weight,
Integer maxLoad) {
/** Peer kind spawned by {@link dev.ltms.bridged.worker.ClaudeCodeLauncher} (the default). */
public static final String KIND_CLAUDE_CODE = "claude-code";
/** Peer kind spawned by the opencode adapter (CB-402). */
public static final String KIND_OPENCODE = "opencode";
/**
* Peer kind spawned by the Codex adapter (CB-528). Like {@link #KIND_OPENCODE} it carries
* its own argv and never inherits the Claude binary, and it sits outside the
* {@code ANTHROPIC_BASE_URL} subscription guard because Codex has no such seam.
*/
public static final String KIND_CODEX = "codex";
public Worker {
argv = (argv == null || argv.isEmpty()) ? List.of("claude") : List.copyOf(argv);
// A claude-code worker defaults its launch command to `claude`; other kinds carry their own
// argv (e.g. `opencode`) and must not inherit the Claude binary — so only default when unset
// AND this is the claude-code kind.
String k = (kind == null || kind.isBlank()) ? KIND_CLAUDE_CODE : kind.toLowerCase();
argv = (argv == null || argv.isEmpty())
? (KIND_CLAUDE_CODE.equals(k) ? List.of("claude") : List.of(k))
: List.copyOf(argv);
kind = k;
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_WORKER_TOKEN" : tokenEnv;
placement = (placement == null || placement.isBlank()) ? "tab" : placement.toLowerCase();
workspace = (workspace == null || workspace.isBlank()) ? "bridged-workers" : workspace;
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? "worker: {profile} #{n}" : tabLabel;
// CB-525: .mcp.json is deliberately NOT here. Replicating the primary's MCP config gave a
// worker the primary's IDE servers, which are bound to the primary's checkout — so its
// navigation returned paths outside its own worktree. GitWorktrees now neutralizes that
// file instead; a worker's tools are whatever its launcher mounts.
parityOverlay = (parityOverlay == null || parityOverlay.isEmpty())
? List.of(".claude/settings.local.json", ".env", ".envrc")
: List.copyOf(parityOverlay);
// gitTokenEnv stays null when unset (opt-in). gitHostEnv defaults so operators enabling
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
gitHostEnv = (gitHostEnv == null || gitHostEnv.isBlank()) ? "GITEA_HOST" : gitHostEnv;
env = (env == null) ? Map.of() : Map.copyOf(env);
weight = (weight == null || weight <= 0.0f) ? 1.0f : weight;
maxLoad = (maxLoad == null || maxLoad <= 0) ? null : maxLoad;
}
/**
* Backward-compatible constructor without the CB-302 git-forge fields — the worker is
* granted no PR-create token (push over SSH is unaffected). Keeps pre-CB-302 call sites
* (and any {@code workers:} YAML that omits the git keys) working unchanged.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, null, null, null, null, null, null);
}
/**
* Backward-compatible constructor with the CB-302 git-forge fields but no explicit peer
* {@code kind} — defaults to {@link #KIND_CLAUDE_CODE}. Keeps pre-CB-402 call sites working.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, null, null, null, null);
}
/**
* Backward-compatible constructor without the CB-511 {@code env:} passthrough — the worker
* gets the daemon's PATH and nothing else. Keeps pre-CB-511 call sites working.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, null, null, null);
}
/** A copy with {@code profile} set — used to default a profile to its {@code workers} key. */
public Worker withProfile(String p) {
return new Worker(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad);
}
/** True when this profile is served by the Claude Code adapter (the default kind). */
public boolean isClaudeCode() {
return KIND_CLAUDE_CODE.equals(kind);
}
/** True when this profile is served by the opencode adapter (CB-402). */
public boolean isOpenCode() {
return KIND_OPENCODE.equals(kind);
}
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
public boolean hasGitToken() {
return gitTokenEnv != null && !gitTokenEnv.isBlank();
}
/** True when workers should land in their own tab in the worker space. */
@@ -71,6 +225,11 @@ public record BridgedConfig(
return "tab".equals(placement);
}
/** True when the bridge MCP should be mounted into a spawned worker (via launch flags). */
public boolean hasMcp() {
return mcpUrl != null && !mcpUrl.isBlank();
}
/**
* Render {@link #tabLabel} for the {@code n}-th worker (substitutes
* {@code {profile}}/{@code {model}}/{@code {n}}), so sibling worker tabs are distinct.
@@ -83,6 +242,102 @@ public record BridgedConfig(
}
}
/**
* Session lifecycle limits. All knobs are opt-in: {@code null} or {@code 0} disables the
* feature so existing configs keep the previous behaviour.
*
* @param idleTtlSeconds max seconds a {@code READY}/{@code DONE} session may sit idle
* before it is reaped ({@code null} → disabled)
* @param contextCap max delegated turns a session serves before force-release
* ({@code null} → disabled)
* @param drainTimeoutSeconds seconds to wait for {@code BUSY} sessions to finish before
* forced teardown on shutdown (default 5 when unset)
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Lifecycle(Integer idleTtlSeconds, Integer contextCap, Integer drainTimeoutSeconds) {
}
/**
* External AMQP broker for durable, cross-restart reply delivery (CB-307 Stage 2). Its mere
* presence swaps the in-memory {@code ReplyInbox} for the AMQP-backed adapter; absent, bridged
* stays soft-state. Production default is LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1
* and is a URI-only swap.
*
* @param uri AMQP connection URI, e.g. {@code amqp://guest:guest@127.0.0.1:5672/}. Blank/{@code null}
* ⇒ the broker block is treated as absent (in-memory adapter).
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Broker(String uri) {
/** True when a usable broker URI is configured (an empty block does not enable AMQP). */
public boolean isConfigured() {
return uri != null && !uri.isBlank();
}
}
/**
* Optional pinned primary terminal config (CB-307). When present with a non-blank
* {@code terminal}, the bridge uses this as the primary's herdr identity instead of
* deriving it from the MCP connection. It feeds two consumers: the push loop (where to nudge
* when replies land), and caller resolution — a caller whose connection maps to this pane is
* the primary, where the pane match would otherwise classify it as a worker. Pin it when the
* primary runs <em>inside</em> a herdr pane; it also helps off-host or non-herdr primaries,
* where connection-derived identity is unavailable and only the nudge target matters.
*
* @param terminal the primary's herdr {@code terminal_id} ({@code null}/blank → derive)
* @param pushReminders max reminder nudges before giving up (default 5)
* @param pushBackoffMs delay between reminders in ms (default 15000)
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Primary(String terminal, Integer pushReminders, Integer pushBackoffMs) {
/** @return configured reminder cap, or 5 */
public int remindersOrDefault() {
return pushReminders != null ? pushReminders : 5;
}
/** @return configured backoff in ms, or 15000 */
public long backoffMsOrDefault() {
return pushBackoffMs != null ? pushBackoffMs.longValue() : 15_000L;
}
}
/**
* API authentication (CB-501). Governs how a caller that is <em>not</em> an on-host worker
* pane proves it is the primary.
*
* <p>Worker identity never depends on this block: a loopback peer PID that maps to a herdr
* pane is unforgeable and is always honoured (see
* {@link dev.ltms.bridged.mcp.ConnectionIdentity}). This only decides what happens for
* <em>everyone else</em>.
*
* @param mode {@code "loopback-trust"} (default) — any loopback caller that is not a known
* worker is the primary, no credential needed; this is the historical
* behaviour, now chosen explicitly rather than implied. {@code "token"} — such
* a caller must present {@code Authorization: Bearer <token>} or it is
* {@code ANONYMOUS} and authorized for nothing.
* @param tokenEnv name of the host env var holding the bearer token; the literal value is
* never stored in config. Defaults to {@code BRIDGED_API_TOKEN}. Only read
* when {@code mode} is {@code token}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Auth(String mode, String tokenEnv) {
/** Historical behaviour: loopback non-worker ⇒ primary, no credential. */
public static final String MODE_LOOPBACK_TRUST = "loopback-trust";
/** A non-worker caller must present a valid bearer token to be the primary. */
public static final String MODE_TOKEN = "token";
public Auth {
mode = (mode == null || mode.isBlank()) ? MODE_LOOPBACK_TRUST : mode.toLowerCase();
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_API_TOKEN" : tokenEnv;
}
/** True when a bearer token is required of every non-worker caller. */
public boolean tokenMode() {
return MODE_TOKEN.equals(mode);
}
}
/**
* Subscription boundary. Only these hosts may back a worker's
* {@code ANTHROPIC_BASE_URL}; the primary must carry none.
@@ -100,6 +355,44 @@ public record BridgedConfig(
}
}
/**
* The effective worker profiles, keyed by profile name. Prefers the {@code workers} map (each
* value's {@code profile} defaulted to its key); falls back to the legacy singular {@code worker}
* (keyed by its own profile). Empty if neither is configured.
*/
public Map<String, Worker> workerProfiles() {
if (workers != null && !workers.isEmpty()) {
Map<String, Worker> out = new LinkedHashMap<>();
workers.forEach((name, w) -> out.put(name,
(w.profile() == null || w.profile().isBlank()) ? w.withProfile(name) : w));
// Deliberately NOT Map.copyOf: its iteration order is salted per JVM run, which would
// discard the YAML definition order built above. Placement tie-breaks on candidate
// order (see WeightedRoundRobinPolicy), so losing it makes equal-weight placement
// non-reproducible across restarts. Unmodifiable-wrap instead of copy-and-scramble.
return Collections.unmodifiableMap(out);
}
if (worker != null) {
String name = (worker.profile() == null || worker.profile().isBlank()) ? "default" : worker.profile();
return Map.of(name, worker);
}
return Map.of();
}
/**
* The profile a no-argument spawn uses: {@code defaultWorker} if set, else the legacy single
* {@code worker}'s profile, else the sole/first configured profile, else {@code null}.
*/
public String defaultProfile() {
if (defaultWorker != null && !defaultWorker.isBlank()) {
return defaultWorker;
}
if (worker != null && worker.profile() != null && !worker.profile().isBlank()) {
return worker.profile();
}
Map<String, Worker> p = workerProfiles();
return p.isEmpty() ? null : p.keySet().iterator().next();
}
private static final ObjectMapper YAML = new ObjectMapper(new YAMLFactory());
/** Load and validate config from {@code path}. */
@@ -116,6 +409,47 @@ public record BridgedConfig(
public BridgedConfig withDefaults() {
Bind b = bind != null ? bind : new Bind(null, 0);
Guard g = guard != null ? guard : new Guard(List.of());
return new BridgedConfig(b, herdrSocket, worker, g);
Lifecycle l = lifecycle != null ? lifecycle : new Lifecycle(null, null, null);
Integer timeout = (spawnReadyTimeoutMs != null) ? spawnReadyTimeoutMs : 20000;
Integer pollMs = (spawnReadyPollMs != null) ? spawnReadyPollMs : 300;
Auth a = auth != null ? auth : new Auth(null, null);
String placementOrDefault = (placement != null && !placement.isBlank()) ? placement : "fixed";
// broker is left as-is: null (or an empty/blank uri) keeps the in-memory soft-state inbox.
// primary is left as-is: null defaults to connection-derived identity.
return new BridgedConfig(b, herdrSocket, worker, workers, defaultWorker, g, worktreeRoot, l, timeout, pollMs, broker, primary, placementOrDefault, a);
}
/**
* Reject a configuration whose network exposure outruns its authentication (CB-501).
*
* <p>{@code loopback-trust} means "any caller that is not a known worker pane is the primary" —
* safe only because the OS refuses non-local connections to a loopback bind. Widen
* {@code bind.host} without switching to {@code token} mode and that sentence becomes "any
* client that can reach this port is the primary", which is the most privileged role on the
* bus. Rather than document the hazard, make it unrepresentable: fail fast at startup.
*
* @throws IllegalStateException when a non-loopback bind is paired with {@code loopback-trust}
*/
public void validateAuthExposure() {
String host = bind().host();
if (isLoopbackBind(host) || auth().tokenMode()) {
return;
}
throw new IllegalStateException(
"refusing to start: bind.host=" + host + " is not loopback, but auth.mode="
+ auth().mode() + ". A non-loopback bind treats every unauthenticated "
+ "caller as the primary (spawn/stop/send/drain on any session). Set "
+ "auth.mode: token (with auth.tokenEnv) before exposing this port, or "
+ "bind to 127.0.0.1 and put a reverse proxy in front.");
}
/** True for the loopback addresses and the unspecified-but-local forms we treat as same-host. */
private static boolean isLoopbackBind(String host) {
if (host == null || host.isBlank()) {
return true; // Bind's own default is 127.0.0.1
}
String h = host.trim().toLowerCase();
return h.equals("127.0.0.1") || h.equals("::1") || h.equals("localhost")
|| h.startsWith("127.");
}
}
@@ -15,6 +15,10 @@ import com.fasterxml.jackson.databind.JsonNode;
* @param agentType detected agent kind, e.g. {@code "claude"}, or {@code null} before herdr
* has detected it (the start-time shape)
* @param status current lifecycle state
* @param name the unique label the agent was started with — for a bridge worker this is
* {@code claude-<profile>-<nonce>-<seq>} (CB-117 keys orphan reaping on the
* nonce); {@code null} for agents the bridge did not start, e.g. a user's own
* Claude session
*/
public record Agent(
String terminalId,
@@ -23,7 +27,8 @@ public record Agent(
String tabId,
String sessionId,
String agentType,
AgentStatus status) {
AgentStatus status,
String name) {
/** Project a herdr {@code agent} node. Tolerates the start-time shape (no session yet). */
public static Agent from(JsonNode a) {
@@ -42,6 +47,7 @@ public record Agent(
a.path("tab_id").asText(null),
sessionId,
type,
AgentStatus.fromWire(a.path("agent_status").asText(null)));
AgentStatus.fromWire(a.path("agent_status").asText(null)),
a.path("name").asText(null));
}
}
@@ -6,12 +6,18 @@ import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
/**
* Domain layer over herdr's native {@code agent.*} namespace — the worker south side.
* Chosen in the CB-102 spike over the pane + {@code send_text} fallback because
* {@code agent.start} takes a first-class {@code env} map (clean, guard-checked
* subscription injection) and herdr tracks each worker's Claude session UUID itself.
* Chosen in the CB-102 spike over the pane + {@code send_text} fallback because herdr
* tracks each worker's Claude session UUID itself.
*
* <p>Ported to herdr protocol 19 (herdr 0.8.0, CB-521): {@code agent.start} now starts a
* <em>supported</em> agent ({@code kind}) into an <em>existing</em> pane, so the worker's
* {@code env}/{@code cwd} move to pane creation ({@code tab.create}/{@code pane.split} — see
* {@link WorkspaceControl}), and {@code agent.send} is replaced by {@code agent.prompt}
* (which submits in one call) plus {@code agent.send_keys} for the raw Enter nudge.
*
* <p>Every method is one herdr call through the injected {@link HerdrClient}, so this
* layer is unit-testable with a fake and contract-tested against a live daemon.
@@ -20,42 +26,100 @@ public final class AgentControl {
private final HerdrClient herdr;
/**
* Protocol 19 dropped {@code terminal_id} as an {@code agent.*} target — herdr now resolves
* targets by pane id or agent name only, while the bridge keys every session on the terminal.
* This caches the terminal→pane mapping (stable for a worker's lifetime) so callers keep
* addressing agents by terminal; entries are invalidated on {@code agent_not_found}.
*/
private final Map<String, String> paneByTerminal = new ConcurrentHashMap<>();
public AgentControl(HerdrClient herdr) {
this.herdr = herdr;
}
/**
* Spawn an agent. {@code env} is applied to the process environment verbatim — this
* is where a worker's {@code ANTHROPIC_BASE_URL} lives, and the ONLY place it should.
*
* @param name label/kind for herdr status detection (e.g. {@code "claude"})
* @param argv launch command, e.g. {@code ["claude"]}
* @param env process environment additions ({@code ANTHROPIC_BASE_URL}, token, …)
*/
public Agent start(String name, List<String> argv, Map<String, String> env) {
return start(name, argv, env, null);
/** One agent-targeted call, translating a terminal id to its pane id (retrying once fresh). */
private JsonNode agentCall(String method, String target, Map<String, Object> extra) {
String resolved = resolveTarget(target);
try {
return herdr.call(method, withTarget(resolved, extra));
} catch (HerdrException e) {
if (!"agent_not_found".equals(e.code()) || resolved.equals(target)) throw e;
paneByTerminal.remove(target); // the cached pane went away — re-resolve once
String fresh = resolveTarget(target);
if (fresh.equals(resolved)) throw e;
return herdr.call(method, withTarget(fresh, extra));
}
}
private static Map<String, Object> withTarget(String target, Map<String, Object> extra) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("target", target);
m.putAll(extra);
return m;
}
/** The pane id behind a terminal-id target, or the target verbatim for pane ids / names. */
private String resolveTarget(String target) {
if (target == null || !target.startsWith("term_")) {
return target;
}
String cached = paneByTerminal.get(target);
if (cached != null) {
return cached;
}
for (JsonNode a : herdr.call("agent.list").path("agents")) {
if (target.equals(a.path("terminal_id").asText(null))) {
String pane = a.path("pane_id").asText(null);
if (pane != null) {
paneByTerminal.put(target, pane);
return pane;
}
}
}
return target; // unknown terminal — let herdr report it against the original target
}
/**
* Spawn an agent into a specific tab. With a non-null {@code tabId} the worker lands
* in that tab (the placement policy's dedicated worker tab); with {@code null} herdr
* splits the currently-focused tab (legacy pane placement).
* Start an agent into {@code paneId}, which must be sitting at its interactive shell prompt —
* the seed pane of a freshly-created worker tab, or a fresh split. The pane's shell already
* carries the worker's env ({@code ANTHROPIC_BASE_URL}, token, …) and cwd from pane creation;
* herdr resolves the executable from {@code kind} and waits (its default timeout) until the
* agent is detected and ready for input.
*
* @param name unique label for this agent ({@code <kind>-<profile>-<nonce>-<seq>})
* @param kind supported agent kind and canonical executable, e.g. {@code "claude"},
* {@code "opencode"}
* @param args extra arguments after the executable, e.g. {@code --mcp-config …}
* @param paneId the pane to start the agent in
*/
public Agent start(String name, List<String> argv, Map<String, String> env, String tabId) {
public Agent start(String name, String kind, List<String> args, String paneId) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("name", name);
params.put("argv", argv);
params.put("env", env);
if (tabId != null) {
params.put("tab_id", tabId);
}
params.put("kind", kind);
params.put("pane_id", paneId);
params.put("args", args);
JsonNode result = herdr.call("agent.start", params);
return Agent.from(result.get("agent"));
}
/** Deliver {@code text} to an agent (its next prompt input). */
/**
* Deliver {@code text} to an agent as its next prompt <em>and submit it</em> — herdr's
* {@code agent.prompt} pastes the text (embedded newlines preserved verbatim) and submits it
* in the same call, replacing the pre-protocol-19 two-event {@code agent.send} dance.
*/
public void send(String target, String text) {
herdr.call("agent.send", Map.of("target", target, "text", text));
agentCall("agent.prompt", target, Map.of("text", text));
}
/**
* Re-send the submit keystroke (Enter) to {@code target}. The submit that accompanies a
* delivery can race the paste — especially right as the worker's TUI becomes interactive —
* leaving the text unsubmitted; the injector nudges it with this until the worker actually
* picks up (CB-113).
*/
public void submit(String target) {
agentCall("agent.send_keys", target, Map.of("keys", List.of("enter")));
}
/**
@@ -64,13 +128,13 @@ public final class AgentControl {
* @param source one of {@code visible|recent|recent_unwrapped|detection}
*/
public String read(String target, String source) {
JsonNode result = herdr.call("agent.read", Map.of("target", target, "source", source));
JsonNode result = agentCall("agent.read", target, Map.of("source", source));
return result.path("read").path("text").asText("");
}
/** Current agent record (status, session UUID, pane). */
public Agent get(String target) {
return Agent.from(herdr.call("agent.get", Map.of("target", target)).get("agent"));
return Agent.from(agentCall("agent.get", target, Map.of()).get("agent"));
}
/** Just the lifecycle status — what the status-gated injector checks before send. */
@@ -2,28 +2,38 @@ package dev.ltms.bridged.herdr;
/**
* A herdr agent's lifecycle state, as reported by {@code agent_status}. Drives the
* status-gated injector: a worker is safe to inject into only when {@link #IDLE} or
* {@link #BLOCKED}, never mid-turn ({@link #WORKING}).
* status-gated injector: a worker is safe to inject into only when {@link #IDLE},
* {@link #BLOCKED}, or {@link #DONE}, never mid-turn ({@link #WORKING}).
*/
public enum AgentStatus {
IDLE,
WORKING,
BLOCKED,
/**
* The worker has finished its turn and is settled at an idle prompt. herdr emits this
* (observed live alongside {@code idle}) as a turn-complete marker; earlier code mapped the
* unrecognized string to {@link #UNKNOWN}, which both wedged delivery (not {@link #injectable})
* and mis-fired the CB-109 stall-failure on a worker that had actually answered. It is a
* turn-boundary equivalent to {@link #IDLE}: injectable, and a {@code working → done} edge is a
* real completion.
*/
DONE,
UNKNOWN;
/** Map herdr's wire string ({@code idle|working|blocked|unknown}) to the enum. */
/** Map herdr's wire string ({@code idle|working|blocked|done|unknown}) to the enum. */
public static AgentStatus fromWire(String s) {
if (s == null) return UNKNOWN;
return switch (s.toLowerCase()) {
case "idle" -> IDLE;
case "working" -> WORKING;
case "blocked" -> BLOCKED;
case "done" -> DONE;
default -> UNKNOWN;
};
}
/** Whether {@code bridged} may inject a message now without stepping on a live turn. */
public boolean injectable() {
return this == IDLE || this == BLOCKED;
return this == IDLE || this == BLOCKED || this == DONE;
}
}
@@ -0,0 +1,59 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import java.util.Map;
/**
* Resolves which herdr pane a process belongs to — the herdr half of connection-based MCP
* identity (CB-105). Given the PID that opened an MCP connection, {@link #terminalForPid} finds
* the agent pane whose process tree contains it, so {@code bridged} can tell <em>which worker</em>
* is calling without the worker sending anything spoofable.
*
* <p>herdr owns the PID→pane truth: {@code pane.process_info} reports each pane's {@code shell_pid}
* and foreground process PIDs. This scans agent panes; a spawn-time {@code pid→terminal} cache is
* the obvious optimization once wired into {@code ClaudeCodeLauncher}.
*/
public final class PaneLocator {
private final HerdrClient herdr;
public PaneLocator(HerdrClient herdr) {
this.herdr = herdr;
}
/**
* The {@code terminal_id} of the agent pane whose process tree contains {@code pid}, or
* {@code null} if no agent pane owns it (e.g. the caller is the primary, or off-host).
*/
public String terminalForPid(long pid) {
if (pid <= 0) {
return null;
}
for (JsonNode pane : herdr.call("pane.list", Map.of()).path("panes")) {
String paneId = pane.path("pane_id").asText(null);
if (paneId != null && paneOwnsPid(paneId, pid)) {
return pane.path("terminal_id").asText(null);
}
}
return null;
}
private boolean paneOwnsPid(String paneId, long pid) {
JsonNode info;
try {
info = herdr.call("pane.process_info", Map.of("pane_id", paneId)).path("process_info");
} catch (HerdrException e) {
return false; // pane vanished mid-scan — just skip it
}
if (info.path("shell_pid").asLong(-1) == pid) {
return true;
}
for (JsonNode p : info.path("foreground_processes")) {
if (p.path("pid").asLong(-1) == pid) {
return true;
}
}
return false;
}
}
@@ -23,9 +23,9 @@ public record Tab(String tabId, String workspaceId, String label, int paneCount)
}
/**
* A freshly-created tab together with the placeholder shell pane herdr seeds it with.
* The caller starts the worker into {@link #tab()} then closes {@link #rootPaneId()} so
* only the worker pane remains.
* A freshly-created tab together with the shell pane herdr seeds it with. Under protocol 19
* the caller starts the worker <em>into</em> {@link #rootPaneId()} — the seed pane's shell
* carries the worker's cwd and env from {@code tab.create}, and becomes the worker pane.
*/
public record Created(Tab tab, String rootPaneId) {
/** Project a {@code tab_created} result ({@code {tab, root_pane}}). */
@@ -5,6 +5,7 @@ import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
@@ -65,13 +66,37 @@ public final class WorkspaceControl {
}
/**
* A brand-new tab in {@code workspaceId} plus the placeholder shell pane herdr seeds
* it with. Start the worker into the tab, then {@code pane.close} the root pane so the
* tab holds only the worker.
* A brand-new tab in {@code workspaceId} plus the shell pane herdr seeds it with. Under
* protocol 19 that seed pane is where the worker <em>starts</em>: its shell carries
* {@code cwd} and {@code env} (the worker's {@code ANTHROPIC_BASE_URL} — this is the
* subscription-injection seam now), and {@code agent.start} launches the agent into it.
*/
public Tab.Created createTab(String workspaceId) {
JsonNode result = herdr.call("tab.create", Map.of("workspace_id", workspaceId));
return Tab.Created.from(result);
public Tab.Created createTab(String workspaceId, String cwd, Map<String, String> env) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("workspace_id", workspaceId);
if (cwd != null && !cwd.isBlank()) {
params.put("cwd", cwd);
}
if (env != null && !env.isEmpty()) {
params.put("env", env);
}
return Tab.Created.from(herdr.call("tab.create", params));
}
/**
* Split the currently-focused tab and return the new pane's id — the legacy pane placement's
* seed pane, carrying {@code cwd} and {@code env} exactly as {@link #createTab}'s does.
*/
public String splitPane(String cwd, Map<String, String> env) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("direction", "right");
if (cwd != null && !cwd.isBlank()) {
params.put("cwd", cwd);
}
if (env != null && !env.isEmpty()) {
params.put("env", env);
}
return herdr.call("pane.split", params).path("pane").path("pane_id").asText(null);
}
/** Give a worker's tab a human label in the tab bar. */
@@ -0,0 +1,256 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.msg.Rendezvous;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
/**
* The CB-106 completion fallback: bridges the {@link Injector}'s turn-completion signal to the
* {@link Rendezvous} so a blocking {@code bridge_send} resolves even when the worker finishes its
* task without ever calling {@code bridge_reply} — the common case for a real delegated coding task.
*
* <p>On a confirmed {@code working → idle} boundary it scrapes the worker's recent transcript and
* resolves the awaiting send with that tail (a {@link Rendezvous.Kind#COMPLETION} resolution, so the
* caller can tell a scrape from a structured reply). It scrapes only when a send is actually waiting
* — a fleet worker's own turns, or a send that already timed out, cost no herdr traffic. An explicit
* {@code bridge_reply} that raced in first wins; {@link Rendezvous#resolveCompletion} is then a no-op.
*
* <p>It also handles the CB-109 stall signal ({@link #onTurnFailed}): a worker that ran a turn then
* wedged in an {@code unknown} state resolves the send as a failure (with the error screen as
* context) rather than leaving it to time out.
*
* <p>The scrape is cleaned to the last {@code ⏺} assistant block (stripping TUI chrome) and guarded
* against misattribution (CB-115): the pane content is baselined on delivery ({@link #onDelivered}),
* and a completion whose scrape is unchanged from that baseline — the previous turn's wind-down
* sampled as this turn's boundary on a rapid back-to-back send — is suppressed rather than resolving
* the send with a stale answer.
*
* <p><strong>Waiter-specific resolution (CB-116).</strong> On delivery we also capture the exact
* {@link Rendezvous} waiter this turn belongs to, and the completion/failure fallbacks resolve
* <em>that</em> waiter — never "whatever send is waiting now". A completion fallback runs on a virtual
* thread and can land after the worker's {@code bridge_reply} already resolved the turn and the
* <em>next</em> send opened its own waiter on the same session; resolving the current waiter would
* then deliver turn N's stale scrape as turn N+1's answer. Targeting the captured waiter makes a late
* completion a harmless no-op (its waiter is already done) instead of a cross-turn stale reply.
*
* <p>Wired as the {@link Injector}'s {@link TurnListener}; the handlers hand off to a virtual thread
* so the scrape's herdr round-trip never stalls the status poller. The captured waiter is read on the
* poller thread (before any next-turn delivery can overwrite it) and passed into the virtual thread.
*/
public final class CompletionResolver implements TurnListener {
private static final Logger log = LoggerFactory.getLogger(CompletionResolver.class);
/**
* herdr {@code agent.read} source for the completion scrape. {@code recent} returns the tail of
* the transcript (the worker's last output), which is what a delegator wants when the worker
* didn't structure a reply.
*/
static final String SCRAPE_SOURCE = "recent";
/** Cap the scraped tail so a long transcript can't return an unbounded blob. */
static final int MAX_SCRAPE_CHARS = 4000;
private final AgentControl agents;
private final Rendezvous rendezvous;
/**
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
* delivering send opened, plus the assistant block present when it was delivered.
*
* <p>The {@code waiter} is what makes a late fallback safe (CB-116): we resolve it, not "whoever
* is waiting now", so a completion that fires after the next send has opened its own waiter is a
* no-op rather than a cross-turn stale reply. The {@code baseline} is the CB-115 staleness
* reference: a completion scrape equal to it means the worker produced no new output (the previous
* turn's wind-down sampled as this boundary), so it is suppressed. Overwritten on each delivery;
* cleared when the turn resolves. Package-private so tests can capture and replay a specific turn.
*/
record InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline) {
}
private final ConcurrentHashMap<String, InFlight> inFlight = new ConcurrentHashMap<>();
public CompletionResolver(AgentControl agents, Rendezvous rendezvous) {
this.agents = agents;
this.rendezvous = rendezvous;
}
@Override
public void onDelivered(String target) {
// Capture the exact waiter this turn belongs to (CB-116) and snapshot the pane's pre-turn
// content — what it shows *before* the just-delivered turn produces output — as the staleness
// reference (CB-115). Done synchronously (like the delivering send itself) so both are in
// place before this turn's completion can fire.
captureBaseline(target);
}
/** Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link #onDelivered}). */
void captureBaseline(String target) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null) {
inFlight.remove(target); // no send is waiting on this delivery — nothing to resolve later
return;
}
String baseline;
try {
// Clip to the same cap resolve() applies to the tail (line ~134): the CB-115 misattribution
// guard compares baseline.equals(tail), so both sides must be the same capped representation.
// An unclipped baseline vs a clipped tail would never match for a >MAX_SCRAPE_CHARS block,
// defeating the guard and letting a stale completion resolve the send.
baseline = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
} catch (RuntimeException e) {
baseline = null; // fail open: no baseline ⇒ no suppression
log.debug("delivery baseline for {} failed: {}", target, e.getMessage());
}
inFlight.put(target, new InFlight(waiter, baseline));
}
/** The turn currently baselined for {@code target}, or {@code null} — a test hook for the captureBaseline path. */
InFlight inFlight(String target) {
return inFlight.get(target);
}
@Override
public void onTurnComplete(String target) {
// Read the in-flight turn on the poller thread — before any next-turn delivery can overwrite
// it — then off-load the scrape (a herdr round-trip we must not block polling on) to a vthread.
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("completion-" + target).start(() -> resolve(target, turn));
}
@Override
public void onTurnFailed(String target) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn));
}
/** Synchronous resolve (the unit-testable core of {@link #onTurnComplete}). */
void resolve(String target, InFlight turn) {
CompletableFuture<Rendezvous.Resolution> waiter = turn == null ? null : turn.waiter();
if (waiter == null || waiter.isDone()) {
// Nobody is blocked on THIS turn (it had no send, or its bridge_reply already won). Skip
// the scrape; resolving the current waiter here would be the CB-116 cross-turn stale reply.
inFlight.remove(target, turn);
return;
}
String tail;
boolean scrapeFailed = false;
try {
tail = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
} catch (RuntimeException e) {
// The worker finished but we couldn't read its screen — still resolve the send so the
// caller unblocks; an empty tail beats hanging until the caller's timeout.
log.warn("completion scrape for {} failed; resolving with an empty tail: {}",
target, e.getMessage());
tail = "";
scrapeFailed = true;
}
// Misattribution guard (CB-115): if the scrape is byte-identical to the pane content at
// delivery, this turn produced no new output — the boundary belongs to the previous turn's
// wind-down (common on rapid back-to-back sends). Suppress rather than resolve the send with
// a stale answer; the real bridge_reply (or a later genuine completion) resolves it instead.
// A scrape that failed to read is exempt — an empty tail there is "couldn't see", not "no change".
String baseline = turn.baseline();
if (!scrapeFailed && baseline != null && baseline.equals(tail)) {
log.debug("suppressing misattributed completion for {} (no output change since delivery)",
target);
return; // keep the in-flight record: a later genuine completion still needs it
}
if (rendezvous.resolveCompletion(waiter, tail)) {
inFlight.remove(target, turn);
log.debug("resolved send to {} via turn-completion fallback ({} chars scraped)",
target, tail.length());
}
}
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
void fail(String target, InFlight turn) {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fall back to the currently-registered waiter (unambiguous — that send never completed, so
// no next turn exists to confuse it with).
CompletableFuture<Rendezvous.Resolution> waiter =
turn != null ? turn.waiter() : rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
inFlight.remove(target, turn); // nobody blocked on this worker — nothing to fail
return;
}
String reason;
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
}
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.debug("failed send to {} via turn-stall fallback", target);
}
}
private static String clip(String s) {
if (s == null) return "";
String trimmed = s.strip();
return trimmed.length() <= MAX_SCRAPE_CHARS
? trimmed
: trimmed.substring(trimmed.length() - MAX_SCRAPE_CHARS);
}
/**
* Extract the last assistant message from a raw Claude Code pane scrape (CB-115). Claude Code
* prefixes each assistant turn with {@code ⏺}; the delegator wants that answer, not the TUI
* chrome around it. Take everything from the final {@code ⏺} onward and stop at the <em>first</em>
* hard interface boundary below it — the spinner/status line, input box, {@code ❯} prompt (which
* may echo the <em>next</em> turn's text), footer, or tips/warnings. Stopping at the first
* boundary (rather than trimming only trailing chrome) is what keeps a following turn's echoed
* prompt out of this reply. Blank lines are not boundaries, so a multi-paragraph answer survives;
* trailing blanks are trimmed at the end. With no {@code ⏺} marker (an unusual render) the whole
* text is scanned the same way, so we never lose the reply.
*
* <p>Package-private and pure so it is unit-testable without herdr.
*/
static String lastAssistantBlock(String raw) {
if (raw == null || raw.isBlank()) return "";
int marker = raw.lastIndexOf('⏺');
String block = marker >= 0 ? raw.substring(marker + 1) : raw;
StringBuilder out = new StringBuilder();
int kept = 0;
for (String line : block.split("\n", -1)) {
if (isBoundary(line)) break; // first TUI boundary ends the assistant message
if (kept++ > 0) out.append('\n');
out.append(line);
}
return out.toString().strip();
}
/**
* A hard TUI boundary line that marks the end of an assistant message and the start of interface
* chrome (input box, prompt, spinner, footer, tips/warnings). Blank lines are <em>not</em>
* boundaries — an answer may contain them — so they are kept and trimmed only if trailing.
*/
private static boolean isBoundary(String line) {
String t = line.strip();
if (t.isEmpty()) return false;
// A horizontal rule / all box-drawing separators (e.g. "──────").
if (t.chars().allMatch(c -> c == '─' || c == '—' || c == '━' || c == '═' || c == '-')) {
return true;
}
String lower = t.toLowerCase();
return t.startsWith("╭") || t.startsWith("│") || t.startsWith("╰") || t.startsWith("┌")
|| t.startsWith("└") || t.startsWith("❯") || t.startsWith("⏵")
|| t.startsWith("⎿") || t.startsWith("⚠")
// Status/spinner lines Claude Code renders below a settled or in-flight turn,
// e.g. "✻ Baked for 21s", "✶ Forming…".
|| t.startsWith("✻") || t.startsWith("✳") || t.startsWith("✽") || t.startsWith("·")
|| t.startsWith("●") || t.startsWith("◐") || t.startsWith("✢") || t.startsWith("✶")
|| lower.contains("auto mode") || lower.contains("for shortcuts")
|| lower.contains("esc to interrupt") || lower.contains("bypass permissions");
}
}
@@ -12,6 +12,8 @@ import java.util.List;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Consumer;
import java.util.function.Predicate;
import java.util.stream.Collectors;
/**
@@ -32,6 +34,14 @@ import java.util.stream.Collectors;
* a transient {@code unknown} — counts as a real pickup, so a detection glitch can't prematurely
* release the latch. Perfectly reliable turn boundaries require a herdr {@code events.subscribe}
* stream; that is the intended upgrade and would replace only the sampling, not this queue.
*
* <p><strong>Turn completion (CB-106).</strong> Beyond delivery, the injector reports when a
* delegated turn <em>finishes</em>: after a delivery is picked up (a real {@code working} sample),
* the next injectable sample is a confirmed {@code working → idle} boundary and fires
* {@link TurnListener#onTurnComplete}. Completion is only ever synthesized from a <em>confirmed</em>
* turn — the pickup-grace path (a turn too fast to sample) unwedges the queue but does not fire
* completion, since without a sampled {@code working} there is no trustworthy "the worker just
* finished the task" signal to act on.
*/
public final class Injector {
@@ -45,11 +55,68 @@ public final class Injector {
*/
private static final int PICKUP_GRACE_POLLS = 8;
/**
* How many consecutive {@code unknown} samples while a delegation is outstanding before we
* declare it stalled and fire {@link TurnListener#onTurnFailed} (CB-109). A worker wedged in a
* state herdr can't classify (e.g. an API-error screen) stays {@code unknown} indefinitely and
* would otherwise never resolve; any {@code working}/{@code idle} sample resets the streak, so a
* transient detection glitch cannot trip it. At the 250ms poll interval this is ~30s — far longer
* than any real detection blip, and still vastly better than the async send's timeout.
*/
private static final int TURN_STALL_GRACE_POLLS = 120;
/**
* How many consecutive injectable samples a queued-but-undelivered message may wait on the
* {@link #ready} gate before we give up and fail it (CB-114). The gate holds a message out of a
* worker's boot window (herdr reports {@code idle} while its Claude is still starting), but a
* worker whose Claude crashes during boot — or never connects the bridge MCP — stays "idle and
* not ready" forever: {@link #ready} never accepts it, the message is never delivered, and the
* target would be polled indefinitely with its caller's future never completing. After this
* grace the queued messages are failed and the target released. At the 250ms poll interval this
* is ~60s — deliberately longer than {@link #TURN_STALL_GRACE_POLLS}, since a first boot (spawn
* + model load + MCP connect) legitimately takes longer than an in-turn detection blip.
*/
private static final int READINESS_GRACE_POLLS = 240;
private final AgentControl agents;
private final TurnListener turnListener;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
private final Consumer<String> forget; // CB-114: clear a gone worker's readiness/presence
private final ConcurrentHashMap<String, Target> targets = new ConcurrentHashMap<>();
/** Delivery only; completion signalling is a no-op and every target is treated as available. */
public Injector(AgentControl agents) {
this(agents, TurnListener.NOOP);
}
/** Delivery plus turn-completion signalling (CB-106); every target is treated as available. */
public Injector(AgentControl agents, TurnListener turnListener) {
this(agents, turnListener, _ -> true);
}
/**
* Delivery, completion signalling (CB-106), and a readiness gate (CB-113): a message is delivered
* only when {@code ready} accepts the target — i.e. the worker's Claude has connected the bridge
* MCP. This holds the first delivery out of the worker's boot window, where herdr already reports
* {@code idle} but the TUI would drop an injected paste.
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready) {
this(agents, turnListener, ready, _ -> {
});
}
/**
* Delivery, completion signalling (CB-106), a readiness gate (CB-113), and readiness cleanup
* (CB-114): {@code forget} is invoked with a target when its worker is gone — dropped
* (pane crash) or timed out on the readiness gate — so its stale presence/readiness is cleared
* and does not linger past the worker's life.
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
this.agents = agents;
this.turnListener = turnListener;
this.ready = ready;
this.forget = forget;
}
/** A pending message and the future that completes when it has been delivered. */
@@ -59,8 +126,12 @@ public final class Injector {
/** Per-worker delivery state, guarded by its own monitor (single writer per worker). */
private static final class Target {
final Deque<Pending> queue = new ArrayDeque<>();
boolean awaitingPickup; // sent a message, waiting for the worker to pick it up
boolean awaitingPickup; // sent a message, waiting for the worker to pick it up
int injectableSincePickup; // consecutive injectable samples while awaitingPickup
boolean awaitingCompletion; // a delivered message's turn is not yet known-complete
boolean turnObserved; // saw a real `working` sample since that delivery (turn ran)
int unknownSinceTurn; // consecutive `unknown` samples while a delegation is outstanding (CB-109)
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
synchronized void add(Pending p) {
queue.add(p);
@@ -99,63 +170,154 @@ public final class Injector {
Pending sent = null;
RuntimeException sendError = null;
boolean turnCompleted = false;
boolean turnFailed = false;
boolean resubmit = false;
List<Pending> notReady = null; // queued messages failed because the worker never became ready
synchronized (t) {
if (status == AgentStatus.WORKING) {
// Definitive pickup: the worker is busy on our last message.
// Definitive pickup: the worker is busy on our last message, and (if a delivery is
// outstanding) a real turn is now confirmed to be running.
t.awaitingPickup = false;
t.injectableSincePickup = 0;
t.unknownSinceTurn = 0;
t.notReadySincePoll = 0;
if (t.awaitingCompletion) t.turnObserved = true;
} else if (status.injectable()) { // IDLE or BLOCKED
if (t.awaitingPickup && ++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
// Pickup edge was never sampled (turn faster than the poll, or status lag).
// The worker has plainly moved on — release the latch rather than wedge.
t.awaitingPickup = false;
t.injectableSincePickup = 0;
t.unknownSinceTurn = 0;
if (t.awaitingPickup) {
if (++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
// Pickup edge was never sampled (turn faster than the poll, or status lag).
// Release the latch rather than wedge — and give up on synthesizing a
// completion for this message, since without a confirmed `working` we cannot
// trust that a task-processing turn actually ran.
t.awaitingPickup = false;
t.injectableSincePickup = 0;
t.awaitingCompletion = false;
t.turnObserved = false;
} else {
// Delivered but still idle → the worker hasn't picked it up; the submit
// keystroke likely raced the paste (esp. right as the TUI became ready).
// Re-nudge Enter (CB-113) until the worker starts (WORKING) or the grace ends.
resubmit = true;
}
}
if (!t.awaitingPickup) {
Pending p = t.queue.peek();
if (p != null) {
try {
agents.send(target, p.text());
t.queue.poll();
t.awaitingPickup = true;
t.injectableSincePickup = 0;
sent = p;
} catch (RuntimeException e) {
// Delivery failed at herdr; drop the poisoned message and surface it
// rather than blocking the queue behind it.
t.queue.poll();
sent = p;
sendError = e;
// A confirmed turn (a `working` sample was seen) that has now returned to idle is
// a trustworthy `working → idle` completion boundary.
if (t.awaitingCompletion && t.turnObserved) {
t.awaitingCompletion = false;
t.turnObserved = false;
turnCompleted = true;
}
// Deliver the next queued message only once the prior turn is fully settled, so a
// completion is never confused with the pickup of the following message — and only
// once the worker is available (CB-113), so we never paste into its boot window.
if (!t.awaitingCompletion) {
Pending p = t.queue.peek();
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
try {
agents.send(target, p.text());
t.queue.poll();
t.awaitingPickup = true;
t.awaitingCompletion = true;
t.turnObserved = false;
t.injectableSincePickup = 0;
sent = p;
} catch (RuntimeException e) {
// Delivery failed at herdr; drop the poisoned message and surface it
// rather than blocking the queue behind it.
t.queue.poll();
sent = p;
sendError = e;
}
} else if (p != null && ++t.notReadySincePoll >= READINESS_GRACE_POLLS) {
// The worker has been idle-but-not-ready for the whole grace: its Claude
// never connected the bridge MCP (crashed during boot, or wedged on a
// startup prompt). The readiness gate would hold this message forever, so
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
t.queue.clear();
t.notReadySincePoll = 0;
}
}
}
} else {
// UNKNOWN (or any other non-injectable, non-working): not a safe window nor a
// reliable pickup signal, so we never deliver or release the pickup latch here. But
// an outstanding delegation whose worker has gone unresponsive — stuck in a state
// herdr can't classify (CB-109) — will never yield a working→idle boundary. After a
// sustained streak, declare it failed so the awaiting send resolves rather than
// riding out the async timeout. (This also frees a delivery that wedged before it
// was ever picked up, which the injectable-only pickup grace could never release.)
if (t.awaitingCompletion && ++t.unknownSinceTurn >= TURN_STALL_GRACE_POLLS) {
t.awaitingPickup = false;
t.awaitingCompletion = false;
t.turnObserved = false;
t.unknownSinceTurn = 0;
turnFailed = true;
}
}
// UNKNOWN (and any other non-injectable, non-working): do nothing — neither a safe
// window nor a reliable pickup signal, so we must not deliver or release the latch.
// Reclaim the entry once the worker is fully quiescent, so the map cannot grow without
// bound across many short-lived workers.
if (t.queue.isEmpty() && !t.awaitingPickup) {
// Reclaim the entry once the worker is fully quiescent (nothing queued, no pickup or
// completion awaited), so the map cannot grow without bound across short-lived workers.
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion) {
targets.remove(target, t);
}
}
// Fire listeners / herdr calls after releasing the monitor so nothing runs on the poller
// thread while it holds the target lock.
if (resubmit) {
try {
agents.submit(target); // nudge a raced Enter so the pending paste submits
} catch (RuntimeException e) {
log.debug("resubmit to {} failed (will retry next poll): {}", target, e.getMessage());
}
}
if (notReady != null) {
// Worker never became available: forget its (never-set) readiness, unblock every queued
// caller, and route the awaiting send through the same failure path as a stalled turn so
// a blocking or async waiter resolves WORKER_FAILED rather than riding out the timeout.
forget.accept(target);
RuntimeException cause = new IllegalStateException(
target + " never became available (no bridge MCP connection within the boot window)");
for (Pending p : notReady) {
p.delivered().completeExceptionally(cause);
}
turnListener.onTurnFailed(target);
}
if (turnCompleted) {
turnListener.onTurnComplete(target);
}
if (turnFailed) {
turnListener.onTurnFailed(target);
}
if (sent != null) {
if (sendError != null) {
log.warn("inject to {} failed, dropped message: {}", target, sendError.getMessage());
sent.delivered().completeExceptionally(sendError);
} else {
// Baseline the pane's pre-turn content so a misattributed completion (no new output)
// can't resolve this send with the previous turn's stale answer (CB-115).
turnListener.onDelivered(target);
sent.delivered().complete(null);
}
}
}
/** Targets the poller must keep sampling: those with a queued message or an awaited pickup. */
/**
* Targets the poller must keep sampling: those with a queued message, an awaited pickup, or an
* awaited turn completion (so the {@code working → idle} boundary is observed).
*/
public Set<String> activeTargets() {
return targets.entrySet().stream()
.filter(e -> {
synchronized (e.getValue()) {
return !e.getValue().queue.isEmpty() || e.getValue().awaitingPickup;
Target t = e.getValue();
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion;
}
})
.map(java.util.Map.Entry::getKey)
@@ -163,20 +325,31 @@ public final class Injector {
}
/**
* Forget a target whose worker is gone, failing every still-queued message so awaiting
* callers unblock instead of hanging forever. Futures are completed after the monitor is
* released.
* Forget a target whose worker is gone, failing every still-queued message so awaiting callers
* unblock instead of hanging forever. If a message had already been <em>delivered</em> but its
* turn was not yet resolved (CB-110 — the worker vanished mid-turn, e.g. its pane crashed), fire
* {@link TurnListener#onTurnFailed} for it: a delivered message is no longer in the queue, so
* failing queued waiters alone would leave that send's rendezvous hanging until the async
* timeout. Futures and listeners are completed after the monitor is released.
*/
public void drop(String target, Throwable cause) {
Target t = targets.remove(target);
if (t == null) return;
List<Pending> pending;
boolean hadDeliveredTurn;
synchronized (t) {
pending = new ArrayList<>(t.queue);
t.queue.clear();
hadDeliveredTurn = t.awaitingCompletion;
t.awaitingCompletion = false;
t.awaitingPickup = false;
}
forget.accept(target); // the worker is gone — clear its readiness/presence too (CB-114)
for (Pending p : pending) {
p.delivered().completeExceptionally(cause);
}
if (hadDeliveredTurn) {
turnListener.onTurnFailed(target);
}
}
}
@@ -23,13 +23,20 @@ public final class StatusPoller {
private final AgentControl agents;
private final Injector injector;
private final StatusRefiner refiner;
private final long intervalMillis;
private volatile boolean running;
private Thread thread;
public StatusPoller(AgentControl agents, Injector injector, long intervalMillis) {
this(agents, injector, new StatusRefiner(agents), intervalMillis);
}
public StatusPoller(AgentControl agents, Injector injector, StatusRefiner refiner,
long intervalMillis) {
this.agents = agents;
this.injector = injector;
this.refiner = refiner;
this.intervalMillis = intervalMillis;
}
@@ -47,7 +54,9 @@ public final class StatusPoller {
for (String target : active) {
if (!running) return;
try {
AgentStatus status = agents.status(target);
// herdr's agent_status can misreport a settled worker as `unknown`; refine it
// against the pane content before it drives delivery/completion (CB-115).
AgentStatus status = refiner.refine(target, agents.status(target));
injector.onStatus(target, status);
} catch (HerdrException e) {
// The worker's agent is gone — stop trying and unblock its waiters.
@@ -0,0 +1,89 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
/**
* Refines an unreliable {@link AgentStatus#UNKNOWN} into a real state by reading the worker's
* terminal content (CB-115).
*
* <p>Some workers' panes are misclassified by herdr as {@code unknown} even when the worker is
* plainly settled at an idle prompt (empty {@code ❯}, "auto mode on" footer, a completed
* {@code ⏺} answer above). Left as {@code UNKNOWN} that both <em>wedges delivery</em> — the
* status-gated {@link Injector} only injects into an {@link AgentStatus#injectable} worker — and
* <em>mis-fires the CB-109 stall failure</em> on a worker that has actually answered. herdr's
* {@code agent_status} is a heuristic; the pane content is the ground truth.
*
* <p>The refinement only ever runs on a raw {@code UNKNOWN} sample (every other status is trusted
* as-is), so a healthy worker adds zero extra herdr traffic; a persistently-{@code unknown} worker
* costs one extra {@code agent.read} per poll while it has work outstanding. Classification is
* deliberately conservative — it upgrades {@code UNKNOWN} to {@link AgentStatus#WORKING} or
* {@link AgentStatus#IDLE} only on a clear signal, and leaves a genuinely unclassifiable screen
* (e.g. a wedged error state) as {@code UNKNOWN} so the CB-109 stall path can still fail it.
*/
public final class StatusRefiner {
private static final Logger log = LoggerFactory.getLogger(StatusRefiner.class);
/**
* herdr {@code agent.read} source used to inspect the pane. {@code detection} is the region
* herdr itself uses for status detection (the prompt/footer tail), which is exactly what we
* need to tell "idle at prompt" from "mid-turn".
*/
static final String PROBE_SOURCE = "detection";
private final AgentControl agents;
public StatusRefiner(AgentControl agents) {
this.agents = agents;
}
/**
* Return a trustworthy status for {@code target}. Any non-{@code UNKNOWN} {@code raw} is returned
* unchanged; an {@code UNKNOWN} triggers a pane read and content classification. A read failure
* leaves it {@code UNKNOWN} (the safe default: no delivery, and the stall path still applies).
*/
public AgentStatus refine(String target, AgentStatus raw) {
if (raw != AgentStatus.UNKNOWN) return raw;
String pane;
try {
pane = agents.read(target, PROBE_SOURCE);
} catch (RuntimeException e) {
log.debug("status refine read for {} failed; leaving UNKNOWN: {}", target, e.getMessage());
return AgentStatus.UNKNOWN;
}
AgentStatus refined = classify(pane);
if (refined != AgentStatus.UNKNOWN) {
log.debug("refined {} from UNKNOWN to {} via pane content", target, refined);
}
return refined;
}
/**
* Classify a Claude Code TUI pane tail. Package-private and pure so it is unit-testable without
* herdr.
*
* <ul>
* <li>An active-generation marker ({@code esc to interrupt}) ⇒ {@link AgentStatus#WORKING} —
* never inject here.</li>
* <li>Otherwise, an interactive input prompt with no active-turn marker ({@code ❯}, the
* {@code │ >} input box, or the idle {@code auto mode} / shortcuts footer) ⇒
* {@link AgentStatus#IDLE} — settled and safe to inject / a completed turn.</li>
* <li>Anything else (blank, or an unrecognizable screen) ⇒ {@link AgentStatus#UNKNOWN}.</li>
* </ul>
*/
static AgentStatus classify(String pane) {
if (pane == null || pane.isBlank()) return AgentStatus.UNKNOWN;
String lower = pane.toLowerCase();
// Claude Code shows "(esc to interrupt)" only while a turn is actively generating.
if (lower.contains("esc to interrupt")) return AgentStatus.WORKING;
// A settled, ready input prompt with no active-turn marker = idle-at-prompt.
boolean readyPrompt = pane.contains("❯")
|| pane.contains("│ >")
|| lower.contains("auto mode on")
|| lower.contains("? for shortcuts");
return readyPrompt ? AgentStatus.IDLE : AgentStatus.UNKNOWN;
}
}
@@ -0,0 +1,39 @@
package dev.ltms.bridged.inject;
/**
* Notified when a worker's delegated turn is observed to complete — a confirmed
* {@code WORKING → IDLE} transition after a delivery. This is the CB-106 completion signal the
* {@code CompletionResolver} uses to resolve a blocked send whose worker never called
* {@code bridge_reply}. Kept as a seam so the {@link Injector} needs no dependency on the message
* layer and stays unit-testable with a capturing fake.
*/
@FunctionalInterface
public interface TurnListener {
/** A worker's delegated turn finished (worker returned to idle after visibly working). */
void onTurnComplete(String target);
/**
* A worker that visibly ran a delegated turn then wedged in a non-idle, non-working state
* (CB-109) — e.g. an error screen herdr classifies as {@code unknown} — so no
* {@code working → idle} completion boundary will ever arrive. A default no-op keeps this a
* functional interface; the completion resolver overrides it to fail the awaiting send.
*/
default void onTurnFailed(String target) {
}
/**
* A message was just delivered into {@code target}'s pane (CB-115). Fired so the completion
* resolver can snapshot the pane's pre-turn content: a later {@link #onTurnComplete} whose
* scrape is unchanged from this baseline is a <em>misattributed</em> boundary (e.g. the prior
* turn's wind-down sampled as this turn's completion on rapid back-to-back sends) and must not
* resolve the send with the previous turn's stale answer. A default no-op keeps the interface
* functional for callers that don't scrape.
*/
default void onDelivered(String target) {
}
/** No-op default for callers that only need delivery, not completion signalling. */
TurnListener NOOP = _ -> {
};
}
@@ -0,0 +1,38 @@
package dev.ltms.bridged.inject;
import java.util.concurrent.ConcurrentHashMap;
import java.util.Set;
/**
* Tracks which workers are <em>available</em> — their Claude has booted and connected its MCP client
* to the bridge (CB-113). This is the reliable readiness signal, unlike herdr's {@code agent_status},
* which reports {@code idle} for a worker whose Claude is still booting. Delivering into that boot
* window pastes into a not-yet-ready TUI (the text is lost) and wedges the worker's delivery state,
* so the {@link Injector} holds the first delivery until the worker is present here.
*
* <p>Populated from the MCP transport: any MCP request whose connection resolves to a worker terminal
* marks that worker present (its {@code initialize} is the first such contact). A worker that never
* mounts the bridge MCP is never marked present — its sends stay queued until they time out, which is
* correct (it could not have replied anyway).
*/
public class WorkerPresence {
private final Set<String> present = ConcurrentHashMap.newKeySet();
/** Record that {@code terminal}'s worker has connected its MCP client (is available). */
public void markPresent(String terminal) {
if (terminal != null && !terminal.isBlank()) {
present.add(terminal);
}
}
/** Whether {@code terminal}'s worker is available (has been seen on the bridge MCP). */
public boolean isPresent(String terminal) {
return present.contains(terminal);
}
/** Forget a torn-down worker so its terminal id does not linger as "present". */
public void forget(String terminal) {
present.remove(terminal);
}
}
@@ -0,0 +1,797 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.AuditLog;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.auth.Role;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.peer.PeerLauncher;
import io.modelcontextprotocol.common.McpTransportContext;
import io.modelcontextprotocol.json.McpJsonMapper;
import io.modelcontextprotocol.json.jackson3.JacksonMcpJsonMapperSupplier;
import io.modelcontextprotocol.server.McpServer;
import io.modelcontextprotocol.server.McpSyncServer;
import io.modelcontextprotocol.server.McpSyncServerExchange;
import io.modelcontextprotocol.server.transport.HttpServletStreamableServerTransportProvider;
import io.modelcontextprotocol.spec.McpSchema;
import com.fasterxml.jackson.databind.ObjectMapper;
import jakarta.servlet.http.HttpServlet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.function.Function;
import java.util.stream.Collectors;
/**
* The MCP SERVER face (CB-105): a Streamable-HTTP MCP server whose tools are <em>thin adapters</em>
* over the same {@link MessageService}/{@link Rendezvous} the REST routes use — so the two are
* validated by parity, not by re-implementing behaviour. The primary Opus calls {@code bridge_send}
* / {@code bridge_status}; the worker calls {@code bridge_reply}.
*
* <p>Beyond delegation the primary also manages the fleet here (CB-108): {@code bridge_spawn} /
* {@code bridge_list} / {@code bridge_stop} drive the {@link PeerLauncher} SPI so a worker's whole
* lifecycle is managed through MCP, with each adapter's subscription boundary enforced inside it.
*
* <p>The tool <em>logic</em> lives in package-private static methods returning a
* {@link McpSchema.CallToolResult}, so it is unit-testable without standing up the HTTP transport;
* the SDK owns the wire protocol. Mount {@link #servlet()} at {@code /mcp} on the daemon's Jetty.
*/
public final class BridgeMcp {
private static final long DEFAULT_TIMEOUT_MS = 25_000;
private static final long MAX_TIMEOUT_MS = 120_000;
// bridge_ask blocks the WORKER's own MCP call, which its client caps near 60s — default under
// that so the bridge returns a clean timeout before the client severs the call (CB-205).
private static final long ASK_DEFAULT_TIMEOUT_MS = 55_000;
private static final long ASK_MAX_TIMEOUT_MS = 115_000;
private static final ObjectMapper MAPPER = new ObjectMapper(); // worker-view JSON projections
/** Transport-context key under which the extractor stashes the resolved caller identity. */
static final String CALLER_TERMINAL = "callerTerminal";
/** Transport-context key under which the extractor stashes the caller's PID (for cwd inherit). */
static final String CALLER_PID = "callerPid";
/** Transport-context key under which the extractor stashes the resolved {@link Role} (CB-501). */
static final String CALLER_ROLE = "callerRole";
private final HttpServletStreamableServerTransportProvider transport;
private final McpSyncServer server;
private final CallerResolver authz; // CB-501: null → authorization not enforced (legacy)
private final Metrics metrics; // CB-502: null → auth failures not counted
/**
* Legacy constructor — no authorization. Retained so existing tests exercise tool behaviour
* without an auth fixture.
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
PrimaryRegistry primaryRegistry) {
this(messages, workers, sessions, identity, presence, primaryRegistry, null, null);
}
/**
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
* Jetty's context handler and never passes through Javalin's {@code before}
* filter, so the REST guard does not cover it.
* @param metrics registry for auth-failure counting; may be {@code null}
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
PrimaryRegistry primaryRegistry, CallerResolver callers, Metrics metrics) {
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
.jsonMapper(json)
.mcpEndpoint("/mcp")
// Resolve the caller from the connection (peer PID → herdr pane) in one lookup: the
// worker terminal for bridge_reply (no spoofable arg), and the PID so bridge_spawn can
// inherit the primary's cwd (CB-112). Any contact from a worker marks it available
// (CB-113) — its MCP initialize is the reliable "the agent is up" signal.
.contextExtractor(req -> {
// One resolution per call, shared with the REST surface via CallerResolver so
// the two paths cannot drift on who a caller is.
Principal p = callers != null
? callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization"))
: legacyPrincipal(identity, req.getRemoteAddr(), req.getRemotePort());
presence.markPresent(p.terminal()); // no-op for the primary (null terminal)
return McpTransportContext.create(Map.of(
CALLER_TERMINAL, orEmpty(p.terminal()),
CALLER_PID, Long.toString(p.pid()),
CALLER_ROLE, p.role().name()));
})
.build();
this.server = McpServer.sync(transport)
.serverInfo("bridge", "0.1.0")
.capabilities(McpSchema.ServerCapabilities.builder().tools(true).build())
.toolCall(sendTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SEND,
str(req.arguments(), "sessionId"));
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
Map<String, Object> a = req.arguments();
String turnId = str(a, "turnId");
if (turnId != null && !turnId.isBlank()) {
// Answering a worker's bridge_ask (CB-205): resolve its blocked question and
// block for the worker's reply as it resumes the same turn.
return answer(messages, turnId, str(a, "content"), timeoutMs(a));
}
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, str(a, "sessionId"), str(a, "content"))
: send(messages, str(a, "sessionId"), str(a, "content"), timeoutMs(a));
})
// bridge_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
.toolCall(replyTool(), (exchange, req) -> {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.REPLY, self);
if (denied != null) return denied;
return reply(messages, self, str(req.arguments(), "content"));
})
// bridge_ask (CB-205): a worker's mid-turn question — identity from the CONNECTION.
.toolCall(askTool(), (exchange, req) -> {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.ASK, self);
if (denied != null) return denied;
return ask(messages, self, str(req.arguments(), "question"), timeoutMs(req.arguments()));
})
.toolCall(statusTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return status(messages, str(req.arguments(), "sessionId"));
})
.toolCall(pollTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
Map<String, Object> a = req.arguments();
return poll(messages, str(a, "ticket"), str(a, "target"));
})
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
// Acking removes a reply from the inbox, so it is a drain, not a read.
.toolCall(ackTool(), (exchange, req) -> {
Map<String, Object> a = req.arguments();
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.DRAIN, str(a, "target"));
if (denied != null) return denied;
return ack(messages, str(a, "target"), str(a, "msgId"));
})
// Fleet management (CB-108): spawn/list/stop over ClaudeCodeLauncher.
.toolCall(spawnTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SPAWN, null);
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
Map<String, Object> a = req.arguments();
// CB-112: worker inherits the primary's cwd unless the call pins one.
// CB-301: carry the caller's identity as the session owner (null for the primary).
// CB-301-ext: optional isolated worktree for parallel implementers.
String callerCwd = identity.cwdForPid(callerPid(exchange));
return spawn(sessions, str(a, "profile"), str(a, "cwd"), callerCwd,
callerTerminal(exchange), worktreeRequest(a));
})
.toolCall(listTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listWorkers(workers, sessions);
})
.toolCall(stopTool(), (exchange, req) -> {
String paneId = str(req.arguments(), "paneId");
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.STOP, paneId);
if (denied != null) return denied;
return stop(sessions, paneId);
})
.toolCall(profilesTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return profiles(workers);
})
.toolCall(whoamiTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return whoami(principal(exchange), sessions);
})
.build();
this.authz = callers;
this.metrics = metrics;
}
/**
* Pre-CB-501 identity: worker if the connection maps to a pane, otherwise the primary. Used
* only by the legacy constructor, where authorization is not enforced anyway.
*/
private static Principal legacyPrincipal(ConnectionIdentity identity, String addr, int port) {
ConnectionIdentity.Caller c = identity.resolve(addr, port);
return c.terminal() != null
? Principal.worker(c.terminal(), c.pid())
: Principal.primary(c.pid());
}
/** The caller reconstructed from the transport context. */
private static Principal principal(McpSyncServerExchange exchange) {
return principalFrom(exchange.transportContext().get(CALLER_ROLE),
callerTerminal(exchange), callerPid(exchange));
}
/**
* Rebuild a {@link Principal} from the three values the context extractor stashed.
*
* <p>Split out from {@link #principal(McpSyncServerExchange)} so the identity rules are
* reachable without an {@code McpSyncServerExchange} — that is an SDK type this project has no
* mocking library to fabricate, which is why this logic had no test at all until CB-513.
*
* @param role the stashed {@link Role} name, or {@code null} on the legacy path
* @param terminal the worker terminal, or {@code null} for a non-worker
* @param pid the calling pid, or {@code -1}
*/
static Principal principalFrom(Object role, String terminal, long pid) {
if (role == null) {
// No role stashed (legacy path): fall back to the historical interpretation.
return terminal != null ? Principal.worker(terminal, pid) : Principal.primary(pid);
}
return new Principal(Role.valueOf(role.toString()), terminal, pid);
}
/**
* Gate a tool call on the CB-505 table. Returns {@code null} when the call may proceed, or the
* error result to return when it may not.
*/
private McpSchema.CallToolResult deny(McpSyncServerExchange exchange, Authz.Action action,
String target) {
return denyFor(principal(exchange), action, target);
}
/**
* The policy half of {@link #deny}: everything except pulling the caller out of the MCP
* exchange. Kept separate so the authorization decision — the actual control — is unit-testable
* without fabricating an SDK {@code McpSyncServerExchange}.
*
* <p>This surface exists because the enforcement was previously unreachable from a test: no
* test constructs a {@code BridgeMcp}, so the whole MCP-side gate ran zero times in the suite
* while the REST-side equivalent had ten tests. A security control nothing exercises is a
* claim, not a control.
*
* @return {@code null} when the call may proceed, or the error result to return when it may not
*/
McpSchema.CallToolResult denyFor(Principal caller, Authz.Action action, String target) {
// The enforcement switch lives HERE rather than in the exchange-facing wrapper: any future
// tool that calls this directly must not be able to skip the gate by accident.
if (authz == null) {
return null; // legacy constructor: authorization not enforced
}
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
return null;
}
String reason = Authz.isUnauthenticated(caller) ? "unauthenticated" : "forbidden";
AuditLog.denied(caller, action, target, reason);
if (metrics != null) {
metrics.inc(BridgedMetrics.AUTH_FAILURES, "reason", reason);
}
return error(reason + ": " + caller.describe() + " may not " + action);
}
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
private static String callerTerminal(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_TERMINAL);
String s = v == null ? null : v.toString();
return (s == null || s.isBlank()) ? null : s;
}
/** The caller's PID resolved from this call's connection, or {@code -1} if unknown. */
private static long callerPid(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_PID);
try {
return v == null ? -1 : Long.parseLong(v.toString());
} catch (NumberFormatException e) {
return -1;
}
}
private static String orEmpty(String s) {
return s == null ? "" : s;
}
/** The Streamable-HTTP servlet to mount at {@code /mcp} on the daemon's Jetty. */
public HttpServlet servlet() {
return transport;
}
/** Graceful shutdown of the MCP server. */
public void close() {
server.closeGracefully();
}
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
/** {@code bridge_send}: delegate {@code content} to a worker session and block for its reply. */
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content, Long timeoutMs) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
try {
return formatReply(messages.send(sessionId, content, timeout), timeout);
} catch (HerdrException e) {
return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
}
}
/**
* {@code bridge_send} carrying a {@code turnId}: the primary's answer to a worker's
* {@code bridge_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as
* it resumes the same turn — surfaced to the primary identically to a normal send.
*/
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs) {
if (isBlank(turnId) || isBlank(content)) {
return error("turnId and content are required to answer a worker's question");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
return formatReply(messages.answer(turnId, content, timeout), timeout);
}
/**
* {@code bridge_ask} (CB-205): a worker pauses its delegated turn to ask the primary, blocking
* until the primary answers. The worker is identified by its connection ({@code callerTerminal}),
* never an argument — a {@code null} means the caller is not a known worker.
*/
static McpSchema.CallToolResult ask(MessageService messages, String callerTerminal, String question, Long timeoutMs) {
if (callerTerminal == null) {
return error("bridge_ask is for workers only — could not identify the calling worker "
+ "from the connection");
}
if (isBlank(question)) {
return error("question is required");
}
long timeout = Math.clamp(timeoutMs == null ? ASK_DEFAULT_TIMEOUT_MS : timeoutMs, 1, ASK_MAX_TIMEOUT_MS);
MessageService.AskResult r = messages.ask(callerTerminal, question, timeout);
return switch (r.outcome()) {
case ANSWERED -> text(r.answer());
case NO_WAITER -> error("no primary is awaiting this turn — bridge_ask only works while a "
+ "bridge_send delegation is open to answer it");
case TIMED_OUT -> text("[no answer within " + timeout + "ms — the primary did not respond; "
+ "proceed on your best judgement, then call bridge_reply to end the turn]");
};
}
/** Render a {@link MessageService.Reply} as a tool result — shared by {@link #send} and {@link #answer}. */
private static McpSchema.CallToolResult formatReply(MessageService.Reply r, long timeout) {
return switch (r.outcome()) {
case REPLIED -> text(r.text());
// The worker's turn finished but it never called bridge_reply — hand back the scraped
// transcript tail, flagged so the primary knows it isn't a structured reply.
case COMPLETED_UNREPLIED -> text(
"[worker finished without a structured bridge_reply — transcript tail follows]\n" + r.text());
// The worker ran the turn then wedged (CB-109) — surface the error context.
case WORKER_FAILED -> text("[worker failed — turn ended in an unrecoverable state]\n" + r.text());
// The worker paused mid-turn to ask (CB-205) — tell the primary how to answer in-turn.
case QUESTION -> text("[question] the worker paused to ask before it can finish:\n" + r.text()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + r.turnId()
+ "\" and content set to your answer; the worker resumes the same turn.");
case STALE_TURN -> error("that question is no longer open — it timed out or was already "
+ "answered (turnId stale)");
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> text("[no reply within " + timeout + "ms — worker "
+ r.outcome().name().toLowerCase().replace("timed_out_", "") + "; retry or poll status]");
};
}
/**
* {@code bridge_send} with {@code wait:false}: delegate {@code content} and return a ticket
* immediately (fire-and-poll), so a long task isn't cut off by the caller's MCP call timeout.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
String ticket = messages.sendAsync(sessionId, content);
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
}
/** {@code bridge_poll}: check an async delegation by ticket, or drain a worker's inbox by target. */
static McpSchema.CallToolResult poll(MessageService messages, String ticket, String target) {
if (!isBlank(target)) {
var replies = messages.drainReplies(target);
if (replies.isEmpty()) {
return text("[]");
}
return text(json(replies));
}
if (isBlank(ticket)) {
return error("ticket (or target) is required");
}
MessageService.TaskView v = messages.poll(ticket);
if (v == null) {
return error("unknown ticket: " + ticket + " (never issued, or expired)");
}
return switch (v.phase()) {
case DONE -> text(v.replySource() != null && v.replySource().equals("transcript")
? "[done — worker finished without a structured bridge_reply; transcript tail follows]\n" + v.reply()
: v.reply());
case PENDING -> text("[pending — " + v.detail() + "]");
case FAILED -> text("[failed — " + v.detail() + "]");
};
}
/**
* {@code bridge_reply}: the worker returns its structured answer, resolving the awaiting send
* or — when no send is open — queueing the reply in the inbox for later drain (CB-307).
* {@code callerTerminal} is resolved from the connection (never an argument); a {@code null}
* means the caller is not a known worker (e.g. the primary called it by mistake).
*/
static McpSchema.CallToolResult reply(MessageService messages, String callerTerminal, String content) {
if (callerTerminal == null) {
return error("bridge_reply is for workers only — could not identify the calling worker "
+ "from the connection");
}
if (content == null) {
return error("content is required");
}
messages.reply(callerTerminal, content);
return text("delivered");
}
/** {@code bridge_ack}: acknowledge (remove) a specific reply from the inbox. */
static McpSchema.CallToolResult ack(MessageService messages, String target, String msgId) {
if (isBlank(target) || isBlank(msgId)) {
return error("target and msgId are required");
}
messages.ackReply(target, msgId);
return text("acknowledged " + msgId);
}
/** {@code bridge_status}: the live lifecycle status of a worker session. */
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
if (isBlank(sessionId)) {
return error("sessionId is required");
}
try {
return text(messages.status(sessionId).name().toLowerCase());
} catch (HerdrException e) {
return error("herdr error for session " + sessionId + ": " + e.getMessage());
}
}
/**
* {@code bridge_whoami}: the caller's own identity, as the daemon already resolved it.
*
* <p>Every other tool <em>consumes</em> this identity — the authorization gate, the reply
* rendezvous, the cwd inherit — but none reported it, so an agent had to infer its own role
* from side channels the daemon does not control: a charter string in its system prompt, the
* name its MCP mount happens to carry, or {@code ANTHROPIC_BASE_URL} (which Claude-model
* workers do not set). The failure mode of guessing is asymmetric and silent: a primary that
* mistakes itself for a worker is refused by {@link Authz} and learns immediately, while a
* worker that mistakes itself for the primary ends its turn without {@code bridge_reply} and
* the sender simply receives nothing. This tool removes the guess.
*
* <p>For a worker the session registry adds what it knows about that session. A worker the
* registry has no record of — one that outlived a daemon restart — still gets its role and
* {@code sessionId}, which is the load-bearing part.
*/
static McpSchema.CallToolResult whoami(Principal caller, SessionManager sessions) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("role", caller.role().name().toLowerCase());
if (!caller.isWorker()) {
return text(json(m));
}
m.put("sessionId", caller.terminal());
sessions.roster().stream()
.filter(s -> caller.terminal().equals(s.terminalId()))
.findFirst()
.ifPresent(s -> {
m.put("paneId", s.paneId());
m.put("profile", s.profile());
m.put("state", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
if (s.ownerTerminal() != null) {
m.put("owner", s.ownerTerminal());
}
});
return text(json(m));
}
// --- fleet management logic (CB-108 / CB-301) --------------------------------------------
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
return spawn(sessions, profile, null, null, null, null);
}
/**
* {@code bridge_spawn}: launch a guard-checked worker for {@code profile} (blank → the default
* profile) and return its session id + pane id. The worker's cwd is {@code requestedCwd} if given,
* else the profile's config, else {@code callerCwd} (the primary's directory), else the daemon's.
* CB-301: the session is registered with {@code ownerTerminal} as its owner.
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
*/
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile,
String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest worktreeRequest) {
try {
WorkerSession worker = sessions.acquire(isBlank(profile) ? null : profile,
requestedCwd, callerCwd, ownerTerminal, worktreeRequest);
return text(json(workerView(worker)));
} catch (GuardException e) {
return error("subscription boundary: " + e.getMessage());
} catch (IllegalArgumentException e) {
return error(e.getMessage()); // unknown / no-default profile
} catch (PeerUnreachableException e) {
return error("spawn timed out — worker pane never reached injectable state: " + e.getMessage());
} catch (HerdrException e) {
return error("herdr error spawning worker: " + e.getMessage());
}
}
/** Build a {@link WorktreeRequest} from {@code bridge_spawn}'s optional {@code worktree}/{@code ticket} args. */
private static WorktreeRequest worktreeRequest(Map<String, Object> a) {
Object w = a.get("worktree");
if (w == null || Boolean.FALSE.equals(w)) {
return null;
}
String ticket = str(a, "ticket");
if (w instanceof String s) {
if (s.isBlank() || "false".equalsIgnoreCase(s)) {
return null;
}
if ("true".equalsIgnoreCase(s)) {
if (isBlank(ticket)) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return new WorktreeRequest(s, null);
}
if (w instanceof Boolean b && b) {
if (isBlank(ticket)) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return null;
}
/** {@code bridge_profiles}: the configured worker profiles and the default. */
static McpSchema.CallToolResult profiles(PeerLauncher workers) {
return text(json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
}
/**
* {@code bridge_list}: bridge-owned roster merged with live herdr status. CB-519 decoupled the
* registry key (a host-unique id) from the herdr pane coordinate, so the join is on the
* terminal id, which both the session and the live agent carry.
*/
static McpSchema.CallToolResult listWorkers(PeerLauncher workers, SessionManager sessions) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
return text(json(Map.of("workers", out)));
} catch (HerdrException e) {
return error("herdr error listing workers: " + e.getMessage());
}
}
/** {@code bridge_stop}: tear a worker down by its pane id. */
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
if (isBlank(paneId)) {
return error("paneId is required");
}
try {
sessions.release(paneId);
return text("stopped " + paneId);
} catch (HerdrException e) {
return error("herdr error stopping " + paneId + ": " + e.getMessage());
}
}
/** CB-301 projection from the authoritative session registry. */
private static Map<String, Object> workerView(WorkerSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", s.terminalId());
m.put("paneId", s.paneId());
m.put("status", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
return m;
}
private static String json(Object o) {
try {
return MAPPER.writeValueAsString(o);
} catch (Exception e) {
return String.valueOf(o);
}
}
// --- tool schemas --------------------------------------------------------------------------
private static McpSchema.Tool sendTool() {
return tool("bridge_send",
"Delegate a task to a worker session. By default blocks until the worker replies and "
+ "returns its reply (or a 'still working / queued' note on timeout). Pass wait:false "
+ "for a long task to return a ticket immediately, then poll it with bridge_poll. To "
+ "answer a worker's bridge_ask, pass its turnId (with content) instead of sessionId.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id (herdr terminal_id) to delegate to"),
"content", stringProp("The task/message to send to the worker (or your answer, with turnId)"),
"timeoutMs", Map.of("type", "integer", "description", "Max ms to wait for a reply (blocking mode)"),
"wait", Map.of("type", "boolean",
"description", "Block for the reply (default true); false returns a ticket to poll"),
"turnId", stringProp("When answering a worker's bridge_ask, its question turnId — "
+ "routes your answer back into the same turn (omit for a normal delegation)")),
List.of("content")));
}
private static McpSchema.Tool askTool() {
// No target/session arg — the worker's identity is resolved from the connection.
return tool("bridge_ask",
"Pause your current delegated turn to ask the primary a question, blocking until it "
+ "answers — then resume the same turn with the answer. Use this when only the "
+ "primary has a decision or detail you need to continue. You do not address the "
+ "primary; identity is resolved from your connection.",
objectSchema(Map.of(
"question", stringProp("The question to put to the primary"),
"timeoutMs", Map.of("type", "integer",
"description", "Max ms to wait for the primary's answer")),
List.of("question")));
}
private static McpSchema.Tool pollTool() {
return tool("bridge_poll",
"Check an async delegation (a bridge_send with wait:false) by its ticket: "
+ "pending, done (with the worker's reply), or failed. When target (a worker "
+ "session id) is present instead of ticket, drain that worker's inbox of "
+ "replies delivered when no send was open.",
objectSchema(Map.of(
"ticket", stringProp("The ticket returned by bridge_send wait:false"),
"target", stringProp("Worker session id to drain pending replies from (optional)")),
List.of()));
}
private static McpSchema.Tool ackTool() {
return tool("bridge_ack",
"Acknowledge (remove) a specific reply from a worker's inbox. Use when the primary "
+ "has processed a reply and wants to confirm it, leaving other pending replies "
+ "in the inbox for later drain.",
objectSchema(Map.of(
"target", stringProp("Worker session id whose inbox to ack from"),
"msgId", stringProp("The message id to acknowledge")),
List.of("target", "msgId")));
}
private static McpSchema.Tool spawnTool() {
return tool("bridge_spawn",
"Spawn a new off-subscription worker session. Pass a profile (from bridge_profiles) to "
+ "pick the backend, or omit it for the default. The worker opens your current "
+ "directory by default; pass cwd to pin a different one. Pass worktree:true (with "
+ "ticket) or worktree:<ticket-slug> to provision an isolated git worktree. "
+ "Returns the worker's sessionId (use with bridge_send) and paneId (use with bridge_stop).",
objectSchema(Map.of(
"profile", stringProp("Worker profile to spawn (omit for the default profile)"),
"cwd", stringProp("Working directory for the worker (omit to inherit yours)"),
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
"ticket", stringProp("Ticket slug when worktree:true")),
List.of()));
}
private static McpSchema.Tool profilesTool() {
return tool("bridge_profiles",
"List the configured worker profiles (backends) and which one bridge_spawn uses by default.",
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool listTool() {
return tool("bridge_list",
"List the worker sessions the bridge tracks — each with its sessionId, paneId, profile, "
+ "state, optional worktree/branch/owner, and live herdr status.",
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool stopTool() {
return tool("bridge_stop",
"Tear down a worker session by its paneId (from bridge_spawn or bridge_list).",
objectSchema(Map.of(
"paneId", stringProp("The worker's paneId to stop")),
List.of("paneId")));
}
private static McpSchema.Tool replyTool() {
// No session/target arg — the worker's identity is resolved from the connection.
return tool("bridge_reply",
"Return your structured answer for the task you were delegated, "
+ "resolving the caller's blocked bridge_send.",
objectSchema(Map.of(
"content", stringProp("Your reply/answer")),
List.of("content")));
}
private static McpSchema.Tool statusTool() {
return tool("bridge_status",
"Get the live lifecycle status (idle/working/blocked/unknown) of a worker session.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id to query")),
List.of("sessionId")));
}
private static McpSchema.Tool whoamiTool() {
return tool("bridge_whoami",
"Report who YOU are on the bridge — your role is resolved from your connection "
+ "(unforgeable), never from anything you claim. Returns role 'primary' (you "
+ "orchestrate: spawn/send/stop, and you must never call bridge_reply) or "
+ "'worker' (you were delegated to: you must end every turn with exactly one "
+ "bridge_reply, and cannot spawn or send), plus your own sessionId, profile, "
+ "worktree and branch when you are a worker. Call this first when following "
+ "role-conditional instructions rather than guessing your role.",
objectSchema(Map.of(), List.of()));
}
// --- small helpers -------------------------------------------------------------------------
// The SDK 2.0.0 deprecates its own Tool builders without a stable replacement — isolate it here.
@SuppressWarnings("deprecation")
private static McpSchema.Tool tool(String name, String description, Map<String, Object> inputSchema) {
return McpSchema.Tool.builder(name).description(description).inputSchema(inputSchema).build();
}
private static Map<String, Object> objectSchema(Map<String, Object> properties, List<String> required) {
return Map.of("type", "object", "properties", properties, "required", required);
}
private static Map<String, Object> stringProp(String description) {
return Map.of("type", "string", "description", description);
}
private static McpSchema.CallToolResult text(String s) {
return McpSchema.CallToolResult.builder().addTextContent(s == null ? "" : s).build();
}
private static McpSchema.CallToolResult error(String s) {
return McpSchema.CallToolResult.builder().addTextContent(s).isError(true).build();
}
private static String str(Map<String, Object> args, String key) {
Object v = args.get(key);
return v == null ? null : v.toString();
}
private static Long timeoutMs(Map<String, Object> args) {
Object v = args.get("timeoutMs");
return v instanceof Number n ? n.longValue() : null;
}
private static long clamp(long ms) {
return Math.clamp(ms, 1, MAX_TIMEOUT_MS);
}
private static boolean isBlank(String s) {
return s == null || s.isBlank();
}
}
@@ -0,0 +1,66 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.herdr.PaneLocator;
/**
* Resolves <em>who is calling</em> an MCP tool from the connection alone — the anti-spoofing
* identity model of the MCP contract. It ties the connection's loopback peer PID (from the OS)
* to a herdr agent pane (from herdr), yielding the caller's worker {@code terminal_id}. A caller
* that maps to no worker pane — the primary, or an off-host client — resolves to {@code null}.
*
* <p>Both sources are authoritative and unforgeable: the OS reports the real connecting PID, and
* herdr owns the PID→pane mapping. A worker cannot claim to be another worker, nor the primary.
* Single-host only (the herd shares the {@code bridged} host); the token path is the split-host
* fallback.
*/
public final class ConnectionIdentity {
private final PaneLocator panes;
private final PeerPidLookup pids;
private final ProcessCwdLookup cwds;
/** Identity only (no cwd resolution — {@link #cwdForPid} returns {@code null}). */
public ConnectionIdentity(PaneLocator panes, PeerPidLookup pids) {
this(panes, pids, _ -> null);
}
/** Identity plus cwd resolution (CB-112 — inherit the primary's directory on spawn). */
public ConnectionIdentity(PaneLocator panes, PeerPidLookup pids, ProcessCwdLookup cwds) {
this.panes = panes;
this.pids = pids;
this.cwds = cwds;
}
/**
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
* primary / an off-host client) and its {@code pid} (or {@code -1} if not resolvable).
*/
public record Caller(String terminal, long pid) {
}
/** Resolve the caller's terminal and PID from one peer-PID lookup. */
public Caller resolve(String remoteAddr, int remotePort) {
if (!isLoopback(remoteAddr)) {
return new Caller(null, -1); // only same-host callers can be workers
}
long pid = pids.pidForLocalPort(remotePort);
return new Caller(panes.terminalForPid(pid), pid);
}
/**
* The calling worker's {@code terminal_id}, or {@code null} if the caller is not a known
* on-host worker (treat as the primary).
*/
public String callerTerminal(String remoteAddr, int remotePort) {
return resolve(remoteAddr, remotePort).terminal();
}
/** The working directory of {@code pid} (the primary's cwd on an MCP spawn), or {@code null}. */
public String cwdForPid(long pid) {
return pid > 0 ? cwds.cwdForPid(pid) : null;
}
private static boolean isLoopback(String addr) {
return "127.0.0.1".equals(addr) || "::1".equals(addr) || "0:0:0:0:0:0:0:1".equals(addr);
}
}
@@ -0,0 +1,58 @@
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.BufferedReader;
import java.io.InputStreamReader;
import java.nio.charset.StandardCharsets;
import java.util.concurrent.TimeUnit;
/**
* {@link PeerPidLookup} via {@code lsof} (present on macOS and Linux). For a loopback TCP source
* {@code port}, both the client and this daemon appear on that port — so we exclude our own PID
* and take the other end, which is the calling process.
*/
public final class LsofPeerPidLookup implements PeerPidLookup {
private static final Logger log = LoggerFactory.getLogger(LsofPeerPidLookup.class);
private final long selfPid = ProcessHandle.current().pid();
@Override
public long pidForLocalPort(int port) {
try {
Process p = new ProcessBuilder("lsof", "-nP", "-FpP", "-iTCP:" + port)
.redirectErrorStream(true).start();
long found = -1;
try (BufferedReader r = new BufferedReader(
new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
long current = -1;
String line;
// -Fp emits records: a 'p<pid>' line, then the ports/files under that pid.
while ((line = r.readLine()) != null) {
if (line.startsWith("p")) {
current = parse(line.substring(1));
} else if (current > 0 && current != selfPid) {
found = current; // first process on this port that isn't us = the client
}
}
}
if (!p.waitFor(2, TimeUnit.SECONDS)) {
p.destroyForcibly();
}
return found;
} catch (Exception e) {
log.debug("lsof peer-pid lookup for port {} failed: {}", port, e.getMessage());
return -1;
}
}
private static long parse(String s) {
try {
return Long.parseLong(s.trim());
} catch (NumberFormatException e) {
return -1;
}
}
}
@@ -0,0 +1,48 @@
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.BufferedReader;
import java.io.InputStreamReader;
import java.nio.charset.StandardCharsets;
import java.util.concurrent.TimeUnit;
/**
* {@link ProcessCwdLookup} via {@code lsof} (present on macOS and Linux): {@code lsof -a -p <pid>
* -d cwd -Fn} prints the process's cwd on the {@code n…} line. Used to inherit the primary's
* working directory for a spawned worker (CB-112).
*/
public final class LsofProcessCwdLookup implements ProcessCwdLookup {
private static final Logger log = LoggerFactory.getLogger(LsofProcessCwdLookup.class);
@Override
public String cwdForPid(long pid) {
if (pid <= 0) {
return null;
}
try {
Process p = new ProcessBuilder("lsof", "-a", "-p", Long.toString(pid), "-d", "cwd", "-Fn")
.redirectErrorStream(true).start();
String cwd = null;
try (BufferedReader r = new BufferedReader(
new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
String line;
while ((line = r.readLine()) != null) {
if (line.startsWith("n")) { // 'n<path>' is the file-name field for the cwd fd
cwd = line.substring(1);
break;
}
}
}
if (!p.waitFor(2, TimeUnit.SECONDS)) {
p.destroyForcibly();
}
return (cwd == null || cwd.isBlank()) ? null : cwd;
} catch (Exception e) {
log.debug("lsof cwd lookup for pid {} failed: {}", pid, e.getMessage());
return null;
}
}
}
@@ -0,0 +1,13 @@
package dev.ltms.bridged.mcp;
/**
* Resolves the OS PID that owns a loopback TCP source port — the OS half of connection-based MCP
* identity. Java exposes no peer PID for a TCP socket, so this shells out. Injectable so
* {@link ConnectionIdentity} is testable without a real connection.
*/
@FunctionalInterface
public interface PeerPidLookup {
/** The PID whose socket has local (source) {@code port} on loopback, or {@code -1} if unknown. */
long pidForLocalPort(int port);
}
@@ -0,0 +1,68 @@
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Optional;
import java.util.concurrent.atomic.AtomicReference;
/**
* Single-slot, thread-safe registry for the primary's herdr {@code terminal_id}.
*
* <p>Populated from the caller terminal of orchestration-side MCP tools
* ({@code bridge_send}, {@code bridge_spawn}) — tools that only the primary calls.
* A pinned terminal (from config) seeds the registry at construction and makes
* subsequent {@link #record(String)} calls no-ops.
*
* <p>The push loop ({@code ReplyPushLoop}) uses {@link #isKnown()} to decide
* whether active nudging is possible; an empty registry means the primary is
* off-host or non-herdr and delivery falls back to pull.
*/
public final class PrimaryRegistry {
private static final Logger log = LoggerFactory.getLogger(PrimaryRegistry.class);
private final AtomicReference<String> terminal = new AtomicReference<>();
private final boolean pinned;
/**
* @param pinnedTerminal an optional pinned terminal from config ({@code null}/blank = unpinned)
*/
public PrimaryRegistry(String pinnedTerminal) {
if (pinnedTerminal != null && !pinnedTerminal.isBlank()) {
this.terminal.set(pinnedTerminal);
this.pinned = true;
log.info("primary terminal pinned: {}", pinnedTerminal);
} else {
this.pinned = false;
}
}
/**
* Record a terminal_id. No-op when:
* <ul>
* <li>the registry is pinned (config override),
* <li>{@code terminalId} is {@code null} or blank (non-herdr caller).
* </ul>
*/
public void record(String terminalId) {
if (pinned) return;
if (terminalId == null || terminalId.isBlank()) return;
String prev = terminal.getAndSet(terminalId);
if (prev == null) {
log.debug("primary terminal learned: {}", terminalId);
} else if (!prev.equals(terminalId)) {
log.debug("primary terminal changed: {} -> {}", prev, terminalId);
}
}
/** The known primary terminal, or empty if not yet learned (and not pinned). */
public Optional<String> primaryTerminal() {
return Optional.ofNullable(terminal.get());
}
/** {@code true} once a terminal has been recorded (or was pinned at construction). */
public boolean isKnown() {
return terminal.get() != null;
}
}
@@ -0,0 +1,13 @@
package dev.ltms.bridged.mcp;
/**
* Resolves a process's current working directory from its PID — the OS half of CB-112's
* "a worker inherits the primary's directory." Injectable so {@link ConnectionIdentity} stays
* testable without shelling out.
*/
@FunctionalInterface
public interface ProcessCwdLookup {
/** The working directory of {@code pid}, or {@code null} if unknown. */
String cwdForPid(long pid);
}
@@ -0,0 +1,96 @@
package dev.ltms.bridged.metrics;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import java.util.LinkedHashMap;
import java.util.Map;
/**
* The daemon's metric definitions (CB-502) — one place where every series is named, described, and
* (for gauges) bound to live state.
*
* <p>The set is deliberately small: each series maps to a failure mode this project has actually
* hit, not to whatever was easy to count. The two worth watching in practice are
* {@code bridged_sends_total{outcome="completion_fallback"}} — a rising share means turn detection
* is degrading, the CB-115/116/118 failure family — and
* {@code bridged_push_nudges_total{outcome="exhausted"}}, which means the primary stopped draining
* its inbox and CB-307's active push gave up.
*/
public final class BridgedMetrics {
/** Counter: delegated sends by terminal outcome. */
public static final String SENDS = "bridged_sends_total";
/** Counter: worker replies by the path that carried them (rendezvous vs stranded-to-inbox). */
public static final String REPLIES = "bridged_replies_total";
/** Counter: push-loop nudges to the primary, by outcome. */
public static final String PUSH_NUDGES = "bridged_push_nudges_total";
/** Counter: spawn attempts by peer kind and outcome. */
public static final String SPAWNS = "bridged_spawns_total";
/** Counter: herdr socket calls by method and outcome. */
public static final String HERDR_CALLS = "bridged_herdr_calls_total";
/** Counter: rejected requests by reason (CB-501). */
public static final String AUTH_FAILURES = "bridged_auth_failures_total";
/** Gauge: session census by lifecycle state. */
public static final String SESSIONS = "bridged_sessions";
/** Gauge: undrained replies held per target. */
public static final String INBOX_DEPTH = "bridged_inbox_depth";
private BridgedMetrics() {
}
/**
* Build the registry with its help text and live gauges bound.
*
* @param sessions the authoritative session registry (census gauge)
* @param inbox the reply inbox; only used for a depth gauge when it can be inspected
*/
public static Metrics create(SessionManager sessions, ReplyInbox inbox) {
Metrics m = new Metrics();
m.describe(SENDS, "counter",
"Delegated sends by terminal outcome (replied|completion_fallback|timeout|failed).");
m.describe(REPLIES, "counter",
"Worker replies by delivery path (rendezvous=resolved an open send, inbox=stranded and held).");
m.describe(PUSH_NUDGES, "counter",
"CB-307 push-loop nudges to the primary (delivered|exhausted).");
m.describe(SPAWNS, "counter",
"Worker spawn attempts by peer kind and outcome (ready|timeout|guard_rejected).");
m.describe(HERDR_CALLS, "counter",
"herdr socket calls by method and outcome — the dependency everything else rests on.");
m.describe(AUTH_FAILURES, "counter",
"Requests refused by CB-501/505 (unauthenticated|forbidden).");
m.describe(SESSIONS, "gauge",
"Registered worker sessions by lifecycle state.");
m.describe(INBOX_DEPTH, "gauge",
"Replies held for a target that the primary has not drained. Steady state is 0; "
+ "a target stuck above 0 means CB-307 delivery is not completing.");
// One gauge per state so a scrape shows the whole census even when a state is empty —
// an absent series and a zero series read very differently on a dashboard.
for (WorkerSession.State state : WorkerSession.State.values()) {
String label = state.name().toLowerCase();
m.gauge(SESSIONS, () -> countIn(sessions, state), "state", label);
}
// Depth is per live session, so the label set is only known at scrape time. peek() is the
// port's non-destructive read — scraping metrics must never ack a reply out of the inbox.
m.collector(INBOX_DEPTH, "target", () -> {
Map<String, Number> depths = new LinkedHashMap<>();
for (WorkerSession s : sessions.roster()) {
String target = s.terminalId();
if (target == null) {
continue;
}
depths.put(target, inbox.peek(target).size());
}
return depths;
});
return m;
}
private static long countIn(SessionManager sessions, WorkerSession.State state) {
return sessions.roster().stream().filter(s -> s.state() == state).count();
}
}
@@ -0,0 +1,181 @@
package dev.ltms.bridged.metrics;
import java.util.Map;
import java.util.NavigableMap;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ConcurrentSkipListMap;
import java.util.concurrent.atomic.LongAdder;
import java.util.function.Supplier;
/**
* The daemon's metric registry and Prometheus text renderer (CB-502).
*
* <p>Deliberately dependency-free. The roadmap's tech-stack table specified Micrometer, but this
* pom already carries an unusual reconciliation burden (a hand-pinned {@code jackson-annotations}
* to make the MCP SDK's Jackson 3 coexist with our Jackson 2, a Jetty BOM import to stop version
* skew, and four documented accepted-CVE advisories), and the dependency CVE gate this project
* mandates could not be run when this landed. The metric set is small and fully known, and
* Prometheus text exposition is a stable, well-specified format — so the registry is ~100 lines
* here instead of a new transitive tree. {@code GET /metrics} is the swap seam if Micrometer's
* ecosystem is ever wanted.
*
* <p>Thread-safe: counters are {@link LongAdder} (built for contended increment), gauges are
* supplier-backed so they read live state at scrape time rather than needing to be pushed.
*/
public final class Metrics {
/** Counter series, keyed by the fully-rendered {@code name{labels}} sample id. */
private final NavigableMap<String, LongAdder> counters = new ConcurrentSkipListMap<>();
/** Gauge series, evaluated at scrape time. */
private final NavigableMap<String, Supplier<Number>> gauges = new ConcurrentSkipListMap<>();
/** Gauge families whose label set is only known at scrape time, keyed by metric name. */
private final NavigableMap<String, Collector> collectors = new ConcurrentSkipListMap<>();
/** HELP/TYPE metadata, keyed by bare metric name. */
private final Map<String, String[]> meta = new ConcurrentHashMap<>();
/** A gauge family whose series are discovered per scrape (one label, many values). */
private record Collector(String labelName, Supplier<Map<String, Number>> samples) {
}
/** Declare a metric's help text and type once, so the exposition carries HELP/TYPE lines. */
public Metrics describe(String name, String type, String help) {
meta.put(name, new String[]{type, help});
return this;
}
/** Increment a counter by one. */
public void inc(String name, String... labelPairs) {
add(name, 1, labelPairs);
}
/** Increment a counter by {@code delta}. */
public void add(String name, long delta, String... labelPairs) {
counters.computeIfAbsent(sample(name, labelPairs), _ -> new LongAdder()).add(delta);
}
/**
* Register a live gauge. The supplier is called at scrape time, so it reflects current state
* (session census, inbox depth) without anything having to remember to update it.
*/
public void gauge(String name, Supplier<Number> value, String... labelPairs) {
gauges.put(sample(name, labelPairs), value);
}
/**
* Register a gauge family whose label values are not known up front — inbox depth per target,
* for instance, where the set of targets changes as workers come and go. The supplier returns
* {@code labelValue → value} and is evaluated once per scrape.
*/
public void collector(String name, String labelName, Supplier<Map<String, Number>> samples) {
collectors.put(name, new Collector(labelName, samples));
}
/** Current value of a counter series — for assertions in tests. */
public long count(String name, String... labelPairs) {
LongAdder a = counters.get(sample(name, labelPairs));
return a == null ? 0 : a.sum();
}
/**
* Render the Prometheus text exposition format (version 0.0.4): optional {@code # HELP} and
* {@code # TYPE} lines per metric family, then one line per sample.
*/
public String render() {
StringBuilder out = new StringBuilder(1024);
String lastFamily = null;
for (Map.Entry<String, LongAdder> e : counters.entrySet()) {
lastFamily = emitHeader(out, e.getKey(), lastFamily);
out.append(e.getKey()).append(' ').append(e.getValue().sum()).append('\n');
}
for (Map.Entry<String, Supplier<Number>> e : gauges.entrySet()) {
lastFamily = emitHeader(out, e.getKey(), lastFamily);
Number v;
try {
v = e.getValue().get();
} catch (RuntimeException ex) {
continue; // a broken gauge must never break the whole scrape
}
if (v == null) {
continue;
}
out.append(e.getKey()).append(' ').append(format(v)).append('\n');
}
for (Map.Entry<String, Collector> e : collectors.entrySet()) {
Map<String, Number> samples;
try {
samples = e.getValue().samples().get();
} catch (RuntimeException ex) {
continue; // a broken collector must never break the whole scrape
}
if (samples == null || samples.isEmpty()) {
continue;
}
lastFamily = emitHeader(out, e.getKey(), lastFamily);
// Sort so repeated scrapes are byte-stable and diffable.
new java.util.TreeMap<>(samples).forEach((label, v) -> {
if (v != null) {
out.append(sample(e.getKey(), e.getValue().labelName(), label))
.append(' ').append(format(v)).append('\n');
}
});
}
return out.toString();
}
/** Emit HELP/TYPE when the sample starts a new metric family; returns the current family. */
private String emitHeader(StringBuilder out, String sampleId, String lastFamily) {
String family = familyOf(sampleId);
if (family.equals(lastFamily)) {
return lastFamily;
}
String[] m = meta.get(family);
if (m != null) {
out.append("# HELP ").append(family).append(' ').append(m[1]).append('\n');
out.append("# TYPE ").append(family).append(' ').append(m[0]).append('\n');
}
return family;
}
private static String familyOf(String sampleId) {
int brace = sampleId.indexOf('{');
return brace < 0 ? sampleId : sampleId.substring(0, brace);
}
/** Whole numbers render without a decimal point; everything else as-is. */
private static String format(Number v) {
double d = v.doubleValue();
return (d == Math.rint(d) && !Double.isInfinite(d))
? Long.toString((long) d)
: Double.toString(d);
}
/** Build the {@code name{k="v",k2="v2"}} sample id; labels are sorted for stable output. */
private static String sample(String name, String... labelPairs) {
if (labelPairs == null || labelPairs.length == 0) {
return name;
}
if (labelPairs.length % 2 != 0) {
throw new IllegalArgumentException("labels must be key/value pairs, got " + labelPairs.length);
}
NavigableMap<String, String> sorted = new java.util.TreeMap<>();
for (int i = 0; i < labelPairs.length; i += 2) {
sorted.put(labelPairs[i], labelPairs[i + 1] == null ? "" : labelPairs[i + 1]);
}
StringBuilder sb = new StringBuilder(name.length() + 16 * sorted.size());
sb.append(name).append('{');
boolean first = true;
for (Map.Entry<String, String> e : sorted.entrySet()) {
if (!first) {
sb.append(',');
}
first = false;
sb.append(e.getKey()).append("=\"").append(escapeLabel(e.getValue())).append('"');
}
return sb.append('}').toString();
}
/** Label values are escaped per the exposition format: backslash, quote, newline. */
private static String escapeLabel(String v) {
return v.replace("\\", "\\\\").replace("\"", "\\\"").replace("\n", "\\n");
}
}
@@ -0,0 +1,242 @@
package dev.ltms.bridged.msg;
import com.rabbitmq.client.AMQP;
import com.rabbitmq.client.Channel;
import com.rabbitmq.client.Connection;
import com.rabbitmq.client.ConnectionFactory;
import com.rabbitmq.client.DeliverCallback;
import com.rabbitmq.client.Recoverable;
import com.rabbitmq.client.RecoveryListener;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.concurrent.ConcurrentHashMap;
/**
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
* port {@link InMemoryReplyInbox} implements as soft state.
*
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target has a durable
* queue {@code agent.<target>.inbox}. The gateway that owns the target starts a manual-ack consumer
* ({@link #own}) that pulls persistent messages off that queue into an in-memory <em>held</em> map
* (keyed by {@code msgId}) but does <em>not</em> ack them. {@link #peek} returns that snapshot;
* {@link #ack} acks the broker delivery-tag and drops the entry. Because messages stay unacked until
* the owning gateway actually drains them, a crash (or a {@code java -jar} bounce) before caller-ack
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
* adapter cannot give, with the port contract preserved.
*
* <p><strong>Ownership is explicit.</strong> {@link #own} declares the queue and starts the consumer;
* {@link #release} cancels it. {@link #publish} sends to the queue but does <em>not</em> imply ownership
* and does not attach a consumer. This split is required by CB-308 federation, where one gateway may
* publish to an agent owned by another gateway; in that case the publisher must not compete for
* deliveries.
*
* <p><strong>Dedup.</strong> The consumer keys the held map by {@code msgId}; a redelivered duplicate
* (at-least-once, or a producer double-publish) is acked-and-dropped on arrival, so it never
* double-queues.
*
* <p><strong>Visibility.</strong> Unlike the in-memory adapter, publish → broker → consumer is
* asynchronous, so a {@link #peek} immediately after {@link #publish} may not yet see the message
* (broker delivery latency). Callers that need the reply drained poll (as the primary already does);
* the contract test waits for visibility. This is inherent to broker-backed delivery, not a defect.
*
* <p>The default deploy targets LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1 (URI-only swap),
* so the {@code @Tag("contract")} integration test runs against a RabbitMQ container.
*/
public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
private static final Logger log = LoggerFactory.getLogger(AmqpReplyInbox.class);
private static final String QUEUE_PREFIX = "agent.";
private static final String QUEUE_SUFFIX = ".inbox";
private final Connection connection;
private final Channel channel;
/** All channel operations (publish/declare/ack/cancel) serialize on this — a Channel is not thread-safe. */
private final Object channelLock = new Object();
/** target → (msgId → held delivery). Per-target map is guarded by synchronizing on itself. */
private final ConcurrentHashMap<String, LinkedHashMap<String, Held>> held = new ConcurrentHashMap<>();
/** Targets whose queue is declared and consumer is running, mapped to their broker consumer tag. */
private final ConcurrentHashMap<String, String> consumerTags = new ConcurrentHashMap<>();
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
private record Held(long deliveryTag, InboxMessage message) {}
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) and open the inbox. */
public static AmqpReplyInbox open(String uri) {
try {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"));
} catch (Exception e) {
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
}
}
/** Wrap an already-open connection (injection seam for the contract test). */
AmqpReplyInbox(Connection connection) {
this.connection = connection;
try {
this.channel = connection.createChannel();
} catch (IOException e) {
throw new IllegalStateException("cannot open AMQP channel", e);
}
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
// repopulates it with valid tags (dedup by msgId still prevents any double-queue).
if (connection instanceof Recoverable recoverable) {
recoverable.addRecoveryListener(new RecoveryListener() {
@Override
public void handleRecovery(Recoverable recoverable) {
held.clear();
log.info("AMQP connection recovered; cleared held replies for fresh redelivery");
}
@Override
public void handleRecoveryStarted(Recoverable recoverable) {
// no-op: we act once recovery completes
}
});
}
}
@Override
public void own(String target) {
synchronized (channelLock) {
if (consumerTags.containsKey(target)) {
return; // already owning this target
}
String queue = queueName(target);
try {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
String tag = channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
consumerTags.put(target, tag);
log.debug("AMQP inbox owns queue {} for target {}", queue, target);
} catch (IOException e) {
throw new IllegalStateException("cannot own queue " + queue, e);
}
}
}
@Override
public void release(String target) {
synchronized (channelLock) {
String tag = consumerTags.remove(target);
held.remove(target); // stale delivery tags must not survive release
if (tag == null) {
return;
}
try {
channel.basicCancel(tag);
} catch (IOException e) {
throw new IllegalStateException("cannot cancel consumer for " + target, e);
}
}
}
@Override
public void publish(String target, String msgId, String content) {
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
.messageId(msgId)
.deliveryMode(2) // persistent — survives a broker restart
.contentType("text/plain")
.build();
try {
synchronized (channelLock) {
channel.basicPublish("", queueName(target), props, content.getBytes(StandardCharsets.UTF_8));
}
} catch (IOException e) {
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
}
}
@Override
public List<InboxMessage> peek(String target) {
var perTarget = held.get(target);
if (perTarget == null) {
return List.of();
}
synchronized (perTarget) {
return perTarget.values().stream().map(Held::message).toList();
}
}
@Override
public void ack(String target, String msgId) {
var perTarget = held.get(target);
if (perTarget == null) {
return;
}
Held h;
synchronized (perTarget) {
h = perTarget.remove(msgId);
}
if (h == null) {
return; // never held (or already acked) — no-op
}
try {
synchronized (channelLock) {
channel.basicAck(h.deliveryTag(), false);
}
} catch (IOException e) {
// Ack didn't reach the broker: restore the entry so a later ack (or a redelivery after
// reconnect) can retry. Keeps the at-least-once contract — a reply is never silently lost.
synchronized (perTarget) {
perTarget.putIfAbsent(msgId, h);
}
throw new IllegalStateException("cannot ack reply " + msgId + " on " + queueName(target), e);
}
}
private DeliverCallback deliverCallback(String target) {
return (_, delivery) -> {
String msgId = delivery.getProperties().getMessageId();
long tag = delivery.getEnvelope().getDeliveryTag();
if (msgId == null || msgId.isBlank()) {
msgId = Long.toHexString(tag); // synthesize an id so dedup still has a key
}
String content = new String(delivery.getBody(), StandardCharsets.UTF_8);
var perTarget = held.computeIfAbsent(target, _ -> new LinkedHashMap<>());
boolean duplicate;
synchronized (perTarget) {
if (perTarget.containsKey(msgId)) {
duplicate = true;
} else {
perTarget.put(msgId, new Held(tag, new InboxMessage(msgId, target, content)));
duplicate = false;
}
}
if (duplicate) {
// Redelivered duplicate: ack the new tag and drop it so the broker stops resending.
synchronized (channelLock) {
channel.basicAck(tag, false);
}
}
};
}
private static String queueName(String target) {
return QUEUE_PREFIX + target + QUEUE_SUFFIX;
}
@Override
public void close() {
try {
channel.close();
} catch (Exception e) {
log.debug("AMQP channel close: {}", e.toString());
}
try {
connection.close();
} catch (Exception e) {
log.debug("AMQP connection close: {}", e.toString());
}
}
}
@@ -0,0 +1,74 @@
package dev.ltms.bridged.msg;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
/**
* Soft-state {@link ReplyInbox} backed by a {@link ConcurrentHashMap} keyed by target session.
* Per-target FIFO ordering (insertion order via {@link LinkedHashMap}). Dedup by {@code msgId}
* within a target. Thread-safe for concurrent publish vs. drain.
*
* <p><strong>Ownership is explicit.</strong> {@link #own} marks a target as locally owned so that
* {@link #peek} and {@link #ack} operate on it; {@link #publish} works whether or not the target is
* owned. {@link #release} clears the local snapshot. This mirrors the AMQP adapter's contract so the
* non-broker path stays interchangeable.
*
* <p><strong>This is soft-state, NOT persistence.</strong> Lost on a {@code java -jar} bounce — that
* is correct and consistent with "bridged stays soft-state." The Stage-2 AMQP adapter replaces this.
*/
public final class InMemoryReplyInbox implements ReplyInbox {
private final ConcurrentHashMap<String, LinkedHashMap<String, InboxMessage>> store = new ConcurrentHashMap<>();
private final Set<String> owned = ConcurrentHashMap.newKeySet();
@Override
public void own(String target) {
owned.add(target);
}
@Override
public void release(String target) {
owned.remove(target);
store.remove(target);
}
@Override
public void publish(String target, String msgId, String content) {
var perTarget = store.computeIfAbsent(target, _ -> new LinkedHashMap<>());
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
perTarget.putIfAbsent(msgId, new InboxMessage(msgId, target, content));
}
}
@Override
public List<InboxMessage> peek(String target) {
if (!owned.contains(target)) {
return List.of();
}
var perTarget = store.get(target);
if (perTarget == null) {
return List.of();
}
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
return List.copyOf(perTarget.values());
}
}
@Override
public void ack(String target, String msgId) {
if (!owned.contains(target)) {
return;
}
var perTarget = store.get(target);
if (perTarget != null) {
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
perTarget.remove(msgId);
}
}
}
}
@@ -0,0 +1,527 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.UUID;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.CompletionException;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
import java.util.concurrent.atomic.AtomicLong;
import java.util.concurrent.locks.ReentrantLock;
/**
* The blocking delegation feature (CB-104): deliver {@code content} into a worker and block until
* the worker returns a <em>structured reply</em> via {@code bridge_reply} (the {@link Rendezvous}),
* then hand that reply back. Delivery is the {@link Injector}'s job (the background poller sends it
* when the worker is injectable); this service never drives the injector or scrapes the terminal —
* completion is the worker's explicit reply, not a guess about {@code agent_status}.
*
* <p>Sends are serialized per session so exactly one reply can be outstanding per worker, which is
* what lets a reply map unambiguously to its send (no cross-talk between concurrent callers).
*
* <p>If the worker never replies within the timeout, the caller gets a typed "still working" /
* "queued" outcome — the message may still be mid-flight. A finished-but-unreplied turn is caught
* by the CB-106 completion fallback (see {@link Rendezvous#resolveCompletion}).
*
* <p><strong>Async fire-and-poll (CB-107).</strong> A caller's MCP client caps a blocking call at
* ~60s, but a real delegated task runs for minutes. {@link #sendAsync} therefore runs the same
* blocking {@link #send} on a background virtual thread and hands back a <em>ticket</em> the caller
* polls with {@link #poll}. The blocking and async paths share one code path (and the same per-target
* serialization), so async inherits the reply + completion resolution behaviour for free.
*/
public final class MessageService {
private static final Logger log = LoggerFactory.getLogger(MessageService.class);
/**
* The window a fire-and-poll send waits for resolution — generous, since no caller is blocked on
* it; a real delegated task resolves (reply or completion) well within this, and only a genuinely
* hung worker rides it out.
*/
private static final long ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L;
/** How long a finished (terminal) ticket is retained for polling before it is pruned. */
private static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
/** Outcome of a blocking send. */
public enum Outcome {
/** The worker called {@code bridge_reply}; {@code text} holds the structured answer. */
REPLIED,
/**
* The worker's delegated turn finished without a {@code bridge_reply} (CB-106 fallback);
* {@code text} is the scraped transcript tail rather than a structured answer.
*/
COMPLETED_UNREPLIED,
/**
* The worker ran the turn then wedged in an unrecoverable state (CB-109); {@code text} is the
* failure context (e.g. the error screen). Terminal, but not a successful completion.
*/
WORKER_FAILED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with
* {@link #answer(String, String, long)} and the turn resumes.
*/
QUESTION,
/** Timed out after the message was delivered — the worker is still working. */
TIMED_OUT_WORKING,
/** Timed out before delivery — the message is still queued for the worker. */
TIMED_OUT_QUEUED,
/** Another send to this session was in flight for the whole window. */
BUSY,
/**
* An answer ({@link #answer(String, String, long)}) referenced a {@code turnId} that is no
* longer open — the worker's {@code bridge_ask} already timed out or was answered.
*/
STALE_TURN
}
/**
* @param outcome how the send ended (or paused)
* @param text the worker's answer when {@link #completed()} (a structured {@code bridge_reply}
* for {@link Outcome#REPLIED}, a scraped transcript tail for
* {@link Outcome#COMPLETED_UNREPLIED}), or the question for {@link Outcome#QUESTION},
* else {@code null}
* @param turnId correlation id for a {@link Outcome#QUESTION} (answered via
* {@link #answer(String, String, long)}), else {@code null}
*/
public record Reply(Outcome outcome, String text, String turnId) {
/** A reply with no correlation id (the common terminal outcomes). */
public Reply(Outcome outcome, String text) {
this(outcome, text, null);
}
/** Whether the worker's turn actually finished with an answer (replied or scraped). */
public boolean completed() {
return outcome == Outcome.REPLIED || outcome == Outcome.COMPLETED_UNREPLIED;
}
}
/** How a worker's {@code bridge_ask} (CB-205) resolved. */
public enum AskOutcome {
/** The primary answered; {@link AskResult#answer} carries it. */
ANSWERED,
/** No delegation was open to surface the question to — the worker has no one to ask. */
NO_WAITER,
/** The primary did not answer within the window. */
TIMED_OUT
}
/** The outcome of a worker's {@code bridge_ask}: how it resolved and (if answered) the answer. */
public record AskResult(AskOutcome outcome, String answer) {
}
/** Lifecycle phase of an async delegation ticket. */
public enum Phase {
/** Delegated and in flight — queued for the worker or being worked. */
PENDING,
/** The worker's turn finished; {@link TaskView#reply} holds the answer. */
DONE,
/** The delegation could not complete (timed out, worker gone, or busy). */
FAILED
}
/**
* A poll snapshot of an async delegation.
*
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, else {@code null}
* @param replySource {@code "reply"} (structured {@code bridge_reply}) or {@code "transcript"}
* (completion scrape) when {@link Phase#DONE}, else {@code null}
* @param detail a human note (live worker status while pending, or the failure reason)
*/
public record TaskView(String ticket, Phase phase, String reply, String replySource, String detail) {
}
/** An in-flight or finished async delegation, keyed by its ticket. */
private record Task(String target, CompletableFuture<Reply> future, long createdNanos) {
}
private final AgentControl agents;
private final Injector injector;
private final Rendezvous rendezvous;
private final ReplyInbox inbox;
private final ReplyPushLoop pushLoop;
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
Thread.ofVirtual().name("bridge-async-", 0).factory());
/**
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
*
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
* branch ({@link #reply}) so it can nudge the primary to drain the inbox
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop) {
this(agents, injector, rendezvous, inbox, pushLoop, null);
}
/**
* As above, with a metric registry (CB-502). Instrumenting here rather than at the REST and MCP
* edges means both surfaces are counted by one piece of code and cannot drift.
*
* @param metrics nullable — when null, nothing is recorded
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop, Metrics metrics) {
this.agents = agents;
this.injector = injector;
this.rendezvous = rendezvous;
this.inbox = inbox;
this.pushLoop = pushLoop;
this.metrics = metrics;
}
/** Create with an explicit {@link ReplyInbox} and no push loop. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox) {
this(agents, injector, rendezvous, inbox, null);
}
/** Backward-compatible constructor that uses a default {@link InMemoryReplyInbox}. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous) {
this(agents, injector, rendezvous, new InMemoryReplyInbox());
}
/** Current lifecycle status of a worker (the {@code GET /sessions/{id}/status} surface). */
public AgentStatus status(String target) {
return agents.status(target);
}
/**
* Route a worker's explicit {@code bridge_reply}: resolve an open send, or queue it in the
* inbox if no send is currently open. Unlike the bare {@link Rendezvous#resolve}, a no-waiter
* result is <em>not</em> a failure — the reply is held for later drain.
*
* <p><strong>Do NOT use this for mid-turn questions.</strong> {@code bridge_ask} /
* {@link Rendezvous#resolveQuestion} must keep today's {@code NO_WAITER} behaviour — questions
* are interactive and must never be queued.
*
* @return always {@code true} — the reply either resolved a live send or was queued
*/
public boolean reply(String session, String content) {
if (rendezvous.resolve(session, content)) {
count(BridgedMetrics.REPLIES, "path", "rendezvous");
return true; // a live send took it — unchanged fast path
}
inbox.publish(session, UUID.randomUUID().toString(), content);
// A rising inbox share is the signal CB-307 exists to make visible: the worker finished but
// nobody was waiting, so delivery now depends on the push loop and a drain.
count(BridgedMetrics.REPLIES, "path", "inbox");
if (pushLoop != null) {
pushLoop.onReplyQueued(session);
}
return true; // held, not lost
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
if (metrics != null) {
metrics.inc(name, labels);
}
}
/** Count a send's terminal outcome and pass the reply through unchanged. */
private Reply recorded(Reply r) {
String label = sendOutcomeLabel(r.outcome());
if (label != null) {
count(BridgedMetrics.SENDS, "outcome", label);
}
return r;
}
/** Map a terminal send outcome to its metric label, or {@code null} for non-terminal ones. */
private static String sendOutcomeLabel(Outcome o) {
return switch (o) {
case REPLIED -> "replied";
case COMPLETED_UNREPLIED -> "completion_fallback";
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> "timeout";
case WORKER_FAILED -> "failed";
case STALE_TURN, QUESTION -> null; // not a completed delegation
};
}
/**
* Abandon any send still waiting on {@code target} because its session has gone away (CB-516).
*
* <p>Without this, tearing a worker down left its rendezvous waiter open: a blocking
* {@code bridge_send} kept blocking, and an async one kept reporting {@code PENDING} until
* {@link #ASYNC_TIMEOUT_MS} — thirty minutes — even though the worker provably no longer
* existed and the delegation could never complete. Worse, {@code poll} already had the evidence
* (it calls {@code liveStatus} to build its detail string and gets back {@code "unknown"}) and
* reported {@code PENDING} anyway.
*
* <p>Resolving the waiter as a failure — rather than letting it time out — also means the
* outcome is counted, so a torn-down delegation stops being invisible to {@code /metrics}.
*
* @return true if a live waiter was failed
*/
public boolean abandon(String target, String reason) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
return false; // nobody is blocked on this worker — nothing to abandon
}
boolean failed = rendezvous.resolveFailure(waiter, reason);
if (failed) {
log.debug("abandoned send to {}: {}", target, reason);
}
return failed;
}
/**
* Acknowledge a specific reply by {@code msgId} for {@code target}. Removes it from the inbox
* so that a subsequent drain or peek no longer returns it.
*/
public void ackReply(String target, String msgId) {
inbox.ack(target, msgId);
}
/**
* Drain (peek + ack) all pending inbox replies for {@code target}. At-least-once: returns the
* messages and acknowledges them; an in-flight failure between returning and the caller
* processing them re-surfaces them on a subsequent drain (the ack is local).
*
* @return the drained messages, newest last (FIFO); empty list if none
*/
public List<ReplyInbox.InboxMessage> drainReplies(String target) {
var messages = inbox.peek(target);
for (var msg : messages) {
inbox.ack(target, msg.msgId());
}
return messages;
}
/**
* Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
*/
public Reply send(String target, String content, long timeoutMillis) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
}
try {
CompletableFuture<Void> delivered = injector.enqueue(target, content);
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
} finally {
rendezvous.close(target, reply);
}
} finally {
lock.unlock();
}
}
/**
* A worker's mid-turn question (CB-205 reverse rendezvous): surface {@code question} to the
* primary by resolving its open blocking {@code bridge_send}, then block this (worker) call until
* the primary answers via {@link #answer} or {@code timeoutMillis} elapses. Identity is the
* worker's own session — it does not address the primary.
*
* <p>Returns {@link AskOutcome#NO_WAITER} when no delegation is open to surface the question to
* (nothing to answer it), {@link AskOutcome#ANSWERED} with the primary's answer, or
* {@link AskOutcome#TIMED_OUT} if the primary stayed silent. The worker resumes its turn either
* way — an answered ask hands back the answer; an unanswered one leaves it to proceed alone.
*/
public AskResult ask(String workerSession, String question, long timeoutMillis) {
Rendezvous.AskTicket ticket = rendezvous.openAsk(workerSession);
// Only the freshly-opening caller surfaces the question; a coalesced duplicate simply blocks on
// the shared answer future that the fresh owner is already responsible for.
if (ticket.fresh()) {
// Register the reverse waiter first, then surface the question — so the answer, which can
// arrive the instant the primary reacts, always finds an open waiter to resolve.
if (!rendezvous.resolveQuestion(workerSession, question, ticket.turnId())) {
rendezvous.closeAsk(ticket.turnId());
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
}
}
try {
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
return new AskResult(AskOutcome.ANSWERED, answer);
} catch (TimeoutException e) {
log.debug("bridge_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
return new AskResult(AskOutcome.TIMED_OUT, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting the primary's answer for " + workerSession, e);
} finally {
// Only the fresh owner tears down the shared turn; a duplicate must leave it open.
if (ticket.fresh()) {
rendezvous.closeAsk(ticket.turnId());
}
}
}
/**
* The primary's answer to a worker's {@code bridge_ask} (CB-205): resolve the worker's blocked
* question identified by {@code turnId}, then — like a fresh {@link #send} — block for the worker's
* eventual {@code bridge_reply} as it finishes the resumed turn. The worker session is derived from
* {@code turnId}, never a caller argument.
*
* <p>Unlike {@link #send} this does not re-inject through the {@link Injector}: the worker is
* mid-turn (already picked up), so the answer flows back through its own open {@code bridge_ask}
* call, not a new status-gated delivery. The forward waiter is opened <em>before</em> the worker
* is unblocked so a reply that lands the instant it resumes is not lost.
*/
public Reply answer(String turnId, String content, long timeoutMillis) {
String workerSession = rendezvous.askSession(turnId);
if (workerSession == null) {
return new Reply(Outcome.STALE_TURN, null); // the ask lapsed (timed out or already answered)
}
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(workerSession, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null);
}
try {
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession);
if (!rendezvous.answerAsk(turnId, content)) {
rendezvous.close(workerSession, reply);
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
}
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
} catch (TimeoutException e) {
// The worker resumed but hasn't replied yet — no completion fallback arms an answered
// turn (it never re-entered the injector), so a silent worker rides out the window.
return new Reply(Outcome.TIMED_OUT_WORKING, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + workerSession, e);
} finally {
rendezvous.close(workerSession, reply);
}
} finally {
lock.unlock();
}
}
/**
* Fire-and-poll variant of {@link #send}: deliver {@code content} to {@code target} on a
* background virtual thread and return immediately with a ticket to {@link #poll}. This is how a
* long task is delegated without tripping the caller's MCP client call timeout.
*
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content) {
String ticket = "task-" + ticketSeq.incrementAndGet();
CompletableFuture<Reply> future =
CompletableFuture.supplyAsync(() -> send(target, content, ASYNC_TIMEOUT_MS), asyncExecutor);
tasks.put(ticket, new Task(target, future, System.nanoTime()));
pruneTerminalTickets();
log.debug("async send {} -> {}", ticket, target);
return ticket;
}
/**
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket;
* otherwise a {@link Phase#PENDING} view (with the live worker status as detail), a
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason.
*/
public TaskView poll(String ticket) {
Task task = tasks.get(ticket);
if (task == null) {
return null;
}
CompletableFuture<Reply> f = task.future();
if (!f.isDone()) {
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target()));
}
Reply r;
try {
r = f.getNow(null);
} catch (CompletionException | java.util.concurrent.CancellationException e) {
Throwable cause = (e instanceof CompletionException ce && ce.getCause() != null) ? ce.getCause() : e;
return new TaskView(ticket, Phase.FAILED, null, null, cause.getMessage());
}
if (r.completed()) {
String source = r.outcome() == Outcome.REPLIED ? "reply" : "transcript";
return new TaskView(ticket, Phase.DONE, r.text(), source, null);
}
// A wedged worker (CB-109) carries the error context as its reason; the timeout/busy
// outcomes carry none, so fall back to the outcome name.
String detail = r.outcome() == Outcome.WORKER_FAILED && r.text() != null
? r.text()
: "no reply — " + r.outcome().name().toLowerCase();
return new TaskView(ticket, Phase.FAILED, null, null, detail);
}
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
private String liveStatus(String target) {
try {
return agents.status(target).name().toLowerCase();
} catch (RuntimeException e) {
return "unknown";
}
}
/** Drop finished tickets older than the TTL so the registry cannot grow without bound. */
private void pruneTerminalTickets() {
long cutoff = System.nanoTime() - TICKET_TTL_NANOS;
tasks.values().removeIf(t -> t.future().isDone() && t.createdNanos() < cutoff);
}
/** Release the async executor. */
public void close() {
asyncExecutor.shutdown();
}
/** Map a rendezvous {@link Rendezvous.Kind} onto its send {@link Outcome} (shared by send/answer). */
private static Outcome outcomeOf(Rendezvous.Kind kind) {
return switch (kind) {
case REPLY -> Outcome.REPLIED;
case COMPLETION -> Outcome.COMPLETED_UNREPLIED;
case FAILED -> Outcome.WORKER_FAILED;
case QUESTION -> Outcome.QUESTION;
};
}
private static boolean tryLock(ReentrantLock lock, long millis) {
try {
return lock.tryLock(Math.max(0, millis), TimeUnit.MILLISECONDS);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting the session send lock", e);
}
}
private static long remainingMillis(long deadlineNanos) {
return (deadlineNanos - System.nanoTime()) / 1_000_000L;
}
}
@@ -0,0 +1,216 @@
package dev.ltms.bridged.msg;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicLong;
/**
* The reply rendezvous: where a blocking {@code bridge_send} awaits how the worker's delegated turn
* ends. The sending (primary) request thread {@link #open}s a waiter; it is resolved either by the
* worker's explicit {@code bridge_reply} ({@link #resolve}, arriving on a different thread via
* {@code POST /sessions/{id}/reply}) or — the CB-106 fallback — by the injector observing the
* worker's delegated turn return to idle without a reply ({@link #resolveCompletion}).
*
* <p>At most one waiter per session — {@link MessageService} serializes sends per session, so a
* resolution maps unambiguously to the one outstanding send and cannot be captured by another.
*
* <p><strong>Waiter identity (CB-116).</strong> The completion/failure fallbacks run asynchronously
* and can fire <em>after</em> the turn they belong to has already been resolved by an explicit reply
* and a <em>next</em> send has opened its own waiter on the same session. Resolving "whatever waiter
* is registered now" would then land turn N's stale scrape on turn N+1's send. So those fallbacks
* resolve a <em>specific</em> {@link CompletableFuture} captured when their turn was delivered
* ({@link #resolveCompletion(CompletableFuture, String)} /
* {@link #resolveFailure(CompletableFuture, String)}): a no-op if that waiter was already resolved,
* and it can never touch a later send's waiter.
*/
public final class Rendezvous {
/** How a delegated turn ended (or paused). */
public enum Kind {
/** The worker called {@code bridge_reply} with a structured answer. */
REPLY,
/** The worker's turn finished without a {@code bridge_reply}; {@code text} is a scrape. */
COMPLETION,
/** The worker ran the turn then wedged (CB-109); {@code text} is the failure context. */
FAILED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205 reverse rendezvous);
* {@code text} is the question and {@code turnId} correlates the primary's answer back to
* the worker's blocked {@code bridge_ask}. Not terminal — the turn resumes after the answer.
*/
QUESTION
}
/**
* The resolved outcome of a send: its {@link Kind}, the associated text, and — only for
* {@link Kind#QUESTION} — the {@code turnId} the primary answers with (else {@code null}).
*/
public record Resolution(Kind kind, String text, String turnId) {
/** A terminal resolution (reply / completion / failure) with no correlation id. */
public Resolution(Kind kind, String text) {
this(kind, text, null);
}
}
/** A worker's open mid-turn question: the worker session it belongs to and the answer future. */
private record AskWaiter(String session, CompletableFuture<String> answer) {
}
/**
* Handle to a reverse-rendezvous turn: the {@code turnId}, its answer future, and whether this
* call freshly opened it (versus coalescing onto an already-open ask).
*/
public record AskTicket(String turnId, CompletableFuture<String> answer, boolean fresh) {
}
private final ConcurrentHashMap<String, CompletableFuture<Resolution>> waiters = new ConcurrentHashMap<>();
/** Reverse rendezvous (CB-205): worker questions awaiting the primary's answer, keyed by {@code turnId}. */
private final ConcurrentHashMap<String, AskWaiter> asks = new ConcurrentHashMap<>();
private final AtomicLong askSeq = new AtomicLong();
/** Per-session index of the currently-open ask, so duplicate bridge_ask calls coalesce onto one turn. */
private final ConcurrentHashMap<String, String> openAsksBySession = new ConcurrentHashMap<>();
/**
* Register a waiter for {@code session} — the await side of the public {@code resolve*} methods.
* The caller must hold that session's send lock.
*/
public CompletableFuture<Resolution> open(String session) {
CompletableFuture<Resolution> waiter = new CompletableFuture<>();
waiters.put(session, waiter);
return waiter;
}
/** Remove {@code waiter} for {@code session} (only if it is still the registered one). */
void close(String session, CompletableFuture<Resolution> waiter) {
waiters.remove(session, waiter);
}
/** Whether a send is currently awaiting a resolution for {@code session}. */
public boolean isWaiting(String session) {
return waiters.containsKey(session);
}
/**
* The waiter currently registered for {@code session}, or {@code null} if none is waiting. The
* completion/failure fallbacks capture this at delivery time so they can later resolve that exact
* send (see the CB-116 note above) rather than whichever send happens to be waiting when they fire.
*/
public CompletableFuture<Resolution> currentWaiter(String session) {
return waiters.get(session);
}
/**
* Resolve the send awaiting on {@code session} with the worker's explicit reply {@code content}.
*
* @return {@code true} if a waiter was resolved; {@code false} if none was waiting (a late or
* spurious reply — e.g. the send already timed out)
*/
public boolean resolve(String session, String content) {
return complete(session, new Resolution(Kind.REPLY, content));
}
// --- reverse rendezvous (CB-205 bridge_ask) ------------------------------------------------
/**
* Open a reverse-rendezvous waiter for a worker's mid-turn question. If {@code session} already has
* an open ask, coalesce onto it (same {@code turnId}, same answer future). Otherwise atomically mint
* a fresh {@code turnId}, register it in both the per-turn and per-session indexes, and hand it back
* marked fresh. The caller then {@link #resolveQuestion surfaces the question} to the primary and
* blocks on the returned future until the primary {@link #answerAsk answers}.
*/
public AskTicket openAsk(String session) {
while (true) {
AskWaiter[] minted = { null };
String turnId = openAsksBySession.computeIfAbsent(session, _ -> {
String newTurnId = session + "#" + askSeq.incrementAndGet();
CompletableFuture<String> answer = new CompletableFuture<>();
AskWaiter waiter = new AskWaiter(session, answer);
asks.put(newTurnId, waiter);
minted[0] = waiter;
return newTurnId;
});
if (minted[0] != null) {
return new AskTicket(turnId, minted[0].answer(), true);
}
AskWaiter existing = asks.get(turnId);
if (existing != null) {
return new AskTicket(turnId, existing.answer(), false);
}
// A close raced and removed the waiter after we read the turnId; clear the stale index entry
// and retry so a fresh ask is always backed by a registered waiter.
openAsksBySession.remove(session, turnId);
}
}
/**
* Surface a worker's mid-turn {@code question} by resolving the primary's open {@code bridge_send}
* with a {@link Kind#QUESTION} carrying {@code turnId}. Same session-keyed semantics as
* {@link #resolve}: the one outstanding send for {@code session} unblocks with the question.
*
* @return {@code true} if a send was awaiting (the question reached the primary); {@code false}
* if none was (no delegation is open to answer it)
*/
public boolean resolveQuestion(String session, String question, String turnId) {
return complete(session, new Resolution(Kind.QUESTION, question, turnId));
}
/** The worker session an outstanding ask {@code turnId} belongs to, or {@code null} if unknown/lapsed. */
public String askSession(String turnId) {
AskWaiter w = asks.get(turnId);
return w == null ? null : w.session();
}
/**
* Resolve a worker's blocked {@code bridge_ask} with the primary's {@code answer}, unblocking it
* to resume its turn.
*
* @return {@code true} if the ask was still open and got the answer; {@code false} if the
* {@code turnId} is unknown or the ask already lapsed (timed out / was answered)
*/
public boolean answerAsk(String turnId, String answer) {
AskWaiter w = asks.get(turnId);
return w != null && w.answer().complete(answer);
}
/** Drop a reverse-rendezvous turn once its {@code bridge_ask} has resolved (answered or lapsed). */
public void closeAsk(String turnId) {
AskWaiter w = asks.get(turnId);
if (w == null) {
return;
}
// Remove the session index first and only if it still points to this turn, so a concurrent
// fresh ask cannot inherit a waiter we are about to drop.
openAsksBySession.remove(w.session(), turnId);
asks.remove(turnId);
}
/**
* Resolve a specific captured {@code waiter} as a completion (the delegated turn finished with no
* {@code bridge_reply}); {@code text} is the scraped transcript tail. The waiter is the one
* captured when this turn was delivered, so a late completion for turn N cannot land on turn N+1's
* send (CB-116). A no-op if that waiter was already resolved — a raced {@code bridge_reply} wins.
*
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
*/
public boolean resolveCompletion(CompletableFuture<Resolution> waiter, String text) {
return waiter != null && waiter.complete(new Resolution(Kind.COMPLETION, text));
}
/**
* Resolve a specific captured {@code waiter} as a failure — the worker ran the turn but wedged in
* an unrecoverable state (CB-109); {@code reason} is the failure context (e.g. the error screen).
* Like {@link #resolveCompletion(CompletableFuture, String)} it targets the exact captured send
* (CB-116). A no-op if that waiter was already resolved — first resolution wins.
*
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
*/
public boolean resolveFailure(CompletableFuture<Resolution> waiter, String reason) {
return waiter != null && waiter.complete(new Resolution(Kind.FAILED, reason));
}
private boolean complete(String session, Resolution resolution) {
CompletableFuture<Resolution> waiter = waiters.get(session);
return waiter != null && waiter.complete(resolution);
}
}
@@ -0,0 +1,51 @@
package dev.ltms.bridged.msg;
import java.util.List;
/**
* Holds terminal worker→primary replies that arrive with no live send to resolve, keyed by worker
* session (target), until the primary drains them. Soft-state in Stage 1 (in-memory, lost on restart);
* the Stage 2 AMQP adapter implements the same contract with cross-restart durability.
*
* <p><strong>This interface is the port.</strong> {@link InMemoryReplyInbox} is the Stage-1 adapter;
* an AMQP-backed adapter (Stage 2) must implement the same contract (idempotent publish, FIFO peek,
* at-least-once ack).
*
* <p><strong>Ownership is explicit.</strong> A gateway {@link #own owns} the inbox for each agent it
* spawned; only the owner consumes and drains it. {@link #publish} sends a reply to the target's
* inbox but does <em>not</em> imply ownership or start a consumer. This separation is required by
* CB-308 federation, where one gateway may publish to an agent owned by another gateway.
*/
public interface ReplyInbox {
/** A queued reply: an idempotency id, the worker session it came from, and the reply text. */
record InboxMessage(String msgId, String target, String content) {}
/**
* Start owning (consuming) the inbox for {@code target}. Idempotent: multiple calls for the same
* target are no-ops. The owner is the only gateway that may {@link #peek} and {@link #ack} replies
* for this target.
*/
void own(String target);
/**
* Stop owning (consuming) the inbox for {@code target}. Idempotent. Any replies held locally but
* not yet acked are dropped from the local snapshot; the underlying durable queue keeps
* unacked messages for redelivery when the target is re-owned.
*/
void release(String target);
/**
* Queue {@code content} from worker {@code target} under {@code msgId}. Idempotent: publishing an
* already-present {@code msgId} for {@code target} is a no-op (dedup), so an at-least-once Stage-2
* redelivery cannot double-queue. Publishing does <em>not</em> imply ownership and must not start a
* consumer.
*/
void publish(String target, String msgId, String content);
/** Non-destructive snapshot of pending replies for {@code target} (FIFO), empty list if none. */
List<InboxMessage> peek(String target);
/** Remove the reply {@code msgId} for {@code target} once the primary has taken it. No-op if absent. */
void ack(String target, String msgId);
}
@@ -0,0 +1,176 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
/**
* Mechanism (b) of CB-307: a dedicated, status-gated push loop that nudges the primary's own
* herdr pane when a worker reply lands with no live {@code bridge_send} to resolve it.
*
* <p>The loop is triggered by {@link #onReplyQueued(String)} (called from
* {@link MessageService#reply} after the durable inbox publish). It checks four conditions
* at each tick via {@link #decide(String, int)}, then either injects a drain nudge,
* waits for the primary to become injectable, or stops reminding.
*
* <p>Bounded: at most {@link #maxReminders} nudges per target, with a configurable backoff
* between them. The reply is never lost — the durable inbox is the backstop.
*/
public final class ReplyPushLoop {
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
private final ReplyInbox inbox;
private final ScheduledExecutorService scheduler;
private final int maxReminders;
private final long backoffMs;
private final Metrics metrics; // CB-512: nullable — no registry in unit tests
/** Track targets that have an active schedule. */
private final ConcurrentHashMap<String, Boolean> activeTargets = new ConcurrentHashMap<>();
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
int maxReminders, long backoffMs) {
this(primaryRegistry, agents, inbox, scheduler, maxReminders, backoffMs, null);
}
/** As above, with a metric registry (CB-512) so push outcomes are counted. */
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
int maxReminders, long backoffMs, Metrics metrics) {
this.primaryRegistry = primaryRegistry;
this.agents = agents;
this.inbox = inbox;
this.scheduler = scheduler;
this.maxReminders = maxReminders;
this.backoffMs = backoffMs;
this.metrics = metrics;
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
if (metrics != null) {
metrics.inc(name, labels);
}
}
// --- decision logic (package-private for unit-testing) -------------------------------------
/** The action the loop should take for a target at the given reminder count. */
enum Action { INJECT, WAIT_BUSY, STOP }
/**
* Pure decision function: examine the current state and return what the loop should do.
*
* @param target the worker session (target terminal id)
* @param reminderCount how many nudges have been sent so far for this target
* @return the action the caller should take
*/
Action decide(String target, int reminderCount) {
if (!primaryRegistry.isKnown()) {
log.debug("push: primary unknown, stopping reminder for {}", target);
return Action.STOP;
}
if (inbox.peek(target).isEmpty()) {
log.debug("push: inbox empty for {}, stopping reminder", target);
return Action.STOP;
}
if (reminderCount >= maxReminders) {
log.debug("push: reminder cap ({}) reached for {}, stopping", maxReminders, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted");
return Action.STOP;
}
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
AgentStatus status;
try {
status = agents.status(primaryTerminal);
} catch (RuntimeException e) {
log.debug("push: status check failed for primary {}, will retry", primaryTerminal, e);
return Action.WAIT_BUSY;
}
if (status.injectable()) {
return Action.INJECT;
}
log.debug("push: primary {} is {} (not injectable), waiting", primaryTerminal, status);
return Action.WAIT_BUSY;
}
// --- public entrypoint ---------------------------------------------------------------------
/**
* Called when a reply is queued for {@code target}. Idempotent per target: a second call while
* a schedule is active is a no-op. The schedule nudges the primary, then schedules a follow-up
* check (reminder on backoff, or re-check on WAIT_BUSY), until the inbox is empty or the cap
* is reached.
*/
public void onReplyQueued(String target) {
if (activeTargets.putIfAbsent(target, Boolean.TRUE) != null) {
log.debug("push: already active for {}, ignoring duplicate trigger", target);
return; // already scheduled
}
log.debug("push: starting reminder loop for {}", target);
scheduleNext(target, 0);
}
/** Execute one loop tick — called on the scheduler thread. */
private void tick(String target, int reminderCount) {
var action = decide(target, reminderCount);
switch (action) {
case INJECT -> {
injectNudge(target, reminderCount);
scheduleNext(target, reminderCount + 1);
}
// Re-check after the configured backoff; the primary may become injectable soon.
case WAIT_BUSY -> scheduleNext(target, reminderCount);
case STOP -> {
activeTargets.remove(target);
log.debug("push: reminder loop ended for {}", target);
}
}
}
/** Send the nudge and log the event. */
private void injectNudge(String target, int reminderCount) {
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
String nudge = NUDGE_FORMAT.formatted(target, target);
try {
agents.send(primaryTerminal, nudge);
log.debug("push: nudge {}/{} sent to primary {} for target {}",
reminderCount + 1, maxReminders, primaryTerminal, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered");
} catch (RuntimeException e) {
log.warn("push: failed to nudge primary {} for target {} (reminder {}/{}): {}",
primaryTerminal, target, reminderCount + 1, maxReminders, e.toString());
}
}
/** Schedule the next tick on the scheduler thread pool. */
private void scheduleNext(String target, int nextReminderCount) {
scheduler.schedule(() -> tick(target, nextReminderCount), backoffMs, TimeUnit.MILLISECONDS);
}
// --- lifecycle -----------------------------------------------------------------------------
/** Shut down the scheduler. Outstanding reminders are cancelled. */
public void stop() {
scheduler.shutdownNow();
activeTargets.clear();
}
/** @see #stop() */
public void close() {
stop();
}
}
@@ -0,0 +1,35 @@
package dev.ltms.bridged.peer;
/**
* Declared capabilities of a {@link PeerLauncher}. The protocol is the union across all
* configured launchers; a verb invoked against a peer that lacks the capability returns a clean
* "unsupported for this peer" rather than a crash. Capabilities keep the protocol honest as peers
* diversify and prevent the core from assuming "every peer is a Claude in a worktree."
*/
public enum Capability {
/**
* The peer supports {@code bridge_ask} rendezvous — pausing its delegated turn to ask
* the primary a question, then resuming once answered. All Claude Code peers support this.
*/
MID_TURN_ASK,
/**
* The peer can open its own PR at the end of an implementation turn (CB-302). Opt-in per
* profile: granted only when the profile carries a git-forge token ({@code gitTokenEnv}).
*/
SELF_PR,
/**
* The peer can run inside a provisioned isolated git worktree. All CLI-based peers support
* this since their cwd is set at spawn time.
*/
WORKTREE,
/**
* The spawner can reconcile orphaned peers on boot — workers that outlived a prior daemon
* process and whose pane ids died with it (CB-117). Claude Code over herdr supports this
* via name-based matching against the herdr agent list.
*/
ORPHAN_REAP
}
@@ -0,0 +1,44 @@
package dev.ltms.bridged.peer;
/**
* An opaque handle returned by {@link PeerLauncher#spawn(SpawnRequest)}. The core routes on
* {@link #id()} (the registry/routing key) and uses {@link #terminalId()} for session tracking;
* launcher-private coordinates beyond these are reachable through the concrete implementation.
*
* <p>A {@link PeerHandle} is returned <em>after</em> the peer process is live — the launcher
* has already completed subscription-guarded env/vfs setup, process start, and placement. The
* handle is a ticket the core exchanges for the running peer, not a lazy/delayed reference.
*/
public interface PeerHandle {
/**
* The registry/routing key — an opaque, launcher-assigned identifier (CB-519). Multiple
* daemon processes may run on one host, so the contract is <em>host-unique</em>, not merely
* process-unique: the herdr-backed launcher mints a fresh UUID per spawn, and a non-herdr
* launcher is likewise expected to return an identifier that cannot collide across processes
* on the same host. This id is the routing key and is deliberately decoupled from any launcher
* transport coordinate (e.g. a herdr pane id), which stays launcher-private. Guaranteed to be
* non-null and unique among live peers on the host.
*/
String id();
/**
* The transport-level session identifier used for message routing and presence tracking.
* For the herdr launcher this is the herdr terminal UUID. A non-herdr launcher may return
* its own analogous identifier, or {@code null} if the concept does not apply.
*/
default String terminalId() {
return null;
}
/**
* The worker profile that spawned this peer, if the launcher resolved one. A launcher that
* performs dynamic profile selection (e.g. CB-518 weighted placement) sets this so the
* session registry records the actual profile rather than the requested/default one.
*
* @return the profile name, or {@code null} when the launcher leaves it unspecified
*/
default String profile() {
return null;
}
}
@@ -0,0 +1,85 @@
package dev.ltms.bridged.peer;
import java.util.List;
import java.util.Set;
/**
* SPI for materializing a connected peer — the only way the bridge core creates or tears down
* a peer process. Every launcher is a first-party, in-tree adapter selected by (future) profile
* config; today's single adapter is the {@code ClaudeCodeLauncher} / Claude Code over herdr.
*
* <p>The core delegates spawn and teardown to this interface without knowing how the peer is set
* up. Environment variables, CLI flags, subscription guards, transport (herdr tab/pane) layout,
* and naming conventions are all adapter-private — the core sees only the returned
* {@link PeerHandle} whose {@code id()} is the registry/routing key.
*
* <p>The interface is a superset of what {@code SessionManager} and {@code Bridged.main} call
* on the concrete launcher today.
*/
public interface PeerLauncher {
/**
* The set of {@link Capability capabilities} this launcher declares. A peer whose profile
* opts into a git-forge token should include {@link Capability#SELF_PR}; the base set for
* the Claude Code herdr adapter is always {@code MID_TURN_ASK, WORKTREE, ORPHAN_REAP}.
*/
Set<Capability> capabilities();
/**
* {@code profileName}/requestedCwd null/blank → default resolution. Returns after the peer
* process is live (env + argv + placement complete). Never returns {@code null}.
*
* @param req the spawn parameters (profile, requested cwd, caller cwd)
* @return a handle whose {@link PeerHandle#id()} is the registry/routing key
* @throws IllegalArgumentException if the profile is unknown and no default is configured
*/
PeerHandle spawn(SpawnRequest req);
/**
* The configured worker profile names — the set of names {@code spawn(profileName)} accepts.
*/
Set<String> profiles();
/**
* The profile a no-argument {@link #spawn(SpawnRequest)} uses, or {@code null} if none is configured.
*/
String defaultProfile();
/**
* Resolve the effective working directory for a spawn {@code req} without actually spawning.
* Resolution order: requestedCwd → profile cwd → callerCwd → daemon cwd.
*
* @return the resolved absolute path, never null/blank
*/
String effectiveCwd(SpawnRequest req);
/**
* The parity-overlay file list for {@code profileName} (default list when unset). Used by
* worktree provisioning to copy config files into the isolated checkout before spawning.
*/
List<String> parityOverlay(String profileName);
/**
* The set of all agents this launcher currently tracks, transport-specific. Each element
* exposes at minimum a pane-like {@code id()} matching this launcher's {@link PeerHandle}
* scheme, plus transport-level status. Callers merge this set with the session registry to
* build a live roster view.
*/
List<?> list();
/**
* Reap orphaned peers left behind by a prior daemon process. Only peers whose naming scheme
* matches this launcher's and whose nonce differs from the current process are eligible.
* Best-effort: a failure to list or to stop any one peer is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
int reapOrphanWorkers();
/**
* Tear a peer down by its registry/routing key ({@link PeerHandle#id()}). Tolerates an
* already-gone peer. Also cleans up launcher-private resources (e.g. empty dedicated tabs)
* when safe to do so.
*/
void stop(String id);
}
@@ -0,0 +1,18 @@
package dev.ltms.bridged.peer;
/**
* Thrown when a {@link PeerLauncher} starts a peer process but the peer
* does not reach an injectable (ready-to-receive) state within the configured
* timeout. The launcher MUST clean up any resources it created (pane, tab)
* before throwing — no orphaned peer or pane is left behind.
*
* <p>This is a spawn-time failure, distinct from a post-spawn disconnect.
* Callers treat this as a clean spawn error (the peer never materialized
* into a usable session), not a mid-life session fault.
*/
public final class PeerUnreachableException extends RuntimeException {
public PeerUnreachableException(String message) {
super(message);
}
}
@@ -0,0 +1,13 @@
package dev.ltms.bridged.peer;
/**
* Parameters for a {@link PeerLauncher#spawn(SpawnRequest)} call — the peer-neutral
* aggregation of what the core knows at delegation time: which profile to use, the caller's
* requested working directory, and the caller's own cwd (to inherit when no other cwd is set).
*
* <p>A null or blank {@code profileName} means "use the launcher's default profile."
* A null or blank {@code requestedCwd} means "inherit from config or caller."
* A null {@code callerCwd} means "the request came from the daemon itself (not a primary)."
*/
public record SpawnRequest(String profileName, String requestedCwd, String callerCwd) {
}
@@ -0,0 +1,22 @@
package dev.ltms.bridged.placement;
/**
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps and
* reachability so that a pre-existing config behaves identically after upgrade.
*/
final class FixedPlacementPolicy implements PlacementPolicy {
@Override
public PlacementCandidate select(PlacementContext ctx) {
String d = ctx.defaultProfile();
if (d != null && !d.isBlank()) {
return new PlacementCandidate(d, null, 1.0f, null);
}
if (!ctx.candidates().isEmpty()) {
PlacementCandidate first = ctx.candidates().getFirst();
return new PlacementCandidate(first.profile(), null, first.weight(), first.maxLoad());
}
throw new PlacementException("no worker profiles configured");
}
}
@@ -0,0 +1,20 @@
package dev.ltms.bridged.placement;
/**
* A profile (and, in CB-308, a host) that can be chosen by a {@link PlacementPolicy}.
*
* <p>Keeping this as a small descriptor rather than a bare profile name lets CB-308 widen
* selection to {@code (host, profile)} pairs without changing the policy interface.
*/
public record PlacementCandidate(String profile, String host, float weight, Integer maxLoad) {
/** A candidate with no explicit host (the single-host default) and the given weight/cap. */
public static PlacementCandidate profile(String profile, float weight, Integer maxLoad) {
return new PlacementCandidate(profile, null, weight, maxLoad);
}
/** A candidate with no explicit host, unit weight, and no cap. */
public static PlacementCandidate profile(String profile) {
return new PlacementCandidate(profile, null, 1.0f, null);
}
}
@@ -0,0 +1,19 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.Set;
import java.util.function.Function;
/**
* Everything a {@link PlacementPolicy} needs to make one selection.
*
* @param defaultProfile profile a {@code fixed} policy should return (may be {@code null})
* @param candidates every configured candidate; the policy filters out those at cap or unreachable
* @param liveCount current live worker count per profile (from the session registry)
* @param unreachable profiles already known to have failed in this spawn attempt
*/
public record PlacementContext(String defaultProfile,
List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable) {
}
@@ -0,0 +1,12 @@
package dev.ltms.bridged.placement;
/**
* Thrown when a {@link PlacementPolicy} has no candidate available. Kept as a distinct type so
* callers can distinguish "no capacity" from a spawn-time transport failure.
*/
public final class PlacementException extends IllegalStateException {
public PlacementException(String message) {
super(message);
}
}
@@ -0,0 +1,41 @@
package dev.ltms.bridged.placement;
/**
* Factory for the built-in placement policies.
*/
public final class PlacementPolicies {
private PlacementPolicies() {
}
/**
* Resolve a policy name from config. Absent/blank values and {@code "fixed"} return the
* backward-compatible fixed policy; unknown names throw.
*/
public static PlacementPolicy fromName(String name) {
String n = (name == null) ? "" : name.toLowerCase();
if (n.isBlank() || "fixed".equals(n)) {
return fixed();
}
if ("weighted".equals(n)) {
return weighted();
}
if ("round-robin".equals(n)) {
return roundRobin();
}
throw new IllegalArgumentException("unknown placement policy '" + name
+ "' — must be one of: fixed, round-robin, weighted");
}
public static PlacementPolicy fixed() {
return new FixedPlacementPolicy();
}
public static PlacementPolicy weighted() {
return new WeightedRoundRobinPolicy();
}
public static PlacementPolicy roundRobin() {
return new RoundRobinPlacementPolicy();
}
}
@@ -0,0 +1,16 @@
package dev.ltms.bridged.placement;
/**
* How {@code bridged} chooses a worker profile when a spawn names none. Implementations are
* deterministic and unit-testable; the caller (the composite launcher) handles failover retries.
*/
public interface PlacementPolicy {
/**
* Pick one candidate from the configured set.
*
* @throws java.lang.IllegalStateException when no candidate is available, with a message naming
* whether every profile is at capacity or unreachable
*/
PlacementCandidate select(PlacementContext ctx);
}
@@ -0,0 +1,65 @@
package dev.ltms.bridged.placement;
import java.util.ArrayList;
import java.util.List;
/**
* Shared filtering and empty-set reporting used by the built-in placement policies.
*/
final class PlacementPolicyUtil {
private PlacementPolicyUtil() {
}
/**
* Candidates that are not known-unreachable and have not reached their maxLoad.
* A {@code null} maxLoad means unlimited.
*/
static List<PlacementCandidate> available(PlacementContext ctx) {
List<PlacementCandidate> out = new ArrayList<>();
for (PlacementCandidate c : ctx.candidates()) {
if (ctx.unreachable().contains(c.profile())) {
continue;
}
Integer cap = c.maxLoad();
if (cap != null) {
int live = ctx.liveCount().apply(c.profile());
if (live >= cap) {
continue;
}
}
out.add(c);
}
return out;
}
/**
* Build a clear exception describing why every candidate was dropped: all at capacity,
* all unreachable, or a mix.
*/
static PlacementException emptyException(PlacementContext ctx) {
int atCap = 0;
int unreachable = 0;
for (PlacementCandidate c : ctx.candidates()) {
Integer cap = c.maxLoad();
if (ctx.unreachable().contains(c.profile())) {
unreachable++;
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
atCap++;
}
}
int total = ctx.candidates().size();
if (total == 0) {
return new PlacementException("no worker profiles configured");
}
if (atCap == total) {
return new PlacementException("all worker profiles are at maxLoad");
}
if (unreachable == total) {
return new PlacementException("all worker profiles are unreachable");
}
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
+ unreachable + " unreachable, " + (total - atCap - unreachable) + " remaining");
}
}
@@ -0,0 +1,23 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.concurrent.atomic.AtomicInteger;
/**
* Deterministic round-robin over the profiles that still have capacity and are not known to be
* unreachable. The index advances only on successful selections so the distribution stays even
* across spawns.
*/
final class RoundRobinPlacementPolicy implements PlacementPolicy {
private final AtomicInteger index = new AtomicInteger(0);
@Override
public synchronized PlacementCandidate select(PlacementContext ctx) {
List<PlacementCandidate> available = PlacementPolicyUtil.available(ctx);
if (available.isEmpty()) {
throw PlacementPolicyUtil.emptyException(ctx);
}
return available.get(index.getAndIncrement() % available.size());
}
}
@@ -0,0 +1,50 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
/**
* Smooth weighted round-robin (nginx-style): for each selection, add the candidate's weight to
* its current score, pick the highest score, then subtract the total weight of all available
* candidates from the winner. Weights are not required to sum to 1.0; only their ratios matter.
*
* <p>The state is per-policy instance and protected by {@code synchronized} so concurrent spawns
* see a consistent, deterministic sequence rather than interleaving updates.
*/
final class WeightedRoundRobinPolicy implements PlacementPolicy {
private final Map<String, Double> current = new ConcurrentHashMap<>();
@Override
public synchronized PlacementCandidate select(PlacementContext ctx) {
List<PlacementCandidate> available = PlacementPolicyUtil.available(ctx);
if (available.isEmpty()) {
throw PlacementPolicyUtil.emptyException(ctx);
}
double total = 0.0;
for (PlacementCandidate c : available) {
total += c.weight();
}
if (total <= 0.0) {
throw new PlacementException("all available profiles have non-positive weight");
}
PlacementCandidate best = null;
double bestScore = Double.NEGATIVE_INFINITY;
for (PlacementCandidate c : available) {
double score = current.merge(c.profile(), (double) c.weight(), (old, add) -> old + add);
if (score > bestScore) {
bestScore = score;
best = c;
}
}
if (best == null) {
throw new PlacementException("no placement candidate could be selected");
}
current.put(best.profile(), current.get(best.profile()) - total);
return best;
}
}
@@ -1,18 +1,34 @@
package dev.ltms.bridged.rest;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.bridged.auth.AuditLog;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.worker.WorkerService;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.peer.PeerLauncher;
import io.javalin.Javalin;
import io.javalin.http.Context;
import jakarta.servlet.http.HttpServlet;
import org.eclipse.jetty.servlet.ServletHolder;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.function.Function;
import java.util.stream.Collectors;
/**
* The REST surface — {@code bridged}'s contract and its testability seam. Every
@@ -25,25 +41,141 @@ import java.util.Map;
*/
public final class BridgedApp {
private final HerdrClient herdr;
private final WorkerService workers;
/** Default blocking window for a message; kept under typical HTTP idle timeouts. */
private static final long DEFAULT_MESSAGE_TIMEOUT_MS = 25_000;
private static final long MAX_MESSAGE_TIMEOUT_MS = 120_000;
/** Blocking window for a worker's bridge_ask (CB-205); the worker's MCP client caps its own call. */
private static final long DEFAULT_ASK_TIMEOUT_MS = 55_000;
private static final long MAX_ASK_TIMEOUT_MS = 115_000;
public BridgedApp(HerdrClient herdr, WorkerService workers) {
/** Context attribute under which the resolved caller is stashed by the auth filter. */
private static final String CALLER = "bridged.caller";
private final HerdrClient herdr;
private final PeerLauncher workers;
private final SessionManager sessions; // CB-301: authoritative session registry
private final MessageService messages;
private final WorkerPresence presence; // CB-113: which workers are MCP-connected (available)
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
private final CallerResolver auth; // CB-501: null → authz not enforced (legacy behaviour)
private final Metrics metrics; // CB-502: null → /metrics not exposed
private final ObjectMapper mapper = new ObjectMapper();
/**
* Legacy constructor — no identity resolution and no authorization, exactly as the REST surface
* behaved before CB-501. Retained so existing acceptance tests keep exercising handler
* behaviour without each needing an auth fixture.
*/
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
HttpServlet mcpServlet) {
this(herdr, workers, sessions, messages, presence, mcpServlet, null, null);
}
/**
* @param auth resolves each request's {@link Principal}; {@code null} disables authorization
* entirely (legacy). {@code main} always supplies one.
* @param metrics registry to instrument and expose at {@code GET /metrics}; {@code null} omits
* the endpoint
*/
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics) {
this.herdr = herdr;
this.workers = workers;
this.sessions = sessions;
this.messages = messages;
this.presence = presence;
this.mcpServlet = mcpServlet;
this.auth = auth;
this.metrics = metrics;
}
/** Wire routes onto a fresh, unstarted Javalin instance. Caller starts it. */
public Javalin build() {
Javalin app = Javalin.create(cfg -> cfg.showJavalinBanner = false);
Javalin app = Javalin.create(cfg -> {
cfg.showJavalinBanner = false;
if (mcpServlet != null) {
// The MCP server shares the daemon's port; Jetty routes /mcp to its servlet.
cfg.jetty.modifyServletContextHandler(h ->
h.addServlet(new ServletHolder(mcpServlet), "/mcp"));
}
});
// CB-501: resolve identity once per request, before any handler. /mcp does NOT pass through
// here — it is a raw servlet on Jetty's context handler — so BridgeMcp enforces separately
// against the same CallerResolver. Any check that lives in only one place is not a control.
if (auth != null) {
app.before(ctx -> ctx.attribute(CALLER,
auth.resolve(ctx.req().getRemoteAddr(), ctx.req().getRemotePort(),
ctx.header("Authorization"))));
}
app.get("/healthz", this::healthz);
if (metrics != null) {
app.get("/metrics", this::metrics);
}
app.get("/sessions", this::sessions);
app.get("/agents", this::agents);
app.post("/workers", this::spawnWorker);
app.get("/workers", this::listWorkers); // CB-304: registry roster + live herdr status
app.get("/profiles", this::profiles); // configured worker profiles
app.post("/workers", this::spawnWorker); // optional ?profile= or {"profile":…}
app.delete("/workers/{paneId}", this::stopWorker);
app.post("/sessions/{id}/message", this::sendMessage); // bridge_send (primary; blocking, wait:false, or answer via turnId)
app.post("/sessions/{id}/reply", this::replyMessage); // bridge_reply (worker)
app.get("/sessions/{id}/replies", this::drainReplies); // drain reply inbox (CB-307)
app.post("/sessions/{id}/ask", this::askMessage); // bridge_ask (worker → primary, CB-205)
app.get("/sessions/{id}/status", this::sessionStatus); // bridge_status
app.get("/tasks/{ticket}", this::taskStatus); // poll an async (wait:false) send
return app;
}
/**
* Gate a handler on the CB-505 authorization table. Returns {@code true} when the request may
* proceed; otherwise writes the error response and returns {@code false}.
*
* <p>401 vs 403 is a real distinction here: 401 means "you presented no usable identity" (a
* credential problem the caller can fix), 403 means "you are authenticated, but this is not
* yours" (a worker reaching for another worker's session, or for orchestration).
*/
private boolean allow(Context ctx, Authz.Action action, String target) {
if (auth == null) {
return true; // legacy: authorization not enforced
}
Principal caller = ctx.attribute(CALLER);
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ && action != Authz.Action.METRICS) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
return true;
}
if (Authz.isUnauthenticated(caller)) {
AuditLog.denied(caller, action, target, "unauthenticated");
countAuthFailure("unauthenticated");
ctx.status(401).json(Map.of("error", "unauthenticated",
"detail", "present Authorization: Bearer <token>"));
} else {
AuditLog.denied(caller, action, target, "forbidden");
countAuthFailure("forbidden");
ctx.status(403).json(Map.of("error", "forbidden",
"detail", caller.describe() + " may not " + action + " on "
+ (target == null ? "this resource" : target)));
}
return false;
}
private void countAuthFailure(String reason) {
if (metrics != null) {
metrics.inc("bridged_auth_failures_total", "reason", reason);
}
}
/** Prometheus scrape endpoint (CB-502). */
private void metrics(Context ctx) {
if (!allow(ctx, Authz.Action.METRICS, null)) {
return;
}
ctx.status(200).contentType("text/plain; version=0.0.4; charset=utf-8").result(metrics.render());
}
/** Liveness + herdr reachability. 200 when herdr answers ping, 503 otherwise. */
private void healthz(Context ctx) {
try {
@@ -63,6 +195,9 @@ public final class BridgedApp {
/** Sessions view derived from herdr {@code workspace.list} (one workspace → one row). */
private void sessions(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
JsonNode result = herdr.call("workspace.list");
List<Map<String, Object>> out = new ArrayList<>();
for (JsonNode w : result.path("workspaces")) {
@@ -78,25 +213,323 @@ public final class BridgedApp {
/** Discovery: every agent herdr tracks, keyed by its Claude session UUID. */
private void agents(Context ctx) {
ctx.status(200).json(Map.of("agents", workers.list().stream().map(BridgedApp::view).toList()));
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
ctx.status(200).json(Map.of("agents",
workers.list().stream().map(Agent.class::cast).map(BridgedApp::view).toList()));
}
/** Spawn a guard-checked worker. 403 if the base_url would breach the subscription boundary. */
/** CB-304: bridge-owned roster merged with live herdr status by paneId. */
private void listWorkers(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
// CB-519: the registry key is a host-unique id, not the pane coordinate — join on terminal.
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
ctx.status(200).json(Map.of("workers", out));
}
/** The configured worker profiles and which one a no-argument spawn uses. */
private void profiles(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
ctx.status(200).json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile()));
}
/**
* Spawn a guard-checked worker. An optional {@code profile} (query param or {@code {"profile":…}}
* body) picks which configured profile; omitted → the default. 403 if the base_url would breach
* the subscription boundary, 400 for an unknown profile.
*/
private void spawnWorker(Context ctx) {
if (!allow(ctx, Authz.Action.SPAWN, null)) {
return;
}
String profile = ctx.queryParam("profile");
String cwd = ctx.queryParam("cwd");
String worktree = ctx.queryParam("worktree");
String ticket = ctx.queryParam("ticket");
if (profile == null || profile.isBlank() || cwd == null || cwd.isBlank()
|| worktree == null || worktree.isBlank()) {
try {
String body = ctx.body();
if (!body.isBlank()) {
JsonNode b = mapper.readTree(body);
if (profile == null || profile.isBlank()) profile = b.path("profile").asText(null);
if (cwd == null || cwd.isBlank()) cwd = b.path("cwd").asText(null);
if (worktree == null || worktree.isBlank()) worktree = b.path("worktree").asText(null);
if (ticket == null || ticket.isBlank()) ticket = b.path("ticket").asText(null);
}
} catch (Exception ignored) {
// A malformed/empty body just means "no overrides" → fall through to defaults.
}
}
WorktreeRequest wt = worktreeRequest(worktree, ticket);
try {
Agent worker = workers.spawn();
// No MCP caller over REST, so callerCwd and ownerTerminal are null.
WorkerSession worker = sessions.acquire(blankToNull(profile), blankToNull(cwd), null, null, wt);
ctx.status(201).json(view(worker));
} catch (GuardException e) {
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
} catch (IllegalArgumentException e) {
ctx.status(400).json(Map.of("error", "unknown_profile", "detail", e.getMessage()));
} catch (PeerUnreachableException e) {
ctx.status(502).json(Map.of("error", "spawn_timeout", "detail", e.getMessage()));
}
}
private static WorktreeRequest worktreeRequest(String worktree, String ticket) {
if (worktree == null || worktree.isBlank() || "false".equalsIgnoreCase(worktree)) {
return null;
}
if ("true".equalsIgnoreCase(worktree)) {
if (ticket == null || ticket.isBlank()) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return new WorktreeRequest(worktree, null);
}
private static String blankToNull(String s) {
return (s == null || s.isBlank()) ? null : s;
}
/** Tear a worker down by pane id. */
private void stopWorker(Context ctx) {
workers.stop(ctx.pathParam("paneId"));
String paneId = ctx.pathParam("paneId");
if (!allow(ctx, Authz.Action.STOP, paneId)) {
return;
}
sessions.release(paneId);
ctx.status(204);
}
/**
* The blocking delegation call (CB-104): inject {@code content} into the worker via the
* status-gated injector and block until the worker returns a structured {@code bridge_reply}.
* Times out with a typed 202 (working / queued / busy) rather than an error — the message may
* still land.
*/
private void sendMessage(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.SEND, id)) {
return;
}
String content;
String turnId;
long timeout;
boolean wait;
try {
JsonNode body = mapper.readTree(ctx.body());
content = body.path("content").asText("");
turnId = body.path("turnId").asText(null);
timeout = body.path("timeoutMs").asLong(DEFAULT_MESSAGE_TIMEOUT_MS);
wait = body.path("wait").asBoolean(true); // default: block for the reply (CB-104)
} catch (Exception e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
if (content.isBlank()) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "content is required"));
return;
}
timeout = Math.clamp(timeout, 1, MAX_MESSAGE_TIMEOUT_MS);
// Answering a worker's bridge_ask (CB-205): always blocks, and derives the worker from turnId.
if (turnId != null && !turnId.isBlank()) {
writeReply(ctx, id, messages.answer(turnId, content, timeout), timeout);
return;
}
if (!wait) {
// Fire-and-poll (CB-107): return a ticket immediately; the caller polls GET /tasks/{ticket}.
String ticket = messages.sendAsync(id, content);
ctx.status(202).json(Map.of("sessionId", id, "ticket", ticket, "status", "accepted"));
return;
}
try {
writeReply(ctx, id, messages.send(id, content, timeout), timeout);
} catch (HerdrException e) {
herdrError(ctx, e);
}
}
/**
* Render a {@link MessageService.Reply} onto the response — shared by a normal send and a
* bridge_ask answer. A structured/scraped completion is 200; a worker's mid-turn question a 202
* (with its {@code turnId}); a stale answer a 409; every other non-terminal outcome a typed 202.
*/
private void writeReply(Context ctx, String id, MessageService.Reply reply, long timeout) {
switch (reply.outcome()) {
case QUESTION -> ctx.status(202).json(Map.of(
"sessionId", id, "status", "question",
"question", reply.text(), "turnId", reply.turnId()));
case STALE_TURN -> ctx.status(409).json(Map.of(
"sessionId", id, "error", "stale_turn",
"detail", "that question is no longer open (timed out or already answered)"));
case REPLIED, COMPLETED_UNREPLIED -> {
// replySource distinguishes a structured bridge_reply from the CB-106 completion
// fallback (a scrape of the worker's transcript when it finished without replying).
String source = reply.outcome() == MessageService.Outcome.REPLIED ? "reply" : "transcript";
ctx.status(200).json(Map.of("sessionId", id, "reply", reply.text(), "replySource", source));
}
default -> ctx.status(202).json(Map.of(
"sessionId", id,
"status", switch (reply.outcome()) {
case TIMED_OUT_WORKING -> "working";
case TIMED_OUT_QUEUED -> "queued";
case BUSY -> "busy";
case WORKER_FAILED -> "failed";
default -> "done"; // unreachable (terminal outcomes handled above)
},
"detail", reply.outcome() == MessageService.Outcome.WORKER_FAILED && reply.text() != null
? reply.text()
: "no reply within " + timeout + "ms; poll status or retry"));
}
}
/**
* A worker's mid-turn question ({@code bridge_ask}, CB-205) — surfaces to the primary's open
* blocking send and blocks until it answers. 200 with the answer, 409 if no delegation is open,
* 202 if the primary stayed silent.
*/
private void askMessage(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.ASK, id)) {
return;
}
String question;
long timeout;
try {
JsonNode body = mapper.readTree(ctx.body());
question = body.path("question").asText("");
timeout = body.path("timeoutMs").asLong(DEFAULT_ASK_TIMEOUT_MS);
} catch (Exception e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
if (question.isBlank()) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "question is required"));
return;
}
timeout = Math.clamp(timeout, 1, MAX_ASK_TIMEOUT_MS);
MessageService.AskResult r = messages.ask(id, question, timeout);
switch (r.outcome()) {
case ANSWERED -> ctx.status(200).json(Map.of("sessionId", id, "answered", true, "answer", r.answer()));
case NO_WAITER -> ctx.status(409).json(Map.of(
"sessionId", id, "error", "no_pending_send",
"detail", "no primary is awaiting this turn to answer a question"));
case TIMED_OUT -> ctx.status(202).json(Map.of(
"sessionId", id, "status", "no_answer",
"detail", "the primary did not answer within " + timeout + "ms"));
}
}
/**
* The worker's structured reply ({@code bridge_reply}) — resolves the blocking send awaiting
* on this session, or queues the reply in the inbox when no send is open (CB-307).
*/
private void replyMessage(Context ctx) {
String id = ctx.pathParam("id");
// The rule that matters: a worker may reply only as itself. Over MCP this was already true
// structurally (identity comes from the connection, never an argument); over REST the path
// id was simply trusted, so this is where the invariant actually gets enforced.
if (!allow(ctx, Authz.Action.REPLY, id)) {
return;
}
String content;
try {
content = mapper.readTree(ctx.body()).path("content").asText("");
} catch (Exception e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
messages.reply(id, content);
ctx.status(200).json(Map.of("sessionId", id, "delivered", true));
}
/**
* Drain the reply inbox for a worker session — peek + ack any replies that arrived when no send
* was open. At-least-once: draining removes them from the inbox so a subsequent read returns
* nothing; an in-flight failure between the drain and the caller's processing re-surfaces them.
*/
private void drainReplies(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.DRAIN, id)) {
return;
}
var replies = messages.drainReplies(id);
ctx.status(200).json(Map.of("sessionId", id, "replies",
replies.stream().map(m -> Map.of(
"msgId", m.msgId(),
"content", m.content())).toList()));
}
/**
* Live lifecycle status of a worker (MCP `bridge_status` wraps this in CB-105), plus its
* <em>readiness</em> (CB-113): {@code ready} is true once the worker's Claude has connected the
* bridge MCP — the reliable "available to receive a task" signal, unlike bare {@code idle}, which
* is also true during boot.
*/
private void sessionStatus(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.READ, id)) {
return;
}
try {
ctx.status(200).json(Map.of(
"sessionId", id,
"status", messages.status(id).name().toLowerCase(),
"ready", presence.isPresent(id)));
} catch (HerdrException e) {
herdrError(ctx, e);
}
}
/** Poll an async (wait:false) delegation by ticket. 404 for an unknown/expired ticket. */
private void taskStatus(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"));
if (v == null) {
ctx.status(404).json(Map.of("error", "unknown_ticket", "detail", "no such task (or it has expired)"));
return;
}
Map<String, Object> body = new LinkedHashMap<>();
body.put("ticket", v.ticket());
body.put("phase", v.phase().name().toLowerCase());
if (v.reply() != null) {
body.put("reply", v.reply());
body.put("replySource", v.replySource());
}
if (v.detail() != null) {
body.put("detail", v.detail());
}
ctx.status(200).json(body);
}
/** Map a herdr failure: unknown target → 404, anything else → 502 (herdr is upstream). */
private static void herdrError(Context ctx, HerdrException e) {
if (e.code() != null && e.code().endsWith("_not_found")) {
ctx.status(404).json(Map.of("error", "session_not_found", "detail", e.getMessage()));
} else {
ctx.status(502).json(Map.of("error", "herdr_error", "detail", e.getMessage()));
}
}
/** Stable JSON projection of an agent (null-safe for the start-time shape). */
private static Map<String, Object> view(Agent a) {
Map<String, Object> m = new LinkedHashMap<>();
@@ -109,4 +542,22 @@ public final class BridgedApp {
m.put("status", a.status().name().toLowerCase());
return m;
}
/** CB-301 projection of an authoritative bridge-owned session. */
private static Map<String, Object> view(WorkerSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("terminalId", s.terminalId());
m.put("paneId", s.paneId());
m.put("profile", s.profile());
m.put("cwd", s.cwd());
m.put("ownerTerminal", s.ownerTerminal());
m.put("state", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
return m;
}
}
@@ -0,0 +1,217 @@
package dev.ltms.bridged.session;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.BufferedReader;
import java.io.IOException;
import java.io.InputStreamReader;
import java.io.UncheckedIOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.security.SecureRandom;
import java.util.List;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.stream.Collectors;
/**
* Production {@link Worktrees} implementation that shells {@code git} via {@link ProcessBuilder}.
* Non-zero exits become {@link WorktreeException}. Worktree directories live under a configurable
* root (default: a sibling {@code .bridged-worktrees} of the repo root) so they are never nested
* inside the primary working tree.
*/
public final class GitWorktrees implements Worktrees {
private static final Logger log = LoggerFactory.getLogger(GitWorktrees.class);
/** Project-level MCP config. Present in the repo, so every worktree checks the primary's out. */
private static final String MCP_CONFIG = ".mcp.json";
/** What {@link #isolateToolSurface} writes: a valid, explicitly empty server map. */
private static final String NEUTRAL_MCP_CONFIG = "{\n \"mcpServers\": {}\n}\n";
private final String configuredRoot;
private final SecureRandom random = new SecureRandom();
private final AtomicLong seq = new AtomicLong();
/** Default constructor: worktree root is derived per-repo as {@code <repoRoot>/../.bridged-worktrees}. */
public GitWorktrees() {
this(null);
}
/** @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of the repo root. */
public GitWorktrees(String configuredRoot) {
this.configuredRoot = configuredRoot;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
String base = (baseRef == null || baseRef.isBlank()) ? "HEAD" : baseRef;
String nonce = nonce();
Path root = resolveRoot(repoRoot);
Path path = root.resolve(nonce);
try {
Files.createDirectories(root);
} catch (IOException e) {
throw new WorktreeException("cannot create worktree root " + root + ": " + e.getMessage(), e);
}
String wt = path.toAbsolutePath().toString();
log.info("adding worktree branch={} path={} base={}", branch, wt, base);
exec("git", "-C", repoRoot, "worktree", "add", wt, "-b", branch, base);
isolateToolSurface(wt);
return wt;
}
/**
* Neutralize the worktree's project MCP config so a worker inherits only the tools its launcher
* mounts (the bridge, via {@code --mcp-config}) — never the primary's.
*
* <p>This is unconditional, and it is not the same job as the parity overlay. The repo's own
* committed {@code .mcp.json} declares the primary's IDE servers, so a fresh checkout mounts them
* whether or not the overlay copies anything; a worker that inherits them navigates and edits
* through tools bound to the <em>primary's</em> IntelliJ project, which silently hands it absolute
* paths outside its own worktree. That is not hypothetical: a CB-523 worker made all 59 of its
* edits in the primary checkout while compiling its worktree, so every build it ran was of code
* that did not contain its changes.
*
* <p>Writing an empty server map (rather than deleting the file) keeps a project-level
* {@code .mcp.json} present and explicit, and the {@code --skip-worktree} bit keeps the
* neutralized copy from ever showing up as a local modification the worker might commit.
*/
private void isolateToolSurface(String worktreePath) {
Path root = Path.of(worktreePath).toAbsolutePath().normalize();
Path mcp = root.resolve(MCP_CONFIG);
try {
Files.writeString(mcp, NEUTRAL_MCP_CONFIG);
} catch (IOException e) {
throw new WorktreeException("cannot neutralize " + MCP_CONFIG + " in the worktree: "
+ e.getMessage(), e);
}
if (isTracked(root, MCP_CONFIG)) {
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", MCP_CONFIG);
}
log.debug("neutralized {} — worker tool surface is launcher-mounted only", MCP_CONFIG);
}
@Override
public void remove(String repoRoot, String worktreePath) {
Path p = Path.of(worktreePath);
if (!Files.exists(p)) {
log.debug("worktree {} already gone — nothing to remove", worktreePath);
return;
}
log.info("removing worktree {}", worktreePath);
exec("git", "-C", repoRoot, "worktree", "remove", "--force", worktreePath);
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
if (overlay == null || overlay.isEmpty()) {
return;
}
Path srcRoot = Path.of(repoRoot).toAbsolutePath().normalize();
Path dstRoot = Path.of(worktreePath).toAbsolutePath().normalize();
for (String rel : overlay) {
Path src = srcRoot.resolve(rel).normalize();
if (!Files.exists(src)) {
log.debug("parity overlay source missing — skipping {}", rel);
continue;
}
Path dst = dstRoot.resolve(rel).normalize();
try {
Files.createDirectories(dst.getParent());
Files.copy(src, dst, StandardCopyOption.REPLACE_EXISTING, StandardCopyOption.COPY_ATTRIBUTES);
log.debug("copied parity overlay {}", rel);
} catch (IOException e) {
throw new WorktreeException("cannot copy overlay " + rel + ": " + e.getMessage(), e);
}
if (isTracked(dstRoot, rel)) {
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", rel);
log.debug("marked overlay --skip-worktree {}", rel);
}
}
}
@Override
public String repoRoot(String cwd) {
String out = exec("git", "-C", cwd, "rev-parse", "--show-toplevel");
return Path.of(out.trim()).toAbsolutePath().normalize().toString();
}
/** Resolve the directory that will hold per-session worktree checkouts. */
private Path resolveRoot(String repoRoot) {
if (configuredRoot != null && !configuredRoot.isBlank()) {
return Path.of(configuredRoot).toAbsolutePath().normalize();
}
Path repo = Path.of(repoRoot).toAbsolutePath().normalize();
return repo.resolveSibling(".bridged-worktrees");
}
private String nonce() {
return String.format("%06x", random.nextInt(1 << 24)) + "-" + seq.incrementAndGet();
}
private boolean isTracked(Path worktreeRoot, String rel) {
return exitCode("git", "-C", worktreeRoot.toString(), "ls-files", "--error-unmatch", rel) == 0;
}
/**
* Run a command and return its stdout. Non-zero exit → {@link WorktreeException} with both
* stdout and stderr (merged by redirectErrorStream).
*/
private String exec(String... command) {
String out;
int code;
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
try (BufferedReader r = new BufferedReader(new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
out = r.lines().collect(Collectors.joining("\n"));
} catch (IOException e) {
p.destroyForcibly();
throw new UncheckedIOException(e);
}
try {
if (!p.waitFor(30, TimeUnit.SECONDS)) {
p.destroyForcibly();
throw new WorktreeException("command timed out: " + String.join(" ", command) + "\n" + out);
}
code = p.exitValue();
} catch (InterruptedException e) {
p.destroyForcibly();
Thread.currentThread().interrupt();
throw new WorktreeException("interrupted waiting for command: " + String.join(" ", command), e);
}
if (code != 0) {
throw new WorktreeException("exit " + code + " for: " + String.join(" ", command)
+ (out.isBlank() ? "" : "\n" + out));
}
return out;
}
private int exitCode(String... command) {
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
try {
if (!p.waitFor(30, TimeUnit.SECONDS)) {
p.destroyForcibly();
throw new WorktreeException("command timed out: " + String.join(" ", command));
}
return p.exitValue();
} catch (InterruptedException e) {
p.destroyForcibly();
Thread.currentThread().interrupt();
throw new WorktreeException("interrupted waiting for command: " + String.join(" ", command), e);
}
}
}
@@ -0,0 +1,504 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Consumer;
import java.util.function.LongSupplier;
/**
* Authoritative in-daemon registry of the worker sessions this {@code bridged} process spawned.
* Delegates spawn/teardown to a {@link PeerLauncher} (which performs subscription-guarded env
* setup and process/materialization) and adds lifecycle tracking, ownership, and deterministic
* teardown on top.
*
* <p>The state machine is intentionally one-shot / no-reuse: every acquired worker is fresh,
* and a finished or released worker is torn down, never pooled. {@link #recycle} is a convenience
* for {@code release + acquire} with a new distinct pane id.
*
* <p>The manager implements {@link TurnListener} so the injector's turn boundaries drive
* {@code READY → BUSY → DONE} (or {@code FAILED}). It exposes a {@link WorkerPresence} view via
* {@link #asPresence()}: any MCP contact from a worker marks it present and simultaneously
* transitions the session {@code SPAWNING → READY}.
*/
public final class SessionManager implements TurnListener {
private static final Logger log = LoggerFactory.getLogger(SessionManager.class);
private final PeerLauncher launcher;
private final Worktrees worktrees;
private final ConcurrentHashMap<String /*paneId*/, WorkerSession> registry = new ConcurrentHashMap<>();
private final WorkerPresence presence;
private final SecureRandom nonceRandom = new SecureRandom();
private final AtomicLong nonceSeq = new AtomicLong();
private final LongSupplier nowNanos;
private final int contextCap;
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** CB-516: notified with a terminalId on every release; no-op until wired. */
private final List<Consumer<String>> releaseListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
public SessionManager(PeerLauncher launcher) {
this(launcher, new GitWorktrees(), System::nanoTime, 0);
}
/** Backward-compatible constructor with an injectable worktree seam. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees) {
this(launcher, worktrees, System::nanoTime, 0);
}
/** Test constructor with an injectable clock. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos) {
this(launcher, worktrees, nowNanos, 0);
}
/** Production constructor with a configured context turn cap. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, int contextCap) {
this(launcher, worktrees, System::nanoTime, contextCap);
}
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos,
int contextCap) {
this.launcher = launcher;
this.worktrees = worktrees;
this.presence = new PresenceBridge(this);
this.nowNanos = nowNanos;
this.contextCap = contextCap;
}
/**
* The single {@link WorkerPresence} view of this manager: it records availability and forwards
* the signal to the {@code SPAWNING → READY} transition. Pass this to the {@code Injector} and
* {@code BridgeMcp} where they previously accepted a plain {@link WorkerPresence}. The same
* instance is returned every call — presence is shared state, so a fresh bridge per call would
* fragment the {@code present} set and lose signals across callers.
*/
public WorkerPresence asPresence() {
return presence;
}
/**
* Spawn a worker and register it as {@link WorkerSession.State#SPAWNING}. The caller's
* identity is recorded as {@code ownerTerminal} ({@code null} for daemon/anon callers).
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal) {
return acquire(profile, requestedCwd, callerCwd, ownerTerminal, null);
}
/**
* Spawn a worker, optionally inside a fresh git worktree. When {@code wt} is non-null the
* worktree is provisioned, parity-overlaid, and its path becomes the worker's cwd. On any
* failure before registration the worktree is removed so no dangling checkout is left.
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
if (wt == null) {
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd);
PeerHandle handle = launcher.spawn(req);
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd));
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
cwd,
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
null,
null);
registry.put(handle.id(), session);
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
notifyAcquired(session.terminalId());
return session;
}
return acquireWithWorktree(profile, requestedCwd, callerCwd, ownerTerminal, wt);
}
/** Tear a worker down by pane id and remove it from the registry. Idempotent. */
public void release(String paneId) {
WorkerSession removed = registry.remove(paneId);
if (removed != null) {
log.debug("releasing session pane={} terminal={} state={}",
removed.paneId(), removed.terminalId(), removed.state());
// CB-516: a send still waiting on this worker can never be answered now. Tell the
// listener BEFORE the pane is torn down, so a blocked caller fails fast with a real
// reason instead of sitting on a rendezvous nothing will ever resolve.
notifyReleased(removed.terminalId());
}
launcher.stop(paneId);
if (removed != null && removed.worktree() != null) {
worktrees.remove(worktrees.repoRoot(removed.cwd()), removed.worktree());
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is acquired
* (CB-520). This is the hook that lets the reply inbox {@code own} a target's queue.
*/
public void onAcquire(Consumer<String> listener) {
if (listener != null) {
acquireListeners.add(listener);
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is released
* (CB-516). Every teardown path funnels through {@link #release}, so one hook covers the REST
* and MCP stop tools, the idle-TTL reaper, {@code recycle}, and shutdown drain alike.
*
* <p>Added rather than injected because {@code MessageService} — one intended listener — is
* constructed after this manager (it needs the injector and rendezvous, which need the session
* presence view this manager exposes). Wiring it at construction would require breaking that
* cycle for one callback.
*/
public void onRelease(Consumer<String> listener) {
if (listener != null) {
releaseListeners.add(listener);
}
}
/** A listener failure must never prevent the acquisition it is reacting to. */
private void notifyAcquired(String terminalId) {
if (terminalId == null) {
return;
}
for (Consumer<String> listener : acquireListeners) {
try {
listener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("acquire listener failed for terminal {}: {}", terminalId, e.toString());
}
}
}
/** A listener failure must never prevent the teardown it is reacting to. */
private void notifyReleased(String terminalId) {
if (terminalId == null) {
return;
}
for (Consumer<String> listener : releaseListeners) {
try {
listener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
}
}
}
private WorkerSession acquireWithWorktree(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
String preResolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
// nor a caller cwd, so taking the first non-blank of those two yielded null and put
// `git -C null` on the command line — an NPE out of ProcessBuilder, surfacing as HTTP 500.
// The non-worktree path always used this chain; only this branch was missed.
String repoRoot = worktrees.repoRoot(
launcher.effectiveCwd(new SpawnRequest(preResolvedProfile, requestedCwd, callerCwd)));
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
String path = null;
PeerHandle handle;
try {
path = worktrees.add(repoRoot, branch, wt.baseRef());
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd));
} catch (RuntimeException e) {
if (path != null) {
try {
worktrees.remove(repoRoot, path);
} catch (RuntimeException cleanup) {
log.warn("failed to clean up worktree {} after spawn error: {}", path, cleanup.getMessage());
}
}
throw e;
}
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
cwd,
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
path,
branch);
registry.put(handle.id(), session);
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
notifyAcquired(session.terminalId());
return session;
}
private String slug(String raw) {
return raw == null ? "ticket" : raw.toLowerCase().replaceAll("[^a-z0-9]+", "-").replaceAll("^-+|-+$", "");
}
private String nonce() {
return String.format("%06x", nonceRandom.nextInt(1 << 24)) + "-" + nonceSeq.incrementAndGet();
}
/**
* The profile to record for a session. A launcher that performed dynamic selection tells us
* the actual profile via {@link PeerHandle#profile()}; otherwise fall back to what the caller
* requested (or the launcher's default for a no-profile spawn).
*/
private String resolveProfile(PeerHandle handle, String requestedProfile) {
String fromHandle = handle.profile();
if (fromHandle != null && !fromHandle.isBlank()) {
return fromHandle;
}
if (requestedProfile != null && !requestedProfile.isBlank()) {
return requestedProfile;
}
return launcher.defaultProfile();
}
/**
* Release the old session and acquire a fresh one with the same profile and working directory.
* The new session is guaranteed to have a pane id distinct from the old one (no-reuse invariant).
*/
public WorkerSession recycle(String paneId) {
WorkerSession old = registry.get(paneId);
if (old == null) {
throw new IllegalArgumentException("no session for paneId " + paneId);
}
release(paneId);
return acquire(old.profile(), old.cwd(), old.cwd(), old.ownerTerminal());
}
/** The session for {@code paneId}, if it is still registered and not released. */
public Optional<WorkerSession> get(String paneId) {
return Optional.ofNullable(registry.get(paneId));
}
/** Bridge-owned roster: all registered sessions (acquired minus released). */
public List<WorkerSession> roster() {
return List.copyOf(registry.values());
}
/**
* CB-304 merged roster+live view. The registry is authoritative for worktree, branch,
* profile, owner, and state; the optional live agent supplies the herdr-reported status.
*/
public static Map<String, Object> rosterView(WorkerSession session, Agent live) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", session.terminalId());
m.put("paneId", session.paneId());
m.put("profile", session.profile());
m.put("state", session.state().name().toLowerCase());
if (session.worktree() != null) {
m.put("worktree", session.worktree());
}
if (session.branch() != null) {
m.put("branch", session.branch());
}
if (session.ownerTerminal() != null) {
m.put("owner", session.ownerTerminal());
}
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
return m;
}
/** Lifecycle hook: worker became available on the bridge MCP. */
void onReady(String terminalId) {
transitionByTerminal(terminalId, WorkerSession.State.SPAWNING, WorkerSession.State.READY);
}
/**
* Lifecycle hook: a message was delivered into the worker — it is now busy on a turn.
* The turn count is bumped and the activity timestamp is refreshed. A {@code DONE} session
* can be re-delivered for multi-turn reuse until it is released.
*/
@Override
public void onDelivered(String target) {
WorkerSession current = findByTerminal(target);
if (current == null) return;
if (current.state() != WorkerSession.State.READY && current.state() != WorkerSession.State.DONE) {
return;
}
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.BUSY).bumpTurn(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} {} -> BUSY turn={}",
target, current.paneId(), current.state(), updated.turnCount());
}
}
/** Lifecycle hook: the worker's delegated turn completed successfully. */
@Override
public void onTurnComplete(String target) {
WorkerSession current = findByTerminal(target);
if (current == null || current.state() != WorkerSession.State.BUSY) return;
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.DONE).withActivity(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} BUSY -> DONE turn={}",
target, current.paneId(), updated.turnCount());
}
if (contextCap > 0 && updated.turnCount() >= contextCap) {
release(current.paneId());
}
}
/** Lifecycle hook: the worker's delegated turn failed. */
@Override
public void onTurnFailed(String target) {
onFailed(target);
}
/** Lifecycle hook: the worker vanished or was dropped mid-life. */
void onFailed(String target) {
WorkerSession current = findByTerminal(target);
if (current == null) return;
if (current.state() == WorkerSession.State.RELEASED) return;
if (replace(current, current.withState(WorkerSession.State.FAILED))) {
log.debug("session marked failed terminal={} pane={}", target, current.paneId());
}
}
/**
* Best-effort reap of sessions that have been idle longer than {@code idleTtlNanos}. Only
* {@code READY} and {@code DONE} sessions are eligible — never a {@code SPAWNING} or
* {@code BUSY} worker. Returns the number of sessions released.
*/
int reapIdle(long idleTtlNanos) {
long now = nowNanos.getAsLong();
int reaped = 0;
for (WorkerSession s : roster()) {
if (s.state() != WorkerSession.State.READY && s.state() != WorkerSession.State.DONE) {
continue;
}
if (now - s.lastActivityAtNanos() > idleTtlNanos) {
release(s.paneId());
reaped++;
}
}
return reaped;
}
/**
* Gracefully drain all registered sessions. For each session that is {@code BUSY}, poll up to
* {@code timeoutNanos} for it to leave {@code BUSY}, then release it regardless. Non-busy
* sessions are released immediately. A failure releasing one session is logged and does not
* abort the rest.
*/
void drainAll(long timeoutNanos) {
long deadline = System.nanoTime() + timeoutNanos;
for (WorkerSession s : roster()) {
try {
if (s.state() == WorkerSession.State.BUSY) {
while (System.nanoTime() < deadline) {
WorkerSession current = registry.get(s.paneId());
if (current == null || current.state() != WorkerSession.State.BUSY) {
break;
}
try {
long remaining = deadline - System.nanoTime();
Thread.sleep(Math.min(TimeUnit.NANOSECONDS.toMillis(remaining), 50));
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
break;
}
}
}
release(s.paneId());
} catch (RuntimeException e) {
log.warn("drain failed for pane={}; continuing with remaining sessions", s.paneId(), e);
}
}
}
/**
* Close this manager by draining all sessions. The timeout comes from configuration when set,
* otherwise a sensible default.
*/
public void close(Integer drainTimeoutSeconds) {
int seconds = (drainTimeoutSeconds != null && drainTimeoutSeconds > 0) ? drainTimeoutSeconds : 5;
drainAll(TimeUnit.SECONDS.toNanos(seconds));
}
/** Number of sessions currently registered. */
public int size() {
return registry.size();
}
/**
* The registered session owning {@code terminalId}, or {@code null} if none does.
*
* <p>A null {@code terminalId} is a normal input, not a caller bug: every lifecycle hook here is
* fed from the MCP transport, where the <em>primary</em> resolves to a {@link
* dev.ltms.bridged.auth.Principal} with no terminal. {@code BridgeMcp} documents that contact as
* a no-op, and {@link dev.ltms.bridged.inject.WorkerPresence#markPresent} honours it — but
* {@code PresenceBridge} then forwards the same null here. Matching on a null id can never
* succeed anyway (a registered session always has a terminal), so answer "no match" rather than
* throwing: an NPE on this path takes down an unrelated tool call for the primary.
*/
private WorkerSession findByTerminal(String terminalId) {
if (terminalId == null) return null;
for (WorkerSession s : registry.values()) {
if (terminalId.equals(s.terminalId())) return s;
}
return null;
}
private void transitionByTerminal(String terminalId, WorkerSession.State from,
WorkerSession.State to) {
WorkerSession current = findByTerminal(terminalId);
if (current == null || current.state() != from) return;
long now = nowNanos.getAsLong();
if (replace(current, current.withState(to).withActivity(now))) {
log.debug("session transitioned terminal={} pane={} {} -> {}",
terminalId, current.paneId(), from, to);
}
}
private boolean replace(WorkerSession expected, WorkerSession updated) {
return registry.replace(expected.paneId(), expected, updated);
}
/** WorkerPresence bridge that also drives the manager's READY transition. */
private static final class PresenceBridge extends WorkerPresence {
private final SessionManager sessions;
PresenceBridge(SessionManager sessions) {
this.sessions = sessions;
}
@Override
public void markPresent(String terminal) {
if (terminal == null || terminal.isBlank()) {
return; // the primary's contact carries no worker terminal — not a readiness signal
}
super.markPresent(terminal);
sessions.onReady(terminal);
}
}
}
@@ -0,0 +1,70 @@
package dev.ltms.bridged.session;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.TimeUnit;
/**
* Periodic virtual-thread reaper that tears down {@code READY}/{@code DONE} sessions which have
* exceeded their idle TTL. Modeled on {@link dev.ltms.bridged.inject.StatusPoller}: a single
* virtual-thread loop, idempotent start/stop, and no {@code ScheduledExecutorService}.
*/
public final class SessionReaper {
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
private final SessionManager sessions;
private final long idleTtlNanos;
private final long intervalMillis;
private volatile boolean running;
private Thread thread;
/** Construct a reaper with the default 5-second polling interval. */
public SessionReaper(SessionManager sessions, long idleTtlSeconds) {
this(sessions, idleTtlSeconds, DEFAULT_INTERVAL_MILLIS);
}
/** Construct a reaper with an explicit polling interval (useful for tests). */
public SessionReaper(SessionManager sessions, long idleTtlSeconds, long intervalMillis) {
this.sessions = sessions;
this.idleTtlNanos = TimeUnit.SECONDS.toNanos(idleTtlSeconds);
this.intervalMillis = intervalMillis;
}
/** Start the reaper loop on a virtual thread. Idempotent. */
public synchronized void start() {
if (running) return;
running = true;
thread = Thread.ofVirtual().name("session-reaper").start(this::loop);
log.info("session reaper started (idle ttl {}s, interval {}ms)",
TimeUnit.NANOSECONDS.toSeconds(idleTtlNanos), intervalMillis);
}
private void loop() {
while (running) {
try {
sessions.reapIdle(idleTtlNanos);
} catch (RuntimeException e) {
log.warn("session reaper iteration failed; continuing", e);
}
sleep();
}
}
private void sleep() {
try {
Thread.sleep(intervalMillis);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
running = false;
}
}
/** Stop the reaper loop. Idempotent. */
public synchronized void stop() {
running = false;
if (thread != null) thread.interrupt();
}
}
@@ -0,0 +1,61 @@
package dev.ltms.bridged.session;
/**
* A bridge-owned worker session — the authoritative in-daemon record of a worker this
* process spawned. Immutable; state transitions are performed by replacing the record in
* {@link SessionManager}'s registry.
*
* @param paneId the host-unique opaque id (CB-519) — the registry key and the argument to
* teardown. Despite the historical name this is the {@link
* dev.ltms.bridged.peer.PeerHandle#id()}, a UUID, and is distinct from the
* launcher-private herdr pane coordinate.
* @param terminalId herdr terminal handle — the {@code target} for send/read/status
* @param profile the worker profile name that spawned this session
* @param cwd the resolved working directory the worker started in
* @param ownerTerminal the caller that requested this worker ({@code null} = daemon/anon)
* @param spawnedAtNanos {@link System#nanoTime()} when the session was registered
* @param lastActivityAtNanos {@link System#nanoTime()} of the most recent lifecycle event
* @param turnCount number of delegated turns that have been delivered to this session
* @param state current lifecycle state in the one-shot FSM
*/
public record WorkerSession(
String paneId,
String terminalId,
String profile,
String cwd,
String ownerTerminal,
long spawnedAtNanos,
long lastActivityAtNanos,
int turnCount,
State state,
String worktree,
String branch) {
/** One-shot worker lifecycle states. */
public enum State {
SPAWNING,
READY,
BUSY,
DONE,
FAILED,
RELEASED
}
/** Return a copy of this session in {@code state}. */
public WorkerSession withState(State state) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch);
}
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
public WorkerSession withActivity(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount, state, worktree, branch);
}
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
public WorkerSession bumpTurn(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount + 1, state, worktree, branch);
}
}
@@ -0,0 +1,12 @@
package dev.ltms.bridged.session;
/** Non-zero exit or I/O failure from a git worktree operation. */
public final class WorktreeException extends RuntimeException {
public WorktreeException(String message) {
super(message);
}
public WorktreeException(String message, Throwable cause) {
super(message, cause);
}
}
@@ -0,0 +1,6 @@
package dev.ltms.bridged.session;
/** Ask {@link SessionManager#acquire} to provision an isolated worktree. null ⇒ run in the shared primary tree. */
public record WorktreeRequest(String ticketSlug, String baseRef) {
// ticketSlug seeds the branch name; baseRef null/blank ⇒ current HEAD of the repo.
}
@@ -0,0 +1,18 @@
package dev.ltms.bridged.session;
import java.util.List;
/** Seam between {@link SessionManager} and git worktree operations. Tests use a recording fake. */
public interface Worktrees {
/** git -C <repoRoot> worktree add <path> -b <branch> <baseRef|HEAD>. Returns the worktree path. */
String add(String repoRoot, String branch, String baseRef);
/** git -C <repoRoot> worktree remove --force <path>. Idempotent (already-gone tolerated). */
void remove(String repoRoot, String worktreePath);
/** Copy each existing overlay path repoRoot→worktree; mark tracked ones --skip-worktree. */
void overlayParity(String repoRoot, String worktreePath, List<String> overlay);
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
String repoRoot(String cwd);
}
@@ -0,0 +1,195 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>Claude Code</strong> — the safe path from a
* delegation request to a running off-subscription Claude.
*
* <p>Everything transport-related (tab/pane placement, the CB-306 spawn-readiness gate, unique
* naming, CB-117 orphan reap, teardown, listing, cwd resolution) lives in the base. This class
* supplies only the two Claude-specific seams:
* <ul>
* <li>the {@code claude} name prefix (so reap matches {@code claude-*} panes, never another
* adapter's), and</li>
* <li>{@link #buildLaunch}, which encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em> touching
* herdr, and mount the bridge MCP + reply charter as inline launch flags. A worker's base_url
* lives in the env map handed to herdr and nowhere else; {@code bridged}'s own environment is
* never mutated, and nothing is written to the worker's profile.</li>
* </ul>
*/
public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "claude";
private final SubscriptionGuard guard;
/**
* Standing instruction appended to the worker's system prompt so it returns its result via
* {@code bridge_reply}. Injected as a launch flag, so nothing is written to the worker's
* profile — it is guidance, and a worker that never replies is caught by the send's timeout.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet. Every message you "
+ "receive arrives through the bridge, and the ONLY channel back to the sender is the "
+ "bridge_reply MCP tool. Text you write in your terminal is NOT sent anywhere — the "
+ "sender cannot see your screen, so an in-terminal answer is silently discarded. "
+ "Therefore you MUST end EVERY turn by calling bridge_reply with `content` set to your "
+ "complete response. This holds for every message without exception — tasks, questions, "
+ "clarifications, acknowledgements, and ordinary back-and-forth conversation. Call "
+ "bridge_reply exactly once, as the final action of your turn, with your full answer in "
+ "`content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so
* existing deployments and tests keep the legacy non-blocking spawn semantics.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, guard, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300));
}
/**
* Production constructor with spawn-ready gate enabled. The gate polls {@code agents.status()}
* until the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs));
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
* {@code sleeper} is never called when the gate is disabled ({@code spawnReadyTimeoutMs == 0}).
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param guard subscription-boundary guard (checked before spawning)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (e.g. {@code () -> Thread.sleep(pollMs)}); it
* already encodes the poll interval, so the 8th positional argument
* (poll ms) is accepted for API symmetry but otherwise unused here
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.guard = guard;
}
/**
* {@inheritDoc}
*
* <p>The spawn sequence encodes the subscription boundary: assert the profile's base_url is on
* the allowlist <em>before</em> any herdr call, then build the worker env with
* {@code ANTHROPIC_*}, the parity-neutral git-forge grant, and the bridge MCP + reply charter
* mounted as inline launch flags.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
String baseUrl = cfg.baseUrl();
guard.assertWorker(baseUrl); // hard stop before we spawn anything
Map<String, String> workerEnv = baseEnv(cfg);
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithBridge(cfg));
}
/**
* The launch argv, plus — when {@code worker.mcpUrl} is set — inline {@code --mcp-config} for
* the bridge server and {@code --append-system-prompt} for the {@link #REPLY_CHARTER}. Neither
* touches the profile's config; both are pure command-line flags. This inline-flag mount is
* Claude Code specific — other adapters mount MCP and instructions their own way.
*/
private List<String> argvWithBridge(BridgedConfig.Worker cfg) {
if (!cfg.hasMcp()) {
return cfg.argv();
}
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
List<String> argv = mutableArgv(cfg.argv());
argv.add("--mcp-config");
argv.add(mcpJson);
argv.add("--append-system-prompt");
argv.add(REPLY_CHARTER);
return argv;
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
// --- CB-117 reap predicate (Claude prefix), kept for direct unit testing -------------------
/**
* Whether {@code name} is a Claude Code bridge worker started by a <em>different</em> process
* than {@code currentNonce}. A thin {@code claude}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -0,0 +1,48 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import java.nio.file.Path;
/**
* Provisions an isolated {@code CODEX_HOME} for one Codex peer (CB-528).
*
* <p>Codex reads <em>everything</em> from {@code CODEX_HOME} — its config, credentials, sessions,
* skills, plugins, and state. Pointing a peer at the operator's own {@code ~/.codex} would give it
* the operator's tool surface and let it write into the operator's session history, which is the
* same class of failure CB-525 exists to prevent on the Claude side. So every peer gets its own
* directory, and this is the seam that builds it.
*
* <p>It is an interface rather than a method on the launcher for two reasons: provisioning is
* filesystem work with its own failure modes (a missing credential is the most common, and it
* surfaces as an opaque {@code 401} from Codex rather than a spawn error), and keeping it separate
* lets the launcher be tested without touching a real home directory.
*
* <p>Three things the implementation must put in the home, because Codex has no launch flag for
* any of them:
* <ul>
* <li>the bridge MCP server, as {@code [mcp_servers.*]} in {@code config.toml};</li>
* <li>the reply charter, as {@code AGENTS.md} — Codex has no {@code --append-system-prompt},
* so the standing instruction has to reach it as a file;</li>
* <li>credentials, since a freshly created home has none and Codex fails closed.</li>
* </ul>
*/
public interface CodexHome {
/**
* Build a fresh, isolated home for a peer launching under {@code cfg} and return its path,
* suitable for the {@code CODEX_HOME} environment variable.
*
* @param cfg the profile being launched; supplies the MCP URL and any bearer-token variable
* @return the provisioned directory
* @throws RuntimeException if the home cannot be provisioned — including when no credential is
* available, which must fail loudly here rather than as a 401 later
*/
Path provision(BridgedConfig.Worker cfg);
/**
* Remove a home previously returned by {@link #provision}. Idempotent: releasing an already
* released or never provisioned path is not an error, because teardown races teardown.
*/
void release(Path home);
}
@@ -0,0 +1,262 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.PlacementCandidate;
import dev.ltms.bridged.placement.PlacementContext;
import dev.ltms.bridged.placement.PlacementPolicies;
import dev.ltms.bridged.placement.PlacementPolicy;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.Collections;
import java.util.EnumSet;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Function;
/**
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
* router in front of one {@link HerdrPeerLauncher} per peer {@code kind} (Claude Code, opencode, …).
* It owns no transport of its own; it dispatches each SPI call to the delegate that owns the profile
* involved, and fans the fleet-wide queries (list/reap/caps/profiles) across all delegates.
*
* <p>Routing rules:
* <ul>
* <li><strong>By profile</strong> — {@link #spawn}, {@link #effectiveCwd}, {@link #parityOverlay}
* resolve the profile (a null/blank name → the global {@link #defaultProfile}) and delegate to
* the single adapter that declares it. Profiles partition cleanly across adapters: the
* constructor rejects a name claimed by two.</li>
* <li><strong>By pane id</strong> — {@link #stop} routes to the adapter that spawned that pane
* (recorded at spawn time). A pane the composite never spawned (only real for a caller that
* hand-rolls an id) falls back to the first delegate; teardown is pane-id addressed and
* tab cleanup is single-occupant guarded, so it is safe either way.</li>
* <li><strong>Fleet-wide</strong> — {@link #reapOrphanWorkers} and {@link #capabilities} fan out
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
* shares one herdr connection and so reports the same global agent set.</li>
* </ul>
*
* <p>CB-518: an unqualified spawn is routed through a {@link PlacementPolicy}. The default
* {@code fixed} policy reproduces the historical default-profile behaviour; {@code weighted} uses
* smooth weighted round-robin with {@code maxLoad} gating. If a chosen profile fails with
* {@link PeerUnreachableException}, the composite advances to the next available candidate and
* retries, bounded by the number of candidates.
*/
public final class CompositePeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(CompositePeerLauncher.class);
private final List<HerdrPeerLauncher> delegates;
private final Map<String, HerdrPeerLauncher> byProfile;
private final String defaultProfile;
/** paneId → the delegate that spawned it, so {@link #stop} tears down through the right adapter. */
private final Map<String, HerdrPeerLauncher> spawnedBy = new ConcurrentHashMap<>();
private final Map<String, BridgedConfig.Worker> profileConfigs;
private final PlacementPolicy placementPolicy;
private final Function<String, Integer> liveCount;
/**
* Backward-compatible constructor: fixed placement, no live-counting. Use this for tests and
* simple wiring; it preserves the pre-CB-518 behaviour exactly.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to (may be null)
* @throws IllegalArgumentException if {@code delegates} is empty or two adapters claim one profile
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates, String defaultProfile) {
this(delegates, defaultProfile, Map.of(), PlacementPolicies.fixed(), name -> 0);
}
/**
* Production constructor with a placement policy and live-worker counter.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to under {@code fixed} policy
* @param profileConfigs all configured worker profiles (used for candidate weights/caps)
* @param placementPolicy which policy governs unqualified spawns
* @param liveCount live worker count per profile (must never return {@code null})
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Worker> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount) {
if (delegates.isEmpty()) {
throw new IllegalArgumentException("at least one peer adapter must be configured");
}
this.delegates = List.copyOf(delegates);
this.defaultProfile = defaultProfile;
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
// differ from one JVM run to the next.
this.profileConfigs = Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs));
this.placementPolicy = placementPolicy;
this.liveCount = liveCount;
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
HerdrPeerLauncher prev = index.putIfAbsent(profile, d);
if (prev != null) {
throw new IllegalArgumentException(
"worker profile '" + profile + "' is claimed by two peer adapters");
}
}
}
// Order-preserving for the same reason, and because profiles() is user-visible (bridge_profiles).
this.byProfile = Collections.unmodifiableMap(index);
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (resolved == null) {
// No profile and no default configured — hand to the first delegate so it raises the
// same "no default" error it would on its own; keeps the SPI contract single-sourced.
return delegates.getFirst();
}
HerdrPeerLauncher d = byProfile.get(resolved);
if (d == null) {
throw new IllegalArgumentException("unknown worker profile: " + resolved);
}
return d;
}
@Override
public PeerHandle spawn(SpawnRequest req) {
String requestedProfile = req.profileName();
if (requestedProfile != null && !requestedProfile.isBlank()) {
// An explicit profile bypasses the policy entirely.
HerdrPeerLauncher d = route(requestedProfile);
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
}
List<PlacementCandidate> candidates = candidates();
Set<String> unreachable = new HashSet<>();
PlacementContext ctx = new PlacementContext(defaultProfile, candidates, liveCount, unreachable);
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
for (int attempt = 0; attempt < maxAttempts; attempt++) {
PlacementCandidate chosen;
try {
chosen = placementPolicy.select(ctx);
} catch (RuntimeException e) {
// No candidate left (all at cap or all unreachable). The policy already threw a clear
// message; do not wrap it in a generic PeerUnreachableException.
throw e;
}
HerdrPeerLauncher d = byProfile.get(chosen.profile());
if (d == null) {
// A configured profile with no adapter is a wiring bug; fail fast.
unreachable.add(chosen.profile());
continue;
}
SpawnRequest routedReq = new SpawnRequest(chosen.profile(), req.requestedCwd(), req.callerCwd());
try {
PeerHandle handle = d.spawn(routedReq);
spawnedBy.put(handle.id(), d);
return handle;
} catch (PeerUnreachableException e) {
log.warn("spawn on profile {} unreachable, will retry next candidate if any: {}",
chosen.profile(), e.getMessage());
unreachable.add(chosen.profile());
// Update the context for the next selection so the policy excludes this profile.
ctx = new PlacementContext(defaultProfile, candidates, liveCount, unreachable);
}
}
throw new PeerUnreachableException(
"no reachable worker profile available after trying " + unreachable.size()
+ " candidate(s): " + String.join(", ", unreachable));
}
/** Build the candidate list from the configured profiles, in definition order. */
private List<PlacementCandidate> candidates() {
List<PlacementCandidate> out = new ArrayList<>();
for (Map.Entry<String, BridgedConfig.Worker> e : profileConfigs.entrySet()) {
BridgedConfig.Worker w = e.getValue();
out.add(new PlacementCandidate(e.getKey(), null, w.weight(), w.maxLoad()));
}
return out;
}
@Override
public String effectiveCwd(SpawnRequest req) {
return route(req.profileName()).effectiveCwd(req);
}
@Override
public List<String> parityOverlay(String profileName) {
return route(profileName).parityOverlay(profileName);
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.remove(id);
if (d == null) {
log.debug("stop({}) — no recorded owner, routing to the first adapter (pane-addressed)", id);
d = delegates.getFirst();
}
d.stop(id);
}
@Override
public Set<String> profiles() {
return byProfile.keySet();
}
@Override
public String defaultProfile() {
return defaultProfile;
}
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
Map<String, Agent> byPane = new LinkedHashMap<>();
for (HerdrPeerLauncher d : delegates) {
for (Agent a : d.list()) {
if (a.paneId() != null) {
byPane.putIfAbsent(a.paneId(), a);
}
}
}
return List.copyOf(byPane.values());
}
@Override
public int reapOrphanWorkers() {
int reaped = 0;
for (HerdrPeerLauncher d : delegates) {
reaped += d.reapOrphanWorkers();
}
return reaped;
}
/** The union of every adapter's capabilities — a capability any adapter offers, the fleet offers. */
@Override
public Set<Capability> capabilities() {
EnumSet<Capability> caps = EnumSet.noneOf(Capability.class);
for (HerdrPeerLauncher d : delegates) {
caps.addAll(d.capabilities());
}
return Set.copyOf(caps);
}
}
@@ -0,0 +1,205 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.util.Comparator;
import java.util.function.Function;
/**
* The default {@link CodexHome} provisioner (CB-528): builds a fresh, isolated {@code CODEX_HOME}
* for one Codex peer under a configurable root (default the JVM temp dir, mirroring how
* {@link OpenCodeLauncher} picks its config root).
*
* <p>Codex reads <em>everything</em> from {@code CODEX_HOME} — config, credentials, sessions,
* skills, plugins, state — so a peer must never be pointed at the operator's own {@code ~/.codex}.
* Provisioning therefore assembles, in one throwaway directory:
* <ul>
* <li>{@code config.toml} registering the bridge as a streamable-HTTP MCP server (with the
* profile's bearer-token env var, when one is configured);</li>
* <li>{@code AGENTS.md} carrying the reply charter — codex has no {@code --append-system-prompt},
* so the standing instruction has to reach the peer as a file;</li>
* <li>{@code auth.json} copied from the operator's real codex home, since a fresh home has no
* credential and codex then fails closed with an opaque {@code 401} mid-run.</li>
* </ul>
*
* <p>The credential is <em>copied</em>, never symlinked: a peer that could write through a symlink
* would be able to modify the operator's real {@code auth.json}. Each peer gets its own on-disk
* copy, so nothing outside the provisioned home is ever touched on write.
*/
public final class DefaultCodexHome implements CodexHome {
/** Writer for the generated {@code AGENTS.md} (kept as a field for unit-test inspection). */
static final String AGENTS_HEADER = "# Standing instruction\n\n";
/** Root under which per-peer Codex homes are created (injectable for tests). */
private final Path root;
/** Host env lookup used to resolve the operator's codex home (injectable for tests). */
private final Function<String, String> env;
/** Production constructor — homes land under the JVM temp dir and the operator's credentials
* are resolved from the real environment ({@code CODEX_HOME}, else {@code ~/.codex}). */
public DefaultCodexHome() {
this(Path.of(System.getProperty("java.io.tmpdir")), System::getenv);
}
/**
* Full testability constructor: an injectable {@code root} (so tests never touch the JVM temp
* dir's real real estate) and an injectable env lookup (so the credential source can be pointed
* at a temp dir instead of the operator's real {@code ~/.codex}).
*
* @param root directory under which per-peer Codex homes are created (must exist)
* @param env host environment lookup, used to resolve the operator's codex home
*/
public DefaultCodexHome(Path root, Function<String, String> env) {
this.root = root;
this.env = env;
}
/** Standing instruction written to {@code AGENTS.md}. Adapted from
* {@link OpenCodeLauncher#REPLY_CHARTER}: codex has no {@code --append-system-prompt}, so the
* only way a codex peer receives the reply contract is as a file in its home. The substance must
* survive — every message arrives through the bridge, terminal text reaches nobody, so the turn
* MUST end with exactly one {@code bridge_reply} carrying the complete answer. Kept as a single
* logical line so it reads the same way the opencode charter does (Markdown wraps it on render). */
private static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet, running under codex. "
+ "Every message you receive arrives through the bridge, and the ONLY channel back to the "
+ "sender is the bridge_reply MCP tool. Text you write in your terminal is NOT sent "
+ "anywhere — the sender cannot see your screen, so an in-terminal answer is silently "
+ "discarded. Therefore you MUST end EVERY turn by calling bridge_reply with `content` set "
+ "to your complete response. This holds for every message without exception — tasks, "
+ "questions, clarifications, acknowledgements, and ordinary back-and-forth conversation. "
+ "Call bridge_reply exactly once, as the final action of your turn, with your full answer "
+ "in `content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
@Override
public Path provision(BridgedConfig.Worker cfg) {
Path home;
try {
home = Files.createTempDirectory(root, "bridged-codex-");
} catch (IOException e) {
throw new UncheckedIOException("cannot create Codex home for profile " + cfg.profile(), e);
}
try {
if (cfg.hasMcp()) {
writeConfig(home, cfg);
}
writeCharter(home);
provisionCredential(home);
return home;
} catch (RuntimeException e) {
// Every failure below is already a RuntimeException (wrapped IOException, or the missing-
// credential IllegalStateException). Best-effort remove the partially-built home so a
// failed spawn does not leak a temp dir, then rethrow.
try {
release(home);
} catch (RuntimeException ignored) {
// A cleanup failure must not mask the provisioning failure that got us here.
}
throw e;
}
}
@Override
public void release(Path home) {
if (home == null || !Files.exists(home)) {
return; // already released, or never provisioned — teardown races teardown
}
try (var stream = Files.walk(home)) {
// Delete deepest-first so directories are empty when their turn comes.
stream.sorted(Comparator.reverseOrder()).forEach(p -> {
try {
Files.deleteIfExists(p);
} catch (IOException e) {
throw new UncheckedIOException("cannot remove " + p, e);
}
});
} catch (IOException e) {
throw new UncheckedIOException("cannot walk " + home + " for release", e);
}
}
/** Write {@code config.toml} registering the bridge as a streamable-HTTP MCP server. */
private void writeConfig(Path home, BridgedConfig.Worker cfg) {
StringBuilder sb = new StringBuilder("[mcp_servers.bridged]\n");
sb.append("url = ").append(tomlString(cfg.mcpUrl())).append('\n');
if (cfg.tokenEnv() != null && !cfg.tokenEnv().isBlank()) {
sb.append("bearer_token_env_var = ").append(tomlString(cfg.tokenEnv())).append('\n');
}
writeString(home.resolve("config.toml"), sb.toString());
}
/** Write {@code AGENTS.md} carrying the reply charter (always — the standing instruction applies
* to every codex peer, not only those that mount the bridge). */
private void writeCharter(Path home) {
writeString(home.resolve("AGENTS.md"), AGENTS_HEADER + REPLY_CHARTER + "\n");
}
/** Copy {@code auth.json} from the operator's real codex home into the fresh home. A freshly
* created home has no credential, and codex then fails closed with an opaque {@code 401} mid-run
* rather than a spawn error — so a missing source is thrown from here, loudly, naming the path. */
private void provisionCredential(Path home) {
Path source = credentialSource();
if (!Files.isRegularFile(source)) {
throw new IllegalStateException("no Codex credential for the peer: looked for "
+ source + " but it does not exist. A fresh CODEX_HOME has no credentials and the "
+ "peer would otherwise fail with an opaque 401 Unauthorized mid-run; provision one "
+ "at that path or mount the operator's codex home before spawning codex peers.");
}
try {
Files.copy(source, home.resolve("auth.json"), StandardCopyOption.COPY_ATTRIBUTES);
} catch (IOException e) {
throw new UncheckedIOException("cannot copy Codex credential from " + source, e);
}
}
/** The operator's codex credential source: {@code CODEX_HOME}/auth.json when the env var is set,
* else {@code ~/.codex}/auth.json. */
private Path credentialSource() {
String codexHome = env.apply("CODEX_HOME");
Path homeDir = (codexHome == null || codexHome.isBlank())
? Path.of(System.getProperty("user.home"), ".codex")
: Path.of(codexHome);
return homeDir.resolve("auth.json");
}
private static void writeString(Path file, String content) {
try {
Files.writeString(file, content);
} catch (IOException e) {
throw new UncheckedIOException("cannot write " + file, e);
}
}
/** Render {@code s} as a TOML basic string, escaping what TOML requires. Values here are
* operator- or config-supplied (a URL, an env-var name), so escaping must be real — a malformed
* config.toml makes codex fail in a way that looks like a network problem. */
private static String tomlString(String s) {
StringBuilder sb = new StringBuilder("\"");
for (int i = 0; i < s.length(); i++) {
char c = s.charAt(i);
switch (c) {
case '\\' -> sb.append("\\\\");
case '"' -> sb.append("\\\"");
case '\n' -> sb.append("\\n");
case '\r' -> sb.append("\\r");
case '\t' -> sb.append("\\t");
default -> {
if (c < 0x20) {
sb.append(String.format("\\u%04X", (int) c));
} else {
sb.append(c);
}
}
}
}
return sb.append('"').toString();
}
}
@@ -0,0 +1,600 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.Collection;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ConcurrentMap;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* Abstract base for {@link PeerLauncher} adapters that materialize a peer as a <em>herdr</em>
* agent (a CLI coding agent running in a herdr tab/pane). It owns everything that is the same
* regardless of <em>which</em> coding agent runs: tab/pane placement, the CB-306 spawn-readiness
* gate, unique naming, CB-117 orphan reap, teardown, {@link #list() listing}, and cwd resolution.
*
* <p>Two seams are peer-specific and supplied by the concrete adapter:
* <ul>
* <li>{@code namePrefix} (constructor arg) — the label prefix ({@code claude}, {@code opencode})
* that drives both unique naming and the orphan-reap pattern, so each adapter reaps only its
* own kind of pane and never another's.</li>
* <li>{@link #buildLaunch(BridgedConfig.Worker)} — the peer-specific env map + argv, including any
* subscription/guard check, MCP mount, and instruction injection. The base never sees how the
* peer is configured; it only places and starts the returned {@link Launch}.</li>
* </ul>
*
* <p>Placement: in the default {@code tab} policy a peer lands in its own tab inside a dedicated
* worker space (found-or-created once, then shared), so peers never split or clutter the user's
* real work spaces. Teardown removes the peer's pane <em>and</em> its now-empty tab, tolerating an
* already-gone peer so a repeated DELETE is harmless.
*/
public abstract class HerdrPeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(HerdrPeerLauncher.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
/**
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
*/
private static final int SHELL_READY_RETRIES = 20;
private final String namePrefix; // label prefix: naming + reap scheme
private final AgentControl agents;
private final WorkspaceControl spaces;
private final Map<String, BridgedConfig.Worker> profiles; // profile name → spawn settings
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
/** Host env lookup (injectable for tests); adapters read it in {@link #buildLaunch}. */
protected final Function<String, String> env;
private final AtomicLong nameSeq = new AtomicLong(); // per-peer counter (also the tab #)
private final long spawnReadyTimeoutMs; // 0 = disable gate (legacy non-blocking spawn)
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
private final Runnable sleeper; // sleep/wait hook (injectable for tests; never real-sleep in unit tests)
// Per-process token mixed into each peer name so a fresh process (nameSeq back at 0) cannot
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
// CB-519: PeerHandle.id() is a host-unique opaque UUID, decoupled from the herdr pane id. The
// routing/registry key is the UUID; the herdr pane id is a launcher-private placement/teardown
// coordinate. This map bridges the two so stop(id) can resolve a host-unique key back to the
// exact pane it must tear down. The pane id is launcher-private (never the routing key) — see
// PeerHandle.id().
private final ConcurrentMap<String, String> paneByAgentId = new ConcurrentHashMap<>();
/**
* @param namePrefix label prefix for this peer kind (drives naming and reap)
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param profiles configured peer profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (never called when the gate is disabled); the poll
* interval is baked into this hook, so the base needs no poll field
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
this.namePrefix = namePrefix;
this.agents = agents;
this.spaces = spaces;
this.profiles = Map.copyOf(profiles);
this.defaultProfile = defaultProfile;
this.env = env;
this.spawnReadyTimeoutMs = spawnReadyTimeoutMs;
this.nowMillis = nowMillis;
this.sleeper = sleeper;
}
// --- adapter seams -------------------------------------------------------------------------
/**
* Build the peer-specific launch for {@code cfg}: the environment map and argv handed to herdr.
* Any subscription/guard check, MCP mount, and instruction injection happen here. The env map
* and argv are adapter-private; the base only places and starts what is returned.
*/
protected abstract Launch buildLaunch(BridgedConfig.Worker cfg);
/** A peer-specific launch: the herdr {@code env} map and {@code argv}. */
protected record Launch(Map<String, String> env, List<String> argv) {
}
// --- profile surface -----------------------------------------------------------------------
/** The configured peer profile names (what {@code spawn(profile)} accepts). */
@Override
public Set<String> profiles() {
return profiles.keySet();
}
/** The parity-overlay file list for {@code profileName} (default list when unset). */
@Override
public List<String> parityOverlay(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
return List.of();
}
BridgedConfig.Worker cfg = profiles.get(name);
return cfg == null ? List.of() : cfg.parityOverlay();
}
/** The profile a no-argument spawn uses, or {@code null} if none is configured. */
@Override
public String defaultProfile() {
return defaultProfile;
}
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
protected Collection<BridgedConfig.Worker> profileConfigs() {
return profiles.values();
}
/** Resolve {@code profileName} (null/blank → default) to its config, or throw with the options. */
protected BridgedConfig.Worker requireProfile(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
throw new IllegalArgumentException("no default worker profile is configured — "
+ "pass a profile; configured: " + profiles.keySet());
}
BridgedConfig.Worker cfg = profiles.get(name);
if (cfg == null) {
throw new IllegalArgumentException("unknown worker profile '" + name
+ "' — configured: " + profiles.keySet());
}
return cfg;
}
// --- spawn ---------------------------------------------------------------------------------
/**
* Spawn a peer. {@code profileName} null/blank → the default profile. The working directory
* (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd}, else the
* profile's configured {@code cwd}, else {@code callerCwd} (the primary's cwd, when the spawn
* came from the primary over MCP), else the daemon's cwd — never assumed to be {@code $HOME}.
* The adapter's {@link #buildLaunch} runs before any herdr call.
*/
protected Agent spawnInternal(String profileName, String requestedCwd, String callerCwd) {
BridgedConfig.Worker cfg = requireProfile(profileName);
Launch launch = buildLaunch(cfg);
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
return cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd);
}
/**
* {@inheritDoc}
*
* <p>Delegates to {@link #spawnInternal} and wraps the resulting herdr {@link Agent} in a
* {@link WorkerHandle} whose {@link PeerHandle#id()} is a fresh <em>host-unique</em> opaque
* UUID (CB-519), deliberately decoupled from the herdr pane id: the id is the registry/routing
* key and must never collide across daemon processes on the same host, while the herdr pane id
* stays a launcher-private placement/teardown coordinate, remembered here so {@link #stop}
* can resolve the host-unique key back to its pane. When {@code spawnReadyTimeoutMs > 0},
* blocks until the peer's herdr status is injectable or the timeout elapses; on timeout the
* pane is closed (no orphan) and a {@link PeerUnreachableException} is thrown.
*/
@Override
public PeerHandle spawn(SpawnRequest req) {
Agent agent = spawnInternal(req.profileName(), req.requestedCwd(), req.callerCwd());
String paneId = agent.paneId();
if (spawnReadyTimeoutMs > 0) {
waitUntilInjectableOrThrow(paneId);
}
// CB-519: the handle id is a host-unique UUID; the herdr pane it maps to stays internal.
String id = UUID.randomUUID().toString();
paneByAgentId.put(id, paneId);
return new WorkerHandle(id, agent.terminalId(), requireProfile(req.profileName()).profile());
}
@Override
public String effectiveCwd(SpawnRequest req) {
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
}
/**
* CB-301: the effective working directory a spawn for {@code profileName} would use, without
* actually spawning.
*/
private String effectiveCwd(String profileName, String requestedCwd, String callerCwd) {
return resolveCwd(requestedCwd, requireProfile(profileName), callerCwd);
}
/**
* CB-112 cwd resolution: spawn arg → profile config → the primary's cwd → the daemon's cwd.
* Never returns {@code null}/blank: {@code "."} (the daemon's own working directory) is the
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
*/
private static String resolveCwd(String requestedCwd, BridgedConfig.Worker cfg, String callerCwd) {
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
}
private static String firstNonBlank(String... values) {
for (String v : values) {
if (v != null && !v.isBlank()) return v;
}
return null;
}
/** Dedicated worker space → own tab (carrying cwd+env) → start the peer into the seed pane. */
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId(), cwd, workerEnv);
log.info("spawning {} profile={} space={} tab={} cwd={}",
namePrefix, cfg.profile(), space.workspaceId(), tab.tab().tabId(), cwd);
Started started;
try {
if (tab.rootPaneId() == null) {
// Protocol 19 starts the agent INTO the seed pane — without one there is nowhere
// to start, and a partial tab would be left behind.
throw new IllegalStateException("tab " + tab.tab().tabId()
+ " had no seed pane in the create response — cannot start a peer in it");
}
started = startUniquelyNamed(cfg, argv, tab.rootPaneId());
} catch (RuntimeException e) {
// The peer never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("failed to close orphaned tab {} after spawn error: {}",
tab.tab().tabId(), cleanup.getMessage());
}
throw e;
}
// The peer is LIVE now, in the seed pane itself (no shell pane to drop — protocol 19).
// Labelling is cosmetic: it must not fail the spawn or orphan the running peer — on error
// we log and still return it so the caller gets its paneId and can tear it down.
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
log.info("{} started pane={} tab={} terminal={}",
namePrefix, started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
return started.agent();
}
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
private void tidy(String what, Runnable step) {
try {
step.run();
} catch (RuntimeException e) {
log.warn("post-start step failed ({}) — peer is running regardless: {}", what, e.getMessage());
}
}
/** Legacy placement: split the currently-focused tab; the peer still starts in {@code cwd}. */
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
namePrefix, cfg.profile(), cwd, argv);
String paneId = spaces.splitPane(cwd, workerEnv);
if (paneId == null) {
throw new IllegalStateException("pane.split returned no pane — cannot start a peer");
}
Agent peer = startUniquelyNamed(cfg, argv, paneId).agent();
log.info("{} started pane={} terminal={}", namePrefix, peer.paneId(), peer.terminalId());
return peer;
}
/** A started peer together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
/**
* Start the peer under a unique herdr agent name. herdr requires each running agent's
* {@code name} to be distinct (a 2nd identical {@code name} fails {@code agent_name_taken}) —
* the exact case that makes multiple peers useful. The name is
* {@code <prefix>-<profile>-<nonce>-<seq>}: {@code seq} distinguishes peers within this process,
* and the per-process {@code nonce} keeps a fresh process (whose {@code seq} restarts at 0) from
* colliding with same-profile peers that outlived a restart. The retry is a belt-and-braces
* backstop for the astronomically unlikely nonce+seq clash; the name is a label only — herdr
* detects kind and status from terminal output, not from it.
*/
private Started startUniquelyNamed(BridgedConfig.Worker cfg, List<String> argv, String paneId) {
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
// argv[0] — the configured executable — is dropped and only the extra args are passed.
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(startAwaitingShellPrompt(name, args, paneId), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("peer name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
/** Start the agent into {@code paneId}, waiting out the seed shell's boot with the sleeper. */
private Agent startAwaitingShellPrompt(String name, List<String> args, String paneId) {
HerdrException busy = null;
for (int attempt = 0; attempt < SHELL_READY_RETRIES; attempt++) {
try {
return agents.start(name, namePrefix, args, paneId);
} catch (HerdrException e) {
if (!"agent_pane_busy".equals(e.code())) throw e;
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
busy = e;
sleeper.run();
}
}
throw busy;
}
// --- discovery + reap ----------------------------------------------------------------------
/** All herdr-tracked agents — discovery for "what peers exist". */
@Override
public List<Agent> list() {
return agents.list();
}
/**
* Reap peer panes left behind by an earlier daemon process (CB-117). herdr keeps a peer's pane
* alive across a daemon restart <em>by design</em>, and that pane's id is held only by its
* spawner — so a peer whose owning process exited before issuing the matching teardown leaks
* with nothing tracking it. On boot we scan herdr for agents whose name matches our
* {@code <prefix>-<profile>-<nonce>-<seq>} scheme with a nonce <em>other</em> than this
* process's {@link #nameNonce}, and tear each one down (its pane and, via {@link #stop}, its
* now-empty dedicated tab). A current-nonce peer is ours and live, so it is left running; a
* user's own session carries no such name and is never touched. A peer from a <em>different</em>
* adapter (different prefix) is likewise never touched. Best-effort: a failed listing, or a
* failure to stop any one peer, is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
@Override
public int reapOrphanWorkers() {
List<Agent> all;
try {
all = agents.list();
} catch (RuntimeException e) {
log.warn("orphan-peer reap skipped — agent.list failed: {}", e.getMessage());
return 0;
}
int reaped = 0;
for (Agent a : all) {
if (!isForeignWorker(namePrefix, a.name(), nameNonce)) continue;
try {
stop(a.paneId());
reaped++;
log.info("reaped orphan {} {} (pane={} tab={}) left by a prior daemon",
namePrefix, a.name(), a.paneId(), a.tabId());
} catch (RuntimeException e) {
log.warn("could not reap orphan {} {} (pane={}): {}",
namePrefix, a.name(), a.paneId(), e.getMessage());
}
}
if (reaped > 0) {
log.info("orphan-peer reap complete — {} stale {} peer(s) removed at startup", reaped, namePrefix);
}
return reaped;
}
/** The {@code <prefix>-<profile>-<nonce>-<seq>} name pattern; group 1 captures the 6-hex nonce. */
static Pattern workerNamePattern(String prefix) {
return Pattern.compile(prefix + "-.*-([0-9a-f]{6})-\\d+");
}
/**
* Whether {@code name} is a peer of kind {@code prefix} started by a <em>different</em> process
* than {@code currentNonce} — the reap predicate (CB-117). True only for the prefix's naming
* scheme with a foreign nonce: a non-peer name, a different adapter's name, or our own live
* nonce is excluded. Pure and package-private so the decision is unit-testable without herdr.
*/
static boolean isForeignWorker(String prefix, String name, String currentNonce) {
String nonce = workerNonce(prefix, name);
return nonce != null && !nonce.equals(currentNonce);
}
/** The 6-hex nonce embedded in a {@code prefix} peer name, or {@code null} if not one. */
static String workerNonce(String prefix, String name) {
if (name == null) return null;
Matcher m = workerNamePattern(prefix).matcher(name);
return m.matches() ? m.group(1) : null;
}
/** This process's peer-name nonce (a label component only; exposed for reaper tests). */
String nameNonce() {
return nameNonce;
}
// --- teardown ------------------------------------------------------------------------------
/**
* Tear a peer down: close the pane, and close its tab <em>only</em> when the peer is that tab's
* sole occupant. The single-pane check is what makes this safe regardless of how the peer was
* placed (or a placement-config change across a restart): a pane-placement peer sitting in one
* of the user's shared tabs has siblings, so its tab is never closed — we only ever remove a
* tab we created to hold one peer.
*
* <p>{@code idOrPane} is the {@link PeerHandle#id()} of a peer this launcher spawned (CB-519's
* host-unique opaque UUID), resolved through {@link #paneByAgentId} to the pane it must tear
* down. An argument that is not one of our ids is treated as a raw herdr pane id — the
* {@link #reapOrphanWorkers() orphan-reap} and spawn-gate-timeout paths, plus any caller that
* passes a pane directly, keep working without an owning id.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed peer) is treated as success; any other failure propagates so a
* genuinely failed teardown is not reported as done.
*/
@Override
public void stop(String idOrPane) {
// Teardown knows only the pane, not which profile spawned it. Attempt tab cleanup when any
// profile uses tab placement (so the bridge may have created a dedicated peer tab); the
// single-occupant check below is what actually protects the user's shared tabs.
String paneId = paneByAgentId.remove(idOrPane);
if (paneId == null) {
paneId = idOrPane; // raw-pane fallback (reap, gate timeout, pane-addressed callers)
}
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) throw e;
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
spaces.closeTab(loc.tabId());
} else if (loc != null) {
log.debug("not closing tab {} — it holds {} panes (not a dedicated peer tab)",
loc.tabId(), loc.tabPaneCount());
}
}
/** Whether any configured profile places peers in their own tab (so tabs may need cleanup). */
private boolean usesTabPlacement() {
return profiles.values().stream().anyMatch(BridgedConfig.Worker::tabPlacement);
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
// --- spawn-readiness gate (CB-306) ---------------------------------------------------------
/**
* Poll {@link AgentControl#status} until the pane reports an injectable state or the configured
* timeout elapses. On timeout, close the pane (self-reap) and throw.
*/
private void waitUntilInjectableOrThrow(String paneId) {
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
while (nowMillis.getAsLong() < deadline) {
if (agents.status(paneId).injectable()) {
log.debug("peer pane={} reached injectable state", paneId);
return;
}
sleeper.run();
}
log.warn("peer pane={} did not become injectable within {}ms — closing", paneId, spawnReadyTimeoutMs);
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " did not reach injectable state within "
+ spawnReadyTimeoutMs + "ms");
}
/** A concrete {@link PeerHandle} wrapping herdr agent coordinates and the profile that spawned it. */
private record WorkerHandle(String id, String terminalId, String profile) implements PeerHandle {
}
// --- shared helpers ------------------------------------------------------------------------
/** Put {@code k → v} only when {@code v} is present (non-null, non-blank). */
protected static void putIfPresent(Map<String, String> m, String k, String v) {
if (v != null && !v.isBlank()) {
m.put(k, v);
}
}
/** Host env lookup that tolerates an unconfigured (null/blank) var name — returns null then. */
protected String resolveEnv(String name) {
return (name == null || name.isBlank()) ? null : env.apply(name);
}
/**
* The parity-neutral git-forge token grant (CB-302): when {@code cfg} opts in via
* {@code gitTokenEnv} and the token resolves, inject {@code GITEA_TOKEN} plus its paired
* {@code GITEA_HOST}. Push over SSH is unaffected; the only incremental grant is PR-create.
* Peer-neutral, so every herdr adapter reuses it unchanged.
*/
protected void applyGitToken(Map<String, String> workerEnv, BridgedConfig.Worker cfg) {
if (!cfg.hasGitToken()) {
return;
}
String gitToken = resolveEnv(cfg.gitTokenEnv());
if (gitToken != null) {
workerEnv.put("GITEA_TOKEN", gitToken);
putIfPresent(workerEnv, "GITEA_HOST", resolveEnv(cfg.gitHostEnv()));
}
}
/** A fresh mutable env map — the conventional starting point for {@link #buildLaunch}. */
/**
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
* {@code env:} entries.
*
* <p>Why this exists: bridged passes herdr an explicit env map, and herdr merges it into
* <em>its own</em> process environment. So before this, a worker inherited whatever PATH the
* herdr server happened to be started with — on this host, one from weeks earlier with no JDK
* and no Maven, which left workers unable to run the build they were being asked to run. The
* worker's toolchain must follow from configuration, not from how a long-lived daemon was
* launched.
*
* <p>Adapter-specific variables are layered on top of this by {@code buildLaunch} and therefore
* win. That ordering is deliberate and load-bearing: it stops a profile's {@code env:} from
* overriding {@code ANTHROPIC_BASE_URL} and slipping past {@link
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
* {@code baseUrl} and nothing else.
*/
protected Map<String, String> baseEnv(BridgedConfig.Worker cfg) {
Map<String, String> workerEnv = new LinkedHashMap<>();
String path = env.apply("PATH");
if (path != null && !path.isBlank()) {
workerEnv.put("PATH", path);
}
if (cfg != null && cfg.env() != null) {
workerEnv.putAll(cfg.env());
}
return workerEnv;
}
/** Defensive copy of {@code argv} plus room to append launch flags. */
protected static List<String> mutableArgv(List<String> argv) {
return new ArrayList<>(argv);
}
/**
* Uninterruptible sleep — the production {@link #sleeper}. Tests supply their own no-op /
* fast-faking sleeper so they never real-sleep.
*/
protected static void sleepUninterruptibly(long ms) {
try {
Thread.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
// preserve the interrupt flag but continue — poll loops should not be aborted by an
// interrupt that was not meant for them.
}
}
}
@@ -0,0 +1,316 @@
package dev.ltms.bridged.worker;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.node.ObjectNode;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>opencode</strong> — an open-source,
* provider-agnostic terminal coding agent. Its whole reason for existing is to prove the
* {@code PeerLauncher} SPI is genuinely provider-neutral: opencode shares none of Claude Code's
* private launch seams, yet reuses every line of shared transport in the base (tab/pane placement,
* the CB-306 readiness gate, unique naming + CB-117 reap, teardown, listing, cwd).
*
* <p>The divergences from {@link ClaudeCodeLauncher}, all confined to {@link #buildLaunch}:
* <ul>
* <li><strong>No subscription boundary.</strong> opencode carries no {@code ANTHROPIC_BASE_URL}
* and there is no {@link dev.ltms.bridged.guard.SubscriptionGuard} — the guard is a
* Claude-private concern, not part of the SPI. opencode reads the operator's own provider
* credentials from its global {@code auth.json}; the bridge injects none.</li>
* <li><strong>File-based MCP mount + instructions.</strong> opencode has no inline
* {@code --mcp-config}/{@code --append-system-prompt}. Instead the bridge writes an ephemeral
* {@code opencode.json} that declares the bridge as a {@code remote} MCP server and lists a
* reply-charter file under {@code instructions}, then points the worker at it with
* {@code OPENCODE_CONFIG}. This is the one place the launcher touches disk — Claude never did.</li>
* <li><strong>Model as a flag.</strong> the {@code provider/model} selector is passed as
* {@code -m}, not an env var.</li>
* <li><strong>{@code opencode} name prefix</strong> so reap matches {@code opencode-*} panes and
* never another adapter's.</li>
* </ul>
*/
public final class OpenCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "opencode";
/** Writer for the generated {@code opencode.json}. */
private static final ObjectMapper JSON = new ObjectMapper();
/**
* Standing instruction written to the charter file and mounted via the config's
* {@code instructions} so the worker returns its result through {@code bridge_reply}. Kept on
* disk (not a launch flag) because opencode's {@code instructions} takes file paths, not inline
* text — the file is regenerated per spawn and never touches the worker's own profile.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet, running under opencode. "
+ "Every message you receive arrives through the bridge, and the ONLY channel back to the "
+ "sender is the bridge_reply MCP tool. Text you write in your terminal is NOT sent "
+ "anywhere — the sender cannot see your screen, so an in-terminal answer is silently "
+ "discarded. Therefore you MUST end EVERY turn by calling bridge_reply with `content` set "
+ "to your complete response. This holds for every message without exception — tasks, "
+ "questions, clarifications, acknowledgements, and ordinary back-and-forth conversation. "
+ "Call bridge_reply exactly once, as the final action of your turn, with your full answer "
+ "in `content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/** Root under which per-spawn opencode config dirs are created (injectable for tests). */
private final Path configRoot;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so it
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300),
defaultConfigRoot());
}
/**
* Production constructor with the spawn-ready gate enabled. Polls {@code agents.status()} until
* the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot());
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply a
* fake clock ({@code nowMillis}), poll-loop wait ({@code sleeper}), and a temp {@code configRoot}
* they can inspect the generated {@code opencode.json}/charter under.
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (encodes the poll interval; never called when the
* gate is disabled)
* @param configRoot existing directory under which per-spawn config dirs are created
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper, Path configRoot) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.configRoot = configRoot;
}
private static Path defaultConfigRoot() {
return Path.of(System.getProperty("java.io.tmpdir"));
}
/**
* {@inheritDoc}
*
* <p>Builds the opencode launch: no {@code ANTHROPIC_*} and no guard (opencode reads its own
* provider credentials); when the profile mounts the bridge MCP, generate an ephemeral
* {@code opencode.json} (remote MCP server + reply-charter instructions) and point the worker at
* it via {@code OPENCODE_CONFIG}; carry the parity-neutral git-forge grant; and select the model
* with {@code -m}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
Map<String, String> workerEnv = baseEnv(cfg);
// A config file is needed for the bridge MCP mount, for a pinned endpoint (CB-508), or both.
if (cfg.hasMcp() || hasCustomProvider(cfg)) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg).toString());
}
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithModel(cfg));
}
/**
* True when this profile pins its own OpenAI-compatible endpoint (CB-508) rather than using
* whatever provider opencode resolves by default.
*
* <p>Note this reuses {@code baseUrl}, the same field the Claude adapter injects as
* {@code ANTHROPIC_BASE_URL} — but it does <em>not</em> go through {@code SubscriptionGuard}.
* That asymmetry is deliberate and safe: the guard exists to stop a worker borrowing the
* primary's Anthropic subscription, and an opencode process has no Anthropic credential path
* at all. Pointing it at a local vLLM cannot leak the subscription.
*/
private static boolean hasCustomProvider(BridgedConfig.Worker cfg) {
return cfg.baseUrl() != null && !cfg.baseUrl().isBlank();
}
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
private List<String> argvWithModel(BridgedConfig.Worker cfg) {
List<String> argv = mutableArgv(cfg.argv());
if (cfg.model() != null && !cfg.model().isBlank()) {
argv.add("-m");
argv.add(cfg.model());
}
return argv;
}
/**
* Write an ephemeral {@code opencode.json} (and the reply-charter file it references) into a
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
*/
private Path writeConfig(BridgedConfig.Worker cfg) {
try {
Path dir = Files.createTempDirectory(configRoot, "bridged-opencode-");
dir.toFile().deleteOnExit();
ObjectNode root = JSON.createObjectNode();
root.put("$schema", "https://opencode.ai/config.json");
if (cfg.hasMcp()) {
Path charter = dir.resolve("reply-charter.md");
Files.writeString(charter, REPLY_CHARTER);
charter.toFile().deleteOnExit();
ObjectNode bridge = root.putObject("mcp").putObject("bridge");
bridge.put("type", "remote");
bridge.put("url", cfg.mcpUrl());
bridge.put("enabled", true);
root.putArray("instructions").add(charter.toAbsolutePath().toString());
}
if (hasCustomProvider(cfg)) {
addCustomProvider(root, cfg);
}
Path cfgFile = dir.resolve("opencode.json");
// Built with Jackson rather than string concatenation: the provider block is nested and
// carries operator-supplied values (URL, model id, api key), so escaping must be real.
Files.writeString(cfgFile, JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root));
cfgFile.toFile().deleteOnExit();
return cfgFile;
} catch (IOException e) {
throw new UncheckedIOException(
"cannot write opencode config for profile " + cfg.profile(), e);
}
}
/**
* Declare a custom OpenAI-compatible provider so the worker talks to a pinned endpoint (a local
* vLLM, say) instead of opencode's default gateway (CB-508).
*
* <p>The provider id comes from the {@code provider/model} selector in {@code model:}, so one
* field drives both the declaration and the {@code -m} flag and they cannot drift apart.
*/
private void addCustomProvider(ObjectNode root, BridgedConfig.Worker cfg) {
String[] parts = splitModelSelector(cfg);
String providerId = parts[0];
String modelId = parts[1];
ObjectNode provider = root.putObject("provider").putObject(providerId);
provider.put("npm", "@ai-sdk/openai-compatible");
provider.put("name", providerId + " (bridged)");
ObjectNode options = provider.putObject("options");
options.put("baseURL", openAiBaseUrl(cfg.baseUrl()));
// vLLM and friends usually ignore the key, but the AI SDK still requires a non-empty one.
String token = resolveEnv(cfg.tokenEnv());
options.put("apiKey", (token == null || token.isBlank()) ? "bridged-local-noauth" : token);
provider.putObject("models").putObject(modelId).put("name", modelId);
}
/**
* Split {@code model:} into its {@code provider} and {@code model} halves. A pinned endpoint
* needs both, so a bare model name is rejected loudly rather than silently falling back to the
* default gateway — a worker quietly talking to the wrong endpoint is the failure this avoids.
*/
private static String[] splitModelSelector(BridgedConfig.Worker cfg) {
String model = cfg.model();
int slash = model == null ? -1 : model.indexOf('/');
if (model == null || model.isBlank() || slash <= 0 || slash == model.length() - 1) {
throw new IllegalArgumentException(
"profile " + cfg.profile() + " sets baseUrl (a pinned opencode endpoint) so"
+ " model: must be \"<provider>/<model>\", e.g."
+ " \"local-vllm/deepseek-v4-flash\"; got "
+ (model == null ? "null" : '"' + model + '"'));
}
return new String[]{model.substring(0, slash), model.substring(slash + 1)};
}
/**
* The OpenAI-compatible base URL for {@code baseUrl}. A bare {@code host:port} gets {@code /v1}
* appended (where these servers put the API); a URL that already carries a path is taken as-is,
* so an endpoint mounted somewhere unusual is still reachable.
*/
private static String openAiBaseUrl(String baseUrl) {
String trimmed = baseUrl.trim();
while (trimmed.endsWith("/")) {
trimmed = trimmed.substring(0, trimmed.length() - 1);
}
int schemeEnd = trimmed.indexOf("://");
String afterScheme = schemeEnd < 0 ? trimmed : trimmed.substring(schemeEnd + 3);
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
// --- CB-117 reap predicate (opencode prefix), kept for direct unit testing -----------------
/**
* Whether {@code name} is an opencode bridge worker started by a <em>different</em> process than
* {@code currentNonce}. A thin {@code opencode}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -1,206 +0,0 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
/**
* Spawns and lists worker sessions — the safe path from a delegation request to a
* running off-subscription Claude.
*
* <p>The spawn sequence encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em>
* touching herdr, and only then {@code agent.start}. A worker's base_url lives in the
* env map handed to herdr and nowhere else; {@code bridged}'s own environment is never
* mutated.
*
* <p>Placement: in the default {@code tab} policy a worker lands in its own tab inside a
* dedicated worker space (found-or-created once, then shared), so workers never split or
* clutter the user's real work spaces. Teardown removes the worker's pane <em>and</em> its
* now-empty tab, tolerating an already-gone worker so a repeated DELETE is harmless.
*/
public final class WorkerService {
private static final Logger log = LoggerFactory.getLogger(WorkerService.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
private final AgentControl agents;
private final WorkspaceControl spaces;
private final SubscriptionGuard guard;
private final BridgedConfig.Worker cfg;
private final Function<String, String> env; // host env lookup (injectable for tests)
private final AtomicLong nameSeq = new AtomicLong(); // per-worker counter (also the tab #)
// Per-process token mixed into each worker name so a fresh process (nameSeq back at 0)
// cannot collide with same-profile workers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
public WorkerService(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
BridgedConfig.Worker cfg, Function<String, String> env) {
this.agents = agents;
this.spaces = spaces;
this.guard = guard;
this.cfg = cfg;
this.env = env;
}
/** Spawn a worker for the configured profile. Guard runs before any herdr call. */
public Agent spawn() {
String baseUrl = cfg.baseUrl();
guard.assertWorker(baseUrl); // hard stop before we spawn anything
Map<String, String> workerEnv = new LinkedHashMap<>();
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
String token = env.apply(cfg.tokenEnv());
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", token);
return cfg.tabPlacement() ? spawnInTab(workerEnv) : spawnAsPane(workerEnv);
}
/** Dedicated worker space → own tab → drop the placeholder shell so only the worker remains. */
private Agent spawnInTab(Map<String, String> workerEnv) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId());
log.info("spawning worker profile={} base_url={} space={} tab={}",
cfg.profile(), cfg.baseUrl(), space.workspaceId(), tab.tab().tabId());
Started started;
try {
started = startUniquelyNamed(workerEnv, tab.tab().tabId());
} catch (RuntimeException e) {
// The worker never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("failed to close orphaned tab {} after spawn error: {}",
tab.tab().tabId(), cleanup.getMessage());
}
throw e;
}
// The worker is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell
// so the tab holds only the worker; label the tab). They must not fail the spawn or
// orphan the running worker — on error we log and still return it so the caller gets
// its paneId and can tear it down.
if (tab.rootPaneId() != null) {
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
} else {
log.warn("tab {} had no seed pane in the create response; worker tab may hold an extra pane",
tab.tab().tabId());
}
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
log.info("worker started pane={} tab={} terminal={}",
started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
return started.agent();
}
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
private void tidy(String what, Runnable step) {
try {
step.run();
} catch (RuntimeException e) {
log.warn("post-start step failed ({}) — worker is running regardless: {}", what, e.getMessage());
}
}
/** Legacy placement: herdr splits the currently-focused tab. */
private Agent spawnAsPane(Map<String, String> workerEnv) {
log.info("spawning worker (pane placement) profile={} base_url={} argv={}",
cfg.profile(), cfg.baseUrl(), cfg.argv());
Agent worker = startUniquelyNamed(workerEnv, null).agent();
log.info("worker started pane={} terminal={}", worker.paneId(), worker.terminalId());
return worker;
}
/** A started worker together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
/**
* Start the worker under a unique herdr agent name. herdr requires each running
* agent's {@code name} to be distinct (a 2nd {@code name:"claude"} fails
* {@code agent_name_taken}) — the exact case that makes multiple workers useful. The name
* is {@code claude-<profile>-<nonce>-<seq>}: {@code seq} distinguishes workers within this
* process, and the per-process {@code nonce} keeps a fresh process (whose {@code seq}
* restarts at 0) from colliding with same-profile workers that outlived a restart. The
* retry is a belt-and-braces backstop for the astronomically unlikely nonce+seq clash;
* the name is a label only — herdr detects kind and status from terminal output, not it.
*/
private Started startUniquelyNamed(Map<String, String> workerEnv, String tabId) {
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = "claude-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(agents.start(name, cfg.argv(), workerEnv, tabId), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("worker name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
/** All herdr-tracked agents — discovery for "what workers exist". */
public List<Agent> list() {
return agents.list();
}
/**
* Tear a worker down by pane id: close the pane, and close its tab <em>only</em> when the
* worker is that tab's sole occupant. The single-pane check is what makes this safe
* regardless of how the worker was placed (or a placement-config change across a restart):
* a pane-placement worker sitting in one of the user's shared tabs has siblings, so its
* tab is never closed — we only ever remove a tab we created to hold one worker.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed worker) is treated as success; any other failure propagates so
* a genuinely failed teardown is not reported as done.
*/
public void stop(String paneId) {
WorkspaceControl.PaneLocation loc = cfg.tabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) throw e;
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
spaces.closeTab(loc.tabId());
} else if (loc != null) {
log.debug("not closing tab {} — it holds {} panes (not a dedicated worker tab)",
loc.tabId(), loc.tabPaneCount());
}
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
private static void putIfPresent(Map<String, String> m, String k, String v) {
if (v != null && !v.isBlank()) {
m.put(k, v);
}
}
}
+29
View File
@@ -5,10 +5,39 @@
</encoder>
</appender>
<!--
CB-505 audit trail. Its own file, deliberately not the app log: privileged actions
(spawn/stop/send/reply/drain) must stay greppable and shippable without dragging DEBUG noise
along. AuditLog emits a complete JSON object including its own ISO-8601 "ts" field, so the
pattern is a bare %msg — a pattern that spliced literal braces around the message would
collide with logback's own variable substitution. Rolls daily, 30 days retained, 100MB cap.
NOTE: records carry who/what/target/outcome only. Message CONTENT is never written here —
this bridge carries source code and prompts, and an audit log that accumulated them would be
a transcript archive rather than a control.
-->
<appender name="AUDIT" class="ch.qos.logback.core.rolling.RollingFileAppender">
<file>logs/audit.log</file>
<rollingPolicy class="ch.qos.logback.core.rolling.SizeAndTimeBasedRollingPolicy">
<fileNamePattern>logs/audit.%d{yyyy-MM-dd}.%i.log</fileNamePattern>
<maxFileSize>10MB</maxFileSize>
<maxHistory>30</maxHistory>
<totalSizeCap>100MB</totalSizeCap>
</rollingPolicy>
<encoder>
<pattern>%msg%n</pattern>
</encoder>
</appender>
<logger name="dev.ltms.bridged" level="DEBUG"/>
<logger name="io.javalin" level="INFO"/>
<logger name="org.eclipse.jetty" level="WARN"/>
<!-- additivity=false keeps the audit stream out of stdout; it is its own record. -->
<logger name="audit" level="INFO" additivity="false">
<appender-ref ref="AUDIT"/>
</logger>
<root level="INFO">
<appender-ref ref="STDOUT"/>
</root>
@@ -0,0 +1,100 @@
package dev.ltms.bridged.auth;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-505 — the audit record's shape.
*
* <p>These exist because the first cut of this feature emitted lines that were <em>not</em> valid
* JSON: the timestamp was spliced on by a logback pattern whose literal braces collided with
* logback's variable substitution. The appender failed to parse, and nothing in the build noticed.
* An audit trail that silently stops being machine-readable is worse than none.
*/
class AuditLogTest {
private final ObjectMapper mapper = new ObjectMapper();
private ListAppender<ILoggingEvent> appender;
private ch.qos.logback.classic.Logger auditLogger;
@BeforeEach
void attach() {
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
auditLogger = ctx.getLogger("audit");
appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
auditLogger.addAppender(appender);
auditLogger.setLevel(Level.INFO);
}
@AfterEach
void detach() {
auditLogger.detachAppender(appender);
}
private JsonNode onlyRecord() throws Exception {
assertEquals(1, appender.list.size(), "exactly one audit line expected");
String line = appender.list.getFirst().getFormattedMessage();
return mapper.readTree(line); // throws if the line is not valid JSON
}
@Test
void anAllowedActionIsRecordedAsValidJson() throws Exception {
AuditLog.allowed(Principal.primary(4242), Authz.Action.SPAWN, "term_a");
JsonNode r = onlyRecord();
assertEquals("PRIMARY", r.path("role").asText());
assertEquals("primary", r.path("actor").asText());
assertEquals(4242, r.path("pid").asLong());
assertEquals("SPAWN", r.path("action").asText());
assertEquals("term_a", r.path("target").asText());
assertEquals("allowed", r.path("outcome").asText());
assertFalse(r.path("ts").asText().isBlank(), "every record carries its own timestamp");
}
@Test
void aDenialRecordsTheReason() throws Exception {
AuditLog.denied(Principal.worker("term_b", 7), Authz.Action.REPLY, "term_a", "forbidden");
JsonNode r = onlyRecord();
assertEquals("WORKER", r.path("role").asText());
assertEquals("worker:term_b", r.path("actor").asText());
assertEquals("denied", r.path("outcome").asText());
assertEquals("forbidden", r.path("reason").asText());
}
@Test
void aNullCallerIsRecordedAsAnonymousRatherThanCrashing() throws Exception {
AuditLog.failed(null, Authz.Action.SEND, null, "herdr unreachable");
JsonNode r = onlyRecord();
assertEquals("ANONYMOUS", r.path("role").asText());
assertTrue(r.path("target").isNull(), "an absent target is JSON null, not the string \"null\"");
assertEquals("failed", r.path("outcome").asText());
}
@Test
void hostileValuesAreEscapedAndCannotForgeAnExtraRecord() throws Exception {
// A target id containing a quote and a newline must not be able to terminate the JSON
// object early and inject a second, attacker-shaped audit line.
AuditLog.denied(Principal.worker("term_a", 1), Authz.Action.REPLY,
"evil\",\"outcome\":\"allowed\"}\n{\"forged\":true", "forbidden");
JsonNode r = onlyRecord();
assertEquals("denied", r.path("outcome").asText(),
"the injected outcome must not override the real one");
assertTrue(r.path("target").asText().contains("forged"),
"the hostile text survives as inert data inside the target field");
}
}
@@ -0,0 +1,80 @@
package dev.ltms.bridged.auth;
import org.junit.jupiter.api.Test;
import static dev.ltms.bridged.auth.Authz.Action.*;
import static org.junit.jupiter.api.Assertions.*;
/** CB-505 — the authorization table, pinned so it cannot drift silently. */
class AuthzTest {
private static final Principal PRIMARY = Principal.primary(100);
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal WORKER_B = Principal.worker("term_b", 300);
private static final Principal ANON = Principal.anonymous();
@Test
void anonymousIsAuthorizedForNothing() {
for (Authz.Action a : Authz.Action.values()) {
assertFalse(Authz.permits(ANON, a, "term_a"),
a + " must be refused to an unauthenticated caller");
}
}
@Test
void aNullCallerIsTreatedAsAnonymous() {
assertFalse(Authz.permits(null, READ, null));
assertTrue(Authz.isUnauthenticated(null));
}
@Test
void orchestrationBelongsToThePrimaryAlone() {
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, SEND, DRAIN}) {
assertTrue(Authz.permits(PRIMARY, a, "term_a"), "the primary orchestrates: " + a);
assertFalse(Authz.permits(WORKER_A, a, "term_a"),
"a worker performing " + a + " would be escalating into the orchestrator role");
}
}
@Test
void aWorkerMayReplyAndAskOnlyAsItself() {
assertTrue(Authz.permits(WORKER_A, REPLY, "term_a"));
assertTrue(Authz.permits(WORKER_A, ASK, "term_a"));
assertFalse(Authz.permits(WORKER_A, REPLY, "term_b"),
"worker A must not be able to reply on worker B's session");
assertFalse(Authz.permits(WORKER_B, ASK, "term_a"),
"worker B must not be able to ask as worker A");
}
@Test
void thePrimaryMayNotForgeAWorkersReply() {
// Not a hypothetical nicety: a forged reply would resolve the rendezvous the primary is
// itself blocked on, corrupting the correlation between a turn and its answer.
assertFalse(Authz.permits(PRIMARY, REPLY, "term_a"));
assertFalse(Authz.permits(PRIMARY, ASK, "term_a"));
}
@Test
void aWorkerWithNoTargetCannotReply() {
assertFalse(Authz.permits(WORKER_A, REPLY, null),
"an absent session id must not satisfy the own-session rule");
}
@Test
void observationIsOpenToBothAuthenticatedRoles() {
assertTrue(Authz.permits(PRIMARY, READ, null));
assertTrue(Authz.permits(WORKER_A, READ, null));
assertTrue(Authz.permits(PRIMARY, METRICS, null));
assertTrue(Authz.permits(WORKER_A, METRICS, null));
}
@Test
void unauthenticatedIsDistinguishedFromMerelyForbidden() {
// Drives the 401-vs-403 split: a missing credential is fixable by the caller, a wrong role
// is not.
assertTrue(Authz.isUnauthenticated(ANON));
assertFalse(Authz.isUnauthenticated(WORKER_A));
assertFalse(Authz.isUnauthenticated(PRIMARY));
}
}
@@ -0,0 +1,143 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-501. The behaviour under test is the inversion of the pre-CB-501 default: failing every
* identity check must yield {@link Role#ANONYMOUS}, not {@code PRIMARY}.
*/
class CallerResolverTest {
private final FakeHerdr herdr = new FakeHerdr();
/** Identity resolving the canned worker pane, keyed off a faked peer-PID lookup. */
private ConnectionIdentity identity(long pid) {
return new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
}
/** A PID that owns a worker pane in the fake. */
private ConnectionIdentity workerIdentity() {
return identity(FakeHerdr.WORKER_PID);
}
/** A PID that owns no pane — i.e. the primary, or any other local process. */
private ConnectionIdentity nonWorkerIdentity() {
return identity(999_999);
}
@Test
void aLoopbackWorkerPaneResolvesToWorkerRegardlessOfAuthMode() {
Principal underTrust = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
Principal underToken = new CallerResolver(workerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, underTrust.role());
assertEquals("term_a", underTrust.terminal());
assertEquals(Role.WORKER, underToken.role(),
"worker identity is unforgeable and must never be token-gated — otherwise enabling "
+ "auth would lock the whole fleet out of bridge_reply");
assertEquals("term_a", underToken.terminal());
}
@Test
void aPinnedPrimaryTerminalResolvesToPrimaryNotWorker() {
// The primary's own session lives in a herdr pane (term_a here). Without the pin the pane
// match wins and the primary is locked out of spawn/send/stop as a misread worker.
Principal p = new CallerResolver(workerIdentity(), false, null, "term_a")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
}
@Test
void aPinnedPrimaryTerminalNeedsNoTokenEvenInTokenMode() {
Principal p = new CallerResolver(workerIdentity(), true, "s3cret", "term_a")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role(),
"the pane mapping is as unforgeable as a worker's — the pin outranks the token path");
}
@Test
void otherPanesRemainWorkersWhenAPinIsSet() {
Principal p = new CallerResolver(workerIdentity(), false, null, "term_someone_else")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertEquals("term_a", p.terminal());
}
/** The pin is optional config, so an absent or whitespace one must change nothing at all. */
@Test
void aBlankPinLeavesWorkerResolutionUntouched() {
assertEquals(Role.WORKER,
new CallerResolver(workerIdentity(), false, null, " ").resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER,
new CallerResolver(workerIdentity(), false, null, null).resolve("127.0.0.1", 42, null).role());
}
@Test
void loopbackTrustTreatsANonWorkerLoopbackCallerAsThePrimary() {
Principal p = new CallerResolver(nonWorkerIdentity()).resolve("127.0.0.1", 99, null);
assertEquals(Role.PRIMARY, p.role(), "the historical behaviour, now an explicit choice");
}
@Test
void tokenModeRefusesANonWorkerCallerThatPresentsNoToken() {
Principal p = new CallerResolver(nonWorkerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 99, null);
assertEquals(Role.ANONYMOUS, p.role(),
"no credential must mean NOTHING, not the most privileged role on the bus");
}
@Test
void tokenModeAcceptsAValidBearerTokenAsThePrimary() {
Principal p = new CallerResolver(nonWorkerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 99, "Bearer s3cret");
assertEquals(Role.PRIMARY, p.role());
}
@Test
void tokenModeRejectsAWrongOrMalformedCredential() {
CallerResolver r = new CallerResolver(nonWorkerIdentity(), true, "s3cret");
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "Bearer wrong").role());
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "s3cret").role(), "scheme required");
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "Bearer ").role(), "empty credential");
assertEquals(Role.ANONYMOUS, r.resolve("127.0.0.1", 99, "Basic s3cret").role(), "wrong scheme");
}
@Test
void theBearerSchemeIsCaseInsensitivePerRfc7235() {
CallerResolver r = new CallerResolver(nonWorkerIdentity(), true, "s3cret");
assertEquals(Role.PRIMARY, r.resolve("127.0.0.1", 99, "bearer s3cret").role());
assertEquals(Role.PRIMARY, r.resolve("127.0.0.1", 99, "BEARER s3cret").role());
}
@Test
void aNonLoopbackCallerIsNeverThePrimaryUnderLoopbackTrust() {
// Defence in depth: startup already refuses this pairing (validateAuthExposure), but if a
// proxy ever forwards a remote peer onto the loopback listener, the resolver must not
// hand it the primary role.
Principal p = new CallerResolver(nonWorkerIdentity()).resolve("10.0.0.7", 99, null);
assertEquals(Role.ANONYMOUS, p.role());
}
@Test
void tokenModeRequiresANonEmptyConfiguredToken() {
ConnectionIdentity id = nonWorkerIdentity();
assertThrows(IllegalArgumentException.class, () -> new CallerResolver(id, true, null));
assertThrows(IllegalArgumentException.class, () -> new CallerResolver(id, true, " "));
}
}
@@ -5,6 +5,7 @@ import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
@@ -46,10 +47,365 @@ class BridgedConfigTest {
assertTrue(cfg.guard().offSubscriptionHosts().isEmpty());
}
@Test
void singleWorkerBecomesAOneEntryProfileMapWithItselfAsDefault(@TempDir Path dir) throws Exception {
Path f = dir.resolve("single.yaml");
Files.writeString(f, """
worker:
profile: ltms-local
baseUrl: http://gx00.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(Set.of("ltms-local"), cfg.workerProfiles().keySet(), "legacy worker → one profile");
assertEquals("ltms-local", cfg.defaultProfile());
}
@Test
void loadsMultipleWorkerProfilesWithADefault(@TempDir Path dir) throws Exception {
Path f = dir.resolve("multi.yaml");
Files.writeString(f, """
workers:
gx10:
baseUrl: http://gx10.gw:8000
argv: ["ccs", "gx10"]
ollama:
baseUrl: http://ollama.ltms.dev
argv: ["ccs", "ollama"]
defaultWorker: gx10
guard:
offSubscriptionHosts: [gx10.gw, ollama.ltms.dev]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(Set.of("gx10", "ollama"), cfg.workerProfiles().keySet());
// Order, not just membership: placement breaks an exact-weight tie on definition order, so a
// hash-ordered map here would make equal-weight placement differ from one restart to the next.
assertEquals(java.util.List.of("gx10", "ollama"),
java.util.List.copyOf(cfg.workerProfiles().keySet()),
"workerProfiles must preserve YAML definition order");
assertEquals("gx10", cfg.defaultProfile());
assertEquals("ollama", cfg.workerProfiles().get("ollama").profile(), "profile defaults to its map key");
assertEquals("http://gx10.gw:8000", cfg.workerProfiles().get("gx10").baseUrl());
}
@Test
void ignoresUnknownKeys(@TempDir Path dir) throws Exception {
Path f = dir.resolve("future.yaml");
Files.writeString(f, "bind:\n port: 8080\nfutureFeature:\n enabled: true\n");
assertDoesNotThrow(() -> BridgedConfig.load(f));
}
@Test
void absentBrokerBlockLeavesInboxSoftState(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-broker.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNull(cfg.broker(), "no broker: block → null → in-memory inbox is selected");
}
@Test
void brokerBlockWithUriEnablesAmqpAdapter(@TempDir Path dir) throws Exception {
Path f = dir.resolve("broker.yaml");
Files.writeString(f, """
bind:
port: 8080
broker:
uri: amqp://guest:guest@127.0.0.1:5672/
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.broker());
assertTrue(cfg.broker().isConfigured(), "a non-blank uri enables the AMQP adapter");
assertEquals("amqp://guest:guest@127.0.0.1:5672/", cfg.broker().uri());
}
@Test
void brokerBlockWithBlankUriStaysSoftState(@TempDir Path dir) throws Exception {
Path f = dir.resolve("broker-blank.yaml");
Files.writeString(f, "bind:\n port: 8080\nbroker:\n uri: \"\"\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.broker());
assertFalse(cfg.broker().isConfigured(), "an empty uri must not enable AMQP");
}
@Test
void absentPrimaryBlockLeavesPrimaryNull(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-primary.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNull(cfg.primary(), "no primary: block → null → connection-derived identity");
}
@Test
void primaryBlockWithTerminalPinsIdentity(@TempDir Path dir) throws Exception {
Path f = dir.resolve("primary-pinned.yaml");
Files.writeString(f, """
bind:
port: 8080
primary:
terminal: term_fixed
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.primary());
assertEquals("term_fixed", cfg.primary().terminal());
}
@Test
void primaryBlockWithBlankTerminalDefaultsToDerived(@TempDir Path dir) throws Exception {
Path f = dir.resolve("primary-blank.yaml");
Files.writeString(f, "bind:\n port: 8080\nprimary:\n terminal: \"\"\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.primary());
assertTrue(cfg.primary().terminal() == null || cfg.primary().terminal().isBlank(),
"a blank terminal in yaml should be treated as absent — null or empty are equivalent");
}
// --- CB-402: peer kind discriminator -------------------------------------------------------
@Test
void workerKindDefaultsToClaudeCodeWhenOmitted(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-absent.yaml");
Files.writeString(f, """
workers:
gx10:
baseUrl: http://gx10.gw:8000
argv: ["ccs", "gx10"]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(BridgedConfig.Worker.KIND_CLAUDE_CODE, cfg.workerProfiles().get("gx10").kind(),
"a worker with no kind: is a claude-code worker (backward compatible)");
}
@Test
void opencodeKindIsNormalizedToLowerCase(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-opencode.yaml");
Files.writeString(f, """
workers:
gemini:
kind: OpenCode
model: google/gemini-2.5-pro
argv: ["opencode"]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(BridgedConfig.Worker.KIND_OPENCODE, cfg.workerProfiles().get("gemini").kind(),
"kind is normalised to lower-case so YAML casing does not matter");
}
@Test
void kindPredicatesReflectTheResolvedKind(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-predicates.yaml");
Files.writeString(f, """
workers:
claude:
baseUrl: http://gx10.gw:8000
gemini:
kind: opencode
model: google/gemini-2.5-pro
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Worker claude = cfg.workerProfiles().get("claude");
BridgedConfig.Worker gemini = cfg.workerProfiles().get("gemini");
assertTrue(claude.isClaudeCode(), "the default-kind worker is claude-code");
assertFalse(claude.isOpenCode(), "a claude-code worker is not opencode");
assertTrue(gemini.isOpenCode(), "the kind: opencode worker is opencode");
assertFalse(gemini.isClaudeCode(), "an opencode worker is not claude-code");
}
@Test
void argvDefaultsToTheKindBinaryWhenUnset(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-argv.yaml");
Files.writeString(f, """
workers:
claude:
baseUrl: http://gx10.gw:8000
gemini:
kind: opencode
model: google/gemini-2.5-pro
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(java.util.List.of("claude"), cfg.workerProfiles().get("claude").argv(),
"a claude-code worker with no argv defaults to the claude binary");
assertEquals(java.util.List.of("opencode"), cfg.workerProfiles().get("gemini").argv(),
"an opencode worker with no argv defaults to the opencode binary, never claude");
}
@Test
void authDefaultsToLoopbackTrustSoExistingConfigsBehaveAsBefore(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-auth-block.yaml");
Files.writeString(f, "bind:\n host: 127.0.0.1\n port: 8765\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.auth(), "auth must default rather than be null");
assertFalse(cfg.auth().tokenMode());
assertEquals("BRIDGED_API_TOKEN", cfg.auth().tokenEnv(), "documented default env var");
assertDoesNotThrow(cfg::validateAuthExposure, "loopback + loopback-trust is the safe pairing");
}
/**
* CB-501's highest-value check. Under loopback-trust, "not a known worker" means "the primary" —
* sound only while the OS refuses remote connections. Widening the bind without token mode
* would silently promote every reachable client to the most privileged role on the bus.
*/
@Test
void aNonLoopbackBindWithoutTokenModeIsRefusedAtStartup(@TempDir Path dir) throws Exception {
Path f = dir.resolve("exposed.yaml");
Files.writeString(f, "bind:\n host: 0.0.0.0\n port: 8765\n");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAuthExposure);
assertTrue(e.getMessage().contains("auth.mode: token"),
"the error must say how to fix it, not just that it refused");
}
@Test
void aNonLoopbackBindIsAllowedOnceTokenModeIsOn(@TempDir Path dir) throws Exception {
Path f = dir.resolve("exposed-with-token.yaml");
Files.writeString(f, """
bind:
host: 0.0.0.0
port: 8765
auth:
mode: token
tokenEnv: MY_TOKEN
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertTrue(cfg.auth().tokenMode());
assertEquals("MY_TOKEN", cfg.auth().tokenEnv());
assertDoesNotThrow(cfg::validateAuthExposure);
}
@Test
void loopbackFormsAreAllRecognised(@TempDir Path dir) throws Exception {
for (String host : new String[]{"127.0.0.1", "localhost", "::1", "127.0.0.53"}) {
Path f = dir.resolve("lb-" + host.replace(':', '_') + ".yaml");
Files.writeString(f, "bind:\n host: \"" + host + "\"\n port: 8765\n");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateAuthExposure(),
host + " is loopback and must not trip the exposure guard");
}
}
/**
* The shipped {@code bridged.example.yaml} must actually parse. Config binds through a plain
* Jackson mapper with {@code ignoreUnknown = true}, so a misspelled key in the example is
* silently dropped and the operator gets a default they did not ask for — exactly how a
* {@code spawn_ready_timeout_ms} typo survived in the example until the CB-5xx wrap-up.
*/
@Test
void shippedExampleConfigParses() {
Path example = Path.of("bridged.example.yaml");
assertTrue(Files.exists(example), "bridged.example.yaml must ship next to the pom");
BridgedConfig cfg = BridgedConfig.load(example);
assertEquals(8765, cfg.bind().port(), "example binds the documented default port");
assertTrue(cfg.workerProfiles().containsKey("gx10"), "example documents the gx10 profile");
assertEquals("gx10", cfg.defaultProfile(), "example's defaultWorker resolves");
assertTrue(cfg.guard().hostSet().contains("gx01.gw"),
"every example profile's base_url host must be in the example allowlist");
}
/**
* Every optional knob the example documents must bind under the exact spelling used there.
* Keep this list in step with {@code bridged.example.yaml}: a rename that updates the record
* but not the example (or vice versa) fails here instead of silently no-op'ing in production.
*/
@Test
void everyOptionalKnobDocumentedInTheExampleBinds(@TempDir Path dir) throws Exception {
Path f = dir.resolve("all-knobs.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
spawnReadyTimeoutMs: 25000
spawnReadyPollMs: 400
worktreeRoot: /tmp/bridged-worktrees
workers:
gx10:
kind: claude-code
baseUrl: http://gx01.gw:8000
configDir: /tmp/ccs/gx10
cwd: /tmp/repo
parityOverlay: [".mcp.json", ".env"]
gitTokenEnv: GITEA_TOKEN
gitHostEnv: GITEA_HOST
weight: 0.5
maxLoad: 2
placement: weighted
lifecycle:
idleTtlSeconds: 300
contextCap: 10
drainTimeoutSeconds: 5
broker:
uri: amqp://guest:guest@127.0.0.1:5672
primary:
terminal: term_abc123
pushReminders: 5
pushBackoffMs: 15000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(25000, cfg.spawnReadyTimeoutMs(), "spawnReadyTimeoutMs is camelCase, not snake_case");
assertEquals(400, cfg.spawnReadyPollMs(), "spawnReadyPollMs is camelCase, not snake_case");
assertEquals("/tmp/bridged-worktrees", cfg.worktreeRoot());
BridgedConfig.Worker w = cfg.workerProfiles().get("gx10");
assertEquals("/tmp/ccs/gx10", w.configDir());
assertEquals("/tmp/repo", w.cwd());
assertEquals(java.util.List.of(".mcp.json", ".env"), w.parityOverlay());
assertTrue(w.hasGitToken(), "gitTokenEnv binds and enables the CB-302 PR grant");
assertEquals("GITEA_HOST", w.gitHostEnv());
assertEquals(0.5f, w.weight(), 0.0001f, "weight binds as a float");
assertEquals(2, w.maxLoad(), "maxLoad binds as an integer");
assertEquals("weighted", cfg.placement(), "placement binds at the top level");
assertEquals(300, cfg.lifecycle().idleTtlSeconds());
assertEquals(10, cfg.lifecycle().contextCap());
assertEquals(5, cfg.lifecycle().drainTimeoutSeconds());
assertEquals("amqp://guest:guest@127.0.0.1:5672", cfg.broker().uri());
assertEquals("term_abc123", cfg.primary().terminal());
assertEquals(5, cfg.primary().remindersOrDefault());
assertEquals(15000L, cfg.primary().backoffMsOrDefault());
}
@Test
void placementDefaultsToFixedForExistingConfigs(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-placement.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
baseUrl: http://gx10.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals("fixed", cfg.placement(), "omitted placement must default to fixed");
}
@Test
void workerWeightAndMaxLoadDefaultSanely(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-weight.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
baseUrl: http://gx10.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Worker w = cfg.workerProfiles().get("gx10");
assertEquals(1.0f, w.weight(), 0.0001f, "absent weight defaults to 1.0");
assertNull(w.maxLoad(), "absent maxLoad defaults to unlimited (null)");
}
}
@@ -11,10 +11,11 @@ import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the {@code agent.*} south side against a REAL herdr, locking in
* the CB-102 spike findings. It spawns a HARMLESS probe command (never {@code claude},
* so no subscription/token involvement), proves the {@code env} map reaches the process
* environment, exercises status/read, and always tears the pane down.
* Contract test for the worker env seam against a REAL herdr. Under protocol 19 (CB-521) the
* env map is injected at PANE CREATION ({@code tab.create}), not {@code agent.start} — and
* {@code agent.start} now only launches supported agent kinds, so this probes the seed pane's
* SHELL directly (never {@code claude}, so no subscription/token involvement) and always tears
* the throwaway space down.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@@ -26,38 +27,30 @@ class AgentControlContractTest {
}
@Test
void startInjectsEnvThenReadAndClose() throws Exception {
void tabCreateInjectsEnvIntoTheSeedShell() throws Exception {
assumeTrue(!noSocket(), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
AgentControl agents = new AgentControl(herdr);
Agent probe = agents.start(
"__contract__",
List.of("bash", "-c", "printf 'PROBE_BASE=[%s]\\n' \"$ANTHROPIC_BASE_URL\"; sleep 20"),
WorkspaceControl spaces = new WorkspaceControl(herdr);
Workspace space = spaces.ensureWorkspace("__bridged_env_contract__");
Tab.Created tab = spaces.createTab(space.workspaceId(), null,
Map.of("ANTHROPIC_BASE_URL", "http://gx00.gw:8000"));
assertNotNull(probe.terminalId());
assertNotNull(probe.paneId());
try {
// Give the shell a moment to print, then confirm env reached the process.
assertNotNull(tab.rootPaneId(), "tab.create must return the seed pane");
Thread.sleep(1000); // let the seed shell reach its prompt
herdr.call("pane.send_input", Map.of(
"pane_id", tab.rootPaneId(),
"text", "printf 'PROBE_BASE=[%s]\\n' \"$ANTHROPIC_BASE_URL\"",
"keys", List.of("enter")));
Thread.sleep(800);
String visible = agents.read(probe.terminalId(), "visible");
String visible = herdr.call("pane.read",
Map.of("pane_id", tab.rootPaneId(), "source", "visible"))
.path("read").path("text").asText("");
assertTrue(visible.contains("PROBE_BASE=[http://gx00.gw:8000]"),
"env map must reach the process; saw: " + visible);
// Status is queryable; the probe appears in the agent list.
assertNotNull(agents.status(probe.terminalId()));
assertTrue(agents.list().stream()
.anyMatch(a -> probe.terminalId().equals(a.terminalId())),
"spawned probe should appear in agent.list");
"env map must reach the seed shell; saw: " + visible);
} finally {
agents.close(probe.paneId());
spaces.closeTab(tab.tab().tabId());
herdr.call("workspace.close", Map.of("workspace_id", space.workspaceId()));
}
// After close the pane is gone.
assertFalse(agents.list().stream()
.anyMatch(a -> probe.terminalId().equals(a.terminalId())),
"closed probe should no longer be listed");
}
}
}
@@ -0,0 +1,74 @@
package dev.ltms.bridged.herdr;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
/** Unit-level behaviour of {@link AgentControl} over a fake herdr (protocol 19). */
class AgentControlTest {
/** The {@code text} of every agent.prompt, in call order. */
@SuppressWarnings("unchecked")
private static List<String> promptTexts(FakeHerdr herdr) {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.toList();
}
@Test
void sendDeliversThePayloadAsOnePromptThatSubmitsItself() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_x", "do the thing");
// agent.prompt pastes AND submits in one call — no separate Enter event to assert.
assertEquals(List.of("do the thing"), promptTexts(herdr),
"exactly one agent.prompt carrying the payload");
}
@Test
void sendPreservesEmbeddedNewlinesVerbatim() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_x", "line1\nline2");
assertEquals(List.of("line1\nline2"), promptTexts(herdr),
"multiline content is delivered verbatim in the single prompt");
}
@Test
void submitNudgesWithAStandaloneEnterKeystroke() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).submit("term_x");
FakeHerdr.Call keys = herdr.lastCall("agent.send_keys");
assertEquals(Map.of("target", "term_x", "keys", List.of("enter")), keys.params(),
"the raced-Enter nudge is a raw send_keys, not a second prompt");
}
@Test
@SuppressWarnings("unchecked")
void aTerminalIdTargetIsTranslatedToItsPaneId() {
// Protocol 19 rejects terminal_id as an agent.* target; the fake's agent.list maps
// term_a to pane w2:p7, and the control layer must address herdr by that pane.
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_a", "hello");
Map<String, Object> prompt = (Map<String, Object>) herdr.lastCall("agent.prompt").params();
assertEquals("w2:p7", prompt.get("target"), "terminal target resolved to the agent's pane id");
}
@Test
@SuppressWarnings("unchecked")
void theTerminalToPaneMappingIsCachedAcrossCalls() {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
agents.send("term_a", "one");
agents.send("term_a", "two");
long lists = herdr.calls.stream().filter(c -> c.method().equals("agent.list")).count();
assertEquals(1, lists, "one agent.list resolution serves every later call to the same terminal");
}
}
@@ -0,0 +1,43 @@
package dev.ltms.bridged.herdr;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Wire mapping and injectability of {@link AgentStatus}, including the CB-115 {@code done} state. */
class AgentStatusTest {
@Test
void mapsTheKnownWireStrings() {
assertEquals(AgentStatus.IDLE, AgentStatus.fromWire("idle"));
assertEquals(AgentStatus.WORKING, AgentStatus.fromWire("working"));
assertEquals(AgentStatus.BLOCKED, AgentStatus.fromWire("blocked"));
assertEquals(AgentStatus.DONE, AgentStatus.fromWire("done"));
}
@Test
void mapsDoneCaseInsensitively() {
assertEquals(AgentStatus.DONE, AgentStatus.fromWire("DONE"));
assertEquals(AgentStatus.DONE, AgentStatus.fromWire("Done"));
}
@Test
void unknownAndNullFallToUnknown() {
assertEquals(AgentStatus.UNKNOWN, AgentStatus.fromWire("unknown"));
assertEquals(AgentStatus.UNKNOWN, AgentStatus.fromWire("something-else"));
assertEquals(AgentStatus.UNKNOWN, AgentStatus.fromWire(null));
}
@Test
void doneIsInjectableLikeIdle() {
// The whole point of CB-115: a finished worker herdr reports as `done` must be deliverable,
// not treated as UNKNOWN (which wedged delivery and mis-fired the stall failure).
assertTrue(AgentStatus.DONE.injectable());
assertTrue(AgentStatus.IDLE.injectable());
assertTrue(AgentStatus.BLOCKED.injectable());
assertFalse(AgentStatus.WORKING.injectable());
assertFalse(AgentStatus.UNKNOWN.injectable());
}
}
@@ -8,23 +8,32 @@ import java.util.List;
/**
* Recording fake {@link HerdrClient} for unit/acceptance tests. Returns canned frames
* captured from the real herdr 0.7.0 daemon and records every call so tests can assert
* both behaviour and that guard-blocked paths never reached herdr.
* matching the real herdr 0.8.0 daemon (protocol 19) and records every call so tests can
* assert both behaviour and that guard-blocked paths never reached herdr.
*/
public final class FakeHerdr implements HerdrClient {
public record Call(String method, Object params) {
}
/** The foreground PID of the one agent pane (term_a) in the canned {@code pane.process_info}. */
public static final long WORKER_PID = 4242;
private final ObjectMapper mapper = new ObjectMapper();
public final List<Call> calls = new ArrayList<>();
private boolean healthy = true;
private final List<String> extraWorkspaces = new ArrayList<>();
private final List<String> extraAgents = new ArrayList<>();
private int agentNameTakenFor = 0;
private int agentPaneBusyFor = 0;
private int workerTabPaneCount = 1;
private String paneCloseErrorCode = null;
private String agentSendErrorCode = null;
private volatile String agentStatus = "idle"; // what agent.get reports
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String readText = "worker transcript tail"; // canned agent.read output
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
private String pinnedStartTerminal;
private String pinnedStartPane;
public FakeHerdr healthy(boolean h) {
this.healthy = h;
@@ -37,6 +46,12 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/** Reject the first {@code n} {@code agent.start} calls with {@code agent_pane_busy}. */
public FakeHerdr agentPaneBusyTimes(int n) {
this.agentPaneBusyFor = n;
return this;
}
/** Make the worker tab (w9:t2) report this many panes in {@code tab.list} (default 1). */
public FakeHerdr withWorkerTabPaneCount(int n) {
this.workerTabPaneCount = n;
@@ -55,12 +70,45 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/** Make {@code agent.send} fail with this herdr error code. */
/** The text {@code agent.read} returns (the CB-106 completion scrape). */
public FakeHerdr readText(String text) {
this.readText = text;
return this;
}
/** Make delivery ({@code agent.prompt} / {@code agent.send_keys}) fail with this error code. */
public FakeHerdr agentSendFailsWith(String code) {
this.agentSendErrorCode = code;
return this;
}
/**
* Force the next {@code n} {@code agent.start} calls to report this terminal/pane coordinate,
* instead of the fake's usual incrementing {@code term_new_n}/{@code w9:pRoot_n}. Lets a test make
* two spawns report the <em>same</em> herdr pane, to prove the host-unique id (CB-519) never
* collides on that coordinate.
*/
public FakeHerdr pinNextStarts(int n, String terminalId, String paneId) {
this.pinnedStarts = n;
this.pinnedStartTerminal = terminalId;
this.pinnedStartPane = paneId;
return this;
}
/**
* Seed a named agent into {@code agent.list} (e.g. an orphaned worker for CB-117 reaper tests).
* The {@code name} carries the worker label the reaper keys on; {@code paneId}/{@code tabId}
* locate its pane for teardown.
*/
public FakeHerdr withAgent(String name, String terminalId, String paneId, String tabId) {
extraAgents.add(("{\"terminal_id\":\"%s\",\"agent\":\"claude\",\"agent_status\":\"idle\","
+ "\"name\":\"%s\",\"agent_session\":{\"kind\":\"id\",\"value\":\"sess-%s\"},"
+ "\"workspace_id\":\"wQ\",\"tab_id\":\"%s\",\"pane_id\":\"%s\"}")
.formatted(terminalId, name, terminalId, tabId, paneId));
return this;
}
/** Seed an additional workspace into {@code workspace.list} (e.g. a pre-existing worker space). */
public FakeHerdr withWorkspace(String id, String label) {
extraWorkspaces.add(("{\"workspace_id\":\"%s\",\"label\":\"%s\",\"focused\":false,"
@@ -85,20 +133,31 @@ public final class FakeHerdr implements HerdrClient {
try {
return switch (method) {
case "ping" -> mapper.readTree(
"{\"type\":\"pong\",\"version\":\"0.7.0\",\"protocol\":14}");
"{\"type\":\"pong\",\"version\":\"0.8.0\",\"protocol\":19}");
case "workspace.list" -> mapper.readTree(("""
{"type":"workspace_list","workspaces":[
{"workspace_id":"w1","label":"dev-mgnl","focused":true,"pane_count":7,"agent_status":"unknown"},
{"workspace_id":"w2","label":"ltms","focused":false,"pane_count":5,"agent_status":"done"}%s]}""")
.formatted(extraWorkspaces.isEmpty() ? "" : "," + String.join(",", extraWorkspaces)));
case "agent.list" -> mapper.readTree("""
case "agent.list" -> mapper.readTree(("""
{"type":"agent_list","agents":[
{"terminal_id":"term_a","agent":"claude","agent_status":"idle",
"agent_session":{"kind":"id","value":"sess-1111"},
"workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}]}""");
case "agent.send" -> {
"workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}%s]}""")
.formatted(extraAgents.isEmpty() ? "" : "," + String.join(",", extraAgents)));
case "agent.prompt" -> {
if (agentSendErrorCode != null) {
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.send failed",
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.prompt failed",
agentSendErrorCode, null);
}
yield mapper.readTree(("""
{"type":"agent_prompted","agent":{"terminal_id":"term_a","agent":"claude",
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentStatus));
}
case "agent.send_keys" -> {
if (agentSendErrorCode != null) {
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.send_keys failed",
agentSendErrorCode, null);
}
yield mapper.readTree("{\"type\":\"ok\"}");
@@ -107,28 +166,65 @@ public final class FakeHerdr implements HerdrClient {
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":"claude",
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentStatus));
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
case "agent.start" -> {
// Protocol 19: kind and pane_id are required — reject like the real daemon.
java.util.Map<?, ?> p = params instanceof java.util.Map<?, ?> m ? m : java.util.Map.of();
for (String required : new String[]{"kind", "pane_id"}) {
if (p.get(required) == null) {
throw new HerdrException(
"herdr error [invalid_request]: invalid request: missing field `"
+ required + "`", "invalid_request", null);
}
}
long starts = calls.stream().filter(c -> c.method().equals("agent.start")).count();
if (starts <= agentNameTakenFor) {
if (starts <= agentPaneBusyFor) {
throw new HerdrException(
"herdr error [agent_pane_busy]: agent target pane is not an available shell",
"agent_pane_busy", null);
}
long busyAdjusted = starts - agentPaneBusyFor;
if (busyAdjusted <= agentNameTakenFor) {
throw new HerdrException(
"herdr error [agent_name_taken]: agent name already used",
"agent_name_taken", null);
}
yield mapper.readTree("""
long n = busyAdjusted - agentNameTakenFor;
boolean pinned = pinnedStarts > 0;
if (pinned) {
pinnedStarts--;
}
// Protocol 19: the agent starts INTO the requested pane, so its pane_id normally
// echoes the param. A pin overrides both coordinates, which is the only way to
// make two spawns report one pane — what CB-519's collision test needs.
String terminal = pinned ? pinnedStartTerminal : ("term_new_" + n);
Object pane = pinned ? pinnedStartPane : p.get("pane_id");
yield mapper.readTree(("""
{"type":"agent_started","agent":{
"terminal_id":"term_new","name":"claude","agent_status":"unknown",
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"w9:pW"}}""");
"terminal_id":"%s","name":"claude","agent_status":"unknown",
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"%s"}}""")
.formatted(terminal, pane));
}
case "pane.split" -> mapper.readTree("""
{"type":"pane_info","pane":{"pane_id":"w1:pSplit","workspace_id":"w1",
"tab_id":"w1:t1"}}""");
case "workspace.create" -> mapper.readTree("""
{"type":"workspace_created",
"workspace":{"workspace_id":"w9","label":"bridged-workers","focused":false,
"pane_count":1,"tab_count":1,"active_tab_id":"w9:t1","agent_status":"unknown"},
"tab":{"tab_id":"w9:t1","workspace_id":"w9","label":"1","pane_count":1},
"root_pane":{"pane_id":"w9:p1","workspace_id":"w9","tab_id":"w9:t1"}}""");
case "tab.create" -> mapper.readTree("""
case "tab.create" -> {
// Each tab gets its own seed pane — under protocol 19 that pane becomes the
// worker pane, so distinct spawns must yield distinct pane ids.
long tabs = calls.stream().filter(c -> c.method().equals("tab.create")).count();
yield mapper.readTree(("""
{"type":"tab_created",
"tab":{"tab_id":"w9:t2","workspace_id":"w9","label":"2","pane_count":1},
"root_pane":{"pane_id":"w9:pRoot","workspace_id":"w9","tab_id":"w9:t2"}}""");
"root_pane":{"pane_id":"w9:pRoot_%d","workspace_id":"w9","tab_id":"w9:t2"}}""")
.formatted(tabs));
}
case "tab.rename" -> mapper.readTree("""
{"type":"tab_info","tab":{"tab_id":"w9:t2","workspace_id":"w9",
"label":"worker: ltms-local","pane_count":1}}""");
@@ -141,6 +237,21 @@ public final class FakeHerdr implements HerdrClient {
case "pane.get" -> mapper.readTree("""
{"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9",
"tab_id":"w9:t2","agent_status":"idle"}}""");
case "pane.list" -> mapper.readTree("""
{"type":"pane_list","panes":[
{"pane_id":"w2:p7","terminal_id":"term_a","workspace_id":"w2","tab_id":"w2:t7","agent":"claude"},
{"pane_id":"w2:p9","terminal_id":"term_shell","workspace_id":"w2","tab_id":"w2:t8"}]}""");
case "pane.process_info" -> {
Object paneId = params instanceof java.util.Map<?, ?> m ? m.get("pane_id") : null;
yield "w2:p7".equals(paneId)
? mapper.readTree(("""
{"type":"pane_process_info","process_info":{"pane_id":"w2:p7","shell_pid":%d,
"foreground_processes":[{"pid":%d,"name":"node","argv0":"claude"}]}}""")
.formatted(WORKER_PID, WORKER_PID))
: mapper.readTree("""
{"type":"pane_process_info","process_info":{"pane_id":"w2:p9","shell_pid":9001,
"foreground_processes":[]}}""");
}
case "pane.close" -> {
if (paneCloseErrorCode != null) {
throw new HerdrException("herdr error [" + paneCloseErrorCode + "]: pane.close failed",
@@ -0,0 +1,48 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import java.nio.file.Files;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for {@link PaneLocator} against a REAL herdr: the PID→pane mapping that
* connection identity rests on. Uses a throwaway tab's seed shell as the probe process
* (protocol 19 removed arbitrary-command agents), and always tears the space down.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@Tag("contract")
class PaneLocatorContractTest {
@Test
void resolvesTheTerminalOwningARealProcessPid() throws Exception {
assumeTrue(Files.exists(UnixSocketHerdrClient.defaultSocketPath()), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
WorkspaceControl spaces = new WorkspaceControl(herdr);
Workspace space = spaces.ensureWorkspace("__bridged_pid_contract__");
Tab.Created tab = spaces.createTab(space.workspaceId(), null, Map.of());
try {
JsonNode info = herdr.call("pane.process_info", Map.of("pane_id", tab.rootPaneId()))
.path("process_info");
long shellPid = info.path("shell_pid").asLong(-1);
assertTrue(shellPid > 0, "seed pane should report a shell pid");
String terminalId = herdr.call("pane.get", Map.of("pane_id", tab.rootPaneId()))
.path("pane").path("terminal_id").asText(null);
assertNotNull(terminalId, "seed pane should carry a terminal_id");
assertEquals(terminalId, new PaneLocator(herdr).terminalForPid(shellPid),
"a real PID must resolve back to its own pane's terminal_id");
} finally {
spaces.closeTab(tab.tab().tabId());
herdr.call("workspace.close", Map.of("workspace_id", space.workspaceId()));
}
}
}
}
@@ -0,0 +1,27 @@
package dev.ltms.bridged.herdr;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/** Unit tests for PID → pane resolution (the herdr half of connection-based MCP identity). */
class PaneLocatorTest {
private final PaneLocator loc = new PaneLocator(new FakeHerdr());
@Test
void resolvesTerminalForAForegroundPid() {
assertEquals("term_a", loc.terminalForPid(FakeHerdr.WORKER_PID));
}
@Test
void nullForAPidInNoPane() {
assertNull(loc.terminalForPid(999_999));
}
@Test
void nullForNonPositivePid() {
assertNull(loc.terminalForPid(0));
assertNull(loc.terminalForPid(-1));
}
}
@@ -5,7 +5,6 @@ import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import java.nio.file.Files;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
@@ -13,9 +12,9 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the placement layer ({@code workspace.*}/{@code tab.*}) against a
* REAL herdr, locking in the "one clean tab per worker" recipe: find-or-create a worker
* space, give the worker its own tab, drop herdr's seed shell so the tab holds only the
* worker, and tear it all down. Uses a HARMLESS probe (never {@code claude}) in a
* REAL herdr, locking in the "one clean tab per worker" recipe under protocol 19: find-or-create
* a worker space, give the worker its own tab, and the SEED pane is where the worker starts —
* the tab holds exactly that one pane from creation. Uses no agent (never {@code claude}) in a
* throwaway space that is fully removed at the end.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
@@ -41,7 +40,6 @@ class WorkspacePlacementContractTest {
void workerGetsOwnCleanTabAndTearsDownCompletely() throws Exception {
assumeTrue(!noSocket(), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
Workspace space = spaces.ensureWorkspace(LABEL);
@@ -49,36 +47,27 @@ class WorkspacePlacementContractTest {
// Idempotent: a second ensure finds the same space, never creates a duplicate.
assertEquals(space.workspaceId(), spaces.ensureWorkspace(LABEL).workspaceId());
Tab.Created tab = spaces.createTab(space.workspaceId());
Agent worker = agents.start(
"__contract__",
List.of("bash", "-c", "sleep 20"),
Map.of(),
tab.tab().tabId());
Tab.Created tab = spaces.createTab(space.workspaceId(), null, Map.of());
try {
// The worker landed in its dedicated tab in the worker space.
assertEquals(tab.tab().tabId(), worker.tabId());
assertEquals(space.workspaceId(), worker.workspaceId());
assertNotNull(tab.rootPaneId(), "tab.create must return the seed pane");
// Drop the seed shell; the tab now holds exactly the worker pane.
agents.close(tab.rootPaneId());
// Protocol 19: the seed pane IS the worker pane — the tab holds exactly it.
spaces.renameTab(tab.tab().tabId(), "worker: contract");
assertEquals(1, paneCount(herdr, space.workspaceId(), tab.tab().tabId()),
"worker tab must hold only the worker pane after the seed shell is dropped");
"worker tab must hold exactly the seed/worker pane");
// Teardown resolves the tab from the pane, and sees it holds exactly one pane.
WorkspaceControl.PaneLocation loc = spaces.locatePane(worker.paneId());
WorkspaceControl.PaneLocation loc = spaces.locatePane(tab.rootPaneId());
assertNotNull(loc);
assertEquals(tab.tab().tabId(), loc.tabId());
assertEquals(1, loc.tabPaneCount(), "worker is the tab's sole occupant");
assertEquals(1, loc.tabPaneCount(), "worker pane is the tab's sole occupant");
} finally {
agents.close(worker.paneId());
spaces.closeTab(tab.tab().tabId());
}
// Tolerant teardown: closing an already-gone tab / reading a gone pane is a no-op.
spaces.closeTab(tab.tab().tabId());
assertNull(spaces.locatePane(worker.paneId()), "closed worker pane must be gone");
assertNull(spaces.locatePane(tab.rootPaneId()), "closed worker pane must be gone");
// Remove the throwaway space entirely so the test leaves no residue.
herdr.call("workspace.close", Map.of("workspace_id", space.workspaceId()));
@@ -0,0 +1,287 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.msg.Rendezvous;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Unit behaviour of the CB-106 completion resolver in isolation from the injector. */
class CompletionResolverTest {
@Test
void skipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
resolver.resolve("term_a", null); // no in-flight turn captured for this target
assertFalse(herdr.called("agent.read"),
"a turn nobody is blocked on must not cost a transcript scrape");
}
@Test
void failSkipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
resolver.fail("term_a", null); // no in-flight turn, and no registered waiter to fall back to
assertFalse(herdr.called("agent.read"),
"a wedge nobody is blocked on must not cost a transcript scrape");
}
@Test
void captureBaselineSkipsTheReadWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ X\n❯ ");
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
resolver.captureBaseline("term_a"); // no send to attribute a later completion to
assertFalse(herdr.called("agent.read"),
"with no waiting send there is no turn to baseline — skip the scrape");
}
// --- CB-115 clean scrape: extract the last assistant block ----------------
@Test
void extractsTheLastAssistantBlockStrippingChrome() {
String raw = """
⏺ Reading the file…
⏺ Done. The bug was an off-by-one in the loop bound.
╭──────────────────────────────────────╮
│ > │
╰──────────────────────────────────────╯
⏵⏵ auto mode on · ? for shortcuts
""";
assertEquals("Done. The bug was an off-by-one in the loop bound.",
CompletionResolver.lastAssistantBlock(raw));
}
@Test
void keepsMultiLineAssistantContent() {
String raw = "⏺ Line one.\nLine two.\n❯ ";
assertEquals("Line one.\nLine two.", CompletionResolver.lastAssistantBlock(raw));
}
@Test
void fallsBackToRawTextWhenThereIsNoMarker() {
String raw = "plain worker output with no glyph";
assertEquals("plain worker output with no glyph", CompletionResolver.lastAssistantBlock(raw));
}
@Test
void blankScrapeYieldsEmpty() {
assertTrue(CompletionResolver.lastAssistantBlock("").isEmpty());
assertTrue(CompletionResolver.lastAssistantBlock(null).isEmpty());
}
@Test
void stripsSpinnerAndRuleChrome() {
String raw = """
⏺ Channel check confirmed — your message got through.
✻ Brewed for 11s
─────────────────────────────────────
""";
assertEquals("Channel check confirmed — your message got through.",
CompletionResolver.lastAssistantBlock(raw));
}
@Test
void cutsANextTurnPromptEchoAndTrailingTipsFromTheBlock() {
// The exact turn-2 leak: the scrape captured the settled answer, then a "✻ Cooked" spinner,
// then the NEXT turn's echoed prompt, then a "✶ Forming…" spinner and trailing tips/warnings
// whose lines (⎿, ⚠) are not themselves chrome-terminated. Stopping at the first boundary
// (the ✻ spinner) is what keeps every one of those interface lines out of the reply.
String raw = """
⏺ Channel confirmed — the bridge reply delivered successfully.
✻ Cooked for 9s
❯ Thanks. Now a small task: what is 17 * 23? Show just the number.
✶ Forming…
⎿ Tip: Name your conversations with /rename
⚠ claude.ai connectors are disabled because ANTHROPIC_API_KEY is set
""";
assertEquals("Channel confirmed — the bridge reply delivered successfully.",
CompletionResolver.lastAssistantBlock(raw));
}
// --- CB-115 misattribution guard: suppress a stale (unchanged) completion -------
@Test
void suppressesACompletionWhoseScrapeIsUnchangedFromDelivery() {
// Rapid back-to-back turn: the pane still shows the PREVIOUS turn's answer when this turn's
// (misattributed) completion boundary fires. The scrape == the delivery baseline, so the
// send must NOT be resolved with the stale answer.
FakeHerdr herdr = new FakeHerdr().readText("⏺ 391\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
// The turn as captured at delivery: its waiter, and the previous turn's answer still on screen.
var turn = new CompletionResolver.InFlight(waiter, "391");
resolver.resolve("term_a", turn); // scrape still "391" == baseline → suppress
assertFalse(waiter.isDone(), "a completion with no output change must not resolve the send");
assertTrue(rendezvous.isWaiting("term_a"), "the send stays waiting for a real reply");
}
@Test
void resolvesACompletionWhoseScrapeChangedSinceDelivery() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ No, 391 = 17 × 23.\n❯ "); // the worker's real answer
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
// Delivery baseline was the previous turn's "391"; the scrape now differs → resolve.
var turn = new CompletionResolver.InFlight(waiter, "391");
resolver.resolve("term_a", turn);
assertTrue(waiter.isDone(), "a completion with new output must resolve the send");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
assertEquals("No, 391 = 17 × 23.", waiter.getNow(null).text());
}
@Test
void suppressesAnUnchangedCompletionEvenWhenTheBlockExceedsTheScrapeCap() {
// The fan-out issue-hunt finding: captureBaseline once stored the RAW (unclipped) assistant
// block while resolve compares against a clip()'d tail. For a block longer than MAX_SCRAPE_CHARS
// the two capped representations differ even when the pane never changed, so the CB-115
// byte-identical guard failed to fire and a stale completion could resolve the send. Both sides
// must clip identically; here an unchanged >cap block on rapid back-to-back turns stays suppressed.
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
resolver.captureBaseline("term_a"); // baseline is the clipped >cap block
var turn = resolver.inFlight("term_a");
assertEquals(CompletionResolver.MAX_SCRAPE_CHARS, turn.baseline().length(),
"the delivery baseline is clipped to the same cap resolve() applies to the tail");
resolver.resolve("term_a", turn); // scrape unchanged → clipped tail == baseline → suppress
assertFalse(waiter.isDone(),
"an unchanged >cap block must still be recognised as stale and suppressed");
assertTrue(rendezvous.isWaiting("term_a"), "the send stays waiting for a real reply");
}
@Test
void resolvesWhenThereIsNoBaseline() {
// No delivery baseline (e.g. the pre-turn read failed) ⇒ never suppress; the completion resolves.
FakeHerdr herdr = new FakeHerdr().readText("⏺ hello\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "with no baseline a completion resolves as before");
assertEquals("hello", waiter.getNow(null).text());
}
@Test
void resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent() {
// The most important branch of the CB-115 guard: a failed read means the resolver could not
// SEE the screen — "couldn't see", not "no change". It must still resolve the send (an empty
// tail beats hanging until the caller's timeout), even though a baseline was captured. The
// baseline here is "" (an empty pane at delivery), so without the !scrapeFailed clause the
// byte-identical guard would wrongly match the empty tail and suppress.
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.read throws HerdrException
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, ""); // empty pane baselined at delivery
resolver.resolve("term_a", turn);
assertTrue(waiter.isDone(),
"a failed scrape must still resolve the send, not hang until the caller's timeout");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
assertEquals("", waiter.getNow(null).text(), "the tail is empty because the screen was unreadable");
}
// --- CB-115/CB-116 fail guard: an already-done or absent waiter is left alone ---------
@Test
void failLeavesAnAlreadyResolvedWaiterUntouchedAndSkipsTheScrape() {
// The send was already resolved (e.g. by the worker's explicit reply) before fail fired.
// fail must not overwrite that value, and must not even scrape the worker — nobody needs it.
FakeHerdr herdr = new FakeHerdr().readText("an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null);
assertTrue(rendezvous.resolveCompletion(waiter, "already replied"));
resolver.fail("term_a", turn);
assertFalse(herdr.called("agent.read"),
"fail must not scrape a waiter that is already done");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"fail must not overwrite the existing resolution");
assertEquals("already replied", waiter.getNow(null).text());
}
@Test
void failFallsBackToTheRegisteredWaiterWhenThereIsNoInFlightTurn() {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fail falls back to the waiter currently registered on the Rendezvous and fails it.
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a"); // send registered, but no captureBaseline ever ran
resolver.fail("term_a", null); // no in-flight turn → fall back to the registered waiter
assertTrue(waiter.isDone(), "fail falls back to the registered waiter when no turn is in flight");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
assertEquals("stuck on an error screen", waiter.getNow(null).text());
}
// --- CB-116 waiter identity: a late completion never crosses into the next turn ---------
@Test
void aLateCompletionForOneTurnNeverResolvesTheNextTurnsWaiter() {
// The cross-turn stale reply the conversation test surfaced: turn N's completion fallback
// fires AFTER turn N was resolved by an explicit bridge_reply and turn N+1 has opened its own
// waiter on the same session. Resolving "whatever is waiting now" would hand turn N's stale
// scrape to turn N+1; targeting turn N's captured waiter makes the late completion a no-op.
FakeHerdr herdr = new FakeHerdr().readText("⏺ turn N answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiterN = rendezvous.open("term_a"); // turn N's send
// The turn as the injector captured it at delivery (waiter + pre-turn baseline).
var turnN = new CompletionResolver.InFlight(waiterN, "an earlier answer");
// Turn N is resolved by the worker's explicit reply.
assertTrue(rendezvous.resolve("term_a", "N replied"));
// Turn N+1's send opens its own waiter on the same session (replacing the registered one).
var waiterN1 = rendezvous.open("term_a");
resolver.resolve("term_a", turnN); // turn N's completion fallback finally fires
assertFalse(waiterN1.isDone(), "turn N's late completion must not resolve turn N+1's waiter");
assertEquals(Rendezvous.Kind.REPLY, waiterN.getNow(null).kind(),
"turn N stays resolved by its own reply");
assertTrue(rendezvous.isWaiting("term_a"), "turn N+1 is still awaiting its own resolution");
}
}
@@ -6,8 +6,10 @@ import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.TimeUnit;
@@ -25,11 +27,15 @@ class InjectorTest {
private final FakeHerdr herdr = new FakeHerdr();
private final Injector injector = new Injector(new AgentControl(herdr));
/** Text of every agent.send, in order. */
/**
* The logical messages delivered, in order. Under protocol 19 each delivery is one
* {@code agent.prompt} carrying the payload (it submits itself); the Enter nudge is a
* separate {@code agent.send_keys} and never appears here.
*/
@SuppressWarnings("unchecked")
private List<String> sent() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.toList();
}
@@ -43,6 +49,45 @@ class InjectorTest {
assertEquals(List.of("hello"), sent());
}
@Test
void holdsDeliveryUntilTheWorkerIsAvailable() {
// CB-113: idle alone is not enough — hold until the worker's MCP is connected (ready).
java.util.Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // idle but not yet available → held out of the boot window
assertEquals(List.of(), sent(), "must not deliver into a not-yet-available worker");
ready.add(T); // the worker's Claude connects the bridge MCP
inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("task"), sent(), "delivers once the worker is available");
}
private long enterKeystrokes() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send_keys"))
.count();
}
@Test
void resubmitsEnterWhenADeliveredMessageIsNotPickedUp() {
// CB-113: the Enter at delivery can race the paste; while the worker stays idle (not picked
// up), the injector re-nudges Enter so the pending paste submits.
injector.enqueue(T, "task");
injector.onStatus(T, AgentStatus.IDLE); // deliver: paste + one Enter
long afterDeliver = enterKeystrokes();
injector.onStatus(T, AgentStatus.IDLE); // still idle → re-nudge Enter
injector.onStatus(T, AgentStatus.IDLE); // and again
assertTrue(enterKeystrokes() > afterDeliver, "an unpicked-up delivery re-nudges Enter");
injector.onStatus(T, AgentStatus.WORKING); // worker finally starts
long atPickup = enterKeystrokes();
injector.onStatus(T, AgentStatus.WORKING);
assertEquals(atPickup, enterKeystrokes(), "no more nudges once the worker has picked up");
}
@Test
void holdsWhileWorkingThenDeliversOnIdle() {
injector.enqueue(T, "later");
@@ -120,17 +165,97 @@ class InjectorTest {
}
@Test
void activeWhileQueuedOrInFlightThenQuietAfterPickup() {
void activeWhileQueuedOrInFlightThenQuietAfterTurnCompletes() {
assertTrue(injector.activeTargets().isEmpty());
injector.enqueue(T, "x");
assertEquals(java.util.Set.of(T), injector.activeTargets(), "active while a message is queued");
assertEquals(Set.of(T), injector.activeTargets(), "active while a message is queued");
injector.onStatus(T, AgentStatus.IDLE); // delivers; still in-flight (awaiting pickup)
assertEquals(java.util.Set.of(T), injector.activeTargets(),
injector.onStatus(T, AgentStatus.IDLE); // delivers; awaiting pickup
assertEquals(Set.of(T), injector.activeTargets(),
"stays active so the poller can observe the worker pick the message up");
injector.onStatus(T, AgentStatus.WORKING); // pickup observed → in-flight cleared
assertTrue(injector.activeTargets().isEmpty(), "quiet once queue is empty and pickup is seen");
injector.onStatus(T, AgentStatus.WORKING); // pickup observed; now awaiting turn completion
assertEquals(Set.of(T), injector.activeTargets(),
"stays active after pickup so the working→idle completion boundary is observed");
injector.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete
assertTrue(injector.activeTargets().isEmpty(), "quiet once the delegated turn has completed");
}
@Test
void firesTurnCompleteOnAConfirmedWorkingThenIdle() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // pickup + turn running
assertEquals(List.of(), completed, "no completion until the turn returns to idle");
inj.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete
assertEquals(List.of(T), completed, "a confirmed working→idle fires exactly one completion");
}
@Test
void doesNotSynthesizeCompletionFromAnUnconfirmedTurn() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
// Deliver, then only ever idle — a `working` sample is never seen. The pickup grace unwedges
// the queue but must NOT invent a completion: without a sampled turn there is no trustworthy
// "the worker finished the task" signal, so the send should fall through to its timeout.
for (int i = 0; i < 15; i++) inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of(), completed, "no completion is synthesized from an unconfirmed turn");
}
/** Captures both turn-lifecycle callbacks so the CB-109 stall path can be asserted. */
private static final class Captor implements TurnListener {
final List<String> completed = new ArrayList<>();
final List<String> failed = new ArrayList<>();
@Override
public void onTurnComplete(String target) {
completed.add(target);
}
@Override
public void onTurnFailed(String target) {
failed.add(target);
}
}
// ~30s of unknown at the 250ms prod poll interval; enough onStatus samples to trip the stall.
private static final int STALL_SAMPLES = 130;
@Test
void failsAnOutstandingDelegationWhoseWorkerWedgesInUnknown() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // worker starts the turn
for (int i = 0; i < STALL_SAMPLES; i++) inj.onStatus(T, AgentStatus.UNKNOWN); // then wedges
assertEquals(List.of(T), cap.failed, "a sustained unknown streak fails the outstanding send");
assertEquals(List.of(), cap.completed, "a wedge is a failure, not a completion");
assertTrue(inj.activeTargets().isEmpty(), "the wedged target is reclaimed, not polled forever");
}
@Test
void aTransientUnknownGlitchNeitherFailsNorBlocksCompletion() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // confirmed turn
for (int i = 0; i < 10; i++) inj.onStatus(T, AgentStatus.UNKNOWN); // brief glitch, well under grace
inj.onStatus(T, AgentStatus.IDLE); // working → idle: the real completion
assertEquals(List.of(), cap.failed, "a short unknown blip must not fail the turn");
assertEquals(List.of(T), cap.completed, "the streak reset, so the turn still completes");
}
@Test
@@ -151,6 +276,86 @@ class InjectorTest {
assertTrue(f.isCompletedExceptionally(), "queued waiters unblock when the worker vanishes");
}
@Test
void dropFailsTheTurnOfADeliveredMessageWhenTheWorkerVanishes() {
// CB-110: the message was delivered (no longer queued), so failing queued waiters alone would
// leave its send hanging. A vanished worker must fail that in-flight turn too.
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // turn running
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), cap.failed, "a worker that vanishes mid-turn fails its in-flight send");
}
@Test
void dropFailsADeliveredTurnThatVanishesBeforePickupIsConfirmed() {
// Delivered but no WORKING sampled yet (awaitingCompletion=true, awaitingPickup still true,
// turnObserved=false) — a distinct state the other two drop tests don't cover. (Gap surfaced
// by an off-sub worker's review of CB-110, delegated through the bridge.)
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver; pickup never confirmed
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), cap.failed, "a delivery that vanishes before pickup still fails its send");
}
// ~60s of idle-but-not-ready at the 250ms prod poll interval; enough to trip the readiness grace.
private static final int READINESS_SAMPLES = 245;
@Test
void failsAQueuedMessageWhoseWorkerNeverBecomesReady() {
// CB-114: herdr keeps reporting the worker idle, but its Claude never connects the bridge MCP,
// so the readiness gate never opens. The message must not be held (and the target polled)
// forever — after the grace it fails, the caller unblocks via the worker-failure path, the
// target is reclaimed, and the never-set presence is cleared.
Captor cap = new Captor();
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), cap, _ -> false, forgotten::add);
CompletableFuture<Void> f = inj.enqueue(T, "task");
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of(), sent(), "a never-ready worker is never delivered to");
assertTrue(f.isCompletedExceptionally(), "the caller's future fails instead of hanging forever");
assertEquals(List.of(T), cap.failed, "the awaiting send resolves through the worker-failure path");
assertEquals(List.of(), cap.completed, "a never-ready worker is a failure, not a completion");
assertEquals(List.of(T), forgotten, "the never-ready worker's presence is cleared");
assertTrue(inj.activeTargets().isEmpty(), "the target is reclaimed, not polled forever");
}
@Test
void aWorkerThatBecomesReadyWithinTheGraceIsDeliveredNormally() {
// The readiness grace must not fail a worker that is merely slow to boot: once it becomes
// available before the grace elapses, delivery proceeds as usual (the counter resets).
Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains, _ -> {
});
inj.enqueue(T, "task");
for (int i = 0; i < 100; i++) inj.onStatus(T, AgentStatus.IDLE); // still booting, well under grace
assertEquals(List.of(), sent());
ready.add(T); // MCP connects before the grace elapses
inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("task"), sent(), "a worker that connects within the grace is delivered to");
}
@Test
void dropClearsWorkerPresence() {
// CB-114 (finding #1): a vanished worker's readiness must be forgotten so a stale entry cannot
// linger past the worker's life (WorkerPresence.forget had no caller before this).
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, forgotten::add);
inj.enqueue(T, "orphan");
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), forgotten, "drop clears the gone worker's presence");
}
@Test
void pollerDeliversToAnIdleWorker() throws Exception {
// End-to-end through the poller: idle worker → message delivered without manual onStatus.
@@ -165,12 +370,13 @@ class InjectorTest {
poller.stop();
}
assertEquals(List.of("via-poller"), idle.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> {
@SuppressWarnings("unchecked")
Map<String, Object> p = (Map<String, Object>) c.params();
return p.get("text").toString();
}).toList());
})
.toList());
}
@Test
@@ -0,0 +1,88 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Content-based refinement of an unreliable {@code UNKNOWN} status (CB-115). */
class StatusRefinerTest {
// --- pure classification -------------------------------------------------
@Test
void classifiesAnIdlePromptAsIdle() {
String pane = """
⏺ All done — the file compiles cleanly.
╭──────────────────────────────────────╮
│ > │
╰──────────────────────────────────────╯
⏵⏵ auto mode on (shift+tab to cycle)
""";
assertEquals(AgentStatus.IDLE, StatusRefiner.classify(pane));
}
@Test
void classifiesABarePromptGlyphAsIdle() {
assertEquals(AgentStatus.IDLE, StatusRefiner.classify("some output\n❯ "));
}
@Test
void classifiesActiveGenerationAsWorking() {
String pane = """
⏺ Working on it…
✳ Thinking… (12s · esc to interrupt)
""";
assertEquals(AgentStatus.WORKING, StatusRefiner.classify(pane));
}
@Test
void anEscToInterruptScreenIsWorkingEvenWithAPromptBox() {
// "esc to interrupt" wins over a prompt box: the turn is still generating.
String pane = "│ > │\n esc to interrupt";
assertEquals(AgentStatus.WORKING, StatusRefiner.classify(pane));
}
@Test
void anUnrecognizableScreenStaysUnknown() {
assertEquals(AgentStatus.UNKNOWN, StatusRefiner.classify("garbled ansi noise with no prompt"));
assertEquals(AgentStatus.UNKNOWN, StatusRefiner.classify(""));
assertEquals(AgentStatus.UNKNOWN, StatusRefiner.classify(null));
}
// --- refine() wiring -----------------------------------------------------
@Test
void refinePassesNonUnknownStatusesThroughWithoutReading() {
FakeHerdr herdr = new FakeHerdr();
StatusRefiner refiner = new StatusRefiner(new AgentControl(herdr));
assertEquals(AgentStatus.WORKING, refiner.refine("term_a", AgentStatus.WORKING));
assertEquals(AgentStatus.IDLE, refiner.refine("term_a", AgentStatus.IDLE));
assertFalse(herdr.called("agent.read"),
"a trusted status must not cost a pane read");
}
@Test
void refineUpgradesUnknownToIdleFromPaneContent() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ answer\n❯ ");
StatusRefiner refiner = new StatusRefiner(new AgentControl(herdr));
assertEquals(AgentStatus.IDLE, refiner.refine("term_a", AgentStatus.UNKNOWN));
assertTrue(herdr.called("agent.read"), "an UNKNOWN must trigger a pane read");
}
@Test
void refineLeavesUnknownWhenContentIsUnclassifiable() {
FakeHerdr herdr = new FakeHerdr().readText("nothing recognizable here");
StatusRefiner refiner = new StatusRefiner(new AgentControl(herdr));
assertEquals(AgentStatus.UNKNOWN, refiner.refine("term_a", AgentStatus.UNKNOWN));
}
}
@@ -0,0 +1,33 @@
package dev.ltms.bridged.inject;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** The CB-113 worker-availability registry. */
class WorkerPresenceTest {
@Test
void tracksPresenceAndForgets() {
WorkerPresence p = new WorkerPresence();
assertFalse(p.isPresent("term_a"), "unseen worker is not available");
p.markPresent("term_a");
assertTrue(p.isPresent("term_a"), "a worker seen on the MCP is available");
p.forget("term_a");
assertFalse(p.isPresent("term_a"), "a torn-down worker is no longer available");
}
@Test
void nullOrBlankMarkIsANoOp() {
WorkerPresence p = new WorkerPresence();
assertDoesNotThrow(() -> {
p.markPresent(null);
p.markPresent(" ");
});
assertFalse(p.isPresent(""), "blank/null contacts (the primary) are never present");
}
}
@@ -0,0 +1,168 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.auth.Role;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-513 — the CB-505 authorization gate on the <strong>MCP</strong> entry path.
*
* <p>Why this file exists: CB-505 claimed authorization is "enforced on both entry paths", and it
* is — but only REST was ever tested ({@code BridgedAppAuthTest}). Coverage showed
* {@code BridgeMcp.deny()}, {@code principal()} and every tool-registration lambda at <em>zero</em>
* executed lines, because no test had ever constructed a {@code BridgeMcp} — the existing
* {@code BridgeMcpTest} calls only the static handler methods. An unexercised security control is
* a claim, not a control.
*
* <p>These tests construct a real {@code BridgeMcp} (which also exercises the constructor and the
* tool wiring) and drive the policy half of the gate directly.
*/
class BridgeMcpAuthzTest {
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private Metrics metrics;
private BridgeMcp mcp;
@AfterEach
void close() {
if (mcp != null) mcp.close();
}
/** A fully wired BridgeMcp on fakes — constructing it is itself part of what is under test. */
private BridgeMcp mcp(boolean enforce) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "tok");
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(),
new InMemoryReplyInbox());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
metrics = BridgedMetrics.create(sessions, new InMemoryReplyInbox());
mcp = new BridgeMcp(messages, workers, sessions, identity, sessions.asPresence(),
new PrimaryRegistry(null),
enforce ? new CallerResolver(identity) : null,
metrics);
return mcp;
}
private static final Principal PRIMARY = Principal.primary(100);
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal ANON = Principal.anonymous();
// --- the table, enforced on THIS path too ---------------------------------------------------
@Test
void primaryMayOrchestrate() {
BridgeMcp m = mcp(true);
for (Authz.Action a : new Authz.Action[]{Authz.Action.SPAWN, Authz.Action.STOP,
Authz.Action.SEND, Authz.Action.DRAIN, Authz.Action.READ}) {
assertNull(m.denyFor(PRIMARY, a, "term_a"), a + " is the primary's to perform");
}
}
@Test
void aWorkerMayNotOrchestrateOverMcp() {
BridgeMcp m = mcp(true);
for (Authz.Action a : new Authz.Action[]{Authz.Action.SPAWN, Authz.Action.STOP,
Authz.Action.SEND, Authz.Action.DRAIN}) {
McpSchema.CallToolResult denied = m.denyFor(WORKER_A, a, "term_a");
assertNotNull(denied, a + " must be refused to a worker");
assertTrue(denied.isError(), "a refusal is returned as an MCP tool error");
}
}
@Test
void aWorkerMayReplyAndAskOnlyAsItself() {
BridgeMcp m = mcp(true);
assertNull(m.denyFor(WORKER_A, Authz.Action.REPLY, "term_a"), "its own session is allowed");
assertNull(m.denyFor(WORKER_A, Authz.Action.ASK, "term_a"));
assertNotNull(m.denyFor(WORKER_A, Authz.Action.REPLY, "term_b"),
"worker A must not reply on worker B's session");
assertNotNull(m.denyFor(WORKER_A, Authz.Action.ASK, "term_b"));
}
@Test
void thePrimaryMayNotForgeAWorkerReplyOverMcp() {
BridgeMcp m = mcp(true);
// A forged reply would resolve the very rendezvous the primary is blocked on.
assertNotNull(m.denyFor(PRIMARY, Authz.Action.REPLY, "term_a"));
assertNotNull(m.denyFor(PRIMARY, Authz.Action.ASK, "term_a"));
}
@Test
void anonymousIsRefusedEverythingAndCountedAsUnauthenticated() {
BridgeMcp m = mcp(true);
McpSchema.CallToolResult denied = m.denyFor(ANON, Authz.Action.READ, null);
assertNotNull(denied, "authenticated as nothing ⇒ authorized for nothing");
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "unauthenticated"));
assertEquals(0, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "forbidden"),
"a missing credential is 401-shaped, not 403-shaped");
}
@Test
void aWrongRoleIsCountedAsForbiddenNotUnauthenticated() {
BridgeMcp m = mcp(true);
assertNotNull(m.denyFor(WORKER_A, Authz.Action.SPAWN, null));
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "forbidden"));
assertEquals(0, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "unauthenticated"),
"the caller IS authenticated — it is just not the right role");
}
@Test
void theLegacyConstructorLeavesTheGateOpen() {
// The 22 pre-existing BridgeMcpTest cases rely on no authorization being enforced.
BridgeMcp m = mcp(false);
assertNull(m.denyFor(ANON, Authz.Action.SPAWN, null),
"no CallerResolver supplied ⇒ authorization not enforced (legacy behaviour)");
}
// --- identity reconstruction from the transport context ------------------------------------
@Test
void principalIsRebuiltFromTheStashedRole() {
assertEquals(Role.WORKER, BridgeMcp.principalFrom("WORKER", "term_a", 7).role());
assertEquals("term_a", BridgeMcp.principalFrom("WORKER", "term_a", 7).terminal());
assertEquals(Role.PRIMARY, BridgeMcp.principalFrom("PRIMARY", null, 7).role());
assertEquals(Role.ANONYMOUS, BridgeMcp.principalFrom("ANONYMOUS", null, -1).role());
}
@Test
void aMissingRoleFallsBackToTheHistoricalInterpretation() {
// Legacy path: no role stashed. A terminal means worker; its absence meant "the primary",
// which is exactly the pre-CB-501 default CB-501 inverted — preserved only here.
assertEquals(Role.WORKER, BridgeMcp.principalFrom(null, "term_a", 7).role());
assertEquals(Role.PRIMARY, BridgeMcp.principalFrom(null, null, 7).role());
}
}
@@ -0,0 +1,416 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.modelcontextprotocol.spec.McpSchema;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
/**
* Parity tests for the MCP tool adapters — they must produce the same outcomes as the REST routes,
* since both drive the same {@link MessageService}/{@link Rendezvous}. The MCP wire protocol itself
* is the SDK's concern; here we test the thin adapter logic directly.
*/
class BridgeMcpTest {
private static final String T = "term_a";
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
private final MessageService messages = new MessageService(agents, new Injector(agents), rendezvous, inbox);
@BeforeEach
void setUp() {
// CB-520: the inbox only peeks/acks targets it owns.
inbox.own(T);
}
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
private static ClaudeCodeLauncher workerService(FakeHerdr h, String baseUrl, Set<String> allow) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", baseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(allow), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> "tok");
}
private static SessionManager sessionManager(FakeHerdr h, String baseUrl, Set<String> allow) {
return new SessionManager(workerService(h, baseUrl, allow));
}
@Test
void sendThenReplyRoundTrips() throws Exception {
// bridge_send blocks; bridge_reply resolves it with the worker's structured answer.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> BridgeMcp.send(messages, "term_a", "review this", 4000L));
// Wait until the send has opened its waiter so the reply resolves it (CB-307: reply now
// queues in the inbox if no waiter is open, which would break the round-trip).
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "send should have opened its waiter");
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "LGTM");
assertEquals("delivered", textOf(reply));
McpSchema.CallToolResult res = send.get(6, TimeUnit.SECONDS);
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("LGTM", textOf(res));
}
@Test
void asyncSendReturnsATicketThenPollReportsTheReply() throws Exception {
// wait:false parity — a ticket is issued, resolved by a reply, and surfaced by bridge_poll.
McpSchema.CallToolResult accepted = BridgeMcp.sendAsync(messages, "term_a", "do it");
assertNotEquals(Boolean.TRUE, accepted.isError());
String out = textOf(accepted);
assertTrue(out.contains("ticket="), out);
String ticket = out.substring(out.indexOf("ticket=") + "ticket=".length()).trim();
// Wait until the send has opened its waiter before replying (CB-307: reply never errors,
// so the old retry-on-error pattern no longer works — it would queue instead of resolve).
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "send should have opened its waiter");
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "async LGTM");
assertEquals("delivered", textOf(reply));
// Poll until the async send completes and reports the reply.
McpSchema.CallToolResult polled = BridgeMcp.poll(messages, ticket, null);
deadline = System.currentTimeMillis() + 3000;
while (!textOf(polled).contains("async LGTM") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(10);
polled = BridgeMcp.poll(messages, ticket, null);
}
assertEquals("async LGTM", textOf(polled));
}
@Test
void pollUnknownTicketIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.poll(messages, "task-999", null);
assertTrue(res.isError());
assertTrue(textOf(res).contains("unknown ticket"));
}
@Test
void sendTimesOutWithAWorkingNote() {
McpSchema.CallToolResult res = BridgeMcp.send(messages, "term_a", "hi", 120L);
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res));
}
@Test
void sendRejectsMissingArgs() {
assertTrue(BridgeMcp.send(messages, null, "hi", null).isError());
assertTrue(BridgeMcp.send(messages, "term_a", " ", null).isError());
}
@Test
void replyWithNoPendingSendIsQueuedNotError() {
// CB-307: a reply with no open send is now queued in the inbox, not an error.
McpSchema.CallToolResult res = BridgeMcp.reply(messages, "term_a", "orphan");
assertNotEquals(Boolean.TRUE, res.isError(), "a queued reply is not an error");
assertEquals("delivered", textOf(res));
// The reply is drainable by target.
var drained = messages.drainReplies("term_a");
assertEquals(1, drained.size());
assertEquals("orphan", drained.getFirst().content());
}
@Test
void bridgePollWithTargetDrainsReplies() {
// A reply with no open send queues it in the inbox.
BridgeMcp.reply(messages, "term_a", "queued-msg");
// bridge_poll with target drains the inbox.
McpSchema.CallToolResult res = BridgeMcp.poll(messages, null, "term_a");
assertNotEquals(Boolean.TRUE, res.isError());
String text = textOf(res);
assertTrue(text.contains("queued-msg"), "the drained reply should appear in the result");
// Second drain returns empty.
McpSchema.CallToolResult empty = BridgeMcp.poll(messages, null, "term_a");
assertEquals("[]", textOf(empty));
}
@Test
void askThenAnswerRoundTrips() throws Exception {
// The primary delegates and blocks; wait until its waiter is open before the worker asks.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> BridgeMcp.send(messages, "term_a", "do X", 5000L));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "the send must be waiting for the ask to surface to");
// The worker asks mid-turn; the call blocks for the primary's answer.
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
() -> BridgeMcp.ask(messages, "term_a", "which config?", 5000L));
// The primary's send unblocks with the question and a turnId to answer on.
McpSchema.CallToolResult q = send.get(6, TimeUnit.SECONDS);
assertNotEquals(Boolean.TRUE, q.isError());
String qt = textOf(q);
assertTrue(qt.contains("[question]"), qt);
String afterMarker = qt.substring(qt.indexOf("turnId=\"") + "turnId=\"".length());
String turnId = afterMarker.substring(0, afterMarker.indexOf('"'));
// The primary answers via bridge_send(turnId); this blocks again for the worker's reply.
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> BridgeMcp.answer(messages, turnId, "config.yaml", 5000L));
// The worker's ask returns the answer — it resumes the same turn.
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
// The resumed worker replies, resolving the answering send (wait for the reopened waiter).
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "the answer should have reopened a waiter");
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "done");
assertEquals("delivered", textOf(reply));
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
}
@Test
void askFromANonWorkerConnectionIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.ask(messages, null, "which config?", 500L);
assertTrue(res.isError());
assertTrue(textOf(res).contains("workers only"), textOf(res));
}
@Test
void answerToAStaleTurnIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.answer(messages, "term_a#999", "too late", 500L);
assertTrue(res.isError());
assertTrue(textOf(res).contains("no longer open"), textOf(res));
}
@Test
void spawnReturnsTheNewWorkersSessionAndPane() {
FakeHerdr h = new FakeHerdr();
SessionManager sm = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
McpSchema.CallToolResult res = BridgeMcp.spawn(sm, null);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"sessionId\":\"term_new_1\""), out);
// CB-519: the "paneId" wire field now carries the host-unique opaque id, not the herdr pane.
WorkerSession s = sm.roster().getFirst();
assertTrue(out.contains("\"paneId\":\"" + s.paneId() + "\""), out);
assertNotEquals("w9:pRoot_1", s.paneId(), "the id is decoupled from the herdr pane coordinate");
assertTrue(out.contains("\"status\":\"spawning\""), out);
}
@Test
void spawnRejectsAnOffAllowlistProfileWithoutTouchingHerdr() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res =
BridgeMcp.spawn(sessionManager(h, "https://api.anthropic.com", Set.of("gx00.gw")), null);
assertTrue(res.isError());
assertTrue(textOf(res).contains("subscription boundary"));
assertFalse(h.called("agent.start"), "the guard must block before any spawn");
}
@Test
void spawnRejectsAnUnknownProfileAsAnError() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res =
BridgeMcp.spawn(sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), "nope");
assertTrue(res.isError());
assertTrue(textOf(res).contains("unknown worker profile"), textOf(res));
}
@Test
void spawnPassesTheRequestedCwdToTheWorker() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "/req/dir", null, null, null);
assertNotEquals(Boolean.TRUE, res.isError());
// Protocol 19: the requested cwd roots the worker's pane at creation (tab.create).
@SuppressWarnings("unchecked")
Map<String, Object> create = (Map<String, Object>) h.lastCall("tab.create").params();
assertEquals("/req/dir", create.get("cwd"));
}
@Test
void profilesListsConfiguredProfilesAndDefault() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.profiles(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("ltms-local"), out);
assertTrue(out.contains("\"default\":\"ltms-local\""), out);
}
@Test
void listReportsTrackedWorkers() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-304", null));
McpSchema.CallToolResult res = BridgeMcp.listWorkers(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"sessionId\":\"" + s.terminalId() + "\""), out);
assertTrue(out.contains("\"paneId\":\"" + s.paneId() + "\""), out);
assertTrue(out.contains("\"profile\":\"ltms-local\""), out);
assertTrue(out.contains("\"state\":\"spawning\""), out);
assertTrue(out.contains("\"worktree\":\"" + s.worktree() + "\""), out);
assertTrue(out.contains("\"branch\":\"" + s.branch() + "\""), out);
assertTrue(out.contains("\"owner\":\"term_primary\""), out);
assertTrue(out.contains("\"liveStatus\":\"unknown\""), out);
}
@Test
void stopTearsDownAWorkerByPane() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.stop(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), "w9:pW");
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("stopped w9:pW", textOf(res));
assertTrue(h.called("pane.close"));
}
@Test
void stopRequiresAPaneId() {
FakeHerdr h = new FakeHerdr();
assertTrue(BridgeMcp.stop(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), " ").isError());
}
@Test
void bridgeAckReturnsConfirmationForValidArgs() {
McpSchema.CallToolResult res = BridgeMcp.ack(messages, "term_a", "msg-1");
assertNotEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("msg-1"), "response should mention the msgId");
}
@Test
void bridgeAckRejectsMissingArgs() {
assertTrue(BridgeMcp.ack(messages, null, "msg-1").isError());
assertTrue(BridgeMcp.ack(messages, "term_a", null).isError());
assertTrue(BridgeMcp.ack(messages, " ", "msg-1").isError());
}
@Test
void bridgeAckRemovesSpecificReply() {
// Queue a reply and capture its msgId.
BridgeMcp.reply(messages, "term_a", "orphan");
var before = messages.drainReplies("term_a");
assertEquals(1, before.size(), "one reply in the inbox");
String msgId = before.getFirst().msgId();
// Publish the same reply again and ack it via bridge_ack surface.
BridgeMcp.reply(messages, "term_a", "orphan-again");
var peeked = messages.drainReplies("term_a");
assertEquals(1, peeked.size(), "one fresh reply in the inbox");
// ackReply works (no-op since published with a different UUID, but callable).
assertDoesNotThrow(() -> messages.ackReply("term_a", msgId));
}
@Test
void statusReportsLiveAgentStatus() {
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
AgentControl blockedAgents = new AgentControl(blocked);
McpSchema.CallToolResult res = BridgeMcp.status(
new MessageService(blockedAgents, new Injector(blockedAgents), rendezvous), "term_a");
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("blocked", textOf(res));
}
// --- bridge_whoami: the caller's own identity, so an agent never has to guess its role -------
@Test
void whoamiReportsThePrimaryAsPrimaryAndNothingElse() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.whoami(
Principal.primary(100), sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"primary\""), out);
// The primary owns no session — leaking a sessionId here would invite it to reply as one.
assertFalse(out.contains("sessionId"), out);
}
@Test
void whoamiReportsAWorkerWithItsRegisteredSession() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-517", null));
McpSchema.CallToolResult res = BridgeMcp.whoami(Principal.worker(s.terminalId(), 200), sessions);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"worker\""), out);
assertTrue(out.contains("\"sessionId\":\"" + s.terminalId() + "\""), out);
assertTrue(out.contains("\"profile\":\"ltms-local\""), out);
assertTrue(out.contains("\"worktree\":\"" + s.worktree() + "\""), out);
assertTrue(out.contains("\"branch\":\"" + s.branch() + "\""), out);
assertTrue(out.contains("\"owner\":\"term_primary\""), out);
}
/**
* A worker the registry has no record of — it outlived a daemon restart — must still learn the
* load-bearing fact. Degrading to "I don't know who you are" would put it back to guessing,
* which is the failure this tool exists to remove.
*/
@Test
void whoamiStillReportsWorkerRoleWhenTheSessionIsUnregistered() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.whoami(Principal.worker("term_orphan", 200),
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"worker\""), out);
assertTrue(out.contains("\"sessionId\":\"term_orphan\""), out);
assertFalse(out.contains("profile"), out); // nothing invented for a session we don't track
}
}
@@ -0,0 +1,46 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/** Connection → caller-identity resolution, with the OS peer-PID lookup faked. */
class ConnectionIdentityTest {
private final FakeHerdr herdr = new FakeHerdr();
private ConnectionIdentity with(PeerPidLookup pids) {
return new ConnectionIdentity(new PaneLocator(herdr), pids);
}
@Test
void resolvesWorkerFromLoopbackPeerPid() {
assertEquals("term_a", with(_ -> FakeHerdr.WORKER_PID).callerTerminal("127.0.0.1", 55555));
}
@Test
void nullForOffHostCaller() {
// A non-loopback peer can't be an on-host worker → treat as primary/unknown.
assertNull(with(_ -> FakeHerdr.WORKER_PID).callerTerminal("10.0.0.9", 55555));
}
@Test
void nullWhenPidOwnsNoPane() {
// e.g. the primary — its PID maps to no worker pane.
assertNull(with(_ -> 999_999).callerTerminal("127.0.0.1", 55555));
}
@Test
void resolvesTheCallersPidAndCwd() {
// CB-112: the primary maps to no pane, but its PID and cwd are still readable.
ConnectionIdentity id = new ConnectionIdentity(
new PaneLocator(herdr), _ -> 999_999, pid -> pid == 999_999 ? "/main/project" : null);
ConnectionIdentity.Caller c = id.resolve("127.0.0.1", 55555);
assertNull(c.terminal(), "the primary owns no worker pane");
assertEquals(999_999, c.pid());
assertEquals("/main/project", id.cwdForPid(c.pid()), "the primary's cwd is resolvable from its PID");
assertNull(id.cwdForPid(-1), "no cwd for an unresolved PID");
}
}
@@ -0,0 +1,108 @@
package dev.ltms.bridged.mcp;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for {@link PrimaryRegistry}: pin vs record, isKnown transitions,
* null/blank guard.
*/
class PrimaryRegistryTest {
@Test
void unpinnedInitiallyUnknown() {
var reg = new PrimaryRegistry(null);
assertFalse(reg.isKnown());
assertTrue(reg.primaryTerminal().isEmpty());
}
@Test
void unpinnedAcceptsBlankAsAbsent() {
var reg = new PrimaryRegistry("");
assertFalse(reg.isKnown());
assertTrue(reg.primaryTerminal().isEmpty());
}
@Test
void pinnedFromConstruction() {
var reg = new PrimaryRegistry("term_fixed");
assertTrue(reg.isKnown());
assertEquals("term_fixed", reg.primaryTerminal().orElseThrow());
}
@Test
void recordWhenUnpinnedSetsTheTerminal() {
var reg = new PrimaryRegistry(null);
reg.record("term_abc");
assertTrue(reg.isKnown());
assertEquals("term_abc", reg.primaryTerminal().orElseThrow());
}
@Test
void recordWithNullDoesNothingWhenUnpinned() {
var reg = new PrimaryRegistry(null);
reg.record(null);
assertFalse(reg.isKnown());
}
@Test
void recordWithBlankDoesNothingWhenUnpinned() {
var reg = new PrimaryRegistry(null);
reg.record(" ");
assertFalse(reg.isKnown());
}
@Test
void recordOverwritesWhenUnpinned() {
var reg = new PrimaryRegistry(null);
reg.record("term_first");
assertEquals("term_first", reg.primaryTerminal().orElseThrow());
reg.record("term_second");
assertEquals("term_second", reg.primaryTerminal().orElseThrow());
}
@Test
void recordIsIgnoredWhenPinned() {
var reg = new PrimaryRegistry("term_pinned");
reg.record("term_other");
assertEquals("term_pinned", reg.primaryTerminal().orElseThrow(), "pinned value must survive record");
}
@Test
void nullRecordIsIgnoredWhenPinned() {
var reg = new PrimaryRegistry("term_pinned");
reg.record(null);
assertTrue(reg.isKnown());
assertEquals("term_pinned", reg.primaryTerminal().orElseThrow());
}
@Test
void blankRecordIsIgnoredWhenPinned() {
var reg = new PrimaryRegistry("term_pinned");
reg.record(" ");
assertTrue(reg.isKnown());
assertEquals("term_pinned", reg.primaryTerminal().orElseThrow());
}
@Test
void isKnownFalseAfterConstructionWithNull() {
var reg = new PrimaryRegistry(null);
assertFalse(reg.isKnown());
}
@Test
void isKnownAfterRecord() {
var reg = new PrimaryRegistry(null);
reg.record("term_x");
assertTrue(reg.isKnown());
}
@Test
void primaryTerminalRoundTrip() {
var reg = new PrimaryRegistry(null);
assertTrue(reg.primaryTerminal().isEmpty());
reg.record("term_found");
assertEquals("term_found", reg.primaryTerminal().get());
}
}
@@ -0,0 +1,127 @@
package dev.ltms.bridged.metrics;
import org.junit.jupiter.api.Test;
import java.util.LinkedHashMap;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/** CB-502 — the zero-dependency Prometheus text renderer. */
class MetricsTest {
@Test
void countersAccumulatePerLabelSet() {
Metrics m = new Metrics();
m.inc("bridged_sends_total", "outcome", "replied");
m.inc("bridged_sends_total", "outcome", "replied");
m.inc("bridged_sends_total", "outcome", "timeout");
assertEquals(2, m.count("bridged_sends_total", "outcome", "replied"));
assertEquals(1, m.count("bridged_sends_total", "outcome", "timeout"));
assertEquals(0, m.count("bridged_sends_total", "outcome", "failed"),
"an untouched series reads as zero, not an error");
}
@Test
void rendersHelpAndTypeOncePerFamily() {
Metrics m = new Metrics();
m.describe("bridged_sends_total", "counter", "Delegated sends by outcome.");
m.inc("bridged_sends_total", "outcome", "replied");
m.inc("bridged_sends_total", "outcome", "timeout");
String out = m.render();
assertEquals(1, countOccurrences(out, "# HELP bridged_sends_total"),
"HELP is per family, not per series");
assertEquals(1, countOccurrences(out, "# TYPE bridged_sends_total counter"));
assertTrue(out.contains("bridged_sends_total{outcome=\"replied\"} 1"));
assertTrue(out.contains("bridged_sends_total{outcome=\"timeout\"} 1"));
}
@Test
void labelsAreSortedSoScrapesAreByteStable() {
Metrics a = new Metrics();
a.inc("m", "b", "2", "a", "1");
Metrics b = new Metrics();
b.inc("m", "a", "1", "b", "2");
assertEquals(a.render(), b.render(), "label order in the call must not change the output");
assertTrue(a.render().contains("m{a=\"1\",b=\"2\"}"));
}
@Test
void gaugesAreEvaluatedAtScrapeTimeNotRegistrationTime() {
Metrics m = new Metrics();
int[] live = {1};
m.gauge("bridged_sessions", () -> live[0], "state", "ready");
assertTrue(m.render().contains("bridged_sessions{state=\"ready\"} 1"));
live[0] = 5;
assertTrue(m.render().contains("bridged_sessions{state=\"ready\"} 5"),
"the gauge must read current state on every scrape");
}
@Test
void aThrowingGaugeDoesNotBreakTheWholeScrape() {
Metrics m = new Metrics();
m.inc("good_total");
m.gauge("bad_gauge", () -> {
throw new IllegalStateException("herdr is down");
});
String out = assertDoesNotThrow(m::render);
assertTrue(out.contains("good_total 1"), "healthy series must still be exported");
assertFalse(out.contains("bad_gauge"), "the broken series is simply absent");
}
@Test
void collectorsDiscoverTheirLabelSetPerScrape() {
Metrics m = new Metrics();
Map<String, Number> depths = new LinkedHashMap<>();
m.collector("bridged_inbox_depth", "target", () -> depths);
assertFalse(m.render().contains("bridged_inbox_depth"), "no targets yet ⇒ no series");
depths.put("term_a", 2);
depths.put("term_b", 0);
String out = m.render();
assertTrue(out.contains("bridged_inbox_depth{target=\"term_a\"} 2"));
assertTrue(out.contains("bridged_inbox_depth{target=\"term_b\"} 0"));
}
@Test
void labelValuesAreEscaped() {
Metrics m = new Metrics();
m.inc("m", "detail", "he said \"hi\"\nand \\left");
String out = m.render();
assertTrue(out.contains("\\\""), "quotes escaped");
assertTrue(out.contains("\\n"), "newlines escaped — a raw one would corrupt the exposition");
assertTrue(out.contains("\\\\"), "backslashes escaped");
}
@Test
void wholeNumberGaugesRenderWithoutADecimalPoint() {
Metrics m = new Metrics();
m.gauge("whole", () -> 3.0);
m.gauge("fractional", () -> 1.5);
String out = m.render();
assertTrue(out.contains("whole 3"), "3.0 should not render as 3.0");
assertTrue(out.contains("fractional 1.5"));
}
@Test
void oddLabelCountIsRejected() {
Metrics m = new Metrics();
assertThrows(IllegalArgumentException.class, () -> m.inc("m", "dangling"));
}
private static int countOccurrences(String haystack, String needle) {
int n = 0;
for (int i = haystack.indexOf(needle); i >= 0; i = haystack.indexOf(needle, i + 1)) {
n++;
}
return n;
}
}
@@ -0,0 +1,192 @@
package dev.ltms.bridged.msg;
import com.rabbitmq.client.AMQP;
import com.rabbitmq.client.Channel;
import com.rabbitmq.client.Connection;
import com.rabbitmq.client.ConnectionFactory;
import org.junit.jupiter.api.BeforeAll;
import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import org.testcontainers.containers.RabbitMQContainer;
import org.testcontainers.junit.jupiter.Container;
import org.testcontainers.junit.jupiter.Testcontainers;
import org.testcontainers.utility.DockerImageName;
import java.util.List;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Contract test for {@link AmqpReplyInbox} against a REAL broker (a RabbitMQ container — the same
* AMQP 0-9-1 the production LavinMQ deploy speaks, URI-only swap). Tagged {@code contract} so it is
* excluded from {@code mvn test}/{@code mvn clean install} (which stay hermetic and need no Docker);
* run it with Docker present via {@code mvn test -Pcontract}.
*
* <p>Two broker modes:
* <ul>
* <li><b>Locally</b> ({@code AMQP_URI} unset): Testcontainers spins a RabbitMQ container. Requires
* a working Docker engine; see the "Running the contract tests" note in
* {@code docs/CB-307-Reliable-Delivery.md} for the {@code api.version} engine-compat pin.</li>
* <li><b>In CI</b> ({@code AMQP_URI} set): a RabbitMQ service container provisions the broker and
* {@code AMQP_URI} points at it, so the contract job needs <em>no</em> Docker on the runner.</li>
* </ul>
*
* <p>It proves the port contract on genuine infrastructure: eventual visibility of a published reply,
* ack removal, msgId dedup, explicit ownership, and — the reason Stage 2 exists — cross-restart
* durability: an unacked reply survives closing the inbox and is redelivered to a fresh connection.
*/
@Tag("contract")
// disabledWithoutDocker=false on purpose: on the CI path (AMQP_URI set) no container is started and
// the class must still RUN against the external broker even though the runner has no Docker — a
// disabled-without-docker check would silently skip the whole contract suite there.
@Testcontainers(disabledWithoutDocker = false)
class AmqpReplyInboxContractTest {
// When a broker is provisioned out-of-band (CI service container), AMQP_URI takes us straight to
// it and we never touch Testcontainers. Unset locally → Testcontainers starts the container below.
private static final String EXTERNAL_URI = System.getenv("AMQP_URI");
private static final RabbitMQContainer BROKER =
new RabbitMQContainer(DockerImageName.parse("rabbitmq:3.13-management"));
// No @Container on BROKER: the JUnit 5 extension would force-start it even when AMQP_URI is set.
// Start it manually only on the local (no-external-broker) path; Ryuk reaps it on JVM exit.
@BeforeAll
static void startBrokerUnlessExternal() {
if (EXTERNAL_URI == null) {
BROKER.start();
}
}
private static String uri() {
if (EXTERNAL_URI != null) {
return EXTERNAL_URI;
}
// guest/guest against the mapped AMQP port. No trailing slash: an empty path is vhost "",
// which does not exist — omitting it selects the default vhost "/".
return "amqp://guest:guest@" + BROKER.getHost() + ":" + BROKER.getAmqpPort();
}
@Test
void publishThenPeekThenAck() throws Exception {
String target = "worker-pub-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.own(target);
inbox.publish(target, "m1", "hello primary");
List<ReplyInbox.InboxMessage> got = awaitPeek(inbox, target);
assertEquals(1, got.size(), "the published reply should be held for drain");
assertEquals("m1", got.getFirst().msgId());
assertEquals(target, got.getFirst().target());
assertEquals("hello primary", got.getFirst().content());
inbox.ack(target, "m1");
assertTrue(inbox.peek(target).isEmpty(), "an acked reply is dropped");
}
}
@Test
void duplicateMsgIdIsNotDoubleQueued() throws Exception {
String target = "worker-dedup-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.own(target);
inbox.publish(target, "dup", "first");
awaitPeek(inbox, target);
inbox.publish(target, "dup", "second"); // same msgId — must be a no-op
// Give any erroneous second delivery time to land, then assert still exactly one.
Thread.sleep(500);
List<ReplyInbox.InboxMessage> got = inbox.peek(target);
assertEquals(1, got.size(), "a repeated msgId must not double-queue");
assertEquals("first", got.getFirst().content(), "the first payload wins");
}
}
@Test
void unackedReplySurvivesRestartAndIsRedelivered() throws Exception {
String target = "worker-durable-" + System.nanoTime();
// First "process life": own, publish, see it held, but crash before acking.
try (AmqpReplyInbox first = AmqpReplyInbox.open(uri())) {
first.own(target);
first.publish(target, "persist-1", "survive me");
assertEquals(1, awaitPeek(first, target).size());
// no ack — simulate a java -jar bounce with the reply still pending
}
// Second "process life": a fresh connection to the same broker must be redelivered the reply.
try (AmqpReplyInbox second = AmqpReplyInbox.open(uri())) {
second.own(target);
List<ReplyInbox.InboxMessage> got = awaitPeek(second, target);
assertEquals(1, got.size(), "an unacked persistent reply is redelivered after restart");
assertEquals("persist-1", got.getFirst().msgId());
assertEquals("survive me", got.getFirst().content());
second.ack(target, "persist-1");
}
// Third life: once acked, it is gone for good — durability is not endless replay.
try (AmqpReplyInbox third = AmqpReplyInbox.open(uri())) {
third.own(target);
Thread.sleep(500);
assertTrue(third.peek(target).isEmpty(), "an acked reply does not come back on the next restart");
}
}
@Test
void publishDoesNotAttachAConsumer() throws Exception {
String target = "worker-pub-no-consumer-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri());
Connection inspect = newConnection()) {
inbox.own(target);
inbox.publish(target, "m1", "published");
awaitPeek(inbox, target); // ensure the owner's consumer received it
try (Channel ch = inspect.createChannel()) {
AMQP.Queue.DeclareOk ok = ch.queueDeclare(queueName(target), true, false, false, null);
assertEquals(1, ok.getConsumerCount(),
"publish must not attach a consumer; only the owner's consumer should exist");
}
}
}
@Test
void releaseCancelsConsumerAndClearsHeld() throws Exception {
String target = "worker-release-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.own(target);
inbox.publish(target, "m1", "release me");
assertEquals(1, awaitPeek(inbox, target).size(), "owned target holds the reply");
inbox.release(target);
assertTrue(inbox.peek(target).isEmpty(),
"release clears the local held snapshot");
}
}
/** Poll peek (broker delivery is async) until a reply for {@code target} appears or ~10s elapse. */
@SuppressWarnings("BusyWait") // deliberate poll for async broker delivery, bounded by the deadline
private static List<ReplyInbox.InboxMessage> awaitPeek(AmqpReplyInbox inbox, String target)
throws InterruptedException {
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(10);
List<ReplyInbox.InboxMessage> msgs = inbox.peek(target);
while (msgs.isEmpty() && System.nanoTime() < deadline) {
Thread.sleep(50);
msgs = inbox.peek(target);
}
return msgs;
}
/** A separate broker connection for inspecting queue state without disturbing the inbox. */
private static Connection newConnection() throws Exception {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri());
return factory.newConnection("contract-inspector");
}
private static String queueName(String target) {
return "agent." + target + ".inbox";
}
}
@@ -0,0 +1,188 @@
package dev.ltms.bridged.msg;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for {@link InMemoryReplyInbox}: publish, peek, ack, dedup, FIFO ordering, thread
* safety under concurrent publish vs. drain, and explicit ownership.
*/
class InMemoryReplyInboxTest {
private final ReplyInbox inbox = new InMemoryReplyInbox();
@BeforeEach
void setUp() {
inbox.own("term_a");
}
@Test
void publishThenPeekReturnsTheMessage() {
inbox.publish("term_a", "m1", "hello");
var msgs = inbox.peek("term_a");
assertEquals(1, msgs.size());
assertEquals("m1", msgs.getFirst().msgId());
assertEquals("term_a", msgs.getFirst().target());
assertEquals("hello", msgs.getFirst().content());
}
@Test
void peekForUnknownTargetReturnsEmpty() {
assertTrue(inbox.peek("no-such-target").isEmpty());
}
@Test
void ackRemovesTheMessage() {
inbox.publish("term_a", "m1", "hello");
inbox.ack("term_a", "m1");
assertTrue(inbox.peek("term_a").isEmpty(), "after ack, the message is gone");
}
@Test
void ackForUnknownMsgIdIsNoOp() {
inbox.publish("term_a", "m1", "hello");
inbox.ack("term_a", "no-such-id"); // no-op
assertEquals(1, inbox.peek("term_a").size(), "the published message is still there");
}
@Test
void ackForUnknownTargetIsNoOp() {
inbox.ack("no-such-target", "m1"); // no-op, should not throw
}
@Test
void dedupByIdempotentMsgId() {
inbox.publish("term_a", "m1", "first");
inbox.publish("term_a", "m1", "second"); // same msgId, different content
var msgs = inbox.peek("term_a");
assertEquals(1, msgs.size(), "dedup: second publish with same msgId is a no-op");
assertEquals("first", msgs.getFirst().content(), "the original content is retained");
}
@Test
void publishesWithDifferentMsgIdsBothAppear() {
inbox.publish("term_a", "m1", "first");
inbox.publish("term_a", "m2", "second");
var msgs = inbox.peek("term_a");
assertEquals(2, msgs.size());
assertEquals("m1", msgs.get(0).msgId());
assertEquals("m2", msgs.get(1).msgId());
}
@Test
void perTargetIsolation() {
inbox.own("term_b");
inbox.publish("term_a", "m1", "for-a");
inbox.publish("term_b", "m2", "for-b");
assertEquals(1, inbox.peek("term_a").size());
assertEquals(1, inbox.peek("term_b").size());
}
@Test
void fifoOrderIsPreserved() {
inbox.publish("term_a", "m1", "first");
inbox.publish("term_a", "m2", "second");
inbox.publish("term_a", "m3", "third");
var msgs = inbox.peek("term_a");
assertEquals(3, msgs.size());
assertEquals("m1", msgs.get(0).msgId());
assertEquals("m2", msgs.get(1).msgId());
assertEquals("m3", msgs.get(2).msgId());
}
@Test
void peekReturnsAnImmutableCopy() {
inbox.publish("term_a", "m1", "hello");
var msgs = inbox.peek("term_a");
assertThrows(UnsupportedOperationException.class, () -> msgs.add(
new ReplyInbox.InboxMessage("x", "term_a", "x")));
}
@Test
void ackRemovesOneMessageLeavesOthers() {
inbox.publish("term_a", "m1", "first");
inbox.publish("term_a", "m2", "second");
inbox.ack("term_a", "m1");
var msgs = inbox.peek("term_a");
assertEquals(1, msgs.size());
assertEquals("m2", msgs.getFirst().msgId());
}
@Test
void concurrentPublishAndDrain() throws Exception {
int msgCount = 100;
ExecutorService exec = Executors.newVirtualThreadPerTaskExecutor();
try {
// Concurrent publishers
var pubDone = new CountDownLatch(msgCount);
for (int i = 0; i < msgCount; i++) {
final int id = i;
exec.submit(() -> {
inbox.publish("term_a", "m" + id, "content-" + id);
pubDone.countDown();
});
}
// Concurrent drainer
AtomicReference<Exception> drainError = new AtomicReference<>();
exec.submit(() -> {
try {
pubDone.await();
for (int i = 0; i < 50; i++) {
var peeked = inbox.peek("term_a");
for (var msg : peeked) {
inbox.ack("term_a", msg.msgId());
}
}
} catch (Exception e) {
drainError.set(e);
}
}).get();
assertNull(drainError.get(), "concurrent drain should not throw");
} finally {
exec.shutdown();
}
}
@Test
void publishWithoutOwnDoesNotClaimOwnership() {
String unowned = "term_unowned";
inbox.publish(unowned, "m1", "hello");
// Without an owner, peek returns nothing — publish did not imply consume.
assertTrue(inbox.peek(unowned).isEmpty(),
"publishing to an unowned target must not make it peekable");
// Owning afterwards makes the already-published message visible.
inbox.own(unowned);
var msgs = inbox.peek(unowned);
assertEquals(1, msgs.size());
assertEquals("m1", msgs.getFirst().msgId());
}
@Test
void peekAndAckAreNoOpsForUnownedTarget() {
assertTrue(inbox.peek("term_not_owned").isEmpty());
inbox.ack("term_not_owned", "m1"); // no-op, should not throw
}
@Test
void releaseStopsConsumingAndClearsHeld() {
inbox.publish("term_a", "m1", "hello");
assertEquals(1, inbox.peek("term_a").size());
inbox.release("term_a");
assertTrue(inbox.peek("term_a").isEmpty(),
"after release, the local snapshot is cleared");
}
@Test
void ownIsIdempotent() {
inbox.own("term_a");
inbox.publish("term_a", "m1", "hello");
assertEquals(1, inbox.peek("term_a").size());
}
}
@@ -0,0 +1,475 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.Injector;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* The message layer's resolution paths (CB-104 reply + CB-106 completion fallback). The turn is
* driven deterministically by feeding {@code onStatus} rather than running a real poller.
*/
class MessageServiceTest {
private static final String T = "term_a";
private final FakeHerdr herdr = new FakeHerdr().readText("BUILD GREEN: 391 files");
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final CompletionResolver completion = new CompletionResolver(agents, rendezvous);
private final Injector injector = new Injector(agents, completion);
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
private final MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
@BeforeEach
void setUp() {
// CB-520: the inbox only peeks/acks targets it owns.
inbox.own(T);
}
/** Run {@code send} on a background thread; the current thread drives the worker's turn. */
private CompletableFuture<MessageService.Reply> sendAsync() {
return CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 5000));
}
private void awaitWaiting() throws InterruptedException {
long deadline = System.currentTimeMillis() + 2000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T), "send should have opened its rendezvous waiter");
}
@Test
void completionFallbackResolvesATurnThatNeverCalledBridgeReply() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
herdr.readText("$ prompt"); // pre-turn pane: no answer yet (baseline reference)
injector.onStatus(T, AgentStatus.IDLE); // deliver the task (baselines the pre-turn content)
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up and works
herdr.readText("BUILD GREEN: 391 files"); // the worker's turn produced new output
injector.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete, no bridge_reply
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, reply.outcome(),
"an unreplied but finished turn resolves via the completion fallback");
assertEquals("BUILD GREEN: 391 files", reply.text(), "the scraped transcript tail is returned");
assertTrue(reply.completed(), "a scraped completion still counts as completed");
}
@Test
void explicitBridgeReplyResolvesAsReplied() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker working
assertTrue(rendezvous.resolve(T, "LGTM ship it"), "an explicit reply resolves the send");
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, reply.outcome());
assertEquals("LGTM ship it", reply.text());
}
@Test
void aWedgedWorkerResolvesTheSendAsFailedWithTheErrorContext() throws Exception {
herdr.readText("API Error: Unable to connect to API (ENOTFOUND)");
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker starts the turn
for (int i = 0; i < 130; i++) injector.onStatus(T, AgentStatus.UNKNOWN); // then wedges (CB-109)
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome());
assertFalse(reply.completed(), "a wedge is terminal but not a successful completion");
assertTrue(reply.text().contains("ENOTFOUND"), "the error screen is carried as the failure reason");
}
@Test
void aWorkerThatVanishesMidTurnResolvesTheSendAsFailed() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker starts the turn
// The worker's pane crashes — the poller sees a *_not_found and drops it (CB-110).
injector.drop(T, new HerdrException("worker gone", "pane_not_found", null));
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome(),
"a delivered send whose worker vanishes fails instead of hanging to the timeout");
assertFalse(reply.completed());
}
// --- bridge_ask reverse rendezvous (CB-205) ------------------------------------------------
@Test
void askSurfacesAsAQuestionAndTheAnswerResumesTheSameTurn() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up, then pauses to ask
// The worker asks mid-turn on its own thread; the call blocks for the primary's answer.
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
// The primary's blocking send unblocks with the question and a turnId to answer on.
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertEquals("which config file?", q.text());
assertNotNull(q.turnId(), "a question carries a turnId to answer on");
// The primary answers via bridge_send(turnId); this blocks again for the worker's reply.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
// The worker's ask returns the answer — it resumes the same turn.
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome());
assertEquals("config.yaml", a.answer());
// The resumed worker finishes with a structured reply, resolving the answering send.
awaitWaiting(); // the answering send has (re)opened its forward waiter
assertTrue(rendezvous.resolve(T, "done"), "the worker's final reply resolves the answering send");
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
assertEquals("done", done.text());
}
@Test
void duplicateAsksFromTheSameSessionCoalesceToOneTurn() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up, then pauses to ask
// A transport retry: two concurrent bridge_ask calls from the same worker session.
CompletableFuture<MessageService.AskResult> ask1 =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
CompletableFuture<MessageService.AskResult> ask2 =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
// The primary's single blocked send surfaces exactly ONE question (one turnId).
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertEquals("which config file?", q.text());
assertNotNull(q.turnId(), "only one turnId should be minted");
// The primary answers that one turnId; both asks unblock with the same answer.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
MessageService.AskResult a1 = ask1.get(5, TimeUnit.SECONDS);
MessageService.AskResult a2 = ask2.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a1.outcome());
assertEquals("config.yaml", a1.answer());
assertEquals(MessageService.AskOutcome.ANSWERED, a2.outcome());
assertEquals("config.yaml", a2.answer());
// The resumed worker finishes with a structured reply, resolving the answering send.
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"), "the worker's final reply resolves the answering send");
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
assertEquals("done", done.text());
}
@Test
void askWithNoOpenDelegationReturnsNoWaiter() {
MessageService.AskResult r = messages.ask(T, "anyone listening?", 500);
assertEquals(MessageService.AskOutcome.NO_WAITER, r.outcome(),
"a question with no blocked send has no primary to answer it");
}
@Test
void askTimesOutWhenThePrimaryNeverAnswers() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
MessageService.AskResult r = messages.ask(T, "still there?", 200); // primary never answers
assertEquals(MessageService.AskOutcome.TIMED_OUT, r.outcome());
// The send itself already unblocked with the question the instant the ask surfaced.
MessageService.Reply q = send.get(2, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
}
@Test
void answeringAnUnknownTurnIsStale() {
MessageService.Reply r = messages.answer(T + "#999", "too late", 500);
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
"an answer to a turn that never existed (or already lapsed) is stale, not a hang");
}
// --- timeout, answer, poll, and lock-contention edges ----------------------------------
@Test
void sendTimesOutBeforeDeliveryIsQueuedNotWorking() {
// Nothing ever delivers the message and nothing resolves the send, so the reply future
// times out with delivery still incomplete — the message is still queued for the worker.
MessageService.Reply r = messages.send(T, "never delivered", 50);
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome(),
"an undelivered send that times out is still queued, not working");
assertNull(r.text());
}
@Test
void sendTimesOutAfterDeliveryIsStillWorking() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 300));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver — the delivered future now completes
injector.onStatus(T, AgentStatus.WORKING); // worker starts but never replies
// No rendezvous.resolve(T, ...) — the reply future rides out its short timeout.
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, r.outcome(),
"a delivered send whose worker never replies times out as still working");
}
@Test
void answerTimesOutWhenTheResumedWorkerNeverReplies() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertNotNull(q.turnId());
// The primary answers, unblocking the worker; but the worker never sends the follow-up
// bridge_reply, so the answering send rides out its short window as still-working.
MessageService.Reply answer = messages.answer(q.turnId(), "config.yaml", 200);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answer.outcome(),
"an answered worker that never replies times out as still working");
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome());
assertEquals("config.yaml", a.answer());
}
@Test
void pollReturnsNullForAnUnknownTicket() {
assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown");
}
@Test
void pollReportsACompletedTicket() throws Exception {
String ticket = messages.sendAsync(T, "long task");
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker works
assertTrue(rendezvous.resolve(T, "async result"), "a reply resolves the async send");
// Wait for the background send to finish and publish a DONE view.
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 2000;
while (view == null || view.phase() != MessageService.Phase.DONE) {
if (System.currentTimeMillis() >= deadline) break;
view = messages.poll(ticket);
//noinspection BusyWait
Thread.sleep(5);
}
assertNotNull(view, "a resolved async send must become DONE");
assertEquals(MessageService.Phase.DONE, view.phase());
assertEquals("async result", view.reply(), "the completed ticket reports the reply");
assertEquals("reply", view.replySource(), "a structured bridge_reply is sourced from 'reply'");
}
@Test
void concurrentSendToSameSessionWhileFirstHoldsItIsBusy() throws Exception {
CompletableFuture<MessageService.Reply> first =
CompletableFuture.supplyAsync(() -> messages.send(T, "first", 5000));
awaitWaiting(); // the first send now holds the session lock, blocked on its reply
// A second send to the SAME session cannot take the lock within its short window.
MessageService.Reply busy = messages.send(T, "second", 100);
assertEquals(MessageService.Outcome.BUSY, busy.outcome(),
"a second send while another holds the session is busy, not a hang");
assertNull(busy.text());
// Release the first send so it resolves cleanly and the test thread is not left pinned.
injector.onStatus(T, AgentStatus.IDLE); // deliver the first message
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up
assertTrue(rendezvous.resolve(T, "first done"), "the first send resolves with a reply");
MessageService.Reply firstReply = first.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, firstReply.outcome());
assertEquals("first done", firstReply.text());
}
// --- CB-307 reply inbox ----------------------------------------------------------------
@Test
void replyQueuesInInboxWhenNoSendIsOpen() {
// No send is open for this session — reply should queue in the inbox.
assertTrue(messages.reply(T, "queued-text"), "reply should succeed (queued)");
var drained = messages.drainReplies(T);
assertEquals(1, drained.size());
assertEquals("queued-text", drained.getFirst().content());
}
@Test
void replyResolvesOpenSendDoesNotQueue() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitUninterruptibly(T);
// An explicit reply resolves the open send.
assertTrue(messages.reply(T, "send-resolved"), "reply should succeed (resolved live send)");
// The inbox should be empty — the reply went to the send, not the inbox.
assertTrue(messages.drainReplies(T).isEmpty(), "no reply in the inbox");
MessageService.Reply r = send.get(3, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, r.outcome());
assertEquals("send-resolved", r.text());
}
@Test
void drainRepliesReturnsAllPendingThenEmptyOnNextCall() {
messages.reply(T, "msg-1");
messages.reply(T, "msg-2");
var first = messages.drainReplies(T);
assertEquals(2, first.size());
var second = messages.drainReplies(T);
assertTrue(second.isEmpty(), "second drain should be empty (acked)");
}
@Test
void aQuestionIsNeverQueuedInTheInbox() {
// No send is open — bridge_ask with no delegation returns NO_WAITER,
// and the question text MUST NOT appear in the reply inbox.
// The inbox is only fed by MessageService.reply(), not by bridge_ask.
MessageService.AskResult r = messages.ask(T, "anyone there?", 500);
assertEquals(MessageService.AskOutcome.NO_WAITER, r.outcome(),
"bridge_ask with no open delegation must return NO_WAITER, never queued");
assertTrue(messages.drainReplies(T).isEmpty(), "questions must never be queued");
}
@Test
void completionFallbackIsNeverQueued() throws Exception {
// The fallback resolves a captured waiter, never the inbox.
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitUninterruptibly(T);
injectDelivery();
// The worker never sends bridge_reply, but the turn completes.
herdr.readText("done-scraped");
completion.onTurnComplete(T); // The fallback arms and resolves the captured waiter.
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, r.outcome());
// The inbox should be empty — the reply went to the captured waiter.
assertTrue(messages.drainReplies(T).isEmpty(), "completion fallback must not queue");
}
// --- helpers ---------------------------------------------------------------------------
/** Like {@link #awaitWaiting()} but rethrows as unchecked. */
private void awaitUninterruptibly(String session) {
try {
awaitWaiting();
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException(e);
}
}
/** Set up a delivered turn so the worker is working, ready for an ask or completion. */
private void injectDelivery() {
herdr.readText("$ prompt"); // pre-turn content baseline
injector.onStatus(T, AgentStatus.IDLE); // deliver the task
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up
}
// --- CB-516: a released session must not leave a send hanging ------------------------------
/**
* The bug this fixes: tearing a worker down left its rendezvous waiter open, so a blocking send
* kept blocking and an async one kept reporting PENDING until the 30-minute async timeout —
* even though the worker provably no longer existed.
*/
@Test
void abandonFailsASendThatIsStillWaitingOnAReleasedSession() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
awaitWaiting();
assertTrue(messages.abandon(T, "session released"), "a live waiter is abandoned");
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, r.outcome(),
"an abandoned send fails rather than riding out its timeout");
assertEquals("session released", r.text(), "the caller is told why");
}
@Test
void abandonIsANoOpWhenNobodyIsWaiting() {
assertFalse(messages.abandon(T, "session released"),
"no open send ⇒ nothing to abandon");
}
@Test
void abandonDoesNotOverwriteAnAlreadyResolvedSend() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
awaitWaiting();
assertTrue(rendezvous.resolve(T, "the real answer"));
assertFalse(messages.abandon(T, "session released"),
"a send already answered by the worker must not be clobbered");
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals("the real answer", r.text());
}
/** The async path is the one that hung: poll must report FAILED, not PENDING forever. */
@Test
void anAbandonedAsyncTaskPollsAsFailedNotPending() throws Exception {
String ticket = messages.sendAsync(T, "long task");
awaitWaiting();
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase());
messages.abandon(T, "session released");
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 3000;
while (System.currentTimeMillis() < deadline) {
view = messages.poll(ticket);
if (view.phase() != MessageService.Phase.PENDING) break;
Thread.sleep(10);
}
assertNotNull(view);
assertEquals(MessageService.Phase.FAILED, view.phase(),
"a delegation whose worker is gone must not keep reporting PENDING");
assertTrue(view.detail() != null && view.detail().contains("released"),
"and the detail says why, rather than 'worker unknown'");
}
}
@@ -0,0 +1,102 @@
package dev.ltms.bridged.msg;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CompletableFuture;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* The reverse rendezvous (CB-205): the {@code bridge_ask} registry that lets a worker pause mid-turn
* to ask the primary. Unit-level — the message-layer round-trip is covered in {@link MessageServiceTest}.
*/
class RendezvousTest {
private static final String W = "term_a";
private final Rendezvous rendezvous = new Rendezvous();
@Test
void openAskMintsAUniqueTurnScopedToItsSessionAndCoalescesDuplicates() {
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
Rendezvous.AskTicket t2 = rendezvous.openAsk(W);
assertEquals(t1.turnId(), t2.turnId(), "duplicate asks from the same session coalesce onto one turn");
assertTrue(t1.fresh(), "the first ask freshly opens the turn");
assertFalse(t2.fresh(), "the coalesced ask rides the existing turn");
assertTrue(t1.turnId().startsWith(W + "#"), "the turnId is scoped to the worker session");
assertEquals(W, rendezvous.askSession(t1.turnId()));
}
@Test
void openAskAfterCloseMintsANewTurn() {
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
rendezvous.closeAsk(t1.turnId());
Rendezvous.AskTicket t2 = rendezvous.openAsk(W);
assertNotEquals(t1.turnId(), t2.turnId(), "after closing, a new ask gets a fresh turnId");
assertTrue(t2.fresh(), "the reopened ask is fresh");
assertEquals(W, rendezvous.askSession(t2.turnId()));
}
@Test
void answerAskCompletesTheWaitersFuture() {
Rendezvous.AskTicket t = rendezvous.openAsk(W);
assertTrue(rendezvous.answerAsk(t.turnId(), "config.yaml"), "answering an open ask succeeds");
assertEquals("config.yaml", t.answer().getNow(null), "the answer reaches the blocked worker");
}
@Test
void answerAskOnAnUnknownTurnIsFalse() {
assertFalse(rendezvous.answerAsk("no-such#1", "x"), "an answer to an unknown turn is a no-op");
}
@Test
void resolveQuestionResolvesAnOpenSendWithTheQuestionKindAndTurnId() {
CompletableFuture<Rendezvous.Resolution> send = rendezvous.open(W);
assertTrue(rendezvous.resolveQuestion(W, "which config?", W + "#7"),
"the question resolves the primary's open send");
Rendezvous.Resolution r = send.getNow(null);
assertEquals(Rendezvous.Kind.QUESTION, r.kind());
assertEquals("which config?", r.text());
assertEquals(W + "#7", r.turnId(), "the turnId rides along so the primary can answer");
}
@Test
void resolveQuestionWithNoOpenSendIsFalse() {
assertFalse(rendezvous.resolveQuestion(W, "anyone?", W + "#1"),
"no blocked send means no primary to surface the question to");
}
@Test
void resolveCompletionTwiceIsANoOpTheSecondTime() {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(W);
assertTrue(rendezvous.resolveCompletion(waiter, "first scrape"), "the first completion resolves");
assertFalse(rendezvous.resolveCompletion(waiter, "second scrape"),
"a second completion on an already-resolved waiter returns false");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
assertEquals("first scrape", waiter.getNow(null).text(),
"the first resolution wins; the stored value is unchanged");
}
@Test
void resolveFailureTwiceIsANoOpTheSecondTime() {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(W);
assertTrue(rendezvous.resolveFailure(waiter, "first reason"), "the first failure resolves");
assertFalse(rendezvous.resolveFailure(waiter, "second reason"),
"a second failure on an already-resolved waiter returns false");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
assertEquals("first reason", waiter.getNow(null).text(),
"the first resolution wins; the stored value is unchanged");
}
@Test
void closeAskRemovesTheTurn() {
Rendezvous.AskTicket t = rendezvous.openAsk(W);
rendezvous.closeAsk(t.turnId());
assertNull(rendezvous.askSession(t.turnId()), "a closed ask is forgotten");
assertFalse(rendezvous.answerAsk(t.turnId(), "late"), "a closed ask can no longer be answered");
}
}
@@ -0,0 +1,304 @@
package dev.ltms.bridged.msg;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.Collections;
import java.util.List;
import java.util.Map;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for {@link ReplyPushLoop}: decision logic, nudge injection, idempotency,
* bounded reminders, and stop conditions.
*
* <p>Uses a {@link RecordingHerdrClient} that synchronizes access to its call list so the
* scheduler thread and test thread never have memory ordering issues. The {@code decide()}
* tests use a simple client with no concurrency concern.
*/
class ReplyPushLoopTest {
private static final String PRIMARY = "term_primary";
private static final String WORKER = "term_worker";
private static final ObjectMapper MAPPER = new ObjectMapper();
private PrimaryRegistry registry;
private AgentControl agents;
private InMemoryReplyInbox inbox;
private ScheduledExecutorService scheduler;
@BeforeEach
void setUp() {
registry = new PrimaryRegistry(PRIMARY);
inbox = new InMemoryReplyInbox();
inbox.own(WORKER); // CB-520: the inbox only peeks/acks targets it owns
scheduler = Executors.newSingleThreadScheduledExecutor();
}
@AfterEach
void tearDown() {
scheduler.shutdownNow();
}
// --- decide() logic ------------------------------------------------------------------------
@Test
void decideWithoutPrimaryIsStop() {
agents = agentWithStatus("idle");
var loop = new ReplyPushLoop(
new PrimaryRegistry(null), agents, inbox, scheduler, 5, 100);
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(WORKER, 0));
}
@Test
void decideWithEmptyInboxIsStop() {
agents = agentWithStatus("idle");
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
}
@Test
void decideAtCapIsStop() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100).decide(WORKER, 2));
}
@Test
void decideUnderCapWithInjectablePrimaryIsInject() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
}
@Test
void decideUnderCapWithBlockedPrimaryIsInject() {
agents = agentWithStatus("blocked");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
"BLOCKED is injectable");
}
@Test
void decideUnderCapWithDonePrimaryIsInject() {
agents = agentWithStatus("done");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
"DONE is injectable");
}
@Test
void decideUnderCapWithBusyPrimaryIsWaitBusy() {
agents = agentWithStatus("working");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
}
@Test
void decideUnderCapWithUnknownPrimaryIsWaitBusy() {
agents = agentWithStatus("unknown");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
}
@Test
void decideStopsAfterInboxIsEmptied() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
inbox.ack(WORKER, "m1");
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
}
// --- onReplyQueued integration -------------------------------------------------------------
@Test
void injectablePrimaryCausesExactlyOneNudge() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
loop(1, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"one nudge (1 agent.prompt call) should have been sent");
// Exactly one nudge = exactly 1 agent.prompt call (it submits itself)
assertEquals(1, rec.sendCount());
assertTrue(rec.sentParams().stream()
.anyMatch(e -> e.getValue().toString().contains("bridge_poll")),
"nudge text should contain bridge_poll");
}
@Test
void onReplyQueuedIsIdempotentPerTarget() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
var loop = loop(1, 100);
loop.onReplyQueued(WORKER);
loop.onReplyQueued(WORKER); // second call — should be a no-op
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"expected exactly one nudge (1 prompt)");
Thread.sleep(200);
assertEquals(1, rec.sendCount(),
"second onReplyQueued must not trigger another nudge");
}
@Test
void sendsUpToCapThenStops() throws Exception {
int cap = 2;
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
rec.sendLatch = new CountDownLatch(cap);
loop(cap, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(5, TimeUnit.SECONDS),
cap + " nudges (" + cap + " prompts) should have fired");
Thread.sleep(300);
assertEquals(cap, rec.sendCount(),
"exactly " + cap + " agent.prompt calls (cap=" + cap + ")");
}
// --- nudge format --------------------------------------------------------------------------
@Test
void nudgeFormatIsCorrect() {
String nudge = ReplyPushLoop.NUDGE_FORMAT.formatted(WORKER, WORKER);
assertTrue(nudge.contains("Worker term_worker"));
assertTrue(nudge.contains("bridge_poll(target=term_worker)"));
}
// --- metrics (CB-512) ----------------------------------------------------------------------
@Test
void successfulNudgeIncrementsDelivered() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
Metrics metrics = new Metrics();
loop(1, 50, metrics).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"one nudge (1 agent.prompt call) should have been sent");
// The delivered count is bumped on the scheduler thread right after the send that releases
// the latch — settle briefly so the counter is published before we read it.
Thread.sleep(200);
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"),
"a successfully sent nudge must count as delivered");
}
@Test
void reminderCapIncrementsExhausted() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
Metrics metrics = new Metrics();
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100, metrics).decide(WORKER, 2));
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
"hitting the reminder cap must count as exhausted");
assertEquals(0, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"));
}
// --- helpers -------------------------------------------------------------------------------
private ReplyPushLoop loop() {
return loop(5, 100);
}
private ReplyPushLoop loop(int maxReminders, long backoffMs) {
return new ReplyPushLoop(registry, agents, inbox, scheduler, maxReminders, backoffMs);
}
private ReplyPushLoop loop(int maxReminders, long backoffMs, Metrics metrics) {
return new ReplyPushLoop(registry, agents, inbox, scheduler, maxReminders, backoffMs, metrics);
}
private static AgentControl agentWithStatus(String status) {
return new AgentControl(new FakeHerdrClient(status));
}
/** Non-recording (single-threaded) fake — safe for decide() tests. */
private static final class FakeHerdrClient implements HerdrClient {
private final String agentStatus;
FakeHerdrClient(String agentStatus) {
this.agentStatus = agentStatus;
}
@Override
public JsonNode call(String method, Object params) {
if ("agent.get".equals(method)) {
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", PRIMARY)
.put("agent_status", agentStatus));
}
return MAPPER.createObjectNode();
}
@Override
public void close() {
}
}
/**
* Thread-safe recording fake that counts agent.prompt calls (protocol 19: one nudge = one
* prompt). Uses synchronized access so the scheduler thread and test thread never race.
*/
private static final class RecordingHerdrClient implements HerdrClient {
private final List<Map.Entry<String, Object>> calls =
Collections.synchronizedList(new ArrayList<>());
volatile CountDownLatch sendLatch = new CountDownLatch(1);
@Override
public JsonNode call(String method, Object params) {
if ("agent.get".equals(method)) {
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", PRIMARY)
.put("agent_status", "idle")); // recording double is always injectable
}
if ("agent.prompt".equals(method)) {
calls.add(Map.entry(method, params));
sendLatch.countDown();
}
return MAPPER.createObjectNode();
}
long sendCount() {
return calls.size();
}
List<Map.Entry<String, Object>> sentParams() {
return List.copyOf(calls);
}
@Override
public void close() {
}
}
private static RecordingHerdrClient recordingClient() {
return new RecordingHerdrClient();
}
}
@@ -0,0 +1,183 @@
package dev.ltms.bridged.placement;
import org.junit.jupiter.api.Test;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for the placement policies. They run with no herdr and no launcher — pure selection
* logic exercised through the descriptor type so CB-308 host expansion will not need to rewrite
* these assertions.
*/
class PlacementPolicyTest {
private static Function<String, Integer> noSessions() {
return name -> 0;
}
private static PlacementContext ctx(List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable) {
return new PlacementContext("b", candidates, liveCount, unreachable);
}
private static PlacementContext ctx(List<PlacementCandidate> candidates,
Function<String, Integer> liveCount) {
return ctx(candidates, liveCount, Set.of());
}
@Test
void fixedReturnsDefaultEvenIfOtherProfilesExist() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = ctx(List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b")), noSessions());
assertEquals("b", policy.select(ctx).profile());
}
@Test
void fixedFallsBackToFirstCandidateWhenNoDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null,
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of());
assertEquals("a", policy.select(ctx).profile());
}
@Test
void fixedThrowsWhenNoProfilesAndNoDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null, List.of(), noSessions(), Set.of());
assertThrows(PlacementException.class, () -> policy.select(ctx));
}
@Test
void roundRobinCyclesThroughAvailableProfiles() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b"),
PlacementCandidate.profile("c"));
assertEquals("a", policy.select(ctx(candidates, noSessions())).profile());
assertEquals("b", policy.select(ctx(candidates, noSessions())).profile());
assertEquals("c", policy.select(ctx(candidates, noSessions())).profile());
assertEquals("a", policy.select(ctx(candidates, noSessions())).profile());
}
@Test
void roundRobinSkipsProfilesAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 2),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = Map.of("a", 2)::get;
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, liveCount)).profile());
}
}
@Test
void roundRobinThrowsWhenAllAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, 1));
Function<String, Integer> liveCount = Map.of("a", 1, "b", 1)::get;
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
@Test
void weightedAlternatesEvenlyWithEqualWeights() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.5f, null),
PlacementCandidate.profile("b", 0.5f, null));
int a = 0, b = 0;
for (int i = 0; i < 100; i++) {
String p = policy.select(ctx(candidates, noSessions())).profile();
if ("a".equals(p)) a++;
else if ("b".equals(p)) b++;
}
assertEquals(50, a, "equal weights should split 50/50");
assertEquals(50, b);
}
@Test
void weightedHoldsThreeToOneRatio() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.75f, null),
PlacementCandidate.profile("b", 0.25f, null));
int a = 0, b = 0;
for (int i = 0; i < 40; i++) {
String p = policy.select(ctx(candidates, noSessions())).profile();
if ("a".equals(p)) a++;
else if ("b".equals(p)) b++;
}
assertEquals(30, a, "0.75/0.25 should yield a 3:1 ratio over a multiple of 4");
assertEquals(10, b);
}
@Test
void weightedSkipsProfileAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = name -> "a".equals(name) ? 1 : 0;
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, liveCount)).profile());
}
}
@Test
void weightedThrowsWhenAllAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, 1));
Function<String, Integer> liveCount = name -> 1;
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
@Test
void weightedThrowsWhenAllUnreachable() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b"));
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, noSessions(), Set.of("a", "b"))));
assertTrue(e.getMessage().contains("unreachable"), e.getMessage());
}
@Test
void mixedExclusionMessageNamesBothReasons() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = name -> "a".equals(name) ? 1 : 0;
Set<String> unreachable = new HashSet<>();
unreachable.add("b");
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount, unreachable)));
assertTrue(e.getMessage().contains("1 at maxLoad"), e.getMessage());
assertTrue(e.getMessage().contains("1 unreachable"), e.getMessage());
}
@Test
void unknownPolicyNameThrows() {
assertThrows(IllegalArgumentException.class, () -> PlacementPolicies.fromName("random"));
}
}

Some files were not shown because too many files have changed in this diff Show More