Compare commits

...

41 Commits

Author SHA1 Message Date
Dai Ha 37b4031ca4 CB-553: enforce maxLoad on explicit-profile spawns (no cap bypass)
CI / build (pull_request) Successful in 56s
CI / contract (pull_request) Successful in 1m16s
2026-08-13 21:12:32 +02:00
ltms 91c9f981c5 Merge CB-548 send ownership and rendezvous hardening
CI / build (push) Successful in 1m16s
CI / contract (push) Successful in 1m17s
2026-08-13 19:44:02 +02:00
Dai Ha dd906526c0 CB-548: guard primary singleton to PRIMARY callers; open waiter before enqueue
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 53s
Fix two handler-level bugs found in PR #21:
- Only PRIMARY callers may update PrimaryRegistry.record (the legacy singleton 'primary'
  fallback for no-delegation inbox nudges). An architect SEND previously recorded its terminal
  as the fallback; the per-target delegation map does not cure the singleton. New
  BridgeMcp.recordPrimarySingleton uses the resolved role (caller.isPrimary()) — named leads
  (PRIMARY) still record, architects never do.
- MessageService.send now opens the rendezvous waiter BEFORE queueing delivery, fixing both the
  enqueue-before-open fast-reply race (a fast reply no longer orphans into the inbox) and
  callback-failure ordering: a throwing onAccepted (public callback) fails the send cleanly with
  no stale waiter and no queued, orphanable message.

Tests: architect SEND vs lead SEND primary-singleton regression; throwing onAccepted leaves no
stale waiter or queued orphan.
2026-08-13 19:42:48 +02:00
Dai Ha ec3001796a CB-548: record delegator ownership only when a send is accepted
Record PrimaryRegistry delegator ownership via a MessageService accepted-delivery
hook (won the session lock + queued delivery), never at bridge_send request time, so
a concurrent sender that times out BUSY cannot steal a live turn's reply routing.
Make Rendezvous.open atomic fail-if-present so a double open trips loudly instead of
replacing the waiter another send is blocked on. Answering a bridge_ask keeps the same
ownership (no rewrite). Adds ownership/rendezvous regression tests.
2026-08-13 19:42:48 +02:00
ltms 86cf4c285a CB-547b: opencode session identity — resume (-s) + post-hoc discovery (#18)
CI / build (push) Successful in 48s
CI / contract (push) Successful in 1m17s
2026-08-13 19:41:06 +02:00
Dai Ha cd18887b69 CB-547b: opencode session identity — resume (-s) + post-hoc session discovery
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m16s
2026-08-13 19:40:17 +02:00
ltms 4637c68295 CB-543: neutralize worktree-hostile tracked configs (#22)
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m14s
2026-08-13 19:39:54 +02:00
Dai Ha 0114bd1fa7 CB-543: neutralize tracked opencode.json and .autoenv in provisioned worktrees
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m26s
2026-08-13 19:30:19 +02:00
ltms 6dee84ca71 Merge CB-548 architect config and authorization foundation
CI / build (push) Successful in 57s
CI / contract (push) Successful in 1m17s
2026-08-13 18:33:11 +02:00
Dai Ha f004a0c654 CB-548: make duplicate-architect-slot detection top-level-only and depth-safe
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m13s
2026-08-13 18:31:29 +02:00
Dai Ha 6123576c68 CB-548: correct architect premise — profile-only slots, registry-owned bindings, dup-key rejection
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m6s
2026-08-13 18:05:53 +02:00
Dai Ha 21cfc09f8e CB-548: config-declared architect slots + Role.ARCHITECT authz
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m6s
2026-08-13 17:28:14 +02:00
ltms 509530e235 CB-547a: peer-neutral session-identity contract + Claude Code adapter (#17)
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m2s
2026-08-13 16:56:32 +02:00
Dai Ha d146a01422 CB-547a: peer-neutral session-identity contract + Claude Code adapter
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m31s
2026-08-13 16:43:47 +02:00
Dai Ha 91332723db Bump wiki to 8c63db5 — the CB-539/CB-542 docs
CI / build (push) Successful in 1m18s
CI / contract (push) Successful in 1m52s
Pairs with 244fbd9. The docs half lived on the submodule's own remote, so the
pointer bump is separate by necessity, not by preference.

The substantive part is not the new 'Run a worker on the subscription' entry but
the correction beside it: 'Give workers a toolchain' asserted flatly that an env:
entry cannot repoint a worker past the SubscriptionGuard. CB-539 made that false,
and the wiki went on claiming it until CB-542 closed the hole. The entry now
states the rule and its one exception together.
2026-08-13 15:32:48 +02:00
ltms 244fbd98a5 Merge CB-539 + CB-542: subscription profiles, with the env: bypass closed
CI / contract (push) Successful in 40s
CI / build (push) Successful in 50s
Verified by the lead in a clean worktree at 5afe8e1 rather than on the worker's
report: Tests run: 474, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS (main
was at 464).

The invariant this lands: there is no configuration in which a worker reaches an
Anthropic endpoint that no guard vetted. Closed at two layers — a fatal, profile-
naming refusal at config load, and a launcher-side strip so it holds for profiles
built in code that never passed validation.

The wiki half is a separate branch on the submodule's own remote; the pointer bump
follows as its own commit on main.
2026-08-13 15:31:51 +02:00
Dai Ha 5afe8e14d9 CB-542: close the subscription env: bypass, and fix the env: env-carries-anthropic docs
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 51s
2026-08-13 15:26:20 +02:00
Dai Ha 73f6b12dd8 CB-539: let a worker profile run on the subscription, explicitly
Add a per-profile subscription: true opt-in that lets a claude-code worker run on
the operator's Claude subscription when there is no off-subscription endpoint for it
(e.g. sonnet on ccs). When set, the launcher injects neither ANTHROPIC_BASE_URL nor
ANTHROPIC_AUTH_TOKEN and skips SubscriptionGuard's base_url requirement for that
profile only, logging a WARN naming the profile. subscription: true alongside a
baseUrl is refused as contradictory. The default (absent/false) keeps today's hard
refusal unchanged; every other profile stays allowlist-checked and SubscriptionGuard
is untouched.
2026-08-13 15:15:14 +02:00
Dai Ha 793f2e7157 Revert "Commit .autoenv" — it wedges every worker spawn
CI / build (push) Successful in 53s
CI / contract (push) Successful in 1m24s
Regression I introduced one commit ago, caught by two consecutive spawn failures
and reproduced end-to-end.

Tracking `.autoenv` means git checks it out into every provisioned worktree. A
worktree is a new path, and autoenv authorizes by path, so the file is always
unauthorized there — autoenv prints "[autoenv] Authorize this file? (y/n/d)" and
blocks on `read` (activate.sh:211-222). The pane's shell sits at that prompt, so
`ccs <profile>` never runs, the peer never becomes injectable, and CB-306's
readiness gate closes the pane after 20s. Symptom is a bare "spawn timed out";
nothing names autoenv, which is what made it worth writing down.

Confirmed rather than inferred: the peer lead's cb-537 worker, spawned BEFORE
f8182e4, has no `.autoenv` in its worktree and is still alive; a throwaway
worktree at HEAD reproduces the prompt on entry.

The file is still worth committing — the reasoning in f8182e4 stands, and it is
recoverable from there. What is missing is the other half: GitWorktrees already
neutralizes `.mcp.json` in a provisioned worktree ("worker tool surface is
launcher-mounted only"), and `.autoenv` needs exactly the same treatment for
exactly the same reason — a worker's environment is launcher-supplied, never
repo-supplied. Re-land it with that, tracked as CB-543.

Rejected the quicker fix of exporting AUTOENV_ASSUME_YES: it auto-approves
arbitrary repo-controlled shell code in every spawned worker, which is a worse
trade than one unset variable.
2026-08-13 15:07:59 +02:00
Dai Ha cf48983046 Merge CB-538: opencode peers launch with --auto
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m30s
A spawned peer has no human at its pane, so an approval prompt is not a pause —
it is a wedge. The agent stops, looks identical to a legitimate mid-turn wait,
and can never reach its bridge_reply, so the delegation dies silently and the
lead learns nothing until the timeout.

Unconditional rather than a per-profile knob, which is the right call: there is
no configuration under which a bridge-spawned opencode worker WANTS to block on
an approval it has no way to answer. opencode's help calls --auto 'dangerous!',
and that warning is written for a human at a terminal; the blast radius here is
already bounded by the layer above — a worker runs in its own git worktree, on
its own branch, off-subscription, and cannot merge. The lead is the gate.

Reviewed by me rather than fanned out: 24 lines across one method and its two
tests, below the threshold where a reviewer pass pays for itself.

argvWithModel is re-signatured to take composed argv instead of building it,
so the two flag-appenders compose rather than each owning construction.
2026-08-13 15:03:59 +02:00
Dai Ha d038516750 Bump wiki to 0a21b49 — multi-lead arc, and opencode's secret supply
CI / build (push) Successful in 53s
CI / contract (push) Successful in 1m48s
Moves the submodule pointer 4320c1c -> 0a21b49, matching the docs to the code in
bc13b8e rather than leaving chapter 11 describing a fleet with one lead in it.

  7b5381b  11, 7: CB-530..536 — leaders:, leadScan:, lead-to-lead messaging,
           lead deliverability, leads in bridge_list; 7's copy of the portable
           block re-synced byte-identically
  0a21b49  12: where opencode's secrets actually come from, and why {env:...}
           is green under Claude Code and broken in a terminal

Both are already pushed to the wiki's own remote, so this pointer resolves for
anyone who clones — the ordering that matters, and the reason the wiki went
first.

This is the pointer only. `wiki/` is a submodule with its own remote and its
content is never committed here; bumping the recorded SHA in its own commit is
how this repo has always recorded a docs update (see "bump wiki to 4320c1c").
2026-08-13 10:44:38 +02:00
Dai Ha f8182e4514 Commit .autoenv — the loader that keeps Claude Code on one secret store
This file has always been designed to be committed and says so in its own header;
it simply never was, so every clone of this workspace has been reconstructing it
by hand or duplicating tokens instead.

It holds no secret. It reads `.secrets/` (gitignored, 0600) and exports three
variables, because Claude Code expands `${VAR}` in `.mcp.json` from the *process*
environment and cannot read a file — so without it, CONTEXT7_TOKEN and the gitea
pair must be duplicated as literals in `.claude/settings.local.json`. opencode
needs none of this: it reads `.secrets/` directly via `{file:...}`.

Verified before committing that no value appears in it, only the three names and
the paths they are read from.

Safe in a worktree by construction: a worktree receives tracked files only, so
`.secrets/` is absent there and the whole block is skipped rather than failing.
Workers are fed by the launcher's env instead — which is where the name mismatch
documented in wiki chapter 12 (GITEA_TOKEN vs GITEA_ACCESS_TOKEN) has to be
reconciled.
2026-08-13 10:44:21 +02:00
Dai Ha bc13b8e92c CB-530..536, CB-538 groundwork: leads as peers, and a fleet that can find itself
Two leads now work as peers rather than one primary plus workers. The arc:

  CB-530/531  lead identity: `leaders:` names panes, `leadScan:` discovers them by
              tab label (LeadTabScanner, TTL-cached, worker spaces excluded).
  CB-532      leads can message each other AND be answered. Principal.leader now
              carries its terminal, so ownsSession() can be true for a lead; the
              "and you must be a worker" conjunct beside it protected nothing.
              Retires `primary:` — reply nudges follow the delegating lead, a
              binding recorded at bridge_send where both halves are known.
  CB-533      ClaudeCodeLauncher passes --model. argv is usually a wrapper
              (`ccs <profile>`) that re-exports its own model family, so
              ANTHROPIC_MODEL alone was silently overruled.
  CB-534      a lead is deliverable. The CB-113 readiness gate only opened for
              terminals in WorkerPresence, which only workers ever enter, so every
              lead->lead send waited out the ~60s grace and failed having never
              been typed. The gate guards a *spawned* peer's boot window; a lead
              is never spawned.
  CB-535      bridge_list returns `leads` alongside `workers`, with `self` on the
              caller's row. An empty worker roster no longer reads as "no peers".
  CB-536      CLAUDE.md: lead<->lead is coordinate-only, never sideways delegation.
              Propagated byte-identically to wiki/7-Use-Cases.md.

MIXED PROVENANCE — recorded deliberately rather than hidden. This tree also carries
in-progress CB-537 (context separation) authored by the peer lead gpt-sol-5.6 and
its worker: Capability.CONTEXT_RESET, SessionManager.clearAfterTurn, and the
Injector/TurnListener/CompletionResolver/launcher changes around it. That work was
done in this shared working tree rather than a worktree, and is entangled with the
above in BridgedConfig.java, Bridged.java and ClaudeCodeLauncher.java, so neither
lead could stage its own half without sweeping in the other's. Committing the whole
green state is the honest resolution; the peer branches from here.

Note for whoever picks CB-537 up: the design in this commit is SUPERSEDED. Both
leads agreed to replace the global `clearAfterTurn` boolean with per-delivery
policy (inherit|fresh|thread) applied PRE-delivery, because a post-turn reset races
by construction — Injector.onStatus clears awaitingCompletion and dequeues the next
message in the same tick. `fresh` is also a correctness guarantee, so an adapter
without a reset capability must refuse it rather than log a no-op.

mvn clean install: Tests run: 464, Failures: 0, Errors: 0, Skipped: 0. BUILD SUCCESS.
2026-08-13 09:38:24 +02:00
Dai Ha fbe79258bf CB-538: opencode peers launch with --auto 2026-08-13 07:19:19 +02:00
Dai Ha ef1e014b41 CB-527: ship the bridge as an installable Claude Code plugin
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m21s
The orchestration contract had no distributable form. Every consuming project
hand-copied a block of CLAUDE.md and hand-wrote an .mcp.json, and we keep a
script whose only job is to notice those copies drifting apart. A plugin is
versioned, installed once, and updates in place.

Ships no credentials, deliberately: every secret is referenced by environment
variable NAME and the value never enters a file, which is what makes the
artifact safe to publish. The setup skill states the two rules that are easy to
get wrong for the right-sounding reasons — the PR token must not be able to
merge (a worker opens, the primary gates), and ANTHROPIC_BASE_URL must never be
set by setup, because mounting the bridge must not move a session off
subscription.

The plugin root is plugin/, not the repo root. An installed plugin's .mcp.json
is a committed file, while this repo's root .mcp.json is local-only and
--skip-worktree; rooting the plugin at the repo would commit the primary's IDE
servers and hand them to every worker — the exact failure CB-525 exists to
prevent.

Scope is client-side setup only. herdr and bridged stay separate services with
their own lifecycles, and the skill refuses to install them rather than guess.
It also refuses to accept /healthz as proof: health reports only that the daemon
can reach herdr, and CB-521 showed it staying green while every spawn failed, so
verification ends with a real spawn.

Both manifests pass `claude plugin validate --strict`.
2026-08-10 19:55:17 +02:00
Dai Ha e3c8393d1b CB-527: retire the ollama profile from the worker choice
The ollama backend is decommissioned, so the example config stops pointing
readers at a dead host and the guard allowlist stops carrying an entry with
no profile behind it — a stale entry there is dead permission, and that list
is the only thing keeping a worker off the primary's subscription.

The second illustrative profile survives as gx11: the example exists to show
`placement: weighted` having something to choose between, and a one-profile
example would quietly stop demonstrating that.

It also moves the CB-523 auto-compact override onto the surviving profile.
That guard had been attached to `ollama` alone, so retiring the profile would
have removed the fleet's only protection against the failure it was written
for — a worker whose prompt is rejected before auto-compact ever fires. The
window belongs on every profile, not on whichever one happened to hit it.
2026-08-10 19:55:17 +02:00
kevin 379e03f9d0 CB-308: fold in the adversarial review; bump wiki to 4320c1c
CI / build (push) Successful in 59s
CI / contract (push) Successful in 1m16s
Four new resolved decisions (§7.7-7.10): turn state split by where the signals
are, with ABANDONED explicitly belt-and-braces over the waiter timeout; the
dual ack model with spawn idempotence by construction (gid stored IN the herdr
pane — the load-bearing detail of the no-ledger position — plus an in-flight
reservation for redelivery during a slow spawn); enforced publish semantics
(confirms + mandatory on a separate channel, return-before-confirm caveat);
queue lifecycle = session lifecycle with .v2 names for the redeclare hazard.

§8 reworked: global id scheme resolved and moved up; control authorization
sharpened into the hard gate on U4 (per-host allowlist beside the peer keys);
key distribution/rotation added. New §9: implementation order, each step
verifiable single-host, U4 gated, U8 last.

Wiki pointer bumped to 4320c1c (chapter 10 same-pass changes).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 22:32:00 +07:00
kevin e13921aa8a CB-308: resolve the cross-host design review; bump wiki to 36bb865
CI / build (push) Successful in 54s
CI / contract (push) Successful in 1m5s
Six decisions recorded in §7, replacing the matching open questions: signed
messages (identity from the key, extending identity-from-connection across the
broker), target-host-owned profiles advertised via presence, repo provisioning
by pinned forge clone, live-only asks with TTL + TOO_LATE notice, spawn-id
dedup on the target, and broker-outage semantics (local unaffected, remote
fails fast, gateway stays soft-state). §8 keeps what is genuinely still open,
with control *authorization* now separated from the resolved *authenticity*.

The broker-level half (U8 broadcast, exclusive consumers, inbox caps, TLS,
schema version, trace id) lands in wiki chapter 10 §10 — pointer bumped
(also picks up 1710a77, chapter 11 Features).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 21:54:30 +07:00
kevin 8a549d8610 Release 1.0.0 — one leader, one host, complete
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m36s
Bump bridged to 1.0.0 and add the release notes: the single-leader,
single-host scope is closed — gateway, lifecycle, two-way delivery,
pluggable peers, auth/authz/audit, supervision, CI. Cross-host
federation (CB-308) is the next major line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 20:58:06 +07:00
Dai Ha 84081b2bd8 CB-526: make a shipped capability undocumentable-by-accident
CI / contract (push) Successful in 49s
CI / build (push) Successful in 1m23s
CLAUDE.md already carries a mandatory before-done checklist ("the prompt is part
of the product"). It covered the instruction surface but not the operator-facing
one, and the result was measurable: CB-506 through CB-525 shipped without a
single wiki mention, while the Roadmap went on claiming Stage 5 was finished.

Discipline is what already failed, so this rides the existing gate rather than
adding a new habit to remember: one more row, firing when a change touches
anything an operator can use, configure, or observe. The row names where the
other two kinds of change go too (contracts to Implementation, coverage to the
Roadmap), so "nothing to document" is a decision the table makes rather than a
default you fall into.

Addendum-only — the canonical block is untouched and still byte-identical to the
wiki template (verified).
2026-08-09 19:38:34 +02:00
Dai Ha defe3365c4 CB-519: re-point pane assertions at the protocol-19 coordinate
CI / build (push) Successful in 1m22s
CI / contract (push) Successful in 2m32s
Rebase integration only, no behaviour change. CB-519's tests named the pane the
pre-protocol-19 fake produced (w9:pW_n); upstream's herdr 0.8.0 port creates the
pane through tab.create and starts the agent into it, so the fake now reports
w9:pRoot_n. Five assertions were therefore counting closes of a pane that never
existed and reading 0.

mvn clean install: Tests run: 399, Failures: 0, Errors: 0 — BUILD SUCCESS
2026-08-09 05:59:14 +02:00
Dai Ha 4e6201ecd1 CB-525: isolate a worker's tool surface to what its launcher mounts
A worker in a provisioned worktree was inheriting the primary's MCP servers by
two independent routes: the repo commits a .mcp.json declaring the IDE servers,
so a fresh checkout mounts them, and the default parity overlay then copied the
primary's own copy over the top.

Those servers are bound to the primary's IntelliJ project, so every path they
hand back points into the primary's checkout. A CB-523 worker made all 59 of its
edits there while running `mvn -f bridged/pom.xml` against its worktree — every
build it ran was of code that did not contain its changes, and it passed. The
worker's own `ls` of the file it had "edited" returned "No such file".

GitWorktrees now neutralizes .mcp.json at provisioning: an explicitly empty
server map, --skip-worktree'd when tracked so it never reads as pending work a
worker might commit. Unconditional, because the overlay was only half the leak.
The bridge itself is unaffected — it reaches a worker through the launcher's
--mcp-config flag, not the project file, so bridge_reply still works.

- BridgedConfig: .mcp.json out of the default parity overlay
- GitWorktrees: isolateToolSurface() on add(), with the rationale in javadoc
- GitWorktreesTest: 4 real-git acceptance tests (2 fail if the call is removed)
- implementer skill: work from $PWD, and quote a green unpiped `mvn clean
  install` from the worktree as the acceptance criterion

mvn clean install: Tests run: 392, Failures: 0, Errors: 0 — BUILD SUCCESS
2026-08-09 05:57:30 +02:00
Dai Ha ccf50f950e CB-524: make worker placement order deterministic across JVM runs
The weighted policy breaks an exact-weight tie on candidate list order
(WeightedRoundRobinPolicy picks the first candidate with a strictly greater
score), and that list comes from CompositePeerLauncher.candidates(), which
iterates profileConfigs. Both that map and BridgedConfig.workerProfiles() were
built with Map.copyOf, whose iteration order is salted per JVM run — so the
"in definition order" contract candidates() documents was not held.

Two consequences. In production, a config with equal weights (ollama 0.5 /
gx10 0.5) placed its first worker on a profile chosen at random on every daemon
restart. In the suite, CompositePeerLauncherTest.failoverRetriesNextCandidate-
WhenProfileIsUnreachable failed roughly one run in four, because whether
profile "a" was tried first depended on the salt.

Preserve definition order at every layer: unmodifiable LinkedHashMap for
workerProfiles(), profileConfigs, and byProfile (which also feeds the
user-visible bridge_profiles listing). The tests build profile maps with an
ordered helper rather than Map.of, which is salted for the same reason.

Guarded by a pair of tests declaring the same two profiles in opposite order
and asserting opposite first attempts, so any order-scrambling implementation
must fail one of them. Verified by mutation: reverting profileConfigs to
Map.copyOf fails 8/8 runs (6 caught by the original test, 2 only by the new
reversed-order one); with the fix, 10/10 fresh JVMs pass, 388 tests green.
2026-08-09 05:57:30 +02:00
Dai Ha a7f0211e2f CB-518: weighted placement policy for worker spawns 2026-08-09 05:57:30 +02:00
Dai Ha 049ce4828c CB-520: split ReplyInbox into explicit own/release and publish halves 2026-08-09 05:57:12 +02:00
Dai Ha c1173346ef CB-521: make the AMQP contract test runnable locally and in CI 2026-08-09 05:56:45 +02:00
Dai Ha 7ace184fe6 CB-519: make PeerHandle.id() a host-unique opaque UUID, decoupled from the herdr pane id 2026-08-09 05:56:45 +02:00
Dai Ha 54b314ace5 CB-522: let the primary run inside a herdr pane
Caller identity resolved any loopback PID that mapped to a herdr pane as a
WORKER, and PaneLocator scans every pane -- not just bridged-spawned ones. A
primary running inside a herdr pane therefore classified itself as a worker and
was refused SPAWN/SEND/STOP, i.e. every orchestration verb it exists to call.

The failure is self-locking: PrimaryRegistry only learns the primary's terminal
from bridge_send/bridge_spawn, the exact calls being refused, so the learned
value can never bootstrap. Only an operator-set pin breaks the cycle.

CallerResolver now consults primary.terminal from config *before* the pane
lookup. Deliberately the pinned value only, never the learned one -- the learned
terminal is populated by the callers this method is itself classifying, so
trusting it would be circular. Config is operator input, never network input, so
this widens no attack surface; bridge_whoami and the authz gate still share one
resolution.

Fixing that exposed a second, older bug. BridgeMcp's context extractor forwards
the caller's terminal into markPresent on every MCP call, documented as "no-op
for the primary (null terminal)". WorkerPresence.markPresent honours that, but
PresenceBridge overrides it and forwards the same null into SessionManager.
onReady -> transitionByTerminal -> findByTerminal, which called
terminalId.equals(...) unguarded. It only reached the scan once the registry was
non-empty, so the primary's first spawn succeeded and every later call NPE'd
with an HTTP 500 -- and it would have fired for ANY primary not living in a
herdr pane, pinned or not.

findByTerminal is now total. That covers onReady, onDelivered, onTurnComplete
and onTurnFailed at once; a null id could never match a registered session
anyway, so "no match" is the honest answer rather than taking down an unrelated
tool call.

Also drops two dead pass-throughs on CallerResolver (cwdForPid, tokenMode) that
IDE inspections flagged -- callers use ConnectionIdentity and BridgedConfig.Auth
directly.

The example config now states that primary.terminal is REQUIRED, not just a
push-loop optimisation, when the primary shares a herdr pane.

mvn clean install: 360 tests, 0 failures. Verified live: daemon restarted on
this jar, bridge_whoami reports primary, and four concurrent worktree spawns --
the exact shape that NPE'd -- now all succeed.
2026-08-09 05:55:12 +02:00
kevin 224b344445 CB-522: resolve the pinned primary.terminal pane as the primary, not a worker
CI / build (push) Successful in 2m54s
A primary running INSIDE a herdr pane was resolved as a worker by the
pane-match rule and refused every orchestration tool — the exact lockout
bridge_whoami surfaced on this deployment. The CB-307 primary.terminal pin
always claimed to replace connection-derived identity but only fed the push
loop; it now short-circuits CallerResolver ahead of the pane→worker rule
(the pane mapping is as unforgeable as a worker's, so no credential needed,
even in token mode). bridged.example.yaml documents the block.

Also guard the presence bridge against the primary's null terminal: the MCP
context extractor marks presence on every request, and the first genuine
primary contact NPEd into the SPAWNING→READY transition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
2026-08-08 21:52:38 +07:00
kevin 0b28b4cb0f CB-521: port the herdr adapter to protocol 19 (herdr 0.8.0)
herdr 0.8.0 redesigned the agent API out from under the daemon: agent.start
now launches a supported kind INTO an existing pane, env/cwd move to pane
creation (tab.create / pane.split — the subscription-boundary seam now),
agent.send is replaced by agent.prompt (self-submitting) plus agent.send_keys
for the Enter nudge, and terminal ids are no longer valid agent.* targets.

- AgentControl: start(name, kind, args, paneId); prompt/send_keys delivery;
  cached terminal→pane target translation (invalidated on agent_not_found).
- WorkspaceControl: tab.create carries cwd+env; pane.split for legacy placement.
- HerdrPeerLauncher: the seed pane IS the worker pane (no drop step); retry
  agent.start while the seed shell boots (agent_pane_busy).
- FakeHerdr and the test suite model protocol 19 (unique seed panes, required
  kind/pane_id, prompt-based delivery); contract tests probe the seed shell
  instead of arbitrary-command agents, which protocol 19 removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
2026-08-08 21:52:25 +07:00
kevin 83129e165c Merge CB-518: state the primary's orchestration as an explicit, ordered flow
CI / build (push) Successful in 1m19s
Turns the primary's half of the bridge charter from a bullet list of
policies into a numbered 0-8 procedure, and splits delegated review out
of the merge step it used to sit beside. Wiki template kept byte-identical
by splicing; pointer bumped to 0c896eb in the same commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw1FosEt3Noix5GG9wJ2r
2026-08-04 22:37:17 +07:00
91 changed files with 7788 additions and 554 deletions
+18
View File
@@ -0,0 +1,18 @@
{
"name": "claude-bridge",
"description": "Tooling for orchestrating a fleet of delegated coding agents through the bridged MCP gateway.",
"owner": {
"name": "LTMS"
},
"plugins": [
{
"name": "claude-bridge",
"source": "./plugin",
"description": "Make a project bridge-ready: mount the bridged MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.1.0",
"author": {
"name": "LTMS"
}
}
]
}
+37 -9
View File
@@ -9,11 +9,13 @@ The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, ho
never merge, never commit `.mcp.json` or `wiki/`) is in **`CLAUDE.md` → Bridge communication →
Worker** and already applies. This skill is only the *implement-and-hand-off procedure*.
You run in an **isolated git worktree on your own branch** — a full peer of the primary (same
repo, `CLAUDE.md`, skills, MCP), differing in the model behind you and the branch you sit on.
You run in an **isolated git worktree on your own branch** — a full peer of the primary (same repo,
`CLAUDE.md`, skills), differing in the model behind you and the branch you sit on. Your MCP surface
is **only what your launcher mounted** (the bridge): the primary's IDE and forge servers are not
yours, and the worktree's `.mcp.json` is deliberately emptied so you cannot inherit them.
The worktree model is documented in [`docs/Worker-Git-Workflow.md`](../../../docs/Worker-Git-Workflow.md).
## 1. Confirm where you are
## 1. Confirm where you are — then never leave
Before touching anything:
@@ -26,13 +28,36 @@ git status # should be clean at the start
Do **all** work here, on this branch. Never `git checkout main`, never rebase onto or push to
`main`. The branch is your isolation — respect it.
**Every path you read, edit, or build is relative to that root.** Work from `$PWD`; if a tool, a
brief, or your own memory hands you an absolute path, check it starts with your worktree root
before you touch it, and stop if it doesn't. An absolute path pointing anywhere else is the
primary's checkout — editing there while building here means **every build you run is of code that
does not contain your changes**, and it passes while your work goes nowhere. This has happened:
a worker made all 59 of its edits in the primary's tree and never noticed.
```bash
test "$(git rev-parse --show-toplevel)" = "$PWD" || cd "$(git rev-parse --show-toplevel)"
```
## 2. Implement
- Implement exactly the scope the lead named. Keep the diff focused; note anything out of scope
in your reply instead of widening it.
- Match the surrounding code's style, naming, and idioms.
- Run whatever build/test you can — `mvn clean install` from the module root. Read its **full**
output; a piped `mvn ... | tail` hides failures.
**Acceptance criterion — a green build, quoted.** Your work is not done until this passes *inside
your worktree*:
```bash
cd "$(git rev-parse --show-toplevel)/bridged" && mvn clean install
echo "exit=$?"
```
Read its **full** output — never pipe it through `tail`/`head`/`grep`, which hide a failure behind
a zero exit. Then quote the real `Tests run: … Failures: … Errors: …` line and the
`BUILD SUCCESS`/`FAILURE` verbatim in your reply. If it does not go green, say so with the actual
error; a failing build honestly reported is a usable result, a claimed-green one is not. You have
no IDE MCP tools, so `mvn` is your only verification — never claim a check you had no way to run.
## 3. Commit
@@ -41,7 +66,8 @@ git add <the files you changed> # explicitly — never `git add -A` / `git a
git commit -m "<ticket>: <clear one-line summary>"
```
`.mcp.json` will show as modified. Leave it — it is `--skip-worktree` and not yours to commit.
`.mcp.json` is neutralized and `--skip-worktree` in your worktree — never `git add` it, and never
"restore" it from the primary's copy. Same for `wiki/` (a submodule with its own remote).
## 4. Push
@@ -84,8 +110,9 @@ The reply is the entire handoff; the lead cannot see your terminal.
```
PR: <html_url from step 5, or "not created: <reason>" + branch name>
branch: <your branch>
files: <the files you changed>
tests: <what you ran and its REAL result — or "not run: <why>">
root: <git rev-parse --show-toplevel — proves you worked in your own worktree>
files: <worktree-relative paths you changed>
build: <the verbatim "Tests run: …" and BUILD SUCCESS/FAILURE lines — or "not run: <why>">
summary: <2-3 lines: what you implemented and any caveat the reviewer needs>
```
@@ -97,7 +124,8 @@ sequenceDiagram
participant G as git / gitea
L->>I: delegated task (you are in a worktree on your branch)
I->>I: implement + build/test here
I->>I: "implement here — every path under $PWD"
I->>I: "mvn clean install in this worktree, unpiped, until green"
I->>G: git commit (never .mcp.json / wiki)
I->>G: git push -u origin HEAD
I->>G: POST /pulls (GITEA_TOKEN) — open PR to main
+172
View File
@@ -0,0 +1,172 @@
---
name: port-to-opencode
description: Make an OpenCode session a first-class participant in a Claude Code workspace — instructions, MCP servers, and secrets — without duplicating config. Use when onboarding opencode to a project that already has CLAUDE.md and .mcp.json, or when an opencode peer needs the same tools and rules as the Claude session.
---
# Porting a Claude Code workspace to OpenCode
**The headline: there is almost nothing to port.** OpenCode reads `CLAUDE.md` natively. The only
artifact you create is one `opencode.json` mapping MCP servers. Do not translate instructions, do
not generate a second rules file, and do not install a sync tool — every one of those makes the
workspace worse.
Everything below was verified against `opencode 1.18.16` and the OpenCode docs.
## 1. Know what you get for free
OpenCode's instruction search order:
```
1. walking up from cwd: AGENTS.md , then CLAUDE.md
2. global: ~/.config/opencode/AGENTS.md
3. Claude Code global: ~/.claude/CLAUDE.md (unless disabled)
```
*"The first matching file wins in each category."*
Consequences that decide the whole procedure:
- **A project `CLAUDE.md` is already read.** No port needed.
- **Your user-level `~/.claude/CLAUDE.md` is already read too.** Global preferences carry over.
- **An `AGENTS.md` in the repo SHADOWS `CLAUDE.md`.** If one exists from a previous Codex port,
**delete it** — otherwise opencode reads the stale translated copy instead of the real rules.
This is the single most likely way to get this wrong.
## 2. Create `opencode.json` for MCP servers only
Project config lives at `opencode.json` in the repo root; the global one is
`~/.config/opencode/opencode.json`. **Configs merge, they do not replace** — so machine-local
servers belong in the global file and shared ones in the project file.
Map each entry from `.mcp.json`:
| `.mcp.json` | `opencode.json` |
|---|---|
| `"type": "http"` / `"sse"` | `"type": "remote"`, `"url"` |
| `"type": "stdio"` | `"type": "local"`, `"command": ["bin", "arg"]` |
| `"command"` + `"args"` | single `"command"` array |
| `"env"` | `"environment"` |
| `"headers"` | `"headers"` |
```json
{
"$schema": "https://opencode.ai/config.json",
"instructions": ["CLAUDE.md"],
"mcp": {
"bridged": { "type": "remote", "url": "http://127.0.0.1:8765/mcp", "enabled": true },
"context7": { "type": "remote", "url": "https://example.dev/mcp", "enabled": true,
"headers": { "Authorization": "Bearer {env:CONTEXT7_TOKEN}" } },
"gitea": { "type": "local", "command": ["gitea-mcp", "-t", "stdio"], "enabled": true,
"environment": { "GITEA_ACCESS_TOKEN": "{env:GITEA_ACCESS_TOKEN}" } }
}
}
```
Set `instructions` explicitly even though `CLAUDE.md` is found anyway — the fallback only applies
while no `AGENTS.md` exists, and being explicit survives someone adding one later.
## 3. Reference secrets, never embed them
OpenCode substitutes at load time, in both `headers` and `environment`:
```
{env:VARIABLE_NAME} value from the environment
{file:~/.secrets/token} value read from a file
```
**No credential ever belongs in `opencode.json`.** With `{env:…}` there is no reason to write one,
which is what makes this file safe to commit — and it must be committable, because a peer running
in a git worktree receives tracked files only.
**But `{env:…}` reads OpenCode's *process* environment — and OpenCode has no env store of its
own.** A Claude Code `env` block in `~/.claude/settings.json` or `.claude/settings.local.json` does
**not** reach it: those files are Claude Code's, and the variables exist only inside processes
Claude Code spawned. Verify with a clean login shell, not the shell your agent hands you:
```bash
env -u MY_TOKEN zsh -lc 'echo "${MY_TOKEN:-NOT IN PROFILE}"'
```
So `{env:…}` only works if something puts the variable in the environment first. Pick the supply
route by who launches opencode:
| Launcher | Route |
|---|---|
| a human, from a terminal | `{file:…}` — see below |
| a spawner (bridge, CI, IDE) | `{env:…}`, with the spawner injecting the variable |
For the human case prefer **`{file:…}` with a workspace-relative path**, kept in a gitignored
`.secrets/` directory beside `opencode.json`:
```json
"headers": { "Authorization": "Bearer {file:.secrets/api-token}" }
```
Verified: opencode resolves relative `{file:}` paths against the project root, so this needs no
shell setup at all — no rc export leaking the secret to every process, no direnv dependency.
**The catch, and state it out loud:** `.secrets/` is gitignored, so a peer running in a git worktree
does **not** get it — worktrees receive tracked files only, the same rule that makes `opencode.json`
itself worth committing. Spawned peers must therefore be fed through `{env:…}` by whatever launches
them. Check the variable *names* match: a spawner often injects under a different name than your
shell uses, and the config has no fallback.
## 4. Do not port machine-local MCP servers
IDE indexes, language servers, editor bridges — anything bound to *your* checkout — stay out of the
project file. Put them in `~/.config/opencode/opencode.json` if you want them personally.
A committed project config reaches every worktree. A worker that mounts servers whose paths point
into the primary's checkout will edit the primary's files while building in its own — every build
passes, every change lands in the wrong tree.
Port the servers the work needs. Leave the rest.
## 5. Verify against the running agent, not the file
A file on disk proves nothing about what the agent loaded.
```bash
opencode run "In one line: state a rule from this project's instructions."
```
The answer must reflect the actual `CLAUDE.md`. If it answers generically, the instructions did not
reach the model and everything after this is built on sand.
Then confirm the tools are mounted:
```bash
opencode mcp list
```
**"connected" does not mean "working".** A stdio server with a missing credential still completes
the MCP handshake and reports green; only a real tool call reveals it. Verified: `gitea` showed
`✓ connected` with no token, then failed the first call with `token is required`. A remote server
is more honest (`⚠ needs authentication`), but do not rely on that difference — **exercise one
authenticated tool per server**:
```bash
opencode run "Call <server>'s <tool>. Report the result or the exact error. One line."
```
Check too that no server you deliberately withheld is present.
## 6. Report
State what you changed, which servers crossed and which you withheld and why, and quote the
verification answer verbatim. If any server failed to connect, say so plainly — a partially mounted
peer is worse than a missing one, because it looks configured.
## Gotchas
- **A leftover `AGENTS.md` silently wins over `CLAUDE.md`.** Check for one before anything else.
- **`opencode.json` is merged, not overridden** — a global entry and a project entry with the same
server name both matter; keep names distinct unless you intend to layer them.
- **`enabled: false`** turns a server off without deleting its config — prefer it over removal when
you may want the server back.
- **OAuth-based servers** store tokens in `~/.local/share/opencode/mcp-auth.json` after
`opencode mcp auth <server>`; that is machine state, never config to commit.
- **Skills and subagents do not port.** OpenCode uses its own agent markdown under
`.opencode/agents/`; `.claude/skills/**` is not read. If a delegation brief tells a peer to load a
skill by name, that instruction has no effect on an opencode peer — spell the procedure out in the
brief, or author the equivalent agent file.
+52
View File
@@ -53,3 +53,55 @@ jobs:
grep -qE "Failures: [1-9]|Errors: [1-9]" "$f" && { echo "===== $f ====="; cat "$f"; }
done
exit 0
# CB-521 — actually run the AMQP contract test in CI, against a REAL broker. The broker is a
# RabbitMQ SERVICE CONTAINER, not Testcontainers-with-Docker: the runner image has no Docker, so
# AmqpReplyInboxContractTest reads AMQP_URI (set below to the service's network alias) and binds
# straight to it — no Docker, no skipped tests. This separation (build job hermetic and
# Docker-free; contract job broker-provided) is deliberate — see the default-excludes/contract
# profiles in bridged/pom.xml. `setup-java` provides the JDK only; Maven is installed separately,
# exactly as in the build job above.
contract:
runs-on: ubuntu-latest
services:
rabbitmq:
image: rabbitmq:3.13 # same AMQP 0-9-1 engine the local Testcontainers fixture uses
env:
RABBITMQ_DEFAULT_USER: guest
RABBITMQ_DEFAULT_PASS: guest
ports:
- 5672:5672
env:
# Service containers are reachable from the job by their network alias on their internal port.
AMQP_URI: amqp://guest:guest@rabbitmq:5672
steps:
- uses: actions/checkout@v4
- name: Set up JDK 25
uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '25'
cache: maven
- name: Install Maven
run: |
apt-get update && apt-get install -y --no-install-recommends maven
mvn -version
# The `contract` profile clears the default-excludes group, so the @Tag("contract") AMQP test
# runs against the RabbitMQ service container (AMQP_URI). Pinned to the one contract test to
# avoid re-running the unit suite already covered by the `build` job.
- name: Contract tests
working-directory: bridged
run: mvn -B -Pcontract test -Dtest=AmqpReplyInboxContractTest
- name: Failing test output
if: failure()
working-directory: bridged
run: |
for f in target/surefire-reports/*.txt; do
[ -f "$f" ] || continue
grep -qE "Failures: [1-9]|Errors: [1-9]" "$f" && { echo "===== $f ====="; cat "$f"; }
done
exit 0
+12
View File
@@ -0,0 +1,12 @@
# Workspace-scoped secrets for tools that read no settings cascade of their own
# (opencode resolves these via {file:.secrets/...} in opencode.json).
.secrets/
# Settings backups inherit the env block — and secrets with it.
.claude/settings.local.json.bak*
# Daemon runtime artefacts. bridged appends its log wherever it is launched from, so both the
# repo root and bridged/ collect one; neither belongs in git.
bridged.out
bridged/bridged.out
logs/
+44 -3
View File
@@ -41,8 +41,9 @@ nothing. Fail toward the recoverable error.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/send/drain are primary-only; reply/ask are
worker-only-and-only-as-itself. A call outside your role is refused, not queued.
cannot act as another session. Spawn/stop/send/drain are lead-only; reply/ask are
only-as-itself — any peer may answer for its own pane, and for no other. A call outside your
role is refused, not queued.
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
re-send because a call looks slow — the bridge delivers when the peer is `idle`/`blocked`.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
@@ -101,13 +102,42 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a worker | `bridge_spawn{profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` · one worker's state: `bridge_status{sessionId}` |
| See the fleet | `bridge_list` → `leads` (your peers) + `workers` · one peer's state: `bridge_status{sessionId}` |
| Delegate (blocking) | `bridge_send{sessionId, content}` |
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
| Answer a worker's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** | `bridge_send{sessionId: <their terminal>, content}` — `bridge_list` → `leads` reports it. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `bridge_reply{content}` — the one case a lead replies |
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
| Tear down | `bridge_stop{paneId}` |
### Lead ↔ lead — coordinate, never delegate
`bridge_list` returns `leads` alongside `workers`; your own row carries `self: true`. Every other row
is a peer — an orchestrator with its own context, its own workers, and its own judgment. An empty
`workers` array means no workers are spawned; it says nothing about peers.
**A lead never assigns a task to another lead.** Work goes to workers — only ever downward, never
sideways. Sending a peer a brief with acceptance criteria is a category error: a brief is a worker's
artefact, and a peer is not yours to task. If a unit needs doing and it falls in your area, spawn a
worker and delegate it yourself; if it falls in the peer's area, say so and let the peer assign it.
The traffic between leads is coordination and nothing else:
1. **Divide the map, not the work.** Agree who owns which area, then each of you assigns inside your
own. Split by **context ownership** — whoever already holds the context owns that area — and say
who takes what, in one message, before either of you starts. Two leads silently working the same
unit is the failure mode here, and neither notices until the merge.
2. **Share findings, hazards, and corrections.** What you have already discovered, what broke, what
the next person will trip on. This is the traffic that actually pays for the channel: it costs one
message and saves a peer a rediscovery.
3. **Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
against the code, and re-run the build. A peer's correction gets the same treatment — right or
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
Being messaged by a peer does not make you its worker: answer with `bridge_reply`, and push back on
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
you.
### Worker — the turn contract
1. **Load the playbook skill the lead named** before doing anything else.
@@ -143,6 +173,8 @@ charter, not here.
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
`reviewer` (scoped review → one structured finding). Name one in every delegation.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace).
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
@@ -168,6 +200,15 @@ Before you call any work done, check the row that matches what you touched:
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
| **anything an operator can use, configure, or observe** — an MCP tool, a `bridged.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
That last row is not bookkeeping. Chapters 1–10 answer *how is this built* and *why this way*;
none of them has a home for *what can it do and how do I turn it on*, so for twenty tickets a
shipped capability landed nowhere and the Roadmap went on claiming the stage was finished. The
*why* line is the one that matters — without it a decision gets re-litigated from scratch a month
later. Internal contract changes go to `wiki/9-Implementation.md` instead; test and coverage work
is a Roadmap line. A change that touches none of the three earns no entry, and that is a normal
outcome rather than an omission.
Then **propagate**: the block in this file and the template in the wiki
([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
+89 -6
View File
@@ -31,6 +31,59 @@ bind:
# mode: token
# tokenEnv: BRIDGED_API_TOKEN
# Optional pinned primary terminal (CB-307). Names the herdr pane the PRIMARY itself runs in:
# a caller whose connection maps to this pane resolves as the primary (no credential needed —
# the pane mapping is as unforgeable as a worker's), and reply nudges are pushed to it.
# REQUIRED when the primary runs inside a herdr pane — without it the pane match reads the
# primary as a worker and refuses spawn/send/stop. Get the id from bridge_whoami; re-pin if
# the primary moves panes.
# primary:
# terminal: term_0123456789abcd
# pushReminders: 5 # max nudges before giving up (default 5)
# pushBackoffMs: 15000 # delay between nudges (default 15000)
# CB-530: MORE THAN ONE LEAD. `primary:` above is singular by construction — every other pane
# resolves as a worker — which is right for one lead driving a fleet and wrong the moment two leads
# (say a Claude lead and an opencode lead) work as peers: the second is silently demoted and refused
# every orchestration call. List each lead's pane here and all of them resolve as leads.
#
# terminal → the ONLY field identity depends on; get it from that session's bridge_whoami
# kind/model → descriptive; they document what runs in the pane and are echoed by bridge_whoami
#
# A lead is never spawned — it pre-exists, which is exactly why it must be named rather than created.
# `bridge_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
# IS a primary for authorization, so nothing that keys on the role breaks.
#
# KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single
# destination for its nudges. If both name the same terminal, the `leaders:` entry wins.
# leaders:
# opus-5.0:
# terminal: term_0123456789abcd
# kind: claude
# gpt-sol-5.6:
# terminal: term_fedcba9876543
# kind: opencode
# model: openai/gpt-5.6-terra
# CB-531: FIND LEADS BY TAB NAME instead of pasting terminal ids. `leaders:` above needs an id that
# only exists once the session is running, so adding a lead is: open a tab, start the agent, ask it
# bridge_whoami, edit this file, restart the daemon. This block replaces all of that with a naming
# convention — label the tab `lead: <name>` when you open it and the pane is recognised on the next
# rescan, with no config edit and no restart. Reopen the tab later and the id changes; the label
# does not.
#
# bridged NEVER writes these labels. It renames worker tabs (see `tabLabel` below) but reads lead
# tabs read-only, so what is in the tab bar is always what you typed. Two things keep the convention
# from being a way to claim leadership: the configured worker spaces are excluded from the scan, so
# nothing bridged places can land in a matching tab; and startup REFUSES a `tabPrefix` that any
# worker `tabLabel` also matches, so the two namespaces cannot overlap by accident.
#
# Opt-in on purpose — this widens who resolves as a lead, so upgrading the daemon must never switch
# it on for you. Absent block = leads come only from `leaders:`/`primary:`, exactly as before.
# leadScan:
# tabPrefix: "lead:" # `lead: opus-5.0` ⇒ a lead named opus-5.0 (case-insensitive; default "lead:")
# intervalSeconds: 10 # rescan cadence, and the worst case before a new tab is recognised
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
@@ -53,7 +106,16 @@ herdrSocket: ~/.config/herdr/herdr.sock
# skills/MCP/hooks. Omit to leave the worker on the host default.
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
# worktree sees the same local config (CB-301-ext). Omit for the default set:
# [.mcp.json, .claude/settings.local.json, .env, .envrc].
# [.claude/settings.local.json, .env, .envrc].
#
# Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher
# mounts — the bridge, and nothing else. Replicating the primary's MCP config
# handed a worker the primary's IDE servers, which are bound to the primary's
# checkout, so its navigation returned paths OUTSIDE its own worktree: one
# worker made all 59 of its edits in the primary tree while compiling its
# worktree, and every build it ran was of code that did not contain them.
# bridged neutralizes a provisioned worktree's .mcp.json for this reason;
# listing it here would copy the primary's back over that.
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
# Opt-in by design — omit and the worker gets no PR-create grant (push over
@@ -89,18 +151,28 @@ workers:
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
weight: 0.5 # relative selection weight for placement: weighted
maxLoad: 2 # max live workers on this profile (omit for unlimited)
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".mcp.json", ".claude/settings.local.json", ".env", ".envrc"]
ollama:
baseUrl: http://ollama.ltms.dev # local/self-hosted; usually no token
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
gx11: # a second backend, so `placement: weighted` has a choice
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}"
mcpUrl: http://127.0.0.1:8765/mcp
argv: ["ccs", "ollama"]
argv: ["ccs", "gx11"]
weight: 0.5
maxLoad: 2
# Pin an auto-compact window BELOW the served model's context ceiling. The global
# ~/.claude/settings.json value is shared by every ccs instance and the primary, so the
# per-profile override belongs here. Equal to the ceiling means auto-compact never fires
# before the server rejects the prompt, which kills a worker mid-turn (CB-523).
env:
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "280000"
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
@@ -144,6 +216,9 @@ workers:
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
placement: weighted
defaultWorker: gx10
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
@@ -152,7 +227,6 @@ guard:
offSubscriptionHosts:
- gx00.gw
- gx01.gw
- ollama.ltms.dev
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
@@ -195,6 +269,15 @@ guard:
# the first orchestration-side MCP call (the normal case). An off-host or
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
# degrades to pull; the reply is still never lost.
#
# REQUIRED (CB-522) if the primary itself runs inside a herdr pane. Caller
# identity resolves a loopback PID to its herdr pane, and PaneLocator scans
# EVERY pane — not just bridged-spawned ones — so such a primary is otherwise
# classified as a WORKER and refused SPAWN/SEND/STOP. That failure is
# self-locking: the learned terminal is populated by the very orchestration
# calls being refused, so only this pinned value can break the cycle. Read the
# id off bridge_whoami (it reports the current terminal even while
# misclassified) and re-pin whenever the primary moves panes.
# pushReminders → max nudges before giving up (default 5)
# pushBackoffMs → delay between nudges in ms (default 15000)
# primary:
+21 -1
View File
@@ -6,7 +6,7 @@
<groupId>dev.ltms</groupId>
<artifactId>bridged</artifactId>
<version>0.1.0-SNAPSHOT</version>
<version>1.0.0</version>
<packaging>jar</packaging>
<name>bridged</name>
@@ -236,6 +236,26 @@
<profile>
<id>contract</id>
<properties><excludedGroups/></properties>
<build>
<plugins>
<plugin>
<artifactId>maven-surefire-plugin</artifactId>
<configuration>
<!-- Docker-engine compat (see "Running the contract tests" in
docs/CB-307-Reliable-Delivery.md): Testcontainers 1.20.4's docker-java
client defaults to Docker API 1.32 when no version is set, but modern
engines (OrbStack on this dev host, min 1.40) reject that as too old —
which surfaces as "Could not find a valid Docker environment". Pinning
api.version=1.43 works on OrbStack and Docker 24+, and is overridable
per-host via -Dapi.version. Only active under -Pcontract, so the
default hermetic build never sets it. -->
<systemPropertyVariables>
<api.version>1.43</api.version>
</systemPropertyVariables>
</configuration>
</plugin>
</plugins>
</build>
</profile>
</profiles>
</project>
@@ -5,6 +5,7 @@ import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.LeadTabScanner;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
import dev.ltms.bridged.herdr.WorkspaceControl;
@@ -13,6 +14,7 @@ import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.auth.ArchitectRegistry;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
@@ -33,6 +35,7 @@ import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.session.SessionReaper;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.placement.PlacementPolicies;
import dev.ltms.bridged.worker.CompositePeerLauncher;
import dev.ltms.bridged.worker.HerdrPeerLauncher;
import dev.ltms.bridged.worker.OpenCodeLauncher;
@@ -45,7 +48,15 @@ import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.stream.Collectors;
/**
* {@code bridged} entry point. Wires the real herdr socket client to the REST app and
@@ -76,6 +87,13 @@ public final class Bridged {
// refuses remote connections to a loopback socket. This throws rather than warns so the
// dangerous configuration cannot be reached by ignoring a log line.
cfg.validateAuthExposure();
cfg.validateLeadScan();
// CB-542: a subscription:true profile whose env: reseats ANTHROPIC_BASE_URL/AUTH_TOKEN would
// reach an unguarded endpoint (the launcher skips SubscriptionGuard for it). Refuse at load.
cfg.validateSubscriptionProfiles();
// CB-548: every architect slot must name a configured workers: profile — the strong-model
// backend the future spawn lifecycle would read. A stale reference dies here, not later.
cfg.validateArchitects();
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
@@ -111,7 +129,13 @@ public final class Bridged {
opencodeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
}
PeerLauncher workers = new CompositePeerLauncher(adapters, cfg.defaultProfile());
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
PeerLauncher workers = new CompositePeerLauncher(
adapters,
cfg.defaultProfile(),
cfg.workerProfiles(),
PlacementPolicies.fromName(cfg.placement()),
profileName -> liveCountRef.get().apply(profileName));
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
@@ -135,7 +159,12 @@ public final class Bridged {
&& cfg.lifecycle().contextCap() > 0) {
contextCap = cfg.lifecycle().contextCap();
}
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()), contextCap);
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()),
System::nanoTime, contextCap, clearAfterTurn);
liveCountRef.set(profileName -> (int) sessions.roster().stream()
.filter(s -> profileName.equals(s.profile()))
.count());
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
final SessionReaper reaper;
@@ -148,6 +177,45 @@ public final class Bridged {
reaper = null;
}
// CB-530: every pane the config names as a lead, merged from `leaders:` and the legacy
// singular pin. PrimaryRegistry below still tracks ONE terminal — it addresses the push
// loop's nudges, which need a single destination — so it keeps the legacy pin.
Map<String, String> leadTerminals = cfg.leaderTerminals();
if (leadTerminals.size() > 1) {
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
}
// CB-531: on top of the static registry, discover leads by the tab labels the operator
// writes. Opt-in, so a config with no `leadScan:` block resolves exactly as it did under
// CB-530 — the supplier is then a constant and never touches herdr.
final Supplier<Map<String, String>> leads;
if (cfg.leadScan() != null) {
var scan = cfg.leadScan();
Set<String> workerSpaces = cfg.workerProfiles().values().stream()
.map(BridgedConfig.Worker::workspace)
.filter(Objects::nonNull)
.collect(Collectors.toSet());
leads = new LeadTabScanner(herdr, scan.tabPrefix(), workerSpaces, leadTerminals,
TimeUnit.SECONDS.toNanos(scan.intervalSeconds()), System::nanoTime);
log.info("lead scan: tabs labelled '{}…' host a lead (rescan every {}s, worker spaces {} "
+ "excluded)",
scan.tabPrefix(), scan.intervalSeconds(), workerSpaces);
} else {
leads = () -> leadTerminals;
}
// CB-548: config-declared architect slots. Config supplies only the stable name → profile
// map; the terminal → slot binding is owned by the registry and is empty at startup, so no
// pane resolves to an architect until the later spawn lifecycle binds one. The registry is
// what CallerResolver resolves against and what that lifecycle will read profiles from;
// nothing here spawns a slot.
ArchitectRegistry architects = new ArchitectRegistry(
cfg.architects() == null ? Map.of() : cfg.architects());
if (!architects.slots().isEmpty()) {
log.info("architect slots: {} configured {} — none bound yet (a slot is idle until the "
+ "spawn lifecycle binds a live terminal to it)",
architects.slots().size(), architects.slots().keySet());
}
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
@@ -163,6 +231,17 @@ public final class Bridged {
sessions.onTurnComplete(target);
}
@Override
public boolean hasPostTurnAction(String target) {
return sessions.hasPostTurnAction(target);
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
completion.resolveBeforePostAction(target);
return sessions.onTurnCompleteWithPostAction(target);
}
@Override
public void onDelivered(String target) {
completion.onDelivered(target);
@@ -175,7 +254,8 @@ public final class Bridged {
sessions.onTurnFailed(target);
}
};
Injector injector = new Injector(agents, turnListener, presence::isPresent, presence::forget);
Injector injector = new Injector(agents, turnListener, deliverableTo(presence, leads),
presence::forget);
StatusPoller poller = new StatusPoller(agents, injector, INJECT_POLL_MILLIS);
poller.start();
@@ -191,9 +271,19 @@ public final class Bridged {
log.info("reply inbox: in-memory (soft-state)");
}
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
PrimaryRegistry primaryRegistry = new PrimaryRegistry(
cfg.primary() != null ? cfg.primary().terminal() : null);
// The pin also feeds CallerResolver below: a primary running inside a herdr pane would
// otherwise resolve as a worker and be refused every orchestration tool.
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
// CB-532: `primary.terminal` is superseded and no longer needed for either of its jobs —
// identity comes from `leaders:`/`leadScan:`, and reply nudges now follow the delegating
// lead. Say so once at startup rather than leaving a redundant pin to look load-bearing.
if (pinnedPrimaryTerminal != null && !pinnedPrimaryTerminal.isBlank()) {
log.warn("primary.terminal is DEPRECATED (CB-532) and can be deleted: identity now comes "
+ "from leaders:/leadScan:, and reply nudges follow the lead that delegated. "
+ "It still works, and is still the fallback nudge destination when a restart "
+ "has lost the delegation map. Its pushReminders/pushBackoffMs stay valid.");
}
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
// open bridge_send. Uses its own lightweight scheduled executor, separate from the injector.
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
@@ -209,11 +299,17 @@ public final class Bridged {
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
sessions.onAcquire(replyInbox::own);
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
sessions.onRelease(terminal ->
messages.abandon(terminal, "the worker session was released before it replied"));
sessions.onRelease(terminal -> {
messages.abandon(terminal, "the worker session was released before it replied");
replyInbox.release(terminal);
primaryRegistry.forgetDelegation(terminal); // CB-532: don't leak the lead binding
});
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
@@ -229,11 +325,13 @@ public final class Bridged {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting bridged");
}
callers = new CallerResolver(identity, true, token);
callers = CallerResolver.withLeadsAndArchitects(identity, true, token, leads,
architects::snapshot);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = new CallerResolver(identity);
callers = CallerResolver.withLeadsAndArchitects(identity, false, null, leads,
architects::snapshot);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
@@ -268,6 +366,28 @@ public final class Bridged {
cfg.bind().host(), cfg.bind().port(), socket);
}
/**
* The {@link Injector}'s readiness gate (CB-534): a target is deliverable if it is a worker whose
* agent has connected the bridge MCP, <em>or</em> a lead.
*
* <p>The gate exists for one reason — to hold a delivery out of a <em>spawned</em> worker's boot
* window, where herdr already reports {@code idle} but the TUI would drop an injected paste. That
* hazard is a property of spawning. A lead is never spawned: the operator started it and named it
* (or labelled its tab) only once it was up, so there is no boot window to guard.
*
* <p>A lead is also never enrolled in {@link WorkerPresence} — {@code BridgeMcp} marks presence
* only for a worker, deliberately, since that map doubles as the worker roster's availability
* signal and a lead counted there would show up as an available worker. So without the second
* disjunct a lead is permanently un-deliverable: every lead→lead send sat on the gate for
* {@code READINESS_GRACE_POLLS} (~60s) and then failed having never been typed into the pane.
*
* <p>The lead set is read through the supplier on each call rather than snapshotted, so a lead
* discovered by {@code leadScan} after startup becomes deliverable without a restart.
*/
static Predicate<String> deliverableTo(WorkerPresence presence, Supplier<Map<String, String>> leads) {
return target -> presence.isPresent(target) || leads.get().containsKey(target);
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
*
@@ -0,0 +1,142 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.config.BridgedConfig;
import java.util.HashMap;
import java.util.Map;
/**
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
* profile it points at, plus the <em>live</em> bindings from a live architect's herdr terminal to
* its slot.
*
* <p>Two halves, split by who owns each:
* <ul>
* <li><b>slots</b> — configured once, keyed by the gateway-local unique name; each carries the
* {@code profile} reference the spawn lifecycle reads when it stands the slot up. A read-only
* snapshot taken at construction.</li>
* <li><b>terminal bindings</b> — owned by this registry and initially <em>empty</em>. Config
* declares no architect terminal, so at startup every slot is idle and nothing resolves to an
* architect; a session only becomes one when the spawn lifecycle {@linkplain #bind(String,
* String) binds} its terminal to a slot. {@link CallerResolver} reads this through
* {@link #snapshot()} to turn a pane into an {@link Role#ARCHITECT}.</li>
* </ul>
*
* <p>Spawning/lifecycle is deliberately a separate unit: this class only owns the bindings and
* exposes the map the resolver resolves against plus the profile lookup lifecycle will call.
* Nothing here creates or manages an architect session.
*/
public final class ArchitectRegistry {
private final Map<String, BridgedConfig.Architect> slots;
/** Live {@code terminal_id → slot name}; guarded by {@code this}. */
private final Map<String, String> terminalToSlot = new HashMap<>();
public ArchitectRegistry(Map<String, BridgedConfig.Architect> slots) {
this.slots = slots == null ? Map.of() : Map.copyOf(slots);
}
/** The configured slots, keyed by gateway-local unique name. Unmodifiable snapshot. */
public Map<String, BridgedConfig.Architect> slots() {
return slots;
}
/**
* An immutable copy of the live {@code terminal_id → slot name} bindings.
*
* <p>Passed to {@link CallerResolver} as the source of architect identity, and what
* {@code bridge_whoami}/the roster will read to say which slot a pane hosts. Empty until the
* spawn lifecycle binds a slot.
*/
public Map<String, String> snapshot() {
synchronized (terminalToSlot) {
return Map.copyOf(terminalToSlot);
}
}
/** The slot a live terminal is bound to, or {@code null} if it is not an architect slot. */
public String slotForTerminal(String terminal) {
if (terminal == null) {
return null;
}
synchronized (terminalToSlot) {
return terminalToSlot.get(terminal);
}
}
/**
* The strong-model profile a slot runs under — what the spawn lifecycle reads.
*
* @return the slot's configured {@code profile}, or {@code null} if the slot is unknown or
* declares none
*/
public String profileForSlot(String slotName) {
BridgedConfig.Architect a = slots.get(slotName);
return (a == null || a.profile() == null) ? null : a.profile();
}
/** True when {@code slotName} is a configured architect slot. */
public boolean isSlot(String slotName) {
return slots.containsKey(slotName);
}
/**
* Bind {@code terminal} to {@code slot} (CB-548).
*
* <p>The spawn lifecycle calls this when it stands a slot up. The bind is atomic and preserves
* the two cardinality invariants: a terminal may occupy at most one slot, and a slot may host at
* most one terminal. Binding the same terminal to the same slot again is a harmless no-op.
*
* @param slot a configured slot name, or the bind is refused
* @param terminal the pane that will act as this architect
* @return {@code true} if the binding is now {@code terminal → slot}; {@code false} if it was
* refused — an unknown slot, a terminal already bound to a different slot, or a slot
* already hosting a different terminal
*/
public boolean bind(String slot, String terminal) {
if (slot == null || terminal == null || terminal.isBlank()) {
return false;
}
synchronized (terminalToSlot) {
if (!isSlot(slot)) {
return false; // unknown slot — nothing to bind to
}
String existingSlot = terminalToSlot.get(terminal);
if (existingSlot != null) {
return slot.equals(existingSlot); // already this slot (idempotent) or a different one
}
if (terminalToSlot.containsValue(slot)) {
return false; // slot already hosts a terminal — no second one
}
terminalToSlot.put(terminal, slot);
return true;
}
}
/**
* Compare-safe unbind of {@code expectedTerminal} from {@code slot} (CB-548).
*
* <p>The spawn lifecycle calls this when it tears a slot down. Only the exact binding
* {@code expectedTerminal → slot} is removed; if that terminal was since rebound to a different
* slot (or the slot to a different terminal), the call is a no-op returning {@code false} — a
* stale unbind must never remove a replacement.
*
* @param slot the slot the caller believes the terminal is bound to
* @param expectedTerminal the terminal it expects to be bound there
* @return {@code true} if {@code expectedTerminal → slot} was removed; {@code false} if nothing
* was (no such binding, or the binding had already moved)
*/
public boolean unbind(String slot, String expectedTerminal) {
if (slot == null || expectedTerminal == null) {
return false;
}
synchronized (terminalToSlot) {
String current = terminalToSlot.get(expectedTerminal);
if (current == null || !slot.equals(current)) {
return false; // absent, or a replacement/moved binding — leave it in place
}
terminalToSlot.remove(expectedTerminal);
return true;
}
}
}
@@ -46,18 +46,28 @@ public final class Authz {
return false; // authenticated as nothing ⇒ authorized for nothing
}
return switch (action) {
// Orchestration is the primary's alone. A worker driving spawn/stop/send would be a
// worker escalating into the orchestrator role.
case SPAWN, STOP, SEND, DRAIN -> caller.isPrimary();
// Fleet lifecycle is the primary's alone — spawn, stop, drain. An architect
// deliberately does NOT get these (CB-548), so it cannot tear down or stand up workers
// even though it coordinates them; and a worker driving any of these would be a worker
// escalating into the orchestrator role.
case SPAWN, STOP, DRAIN -> caller.isPrimary();
// The load-bearing rule: a worker acts only as itself. The primary is deliberately
// excluded — a reply/ask is a worker's own turn output, and letting the primary forge
// one would corrupt the rendezvous correlation it is itself waiting on.
// Delivering a turn is open to the primary and the architect: an architect delegates
// to workers (that is the role's point) but still has no lifecycle rights. A worker is
// excluded — sending would be it escalating.
case SEND -> caller.isPrimary() || caller.isArchitect();
// The load-bearing rule: a caller acts only as the pane it occupies. CB-532 widened who
// that can be — a lead answering another lead is replying for its OWN terminal, which
// this already permits — while the rule itself is unchanged, and is what stops anyone
// forging a reply for a rendezvous someone else is waiting on. An architect's own pane
// passes through the same check, so it can answer a funnel that delegated to it. An
// unnamed primary (token/loopback, no pane) owns nothing and is still excluded.
case REPLY, ASK -> caller.ownsSession(targetSession);
// Observation is open to both authenticated roles: a worker legitimately polls its own
// Observation is open to every authenticated role: a worker legitimately polls its own
// status, and the roster carries no secrets.
case READ, METRICS -> caller.isPrimary() || caller.isWorker();
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
};
}
@@ -4,6 +4,8 @@ import dev.ltms.bridged.mcp.ConnectionIdentity;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
import java.util.Map;
import java.util.function.Supplier;
/**
* Resolves every caller to a {@link Principal}, for both entry paths into the core (CB-501).
@@ -15,7 +17,18 @@ import java.security.MessageDigest;
*
* <p><strong>Resolution order</strong> — connection identity first, token second, nothing third:
* <ol>
* <li>A loopback peer PID that maps to a herdr worker pane ⇒ {@link Role#WORKER}. This is
* <li>A loopback peer PID that maps to a pane named by {@code leaders:}, by the legacy
* {@code primary.terminal} pin, or by an operator-labelled lead tab (CB-307, CB-530, CB-531)
* ⇒ {@link Role#PRIMARY}, carrying that lead's
* name. The pane mapping is as unforgeable as a worker's, and the config explicitly names
* that pane as a lead's own — without this rule a lead running <em>inside</em> a herdr pane
* is misread as a worker and locked out of orchestration. More than one pane may be named,
* so two leads can work as peers rather than one being demoted.</li>
* <li>A loopback peer PID that maps to a pane bound to a CB-548 architect slot ⇒
* {@link Role#ARCHITECT}, carrying the slot name. Just unforgeable as a worker's, and
* resolved from the <em>live</em> terminal→slot binding (never a request argument), before
* the generic worker fallback.</li>
* <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#WORKER}. This is
* unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured
* regardless of auth mode, so enabling auth never breaks the fleet.</li>
* <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li>
@@ -29,18 +42,115 @@ public final class CallerResolver {
private final ConnectionIdentity identity;
private final boolean tokenMode;
private final byte[] expectedToken; // null unless tokenMode
/**
* terminal_id → lead name; empty when nothing is pinned. CB-530.
*
* <p>A supplier rather than a map because the registry is no longer fixed at startup: CB-531
* discovers leads by scanning herdr for operator-labelled tabs, so a lead that opens its tab
* after the daemon booted must still be recognised. Consulted per resolve; the scanner behind
* it is TTL-cached, so this is a map lookup in the common case.
*/
private final Supplier<Map<String, String>> leadTerminals;
/**
* terminal_id → architect slot name; empty when nothing is configured. CB-548.
*
* <p>Like {@link #leadTerminals}, a supplier rather than a fixed map, so a binding injected
* after startup — when the later spawn lifecycle establishes a live architect session, or an
* operator pins one — takes effect without a restart. Consulted per resolve; today's wiring
* in {@code Bridged} reads a constant from config, which is the degenerate live case.
*/
private final Supplier<Map<String, String>> architectTerminals;
/** Loopback-trust resolver: no token required, historical behaviour. */
public CallerResolver(ConnectionIdentity identity) {
this(identity, false, null);
this(identity, false, null, Map.of());
}
/** As {@link #CallerResolver(ConnectionIdentity, boolean, String, Map)} with no leads pinned. */
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
this(identity, tokenMode, token, Map.of());
}
/**
* @param identity connection-based worker identification
* @param tokenMode when true, a non-worker caller must present a valid bearer token
* @param token the expected bearer token; required (non-blank) when {@code tokenMode}
* Single-pin form for the legacy {@code primary.terminal}-only configuration — one lead, named
* {@code primary}.
*
* <p>A static factory rather than a fourth constructor overload on purpose: {@code String} and
* {@code Map} overloads are ambiguous for a literal {@code null} argument, which is a compile
* error at the call site and exactly the shape "unpinned" is written in.
*
* @param pinnedPrimaryTerminal the primary's own herdr {@code terminal_id}
* ({@code null}/blank = unpinned)
*/
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
public static CallerResolver pinnedTo(ConnectionIdentity identity, boolean tokenMode,
String token, String pinnedPrimaryTerminal) {
return new CallerResolver(identity, tokenMode, token,
pinnedPrimaryTerminal == null || pinnedPrimaryTerminal.isBlank()
? Map.of() : Map.of(pinnedPrimaryTerminal, "primary"));
}
/**
* @param identity connection-based worker identification
* @param tokenMode when true, a non-worker caller must present a valid bearer token
* @param token the expected bearer token; required (non-blank) when {@code tokenMode}
* @param leadTerminals herdr {@code terminal_id} → lead name for every configured lead
* (CB-530). A caller resolving to one of these panes is that lead — a
* {@link Role#PRIMARY} — rather than a worker. Empty = nothing pinned,
* so every pane resolves as a worker.
*/
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Map<String, String> leadTerminals) {
this(identity, tokenMode, token, fixed(leadTerminals), null);
}
/**
* Map-form of both registries (CB-548): lead terminals and the initial architect terminal
* bindings, each snapshotted at construction (a handed-over map is not offered as live state).
*/
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Map<String, String> leadTerminals,
Map<String, String> architectTerminals) {
this(identity, tokenMode, token, fixed(leadTerminals), fixed(architectTerminals));
}
/**
* Live-registry form: {@code leadTerminals} is consulted on every resolve, so leads discovered
* after startup (CB-531's tab scan) take effect without a restart.
*
* <p>A static factory rather than a fourth constructor overload, for the same reason as
* {@link #pinnedTo}: {@code Map} and {@code Supplier} overloads are ambiguous for a literal
* {@code null}.
*/
public static CallerResolver withLeads(ConnectionIdentity identity, boolean tokenMode,
String token,
Supplier<Map<String, String>> leadTerminals) {
return new CallerResolver(identity, tokenMode, token, leadTerminals, null);
}
/**
* Live-registry form for both {@code leadTerminals} and the CB-548 architect registry: both
* are consulted on every resolve, so a slot binding injected after startup takes effect
* without a restart.
*
* <p>A static factory rather than a constructor overload, for the same reason as
* {@link #pinnedTo}: too many {@code Map}/{@code Supplier} combinations to make {@code null}
* unambiguous.
*/
public static CallerResolver withLeadsAndArchitects(ConnectionIdentity identity,
boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals) {
return new CallerResolver(identity, tokenMode, token, leadTerminals, architectTerminals);
}
private static Supplier<Map<String, String>> fixed(Map<String, String> leadTerminals) {
Map<String, String> snapshot = leadTerminals == null ? Map.of() : Map.copyOf(leadTerminals);
return () -> snapshot;
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals) {
if (tokenMode && (token == null || token.isBlank())) {
throw new IllegalArgumentException(
"auth.mode=token requires a non-empty token; check that the env var named by "
@@ -49,6 +159,32 @@ public final class CallerResolver {
this.identity = identity;
this.tokenMode = tokenMode;
this.expectedToken = tokenMode ? token.getBytes(StandardCharsets.UTF_8) : null;
this.leadTerminals = leadTerminals == null ? Map::of : leadTerminals;
this.architectTerminals = architectTerminals == null ? Map::of : architectTerminals;
}
/**
* The currently-recognised leads, {@code terminal_id → name} (CB-535).
*
* <p>Deliberately read from the same supplier {@link #resolve} consults, rather than from a
* second copy handed to the roster: a lead that is <em>listed</em> but would not <em>resolve</em>
* (or the reverse) is an address a peer cannot actually reach, and the two answers drifting apart
* is precisely the confusion this exists to end. Live, so a lead discovered by the tab scan after
* startup appears without a restart.
*/
public Map<String, String> leads() {
return leadTerminals.get();
}
/**
* The currently-recognised architect slots, {@code terminal_id → slot name} (CB-548).
*
* <p>Read from the same supplier {@link #resolve} consults, so a slot that is <em>listed</em>
* here but would not <em>resolve</em> (or the reverse) cannot drift apart. Live for the same
* reason as {@link #leads()}.
*/
public Map<String, String> architects() {
return architectTerminals.get();
}
/**
@@ -61,6 +197,21 @@ public final class CallerResolver {
public Principal resolve(String remoteAddr, int remotePort, String authorizationHeader) {
ConnectionIdentity.Caller c = identity.resolve(remoteAddr, remotePort);
if (c.terminal() != null) {
String lead = leadTerminals.get().get(c.terminal());
if (lead != null) {
// The config names this pane as a lead's own. The pane mapping is exactly as
// unforgeable as a worker's, so it outranks the token path — no credential needed.
// Checked before the architect registry so a pane named in BOTH is still the lead
// (CB-548 preserves every existing leader behaviour).
return Principal.leader(lead, c.terminal(), c.pid());
}
String slot = architectTerminals.get().get(c.terminal());
if (slot != null) {
// The config/live binding names this pane as an architect slot's own. Same
// unforgeable pane mapping; the live binding, never a request argument, decides.
// Checked before the generic worker fallback, per the CB-548 precedence order.
return Principal.architect(slot, c.terminal(), c.pid());
}
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
}
@@ -76,16 +227,6 @@ public final class CallerResolver {
return isLoopback(remoteAddr) ? Principal.primary(c.pid()) : Principal.anonymous();
}
/** The working directory of the calling process (CB-112 spawn cwd inheritance), or {@code null}. */
public String cwdForPid(long pid) {
return identity.cwdForPid(pid);
}
/** True when auth requires a bearer token of non-worker callers. */
public boolean tokenMode() {
return tokenMode;
}
private boolean presentedTokenMatches(String authorizationHeader) {
String presented = bearerValue(authorizationHeader);
if (presented == null) {
@@ -5,30 +5,82 @@ package dev.ltms.bridged.auth;
* identifies which worker it is (CB-501).
*
* @param role what this caller is authorized to act as
* @param terminal the worker's herdr terminal id; {@code null} for {@code PRIMARY}/{@code ANONYMOUS}
* @param terminal the herdr {@code terminal_id} of the pane this caller occupies — a worker's, or
* (since CB-532) a named lead's; {@code null} for an unnamed primary resolved off a
* token or loopback trust, and for {@code ANONYMOUS}
* @param pid the connecting process id, or {@code -1} when not resolvable (audit context)
* @param name for a lead resolved from the CB-530 {@code leaders:} registry, which lead it is;
* for an architect resolved from the CB-548 {@code architects:} registry, which
* slot it occupies; {@code null} for every other caller, including an unnamed primary
*/
public record Principal(Role role, String terminal, long pid) {
public record Principal(Role role, String terminal, long pid, String name) {
/**
* Three-arg form for the callers that have no name to carry (workers, anonymous, and the
* token/loopback primary paths). Kept so adding CB-530's {@code name} did not churn every
* construction site — and so a reconstructed principal without a stashed name still works.
*/
public Principal(Role role, String terminal, long pid) {
this(role, terminal, pid, null);
}
/** A caller authenticated as nothing — the default when no check establishes anything else. */
public static Principal anonymous() {
return new Principal(Role.ANONYMOUS, null, -1);
}
/** The orchestrating session. */
/** The orchestrating session, unnamed (token or loopback-trust path). */
public static Principal primary(long pid) {
return new Principal(Role.PRIMARY, null, pid);
}
/**
* A named lead from the {@code leaders:} registry (CB-530).
*
* <p>Carries {@link Role#PRIMARY}: a lead <em>is</em> a primary as far as authorization goes,
* so every existing {@code isPrimary()} gate keeps working unchanged and the role table needed
* no new entry. The name is reporting only — it lets {@code bridge_whoami} say <em>which</em>
* lead is asking once more than one is configured.
*
* <p><strong>CB-532: a lead now carries the terminal it was matched by.</strong> Under CB-530 it
* deliberately did not, because {@code terminal} meant "which worker pane" everywhere and a
* non-null one would have enrolled the lead in the worker presence map. That reading was what
* made a lead unaddressable: {@link #ownsSession} could never be true for it, so
* {@code bridge_reply} was refused and one lead could send to another but never be answered.
* The terminal now means "which pane is this caller", the presence map keys on
* {@link #isWorker()} instead, and a lead is a peer that can both send and receive.
*/
public static Principal leader(String name, String terminal, long pid) {
return new Principal(Role.PRIMARY, terminal, pid, name);
}
/** A worker peer, identified by its herdr pane. */
public static Principal worker(String terminal, long pid) {
return new Principal(Role.WORKER, terminal, pid);
}
/**
* An architect (CB-548), identified by the slot it occupies and the pane bound to it.
*
* <p>Carries {@link Role#ARCHITECT}. {@code slotName} is reporting only — it lets
* {@code bridge_whoami} say <em>which</em> architect slot is asking, and it is the key the
* (future) spawn lifecycle reads a profile back from. Identity is the {@code terminal}: like a
* worker's it comes from the connection and the live terminal→slot binding, so
* {@code ownsSession} works exactly as it does for a worker — an architect acts as its own
* pane and no other.
*/
public static Principal architect(String slotName, String terminal, long pid) {
return new Principal(Role.ARCHITECT, terminal, pid, slotName);
}
public boolean isPrimary() {
return role == Role.PRIMARY;
}
public boolean isArchitect() {
return role == Role.ARCHITECT;
}
public boolean isWorker() {
return role == Role.WORKER;
}
@@ -39,18 +91,26 @@ public record Principal(Role role, String terminal, long pid) {
/**
* Whether this caller may act <em>as</em> {@code sessionId} — the "own session only" rule that
* keeps one worker from replying or asking on another's behalf. Only a worker can own a
* session, and only its own.
* keeps one peer from replying or asking on another's behalf.
*
* <p>The rule is about <em>identity</em>, not rank: a caller may act as the pane it demonstrably
* occupies, and as no other. CB-532 dropped the extra {@code isWorker()} conjunct that used to
* be here. It was not what enforced the rule — {@code terminal.equals(sessionId)} is, and that
* terminal comes from the connection, so it cannot be forged either way. All the conjunct did
* was make a lead permanently unable to answer anyone, since a lead's terminal was null and a
* lead is not a worker. A caller with no terminal at all (an off-host or token-authenticated
* primary) still owns nothing, which is the case the null check covers.
*/
public boolean ownsSession(String sessionId) {
return isWorker() && terminal != null && terminal.equals(sessionId);
return terminal != null && terminal.equals(sessionId);
}
/** Short, non-sensitive description for audit lines and error details. */
public String describe() {
return switch (role) {
case WORKER -> "worker:" + terminal;
case PRIMARY -> "primary";
case ARCHITECT -> "architect:" + name;
case PRIMARY -> name == null ? "primary" : "leader:" + name;
case ANONYMOUS -> "anonymous";
};
}
@@ -24,6 +24,16 @@ public enum Role {
*/
WORKER,
/**
* A config-declared architect slot (CB-548): a gateway-local named session on a strong-model
* profile that coordinates and delegates turns but does not own the fleet. Unforgeable like a
* worker's — derived from the connection's pane and the live terminal→slot binding, never from
* a request argument. May {@code SEND} a turn, {@code REPLY}/{@code ASK} only as its own pane,
* and {@code READ}/{@code METRICS}; may <em>not</em> {@code SPAWN}/{@code STOP}/{@code DRAIN}
* (those stay the primary's, to keep lifecycle in one pair of hands).
*/
ARCHITECT,
/** Authenticated as nothing. Authorized for nothing but {@code /healthz}. */
ANONYMOUS
}
@@ -1,13 +1,19 @@
package dev.ltms.bridged.config;
import com.fasterxml.jackson.annotation.JsonIgnoreProperties;
import com.fasterxml.jackson.core.JsonParser;
import com.fasterxml.jackson.core.JsonToken;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.dataformat.yaml.YAMLFactory;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Collections;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
@@ -16,7 +22,8 @@ import java.util.Set;
/**
* {@code bridged} configuration, loaded from a YAML file (see
* {@code bridged.example.yaml}). Unknown keys are ignored so config can grow ahead
* of the code.
* of the code — but an unknown <em>top-level</em> key is logged as a WARN at load (CB-530), because
* silently dropping a whole block is indistinguishable from honouring it.
*
* @param bind REST/MCP listen host:port
* @param herdrSocket path to herdr's Unix socket ({@code null} → client default)
@@ -36,6 +43,18 @@ import java.util.Set;
* @param primary optional pinned primary terminal config ({@code null} → derived from connection);
* a non-blank {@code terminal} seeds {@code PrimaryRegistry} and prevents
* connection-derived overrides, CB-307
* @param leaders named panes that orchestrate rather than are orchestrated (CB-530), keyed by
* lead name; supersedes the singular {@code primary} pin, which stays honoured.
* See {@link #leaderTerminals()} for how the two merge
* @param architects CB-548 architect slots, keyed by gateway-local unique slot name; each points
* at a strong-model profile the future spawn lifecycle reads back. An architect
* is <em>not</em> recognised like a lead: config declares the slots only, and a
* live session becomes an architect when the spawn lifecycle binds its terminal
* to a slot. Nothing here spawns a slot.
* @param leadScan opt-in discovery of leads by tab label (CB-531); {@code null} ⇒ no scanning,
* and only {@code leaders:}/{@code primary:} name a lead
* @param placement how to choose a worker profile for an unqualified spawn:
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
* historical behaviour), CB-501
*/
@@ -53,6 +72,10 @@ public record BridgedConfig(
Integer spawnReadyPollMs,
Broker broker,
Primary primary,
Map<String, Leader> leaders,
Map<String, Architect> architects,
LeadScan leadScan,
String placement,
Auth auth) {
@JsonIgnoreProperties(ignoreUnknown = true)
@@ -93,12 +116,29 @@ public record BridgedConfig(
* (minimal-grant default — push over SSH stays free, PR-create is opt-in)
* @param gitHostEnv name of the host env var holding the forge host (default {@code GITEA_HOST});
* injected as {@code GITEA_HOST} <em>only</em> when {@code gitTokenEnv} is set
* @param weight relative selection weight for {@code placement: weighted}. Absent or
* non-positive ⇒ 1.0. Weights are normalised by the policy, so they need
* not sum to 1.0.
* @param maxLoad max live workers allowed on this profile at one time; absent or
* non-positive ⇒ unlimited. Live means any session the registry still owns
* (acquired and not yet released), in any state.
* @param kind which peer launcher spawns this profile: {@code "claude-code"} (default —
* the {@link dev.ltms.bridged.worker.ClaudeCodeLauncher}) or {@code "opencode"}.
* The {@code CompositePeerLauncher} routes {@code spawn}/reap by this value, so
* each adapter drives only its own kind. Normalised to lower-case; blank ⇒ the
* default. It selects the adapter, not the transport — placement, tabs, cwd, and
* the readiness gate are kind-independent and stay in the shared base.
* @param subscription {@code true} to run this profile's workers on the operator's Claude
* subscription, on purpose (CB-539). When set, the launcher neither requires
* nor injects {@code ANTHROPIC_BASE_URL} / {@code ANTHROPIC_AUTH_TOKEN}, and the
* {@code SubscriptionGuard} base_url requirement is skipped <em>for this
* profile only</em>. Absent/{@code false} (the default) keeps today's hard
* refusal: a claude-code profile with no base_url may not spawn, because
* spawning one would bill the subscription. Mutually exclusive with
* {@code baseUrl} — setting both is a configuration error (the two state
* opposite intents). For the same reason, an {@code env:} entry naming
* {@code ANTHROPIC_BASE_URL} or {@code ANTHROPIC_AUTH_TOKEN} is refused at
* config load (CB-542): on the subscription path no guard would vet it.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Worker(String profile, String baseUrl, String model,
@@ -108,7 +148,10 @@ public record BridgedConfig(
List<String> parityOverlay,
String gitTokenEnv, String gitHostEnv,
String kind,
Map<String, String> env) {
Map<String, String> env,
Float weight,
Integer maxLoad,
Boolean subscription) {
/** Peer kind spawned by {@link dev.ltms.bridged.worker.ClaudeCodeLauncher} (the default). */
public static final String KIND_CLAUDE_CODE = "claude-code";
@@ -128,13 +171,20 @@ public record BridgedConfig(
placement = (placement == null || placement.isBlank()) ? "tab" : placement.toLowerCase();
workspace = (workspace == null || workspace.isBlank()) ? "bridged-workers" : workspace;
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? "worker: {profile} #{n}" : tabLabel;
// CB-525: .mcp.json is deliberately NOT here. Replicating the primary's MCP config gave a
// worker the primary's IDE servers, which are bound to the primary's checkout — so its
// navigation returned paths outside its own worktree. GitWorktrees now neutralizes that
// file instead; a worker's tools are whatever its launcher mounts.
parityOverlay = (parityOverlay == null || parityOverlay.isEmpty())
? List.of(".mcp.json", ".claude/settings.local.json", ".env", ".envrc")
? List.of(".claude/settings.local.json", ".env", ".envrc")
: List.copyOf(parityOverlay);
// gitTokenEnv stays null when unset (opt-in). gitHostEnv defaults so operators enabling
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
gitHostEnv = (gitHostEnv == null || gitHostEnv.isBlank()) ? "GITEA_HOST" : gitHostEnv;
env = (env == null) ? Map.of() : Map.copyOf(env);
weight = (weight == null || weight <= 0.0f) ? 1.0f : weight;
maxLoad = (maxLoad == null || maxLoad <= 0) ? null : maxLoad;
subscription = (subscription != null && subscription) ? Boolean.TRUE : Boolean.FALSE;
}
/**
@@ -147,7 +197,7 @@ public record BridgedConfig(
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, null, null, null);
mcpUrl, cwd, parityOverlay, null, null, null, null, null, null, null);
}
/**
@@ -159,7 +209,7 @@ public record BridgedConfig(
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, null, null);
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, null, null, null, null, null);
}
/**
@@ -172,13 +222,13 @@ public record BridgedConfig(
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, null);
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, null, null, null, null);
}
/** A copy with {@code profile} set — used to default a profile to its {@code workers} key. */
public Worker withProfile(String p) {
return new Worker(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env);
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, subscription);
}
/** True when this profile is served by the Claude Code adapter (the default kind). */
@@ -191,11 +241,48 @@ public record BridgedConfig(
return KIND_OPENCODE.equals(kind);
}
/**
* True when this profile runs on the operator's Claude subscription, on purpose (CB-539).
* Absent/{@code false} (the default) keeps the hard refusal: a claude-code profile with no
* base_url may not spawn, because doing so would bill the subscription.
*/
public boolean isSubscription() {
return Boolean.TRUE.equals(subscription);
}
/**
* Backward-compatible constructor without the CB-539 subscription flag — the worker stays on
* the off-subscription boundary (the default). Keeps pre-CB-539 call sites (and any YAML
* that omits the flag) compiling and behaving identically.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind, Map<String, String> env, Float weight, Integer maxLoad) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, null);
}
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
public boolean hasGitToken() {
return gitTokenEnv != null && !gitTokenEnv.isBlank();
}
/**
* True when this profile's {@code env:} block names a worker-side Anthropic binding variable.
* Those two keys are the adapter's, never the operator's: on a claude-code worker
* {@code ANTHROPIC_BASE_URL} is the endpoint the {@code SubscriptionGuard} vetted, and
* {@code ANTHROPIC_AUTH_TOKEN} is injected from {@code tokenEnv}. An {@code env:} entry for
* either is a bypass vector — it is what the subscription path would otherwise let survive
* unguarded — so it is rejected at config load (see
* {@link BridgedConfig#validateSubscriptionProfiles()}).
*/
public boolean envCarriesAnthropicBinding() {
return env != null
&& (env.containsKey("ANTHROPIC_BASE_URL") || env.containsKey("ANTHROPIC_AUTH_TOKEN"));
}
/** True when workers should land in their own tab in the worker space. */
public boolean tabPlacement() {
return "tab".equals(placement);
@@ -228,9 +315,12 @@ public record BridgedConfig(
* ({@code null} → disabled)
* @param drainTimeoutSeconds seconds to wait for {@code BUSY} sessions to finish before
* forced teardown on shutdown (default 5 when unset)
* @param clearAfterTurn whether a reusable worker discards its conversation context after
* every completed delegated turn (default false)
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Lifecycle(Integer idleTtlSeconds, Integer contextCap, Integer drainTimeoutSeconds) {
public record Lifecycle(Integer idleTtlSeconds, Integer contextCap, Integer drainTimeoutSeconds,
boolean clearAfterTurn) {
}
/**
@@ -254,8 +344,11 @@ public record BridgedConfig(
/**
* Optional pinned primary terminal config (CB-307). When present with a non-blank
* {@code terminal}, the bridge uses this as the primary's herdr identity instead of
* deriving it from the MCP connection. Useful when the primary runs off-host or in a
* non-herdr terminal where connection-derived identity is unavailable.
* deriving it from the MCP connection. It feeds two consumers: the push loop (where to nudge
* when replies land), and caller resolution — a caller whose connection maps to this pane is
* the primary, where the pane match would otherwise classify it as a worker. Pin it when the
* primary runs <em>inside</em> a herdr pane; it also helps off-host or non-herdr primaries,
* where connection-derived identity is unavailable and only the nudge target matters.
*
* @param terminal the primary's herdr {@code terminal_id} ({@code null}/blank → derive)
* @param pushReminders max reminder nudges before giving up (default 5)
@@ -274,6 +367,107 @@ public record BridgedConfig(
}
}
/**
* One entry of the CB-530 {@code leaders:} registry — a pane that orchestrates rather than one
* that is orchestrated.
*
* <p>Why a registry and not a second {@code primary:}: {@code primary.terminal} is singular by
* construction, so a session in any other pane resolves as a worker. That is correct while one
* lead drives a fleet, and wrong the moment two leads (say an Opus lead and an opencode lead)
* work as peers — the second is silently demoted and refused every orchestration call.
*
* <p>{@code kind} and {@code model} are descriptive only at this stage: they document what runs
* in the pane and are reported back by {@code bridge_whoami}. Nothing spawns a lead — a lead
* pre-exists, which is precisely why it must be recognised by configuration rather than created.
*
* @param terminal the lead's herdr {@code terminal_id}; the only field identity depends on
* @param kind which agent runs there ({@code claude}, {@code opencode}, …); descriptive
* @param model the model or selector it runs, for operators reading the roster; descriptive
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Leader(String terminal, String kind, String model) {
}
/**
* One entry of the CB-548 {@code architects:} registry — a gateway-local named slot that points
* at a strong-model profile.
*
* <p>A slot is <em>declared</em>, not recognised: config names the slot and the profile it runs,
* and nothing else. Unlike a lead (which config pins by herdr {@code terminal_id} and is
* recognised at startup), an architect slot is idle at boot — config supplies no terminal, so no
* session resolves to one until the spawn lifecycle binds a live terminal to the slot. The
* stable name + profile pair is the only config-time identity; live identity is defined purely
* by the runtime {@link dev.ltms.bridged.auth.ArchitectRegistry} binding.
*
* <p>Why a {@code profile} reference: an architect is meant to run a strong model, and the slot
* records which {@code workers:} profile that is — the value the future spawn lifecycle reads.
* It must name a configured profile, enforced by {@link #validateArchitects()} (a stale or
* typo'd reference fails at startup rather than silently spawning the wrong backend later).
*
* @param profile the name of the strong-model {@code workers:} profile this slot runs;
* required and validated against {@link #workerProfiles()}
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Architect(String profile) {
}
/**
* Discover leads by tab label instead of by pasted {@code terminal_id} (CB-531).
*
* <p>Why: a lead is not spawned, so its {@code terminal_id} exists only once a human has opened
* the tab and started the agent — which makes {@code leaders:} a three-step ritual (start it,
* ask it its id, edit config, restart) repeated per lead. Naming the tab is one step, done at
* the moment the operator is already there. The convention also survives what an id does not:
* close the tab and reopen it and the id changes, while the label is retyped as-is.
*
* <p>Deliberately opt-in ({@code null} ⇒ off). Turning it on widens who resolves as
* {@link dev.ltms.bridged.auth.Role#PRIMARY}, and a config that never asked for it must not
* acquire that by upgrading the daemon.
*
* <p>bridged never writes these labels — see {@link dev.ltms.bridged.herdr.LeadTabScanner} for
* why that one-way direction is what keeps the convention trustworthy.
*
* @param tabPrefix label prefix marking a lead's tab, matched case-insensitively; the
* remainder is the lead's name ({@code "lead: opus-5.0"} → {@code
* opus-5.0}). Default {@code "lead:"}
* @param intervalSeconds how long a scan is cached before herdr is asked again; also the worst
* case before a newly-labelled tab is recognised. Default 10
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record LeadScan(String tabPrefix, Integer intervalSeconds) {
public LeadScan {
tabPrefix = (tabPrefix == null || tabPrefix.isBlank()) ? "lead:" : tabPrefix.strip();
intervalSeconds = (intervalSeconds == null || intervalSeconds <= 0) ? 10 : intervalSeconds;
}
}
/**
* The terminal → lead-name map that {@link dev.ltms.bridged.auth.CallerResolver} resolves
* against, merging the {@code leaders:} registry with the legacy singular {@code primary:} pin.
*
* <p>Precedence: an explicit {@code leaders:} entry wins over the {@code primary:} pin for the
* same terminal. The pin is the older, less expressive spelling of the same fact, so when both
* name a pane the named entry is the one an operator meant. The pin is still honoured on its
* own — a config carrying only {@code primary:} behaves exactly as it did before CB-530.
*
* @return an unmodifiable map, empty when neither block is configured (nothing is pinned, and
* every pane therefore resolves as a worker — the pre-CB-307 behaviour)
*/
public Map<String, String> leaderTerminals() {
Map<String, String> byTerminal = new LinkedHashMap<>();
if (leaders != null) {
leaders.forEach((name, leader) -> {
if (leader != null && leader.terminal() != null && !leader.terminal().isBlank()) {
byTerminal.put(leader.terminal(), name);
}
});
}
if (primary != null && primary.terminal() != null && !primary.terminal().isBlank()) {
byTerminal.putIfAbsent(primary.terminal(), "primary");
}
return Collections.unmodifiableMap(byTerminal);
}
/**
* API authentication (CB-501). Governs how a caller that is <em>not</em> an on-host worker
* pane proves it is the primary.
@@ -338,7 +532,11 @@ public record BridgedConfig(
Map<String, Worker> out = new LinkedHashMap<>();
workers.forEach((name, w) -> out.put(name,
(w.profile() == null || w.profile().isBlank()) ? w.withProfile(name) : w));
return Map.copyOf(out);
// Deliberately NOT Map.copyOf: its iteration order is salted per JVM run, which would
// discard the YAML definition order built above. Placement tie-breaks on candidate
// order (see WeightedRoundRobinPolicy), so losing it makes equal-weight placement
// non-reproducible across restarts. Unmodifiable-wrap instead of copy-and-scramble.
return Collections.unmodifiableMap(out);
}
if (worker != null) {
String name = (worker.profile() == null || worker.profile().isBlank()) ? "default" : worker.profile();
@@ -362,29 +560,192 @@ public record BridgedConfig(
return p.isEmpty() ? null : p.keySet().iterator().next();
}
private static final Logger log = LoggerFactory.getLogger(BridgedConfig.class);
private static final ObjectMapper YAML = new ObjectMapper(new YAMLFactory());
/**
* Top-level keys this version understands. Used only to warn about the rest — see
* {@link #warnUnknownTopLevelKeys}. Keep in step with the record components.
*/
private static final Set<String> KNOWN_TOP_LEVEL_KEYS = Set.of(
"bind", "herdrSocket", "worker", "workers", "defaultWorker", "guard", "worktreeRoot",
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "leaders",
"architects", "leadScan", "placement", "auth");
/** Load and validate config from {@code path}. */
public static BridgedConfig load(Path path) {
try {
BridgedConfig cfg = YAML.readValue(Files.readString(path), BridgedConfig.class);
String yaml = Files.readString(path);
warnUnknownTopLevelKeys(yaml, path);
rejectDuplicateArchitectSlots(yaml);
BridgedConfig cfg = YAML.readValue(yaml, BridgedConfig.class);
return cfg.withDefaults();
} catch (IOException e) {
throw new UncheckedIOException("cannot read bridged config at " + path, e);
}
}
/**
* Reject an {@code architects:} registry whose slot names repeat (CB-548).
*
* <p>The registry is a {@code Map} keyed by slot name, so by the time it is read duplicate keys
* have already collapsed last-wins — a duplicated slot name would silently drop one slot and the
* daemon would never know. Jackson's YAML parser does not fail on duplicate mapping keys by
* default, so duplicates are caught here, at parse time, before the map is built. Only the
* <em>top-level</em> {@code architects:} block is considered, and only its direct child keys (the
* slot names) — a nested field elsewhere, even one also named {@code architects:}, is ignored, so
* parsing of the rest of the config is unaffected.
*
* @throws IllegalStateException when two {@code architects:} entries share a slot name, naming it
*/
static void rejectDuplicateArchitectSlots(String yaml) {
try (JsonParser p = YAML.createParser(yaml)) {
if (p.nextToken() != JsonToken.START_OBJECT) {
return; // not a mapping at top level — readValue reports the malformed file
}
// Scan the TOP-LEVEL mapping only. Every other field's value (however deep, including
// any nested field also literally named "architects") is consumed whole by skipValue, so
// the loop below can only ever see the top-level field names — a nested `architects:` can
// neither suppress the real block nor be misread as one.
JsonToken t;
while ((t = p.nextToken()) != null && t != JsonToken.END_OBJECT) {
if (t == JsonToken.FIELD_NAME) {
String name = p.getCurrentName();
JsonToken value = p.nextToken();
if ("architects".equals(name)) {
if (value == JsonToken.START_OBJECT) {
rejectDuplicateChildSlotKeys(p);
}
return; // the single top-level architects block is handled; nothing more to check
}
skipValue(p, value);
}
}
} catch (IOException e) {
// Not a duplicate-name condition — let readValue report the malformed file itself.
}
}
/**
* Reject a duplicated <em>direct child</em> key of the (already-positioned) {@code architects:}
* mapping — i.e. a duplicated {@code slot name}.
*
* <p>Each slot's value is consumed whole by {@link #skipValue}, so a duplicated field <em>inside</em>
* a slot (e.g. two {@code profile:} keys, or a duplicate nested {@code architects:}) is never seen
* here and cannot masquerade as a duplicated slot name.
*
* @throws IllegalStateException when two {@code architects:} entries share a slot name, naming it
*/
private static void rejectDuplicateChildSlotKeys(JsonParser p) throws IOException {
Set<String> seen = new HashSet<>();
JsonToken t;
while ((t = p.nextToken()) != null && t != JsonToken.END_OBJECT) {
if (t == JsonToken.FIELD_NAME) {
if (!seen.add(p.getCurrentName())) {
throw new IllegalStateException("refusing to start: duplicate architect slot name '"
+ p.getCurrentName() + "' — slot names must be unique; a later entry would "
+ "silently overwrite the earlier one");
}
skipValue(p, p.nextToken()); // the slot's entire value
}
}
}
/**
* Consume the whole value that starts at {@code start}, including every nested structure, and
* leave the parser positioned just past it. Used so depth is handled structurally rather than by
* a heuristic — a nested field is never interpreted as a top-level {@code architects:}.
*/
private static void skipValue(JsonParser p, JsonToken start) throws IOException {
switch (start) {
case START_OBJECT: {
JsonToken t;
while ((t = p.nextToken()) != null && t != JsonToken.END_OBJECT) {
if (t == JsonToken.FIELD_NAME) {
skipValue(p, p.nextToken());
}
}
return;
}
case START_ARRAY: {
JsonToken t;
while ((t = p.nextToken()) != null && t != JsonToken.END_ARRAY) {
skipValue(p, t);
}
return;
}
default:
// A scalar (VALUE_* / VALUE_NULL) is already fully consumed by the nextToken that
// returned it — nothing further to skip.
}
}
/**
* Log a WARN naming any top-level key this version does not understand (CB-530).
*
* <p>Why this exists: every record here is {@code @JsonIgnoreProperties(ignoreUnknown = true)},
* which is deliberate — config must be allowed to grow ahead of the code, and a rolled-back
* daemon must still start. The cost is that a whole block can be written, parsed, dropped, and
* never mentioned again. That is exactly how a hand-written {@code leaders:} registry came to
* look configured while being inert: the daemon started, nothing complained, and the only way
* to discover it was reading the config class.
*
* <p>A warning rather than a failure, on purpose. Failing closed would turn "the config names
* something this build has not learned yet" into a daemon that will not boot — which is the
* forward-compatibility this annotation was chosen to preserve. Loud, not fatal.
*/
private static void warnUnknownTopLevelKeys(String yaml, Path path) {
List<String> unknown = unknownTopLevelKeys(yaml);
if (!unknown.isEmpty()) {
log.warn("{}: ignoring unknown top-level config key(s) {} — this build does not "
+ "understand them, so they have NO effect. Check for a typo, or a "
+ "feature not in this version.",
path, unknown);
}
}
/**
* The top-level keys in {@code yaml} that this build does not understand, sorted. Package-private
* so the guardrail is asserted directly rather than through a log appender.
*
* @return empty when everything is known, or when {@code yaml} is not a mapping at all (a
* malformed file is {@code readValue}'s error to report, not this method's)
*/
static List<String> unknownTopLevelKeys(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return List.of();
}
if (raw == null) {
return List.of();
}
return raw.keySet().stream()
.map(String::valueOf)
.filter(k -> !KNOWN_TOP_LEVEL_KEYS.contains(k))
.sorted()
.toList();
}
/** Fill in nested defaults so callers never see nulls for structural fields. */
public BridgedConfig withDefaults() {
Bind b = bind != null ? bind : new Bind(null, 0);
Guard g = guard != null ? guard : new Guard(List.of());
Lifecycle l = lifecycle != null ? lifecycle : new Lifecycle(null, null, null);
Lifecycle l = lifecycle != null ? lifecycle : new Lifecycle(null, null, null, false);
Integer timeout = (spawnReadyTimeoutMs != null) ? spawnReadyTimeoutMs : 20000;
Integer pollMs = (spawnReadyPollMs != null) ? spawnReadyPollMs : 300;
Auth a = auth != null ? auth : new Auth(null, null);
String placementOrDefault = (placement != null && !placement.isBlank()) ? placement : "fixed";
// broker is left as-is: null (or an empty/blank uri) keeps the in-memory soft-state inbox.
// primary is left as-is: null defaults to connection-derived identity.
return new BridgedConfig(b, herdrSocket, worker, workers, defaultWorker, g, worktreeRoot, l, timeout, pollMs, broker, primary, a);
// leadScan is left as-is: null is "off", and LeadScan's own compact constructor defaults the
// fields of a block that IS present. Defaulting it here would switch the feature on for
// every config that never mentioned it.
// architects is left as-is: null is "none configured", and Architect's fields have no
// defaults to fill. Defaulting it here would change nothing, so leave the call natural.
return new BridgedConfig(b, herdrSocket, worker, workers, defaultWorker, g, worktreeRoot, l, timeout, pollMs, broker, primary, leaders, architects, leadScan, placementOrDefault, a);
}
/**
@@ -411,6 +772,125 @@ public record BridgedConfig(
+ "bind to 127.0.0.1 and put a reverse proxy in front.");
}
/**
* Reject a lead-scan convention that a worker tab would also satisfy (CB-531).
*
* <p>The scan reads a tab label and concludes "a lead lives here". bridged also <em>writes</em>
* tab labels — every worker gets {@code tabLabel} rendered into its tab. Choose a
* {@code leadScan.tabPrefix} that a worker template matches and the daemon starts labelling its
* own workers as leads, promoting the entire fleet to {@link dev.ltms.bridged.auth.Role#PRIMARY}
* with no message and no diff. The worker-space exclusion in
* {@link dev.ltms.bridged.herdr.LeadTabScanner} already blocks the realistic path, but defence
* that depends on one workspace label holding is not defence enough for a privilege boundary.
*
* <p>Fatal rather than a warning, unlike {@link #warnUnknownTopLevelKeys}: an unknown key means
* a feature does nothing, while this means a feature does the opposite of what it says.
*
* @throws IllegalStateException when any worker profile's {@code tabLabel} starts with the
* configured lead prefix
*/
public void validateLeadScan() {
if (leadScan == null) {
return;
}
String prefix = leadScan.tabPrefix();
List<String> clashing = workerProfiles().entrySet().stream()
.filter(e -> e.getValue().tabLabel() != null
&& e.getValue().tabLabel().strip()
.regionMatches(true, 0, prefix, 0, prefix.length()))
.map(Map.Entry::getKey)
.sorted()
.toList();
if (clashing.isEmpty()) {
return;
}
throw new IllegalStateException(
"refusing to start: leadScan.tabPrefix=\"" + prefix + "\" also matches the tabLabel "
+ "of worker profile(s) " + clashing + ". Every worker spawned under them "
+ "would be read back as a lead and granted spawn/stop/send on the whole "
+ "fleet. Change one of the two so worker tabs and lead tabs cannot be "
+ "confused.");
}
/**
* Reject a subscription profile whose {@code env:} block tries to reseat the Anthropic binding
* (CB-542).
*
* <p>Why this must be fatal rather than sanitised: {@code subscription: true} deliberately
* stops the launcher from writing {@code ANTHROPIC_BASE_URL}/{@code ANTHROPIC_AUTH_TOKEN} and
* skips the {@code SubscriptionGuard} for that profile. But the profile's {@code env:} map is
* layered into the worker environment separately, so an {@code ANTHROPIC_BASE_URL} sitting
* there would survive into the worker having passed no guard at all — {@code subscription: true}
* plus an {@code env:} repoint is a contradiction just like {@code subscription: true} plus a
* {@code baseUrl}. The launcher also hard-strips these two keys from the worker env as a
* belt-and-braces measure; this method is the loud, load-time refusal so the operator is told
* about the mistake instead of having it silently cleaned up.
*
* @throws IllegalStateException when any subscription profile's {@code env:} names
* {@code ANTHROPIC_BASE_URL} or {@code ANTHROPIC_AUTH_TOKEN},
* naming the profile and the offending key(s)
*/
public void validateSubscriptionProfiles() {
List<String> bad = new java.util.ArrayList<>();
workerProfiles().forEach((name, w) -> {
if (w.isSubscription() && w.envCarriesAnthropicBinding()) {
List<String> keys = w.env().keySet().stream()
.filter(k -> k.equals("ANTHROPIC_BASE_URL") || k.equals("ANTHROPIC_AUTH_TOKEN"))
.sorted()
.toList();
bad.add("worker profile '" + name + "' carries " + keys + " in env: — "
+ "subscription: true forbids ANTHROPIC_BASE_URL / ANTHROPIC_AUTH_TOKEN there, "
+ "because on the subscription no guard vets them (they would repoint the "
+ "worker past the SubscriptionGuard). Remove them from env: (or drop "
+ "subscription: true).");
}
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
}
/**
* Reject an architect slot whose profile reference does not resolve (CB-548).
*
* <p>An architect's {@code profile} is the strong-model {@code workers:} profile the future
* spawn lifecycle will read to stand the slot up. A reference that names no configured profile
* is a typo or a stale config — and unlike a worker spawn (which fails loudly at its call site
* when it cannot resolve), an architect slot fails only when something later tries to use it.
* This config class <em>does</em> have access to {@link #workerProfiles()}, so the reference
* is validated at startup and the mistake is named then, not discovered months later by a
* spawn that quietly has no backend to use.
*
* <p>Slot-name uniqueness needs no check here: the registry is a {@code Map} keyed by name, so
* duplicates are unrepresentable by construction once loaded — and {@link #load(Path)} already
* rejects a duplicated slot name at parse time, before the map collapses.
*
* @throws IllegalStateException when any architect slot is missing or names an unknown profile,
* naming the slot and the offending reference
*/
public void validateArchitects() {
if (architects == null) {
return;
}
Map<String, Worker> profiles = workerProfiles();
List<String> bad = new java.util.ArrayList<>();
architects.forEach((name, arch) -> {
if (arch == null || arch.profile() == null || arch.profile().isBlank()) {
bad.add("architect slot '" + name + "' has no profile: — give it the name of a "
+ "workers: profile (the strong-model backend it runs).");
return;
}
if (!profiles.containsKey(arch.profile())) {
bad.add("architect slot '" + name + "' references profile '" + arch.profile()
+ "', which is not a configured workers: profile (have: " + profiles.keySet()
+ ").");
}
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
}
/** True for the loopback addresses and the unspecified-but-local forms we treat as same-host. */
private static boolean isLoopbackBind(String host) {
if (host == null || host.isBlank()) {
@@ -6,93 +6,120 @@ import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
/**
* Domain layer over herdr's native {@code agent.*} namespace — the worker south side.
* Chosen in the CB-102 spike over the pane + {@code send_text} fallback because
* {@code agent.start} takes a first-class {@code env} map (clean, guard-checked
* subscription injection) and herdr tracks each worker's Claude session UUID itself.
* Chosen in the CB-102 spike over the pane + {@code send_text} fallback because herdr
* tracks each worker's Claude session UUID itself.
*
* <p>Ported to herdr protocol 19 (herdr 0.8.0, CB-521): {@code agent.start} now starts a
* <em>supported</em> agent ({@code kind}) into an <em>existing</em> pane, so the worker's
* {@code env}/{@code cwd} move to pane creation ({@code tab.create}/{@code pane.split} — see
* {@link WorkspaceControl}), and {@code agent.send} is replaced by {@code agent.prompt}
* (which submits in one call) plus {@code agent.send_keys} for the raw Enter nudge.
*
* <p>Every method is one herdr call through the injected {@link HerdrClient}, so this
* layer is unit-testable with a fake and contract-tested against a live daemon.
*/
public final class AgentControl {
/**
* The keystroke that submits a prompt in the Claude Code TUI: a carriage return (Enter).
* It must be delivered as its <em>own</em> {@code agent.send} call — herdr delivers a message's
* text as a bracketed paste, and a {@code "\r"} appended to that same text is swallowed as
* literal newline content, not a submit. Sent as a separate keystroke event it lands outside
* the paste and submits. (A bare {@code "\n"} inserts a newline either way.) Verified live
* against Claude Code v2.1.210: an injected task stayed unsubmitted with {@code "text\r"} in
* one call, and submitted the instant a standalone {@code "\r"} was sent.
*/
static final String SUBMIT_KEY = "\r";
private final HerdrClient herdr;
/**
* Protocol 19 dropped {@code terminal_id} as an {@code agent.*} target — herdr now resolves
* targets by pane id or agent name only, while the bridge keys every session on the terminal.
* This caches the terminal→pane mapping (stable for a worker's lifetime) so callers keep
* addressing agents by terminal; entries are invalidated on {@code agent_not_found}.
*/
private final Map<String, String> paneByTerminal = new ConcurrentHashMap<>();
public AgentControl(HerdrClient herdr) {
this.herdr = herdr;
}
/** One agent-targeted call, translating a terminal id to its pane id (retrying once fresh). */
private JsonNode agentCall(String method, String target, Map<String, Object> extra) {
String resolved = resolveTarget(target);
try {
return herdr.call(method, withTarget(resolved, extra));
} catch (HerdrException e) {
if (!"agent_not_found".equals(e.code()) || resolved.equals(target)) throw e;
paneByTerminal.remove(target); // the cached pane went away — re-resolve once
String fresh = resolveTarget(target);
if (fresh.equals(resolved)) throw e;
return herdr.call(method, withTarget(fresh, extra));
}
}
private static Map<String, Object> withTarget(String target, Map<String, Object> extra) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("target", target);
m.putAll(extra);
return m;
}
/** The pane id behind a terminal-id target, or the target verbatim for pane ids / names. */
private String resolveTarget(String target) {
if (target == null || !target.startsWith("term_")) {
return target;
}
String cached = paneByTerminal.get(target);
if (cached != null) {
return cached;
}
for (JsonNode a : herdr.call("agent.list").path("agents")) {
if (target.equals(a.path("terminal_id").asText(null))) {
String pane = a.path("pane_id").asText(null);
if (pane != null) {
paneByTerminal.put(target, pane);
return pane;
}
}
}
return target; // unknown terminal — let herdr report it against the original target
}
/**
* Spawn an agent. {@code env} is applied to the process environment verbatim — this
* is where a worker's {@code ANTHROPIC_BASE_URL} lives, and the ONLY place it should.
* Start an agent into {@code paneId}, which must be sitting at its interactive shell prompt —
* the seed pane of a freshly-created worker tab, or a fresh split. The pane's shell already
* carries the worker's env ({@code ANTHROPIC_BASE_URL}, token, …) and cwd from pane creation;
* herdr resolves the executable from {@code kind} and waits (its default timeout) until the
* agent is detected and ready for input.
*
* @param name label/kind for herdr status detection (e.g. {@code "claude"})
* @param argv launch command, e.g. {@code ["claude"]}
* @param env process environment additions ({@code ANTHROPIC_BASE_URL}, token, …)
* @param name unique label for this agent ({@code <kind>-<profile>-<nonce>-<seq>})
* @param kind supported agent kind and canonical executable, e.g. {@code "claude"},
* {@code "opencode"}
* @param args extra arguments after the executable, e.g. {@code --mcp-config …}
* @param paneId the pane to start the agent in
*/
public Agent start(String name, List<String> argv, Map<String, String> env) {
return start(name, argv, env, null);
}
/** Spawn an agent into {@code tabId} at herdr's default cwd. */
public Agent start(String name, List<String> argv, Map<String, String> env, String tabId) {
return start(name, argv, env, tabId, null);
}
/**
* Spawn an agent. With a non-null {@code tabId} the worker lands in that tab (the placement
* policy's dedicated worker tab); with {@code null} herdr splits the currently-focused tab
* (legacy pane placement). A non-blank {@code cwd} sets the worker process's working directory —
* {@code agent.start} honours {@code cwd} directly (an agent pane does <em>not</em> inherit the
* tab's or workspace's cwd, so this is the only way to root a worker in the primary's directory;
* CB-112).
*/
public Agent start(String name, List<String> argv, Map<String, String> env, String tabId, String cwd) {
public Agent start(String name, String kind, List<String> args, String paneId) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("name", name);
params.put("argv", argv);
params.put("env", env);
if (tabId != null) {
params.put("tab_id", tabId);
}
if (cwd != null && !cwd.isBlank()) {
params.put("cwd", cwd);
}
params.put("kind", kind);
params.put("pane_id", paneId);
params.put("args", args);
JsonNode result = herdr.call("agent.start", params);
return Agent.from(result.get("agent"));
}
/**
* Deliver {@code text} to an agent as its next prompt <em>and submit it</em> — two keystroke
* events: the message (a bracketed paste, so any embedded newlines are preserved verbatim),
* then a standalone {@link #SUBMIT_KEY} (Enter) that actually submits it. Without the second
* event the text just sits in the worker's input box, never processed (see {@link #SUBMIT_KEY}).
* Deliver {@code text} to an agent as its next prompt <em>and submit it</em> — herdr's
* {@code agent.prompt} pastes the text (embedded newlines preserved verbatim) and submits it
* in the same call, replacing the pre-protocol-19 two-event {@code agent.send} dance.
*/
public void send(String target, String text) {
herdr.call("agent.send", Map.of("target", target, "text", text));
herdr.call("agent.send", Map.of("target", target, "text", SUBMIT_KEY));
agentCall("agent.prompt", target, Map.of("text", text));
}
/**
* Re-send the submit keystroke (Enter) to {@code target}. The Enter that accompanies a delivery
* can race the paste — especially right as the worker's TUI becomes interactive — leaving the
* text unsubmitted; the injector nudges it with this until the worker actually picks up (CB-113).
* Re-send the submit keystroke (Enter) to {@code target}. The submit that accompanies a
* delivery can race the paste — especially right as the worker's TUI becomes interactive —
* leaving the text unsubmitted; the injector nudges it with this until the worker actually
* picks up (CB-113).
*/
public void submit(String target) {
herdr.call("agent.send", Map.of("target", target, "text", SUBMIT_KEY));
agentCall("agent.send_keys", target, Map.of("keys", List.of("enter")));
}
/**
@@ -101,13 +128,13 @@ public final class AgentControl {
* @param source one of {@code visible|recent|recent_unwrapped|detection}
*/
public String read(String target, String source) {
JsonNode result = herdr.call("agent.read", Map.of("target", target, "source", source));
JsonNode result = agentCall("agent.read", target, Map.of("source", source));
return result.path("read").path("text").asText("");
}
/** Current agent record (status, session UUID, pane). */
public Agent get(String target) {
return Agent.from(herdr.call("agent.get", Map.of("target", target)).get("agent"));
return Agent.from(agentCall("agent.get", target, Map.of()).get("agent"));
}
/** Just the lifecycle status — what the status-gated injector checks before send. */
@@ -0,0 +1,164 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Set;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* Discovers which panes host a lead by scanning herdr for tabs the operator labelled by convention
* (CB-531), and hands {@link dev.ltms.bridged.auth.CallerResolver} the resulting
* {@code terminal_id → lead name} map.
*
* <p><strong>Why scan at all.</strong> A lead is never spawned — a human opens a tab and starts an
* agent in it — so the daemon cannot learn a lead's {@code terminal_id} at creation time the way it
* does a worker's. CB-530 solved that by having the operator paste each id into {@code leaders:},
* which works but costs a config edit and a daemon restart per lead, and the id is only obtainable
* by first starting the session and asking it. Scanning closes that loop: label the tab, and the
* pane is recognised on the next resolve.
*
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
* anything a pane could take for itself. Three properties keep that honest:
* <ol>
* <li>bridged never renames a lead tab. The operator's label is read-only input, so what is in
* the tab bar is always what the human wrote — no round-trip where the daemon's own rename
* becomes the evidence for its next decision.</li>
* <li>Worker spaces are excluded wholesale ({@code excludedWorkspaceLabels}), so a worker cannot
* become a lead by being placed — as a split, say — inside a matching tab.</li>
* <li>A worker cannot rename a tab: {@code tab.rename} is reachable only through
* {@link WorkspaceControl}, which no {@code bridge_*} tool exposes. The label is writable by
* the human at the terminal and by nobody the bridge is defending against.</li>
* </ol>
* The remaining hazard is an <em>operator</em> one — a worker {@code tabLabel} template that
* happens to start with the same prefix would promote the whole fleet — and that is refused at
* startup by {@code BridgedConfig.validateLeadScan} rather than documented here.
*
* <p><strong>Caching.</strong> {@link #get()} is on the request path (every resolve), so the scan
* is TTL-cached and a stale-but-valid map is preferred to a herdr round-trip. A failed scan keeps
* the previous answer instead of emptying it — a herdr hiccup must not silently demote a live lead
* mid-session.
*/
public final class LeadTabScanner implements Supplier<Map<String, String>> {
private static final Logger log = LoggerFactory.getLogger(LeadTabScanner.class);
private final HerdrClient herdr;
private final String tabPrefix;
private final Set<String> excludedWorkspaceLabels;
private final Map<String, String> configuredLeads;
private final long ttlNanos;
private final LongSupplier clock;
private Map<String, String> cached;
private long scannedAtNanos;
private boolean everScanned;
/**
* @param herdr the herdr client to query ({@code workspace.list},
* {@code tab.list}, {@code pane.list} — all read-only)
* @param tabPrefix a tab whose label starts with this (case-insensitively) hosts a
* lead; the rest of the label, trimmed, is the lead's name
* @param excludedWorkspaceLabels workspaces never scanned — the configured worker spaces
* @param configuredLeads the static {@code leaders:}/{@code primary:} registry, merged
* over every scan result. Explicit config outranks the
* convention, and survives a scan that cannot run at all
* @param ttlNanos how long a scan result is reused before the next one
* @param clock nanosecond time source ({@code System::nanoTime} in production)
*/
public LeadTabScanner(HerdrClient herdr, String tabPrefix, Set<String> excludedWorkspaceLabels,
Map<String, String> configuredLeads, long ttlNanos, LongSupplier clock) {
this.herdr = herdr;
this.tabPrefix = tabPrefix == null || tabPrefix.isBlank() ? "lead:" : tabPrefix.strip();
this.excludedWorkspaceLabels = excludedWorkspaceLabels == null
? Set.of() : Set.copyOf(excludedWorkspaceLabels);
this.configuredLeads = configuredLeads == null ? Map.of() : Map.copyOf(configuredLeads);
this.ttlNanos = ttlNanos;
this.clock = clock;
this.cached = this.configuredLeads;
}
/**
* The current {@code terminal_id → lead name} map, rescanning when the cache has expired.
*
* <p>Synchronized so a burst of concurrent calls produces one scan rather than one each; a scan
* is a handful of RPCs over a Unix socket and is rate-limited to one per TTL.
*/
@Override
public synchronized Map<String, String> get() {
long now = clock.getAsLong();
if (everScanned && now - scannedAtNanos < ttlNanos) {
return cached;
}
// Stamp before scanning, not after: a herdr that is down must cost one attempt per TTL, not
// one per request.
scannedAtNanos = now;
everScanned = true;
try {
Map<String, String> fresh = scan();
if (!fresh.equals(cached)) {
log.info("lead panes: {}", fresh);
}
cached = fresh;
} catch (HerdrException e) {
log.warn("lead-tab scan failed, keeping the {} lead(s) already known: {}",
cached.size(), e.getMessage());
}
return cached;
}
/** One full pass: labelled tabs → their panes → those panes' terminals. */
private Map<String, String> scan() {
Map<String, String> nameByTab = new LinkedHashMap<>();
for (JsonNode w : herdr.call("workspace.list").path("workspaces")) {
Workspace ws = Workspace.from(w);
if (ws.workspaceId() == null || excludedWorkspaceLabels.contains(ws.label())) {
continue;
}
for (JsonNode t : herdr.call("tab.list", Map.of("workspace_id", ws.workspaceId())).path("tabs")) {
Tab tab = Tab.from(t);
String name = leadNameOf(tab.label());
if (name != null && tab.tabId() != null) {
nameByTab.put(tab.tabId(), name);
}
}
}
Map<String, String> byTerminal = new LinkedHashMap<>();
if (!nameByTab.isEmpty()) {
// One pane.list for every tab: panes carry tab_id, so the join is local.
for (JsonNode p : herdr.call("pane.list", Map.of()).path("panes")) {
String name = nameByTab.get(p.path("tab_id").asText(null));
String terminal = p.path("terminal_id").asText(null);
if (name != null && terminal != null && !terminal.isBlank()) {
byTerminal.put(terminal, name);
}
}
}
byTerminal.putAll(configuredLeads); // an explicit pin outranks a label
return Collections.unmodifiableMap(byTerminal);
}
/**
* The lead name a tab label declares, or {@code null} if it declares none.
*
* <p>{@code "lead: opus-5.0"} → {@code "opus-5.0"}. A bare {@code "lead:"} names nobody and is
* rejected: an unnamed lead would resolve as {@code PRIMARY} with nothing to attribute it to.
*/
private String leadNameOf(String label) {
if (label == null) {
return null;
}
String l = label.strip();
if (!l.regionMatches(true, 0, tabPrefix, 0, tabPrefix.length())) {
return null;
}
String name = l.substring(tabPrefix.length()).strip();
return name.isEmpty() ? null : name;
}
}
@@ -23,9 +23,9 @@ public record Tab(String tabId, String workspaceId, String label, int paneCount)
}
/**
* A freshly-created tab together with the placeholder shell pane herdr seeds it with.
* The caller starts the worker into {@link #tab()} then closes {@link #rootPaneId()} so
* only the worker pane remains.
* A freshly-created tab together with the shell pane herdr seeds it with. Under protocol 19
* the caller starts the worker <em>into</em> {@link #rootPaneId()} — the seed pane's shell
* carries the worker's cwd and env from {@code tab.create}, and becomes the worker pane.
*/
public record Created(Tab tab, String rootPaneId) {
/** Project a {@code tab_created} result ({@code {tab, root_pane}}). */
@@ -5,6 +5,7 @@ import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
@@ -65,13 +66,37 @@ public final class WorkspaceControl {
}
/**
* A brand-new tab in {@code workspaceId} plus the placeholder shell pane herdr seeds it with.
* Start the worker into the tab, then {@code pane.close} the root pane so the tab holds only the
* worker. (The worker's own cwd is set on {@code agent.start}, not here — an {@code agent.start}
* pane does not inherit the tab's cwd; see {@code AgentControl.start}.)
* A brand-new tab in {@code workspaceId} plus the shell pane herdr seeds it with. Under
* protocol 19 that seed pane is where the worker <em>starts</em>: its shell carries
* {@code cwd} and {@code env} (the worker's {@code ANTHROPIC_BASE_URL} — this is the
* subscription-injection seam now), and {@code agent.start} launches the agent into it.
*/
public Tab.Created createTab(String workspaceId) {
return Tab.Created.from(herdr.call("tab.create", Map.of("workspace_id", workspaceId)));
public Tab.Created createTab(String workspaceId, String cwd, Map<String, String> env) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("workspace_id", workspaceId);
if (cwd != null && !cwd.isBlank()) {
params.put("cwd", cwd);
}
if (env != null && !env.isEmpty()) {
params.put("env", env);
}
return Tab.Created.from(herdr.call("tab.create", params));
}
/**
* Split the currently-focused tab and return the new pane's id — the legacy pane placement's
* seed pane, carrying {@code cwd} and {@code env} exactly as {@link #createTab}'s does.
*/
public String splitPane(String cwd, Map<String, String> env) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("direction", "right");
if (cwd != null && !cwd.isBlank()) {
params.put("cwd", cwd);
}
if (env != null && !env.isEmpty()) {
params.put("env", env);
}
return herdr.call("pane.split", params).path("pane").path("pane_id").asText(null);
}
/** Give a worker's tab a human label in the tab bar. */
@@ -122,6 +122,15 @@ public final class CompletionResolver implements TurnListener {
Thread.ofVirtual().name("completion-" + target).start(() -> resolve(target, turn));
}
/**
* Resolve the completed turn before adapter housekeeping can erase its rendered output. This is
* intentionally synchronous and used only when a post-turn context reset is enabled; the normal
* path remains off-loaded so polling is not blocked by a scrape.
*/
public void resolveBeforePostAction(String target) {
resolve(target, inFlight.get(target));
}
@Override
public void onTurnFailed(String target) {
InFlight turn = inFlight.get(target);
@@ -132,6 +132,10 @@ public final class Injector {
boolean turnObserved; // saw a real `working` sample since that delivery (turn ran)
int unknownSinceTurn; // consecutive `unknown` samples while a delegation is outstanding (CB-109)
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
boolean postTurnPending; // completion observed; adapter housekeeping has not started yet
boolean awaitingPostTurnPickup;
boolean postTurnObserved;
int injectableSincePostTurnPickup;
synchronized void add(Pending p) {
queue.add(p);
@@ -173,9 +177,15 @@ public final class Injector {
boolean turnCompleted = false;
boolean turnFailed = false;
boolean resubmit = false;
boolean startPostTurn = false;
List<Pending> notReady = null; // queued messages failed because the worker never became ready
synchronized (t) {
if (status == AgentStatus.WORKING) {
if (t.awaitingPostTurnPickup) {
t.awaitingPostTurnPickup = false;
t.injectableSincePostTurnPickup = 0;
t.postTurnObserved = true;
}
// Definitive pickup: the worker is busy on our last message, and (if a delivery is
// outstanding) a real turn is now confirmed to be running.
t.awaitingPickup = false;
@@ -185,6 +195,16 @@ public final class Injector {
if (t.awaitingCompletion) t.turnObserved = true;
} else if (status.injectable()) { // IDLE or BLOCKED
t.unknownSinceTurn = 0;
if (t.awaitingPostTurnPickup) {
if (++t.injectableSincePostTurnPickup >= PICKUP_GRACE_POLLS) {
t.awaitingPostTurnPickup = false;
t.injectableSincePostTurnPickup = 0;
} else {
resubmit = true;
}
} else if (t.postTurnObserved) {
t.postTurnObserved = false;
}
if (t.awaitingPickup) {
if (++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
// Pickup edge was never sampled (turn faster than the poll, or status lag).
@@ -209,11 +229,16 @@ public final class Injector {
t.awaitingCompletion = false;
t.turnObserved = false;
turnCompleted = true;
if (turnListener.hasPostTurnAction(target)) {
t.postTurnPending = true;
startPostTurn = true;
}
}
// Deliver the next queued message only once the prior turn is fully settled, so a
// completion is never confused with the pickup of the following message — and only
// once the worker is available (CB-113), so we never paste into its boot window.
if (!t.awaitingCompletion) {
if (!t.awaitingCompletion && !t.postTurnPending
&& !t.awaitingPostTurnPickup && !t.postTurnObserved) {
Pending p = t.queue.peek();
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
@@ -263,7 +288,8 @@ public final class Injector {
// Reclaim the entry once the worker is fully quiescent (nothing queued, no pickup or
// completion awaited), so the map cannot grow without bound across short-lived workers.
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion) {
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion
&& !t.postTurnPending && !t.awaitingPostTurnPickup && !t.postTurnObserved) {
targets.remove(target, t);
}
}
@@ -290,7 +316,21 @@ public final class Injector {
turnListener.onTurnFailed(target);
}
if (turnCompleted) {
turnListener.onTurnComplete(target);
if (startPostTurn) {
boolean started = turnListener.onTurnCompleteWithPostAction(target);
synchronized (t) {
t.postTurnPending = false;
if (started) {
t.awaitingPostTurnPickup = true;
t.injectableSincePostTurnPickup = 0;
}
if (t.queue.isEmpty() && !t.awaitingPostTurnPickup) {
targets.remove(target, t);
}
}
} else {
turnListener.onTurnComplete(target);
}
}
if (turnFailed) {
turnListener.onTurnFailed(target);
@@ -317,7 +357,8 @@ public final class Injector {
.filter(e -> {
synchronized (e.getValue()) {
Target t = e.getValue();
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion;
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion
|| t.postTurnPending || t.awaitingPostTurnPickup || t.postTurnObserved;
}
})
.map(java.util.Map.Entry::getKey)
@@ -343,6 +384,9 @@ public final class Injector {
hadDeliveredTurn = t.awaitingCompletion;
t.awaitingCompletion = false;
t.awaitingPickup = false;
t.postTurnPending = false;
t.awaitingPostTurnPickup = false;
t.postTurnObserved = false;
}
forget.accept(target); // the worker is gone — clear its readiness/presence too (CB-114)
for (Pending p : pending) {
@@ -13,6 +13,25 @@ public interface TurnListener {
/** A worker's delegated turn finished (worker returned to idle after visibly working). */
void onTurnComplete(String target);
/**
* Whether completion must pause ordinary delivery while adapter-specific housekeeping starts.
* This is queried before the injector considers the next queued message, closing the same-tick
* ordering gap. Implementations must not mutate state here.
*/
default boolean hasPostTurnAction(String target) {
return false;
}
/**
* Complete the delegated turn and start its non-turn housekeeping operation.
*
* @return true when the injector must observe that operation settle before the next delivery
*/
default boolean onTurnCompleteWithPostAction(String target) {
onTurnComplete(target);
return false;
}
/**
* A worker that visibly ran a delegated turn then wedged in a non-idle, non-working state
* (CB-109) — e.g. an error screen herdr classifies as {@code unknown} — so no
@@ -65,6 +65,8 @@ public final class BridgeMcp {
static final String CALLER_PID = "callerPid";
/** Transport-context key under which the extractor stashes the resolved {@link Role} (CB-501). */
static final String CALLER_ROLE = "callerRole";
/** Transport-context key for the lead's configured name, when the caller is one (CB-530). */
static final String CALLER_NAME = "callerName";
private final HttpServletStreamableServerTransportProvider transport;
private final McpSyncServer server;
@@ -106,11 +108,15 @@ public final class BridgeMcp {
? callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization"))
: legacyPrincipal(identity, req.getRemoteAddr(), req.getRemotePort());
presence.markPresent(p.terminal()); // no-op for the primary (null terminal)
// CB-532: guard on the ROLE, not on the terminal being null. A named lead now
// carries its pane too, and enrolling a lead in the worker presence map would
// have it counted as an available worker.
if (p.isWorker()) presence.markPresent(p.terminal());
return McpTransportContext.create(Map.of(
CALLER_TERMINAL, orEmpty(p.terminal()),
CALLER_PID, Long.toString(p.pid()),
CALLER_ROLE, p.role().name()));
CALLER_ROLE, p.role().name(),
CALLER_NAME, orEmpty(p.name())));
})
.build();
this.server = McpServer.sync(transport)
@@ -121,18 +127,31 @@ public final class BridgeMcp {
str(req.arguments(), "sessionId"));
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
// CB-548: only a PRIMARY caller may claim the legacy singleton "primary" fallback.
// An architect delegates as its own pane but must never become the fallback that
// no-delegation inbox nudges target as if it were the primary (the per-target
// delegation map does not cure the singleton).
recordPrimarySingleton(primaryRegistry, caller, principal(exchange));
Map<String, Object> a = req.arguments();
String target = str(a, "sessionId");
String content = str(a, "content");
String turnId = str(a, "turnId");
if (turnId != null && !turnId.isBlank()) {
// Answering a worker's bridge_ask (CB-205): resolve its blocked question and
// block for the worker's reply as it resumes the same turn.
return answer(messages, turnId, str(a, "content"), timeoutMs(a));
// block for the worker's reply as it resumes the same turn. This is the same
// delegation, so ownership is left untouched (CB-548) — never re-recorded.
return answer(messages, turnId, content, timeoutMs(a));
}
// CB-548: delegator ownership (which lead's reply nudge this worker routes to,
// CB-532) is recorded only once the send is ACCEPTED — MessageService has won the
// session lock and queued delivery — via the accepted-delivery callback, never at
// request time. A concurrent sender that times out BUSY therefore cannot steal a
// live turn's reply routing without ever owning the turn.
Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, caller);
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, str(a, "sessionId"), str(a, "content"))
: send(messages, str(a, "sessionId"), str(a, "content"), timeoutMs(a));
? sendAsync(messages, target, content, onAccepted)
: send(messages, target, content, timeoutMs(a), onAccepted);
})
// bridge_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
@@ -173,7 +192,9 @@ public final class BridgeMcp {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SPAWN, null);
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
// SPAWN is already auth-gated to PRIMARY (architects can never call it), but
// enforce the same invariant here: only a PRIMARY may claim the legacy singleton.
recordPrimarySingleton(primaryRegistry, caller, principal(exchange));
Map<String, Object> a = req.arguments();
// CB-112: worker inherits the primary's cwd unless the call pins one.
// CB-301: carry the caller's identity as the session owner (null for the primary).
@@ -185,7 +206,9 @@ public final class BridgeMcp {
.toolCall(listTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listWorkers(workers, sessions);
return listFleet(workers, sessions,
callers == null ? Map.of() : callers.leads(),
callerTerminal(exchange));
})
.toolCall(stopTool(), (exchange, req) -> {
String paneId = str(req.arguments(), "paneId");
@@ -222,7 +245,7 @@ public final class BridgeMcp {
/** The caller reconstructed from the transport context. */
private static Principal principal(McpSyncServerExchange exchange) {
return principalFrom(exchange.transportContext().get(CALLER_ROLE),
callerTerminal(exchange), callerPid(exchange));
callerTerminal(exchange), callerPid(exchange), callerName(exchange));
}
/**
@@ -237,11 +260,16 @@ public final class BridgeMcp {
* @param pid the calling pid, or {@code -1}
*/
static Principal principalFrom(Object role, String terminal, long pid) {
return principalFrom(role, terminal, pid, null);
}
/** As {@link #principalFrom(Object, String, long)}, carrying a lead's name (CB-530). */
static Principal principalFrom(Object role, String terminal, long pid, String name) {
if (role == null) {
// No role stashed (legacy path): fall back to the historical interpretation.
return terminal != null ? Principal.worker(terminal, pid) : Principal.primary(pid);
}
return new Principal(Role.valueOf(role.toString()), terminal, pid);
return new Principal(Role.valueOf(role.toString()), terminal, pid, name);
}
/**
@@ -285,6 +313,24 @@ public final class BridgeMcp {
return error(reason + ": " + caller.describe() + " may not " + action);
}
/**
* Update the legacy singleton "primary" fallback used for no-delegation inbox nudges (CB-548).
*
* <p>Only {@link Role#PRIMARY} callers — the unnamed primary and named leads alike — may claim
* it. An architect delegates as its own pane but must never become the fallback: the per-target
* delegation map ({@code PrimaryRegistry#recordDelegation}) does not cure the singleton, so an
* architect left here would draw nudges that belong to a primary. The decision uses the resolved
* role, never name/kind sniffing. A null {@code caller} (legacy/no-auth path) records nothing.
*
* <p>Split out of the tool handlers so the guard is unit-testable without fabricating an SDK
* {@code McpSyncServerExchange} (same pattern as {@link #denyFor}/{@link #principalFrom}).
*/
static void recordPrimarySingleton(PrimaryRegistry registry, String callerTerminal, Principal caller) {
if (caller != null && caller.isPrimary()) {
registry.record(callerTerminal);
}
}
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
private static String callerTerminal(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_TERMINAL);
@@ -292,6 +338,13 @@ public final class BridgeMcp {
return (s == null || s.isBlank()) ? null : s;
}
/** The lead name resolved from this call's connection, or {@code null} (CB-530). */
private static String callerName(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_NAME);
String s = v == null ? null : v.toString();
return (s == null || s.isBlank()) ? null : s;
}
/** The caller's PID resolved from this call's connection, or {@code -1} if unknown. */
private static long callerPid(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_PID);
@@ -320,12 +373,22 @@ public final class BridgeMcp {
/** {@code bridge_send}: delegate {@code content} to a worker session and block for its reply. */
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content, Long timeoutMs) {
return send(messages, sessionId, content, timeoutMs, null);
}
/**
* As {@link #send(MessageService, String, String, Long)}, wiring an accepted-delivery hook
* (CB-548): {@code onAccepted} records delegator ownership the instant the send is accepted, so
* a BUSY interloper never claims a turn it did not win. {@code null} disables recording.
*/
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content,
Long timeoutMs, Runnable onAccepted) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
try {
return formatReply(messages.send(sessionId, content, timeout), timeout);
return formatReply(messages.send(sessionId, content, timeout, onAccepted), timeout);
} catch (HerdrException e) {
return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
}
@@ -394,10 +457,19 @@ public final class BridgeMcp {
* immediately (fire-and-poll), so a long task isn't cut off by the caller's MCP call timeout.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content) {
return sendAsync(messages, sessionId, content, null);
}
/**
* As {@link #sendAsync(MessageService, String, String)}, wiring the accepted-delivery hook
* (CB-548) so an async flooding send records delegator ownership exactly once it is accepted.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content,
Runnable onAccepted) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
String ticket = messages.sendAsync(sessionId, content);
String ticket = messages.sendAsync(sessionId, content, onAccepted);
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
}
@@ -484,7 +556,29 @@ public final class BridgeMcp {
static McpSchema.CallToolResult whoami(Principal caller, SessionManager sessions) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("role", caller.role().name().toLowerCase());
if (caller.isArchitect()) {
// CB-548: the role reads "architect"; the name is the gateway-local slot the pane is
// bound to, and the pane itself so a peer knows where to reach it.
if (caller.name() != null) {
m.put("architect", caller.name());
}
if (caller.terminal() != null) {
m.put("sessionId", caller.terminal());
}
return text(json(m));
}
if (!caller.isWorker()) {
// CB-530: which lead, once more than one pane is configured as one. `role` deliberately
// still reads "primary" — the fallback ladder in CLAUDE.md keys on it, and a lead IS a
// primary for authorization; the name is additive so no existing reader breaks.
if (caller.name() != null) {
m.put("leader", caller.name());
}
// CB-532: a lead's own pane, so it can tell a peer where to reach it — and so an
// operator can read off which tab hosts which lead without going to herdr.
if (caller.terminal() != null) {
m.put("sessionId", caller.terminal());
}
return text(json(m));
}
m.put("sessionId", caller.terminal());
@@ -575,22 +669,66 @@ public final class BridgeMcp {
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
}
/** {@code bridge_list}: bridge-owned roster merged with live herdr status by paneId. */
static McpSchema.CallToolResult listWorkers(PeerLauncher workers, SessionManager sessions) {
/**
* {@code bridge_list}: the whole fleet — {@code leads} and {@code workers} — each merged with
* live herdr status. CB-519 decoupled the registry key (a host-unique id) from the herdr pane
* coordinate, so the join is on the terminal id, which both the session and the live agent carry.
*
* <p>CB-535 added the {@code leads} half. Until then this listed the worker roster alone, and a
* lead asking "who else is here?" got an empty array — which reads as <em>no peers</em> but
* actually means <em>no workers spawned</em>. There was no way at all for a lead to learn a
* peer's address; it had to be carried across by a human. Both halves are reported even when a
* half is empty, so an empty {@code workers} can no longer be mistaken for an empty fleet.
*
* <p>Leads are drawn from the resolver rather than from a second registry, so an address listed
* here is one that would actually resolve as a lead — see {@link CallerResolver#leads()}. The
* caller's own row is flagged {@code "self": true}: a peer needs to tell its own pane apart from
* a peer's, and the alternative is every lead calling {@code bridge_whoami} to subtract itself.
*
* @param leads terminal_id → lead name, live from the resolver
* @param selfTerm the calling pane's terminal id, or blank for a caller with no pane
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions,
Map<String, String> leads, String selfTerm) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.paneId() != null)
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.paneId())))
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> leadRows = leads.entrySet().stream()
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm))
.toList();
return text(json(Map.of("workers", out)));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
return text(json(Map.of("leads", leadRows, "workers", out)));
} catch (HerdrException e) {
return error("herdr error listing workers: " + e.getMessage());
return error("herdr error listing the fleet: " + e.getMessage());
}
}
/**
* One lead's row: its address, its name, and whether it can be reached right now.
*
* <p>{@code status} is herdr's live view, and {@code unknown} when herdr is not tracking that
* pane as an agent — the honest answer, and the one that matters: a lead whose pane herdr cannot
* see is a lead a {@code bridge_send} cannot be typed into. It is reported rather than hidden,
* because a peer that has gone unreachable is exactly what the sender needs to know.
*/
private static Map<String, Object> leadView(String terminal, String name, Agent live,
String selfTerm) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", terminal);
m.put("name", name);
m.put("status", live == null || live.status() == null
? "unknown" : live.status().name().toLowerCase());
if (terminal.equals(selfTerm)) {
m.put("self", true);
}
return m;
}
/** {@code bridge_stop}: tear a worker down by its pane id. */
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
if (isBlank(paneId)) {
@@ -706,8 +844,13 @@ public final class BridgeMcp {
private static McpSchema.Tool listTool() {
return tool("bridge_list",
"List the worker sessions the bridge tracks — each with its sessionId, paneId, profile, "
+ "state, optional worktree/branch/owner, and live herdr status.",
"List the whole fleet the bridge tracks, in two parts. 'leads' are your PEERS — other "
+ "orchestrators, each with its sessionId (the address to bridge_send to), "
+ "name, live status, and 'self': true on your own row; this is how you "
+ "discover a peer lead without being told its address. 'workers' are the "
+ "sessions delegated to — each with sessionId, paneId, profile, state, "
+ "optional worktree/branch/owner, and live herdr status. An empty 'workers' "
+ "means no workers are spawned; it says nothing about peers.",
objectSchema(Map.of(), List.of()));
}
@@ -720,10 +863,12 @@ public final class BridgeMcp {
}
private static McpSchema.Tool replyTool() {
// No session/target arg — the worker's identity is resolved from the connection.
// No session/target arg — the caller's identity is resolved from the connection.
return tool("bridge_reply",
"Return your structured answer for the task you were delegated, "
+ "resolving the caller's blocked bridge_send.",
"Return your structured answer for a message you were sent, resolving the sender's "
+ "blocked bridge_send. A worker MUST end every delegated turn with exactly "
+ "one of these. A lead uses it only to answer another lead that messaged "
+ "it — never to answer a worker, whose turn it is not.",
objectSchema(Map.of(
"content", stringProp("Your reply/answer")),
List.of("content")));
@@ -741,11 +886,14 @@ public final class BridgeMcp {
return tool("bridge_whoami",
"Report who YOU are on the bridge — your role is resolved from your connection "
+ "(unforgeable), never from anything you claim. Returns role 'primary' (you "
+ "orchestrate: spawn/send/stop, and you must never call bridge_reply) or "
+ "'worker' (you were delegated to: you must end every turn with exactly one "
+ "bridge_reply, and cannot spawn or send), plus your own sessionId, profile, "
+ "worktree and branch when you are a worker. Call this first when following "
+ "role-conditional instructions rather than guessing your role.",
+ "orchestrate: spawn/send/stop; reply ONLY to answer a peer lead that "
+ "messaged you, never to answer a worker), 'architect' (you delegate turns "
+ "and reply/ask as your own pane, but cannot spawn/stop/drain), or 'worker' "
+ "(you were delegated to: you must end every turn with exactly one "
+ "bridge_reply, and cannot spawn or send), plus 'leader'/'architect' naming "
+ "which one you are, your own sessionId, and profile/worktree/branch when "
+ "you are a worker. Call this first when following role-conditional "
+ "instructions rather than guessing.",
objectSchema(Map.of(), List.of()));
}
@@ -4,6 +4,7 @@ import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Optional;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicReference;
/**
@@ -25,6 +26,15 @@ public final class PrimaryRegistry {
private final AtomicReference<String> terminal = new AtomicReference<>();
private final boolean pinned;
/**
* CB-532: worker terminal → the lead that delegated to it. The single slot above answers "who is
* THE primary", a question with no correct answer once two leads orchestrate the same fleet:
* whichever called {@code bridge_send} first captured every nudge, including nudges for the
* other lead's delegations. This map answers the question that actually matters — "who is
* waiting on THIS worker" — and is what lets {@code primary.terminal} be retired.
*/
private final ConcurrentHashMap<String, String> leadByTarget = new ConcurrentHashMap<>();
/**
* @param pinnedTerminal an optional pinned terminal from config ({@code null}/blank = unpinned)
*/
@@ -56,6 +66,45 @@ public final class PrimaryRegistry {
}
}
/**
* Record that {@code leadTerminal} owns the accepted delegation of worker {@code target} (CB-532).
*
* <p>Called from the {@code MessageService} accepted-delivery hook — only after a send has won
* the session's send lock and queued delivery — where both halves are known (CB-548). It is
* deliberately <em>not</em> called at {@code bridge_send} request time: a concurrent sender that
* times out {@code BUSY} must not steal a live delegation's reply routing without ever owning
* the turn. Last writer wins — if a second lead's later send is accepted, replies follow the
* lead that most recently delegated to it, which is the one waiting.
*/
public void recordDelegation(String target, String leadTerminal) {
if (target == null || target.isBlank() || leadTerminal == null || leadTerminal.isBlank()) {
return;
}
leadByTarget.put(target, leadTerminal);
}
/** Forget a worker's delegating lead — call on release, so a torn-down session leaks nothing. */
public void forgetDelegation(String target) {
if (target != null) {
leadByTarget.remove(target);
}
}
/**
* Where a nudge about {@code target}'s reply should go: the lead that delegated to it, falling
* back to the single known primary.
*
* <p>The fallback matters after a daemon restart, which loses the map while the durable inbox
* keeps the reply. With one lead the fallback is unambiguous and correct. With several and no
* recorded delegation there is no right answer, so this returns empty rather than guessing —
* delivery degrades to pull, which is exactly what the durable inbox is for, instead of
* interrupting the wrong lead with someone else's result.
*/
public Optional<String> nudgeTargetFor(String target) {
String lead = target == null ? null : leadByTarget.get(target);
return lead != null ? Optional.of(lead) : Optional.ofNullable(terminal.get());
}
/** The known primary terminal, or empty if not yet learned (and not pinned). */
public Optional<String> primaryTerminal() {
return Optional.ofNullable(terminal.get());
@@ -14,25 +14,30 @@ import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
/**
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
* port {@link InMemoryReplyInbox} implements as soft state.
*
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target owns a durable
* queue {@code agent.<target>.inbox}. A manual-ack consumer pulls persistent messages off that queue
* into an in-memory <em>held</em> map (keyed by {@code msgId}) but does <em>not</em> ack them.
* {@link #peek} returns that snapshot; {@link #ack} acks the broker delivery-tag and drops the entry.
* Because messages stay unacked until the primary actually drains them, a crash (or a {@code java -jar}
* bounce) before caller-ack leaves them on the broker — it redelivers on reconnect. That is the
* durability the in-memory adapter cannot give, with the port contract preserved.
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target has a durable
* queue {@code agent.<target>.inbox}. The gateway that owns the target starts a manual-ack consumer
* ({@link #own}) that pulls persistent messages off that queue into an in-memory <em>held</em> map
* (keyed by {@code msgId}) but does <em>not</em> ack them. {@link #peek} returns that snapshot;
* {@link #ack} acks the broker delivery-tag and drops the entry. Because messages stay unacked until
* the owning gateway actually drains them, a crash (or a {@code java -jar} bounce) before caller-ack
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
* adapter cannot give, with the port contract preserved.
*
* <p><strong>Ownership is explicit.</strong> {@link #own} declares the queue and starts the consumer;
* {@link #release} cancels it. {@link #publish} sends to the queue but does <em>not</em> imply ownership
* and does not attach a consumer. This split is required by CB-308 federation, where one gateway may
* publish to an agent owned by another gateway; in that case the publisher must not compete for
* deliveries.
*
* <p><strong>Dedup.</strong> The consumer keys the held map by {@code msgId}; a redelivered duplicate
* (at-least-once, or a producer double-publish) is acked-and-dropped on arrival, so it never
* double-queues. {@link #publish} additionally short-circuits an already-held {@code msgId} — a
* fast path; the consumer-side check is the real guarantee.
* double-queues.
*
* <p><strong>Visibility.</strong> Unlike the in-memory adapter, publish → broker → consumer is
* asynchronous, so a {@link #peek} immediately after {@link #publish} may not yet see the message
@@ -51,12 +56,12 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
private final Connection connection;
private final Channel channel;
/** All channel operations (publish/declare/ack) serialize on this — a Channel is not thread-safe. */
/** All channel operations (publish/declare/ack/cancel) serialize on this — a Channel is not thread-safe. */
private final Object channelLock = new Object();
/** target → (msgId → held delivery). Per-target map is guarded by synchronizing on itself. */
private final ConcurrentHashMap<String, LinkedHashMap<String, Held>> held = new ConcurrentHashMap<>();
/** Targets whose queue is declared and consumer is running. */
private final Set<String> consuming = ConcurrentHashMap.newKeySet();
/** Targets whose queue is declared and consumer is running, mapped to their broker consumer tag. */
private final ConcurrentHashMap<String, String> consumerTags = new ConcurrentHashMap<>();
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
private record Held(long deliveryTag, InboxMessage message) {}
@@ -103,16 +108,41 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
@Override
public void publish(String target, String msgId, String content) {
ensureConsuming(target);
var perTarget = held.get(target);
if (perTarget != null) {
synchronized (perTarget) {
if (perTarget.containsKey(msgId)) {
return; // already held — producer-side fast dedup
}
public void own(String target) {
synchronized (channelLock) {
if (consumerTags.containsKey(target)) {
return; // already owning this target
}
String queue = queueName(target);
try {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
String tag = channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
consumerTags.put(target, tag);
log.debug("AMQP inbox owns queue {} for target {}", queue, target);
} catch (IOException e) {
throw new IllegalStateException("cannot own queue " + queue, e);
}
}
}
@Override
public void release(String target) {
synchronized (channelLock) {
String tag = consumerTags.remove(target);
held.remove(target); // stale delivery tags must not survive release
if (tag == null) {
return;
}
try {
channel.basicCancel(tag);
} catch (IOException e) {
throw new IllegalStateException("cannot cancel consumer for " + target, e);
}
}
}
@Override
public void publish(String target, String msgId, String content) {
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
.messageId(msgId)
.deliveryMode(2) // persistent — survives a broker restart
@@ -129,7 +159,6 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
@Override
public List<InboxMessage> peek(String target) {
ensureConsuming(target);
var perTarget = held.get(target);
if (perTarget == null) {
return List.of();
@@ -166,26 +195,6 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
}
/** Declare the durable per-target queue and start its manual-ack consumer, once per target. */
private void ensureConsuming(String target) {
if (consuming.contains(target)) {
return;
}
synchronized (channelLock) {
if (!consuming.add(target)) {
return; // another thread just set it up
}
String queue = queueName(target);
try {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
} catch (IOException e) {
consuming.remove(target);
throw new IllegalStateException("cannot consume queue " + queue, e);
}
}
}
private DeliverCallback deliverCallback(String target) {
return (_, delivery) -> {
String msgId = delivery.getProperties().getMessageId();
@@ -2,6 +2,7 @@ package dev.ltms.bridged.msg;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
/**
@@ -9,12 +10,29 @@ import java.util.concurrent.ConcurrentHashMap;
* Per-target FIFO ordering (insertion order via {@link LinkedHashMap}). Dedup by {@code msgId}
* within a target. Thread-safe for concurrent publish vs. drain.
*
* <p><strong>Ownership is explicit.</strong> {@link #own} marks a target as locally owned so that
* {@link #peek} and {@link #ack} operate on it; {@link #publish} works whether or not the target is
* owned. {@link #release} clears the local snapshot. This mirrors the AMQP adapter's contract so the
* non-broker path stays interchangeable.
*
* <p><strong>This is soft-state, NOT persistence.</strong> Lost on a {@code java -jar} bounce — that
* is correct and consistent with "bridged stays soft-state." The Stage-2 AMQP adapter replaces this.
*/
public final class InMemoryReplyInbox implements ReplyInbox {
private final ConcurrentHashMap<String, LinkedHashMap<String, InboxMessage>> store = new ConcurrentHashMap<>();
private final Set<String> owned = ConcurrentHashMap.newKeySet();
@Override
public void own(String target) {
owned.add(target);
}
@Override
public void release(String target) {
owned.remove(target);
store.remove(target);
}
@Override
public void publish(String target, String msgId, String content) {
@@ -27,6 +45,9 @@ public final class InMemoryReplyInbox implements ReplyInbox {
@Override
public List<InboxMessage> peek(String target) {
if (!owned.contains(target)) {
return List.of();
}
var perTarget = store.get(target);
if (perTarget == null) {
return List.of();
@@ -39,6 +60,9 @@ public final class InMemoryReplyInbox implements ReplyInbox {
@Override
public void ack(String target, String msgId) {
if (!owned.contains(target)) {
return;
}
var perTarget = store.get(target);
if (perTarget != null) {
//noinspection SynchronizationOnLocalVariableOrMethodParameter
@@ -310,6 +310,22 @@ public final class MessageService {
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
*/
public Reply send(String target, String content, long timeoutMillis) {
return send(target, content, timeoutMillis, null);
}
/**
* As {@link #send(String, String, long)}, but with an accepted-delivery hook.
*
* <p>{@code onAccepted} is invoked exactly once, once this send has won {@code target}'s send
* lock and so become the <em>accepted target turn</em> — it runs <em>before</em> delivery is
* queued, so a throwing hook fails the send cleanly (the waiter it already opened is closed and
* nothing is left queued). It is <em>not</em> invoked when the send is {@link Outcome#BUSY}
* (lock never taken). A caller uses this to record that <em>it</em> now owns the delegation's
* reply routing (CB-548: {@code PrimaryRegistry} delegator ownership) — recording only on
* acceptance means a concurrent sender that times out {@code BUSY} can never steal ownership it
* never earned. {@code null} disables the hook.
*/
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
@@ -317,22 +333,36 @@ public final class MessageService {
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
}
try {
CompletableFuture<Void> delivered = injector.enqueue(target, content);
// Open the waiter BEFORE queueing delivery (CB-548). A fast reply — the worker already
// injectable the instant we enqueue — otherwise arrives before the waiter is registered
// and orphans into the inbox while this send blocks to the timeout (the enqueue-before-
// open race). Opening first also means a throwing onAccepted (fired before enqueue) or an
// enqueue failure is safely closed by the finally below: nothing is left queued, and the
// failed send leaves no stale waiter behind.
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
// The send has won the lock; the accepted-delivery hook records delegator ownership
// here (CB-548). It runs BEFORE enqueue so a throwing hook — onAccepted is now a
// public callback — fails the send without queuing a message that would orphan.
if (onAccepted != null) {
onAccepted.run();
}
CompletableFuture<Void> delivered = injector.enqueue(target, content);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
}
} finally {
rendezvous.close(target, reply);
}
@@ -440,9 +470,21 @@ public final class MessageService {
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content) {
return sendAsync(target, content, null);
}
/**
* As {@link #sendAsync(String, String)}, with the accepted-delivery hook of
* {@link #send(String, String, long, Runnable)} — the running {@code send} invokes {@code onAccepted}
* the moment it becomes the accepted target turn, so async flooding records delegator ownership
* exactly as the blocking path does (CB-548).
*
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content, Runnable onAccepted) {
String ticket = "task-" + ticketSeq.incrementAndGet();
CompletableFuture<Reply> future =
CompletableFuture.supplyAsync(() -> send(target, content, ASYNC_TIMEOUT_MS), asyncExecutor);
CompletableFuture<Reply> future = CompletableFuture.supplyAsync(
() -> send(target, content, ASYNC_TIMEOUT_MS, onAccepted), asyncExecutor);
tasks.put(ticket, new Task(target, future, System.nanoTime()));
pruneTerminalTickets();
log.debug("async send {} -> {}", ticket, target);
@@ -74,15 +74,30 @@ public final class Rendezvous {
/**
* Register a waiter for {@code session} — the await side of the public {@code resolve*} methods.
* The caller must hold that session's send lock.
*
* <p>Atomic fail-if-present (CB-548): if a waiter is already registered for {@code session}, an
* {@link IllegalStateException} is thrown rather than replacing the first — so any future
* invariant violation fails loudly instead of silently swapping the waiter another send is
* blocked on. {@code MessageService} serializes sends per session (the send lock), so in correct
* code a double open is impossible; this is a tripwire for the day that no longer holds.
*/
public CompletableFuture<Resolution> open(String session) {
CompletableFuture<Resolution> waiter = new CompletableFuture<>();
waiters.put(session, waiter);
CompletableFuture<Resolution> existing = waiters.putIfAbsent(session, waiter);
if (existing != null) {
throw new IllegalStateException(
"rendezvous double-open for session " + session + " — a waiter is already registered");
}
return waiter;
}
/** Remove {@code waiter} for {@code session} (only if it is still the registered one). */
void close(String session, CompletableFuture<Resolution> waiter) {
/**
* Remove {@code waiter} for {@code session}, only if it is still the registered one. The
* symmetric complement of {@link #open}: a terminal send deregisters its waiter so the next
* send on the session may {@link #open} a fresh one (CB-548 makes double-open an error, so a
* successful {@code open} after a finished turn requires this close to have happened first).
*/
public void close(String session, CompletableFuture<Resolution> waiter) {
waiters.remove(session, waiter);
}
@@ -10,16 +10,36 @@ import java.util.List;
* <p><strong>This interface is the port.</strong> {@link InMemoryReplyInbox} is the Stage-1 adapter;
* an AMQP-backed adapter (Stage 2) must implement the same contract (idempotent publish, FIFO peek,
* at-least-once ack).
*
* <p><strong>Ownership is explicit.</strong> A gateway {@link #own owns} the inbox for each agent it
* spawned; only the owner consumes and drains it. {@link #publish} sends a reply to the target's
* inbox but does <em>not</em> imply ownership or start a consumer. This separation is required by
* CB-308 federation, where one gateway may publish to an agent owned by another gateway.
*/
public interface ReplyInbox {
/** A queued reply: an idempotency id, the worker session it came from, and the reply text. */
record InboxMessage(String msgId, String target, String content) {}
/**
* Start owning (consuming) the inbox for {@code target}. Idempotent: multiple calls for the same
* target are no-ops. The owner is the only gateway that may {@link #peek} and {@link #ack} replies
* for this target.
*/
void own(String target);
/**
* Stop owning (consuming) the inbox for {@code target}. Idempotent. Any replies held locally but
* not yet acked are dropped from the local snapshot; the underlying durable queue keeps
* unacked messages for redelivery when the target is re-owned.
*/
void release(String target);
/**
* Queue {@code content} from worker {@code target} under {@code msgId}. Idempotent: publishing an
* already-present {@code msgId} for {@code target} is a no-op (dedup), so an at-least-once Stage-2
* redelivery cannot double-queue.
* redelivery cannot double-queue. Publishing does <em>not</em> imply ownership and must not start a
* consumer.
*/
void publish(String target, String msgId, String content);
@@ -59,10 +59,10 @@ public final class ReplyPushLoop {
this.metrics = metrics;
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(name, labels);
metrics.inc(BridgedMetrics.PUSH_NUDGES, "outcome", outcome);
}
}
@@ -79,8 +79,13 @@ public final class ReplyPushLoop {
* @return the action the caller should take
*/
Action decide(String target, int reminderCount) {
if (!primaryRegistry.isKnown()) {
log.debug("push: primary unknown, stopping reminder for {}", target);
// CB-532: the destination is per-delegation — the lead that sent this worker its work, not
// "the primary". With two leads orchestrating one fleet the singular question has no right
// answer, and answering it anyway interrupted whichever lead happened to call bridge_send
// first with results it never asked for.
var nudgeTarget = primaryRegistry.nudgeTargetFor(target);
if (nudgeTarget.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}, stopping reminder", target);
return Action.STOP;
}
if (inbox.peek(target).isEmpty()) {
@@ -89,21 +94,21 @@ public final class ReplyPushLoop {
}
if (reminderCount >= maxReminders) {
log.debug("push: reminder cap ({}) reached for {}, stopping", maxReminders, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted");
countNudge("exhausted");
return Action.STOP;
}
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
String leadTerminal = nudgeTarget.get();
AgentStatus status;
try {
status = agents.status(primaryTerminal);
status = agents.status(leadTerminal);
} catch (RuntimeException e) {
log.debug("push: status check failed for primary {}, will retry", primaryTerminal, e);
log.debug("push: status check failed for lead {}, will retry", leadTerminal, e);
return Action.WAIT_BUSY;
}
if (status.injectable()) {
return Action.INJECT;
}
log.debug("push: primary {} is {} (not injectable), waiting", primaryTerminal, status);
log.debug("push: lead {} is {} (not injectable), waiting", leadTerminal, status);
return Action.WAIT_BUSY;
}
@@ -143,16 +148,23 @@ public final class ReplyPushLoop {
/** Send the nudge and log the event. */
private void injectNudge(String target, int reminderCount) {
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
// Re-read rather than threading it down from decide(): the delegating lead can change
// between the decision and the injection, and the nudge should follow the current one.
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.debug("push: lead for {} disappeared before the nudge could be sent", target);
return;
}
String leadTerminal = lead.get();
String nudge = NUDGE_FORMAT.formatted(target, target);
try {
agents.send(primaryTerminal, nudge);
log.debug("push: nudge {}/{} sent to primary {} for target {}",
reminderCount + 1, maxReminders, primaryTerminal, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered");
agents.send(leadTerminal, nudge);
log.debug("push: nudge {}/{} sent to lead {} for target {}",
reminderCount + 1, maxReminders, leadTerminal, target);
countNudge("delivered");
} catch (RuntimeException e) {
log.warn("push: failed to nudge primary {} for target {} (reminder {}/{}): {}",
primaryTerminal, target, reminderCount + 1, maxReminders, e.toString());
log.warn("push: failed to nudge lead {} for target {} (reminder {}/{}): {}",
leadTerminal, target, reminderCount + 1, maxReminders, e.toString());
}
}
@@ -26,10 +26,33 @@ public enum Capability {
*/
WORKTREE,
/**
* The peer can discard its current conversation context without starting a delegated bridge
* turn. The command or API used to do that is adapter-specific.
*/
CONTEXT_RESET,
/**
* The spawner can reconcile orphaned peers on boot — workers that outlived a prior daemon
* process and whose pane ids died with it (CB-117). Claude Code over herdr supports this
* via name-based matching against the herdr agent list.
*/
ORPHAN_REAP
ORPHAN_REAP,
/**
* The peer surfaces the bridge's logical name ({@link SpawnRequest#sessionName()}) in its own
* UI at spawn — e.g. Claude Code's {@code -n} display name, which shows in the prompt box, the
* {@code /resume} picker, and the terminal title. This is what lets an operator tell the
* bridge's worker apart from a user's own session in the same terminal after a restart.
*/
SESSION_NAME,
/**
* The peer can be relaunched onto a prior conversation via that conversation's own session id
* ({@link SpawnRequest#resumeSessionId()}) — e.g. Claude Code's {@code -r}, which adopts the id
* as the agent's own identity rather than starting a new conversation. A launcher that declares
* this mints/resolves that id at spawn and exposes it on the returned
* {@link PeerHandle#agentSessionId()}, so a later resume addresses the same conversation.
*/
SESSION_RESUME
}
@@ -12,9 +12,13 @@ package dev.ltms.bridged.peer;
public interface PeerHandle {
/**
* The registry/routing key — an opaque, launcher-assigned identifier. For the herdr-backed
* launcher this is the herdr pane id; for other launchers it is whatever their transport
* uses. Guaranteed to be non-null and unique among live peers within a single daemon process.
* The registry/routing key — an opaque, launcher-assigned identifier (CB-519). Multiple
* daemon processes may run on one host, so the contract is <em>host-unique</em>, not merely
* process-unique: the herdr-backed launcher mints a fresh UUID per spawn, and a non-herdr
* launcher is likewise expected to return an identifier that cannot collide across processes
* on the same host. This id is the routing key and is deliberately decoupled from any launcher
* transport coordinate (e.g. a herdr pane id), which stays launcher-private. Guaranteed to be
* non-null and unique among live peers on the host.
*/
String id();
@@ -26,4 +30,41 @@ public interface PeerHandle {
default String terminalId() {
return null;
}
/**
* The worker profile that spawned this peer, if the launcher resolved one. A launcher that
* performs dynamic profile selection (e.g. CB-518 weighted placement) sets this so the
* session registry records the actual profile rather than the requested/default one.
*
* @return the profile name, or {@code null} when the launcher leaves it unspecified
*/
default String profile() {
return null;
}
/**
* The bridge's logical name for this session, as assigned at spawn
* ({@link SpawnRequest#sessionName()}). Stable across restarts and meaningful to an operator,
* unlike the transport identifiers above; a launch with no name leaves the peer's display
* identity to the launcher to derive.
*
* @return the bridge-assigned logical session name, or {@code null} if none was assigned
*/
default String sessionName() {
return null;
}
/**
* The peer's OWN session id — the handle that resumes this conversation later (the id a
* later {@link Capability#SESSION_RESUME resume} spawn would pass back). Null when the adapter
* cannot determine it — the contract for an adapter that declines
* {@link Capability#SESSION_RESUME}; an adapter that declares that capability returns this
* non-null for a spawn that requested session identity, because it knows the id before the
* peer has written anything.
*
* @return the peer's own session id, or {@code null} when not determinable
*/
default String agentSessionId() {
return null;
}
}
@@ -82,4 +82,13 @@ public interface PeerLauncher {
* when safe to do so.
*/
void stop(String id);
/**
* Discard the context of the peer identified by {@code id}. Implementations must bypass normal
* bridge delivery/turn accounting. Unsupported peer kinds return {@code false} without sending
* a guessed command.
*
* @return {@code true} when a reset was sent and its status transition must settle before reuse
*/
boolean clearContext(String id);
}
@@ -8,6 +8,17 @@ package dev.ltms.bridged.peer;
* <p>A null or blank {@code profileName} means "use the launcher's default profile."
* A null or blank {@code requestedCwd} means "inherit from config or caller."
* A null {@code callerCwd} means "the request came from the daemon itself (not a primary)."
*
* <p>{@code sessionName} and {@code resumeSessionId} carry the session's durable identity (CB-547a):
* the bridge's LOGICAL name for the session (stable across restarts, meaningful to an operator)
* and the peer's OWN prior session id to resume, respectively. Both are <em>opted in</em> — either
* may be null/blank, in which case the launcher derives a display name and mints a fresh session.
*/
public record SpawnRequest(String profileName, String requestedCwd, String callerCwd) {
public record SpawnRequest(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId) {
/** Back-compat: a spawn with no session identity (fresh session, launcher-derived name). */
public SpawnRequest(String profileName, String requestedCwd, String callerCwd) {
this(profileName, requestedCwd, callerCwd, null, null);
}
}
@@ -0,0 +1,22 @@
package dev.ltms.bridged.placement;
/**
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps and
* reachability so that a pre-existing config behaves identically after upgrade.
*/
final class FixedPlacementPolicy implements PlacementPolicy {
@Override
public PlacementCandidate select(PlacementContext ctx) {
String d = ctx.defaultProfile();
if (d != null && !d.isBlank()) {
return new PlacementCandidate(d, null, 1.0f, null);
}
if (!ctx.candidates().isEmpty()) {
PlacementCandidate first = ctx.candidates().getFirst();
return new PlacementCandidate(first.profile(), null, first.weight(), first.maxLoad());
}
throw new PlacementException("no worker profiles configured");
}
}
@@ -0,0 +1,20 @@
package dev.ltms.bridged.placement;
/**
* A profile (and, in CB-308, a host) that can be chosen by a {@link PlacementPolicy}.
*
* <p>Keeping this as a small descriptor rather than a bare profile name lets CB-308 widen
* selection to {@code (host, profile)} pairs without changing the policy interface.
*/
public record PlacementCandidate(String profile, String host, float weight, Integer maxLoad) {
/** A candidate with no explicit host (the single-host default) and the given weight/cap. */
public static PlacementCandidate profile(String profile, float weight, Integer maxLoad) {
return new PlacementCandidate(profile, null, weight, maxLoad);
}
/** A candidate with no explicit host, unit weight, and no cap. */
public static PlacementCandidate profile(String profile) {
return new PlacementCandidate(profile, null, 1.0f, null);
}
}
@@ -0,0 +1,19 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.Set;
import java.util.function.Function;
/**
* Everything a {@link PlacementPolicy} needs to make one selection.
*
* @param defaultProfile profile a {@code fixed} policy should return (may be {@code null})
* @param candidates every configured candidate; the policy filters out those at cap or unreachable
* @param liveCount current live worker count per profile (from the session registry)
* @param unreachable profiles already known to have failed in this spawn attempt
*/
public record PlacementContext(String defaultProfile,
List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable) {
}
@@ -0,0 +1,12 @@
package dev.ltms.bridged.placement;
/**
* Thrown when a {@link PlacementPolicy} has no candidate available. Kept as a distinct type so
* callers can distinguish "no capacity" from a spawn-time transport failure.
*/
public final class PlacementException extends IllegalStateException {
public PlacementException(String message) {
super(message);
}
}
@@ -0,0 +1,41 @@
package dev.ltms.bridged.placement;
/**
* Factory for the built-in placement policies.
*/
public final class PlacementPolicies {
private PlacementPolicies() {
}
/**
* Resolve a policy name from config. Absent/blank values and {@code "fixed"} return the
* backward-compatible fixed policy; unknown names throw.
*/
public static PlacementPolicy fromName(String name) {
String n = (name == null) ? "" : name.toLowerCase();
if (n.isBlank() || "fixed".equals(n)) {
return fixed();
}
if ("weighted".equals(n)) {
return weighted();
}
if ("round-robin".equals(n)) {
return roundRobin();
}
throw new IllegalArgumentException("unknown placement policy '" + name
+ "' — must be one of: fixed, round-robin, weighted");
}
public static PlacementPolicy fixed() {
return new FixedPlacementPolicy();
}
public static PlacementPolicy weighted() {
return new WeightedRoundRobinPolicy();
}
public static PlacementPolicy roundRobin() {
return new RoundRobinPlacementPolicy();
}
}
@@ -0,0 +1,16 @@
package dev.ltms.bridged.placement;
/**
* How {@code bridged} chooses a worker profile when a spawn names none. Implementations are
* deterministic and unit-testable; the caller (the composite launcher) handles failover retries.
*/
public interface PlacementPolicy {
/**
* Pick one candidate from the configured set.
*
* @throws java.lang.IllegalStateException when no candidate is available, with a message naming
* whether every profile is at capacity or unreachable
*/
PlacementCandidate select(PlacementContext ctx);
}
@@ -0,0 +1,65 @@
package dev.ltms.bridged.placement;
import java.util.ArrayList;
import java.util.List;
/**
* Shared filtering and empty-set reporting used by the built-in placement policies.
*/
final class PlacementPolicyUtil {
private PlacementPolicyUtil() {
}
/**
* Candidates that are not known-unreachable and have not reached their maxLoad.
* A {@code null} maxLoad means unlimited.
*/
static List<PlacementCandidate> available(PlacementContext ctx) {
List<PlacementCandidate> out = new ArrayList<>();
for (PlacementCandidate c : ctx.candidates()) {
if (ctx.unreachable().contains(c.profile())) {
continue;
}
Integer cap = c.maxLoad();
if (cap != null) {
int live = ctx.liveCount().apply(c.profile());
if (live >= cap) {
continue;
}
}
out.add(c);
}
return out;
}
/**
* Build a clear exception describing why every candidate was dropped: all at capacity,
* all unreachable, or a mix.
*/
static PlacementException emptyException(PlacementContext ctx) {
int atCap = 0;
int unreachable = 0;
for (PlacementCandidate c : ctx.candidates()) {
Integer cap = c.maxLoad();
if (ctx.unreachable().contains(c.profile())) {
unreachable++;
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
atCap++;
}
}
int total = ctx.candidates().size();
if (total == 0) {
return new PlacementException("no worker profiles configured");
}
if (atCap == total) {
return new PlacementException("all worker profiles are at maxLoad");
}
if (unreachable == total) {
return new PlacementException("all worker profiles are unreachable");
}
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
+ unreachable + " unreachable, " + (total - atCap - unreachable) + " remaining");
}
}
@@ -0,0 +1,23 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.concurrent.atomic.AtomicInteger;
/**
* Deterministic round-robin over the profiles that still have capacity and are not known to be
* unreachable. The index advances only on successful selections so the distribution stays even
* across spawns.
*/
final class RoundRobinPlacementPolicy implements PlacementPolicy {
private final AtomicInteger index = new AtomicInteger(0);
@Override
public synchronized PlacementCandidate select(PlacementContext ctx) {
List<PlacementCandidate> available = PlacementPolicyUtil.available(ctx);
if (available.isEmpty()) {
throw PlacementPolicyUtil.emptyException(ctx);
}
return available.get(index.getAndIncrement() % available.size());
}
}
@@ -0,0 +1,50 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
/**
* Smooth weighted round-robin (nginx-style): for each selection, add the candidate's weight to
* its current score, pick the highest score, then subtract the total weight of all available
* candidates from the winner. Weights are not required to sum to 1.0; only their ratios matter.
*
* <p>The state is per-policy instance and protected by {@code synchronized} so concurrent spawns
* see a consistent, deterministic sequence rather than interleaving updates.
*/
final class WeightedRoundRobinPolicy implements PlacementPolicy {
private final Map<String, Double> current = new ConcurrentHashMap<>();
@Override
public synchronized PlacementCandidate select(PlacementContext ctx) {
List<PlacementCandidate> available = PlacementPolicyUtil.available(ctx);
if (available.isEmpty()) {
throw PlacementPolicyUtil.emptyException(ctx);
}
double total = 0.0;
for (PlacementCandidate c : available) {
total += c.weight();
}
if (total <= 0.0) {
throw new PlacementException("all available profiles have non-positive weight");
}
PlacementCandidate best = null;
double bestScore = Double.NEGATIVE_INFINITY;
for (PlacementCandidate c : available) {
double score = current.merge(c.profile(), (double) c.weight(), (old, add) -> old + add);
if (score > bestScore) {
bestScore = score;
best = c;
}
}
if (best == null) {
throw new PlacementException("no placement candidate could be selected");
}
current.put(best.profile(), current.get(best.profile()) - total);
return best;
}
}
@@ -225,12 +225,13 @@ public final class BridgedApp {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
// CB-519: the registry key is a host-unique id, not the pane coordinate — join on terminal.
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.paneId() != null)
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.paneId())))
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
ctx.status(200).json(Map.of("workers", out));
}
@@ -27,6 +27,47 @@ public final class GitWorktrees implements Worktrees {
private static final Logger log = LoggerFactory.getLogger(GitWorktrees.class);
/** Project-level MCP config. Present in the repo, so every worktree would otherwise inherit the
* primary's IDE server mounts (a CB-523 worker edited the primary checkout; see the isolation
* javadoc). Neutralized unconditionally. */
private static final String MCP_CONFIG = ".mcp.json";
/** What {@link #isolateToolSurface} writes for {@code .mcp.json}: a valid, explicitly empty server map. */
private static final String NEUTRAL_MCP_CONFIG = "{\n \"mcpServers\": {}\n}\n";
/** OpenCode's repo-level config. Tracked here, so it lands in every worktree; it carries
* {@code {file:.secrets/...}} references to gitignored secrets that never reach a worktree, and
* opencode refuses to start on a dangling reference — so it is neutralized and the worker gets
* only the config its launcher writes via {@code OPENCODE_CONFIG}. */
private static final String OPENCODE_CONFIG = "opencode.json";
/** What {@link #isolateToolSurface} writes for {@code opencode.json}: a valid, empty JSON object. */
private static final String NEUTRAL_OPENCODE_CONFIG = "{}\n";
/** Autoenv's repo-level config. Not tracked today, but re-landing it must stay safe: autoenv
* authorizes by path, so a fresh worktree path is always unauthorized and its interactive prompt
* would block every spawn — neutralize it so it can never be committed. */
private static final String AUTOENV_CONFIG = ".autoenv";
/** What {@link #isolateToolSurface} writes for {@code .autoenv}: a valid, empty env file. */
private static final String NEUTRAL_AUTOENV_CONFIG = "";
/**
* A tracked project config that is hostile in a provisioned worktree, and what to replace it
* with. {@link #file} is the repo-relative path; {@link #stub} is a neutral but VALID payload for
* that file's format — a malformed stub would only trade one crash for another;
* {@link #createIfAbsent} keeps {@code .mcp.json}'s long-standing behaviour of writing its stub
* even when the repo carries no such file, whereas the others are only touched when present.
*/
private record WorktreeHostileConfig(String file, String stub, boolean createIfAbsent) {}
/** The worktree-hostile configs neutralized in every provisioned worktree, in order. */
private static final List<WorktreeHostileConfig> WORKTREE_HOSTILE_CONFIGS = List.of(
new WorktreeHostileConfig(MCP_CONFIG, NEUTRAL_MCP_CONFIG, true),
new WorktreeHostileConfig(OPENCODE_CONFIG, NEUTRAL_OPENCODE_CONFIG, false),
new WorktreeHostileConfig(AUTOENV_CONFIG, NEUTRAL_AUTOENV_CONFIG, false)
);
private final String configuredRoot;
private final SecureRandom random = new SecureRandom();
private final AtomicLong seq = new AtomicLong();
@@ -55,9 +96,59 @@ public final class GitWorktrees implements Worktrees {
String wt = path.toAbsolutePath().toString();
log.info("adding worktree branch={} path={} base={}", branch, wt, base);
exec("git", "-C", repoRoot, "worktree", "add", wt, "-b", branch, base);
isolateToolSurface(wt);
return wt;
}
/**
* Neutralize the worktree's worktree-hostile project configs so a worker inherits only the tools
* and environment its launcher mounts (the bridge via {@code --mcp-config}, the opencode config
* via {@code OPENCODE_CONFIG}) — never the primary's.
*
* <p>This is unconditional, and it is not the same job as the parity overlay. The repo's own
* committed {@code .mcp.json} declares the primary's IDE servers, so a fresh checkout mounts them
* whether or not the overlay copies anything; a worker that inherits them navigates and edits
* through tools bound to the <em>primary's</em> IntelliJ project, which silently hands it absolute
* paths outside its own worktree. That is not hypothetical: a CB-523 worker made all 59 of its
* edits in the primary checkout while compiling its worktree, so every build it ran was of code
* that did not contain its changes. {@code opencode.json} is the same trap one tool over — tracked,
* so it lands in every worktree, referencing gitignored {@code .secrets/} files that never do, and
* opencode refuses to start on the dangling reference. {@code .autoenv} extends the principle to a
* config that is not tracked today: autoenv authorizes by path, so a fresh worktree path is always
* unauthorized and its interactive prompt would block every spawn, so re-landing one must be safe.
*
* <p>Where a config exists it is replaced by a valid neutral stub (an explicitly empty
* map/object, or an empty env file — never a deletion, which would still let a later
* {@code git checkout} restore the hostile copy). The {@code --skip-worktree} bit keeps the
* neutralized copy from ever showing up as a local modification the worker might commit. A config
* the repo does not carry is skipped silently — no stub is invented for a file the repo does not
* have, and one missing file must never fail provisioning.
*/
private void isolateToolSurface(String worktreePath) {
Path root = Path.of(worktreePath).toAbsolutePath().normalize();
for (WorktreeHostileConfig cfg : WORKTREE_HOSTILE_CONFIGS) {
neutralize(root, worktreePath, cfg);
}
}
private void neutralize(Path root, String worktreePath, WorktreeHostileConfig cfg) {
Path target = root.resolve(cfg.file());
if (!Files.exists(target) && !cfg.createIfAbsent()) {
log.debug("{} absent in the worktree — skipping (repo does not carry it)", cfg.file());
return;
}
try {
Files.writeString(target, cfg.stub());
} catch (IOException e) {
throw new WorktreeException("cannot neutralize " + cfg.file() + " in the worktree: "
+ e.getMessage(), e);
}
if (isTracked(root, cfg.file())) {
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", cfg.file());
}
log.debug("neutralized {} — worker tool surface is launcher-mounted only", cfg.file());
}
@Override
public void remove(String repoRoot, String worktreePath) {
Path p = Path.of(worktreePath);
@@ -47,37 +47,46 @@ public final class SessionManager implements TurnListener {
private final AtomicLong nonceSeq = new AtomicLong();
private final LongSupplier nowNanos;
private final int contextCap;
private final boolean clearAfterTurn;
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** CB-516: notified with a terminalId on every release; no-op until wired. */
private volatile Consumer<String> releaseListener = _ -> { };
private final List<Consumer<String>> releaseListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
public SessionManager(PeerLauncher launcher) {
this(launcher, new GitWorktrees(), System::nanoTime, 0);
this(launcher, new GitWorktrees(), System::nanoTime, 0, false);
}
/** Backward-compatible constructor with an injectable worktree seam. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees) {
this(launcher, worktrees, System::nanoTime, 0);
this(launcher, worktrees, System::nanoTime, 0, false);
}
/** Test constructor with an injectable clock. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos) {
this(launcher, worktrees, nowNanos, 0);
this(launcher, worktrees, nowNanos, 0, false);
}
/** Production constructor with a configured context turn cap. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, int contextCap) {
this(launcher, worktrees, System::nanoTime, contextCap);
this(launcher, worktrees, System::nanoTime, contextCap, false);
}
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos,
int contextCap) {
int contextCap) {
this(launcher, worktrees, nowNanos, contextCap, false);
}
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos,
int contextCap, boolean clearAfterTurn) {
this.launcher = launcher;
this.worktrees = worktrees;
this.presence = new PresenceBridge(this);
this.nowNanos = nowNanos;
this.contextCap = contextCap;
this.clearAfterTurn = clearAfterTurn;
}
/**
@@ -110,9 +119,8 @@ public final class SessionManager implements TurnListener {
if (wt == null) {
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd);
PeerHandle handle = launcher.spawn(req);
String resolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
String cwd = launcher.effectiveCwd(req);
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd));
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
handle.id(),
@@ -129,6 +137,7 @@ public final class SessionManager implements TurnListener {
registry.put(handle.id(), session);
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
notifyAcquired(session.terminalId());
return session;
}
return acquireWithWorktree(profile, requestedCwd, callerCwd, ownerTerminal, wt);
@@ -151,18 +160,44 @@ public final class SessionManager implements TurnListener {
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is acquired
* (CB-520). This is the hook that lets the reply inbox {@code own} a target's queue.
*/
public void onAcquire(Consumer<String> listener) {
if (listener != null) {
acquireListeners.add(listener);
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is released
* (CB-516). Every teardown path funnels through {@link #release}, so one hook covers the REST
* and MCP stop tools, the idle-TTL reaper, {@code recycle}, and shutdown drain alike.
*
* <p>Set rather than injected because {@code MessageService} — the intended listener — is
* <p>Added rather than injected because {@code MessageService} — one intended listener — is
* constructed after this manager (it needs the injector and rendezvous, which need the session
* presence view this manager exposes). Wiring it at construction would require breaking that
* cycle for one callback.
*/
public void onRelease(Consumer<String> listener) {
this.releaseListener = (listener == null) ? _ -> { } : listener;
if (listener != null) {
releaseListeners.add(listener);
}
}
/** A listener failure must never prevent the acquisition it is reacting to. */
private void notifyAcquired(String terminalId) {
if (terminalId == null) {
return;
}
for (Consumer<String> listener : acquireListeners) {
try {
listener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("acquire listener failed for terminal {}: {}", terminalId, e.toString());
}
}
}
/** A listener failure must never prevent the teardown it is reacting to. */
@@ -170,16 +205,18 @@ public final class SessionManager implements TurnListener {
if (terminalId == null) {
return;
}
try {
releaseListener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
for (Consumer<String> listener : releaseListeners) {
try {
listener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
}
}
}
private WorkerSession acquireWithWorktree(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
String resolvedProfile = (profile == null || profile.isBlank())
String preResolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
@@ -187,13 +224,13 @@ public final class SessionManager implements TurnListener {
// `git -C null` on the command line — an NPE out of ProcessBuilder, surfacing as HTTP 500.
// The non-worktree path always used this chain; only this branch was missed.
String repoRoot = worktrees.repoRoot(
launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd)));
launcher.effectiveCwd(new SpawnRequest(preResolvedProfile, requestedCwd, callerCwd)));
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
String path = null;
PeerHandle handle;
try {
path = worktrees.add(repoRoot, branch, wt.baseRef());
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(resolvedProfile));
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd));
} catch (RuntimeException e) {
if (path != null) {
@@ -205,12 +242,14 @@ public final class SessionManager implements TurnListener {
}
throw e;
}
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
resolveCwd(path, profile, callerCwd),
cwd,
ownerTerminal,
now,
now,
@@ -221,6 +260,7 @@ public final class SessionManager implements TurnListener {
registry.put(handle.id(), session);
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
notifyAcquired(session.terminalId());
return session;
}
@@ -232,6 +272,22 @@ public final class SessionManager implements TurnListener {
return String.format("%06x", nonceRandom.nextInt(1 << 24)) + "-" + nonceSeq.incrementAndGet();
}
/**
* The profile to record for a session. A launcher that performed dynamic selection tells us
* the actual profile via {@link PeerHandle#profile()}; otherwise fall back to what the caller
* requested (or the launcher's default for a no-profile spawn).
*/
private String resolveProfile(PeerHandle handle, String requestedProfile) {
String fromHandle = handle.profile();
if (fromHandle != null && !fromHandle.isBlank()) {
return fromHandle;
}
if (requestedProfile != null && !requestedProfile.isBlank()) {
return requestedProfile;
}
return launcher.defaultProfile();
}
/**
* Release the old session and acquire a fresh one with the same profile and working directory.
* The new session is guaranteed to have a pane id distinct from the old one (no-reuse invariant).
@@ -306,8 +362,25 @@ public final class SessionManager implements TurnListener {
/** Lifecycle hook: the worker's delegated turn completed successfully. */
@Override
public void onTurnComplete(String target) {
completeTurn(target, false);
}
@Override
public boolean hasPostTurnAction(String target) {
if (!clearAfterTurn) return false;
WorkerSession current = findByTerminal(target);
if (current == null || current.state() != WorkerSession.State.BUSY) return;
return current != null && current.state() == WorkerSession.State.BUSY
&& (contextCap <= 0 || current.turnCount() < contextCap);
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
return completeTurn(target, true);
}
private boolean completeTurn(String target, boolean startContextReset) {
WorkerSession current = findByTerminal(target);
if (current == null || current.state() != WorkerSession.State.BUSY) return false;
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.DONE).withActivity(now);
if (replace(current, updated)) {
@@ -316,6 +389,15 @@ public final class SessionManager implements TurnListener {
}
if (contextCap > 0 && updated.turnCount() >= contextCap) {
release(current.paneId());
return false;
}
if (!startContextReset || !clearAfterTurn) return false;
try {
return launcher.clearContext(current.paneId());
} catch (RuntimeException e) {
log.warn("context reset failed for terminal={} pane={}; continuing without reset: {}",
target, current.paneId(), e.getMessage());
return false;
}
}
@@ -401,7 +483,19 @@ public final class SessionManager implements TurnListener {
return registry.size();
}
/**
* The registered session owning {@code terminalId}, or {@code null} if none does.
*
* <p>A null {@code terminalId} is a normal input, not a caller bug: every lifecycle hook here is
* fed from the MCP transport, where the <em>primary</em> resolves to a {@link
* dev.ltms.bridged.auth.Principal} with no terminal. {@code BridgeMcp} documents that contact as
* a no-op, and {@link dev.ltms.bridged.inject.WorkerPresence#markPresent} honours it — but
* {@code PresenceBridge} then forwards the same null here. Matching on a null id can never
* succeed anyway (a registered session always has a terminal), so answer "no match" rather than
* throwing: an NPE on this path takes down an unrelated tool call for the primary.
*/
private WorkerSession findByTerminal(String terminalId) {
if (terminalId == null) return null;
for (WorkerSession s : registry.values()) {
if (terminalId.equals(s.terminalId())) return s;
}
@@ -423,10 +517,6 @@ public final class SessionManager implements TurnListener {
return registry.replace(expected.paneId(), expected, updated);
}
private String resolveCwd(String requestedCwd, String profileName, String callerCwd) {
return launcher.effectiveCwd(new SpawnRequest(profileName, requestedCwd, callerCwd));
}
/** WorkerPresence bridge that also drives the manager's READY transition. */
private static final class PresenceBridge extends WorkerPresence {
private final SessionManager sessions;
@@ -437,6 +527,9 @@ public final class SessionManager implements TurnListener {
@Override
public void markPresent(String terminal) {
if (terminal == null || terminal.isBlank()) {
return; // the primary's contact carries no worker terminal — not a readiness signal
}
super.markPresent(terminal);
sessions.onReady(terminal);
}
@@ -5,7 +5,10 @@ package dev.ltms.bridged.session;
* process spawned. Immutable; state transitions are performed by replacing the record in
* {@link SessionManager}'s registry.
*
* @param paneId herdr pane handle — the registry key and the argument to teardown
* @param paneId the host-unique opaque id (CB-519) — the registry key and the argument to
* teardown. Despite the historical name this is the {@link
* dev.ltms.bridged.peer.PeerHandle#id()}, a UUID, and is distinct from the
* launcher-private herdr pane coordinate.
* @param terminalId herdr terminal handle — the {@code target} for send/read/status
* @param profile the worker profile name that spawned this session
* @param cwd the resolved working directory the worker started in
@@ -6,11 +6,14 @@ import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.function.Function;
import java.util.function.LongSupplier;
@@ -36,6 +39,8 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "claude";
private static final Logger log = LoggerFactory.getLogger(ClaudeCodeLauncher.class);
private final SubscriptionGuard guard;
/**
@@ -109,24 +114,103 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/**
* {@inheritDoc}
*
* <p>The spawn sequence encodes the subscription boundary: assert the profile's base_url is on
* the allowlist <em>before</em> any herdr call, then build the worker env with
* {@code ANTHROPIC_*}, the parity-neutral git-forge grant, and the bridge MCP + reply charter
* mounted as inline launch flags.
* <p>A legacy spawn with no session identity is a fresh, launcher-derived session — delegate to
* the session-aware form with no name and no resume id.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
return buildLaunch(cfg, null, null);
}
/**
* {@inheritDoc}
*
* <p>The spawn sequence encodes the subscription boundary: assert the profile's base_url is on
* the allowlist <em>before</em> any herdr call, then build the worker env with
* {@code ANTHROPIC_*}, the parity-neutral git-forge grant, and the bridge MCP + reply charter
* mounted as inline launch flags. When the request carries session identity (CB-547a) it is
* applied here — see {@link #applySessionIdentity}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg, String sessionName, String resumeSessionId) {
// CB-539: a profile may deliberately opt into the subscription (subscription: true) when no
// off-subscription endpoint exists for it — e.g. `sonnet` on `ccs`. That profile gets no
// ANTHROPIC_BASE_URL/AUTH_TOKEN (there is nothing to point them at) and the guard's base_url
// requirement is skipped FOR IT ONLY. Every other profile keeps the hard boundary below.
boolean onSubscription = cfg.isSubscription();
String baseUrl = cfg.baseUrl();
guard.assertWorker(baseUrl); // hard stop before we spawn anything
if (onSubscription) {
// NO SILENT CONTRADICTION: subscription:true + a baseUrl state opposite intents; refuse
// loudly rather than pick a winner.
if (baseUrl != null && !baseUrl.isBlank()) {
throw new IllegalStateException("profile '" + cfg.profile()
+ "' sets both subscription: true and a baseUrl ('" + baseUrl + "') — the two "
+ "are contradictory: a subscription profile must not point at an endpoint. "
+ "Drop baseUrl, or drop subscription: true.");
}
// Visible without anyone going looking for it: this worker bills the subscription.
log.warn("spawning profile '{}' on the Claude subscription (subscription: true) — this "
+ "worker WILL bill the operator's subscription", cfg.profile());
} else {
guard.assertWorker(baseUrl); // hard stop before we spawn anything
}
Map<String, String> workerEnv = baseEnv(cfg);
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
if (onSubscription) {
// CB-542 belt-and-braces: on the subscription path no guard vets these two keys, and the
// profile's env: is layered in by baseEnv — so strip any that rode in there. Config load
// already rejects this (loudly, naming the profile); this makes the boundary hold even
// for a profile built in code that never passed through that validation.
workerEnv.remove("ANTHROPIC_BASE_URL");
workerEnv.remove("ANTHROPIC_AUTH_TOKEN");
} else {
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
}
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithBridge(cfg));
// CB-547a: Claude Code can MINT its own session id, so bridged chooses it — a fresh spawn
// gets a UUID we pass as --session-id and return from agentSessionId(), so the resume
// handle is known BEFORE the agent has written anything; a resume spawn adopts its prior
// id via -r and passes no --session-id (the two conflict). Both are injected before the
// model flag so --model keeps outranking the operator's own argv.
// mutableArgv: argvWithBridge may hand back the profile's own (immutable) List.of when it
// has no MCP — session flags must be added into a list we own.
List<String> argv = mutableArgv(argvWithBridge(cfg));
String agentSessionId = applySessionIdentity(argv, sessionName, resumeSessionId);
return new Launch(workerEnv, argvWithModel(argv, cfg), agentSessionId);
}
/**
* Add the Claude-specific session-identity flags to {@code argv} and return the peer's OWN
* session id — the resume handle. A resume request passes the prior id via {@code -r} and
* returns that id; a fresh named session mints a new UUID, passes it via {@code --session-id},
* and returns the mint. The bridge's logical name rides along as {@code -n} when present. When
* <em>no</em> identity is requested (sessionName and resumeSessionId both blank) this adds
* nothing and returns {@code null}, keeping the legacy no-identity launch byte-identical.
*/
private static String applySessionIdentity(List<String> argv, String sessionName, String resumeSessionId) {
boolean resuming = resumeSessionId != null && !resumeSessionId.isBlank();
boolean named = sessionName != null && !sessionName.isBlank();
if (!resuming && !named) {
return null; // no identity requested — keep the legacy launch byte-identical
}
if (named) {
argv.add("-n");
argv.add(sessionName);
}
if (resuming) {
argv.add("-r");
argv.add(resumeSessionId);
return resumeSessionId;
}
String minted = UUID.randomUUID().toString();
argv.add("--session-id");
argv.add(minted);
return minted;
}
/**
@@ -149,6 +233,32 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
return argv;
}
/**
* Pin the model on the command line as well as in {@code ANTHROPIC_MODEL} (CB-533).
*
* <p>The env var alone is not a reliable pin for this adapter, because the argv is usually a
* launcher rather than {@code claude} itself — {@code ["ccs", "<profile>"]} — and {@code ccs}
* exports its profile's own model family ({@code ANTHROPIC_MODEL}, {@code DEFAULT_OPUS/SONNET/
* HAIKU}, {@code CLAUDE_CODE_SUBAGENT_MODEL}) over whatever it inherited. A worker profile that
* set {@code model:} therefore got silently overruled by its own launcher. Claude Code's
* {@code --model} flag outranks the environment, and {@code ccs <profile> [claude-args...]}
* passes trailing arguments through, so the flag survives the wrapper.
*
* <p>Appended last so it also outranks anything in the operator's own {@code argv}. Profiles
* that deliberately leave {@code model:} unset (letting {@code ccs} own model selection, as
* {@code gx10} does) are untouched — this adds nothing when there is nothing to add. This is
* the {@code kind: claude} counterpart of the opencode adapter's {@code -m provider/model}.
*/
private static List<String> argvWithModel(List<String> argv, BridgedConfig.Worker cfg) {
if (cfg.model() == null || cfg.model().isBlank()) {
return argv;
}
List<String> withModel = mutableArgv(argv);
withModel.add("--model");
withModel.add(cfg.model());
return withModel;
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
@@ -170,13 +280,26 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE,
Capability.CONTEXT_RESET, Capability.ORPHAN_REAP,
Capability.SESSION_NAME, Capability.SESSION_RESUME);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
@Override
public boolean clearContext(String id) {
String target = agentTarget(id);
if (target == null) {
return false;
}
// This deliberately bypasses Injector: /clear is housekeeping, not a delegated turn.
agents().send(target, "/clear");
return true;
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
@@ -1,19 +1,30 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.PlacementCandidate;
import dev.ltms.bridged.placement.PlacementContext;
import dev.ltms.bridged.placement.PlacementException;
import dev.ltms.bridged.placement.PlacementPolicies;
import dev.ltms.bridged.placement.PlacementPolicy;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.Collections;
import java.util.EnumSet;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Function;
/**
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
@@ -35,6 +46,12 @@ import java.util.concurrent.ConcurrentHashMap;
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
* shares one herdr connection and so reports the same global agent set.</li>
* </ul>
*
* <p>CB-518: an unqualified spawn is routed through a {@link PlacementPolicy}. The default
* {@code fixed} policy reproduces the historical default-profile behaviour; {@code weighted} uses
* smooth weighted round-robin with {@code maxLoad} gating. If a chosen profile fails with
* {@link PeerUnreachableException}, the composite advances to the next available candidate and
* retries, bounded by the number of candidates.
*/
public final class CompositePeerLauncher implements PeerLauncher {
@@ -47,18 +64,49 @@ public final class CompositePeerLauncher implements PeerLauncher {
/** paneId → the delegate that spawned it, so {@link #stop} tears down through the right adapter. */
private final Map<String, HerdrPeerLauncher> spawnedBy = new ConcurrentHashMap<>();
private final Map<String, BridgedConfig.Worker> profileConfigs;
private final PlacementPolicy placementPolicy;
private final Function<String, Integer> liveCount;
/**
* Backward-compatible constructor: fixed placement, no live-counting. Use this for tests and
* simple wiring; it preserves the pre-CB-518 behaviour exactly.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to (may be null)
* @throws IllegalArgumentException if {@code delegates} is empty or two adapters claim one profile
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates, String defaultProfile) {
this(delegates, defaultProfile, Map.of(), PlacementPolicies.fixed(), name -> 0);
}
/**
* Production constructor with a placement policy and live-worker counter.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to under {@code fixed} policy
* @param profileConfigs all configured worker profiles (used for candidate weights/caps)
* @param placementPolicy which policy governs unqualified spawns
* @param liveCount live worker count per profile (must never return {@code null})
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Worker> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount) {
if (delegates.isEmpty()) {
throw new IllegalArgumentException("at least one peer adapter must be configured");
}
this.delegates = List.copyOf(delegates);
this.defaultProfile = defaultProfile;
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
// differ from one JVM run to the next.
this.profileConfigs = Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs));
this.placementPolicy = placementPolicy;
this.liveCount = liveCount;
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
@@ -69,7 +117,8 @@ public final class CompositePeerLauncher implements PeerLauncher {
}
}
}
this.byProfile = Map.copyOf(index);
// Order-preserving for the same reason, and because profiles() is user-visible (bridge_profiles).
this.byProfile = Collections.unmodifiableMap(index);
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
@@ -89,10 +138,107 @@ public final class CompositePeerLauncher implements PeerLauncher {
@Override
public PeerHandle spawn(SpawnRequest req) {
HerdrPeerLauncher d = route(req.profileName());
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
String requestedProfile = req.profileName();
if (requestedProfile != null && !requestedProfile.isBlank()) {
// An explicit profile bypasses the placement policy, but not the capacity cap: maxLoad
// is documented as an unconditional limit on this profile (BridgedConfig.Worker), and
// the charter makes explicit-profile spawns the normal path — so skipping the check
// here would leave the cap dead config in real operation.
HerdrPeerLauncher d = route(requestedProfile);
enforceMaxLoad(requestedProfile);
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
}
List<PlacementCandidate> candidates = candidates();
Set<String> unreachable = new HashSet<>();
PlacementContext ctx = new PlacementContext(defaultProfile, candidates, liveCount, unreachable);
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
for (int attempt = 0; attempt < maxAttempts; attempt++) {
PlacementCandidate chosen;
try {
chosen = placementPolicy.select(ctx);
} catch (RuntimeException e) {
// No candidate left (all at cap or all unreachable). The policy already threw a clear
// message; do not wrap it in a generic PeerUnreachableException.
throw e;
}
HerdrPeerLauncher d = byProfile.get(chosen.profile());
if (d == null) {
// A configured profile with no adapter is a wiring bug; fail fast.
unreachable.add(chosen.profile());
continue;
}
// CB-547a: route the chosen profile but keep the caller's session identity — dropping it
// here would silently sever the resume handle on every policy-routed spawn.
SpawnRequest routedReq = new SpawnRequest(chosen.profile(), req.requestedCwd(), req.callerCwd(),
req.sessionName(), req.resumeSessionId());
try {
PeerHandle handle = d.spawn(routedReq);
spawnedBy.put(handle.id(), d);
return handle;
} catch (PeerUnreachableException e) {
log.warn("spawn on profile {} unreachable, will retry next candidate if any: {}",
chosen.profile(), e.getMessage());
unreachable.add(chosen.profile());
// Update the context for the next selection so the policy excludes this profile.
ctx = new PlacementContext(defaultProfile, candidates, liveCount, unreachable);
}
}
throw new PeerUnreachableException(
"no reachable worker profile available after trying " + unreachable.size()
+ " candidate(s): " + String.join(", ", unreachable));
}
/**
* Refuse an explicit-profile spawn when the profile is at its {@code maxLoad} cap.
*
* <p>maxLoad is a documented, unconditional capacity limit (see {@code BridgedConfig.Worker#maxLoad}),
* and the charter makes explicit-profile spawns the normal path — so enforcing it only in placement
* ({@link dev.ltms.bridged.placement.PlacementPolicyUtil}) would leave the cap dead config on every
* call that names a profile. Same rule as placement: {@code live >= cap} is at capacity.
*
* <p>Deliberately no fallback to another profile: the caller named {@code profile} for a cost/model
* reason, and silently re-routing a paid-tier (subscription) request elsewhere is worse than
* refusing it. A caller that wants placement should omit the profile and let the policy pick.
*
* <p>Known TOCTOU limitation — documented, not fixed. {@link #liveCount} is read outside any lock and
* {@code SessionManager} registers a session only after {@code launcher.spawn} returns, so two
* genuinely concurrent spawns can both pass this check. The race already exists on the placement
* path. Closing it needs slot reservation in the registry; serializing spawn here would block on
* the readiness gate and is a far worse trade.
*
* @param profile the profile the caller explicitly named
* @throws PlacementException when the profile is at capacity
*/
private void enforceMaxLoad(String profile) {
// Absent config, or a config whose maxLoad normalized to null (non-positive ⇒ unlimited at
// load), means no cap — never cap what wasn't configured.
BridgedConfig.Worker cfg = profileConfigs.get(profile);
Integer cap = (cfg == null) ? null : cfg.maxLoad();
if (cap == null) {
return;
}
int live = liveCount.apply(profile);
if (live >= cap) {
throw new PlacementException("worker profile '" + profile + "' is at maxLoad: " + live
+ " live >= " + cap + " cap; refusing spawn — no fallback to another profile");
}
}
/** Build the candidate list from the configured profiles, in definition order. */
private List<PlacementCandidate> candidates() {
List<PlacementCandidate> out = new ArrayList<>();
for (Map.Entry<String, BridgedConfig.Worker> e : profileConfigs.entrySet()) {
BridgedConfig.Worker w = e.getValue();
out.add(new PlacementCandidate(e.getKey(), null, w.weight(), w.maxLoad()));
}
return out;
}
@Override
@@ -115,6 +261,16 @@ public final class CompositePeerLauncher implements PeerLauncher {
d.stop(id);
}
@Override
public boolean clearContext(String id) {
HerdrPeerLauncher delegate = spawnedBy.get(id);
if (delegate == null) {
log.debug("clearContext({}) ignored — no recorded owning adapter", id);
return false;
}
return delegate.clearContext(id);
}
@Override
public Set<String> profiles() {
return byProfile.keySet();
@@ -21,6 +21,10 @@ import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ConcurrentMap;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
import java.util.function.LongSupplier;
@@ -55,6 +59,13 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
/**
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
*/
private static final int SHELL_READY_RETRIES = 20;
private final String namePrefix; // label prefix: naming + reap scheme
private final AgentControl agents;
private final WorkspaceControl spaces;
@@ -74,6 +85,14 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
// CB-519: PeerHandle.id() is a host-unique opaque UUID, decoupled from the herdr pane id. The
// routing/registry key is the UUID; the herdr pane id is a launcher-private placement/teardown
// coordinate. This map bridges the two so stop(id) can resolve a host-unique key back to the
// exact pane it must tear down. The pane id is launcher-private (never the routing key) — see
// PeerHandle.id().
private final ConcurrentMap<String, String> paneByAgentId = new ConcurrentHashMap<>();
private final AtomicBoolean resetUnsupportedLogged = new AtomicBoolean();
/**
* @param namePrefix label prefix for this peer kind (drives naming and reap)
* @param agents herdr agent control (start, status, close)
@@ -111,8 +130,52 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
*/
protected abstract Launch buildLaunch(BridgedConfig.Worker cfg);
/** A peer-specific launch: the herdr {@code env} map and {@code argv}. */
protected record Launch(Map<String, String> env, List<String> argv) {
/**
* Session-aware variant of {@link #buildLaunch(BridgedConfig.Worker)} (CB-547a). Default
* discards the session identity and delegates to the profile-only form, so an adapter that
* carries no durable peer session (opencode, say) inherits byte-identical behaviour and needs
* no change. An adapter that does (Claude Code) overrides this to mint/resume the id and to
* surface it on the returned {@link Launch#agentSessionId()}.
*
* @param cfg the resolved profile to spawn
* @param sessionName the bridge's logical session name, or null/blank for launcher-derived
* @param resumeSessionId the peer's own prior session id to resume, or null/blank for fresh
*/
protected Launch buildLaunch(BridgedConfig.Worker cfg, String sessionName, String resumeSessionId) {
return buildLaunch(cfg);
}
/** Direct transport access for peer-specific, non-turn control operations. */
protected final AgentControl agents() {
return agents;
}
/** Resolve the public peer id to the launcher's private herdr target. */
protected final String agentTarget(String id) {
return paneByAgentId.get(id);
}
@Override
public boolean clearContext(String id) {
if (resetUnsupportedLogged.compareAndSet(false, true)) {
log.warn("context reset is unsupported for peer kind {}; clearAfterTurn is a no-op",
namePrefix);
}
return false;
}
/**
* A peer-specific launch: the herdr {@code env} map and {@code argv}, plus — for an adapter
* that carries durable session identity (CB-547a) — the peer's OWN session id
* ({@link PeerHandle#agentSessionId()}), known before the peer has written anything. Null for
* a launch that carries no identity.
*/
protected record Launch(Map<String, String> env, List<String> argv, String agentSessionId) {
/** A launch without a discoverable agent session id (an adapter that carries none). */
Launch(Map<String, String> env, List<String> argv) {
this(env, argv, null);
}
}
// --- profile surface -----------------------------------------------------------------------
@@ -162,6 +225,10 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// --- spawn ---------------------------------------------------------------------------------
/** A started peer plus the launch's agent-session id (the resume handle, or null). */
private record Spawned(Agent agent, String agentSessionId) {
}
/**
* Spawn a peer. {@code profileName} null/blank → the default profile. The working directory
* (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd}, else the
@@ -170,31 +237,52 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* The adapter's {@link #buildLaunch} runs before any herdr call.
*/
protected Agent spawnInternal(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd, null, null).agent();
}
/**
* Spawn a peer with session identity (CB-547a). {@code sessionName} and {@code resumeSessionId}
* are threaded from the {@link SpawnRequest} into {@link #buildLaunch(BridgedConfig.Worker,
* String, String)}, and the launch's resolved agent-session id is returned alongside the agent
* so the caller can put it on the {@link PeerHandle}.
*/
protected Spawned spawnInternal(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId) {
BridgedConfig.Worker cfg = requireProfile(profileName);
Launch launch = buildLaunch(cfg);
Launch launch = buildLaunch(cfg, sessionName, resumeSessionId);
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
return cfg.tabPlacement()
Agent agent = cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd);
return new Spawned(agent, launch.agentSessionId());
}
/**
* {@inheritDoc}
*
* <p>Delegates to {@link #spawnInternal} and wraps the resulting herdr {@link Agent} in a
* {@link WorkerHandle} whose {@link PeerHandle#id()} equals the agent's paneId. When
* {@code spawnReadyTimeoutMs > 0}, blocks until the peer's herdr status is injectable or the
* timeout elapses; on timeout the pane is closed (no orphan) and a
* {@link PeerUnreachableException} is thrown.
* {@link WorkerHandle} whose {@link PeerHandle#id()} is a fresh <em>host-unique</em> opaque
* UUID (CB-519), deliberately decoupled from the herdr pane id: the id is the registry/routing
* key and must never collide across daemon processes on the same host, while the herdr pane id
* stays a launcher-private placement/teardown coordinate, remembered here so {@link #stop}
* can resolve the host-unique key back to its pane. When {@code spawnReadyTimeoutMs > 0},
* blocks until the peer's herdr status is injectable or the timeout elapses; on timeout the
* pane is closed (no orphan) and a {@link PeerUnreachableException} is thrown.
*/
@Override
public PeerHandle spawn(SpawnRequest req) {
Agent agent = spawnInternal(req.profileName(), req.requestedCwd(), req.callerCwd());
Spawned spawned = spawnInternal(req.profileName(), req.requestedCwd(), req.callerCwd(),
req.sessionName(), req.resumeSessionId());
Agent agent = spawned.agent();
String paneId = agent.paneId();
if (spawnReadyTimeoutMs > 0) {
waitUntilInjectableOrThrow(paneId);
}
return new WorkerHandle(paneId, agent.terminalId());
// CB-519: the handle id is a host-unique UUID; the herdr pane it maps to stays internal.
String id = UUID.randomUUID().toString();
paneByAgentId.put(id, paneId);
return new WorkerHandle(id, agent.terminalId(), requireProfile(req.profileName()).profile(),
req.sessionName(), spawned.agentSessionId());
}
@Override
@@ -227,17 +315,23 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
return null;
}
/** Dedicated worker space → own tab → start the peer (rooted at {@code cwd}) → drop the shell. */
/** Dedicated worker space → own tab (carrying cwd+env) → start the peer into the seed pane. */
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId());
Tab.Created tab = spaces.createTab(space.workspaceId(), cwd, workerEnv);
log.info("spawning {} profile={} space={} tab={} cwd={}",
namePrefix, cfg.profile(), space.workspaceId(), tab.tab().tabId(), cwd);
Started started;
try {
started = startUniquelyNamed(cfg, workerEnv, argv, tab.tab().tabId(), cwd);
if (tab.rootPaneId() == null) {
// Protocol 19 starts the agent INTO the seed pane — without one there is nowhere
// to start, and a partial tab would be left behind.
throw new IllegalStateException("tab " + tab.tab().tabId()
+ " had no seed pane in the create response — cannot start a peer in it");
}
started = startUniquelyNamed(cfg, argv, tab.rootPaneId());
} catch (RuntimeException e) {
// The peer never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
@@ -250,16 +344,9 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
throw e;
}
// The peer is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell so the
// tab holds only the peer; label the tab). They must not fail the spawn or orphan the
// running peer — on error we log and still return it so the caller gets its paneId and can
// tear it down.
if (tab.rootPaneId() != null) {
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
} else {
log.warn("tab {} had no seed pane in the create response; peer tab may hold an extra pane",
tab.tab().tabId());
}
// The peer is LIVE now, in the seed pane itself (no shell pane to drop — protocol 19).
// Labelling is cosmetic: it must not fail the spawn or orphan the running peer — on error
// we log and still return it so the caller gets its paneId and can tear it down.
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
log.info("{} started pane={} tab={} terminal={}",
@@ -276,12 +363,16 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
}
/** Legacy placement: herdr splits the currently-focused tab; the peer still starts in {@code cwd}. */
/** Legacy placement: split the currently-focused tab; the peer still starts in {@code cwd}. */
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
namePrefix, cfg.profile(), cwd, argv);
Agent peer = startUniquelyNamed(cfg, workerEnv, argv, null, cwd).agent();
String paneId = spaces.splitPane(cwd, workerEnv);
if (paneId == null) {
throw new IllegalStateException("pane.split returned no pane — cannot start a peer");
}
Agent peer = startUniquelyNamed(cfg, argv, paneId).agent();
log.info("{} started pane={} terminal={}", namePrefix, peer.paneId(), peer.terminalId());
return peer;
}
@@ -300,14 +391,16 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* backstop for the astronomically unlikely nonce+seq clash; the name is a label only — herdr
* detects kind and status from terminal output, not from it.
*/
private Started startUniquelyNamed(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String tabId, String cwd) {
private Started startUniquelyNamed(BridgedConfig.Worker cfg, List<String> argv, String paneId) {
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
// argv[0] — the configured executable — is dropped and only the extra args are passed.
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(agents.start(name, argv, workerEnv, tabId, cwd), seq);
return new Started(startAwaitingShellPrompt(name, args, paneId), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("peer name '{}' taken, retrying", name);
@@ -317,6 +410,22 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
throw last;
}
/** Start the agent into {@code paneId}, waiting out the seed shell's boot with the sleeper. */
private Agent startAwaitingShellPrompt(String name, List<String> args, String paneId) {
HerdrException busy = null;
for (int attempt = 0; attempt < SHELL_READY_RETRIES; attempt++) {
try {
return agents.start(name, namePrefix, args, paneId);
} catch (HerdrException e) {
if (!"agent_pane_busy".equals(e.code())) throw e;
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
busy = e;
sleeper.run();
}
}
throw busy;
}
// --- discovery + reap ----------------------------------------------------------------------
/** All herdr-tracked agents — discovery for "what peers exist". */
@@ -398,21 +507,31 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// --- teardown ------------------------------------------------------------------------------
/**
* Tear a peer down by pane id: close the pane, and close its tab <em>only</em> when the peer is
* that tab's sole occupant. The single-pane check is what makes this safe regardless of how the
* peer was placed (or a placement-config change across a restart): a pane-placement peer sitting
* in one of the user's shared tabs has siblings, so its tab is never closed — we only ever
* remove a tab we created to hold one peer.
* Tear a peer down: close the pane, and close its tab <em>only</em> when the peer is that tab's
* sole occupant. The single-pane check is what makes this safe regardless of how the peer was
* placed (or a placement-config change across a restart): a pane-placement peer sitting in one
* of the user's shared tabs has siblings, so its tab is never closed — we only ever remove a
* tab we created to hold one peer.
*
* <p>{@code idOrPane} is the {@link PeerHandle#id()} of a peer this launcher spawned (CB-519's
* host-unique opaque UUID), resolved through {@link #paneByAgentId} to the pane it must tear
* down. An argument that is not one of our ids is treated as a raw herdr pane id — the
* {@link #reapOrphanWorkers() orphan-reap} and spawn-gate-timeout paths, plus any caller that
* passes a pane directly, keep working without an owning id.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed peer) is treated as success; any other failure propagates so a
* genuinely failed teardown is not reported as done.
*/
@Override
public void stop(String paneId) {
// Teardown knows only the paneId, not which profile spawned it. Attempt tab cleanup when any
public void stop(String idOrPane) {
// Teardown knows only the pane, not which profile spawned it. Attempt tab cleanup when any
// profile uses tab placement (so the bridge may have created a dedicated peer tab); the
// single-occupant check below is what actually protects the user's shared tabs.
String paneId = paneByAgentId.remove(idOrPane);
if (paneId == null) {
paneId = idOrPane; // raw-pane fallback (reap, gate timeout, pane-addressed callers)
}
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
@@ -460,8 +579,13 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
+ spawnReadyTimeoutMs + "ms");
}
/** A concrete {@link PeerHandle} wrapping herdr agent coordinates. */
private record WorkerHandle(String id, String terminalId) implements PeerHandle {
/**
* A concrete {@link PeerHandle} wrapping herdr agent coordinates, the profile that spawned it,
* and the session identity the launch resolved (CB-547a): the bridge's logical name and the
* peer's own session id, both null when the spawn carried no identity.
*/
private record WorkerHandle(String id, String terminalId, String profile,
String sessionName, String agentSessionId) implements PeerHandle {
}
// --- shared helpers ------------------------------------------------------------------------
@@ -7,6 +7,8 @@ import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.SpawnRequest;
import java.io.IOException;
import java.io.UncheckedIOException;
@@ -72,6 +74,24 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
/** Root under which per-spawn opencode config dirs are created (injectable for tests). */
private final Path configRoot;
/**
* The current spawn's resume-target session id, threaded from {@link #spawn(SpawnRequest)} to
* {@link #buildLaunch} across the base's {@code spawn -> spawnInternal -> buildLaunch} chain,
* which carries no request. A plain field would race under concurrent spawns (the base supports
* them), so it is thread-local: each spawn captures its own request's id on its own thread, and
* {@code buildLaunch}, synchronous and same-thread, reads exactly that one. Set only around the
* {@code super.spawn} call and cleared in {@code finally}, so a paused/leftover value can never
* bleed into the next spawn.
*/
private final ThreadLocal<String> resumeSessionId = new ThreadLocal<>();
/**
* Session discovery against opencode's on-disk storage ({@link OpenCodeSessionDiscovery}) —
* the one seam that knows opencode's private session-file layout. Its root is injectable for
* tests so they never touch the operator's real {@code ~/.local/share/opencode}.
*/
private final OpenCodeSessionDiscovery discovery;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so it
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
@@ -81,7 +101,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
Function<String, String> env) {
this(agents, spaces, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300),
defaultConfigRoot());
defaultConfigRoot(), defaultDiscoveryRoot());
}
/**
@@ -94,7 +114,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot());
defaultConfigRoot(), defaultDiscoveryRoot());
}
/**
@@ -112,21 +132,31 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* @param sleeper sleep/wait hook (encodes the poll interval; never called when the
* gate is disabled)
* @param configRoot existing directory under which per-spawn config dirs are created
* @param discoveryRoot opencode's on-disk storage root to scan for session records
* (injectable for tests; opencode's layout is matched at
* {@link OpenCodeSessionDiscovery})
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper, Path configRoot) {
LongSupplier nowMillis, Runnable sleeper,
Path configRoot, Path discoveryRoot) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.configRoot = configRoot;
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
}
private static Path defaultConfigRoot() {
return Path.of(System.getProperty("java.io.tmpdir"));
}
/** The default opencode storage root: {@code ~/.local/share/opencode} (the XDG data dir). */
private static Path defaultDiscoveryRoot() {
return Path.of(System.getProperty("user.home"), ".local", "share", "opencode");
}
/**
* {@inheritDoc}
*
@@ -144,7 +174,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg).toString());
}
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithModel(cfg));
return new Launch(workerEnv, argvWithResume(argvWithModel(argvWithAuto(cfg), cfg)));
}
/**
@@ -161,9 +191,42 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
return cfg.baseUrl() != null && !cfg.baseUrl().isBlank();
}
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
private List<String> argvWithModel(BridgedConfig.Worker cfg) {
/**
* The launch argv plus the unconditional {@code --auto} flag, which auto-approves the
* permissions opencode does not explicitly deny. It is unconditional, not a preference: a
* spawned peer has no human at its pane — the bridge spawned it — so one that stops at an
* approval prompt is a wedged agent, indistinguishable from a legitimate mid-turn wait and
* unable to end its turn with {@code bridge_reply}. opencode's own help calls this
* "dangerous!", but the blast radius here is already bounded by design: a worker runs in its
* own git worktree on its own branch, is off-subscription, and cannot merge — the lead is the
* gate.
*/
private List<String> argvWithAuto(BridgedConfig.Worker cfg) {
List<String> argv = mutableArgv(cfg.argv());
argv.add("--auto");
return argv;
}
/**
* The launch argv plus, on a resumed spawn, opencode's {@code -s <id>} flag to continue a prior
* conversation by its session id. {@code -s, --session <id>} resumes an existing session; on a
* fresh spawn (no resume target) no flag is added, letting opencode start a brand-new session.
* The id comes from the current spawn request's {@code resumeSessionId}, threaded per-thread by
* {@link #spawn(SpawnRequest)}.
*/
private List<String> argvWithResume(List<String> argv) {
String id = resumeSessionId.get();
if (id == null || id.isBlank()) {
return argv;
}
List<String> withResume = mutableArgv(argv);
withResume.add("-s");
withResume.add(id);
return withResume;
}
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
private List<String> argvWithModel(List<String> argv, BridgedConfig.Worker cfg) {
if (cfg.model() != null && !cfg.model().isBlank()) {
argv.add("-m");
argv.add(cfg.model());
@@ -270,6 +333,82 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
}
/**
* {@inheritDoc}
*
* <p>adds this adapter's session-identity work around the base's spawn — as opencode cannot be
* told its session id at spawn (see {@link Capability#SESSION_RESUME} vs
* {@link Capability#SESSION_NAME}), identity is only ever adopted after the fact:
* <ul>
* <li>the request's {@code resumeSessionId} is remembered for {@link #buildLaunch} to turn
* into {@code -s <id>}; and</li>
* <li>the returned handle is wrapped so its
* {@link dev.ltms.bridged.peer.PeerHandle#agentSessionId()} performs lazy session
* discovery against opencode's storage (see {@link OpenCodeSessionDiscovery}) — always
* non-blocking, {@code null} until opencode has persisted the session record.</li>
* </ul>
*/
@Override
public PeerHandle spawn(SpawnRequest req) {
resumeSessionId.set(req.resumeSessionId());
try {
PeerHandle inner = super.spawn(req);
return new SessionAwareHandle(inner, discovery, effectiveCwd(req));
} finally {
// Never let a paused/leftover resume id bleed into the next spawn on this thread.
resumeSessionId.remove();
}
}
/**
* A {@link PeerHandle} that delegates everything to the base's worker handle but resolves
* {@link #agentSessionId()} lazily through opencode session discovery. Delegate-only, so the
* base's id/terminalId/profile semantics (CB-519's host-unique routing key, herdr coordinates)
* are untouched — only the opencode-specific identity answer is added. {@code sessionName()}
* stays null: opencode has no display-name seam, so the logical name lives only in the bridge's
* roster (see the SESSION_NAME capability).
*/
private static final class SessionAwareHandle implements PeerHandle {
private final PeerHandle delegate;
private final OpenCodeSessionDiscovery discovery;
private final String cwd;
SessionAwareHandle(PeerHandle delegate, OpenCodeSessionDiscovery discovery, String cwd) {
this.delegate = delegate;
this.discovery = discovery;
this.cwd = cwd;
}
@Override
public String id() {
return delegate.id();
}
@Override
public String terminalId() {
return delegate.terminalId();
}
@Override
public String profile() {
return delegate.profile();
}
@Override
public String sessionName() {
return delegate.sessionName();
}
@Override
public String agentSessionId() {
// Lazy + retried, never a spawn-time blocker: opencode writes the session record only
// when the session is first persisted, so null here is the correct interim answer and
// the caller re-calls later (each call re-scans, picking up a record that has since
// appeared).
return discovery.sessionIdForDirectory(cwd);
}
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
@@ -291,10 +430,15 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE,
Capability.ORPHAN_REAP, Capability.SESSION_RESUME);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
// Deliberately NOT SESSION_NAME: opencode has no display-name flag, so the bridge's logical
// name can't surface in the peer's own UI — declaring the capability would hide that
// asymmetry rather than make it honest. For opencode the name lives only in the bridge's
// roster (see PeerHandle.sessionName() returning null).
return Set.copyOf(caps);
}
@@ -0,0 +1,122 @@
package dev.ltms.bridged.worker;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.stream.Stream;
/**
* Resolves the opencode session id for a bridged worker from opencode's on-disk storage — the
* only place this adapter touches opencode's private layout, and deliberately the <em>only</em>
* class that does.
*
* <p><strong>Why this is isolated behind one seam.</strong> The layout is version-coupled and not a
* stable contract: opencode writes one JSON file per session under
* {@code <storageRoot>/session/<projectID>/<ses_*.json>}, and each record carries a
* {@code "version"} field (e.g. {@code "1.1.31"}), so the exact directory shape, file naming, and
* field names can move between opencode releases. opencode also ships a headless HTTP server that
* may supersede file scanning entirely. Everything this adapter knows about that private storage —
* its shape, naming, and field names — lives here, so a layout change, or a switch to the HTTP
* server, changes exactly one class and nothing in {@link OpenCodeLauncher}.
*
* <p>The determinism that makes this useful is structural, not a guess: every bridged worker runs
* in its own unique git worktree, so the record's {@code directory} (its project root) equals the
* worker's cwd identifies <em>its</em> session unambiguously. We match on {@code directory} rather
* than diffing {@code opencode session list} before/after — that races under concurrent spawns, and
* the CLI listing does not even show the directory.
*
* <p>All reads are best-effort and never throw: a missing or unreadable storage root, a record that
* fails to parse, or a directory with no record yet all yield {@code null}, and the caller (the
* session handle) treats that as "identity not resolved yet" and retries later.
*/
final class OpenCodeSessionDiscovery {
private final Path storageRoot; // e.g. ~/.local/share/opencode (injectable for tests)
private final ObjectMapper json;
OpenCodeSessionDiscovery(Path storageRoot) {
this.storageRoot = storageRoot;
this.json = new ObjectMapper();
}
/**
* The opencode session id whose record references {@code directory} (the worker's cwd), or
* {@code null} when no record matches yet. When several records share the directory — e.g.
* repeated spawns into the same worktree — the <em>most recently modified</em> one wins: it is
* the session the pane most likely corresponds to.
*
* <p>Never throws: a missing {@code storageRoot}, an unreadable/malformed record, or a
* directory that has not been persisted yet all resolve to {@code null} rather than failing a
* spawn. A bridged worker's session record is written lazily (when the session is first
* persisted), so {@code null} here is the normal answer right after the pane is ready, and the
* caller retries later.
*
* @param directory the worker's cwd, as resolved for this spawn
* @return the matching session id, or {@code null} if none is known yet
*/
String sessionIdForDirectory(String directory) {
if (directory == null || directory.isBlank()) {
return null;
}
Path sessionRoot = storageRoot.resolve("session");
if (!Files.isDirectory(sessionRoot)) {
return null;
}
String best = null;
long bestMtime = Long.MIN_VALUE;
try (Stream<Path> projectDirs = Files.list(sessionRoot)) {
for (Path projectDir : projectDirs.filter(Files::isDirectory).toList()) {
try (Stream<Path> records = Files.list(projectDir)) {
for (Path record : records.toList()) {
String id = matchId(record, directory);
if (id == null) {
continue;
}
long mtime = lastModifiedEpochMillis(record);
if (mtime > bestMtime) {
bestMtime = mtime;
best = id;
}
}
} catch (IOException ignored) {
// one project dir unreadable — skip it; another may still match
}
}
} catch (IOException ignored) {
// storage root vanished or became unreadable — "no session known yet"
return null;
}
return best;
}
/**
* The record's session id when it references {@code directory}, else {@code null}. A record
* that is not JSON, lacks {@code id}/{@code directory}, or points at a different directory is
* simply not our session; a malformed one is skipped, never fatal.
*/
private String matchId(Path record, String directory) {
try {
JsonNode node = json.readTree(record.toFile());
JsonNode id = node == null ? null : node.get("id");
JsonNode dir = node == null ? null : node.get("directory");
if (id == null || dir == null || !directory.equals(dir.asText())) {
return null;
}
return id.asText();
} catch (IOException e) {
return null;
}
}
/** The record's last-modified epoch ms, or {@code Long.MIN_VALUE} if unreadable (never wins). */
private static long lastModifiedEpochMillis(Path record) {
try {
return Files.getLastModifiedTime(record).toMillis();
} catch (IOException e) {
return Long.MIN_VALUE;
}
}
}
@@ -0,0 +1,86 @@
package dev.ltms.bridged;
import dev.ltms.bridged.inject.WorkerPresence;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.HashMap;
import java.util.Map;
import java.util.function.Predicate;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-534: the injector's readiness gate must open for a lead as well as for a present worker.
*
* <p>The bug these cover was silent and slow: a lead was never marked present (only workers are), so
* every lead→lead delivery sat on the gate for the full readiness grace and failed ~60s later without
* a keystroke ever reaching the pane.
*/
class BridgedDeliverabilityTest {
private static Supplier<Map<String, String>> leads(Map<String, String> m) {
return () -> m;
}
@Test
@DisplayName("a worker that has connected its MCP is deliverable")
void presentWorkerIsDeliverable() {
WorkerPresence presence = new WorkerPresence();
presence.markPresent("term_worker");
assertTrue(Bridged.deliverableTo(presence, leads(Map.of())).test("term_worker"));
}
@Test
@DisplayName("a worker still in its boot window is held back")
void absentWorkerIsNotDeliverable() {
assertFalse(Bridged.deliverableTo(new WorkerPresence(), leads(Map.of())).test("term_booting"));
}
@Test
@DisplayName("a lead is deliverable without ever being marked present")
void leadIsDeliverableWithoutPresence() {
WorkerPresence presence = new WorkerPresence();
Predicate<String> deliverable =
Bridged.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
assertFalse(presence.isPresent("term_lead"), "a lead is never enrolled in worker presence");
assertTrue(deliverable.test("term_lead"), "…and must be deliverable anyway");
}
@Test
@DisplayName("an unknown terminal is deliverable to neither")
void strangerIsNotDeliverable() {
WorkerPresence presence = new WorkerPresence();
presence.markPresent("term_worker");
assertFalse(Bridged.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")))
.test("term_stranger"));
}
@Test
@DisplayName("a lead discovered after startup becomes deliverable with no restart")
void leadSetIsReadThroughOnEveryCall() {
Map<String, String> discovered = new HashMap<>();
Predicate<String> deliverable = Bridged.deliverableTo(new WorkerPresence(), leads(discovered));
assertFalse(deliverable.test("term_late"));
discovered.put("term_late", "gpt-sol-5.6"); // leadScan picks up a newly labelled tab
assertTrue(deliverable.test("term_late"), "the supplier must be re-read, not snapshotted");
}
@Test
@DisplayName("forgetting a torn-down worker does not strip a lead of its deliverability")
void forgetDoesNotDisarmALead() {
WorkerPresence presence = new WorkerPresence();
Predicate<String> deliverable =
Bridged.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
presence.forget("term_lead"); // the injector's cleanup path runs against every target
assertTrue(deliverable.test("term_lead"));
}
}
@@ -0,0 +1,238 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.config.BridgedConfig;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.Future;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-548 — the architect-slot registry: the config snapshot of slot → profile, and the live
* terminal → slot bindings it owns. The role a binding produces is asserted in
* {@link CallerResolverTest}; this pins the registry object itself — its invariants and their
* thread-safety.
*/
class ArchitectRegistryTest {
private static final Map<String, BridgedConfig.Architect> SLOTS = Map.of(
"lead-designer", new BridgedConfig.Architect("sonnet"),
"reviewer", new BridgedConfig.Architect("gx10"));
private final ArchitectRegistry registry = new ArchitectRegistry(SLOTS);
@Test
void exposesTheConfiguredSlots() {
assertEquals(SLOTS.keySet(), registry.slots().keySet());
assertTrue(registry.isSlot("reviewer"));
assertFalse(registry.isSlot("nope"));
}
@Test
void theSpawnLifecycleReadsTheProfileBackFromASlot() {
assertEquals("sonnet", registry.profileForSlot("lead-designer"));
assertEquals("gx10", registry.profileForSlot("reviewer"));
assertNull(registry.profileForSlot("unknown"), "an unknown slot has no profile");
}
@Test
void startsEmptySoNoTerminalResolvesToAnArchitect() {
assertTrue(registry.snapshot().isEmpty());
assertNull(registry.slotForTerminal("term_design"),
"config declares no architect terminal — nothing is recognised until a bind");
assertNull(registry.slotForTerminal(null), "no terminal ⇒ no slot");
}
// ── bind ──────────────────────────────────────────────────────────────────────────────────
@Test
void bindResolvesTheTerminalToTheSlot() {
assertTrue(registry.bind("lead-designer", "term_design"));
assertEquals("lead-designer", registry.slotForTerminal("term_design"));
assertEquals(Map.of("term_design", "lead-designer"), registry.snapshot());
}
@Test
void bindRefusesAnUnknownSlot() {
assertFalse(registry.bind("nope", "term_x"),
"a slot that is not configured must be refused — bind is not a way to invent one");
assertNull(registry.slotForTerminal("term_x"));
}
@Test
void bindRefusesATerminalInTwoSlots() {
assertTrue(registry.bind("lead-designer", "term_design"));
assertFalse(registry.bind("reviewer", "term_design"),
"a terminal may occupy at most one slot");
assertEquals("lead-designer", registry.slotForTerminal("term_design"),
"the first binding survives the refused second");
}
@Test
void bindRefusesASlotWithTwoTerminals() {
assertTrue(registry.bind("lead-designer", "term_design"));
assertFalse(registry.bind("lead-designer", "term_other"),
"a slot may host at most one terminal");
assertEquals("lead-designer", registry.slotForTerminal("term_design"),
"the first binding survives the refused second");
assertNull(registry.slotForTerminal("term_other"));
}
@Test
void rebindingTheSamePairIsAnIdempotentNoOp() {
assertTrue(registry.bind("lead-designer", "term_design"));
assertTrue(registry.bind("lead-designer", "term_design"),
"the same terminal → slot is harmless to repeat");
assertEquals(1, registry.snapshot().size());
}
// ── unbind ────────────────────────────────────────────────────────────────────────────────
@Test
void unbindRemovesTheExactBinding() {
assertTrue(registry.bind("lead-designer", "term_design"));
assertTrue(registry.unbind("lead-designer", "term_design"));
assertNull(registry.slotForTerminal("term_design"));
assertTrue(registry.snapshot().isEmpty());
}
@Test
void aStaleUnbindDoesNotRemoveAReplacement() {
// Bind, tear down, and stand the slot back up with a NEW terminal.
assertTrue(registry.bind("lead-designer", "term_design"));
registry.unbind("lead-designer", "term_design");
assertTrue(registry.bind("lead-designer", "term_new"));
// A late unbind naming the OLD terminal must not remove the replacement binding.
assertFalse(registry.unbind("lead-designer", "term_design"));
assertEquals("lead-designer", registry.slotForTerminal("term_new"),
"the replacement terminal stays bound");
}
@Test
void aStaleUnbindForATerminalThatMovedSlotsDoesNothing() {
// term_design starts in lead-designer, is torn down, and stands back up in a FREE slot.
assertTrue(registry.bind("lead-designer", "term_design"));
registry.unbind("lead-designer", "term_design");
assertTrue(registry.bind("reviewer", "term_design"));
// Unbinding against the slot it no longer occupies is refused; the new binding is intact.
assertFalse(registry.unbind("lead-designer", "term_design"),
"the old slot must not unbind a terminal that moved elsewhere");
assertEquals("reviewer", registry.slotForTerminal("term_design"));
}
@Test
void unbindOfNothingIsAFalseNoOp() {
assertFalse(registry.unbind("lead-designer", "term_design"),
"nothing was bound, so nothing is removed");
}
// ── snapshot ─────────────────────────────────────────────────────────────────────────────
@Test
void theSnapshotIsAnImmutableCopyNotAliveState() {
assertTrue(registry.bind("lead-designer", "term_design"));
Map<String, String> snap = registry.snapshot();
assertThrows(UnsupportedOperationException.class, () -> snap.put("x", "y"),
"a handed-out snapshot cannot be mutated in place");
// Later binds must not leak into an earlier snapshot.
assertTrue(registry.bind("reviewer", "term_review"));
assertFalse(snap.containsKey("term_review"),
"a snapshot is a point-in-time copy, not a live view");
}
// ── concurrency (CB-548 invariants hold under contention) ─────────────────────────────────
@Test
void concurrentBindsNeverGiveASlotTwoTerminals() throws Exception {
int n = 16;
ExecutorService pool = Executors.newFixedThreadPool(n);
try {
CountDownLatch go = new CountDownLatch(1);
List<Future<Boolean>> results = new ArrayList<>();
for (int i = 0; i < n; i++) {
final String term = "term_" + i; // every thread races for the SAME slot
results.add(pool.submit(() -> {
go.await();
return registry.bind("lead-designer", term);
}));
}
go.countDown();
int won = 0;
for (Future<Boolean> r : results) {
if (r.get()) {
won++;
}
}
assertEquals(1, won, "exactly one terminal may win the sole slot, got " + won);
assertEquals(1, registry.snapshot().size(),
"the slot hosts at most one terminal after the race");
} finally {
pool.shutdownNow();
}
}
@Test
void concurrentBindsNeverPutOneTerminalInTwoSlots() throws Exception {
int n = 16;
ExecutorService pool = Executors.newFixedThreadPool(n);
try {
CountDownLatch go = new CountDownLatch(1);
List<Future<String>> results = new ArrayList<>();
for (int i = 0; i < n; i++) {
final String slot = (i % 2 == 0) ? "lead-designer" : "reviewer"; // all race for ONE terminal
results.add(pool.submit(() -> {
go.await();
return registry.bind(slot, "shared_term")
? registry.slotForTerminal("shared_term") : null;
}));
}
go.countDown();
// Rebinding the same terminal to the same slot is a harmless idempotent true, so count
// winners is not the assertion — agreement is: every thread that reported success must
// have seen the terminal in the SAME slot, never in two at once.
String bound = null;
boolean conflict = false;
for (Future<String> r : results) {
String s = r.get();
if (s != null) {
if (bound == null) {
bound = s;
} else if (!bound.equals(s)) {
conflict = true;
}
}
}
assertFalse(conflict, "a terminal was observed in two slots at once");
assertNotNull(bound, "at least one thread bound the terminal");
assertEquals(1, registry.snapshot().size(),
"the terminal occupies exactly one slot in the final snapshot");
assertEquals(bound, registry.slotForTerminal("shared_term"));
} finally {
pool.shutdownNow();
}
}
/** A handed-over slot map is snapshotted at construction, not offered as live state. */
@Test
void theSlotSnapshotIsFixedByConstruction() {
Map<String, BridgedConfig.Architect> mutable = new HashMap<>(SLOTS);
ArchitectRegistry r = new ArchitectRegistry(mutable);
mutable.put("hijack", new BridgedConfig.Architect("gx10"));
assertFalse(r.isSlot("hijack"), "a handed-over map is not offered as live state");
}
}
@@ -12,6 +12,8 @@ class AuthzTest {
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal WORKER_B = Principal.worker("term_b", 300);
private static final Principal ANON = Principal.anonymous();
private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400);
private static final Principal ARCH_OTHER = Principal.architect("reviewer", "term_review", 500);
@Test
void anonymousIsAuthorizedForNothing() {
@@ -61,6 +63,48 @@ class AuthzTest {
"an absent session id must not satisfy the own-session rule");
}
// ── CB-548: the architect matrix ───────────────────────────────────────────────────────────
@Test
void anArchitectMaySendButNotSpawnStopOrDrain() {
assertTrue(Authz.permits(ARCH_DESIGN, SEND, "term_worker"),
"delegating a turn to a worker IS the architect's job");
assertTrue(Authz.permits(ARCH_DESIGN, SEND, null));
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, DRAIN}) {
assertFalse(Authz.permits(ARCH_DESIGN, a, null),
"an architect must not " + a + " — fleet lifecycle is the primary's alone, so "
+ "a coordinator cannot also stand up or tear down the fleet");
}
}
@Test
void anArchitectMayReplyAndAskOnlyAsItsOwnPane() {
assertTrue(Authz.permits(ARCH_DESIGN, REPLY, "term_design"), "its own pane is its own");
assertTrue(Authz.permits(ARCH_DESIGN, ASK, "term_design"));
assertFalse(Authz.permits(ARCH_DESIGN, REPLY, "term_review"),
"architect 'lead-designer' must not reply on reviewer's pane");
assertFalse(Authz.permits(ARCH_OTHER, ASK, "term_design"),
"reviewer must not ask as lead-designer — no terminal is another's");
assertFalse(Authz.permits(ARCH_DESIGN, REPLY, null),
"an absent target must not pass the own-session rule");
}
@Test
void anArchitectMayReadAndScrapeMetrics() {
assertTrue(Authz.permits(ARCH_DESIGN, READ, null));
assertTrue(Authz.permits(ARCH_DESIGN, METRICS, null));
}
@Test
void anArchitectIsNotCountedAsPrimaryOrWorker() {
assertFalse(Authz.permits(ARCH_DESIGN, SPAWN, null), "not a primary — no lifecycle");
assertFalse(ARCH_DESIGN.isPrimary());
assertFalse(ARCH_DESIGN.isWorker(), "an architect is its own role, not a widened worker");
assertTrue(ARCH_DESIGN.isArchitect());
}
@Test
void observationIsOpenToBothAuthenticatedRoles() {
assertTrue(Authz.permits(PRIMARY, READ, null));
@@ -5,6 +5,8 @@ import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import org.junit.jupiter.api.Test;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/**
@@ -44,6 +46,43 @@ class CallerResolverTest {
assertEquals("term_a", underToken.terminal());
}
@Test
void aPinnedPrimaryTerminalResolvesToPrimaryNotWorker() {
// The primary's own session lives in a herdr pane (term_a here). Without the pin the pane
// match wins and the primary is locked out of spawn/send/stop as a misread worker.
Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_a")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
}
@Test
void aPinnedPrimaryTerminalNeedsNoTokenEvenInTokenMode() {
Principal p = CallerResolver.pinnedTo(workerIdentity(), true, "s3cret", "term_a")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role(),
"the pane mapping is as unforgeable as a worker's — the pin outranks the token path");
}
@Test
void otherPanesRemainWorkersWhenAPinIsSet() {
Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_someone_else")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertEquals("term_a", p.terminal());
}
/** The pin is optional config, so an absent or whitespace one must change nothing at all. */
@Test
void aBlankPinLeavesWorkerResolutionUntouched() {
assertEquals(Role.WORKER,
CallerResolver.pinnedTo(workerIdentity(), false, null, " ").resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER,
CallerResolver.pinnedTo(workerIdentity(), false, null, null).resolve("127.0.0.1", 42, null).role());
}
@Test
void loopbackTrustTreatsANonWorkerLoopbackCallerAsThePrimary() {
Principal p = new CallerResolver(nonWorkerIdentity()).resolve("127.0.0.1", 99, null);
@@ -96,6 +135,260 @@ class CallerResolverTest {
assertEquals(Role.ANONYMOUS, p.role());
}
// ── CB-530: the leaders registry ────────────────────────────────────────────────────────────
// The fake resolves exactly one pane (term_a) from a PID, so "two leads both resolve" is
// asserted at the config layer (BridgedConfigTest#leaderTerminals…). What matters here is that
// resolution is a REGISTRY LOOKUP rather than a single equality test against one pin.
@Test
void aRegisteredLeadPaneResolvesToPrimaryCarryingItsName() {
Principal p = new CallerResolver(workerIdentity(), false, null, Map.of("term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
assertEquals("opus-5.0", p.name(), "whoami must be able to say WHICH lead is asking");
assertEquals("term_a", p.terminal(),
"CB-532: a lead carries the pane it was matched by. Without it ownsSession() can "
+ "never be true for a lead, so it can send to a peer but never answer one");
}
@Test
void aSecondLeadIsRecognisedRatherThanSilentlyDemoted() {
// The regression this feature exists for: with a singular pin, whichever lead was not the
// pin resolved as a worker and was refused every orchestration call.
Principal p = new CallerResolver(workerIdentity(), false, null,
Map.of("term_elsewhere", "gpt-sol-5.6", "term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
assertEquals("opus-5.0", p.name());
}
@Test
void aPaneAbsentFromTheRegistryIsStillAWorker() {
Principal p = new CallerResolver(workerIdentity(), false, null,
Map.of("term_elsewhere", "gpt-sol-5.6"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertEquals("term_a", p.terminal());
assertNull(p.name());
}
@Test
void aRegisteredLeadNeedsNoTokenEvenInTokenMode() {
Principal p = new CallerResolver(workerIdentity(), true, "s3cret", Map.of("term_a", "opus"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role(),
"the pane mapping is as unforgeable as a worker's — it outranks the token path");
assertEquals("opus", p.name());
}
/** The pre-CB-530 spelling must keep working, exactly, including for configs that never migrate. */
@Test
void theLegacySinglePinBehavesAsALeadNamedPrimary() {
Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_a")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
assertEquals("primary", p.name());
}
@Test
void anEmptyRegistryLeavesEveryPaneAWorker() {
Map<String, String> noLeads = null;
assertEquals(Role.WORKER,
new CallerResolver(workerIdentity(), false, null, Map.of())
.resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER,
new CallerResolver(workerIdentity(), false, null, noLeads)
.resolve("127.0.0.1", 42, null).role());
// CB-531: and the same for the live-registry form, whose supplier may also be absent.
assertEquals(Role.WORKER,
CallerResolver.withLeads(workerIdentity(), false, null, null)
.resolve("127.0.0.1", 42, null).role());
}
/** The audit line must distinguish leads once several exist, or a log says nothing useful. */
@Test
void describeNamesTheLeadButStillReadsPrimaryWhenUnnamed() {
assertEquals("leader:opus-5.0", Principal.leader("opus-5.0", "term_a", 1).describe());
assertEquals("primary", Principal.primary(1).describe());
assertEquals("worker:term_a", Principal.worker("term_a", 1).describe());
}
// ── CB-532: a lead is an addressable peer, not only a sender ────────────────────────────────
/**
* The regression this ticket exists for: two leads could both be recognised (CB-530/531) and
* still not converse, because REPLY is gated on ownsSession() and a lead owned nothing.
*/
@Test
void aLeadOwnsItsOwnPaneSoItMayAnswerAPeer() {
Principal lead = new CallerResolver(workerIdentity(), false, null, Map.of("term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertTrue(lead.ownsSession("term_a"));
assertTrue(Authz.permits(lead, Authz.Action.REPLY, "term_a"),
"a lead answering a peer replies for its OWN terminal — the rendezvous the sender "
+ "opened is keyed on exactly that");
assertTrue(Authz.permits(lead, Authz.Action.ASK, "term_a"));
}
@Test
void aLeadStillCannotActAsAnyoneElse() {
Principal lead = new CallerResolver(workerIdentity(), false, null, Map.of("term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertFalse(lead.ownsSession("term_someone_else"));
assertFalse(Authz.permits(lead, Authz.Action.REPLY, "term_someone_else"),
"widening WHO may reply must not widen WHAT they may reply as");
}
/** A primary with no pane — token mode, or off-host — owns nothing and must stay a sender only. */
@Test
void anUnnamedPrimaryWithNoPaneOwnsNothing() {
Principal p = new CallerResolver(nonWorkerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 99, "Bearer s3cret");
assertEquals(Role.PRIMARY, p.role());
assertNull(p.terminal());
assertFalse(p.ownsSession(null), "a null terminal must never match a null session id");
assertFalse(Authz.permits(p, Authz.Action.REPLY, null));
}
@Test
void aLeadKeepsEveryOrchestrationRightItAlreadyHad() {
Principal lead = new CallerResolver(workerIdentity(), false, null, Map.of("term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertTrue(Authz.permits(lead, Authz.Action.SPAWN, null));
assertTrue(Authz.permits(lead, Authz.Action.SEND, "term_worker"));
assertTrue(Authz.permits(lead, Authz.Action.STOP, null));
assertTrue(Authz.permits(lead, Authz.Action.DRAIN, null));
}
/**
* CB-531: the registry is read per resolve, not snapshotted at construction — a lead that
* labels its tab after the daemon booted is recognised without a restart.
*/
@Test
void aLeadRegisteredAfterConstructionIsHonouredWithoutRebuildingTheResolver() {
Map<String, String> live = new java.util.HashMap<>();
CallerResolver r = CallerResolver.withLeads(workerIdentity(), false, null, () -> live);
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
live.put("term_a", "gpt-sol-5.6"); // the scanner sees a newly-labelled tab
Principal p = r.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
assertEquals("gpt-sol-5.6", p.name());
}
/** The map form must stay a snapshot: a caller handing over a map is not offering live state. */
@Test
void theMapFormIsCopiedSoLaterMutationCannotGrantLeadership() {
Map<String, String> mutable = new java.util.HashMap<>();
CallerResolver r = new CallerResolver(workerIdentity(), false, null, mutable);
mutable.put("term_a", "sneaky");
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
}
// ── CB-548: architect slots ─────────────────────────────────────────────────────────────────
@Test
void aBoundArchitectPaneResolvesToArchitectBeforeTheWorkerFallback() {
Principal p = CallerResolver.withLeadsAndArchitects(workerIdentity(), false, null,
Map::of, () -> Map.of("term_a", "lead-designer"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, p.role(),
"a terminal bound to an architect slot is an architect, NOT the generic worker it "
+ "would otherwise resolve to");
assertEquals("lead-designer", p.name(), "whoami must say WHICH slot is asking");
assertEquals("term_a", p.terminal(), "the pane identity is carried so ownsSession works");
}
@Test
void anArchitectNeedsNoTokenEvenInTokenMode() {
Principal p = CallerResolver.withLeadsAndArchitects(workerIdentity(), true, "s3cret",
Map::of, () -> Map.of("term_a", "lead-designer"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, p.role(),
"the pane mapping is as unforgeable as a worker's — it outranks the token path");
}
@Test
void anUnboundPaneStillResolvesAsAWorker() {
Map<String, String> arch = Map.of("term_elsewhere", "reviewer");
Principal p = CallerResolver.withLeadsAndArchitects(workerIdentity(), false, null,
Map::of, () -> arch).resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertNull(p.name());
}
/** CB-548 precedence: lead > architect > worker, so a pane named in BOTH is still a lead. */
@Test
void aLeadWinsOverAnArchitectBindingForTheSamePane() {
Principal p = CallerResolver.withLeadsAndArchitects(workerIdentity(), false, null,
() -> Map.of("term_a", "opus-5.0"), () -> Map.of("term_a", "lead-designer"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role(),
"a pane the config calls a lead must keep resolving as a lead — no behaviour change "
+ "when an architect binding is added to an existing fleet");
assertEquals("opus-5.0", p.name());
}
/** The registry is live, like leads: a binding injected after construction is honoured. */
@Test
void anArchitectBoundAfterConstructionIsHonouredWithoutRebuildingTheResolver() {
Map<String, String> live = new java.util.HashMap<>();
CallerResolver r = CallerResolver.withLeadsAndArchitects(workerIdentity(), false, null,
Map::of, () -> live);
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
live.put("term_a", "lead-designer"); // the later lifecycle binds the slot
assertEquals(Role.ARCHITECT, r.resolve("127.0.0.1", 42, null).role());
assertEquals("lead-designer", r.architects().get("term_a"));
}
@Test
void theArchitectMapFormIsCopiedSoLaterMutationCannotGrantArchitect() {
Map<String, String> mutable = new java.util.LinkedHashMap<>();
CallerResolver r = new CallerResolver(workerIdentity(), false, null, Map.of(), mutable);
mutable.put("term_a", "sneaky");
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
}
@Test
void describeNamesTheArchitectSlot() {
assertEquals("architect:lead-designer",
Principal.architect("lead-designer", "term_a", 1).describe());
}
/** An architect acts only as its own pane — the same ownsSession rule as a worker or lead. */
@Test
void anArchitectOwnsItsOwnPaneAndNoOther() {
Principal arch = CallerResolver.withLeadsAndArchitects(workerIdentity(), false, null,
Map::of, () -> Map.of("term_a", "lead-designer")).resolve("127.0.0.1", 42, null);
assertTrue(arch.ownsSession("term_a"));
assertTrue(Authz.permits(arch, Authz.Action.REPLY, "term_a"));
assertFalse(arch.ownsSession("term_b"));
assertFalse(Authz.permits(arch, Authz.Action.REPLY, "term_b"));
}
@Test
void tokenModeRequiresANonEmptyConfiguredToken() {
ConnectionIdentity id = nonWorkerIdentity();
@@ -1,10 +1,13 @@
package dev.ltms.bridged.config;
import dev.ltms.bridged.auth.ArchitectRegistry;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
@@ -45,6 +48,7 @@ class BridgedConfigTest {
assertEquals(9000, cfg.bind().port());
assertNotNull(cfg.guard(), "guard must default to empty, never null");
assertTrue(cfg.guard().offSubscriptionHosts().isEmpty());
assertFalse(cfg.lifecycle().clearAfterTurn(), "context clearing is opt-in");
}
@Test
@@ -79,6 +83,11 @@ class BridgedConfigTest {
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(Set.of("gx10", "ollama"), cfg.workerProfiles().keySet());
// Order, not just membership: placement breaks an exact-weight tie on definition order, so a
// hash-ordered map here would make equal-weight placement differ from one restart to the next.
assertEquals(java.util.List.of("gx10", "ollama"),
java.util.List.copyOf(cfg.workerProfiles().keySet()),
"workerProfiles must preserve YAML definition order");
assertEquals("gx10", cfg.defaultProfile());
assertEquals("ollama", cfg.workerProfiles().get("ollama").profile(), "profile defaults to its map key");
assertEquals("http://gx10.gw:8000", cfg.workerProfiles().get("gx10").baseUrl());
@@ -91,6 +100,485 @@ class BridgedConfigTest {
assertDoesNotThrow(() -> BridgedConfig.load(f));
}
/**
* CB-530. Unknown keys stay ignored — config must be allowed to run ahead of the code — but they
* must be NAMED at load. A whole block that parses, is dropped, and is never mentioned again is
* indistinguishable from one that works: that is exactly how a hand-written `leaders:` registry
* came to look configured while being inert.
*/
@Test
void unknownTopLevelKeysAreNamedSoADroppedBlockCannotLookLikeAWorkingOne() {
assertEquals(List.of("futureFeature", "leedars"),
BridgedConfig.unknownTopLevelKeys(
"bind:\n port: 8080\nleedars:\n a: b\nfutureFeature: true\n"),
"a typo'd key is the common case and must be reported by name");
}
@Test
void everyKeyThisBuildUnderstandsIsAbsentFromTheUnknownList() {
assertTrue(BridgedConfig.unknownTopLevelKeys("""
bind:
port: 8080
herdrSocket: /tmp/s
workers: {}
defaultWorker: a
guard: {}
worktreeRoot: /tmp
lifecycle: {}
spawnReadyTimeoutMs: 1
spawnReadyPollMs: 1
broker: {}
primary: {}
leaders: {}
architects: {}
leadScan: {}
placement: fixed
auth: {}
""").isEmpty(), "the known-key set must not drift from the record components");
}
@Test
void aMalformedOrEmptyDocumentIsNotReportedAsUnknownKeys() {
assertTrue(BridgedConfig.unknownTopLevelKeys("").isEmpty());
assertTrue(BridgedConfig.unknownTopLevelKeys("just a scalar").isEmpty());
}
// ── CB-531: lead discovery by tab label ─────────────────────────────────────────────────────
@Test
void leadScanIsOffUnlessTheBlockIsPresent(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-scan.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
assertNull(BridgedConfig.load(f).leadScan(),
"turning this on widens who resolves as PRIMARY — upgrading the daemon must not do that");
}
@Test
void leadScanDefaultsItsFieldsWhenTheBlockIsPresentButBare(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bare-scan.yaml");
Files.writeString(f, "bind:\n port: 8080\nleadScan: {}\n");
BridgedConfig.LeadScan scan = BridgedConfig.load(f).leadScan();
assertEquals("lead:", scan.tabPrefix());
assertEquals(10, scan.intervalSeconds());
}
@Test
void leadScanReadsAnExplicitPrefixAndInterval(@TempDir Path dir) throws Exception {
Path f = dir.resolve("scan.yaml");
Files.writeString(f, """
bind:
port: 8080
leadScan:
tabPrefix: "drive:"
intervalSeconds: 30
""");
BridgedConfig.LeadScan scan = BridgedConfig.load(f).leadScan();
assertEquals("drive:", scan.tabPrefix());
assertEquals(30, scan.intervalSeconds());
}
/**
* The hazard the guard exists for: bridged writes worker tab labels and reads lead tab labels.
* Overlap the two and every worker it spawns is read back as a lead.
*/
@Test
void aLeadPrefixThatAWorkerTabLabelAlsoMatchesRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("collide.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
tabLabel: "lead: {profile} #{n}"
leadScan:
tabPrefix: "lead:"
""");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateLeadScan);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
}
@Test
void theDefaultWorkerTabLabelDoesNotCollideWithTheDefaultLeadPrefix(@TempDir Path dir) throws Exception {
Path f = dir.resolve("ok.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
baseUrl: http://gx00.gw:8000
leadScan: {}
""");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateLeadScan());
}
@Test
void theCollisionGuardIsANoOpWhenScanningIsOff(@TempDir Path dir) throws Exception {
Path f = dir.resolve("off.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
tabLabel: "lead: {profile}"
""");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateLeadScan(),
"a label that collides with a convention nobody reads is not a problem");
}
// ── CB-530: the leaders registry ────────────────────────────────────────────────────────────
@Test
void leadersBlockRegistersEveryPaneByName(@TempDir Path dir) throws Exception {
Path f = dir.resolve("leaders.yaml");
Files.writeString(f, """
bind:
port: 8080
leaders:
opus-5.0:
terminal: term_opus
kind: claude
gpt-sol-5.6:
terminal: term_sol
kind: opencode
model: openai/gpt-5.6-terra
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(Set.of("opus-5.0", "gpt-sol-5.6"), cfg.leaders().keySet());
assertEquals("opencode", cfg.leaders().get("gpt-sol-5.6").kind());
assertEquals("openai/gpt-5.6-terra", cfg.leaders().get("gpt-sol-5.6").model());
// The whole point: BOTH panes resolve as leads, so neither is demoted to worker.
assertEquals(Map.of("term_opus", "opus-5.0", "term_sol", "gpt-sol-5.6"),
cfg.leaderTerminals());
}
@Test
void aLegacyPrimaryPinAloneStillRegistersAsALeadNamedPrimary(@TempDir Path dir) throws Exception {
Path f = dir.resolve("legacy-pin.yaml");
Files.writeString(f, "bind:\n port: 8080\nprimary:\n terminal: term_fixed\n");
assertEquals(Map.of("term_fixed", "primary"), BridgedConfig.load(f).leaderTerminals(),
"configs that never migrate must behave exactly as they did before CB-530");
}
@Test
void anExplicitLeadersEntryWinsOverThePinForTheSameTerminal(@TempDir Path dir) throws Exception {
Path f = dir.resolve("both.yaml");
Files.writeString(f, """
bind:
port: 8080
primary:
terminal: term_shared
leaders:
opus-5.0:
terminal: term_shared
""");
assertEquals(Map.of("term_shared", "opus-5.0"), BridgedConfig.load(f).leaderTerminals(),
"the pin is the older spelling of the same fact; the named entry is what was meant");
}
@Test
void bothBlocksTogetherRegisterTheUnionOfTheirTerminals(@TempDir Path dir) throws Exception {
Path f = dir.resolve("union.yaml");
Files.writeString(f, """
bind:
port: 8080
primary:
terminal: term_pinned
leaders:
gpt-sol-5.6:
terminal: term_sol
""");
assertEquals(Map.of("term_pinned", "primary", "term_sol", "gpt-sol-5.6"),
BridgedConfig.load(f).leaderTerminals());
}
@Test
void neitherBlockLeavesNothingRegistered(@TempDir Path dir) throws Exception {
Path f = dir.resolve("none.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
assertTrue(BridgedConfig.load(f).leaderTerminals().isEmpty());
}
/** A lead entry with no terminal identifies nothing — it must not register a null key. */
@Test
void aLeadWithoutATerminalIsNotRegistered(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-terminal.yaml");
Files.writeString(f, """
bind:
port: 8080
leaders:
sketch:
kind: opencode
real:
terminal: term_real
""");
assertEquals(Map.of("term_real", "real"), BridgedConfig.load(f).leaderTerminals());
}
// ── CB-548: the architects registry ────────────────────────────────────────────────────────
@Test
void architectsBlockDeclaresSlotsByNameAndProfileOnly(@TempDir Path dir) throws Exception {
Path f = dir.resolve("architects.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
sonnet:
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
profile: sonnet
reviewer:
profile: sonnet
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(Set.of("lead-designer", "reviewer"), cfg.architects().keySet(),
"slot names are the keys — gateway-local unique by construction");
assertEquals("sonnet", cfg.architects().get("lead-designer").profile(),
"each slot carries its strong-model profile reference");
assertEquals("sonnet", cfg.architects().get("reviewer").profile());
}
@Test
void anArchitectCarriesNoConfigTerminalSoNothingIsRecognisedYet(@TempDir Path dir) throws Exception {
// The corrected CB-548 premise: config declares slots (name + profile) only. A `terminal:`
// key left over from the earlier premise is ignored — an architect is NOT recognised from
// config the way a lead is, so it binds nothing at startup and resolves no architect.
Path f = dir.resolve("arch-stale-terminal.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
sonnet:
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
terminal: term_design
profile: sonnet
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals("sonnet", cfg.architects().get("lead-designer").profile(),
"the profile is still read even when a stray terminal is ignored");
// The registry built from this config owns no bindings: the slot is idle at startup.
ArchitectRegistry r = new ArchitectRegistry(cfg.architects());
assertTrue(r.snapshot().isEmpty());
assertNull(r.slotForTerminal("term_design"),
"a config terminal must not resolve an architect — slots start idle");
}
@Test
void noArchitectsBlockLeavesNothingConfigured(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-arch.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
assertNull(BridgedConfig.load(f).architects(),
"no architects: block ⇒ no architect identity, exactly as before CB-548");
}
@Test
void anArchitectSlotMayResolveToTheSoleProfileWithoutPrivileging(@TempDir Path dir) throws Exception {
// Even a single unqualified worker profile can back an architect slot — the reference is
// by name, not by position, so an explicit name is required.
Path f = dir.resolve("arch-single.yaml");
Files.writeString(f, """
bind:
port: 8080
worker:
profile: ltms-local
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
profile: ltms-local
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertDoesNotThrow(cfg::validateArchitects);
assertEquals("ltms-local", cfg.architects().get("lead-designer").profile());
}
@Test
void anArchitectProfileThatIsNotConfiguredRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("arch-bad-profile.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
profile: sonnet
""");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateArchitects);
assertTrue(e.getMessage().contains("lead-designer"), "the refusal names the slot");
assertTrue(e.getMessage().contains("sonnet"), "the refusal names the offending profile");
}
@Test
void anArchitectSlotMissingAProfileRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("arch-no-profile.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
profile: ""
""");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateArchitects);
assertTrue(e.getMessage().contains("lead-designer"), "the refusal names the slot");
}
@Test
void aValidArchitectRegistryPassesValidation(@TempDir Path dir) throws Exception {
Path f = dir.resolve("arch-ok.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
sonnet:
baseUrl: http://gx10.gw:8000
gx10:
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
profile: sonnet
reviewer:
profile: gx10
""");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateArchitects());
}
@Test
void absentArchitectsBlockPassesValidation(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-arch.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateArchitects());
}
@Test
void duplicateArchitectSlotNamesAreRejectedAtParseTime(@TempDir Path dir) throws Exception {
Path f = dir.resolve("arch-dup.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
sonnet:
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
profile: sonnet
lead-designer:
profile: sonnet
""");
IllegalStateException e =
assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
assertTrue(e.getMessage().contains("lead-designer"),
"the refusal names the duplicated slot, was: " + e.getMessage());
assertTrue(e.getMessage().contains("duplicate architect"),
"the refusal says the slot name is duplicated");
}
@Test
void duplicateKeysOutsideArchitectsAreUnaffected(@TempDir Path dir) throws Exception {
// The duplicate check is scoped to the architects block — a duplicate elsewhere is not this
// guard's concern and must not change parsing of the rest of the config.
Path f = dir.resolve("dup-other.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
sonnet:
baseUrl: http://gx10.gw:8000
sonnet:
baseUrl: http://gx10.gw:8000
""");
// Last-wins for a non-architect duplicate is untouched: only the architects block is walked.
assertEquals(Set.of("sonnet"), BridgedConfig.load(f).workerProfiles().keySet());
}
@Test
void aNestedArchitectsFieldDoesNotSuppressRealDuplicateDetection(@TempDir Path dir) throws Exception {
// A field ALSO named `architects` nested under another block carries its own duplicate and
// sits BEFORE the real top-level block. Only the top-level block is ever inspected: the
// refusal must name the real slot (lead-designer), not the nested one (nested-slot).
Path f = dir.resolve("nested-arch.yaml");
Files.writeString(f, """
bind:
port: 8080
architects:
nested-slot:
k: v
nested-slot:
k: v
architects:
lead-designer:
profile: sonnet
lead-designer:
profile: sonnet
""");
IllegalStateException e =
assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
assertTrue(e.getMessage().contains("lead-designer"),
"the real top-level duplicate must be reported, was: " + e.getMessage());
assertTrue(e.getMessage().contains("duplicate architect"),
"the refusal says the slot name is duplicated");
}
@Test
void nestedDuplicateFieldsInsideASlotAreNotDuplicateSlotNames(@TempDir Path dir) throws Exception {
// A duplicated field nested inside one slot's own value (here inside an ignored `extra:`
// sub-block) is not a duplicate SLOT name — it must not be rejected as one. Only the direct
// child keys of the architects mapping are slot names; whatever is deeper is the slot's
// business and must not masquerade as a duplicate slot.
Path f = dir.resolve("nested-dup-inside-slot.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
sonnet:
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
profile: sonnet
extra:
a: 1
a: 1
""");
BridgedConfig cfg = assertDoesNotThrow(() -> BridgedConfig.load(f));
assertDoesNotThrow(cfg::validateArchitects,
"a nested duplicate inside a slot is not a duplicate slot and must not refuse startup");
assertEquals(Set.of("lead-designer"), cfg.architects().keySet());
assertEquals("sonnet", cfg.architects().get("lead-designer").profile());
}
@Test
void absentBrokerBlockLeavesInboxSoftState(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-broker.yaml");
@@ -333,10 +821,14 @@ class BridgedConfigTest {
parityOverlay: [".mcp.json", ".env"]
gitTokenEnv: GITEA_TOKEN
gitHostEnv: GITEA_HOST
weight: 0.5
maxLoad: 2
placement: weighted
lifecycle:
idleTtlSeconds: 300
contextCap: 10
drainTimeoutSeconds: 5
clearAfterTurn: true
broker:
uri: amqp://guest:guest@127.0.0.1:5672
primary:
@@ -356,13 +848,145 @@ class BridgedConfigTest {
assertEquals(java.util.List.of(".mcp.json", ".env"), w.parityOverlay());
assertTrue(w.hasGitToken(), "gitTokenEnv binds and enables the CB-302 PR grant");
assertEquals("GITEA_HOST", w.gitHostEnv());
assertEquals(0.5f, w.weight(), 0.0001f, "weight binds as a float");
assertEquals(2, w.maxLoad(), "maxLoad binds as an integer");
assertEquals("weighted", cfg.placement(), "placement binds at the top level");
assertEquals(300, cfg.lifecycle().idleTtlSeconds());
assertEquals(10, cfg.lifecycle().contextCap());
assertEquals(5, cfg.lifecycle().drainTimeoutSeconds());
assertTrue(cfg.lifecycle().clearAfterTurn());
assertEquals("amqp://guest:guest@127.0.0.1:5672", cfg.broker().uri());
assertEquals("term_abc123", cfg.primary().terminal());
assertEquals(5, cfg.primary().remindersOrDefault());
assertEquals(15000L, cfg.primary().backoffMsOrDefault());
}
@Test
void placementDefaultsToFixedForExistingConfigs(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-placement.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
baseUrl: http://gx10.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals("fixed", cfg.placement(), "omitted placement must default to fixed");
}
@Test
void workerWeightAndMaxLoadDefaultSanely(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-weight.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
baseUrl: http://gx10.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Worker w = cfg.workerProfiles().get("gx10");
assertEquals(1.0f, w.weight(), 0.0001f, "absent weight defaults to 1.0");
assertNull(w.maxLoad(), "absent maxLoad defaults to unlimited (null)");
}
@Test
void subscriptionFlagBindsAndDefaultsFalse(@TempDir Path dir) throws Exception {
Path f = dir.resolve("subscription.yaml");
Files.writeString(f, """
workers:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
opted:
baseUrl: http://gx10.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertTrue(cfg.workerProfiles().get("sonnet").isSubscription(),
"subscription: true binds as an explicit opt-in");
assertFalse(cfg.workerProfiles().get("opted").isSubscription(),
"a profile without the key stays off-subscription (the default)");
}
// ── CB-542: subscription:true must not smuggle an unguarded endpoint via env: ───────────────
@Test
void aSubscriptionProfileWithAnthropicBaseUrlInEnvIsRejected(@TempDir Path dir) throws Exception {
Path f = dir.resolve("baseUrl.yaml");
Files.writeString(f, """
workers:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
env:
ANTHROPIC_BASE_URL: http://anything-not-on-the-allowlist
""");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class,
cfg::validateSubscriptionProfiles);
assertTrue(e.getMessage().contains("sonnet"), "the refusal names the offending profile");
assertTrue(e.getMessage().contains("ANTHROPIC_BASE_URL"),
"the refusal names the offending key: " + e.getMessage());
}
@Test
void aSubscriptionProfileWithAnthropicAuthTokenInEnvIsRejected(@TempDir Path dir) throws Exception {
Path f = dir.resolve("authToken.yaml");
Files.writeString(f, """
workers:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
env:
ANTHROPIC_AUTH_TOKEN: sk-ant-not-on-any-allowlist
""");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class,
cfg::validateSubscriptionProfiles);
assertTrue(e.getMessage().contains("sonnet"), "the refusal names the offending profile");
assertTrue(e.getMessage().contains("ANTHROPIC_AUTH_TOKEN"),
"the refusal names the offending key: " + e.getMessage());
}
@Test
void aSubscriptionProfileWithACleanEnvPassesValidation(@TempDir Path dir) throws Exception {
Path f = dir.resolve("clean.yaml");
Files.writeString(f, """
workers:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
env:
JAVA_HOME: /opt/jdk
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertDoesNotThrow(cfg::validateSubscriptionProfiles,
"a subscription profile may carry env: — just not the Anthropic binding keys");
}
@Test
void aNonSubscriptionProfileMayCarryAnthropicEnvKeys(@TempDir Path dir) throws Exception {
// The override is only dangerous on the subscription path, where no guard could vet it. A
// plain profile's env: is still overwritten by the launcher's guard-checked value (CB-511).
Path f = dir.resolve("nonsub.yaml");
Files.writeString(f, """
workers:
gx10:
baseUrl: http://gx10.gw:8000
env:
ANTHROPIC_BASE_URL: http://something
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertDoesNotThrow(cfg::validateSubscriptionProfiles,
"only subscription:true profiles are checked — off-subscription ones keep the baseUrl guard");
}
}
@@ -11,10 +11,11 @@ import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the {@code agent.*} south side against a REAL herdr, locking in
* the CB-102 spike findings. It spawns a HARMLESS probe command (never {@code claude},
* so no subscription/token involvement), proves the {@code env} map reaches the process
* environment, exercises status/read, and always tears the pane down.
* Contract test for the worker env seam against a REAL herdr. Under protocol 19 (CB-521) the
* env map is injected at PANE CREATION ({@code tab.create}), not {@code agent.start} — and
* {@code agent.start} now only launches supported agent kinds, so this probes the seed pane's
* SHELL directly (never {@code claude}, so no subscription/token involvement) and always tears
* the throwaway space down.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@@ -26,38 +27,30 @@ class AgentControlContractTest {
}
@Test
void startInjectsEnvThenReadAndClose() throws Exception {
void tabCreateInjectsEnvIntoTheSeedShell() throws Exception {
assumeTrue(!noSocket(), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
AgentControl agents = new AgentControl(herdr);
Agent probe = agents.start(
"__contract__",
List.of("bash", "-c", "printf 'PROBE_BASE=[%s]\\n' \"$ANTHROPIC_BASE_URL\"; sleep 20"),
WorkspaceControl spaces = new WorkspaceControl(herdr);
Workspace space = spaces.ensureWorkspace("__bridged_env_contract__");
Tab.Created tab = spaces.createTab(space.workspaceId(), null,
Map.of("ANTHROPIC_BASE_URL", "http://gx00.gw:8000"));
assertNotNull(probe.terminalId());
assertNotNull(probe.paneId());
try {
// Give the shell a moment to print, then confirm env reached the process.
assertNotNull(tab.rootPaneId(), "tab.create must return the seed pane");
Thread.sleep(1000); // let the seed shell reach its prompt
herdr.call("pane.send_input", Map.of(
"pane_id", tab.rootPaneId(),
"text", "printf 'PROBE_BASE=[%s]\\n' \"$ANTHROPIC_BASE_URL\"",
"keys", List.of("enter")));
Thread.sleep(800);
String visible = agents.read(probe.terminalId(), "visible");
String visible = herdr.call("pane.read",
Map.of("pane_id", tab.rootPaneId(), "source", "visible"))
.path("read").path("text").asText("");
assertTrue(visible.contains("PROBE_BASE=[http://gx00.gw:8000]"),
"env map must reach the process; saw: " + visible);
// Status is queryable; the probe appears in the agent list.
assertNotNull(agents.status(probe.terminalId()));
assertTrue(agents.list().stream()
.anyMatch(a -> probe.terminalId().equals(a.terminalId())),
"spawned probe should appear in agent.list");
"env map must reach the seed shell; saw: " + visible);
} finally {
agents.close(probe.paneId());
spaces.closeTab(tab.tab().tabId());
herdr.call("workspace.close", Map.of("workspace_id", space.workspaceId()));
}
// After close the pane is gone.
assertFalse(agents.list().stream()
.anyMatch(a -> probe.terminalId().equals(a.terminalId())),
"closed probe should no longer be listed");
}
}
}
@@ -7,34 +7,68 @@ import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
/** Unit-level behaviour of {@link AgentControl} over a fake herdr. */
/** Unit-level behaviour of {@link AgentControl} over a fake herdr (protocol 19). */
class AgentControlTest {
/** The {@code text} of every agent.send, in call order. */
/** The {@code text} of every agent.prompt, in call order. */
@SuppressWarnings("unchecked")
private static List<String> sendTexts(FakeHerdr herdr) {
private static List<String> promptTexts(FakeHerdr herdr) {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.toList();
}
@Test
void sendDeliversThePayloadThenAStandaloneSubmitKey() {
void sendDeliversThePayloadAsOnePromptThatSubmitsItself() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_x", "do the thing");
// The Enter must be its own event — appended to the paste it would be swallowed as text.
assertEquals(List.of("do the thing", "\r"), sendTexts(herdr),
"payload paste first, then a separate carriage-return keystroke to submit it");
// agent.prompt pastes AND submits in one call — no separate Enter event to assert.
assertEquals(List.of("do the thing"), promptTexts(herdr),
"exactly one agent.prompt carrying the payload");
}
@Test
void sendPreservesEmbeddedNewlinesAndSubmitsOnlyOnce() {
void sendPreservesEmbeddedNewlinesVerbatim() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_x", "line1\nline2");
assertEquals(List.of("line1\nline2", "\r"), sendTexts(herdr),
"multiline content is delivered verbatim; a single trailing Enter submits it");
assertEquals(List.of("line1\nline2"), promptTexts(herdr),
"multiline content is delivered verbatim in the single prompt");
}
@Test
void submitNudgesWithAStandaloneEnterKeystroke() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).submit("term_x");
FakeHerdr.Call keys = herdr.lastCall("agent.send_keys");
assertEquals(Map.of("target", "term_x", "keys", List.of("enter")), keys.params(),
"the raced-Enter nudge is a raw send_keys, not a second prompt");
}
@Test
@SuppressWarnings("unchecked")
void aTerminalIdTargetIsTranslatedToItsPaneId() {
// Protocol 19 rejects terminal_id as an agent.* target; the fake's agent.list maps
// term_a to pane w2:p7, and the control layer must address herdr by that pane.
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_a", "hello");
Map<String, Object> prompt = (Map<String, Object>) herdr.lastCall("agent.prompt").params();
assertEquals("w2:p7", prompt.get("target"), "terminal target resolved to the agent's pane id");
}
@Test
@SuppressWarnings("unchecked")
void theTerminalToPaneMappingIsCachedAcrossCalls() {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
agents.send("term_a", "one");
agents.send("term_a", "two");
long lists = herdr.calls.stream().filter(c -> c.method().equals("agent.list")).count();
assertEquals(1, lists, "one agent.list resolution serves every later call to the same terminal");
}
}
@@ -8,8 +8,8 @@ import java.util.List;
/**
* Recording fake {@link HerdrClient} for unit/acceptance tests. Returns canned frames
* captured from the real herdr 0.7.0 daemon and records every call so tests can assert
* both behaviour and that guard-blocked paths never reached herdr.
* matching the real herdr 0.8.0 daemon (protocol 19) and records every call so tests can
* assert both behaviour and that guard-blocked paths never reached herdr.
*/
public final class FakeHerdr implements HerdrClient {
@@ -25,11 +25,15 @@ public final class FakeHerdr implements HerdrClient {
private final List<String> extraWorkspaces = new ArrayList<>();
private final List<String> extraAgents = new ArrayList<>();
private int agentNameTakenFor = 0;
private int agentPaneBusyFor = 0;
private int workerTabPaneCount = 1;
private String paneCloseErrorCode = null;
private String agentSendErrorCode = null;
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String readText = "worker transcript tail"; // canned agent.read output
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
private String pinnedStartTerminal;
private String pinnedStartPane;
public FakeHerdr healthy(boolean h) {
this.healthy = h;
@@ -42,6 +46,12 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/** Reject the first {@code n} {@code agent.start} calls with {@code agent_pane_busy}. */
public FakeHerdr agentPaneBusyTimes(int n) {
this.agentPaneBusyFor = n;
return this;
}
/** Make the worker tab (w9:t2) report this many panes in {@code tab.list} (default 1). */
public FakeHerdr withWorkerTabPaneCount(int n) {
this.workerTabPaneCount = n;
@@ -66,12 +76,25 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/** Make {@code agent.send} fail with this herdr error code. */
/** Make delivery ({@code agent.prompt} / {@code agent.send_keys}) fail with this error code. */
public FakeHerdr agentSendFailsWith(String code) {
this.agentSendErrorCode = code;
return this;
}
/**
* Force the next {@code n} {@code agent.start} calls to report this terminal/pane coordinate,
* instead of the fake's usual incrementing {@code term_new_n}/{@code w9:pRoot_n}. Lets a test make
* two spawns report the <em>same</em> herdr pane, to prove the host-unique id (CB-519) never
* collides on that coordinate.
*/
public FakeHerdr pinNextStarts(int n, String terminalId, String paneId) {
this.pinnedStarts = n;
this.pinnedStartTerminal = terminalId;
this.pinnedStartPane = paneId;
return this;
}
/**
* Seed a named agent into {@code agent.list} (e.g. an orphaned worker for CB-117 reaper tests).
@@ -110,7 +133,7 @@ public final class FakeHerdr implements HerdrClient {
try {
return switch (method) {
case "ping" -> mapper.readTree(
"{\"type\":\"pong\",\"version\":\"0.7.0\",\"protocol\":14}");
"{\"type\":\"pong\",\"version\":\"0.8.0\",\"protocol\":19}");
case "workspace.list" -> mapper.readTree(("""
{"type":"workspace_list","workspaces":[
{"workspace_id":"w1","label":"dev-mgnl","focused":true,"pane_count":7,"agent_status":"unknown"},
@@ -122,9 +145,19 @@ public final class FakeHerdr implements HerdrClient {
"agent_session":{"kind":"id","value":"sess-1111"},
"workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}%s]}""")
.formatted(extraAgents.isEmpty() ? "" : "," + String.join(",", extraAgents)));
case "agent.send" -> {
case "agent.prompt" -> {
if (agentSendErrorCode != null) {
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.send failed",
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.prompt failed",
agentSendErrorCode, null);
}
yield mapper.readTree(("""
{"type":"agent_prompted","agent":{"terminal_id":"term_a","agent":"claude",
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentStatus));
}
case "agent.send_keys" -> {
if (agentSendErrorCode != null) {
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.send_keys failed",
agentSendErrorCode, null);
}
yield mapper.readTree("{\"type\":\"ok\"}");
@@ -136,29 +169,62 @@ public final class FakeHerdr implements HerdrClient {
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
case "agent.start" -> {
// Protocol 19: kind and pane_id are required — reject like the real daemon.
java.util.Map<?, ?> p = params instanceof java.util.Map<?, ?> m ? m : java.util.Map.of();
for (String required : new String[]{"kind", "pane_id"}) {
if (p.get(required) == null) {
throw new HerdrException(
"herdr error [invalid_request]: invalid request: missing field `"
+ required + "`", "invalid_request", null);
}
}
long starts = calls.stream().filter(c -> c.method().equals("agent.start")).count();
if (starts <= agentNameTakenFor) {
if (starts <= agentPaneBusyFor) {
throw new HerdrException(
"herdr error [agent_pane_busy]: agent target pane is not an available shell",
"agent_pane_busy", null);
}
long busyAdjusted = starts - agentPaneBusyFor;
if (busyAdjusted <= agentNameTakenFor) {
throw new HerdrException(
"herdr error [agent_name_taken]: agent name already used",
"agent_name_taken", null);
}
long n = starts - agentNameTakenFor;
long n = busyAdjusted - agentNameTakenFor;
boolean pinned = pinnedStarts > 0;
if (pinned) {
pinnedStarts--;
}
// Protocol 19: the agent starts INTO the requested pane, so its pane_id normally
// echoes the param. A pin overrides both coordinates, which is the only way to
// make two spawns report one pane — what CB-519's collision test needs.
String terminal = pinned ? pinnedStartTerminal : ("term_new_" + n);
Object pane = pinned ? pinnedStartPane : p.get("pane_id");
yield mapper.readTree(("""
{"type":"agent_started","agent":{
"terminal_id":"term_new_%d","name":"claude","agent_status":"unknown",
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"w9:pW_%d"}}""")
.formatted(n, n));
"terminal_id":"%s","name":"claude","agent_status":"unknown",
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"%s"}}""")
.formatted(terminal, pane));
}
case "pane.split" -> mapper.readTree("""
{"type":"pane_info","pane":{"pane_id":"w1:pSplit","workspace_id":"w1",
"tab_id":"w1:t1"}}""");
case "workspace.create" -> mapper.readTree("""
{"type":"workspace_created",
"workspace":{"workspace_id":"w9","label":"bridged-workers","focused":false,
"pane_count":1,"tab_count":1,"active_tab_id":"w9:t1","agent_status":"unknown"},
"tab":{"tab_id":"w9:t1","workspace_id":"w9","label":"1","pane_count":1},
"root_pane":{"pane_id":"w9:p1","workspace_id":"w9","tab_id":"w9:t1"}}""");
case "tab.create" -> mapper.readTree("""
case "tab.create" -> {
// Each tab gets its own seed pane — under protocol 19 that pane becomes the
// worker pane, so distinct spawns must yield distinct pane ids.
long tabs = calls.stream().filter(c -> c.method().equals("tab.create")).count();
yield mapper.readTree(("""
{"type":"tab_created",
"tab":{"tab_id":"w9:t2","workspace_id":"w9","label":"2","pane_count":1},
"root_pane":{"pane_id":"w9:pRoot","workspace_id":"w9","tab_id":"w9:t2"}}""");
"root_pane":{"pane_id":"w9:pRoot_%d","workspace_id":"w9","tab_id":"w9:t2"}}""")
.formatted(tabs));
}
case "tab.rename" -> mapper.readTree("""
{"type":"tab_info","tab":{"tab_id":"w9:t2","workspace_id":"w9",
"label":"worker: ltms-local","pane_count":1}}""");
@@ -0,0 +1,262 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-531. A lead is never spawned, so the daemon has to <em>find</em> it: these assert that an
* operator-labelled tab is what makes a pane a lead, and — just as importantly — what does not.
*/
class LeadTabScannerTest {
private static final ObjectMapper MAPPER = new ObjectMapper();
private static final long TTL = TimeUnit.SECONDS.toNanos(10);
/**
* A herdr whose workspace/tab/pane topology is declared per test. Counts calls so the caching
* contract can be asserted, and can be made to fail on demand.
*/
private static final class TopologyHerdr implements HerdrClient {
/** workspace_id → label. */
final Map<String, String> workspaces = new LinkedHashMap<>();
/** tab_id → [workspace_id, label]. */
final Map<String, String[]> tabs = new LinkedHashMap<>();
/** pane_id → [tab_id, terminal_id]. */
final Map<String, String[]> panes = new LinkedHashMap<>();
int calls;
boolean failing;
TopologyHerdr workspace(String id, String label) {
workspaces.put(id, label);
return this;
}
TopologyHerdr tab(String tabId, String workspaceId, String label) {
tabs.put(tabId, new String[]{workspaceId, label});
return this;
}
TopologyHerdr pane(String paneId, String tabId, String terminalId) {
panes.put(paneId, new String[]{tabId, terminalId});
return this;
}
@Override
public JsonNode call(String method, Object params) {
calls++;
if (failing) {
throw new HerdrException("socket closed");
}
List<String> items = new ArrayList<>();
switch (method) {
case "workspace.list" -> {
workspaces.forEach((id, label) -> items.add(
"{\"workspace_id\":\"%s\",\"label\":\"%s\"}".formatted(id, label)));
return read("{\"workspaces\":[%s]}".formatted(String.join(",", items)));
}
case "tab.list" -> {
String ws = String.valueOf(((Map<?, ?>) params).get("workspace_id"));
tabs.forEach((id, t) -> {
if (ws.equals(t[0])) {
items.add(("{\"tab_id\":\"%s\",\"workspace_id\":\"%s\",\"label\":%s,"
+ "\"pane_count\":1}").formatted(id, t[0],
t[1] == null ? "null" : "\"" + t[1] + "\""));
}
});
return read("{\"tabs\":[%s]}".formatted(String.join(",", items)));
}
case "pane.list" -> {
panes.forEach((id, p) -> items.add(
"{\"pane_id\":\"%s\",\"tab_id\":\"%s\",\"terminal_id\":\"%s\"}"
.formatted(id, p[0], p[1])));
return read("{\"panes\":[%s]}".formatted(String.join(",", items)));
}
default -> throw new AssertionError("unexpected herdr call: " + method);
}
}
private static JsonNode read(String json) {
try {
return MAPPER.readTree(json);
} catch (Exception e) {
throw new AssertionError(e);
}
}
@Override
public void close() {
}
}
/**
* The usual shape: one user space with lead tabs, one worker space bridged owns.
*
* <p>Not closed: {@code close()} is a no-op on this fake, and every test needs the handle after
* the scanner is built (to mutate the topology or read {@code calls}).
*/
@SuppressWarnings("resource")
private TopologyHerdr twoLeads() {
return new TopologyHerdr()
.workspace("w1", "main")
.workspace("w9", "bridged-workers")
.tab("w1:t1", "w1", "lead: opus-5.0")
.tab("w1:t2", "w1", "lead: gpt-sol-5.6")
.tab("w1:t3", "w1", "notes")
.tab("w9:t1", "w9", "worker: gx10 #1")
.pane("w1:p1", "w1:t1", "term_opus")
.pane("w1:p2", "w1:t2", "term_gpt")
.pane("w1:p3", "w1:t3", "term_notes")
.pane("w9:p1", "w9:t1", "term_worker");
}
private LeadTabScanner scanner(TopologyHerdr herdr, Map<String, String> configured,
AtomicLong clock) {
return new LeadTabScanner(herdr, "lead:", Set.of("bridged-workers"), configured, TTL,
clock::get);
}
@Test
void everyLabelledTabBecomesALeadNamedByItsLabel() {
Map<String, String> leads = scanner(twoLeads(), Map.of(), new AtomicLong()).get();
assertEquals(Map.of("term_opus", "opus-5.0", "term_gpt", "gpt-sol-5.6"), leads,
"two leads discovered from labels alone — no terminal_id was ever configured");
}
@Test
void anUnlabelledTabContributesNothing() {
assertFalse(scanner(twoLeads(), Map.of(), new AtomicLong()).get().containsKey("term_notes"));
}
/**
* The guard that matters: bridged labels its own worker tabs, so if a worker space were scanned
* a naming accident would promote the fleet. The exclusion is by workspace, not by hoping the
* worker template never collides.
*/
@Test
void aTabInAWorkerSpaceIsNeverALeadEvenWhenItsLabelMatches() {
TopologyHerdr herdr = twoLeads().tab("w9:t2", "w9", "lead: impostor")
.pane("w9:p2", "w9:t2", "term_impostor");
assertFalse(scanner(herdr, Map.of(), new AtomicLong()).get().containsKey("term_impostor"));
}
@Test
void aBarePrefixNamesNobodyAndIsRejected() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
.tab("w1:t1", "w1", "lead:").pane("w1:p1", "w1:t1", "term_a");
assertEquals(Map.of(), scanner(herdr, Map.of(), new AtomicLong()).get(),
"a lead with no name would resolve as PRIMARY with nothing to attribute it to");
}
@Test
void thePrefixMatchesCaseInsensitivelyAndTheNameIsTrimmed() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
.tab("w1:t1", "w1", " LEAD: opus-5.0 ").pane("w1:p1", "w1:t1", "term_a");
assertEquals(Map.of("term_a", "opus-5.0"), scanner(herdr, Map.of(), new AtomicLong()).get());
}
@Test
void everyPaneInALeadTabResolvesAsThatLead() {
// A human may split their own lead tab. Both panes are theirs, so both are that lead —
// nothing bridged placed can land here (see the worker-space test above).
TopologyHerdr herdr = twoLeads().pane("w1:p1b", "w1:t1", "term_opus_split");
assertEquals("opus-5.0", scanner(herdr, Map.of(), new AtomicLong()).get().get("term_opus_split"));
}
@Test
void anExplicitlyConfiguredLeadIsMergedInAndOutranksALabel() {
Map<String, String> configured = Map.of("term_opus", "pinned-name", "term_extra", "from-config");
Map<String, String> leads = scanner(twoLeads(), configured, new AtomicLong()).get();
assertEquals("pinned-name", leads.get("term_opus"), "an explicit pin is the operator's last word");
assertEquals("from-config", leads.get("term_extra"), "a configured lead needs no tab at all");
assertEquals("gpt-sol-5.6", leads.get("term_gpt"));
}
// ── caching ─────────────────────────────────────────────────────────────────────────────────
@Test
void aSecondLookupWithinTheTtlDoesNotTouchHerdr() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
s.get();
int afterFirst = herdr.calls;
clock.addAndGet(TTL - 1);
s.get();
assertEquals(afterFirst, herdr.calls,
"resolve() runs on every request — an un-cached scan would put herdr on that path");
}
@Test
void aTabLabelledAfterStartupIsPickedUpOnceTheTtlExpires() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
assertFalse(s.get().containsKey("term_notes"));
herdr.tab("w1:t3", "w1", "lead: late-arrival"); // the operator renames their tab
clock.addAndGet(TTL);
assertEquals("late-arrival", s.get().get("term_notes"),
"the whole point over `leaders:`: no config edit, no restart");
}
@Test
void aFailedScanKeepsTheLeadsAlreadyKnownRatherThanDemotingThem() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
Map<String, String> before = s.get();
herdr.failing = true;
clock.addAndGet(TTL);
assertEquals(before, s.get(),
"a herdr hiccup must not silently demote a live lead to a worker mid-session");
}
@Test
void aFailedFirstScanStillHonoursTheConfiguredLeads() {
TopologyHerdr herdr = twoLeads();
herdr.failing = true;
Map<String, String> leads = scanner(herdr, Map.of("term_x", "opus-5.0"), new AtomicLong()).get();
assertEquals(Map.of("term_x", "opus-5.0"), leads,
"config-named leads must not depend on herdr answering at all");
}
@Test
void aDownHerdrIsRetriedOncePerTtlNotOncePerRequest() {
TopologyHerdr herdr = twoLeads();
herdr.failing = true;
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
s.get();
int afterFirst = herdr.calls;
s.get();
s.get();
assertEquals(afterFirst, herdr.calls, "the failure path must be rate-limited too");
}
}
@@ -5,16 +5,15 @@ import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import java.nio.file.Files;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the herdr half of connection-based identity against a REAL herdr: spawn a
* harmless probe, read its actual {@code shell_pid} from {@code pane.process_info}, and confirm
* {@link PaneLocator} resolves that PID back to the probe's own {@code terminal_id}.
* Contract test for {@link PaneLocator} against a REAL herdr: the PID→pane mapping that
* connection identity rests on. Uses a throwaway tab's seed shell as the probe process
* (protocol 19 removed arbitrary-command agents), and always tears the space down.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@@ -25,18 +24,24 @@ class PaneLocatorContractTest {
void resolvesTheTerminalOwningARealProcessPid() throws Exception {
assumeTrue(Files.exists(UnixSocketHerdrClient.defaultSocketPath()), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
AgentControl agents = new AgentControl(herdr);
Agent probe = agents.start("__pidprobe__", List.of("bash", "-c", "sleep 20"), Map.of());
WorkspaceControl spaces = new WorkspaceControl(herdr);
Workspace space = spaces.ensureWorkspace("__bridged_pid_contract__");
Tab.Created tab = spaces.createTab(space.workspaceId(), null, Map.of());
try {
JsonNode info = herdr.call("pane.process_info", Map.of("pane_id", probe.paneId()))
JsonNode info = herdr.call("pane.process_info", Map.of("pane_id", tab.rootPaneId()))
.path("process_info");
long shellPid = info.path("shell_pid").asLong(-1);
assertTrue(shellPid > 0, "probe pane should report a shell pid");
assertTrue(shellPid > 0, "seed pane should report a shell pid");
assertEquals(probe.terminalId(), new PaneLocator(herdr).terminalForPid(shellPid),
String terminalId = herdr.call("pane.get", Map.of("pane_id", tab.rootPaneId()))
.path("pane").path("terminal_id").asText(null);
assertNotNull(terminalId, "seed pane should carry a terminal_id");
assertEquals(terminalId, new PaneLocator(herdr).terminalForPid(shellPid),
"a real PID must resolve back to its own pane's terminal_id");
} finally {
agents.close(probe.paneId());
spaces.closeTab(tab.tab().tabId());
herdr.call("workspace.close", Map.of("workspace_id", space.workspaceId()));
}
}
}
@@ -5,7 +5,6 @@ import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import java.nio.file.Files;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
@@ -13,9 +12,9 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the placement layer ({@code workspace.*}/{@code tab.*}) against a
* REAL herdr, locking in the "one clean tab per worker" recipe: find-or-create a worker
* space, give the worker its own tab, drop herdr's seed shell so the tab holds only the
* worker, and tear it all down. Uses a HARMLESS probe (never {@code claude}) in a
* REAL herdr, locking in the "one clean tab per worker" recipe under protocol 19: find-or-create
* a worker space, give the worker its own tab, and the SEED pane is where the worker starts —
* the tab holds exactly that one pane from creation. Uses no agent (never {@code claude}) in a
* throwaway space that is fully removed at the end.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
@@ -41,7 +40,6 @@ class WorkspacePlacementContractTest {
void workerGetsOwnCleanTabAndTearsDownCompletely() throws Exception {
assumeTrue(!noSocket(), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
Workspace space = spaces.ensureWorkspace(LABEL);
@@ -49,36 +47,27 @@ class WorkspacePlacementContractTest {
// Idempotent: a second ensure finds the same space, never creates a duplicate.
assertEquals(space.workspaceId(), spaces.ensureWorkspace(LABEL).workspaceId());
Tab.Created tab = spaces.createTab(space.workspaceId());
Agent worker = agents.start(
"__contract__",
List.of("bash", "-c", "sleep 20"),
Map.of(),
tab.tab().tabId());
Tab.Created tab = spaces.createTab(space.workspaceId(), null, Map.of());
try {
// The worker landed in its dedicated tab in the worker space.
assertEquals(tab.tab().tabId(), worker.tabId());
assertEquals(space.workspaceId(), worker.workspaceId());
assertNotNull(tab.rootPaneId(), "tab.create must return the seed pane");
// Drop the seed shell; the tab now holds exactly the worker pane.
agents.close(tab.rootPaneId());
// Protocol 19: the seed pane IS the worker pane — the tab holds exactly it.
spaces.renameTab(tab.tab().tabId(), "worker: contract");
assertEquals(1, paneCount(herdr, space.workspaceId(), tab.tab().tabId()),
"worker tab must hold only the worker pane after the seed shell is dropped");
"worker tab must hold exactly the seed/worker pane");
// Teardown resolves the tab from the pane, and sees it holds exactly one pane.
WorkspaceControl.PaneLocation loc = spaces.locatePane(worker.paneId());
WorkspaceControl.PaneLocation loc = spaces.locatePane(tab.rootPaneId());
assertNotNull(loc);
assertEquals(tab.tab().tabId(), loc.tabId());
assertEquals(1, loc.tabPaneCount(), "worker is the tab's sole occupant");
assertEquals(1, loc.tabPaneCount(), "worker pane is the tab's sole occupant");
} finally {
agents.close(worker.paneId());
spaces.closeTab(tab.tab().tabId());
}
// Tolerant teardown: closing an already-gone tab / reading a gone pane is a no-op.
spaces.closeTab(tab.tab().tabId());
assertNull(spaces.locatePane(worker.paneId()), "closed worker pane must be gone");
assertNull(spaces.locatePane(tab.rootPaneId()), "closed worker pane must be gone");
// Remove the throwaway space entirely so the test leaves no residue.
herdr.call("workspace.close", Map.of("workspace_id", space.workspaceId()));
@@ -156,6 +156,21 @@ class CompletionResolverTest {
assertEquals("No, 391 = 17 × 23.", waiter.getNow(null).text());
}
@Test
void resolvesSynchronouslyBeforePostTurnContextClearing() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ previous answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
resolver.captureBaseline("term_a");
herdr.readText("⏺ answer that /clear would erase\n❯ ");
resolver.resolveBeforePostAction("term_a");
assertTrue(waiter.isDone(), "the answer is captured before the adapter sends /clear");
assertEquals("answer that /clear would erase", waiter.getNow(null).text());
}
@Test
void suppressesAnUnchangedCompletionEvenWhenTheBlockExceedsTheScrapeCap() {
// The fan-out issue-hunt finding: captureBaseline once stored the RAW (unclipped) assistant
@@ -271,10 +286,12 @@ class CompletionResolverTest {
// The turn as the injector captured it at delivery (waiter + pre-turn baseline).
var turnN = new CompletionResolver.InFlight(waiterN, "an earlier answer");
// Turn N is resolved by the worker's explicit reply.
// Turn N is resolved by the worker's explicit reply, and its send deregisters the waiter.
assertTrue(rendezvous.resolve("term_a", "N replied"));
rendezvous.close("term_a", waiterN); // the sender's finally, before the next turn opens
// Turn N+1's send opens its own waiter on the same session (replacing the registered one).
// Turn N+1's send opens its own waiter on the same session (CB-548: open fails if the
// previous waiter is still registered, so a clean turn deregisters it first as above).
var waiterN1 = rendezvous.open("term_a");
resolver.resolve("term_a", turnN); // turn N's completion fallback finally fires
@@ -28,17 +28,15 @@ class InjectorTest {
private final Injector injector = new Injector(new AgentControl(herdr));
/**
* The logical messages delivered, in order. AgentControl.send emits each delivery as two
* agent.send calls — the payload, then a standalone Enter keystroke ({@code "\r"}) to submit
* it; these tests assert delivery ordering/gating, not the submit event, so drop the bare
* carriage returns.
* The logical messages delivered, in order. Under protocol 19 each delivery is one
* {@code agent.prompt} carrying the payload (it submits itself); the Enter nudge is a
* separate {@code agent.send_keys} and never appears here.
*/
@SuppressWarnings("unchecked")
private List<String> sent() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.filter(t -> !t.equals("\r"))
.toList();
}
@@ -66,11 +64,9 @@ class InjectorTest {
assertEquals(List.of("task"), sent(), "delivers once the worker is available");
}
@SuppressWarnings("unchecked")
private long enterKeystrokes() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> "\r".equals(((Map<String, Object>) c.params()).get("text")))
.filter(c -> c.method().equals("agent.send_keys"))
.count();
}
@@ -200,6 +196,47 @@ class InjectorTest {
assertEquals(List.of(T), completed, "a confirmed working→idle fires exactly one completion");
}
@Test
void postTurnResetSettlesBeforeNextDelegationWithoutBecomingATurn() {
AgentControl agents = new AgentControl(herdr);
class ResetListener implements TurnListener {
int completed;
@Override
public void onTurnComplete(String target) {
completed++;
}
@Override
public boolean hasPostTurnAction(String target) {
return true;
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
completed++;
agents.send(target, "/clear"); // direct housekeeping, never Injector.enqueue
return true;
}
}
ResetListener listener = new ResetListener();
Injector inj = new Injector(agents, listener);
inj.enqueue(T, "first");
inj.enqueue(T, "second");
inj.onStatus(T, AgentStatus.IDLE); // first delegation
inj.onStatus(T, AgentStatus.WORKING);
inj.onStatus(T, AgentStatus.IDLE); // first complete; reset dispatched
assertEquals(List.of("first", "/clear"), sent(),
"same-tick completion must not let the queued delegation overtake reset");
inj.onStatus(T, AgentStatus.WORKING); // reset picked up, but this is not a bridge turn
inj.onStatus(T, AgentStatus.IDLE); // reset settled; second may now deliver
assertEquals(List.of("first", "/clear", "second"), sent());
assertEquals(1, listener.completed,
"reset settlement must not recursively emit another turn completion");
}
@Test
void doesNotSynthesizeCompletionFromAnUnconfirmedTurn() {
List<String> completed = new ArrayList<>();
@@ -374,13 +411,12 @@ class InjectorTest {
poller.stop();
}
assertEquals(List.of("via-poller"), idle.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> {
@SuppressWarnings("unchecked")
Map<String, Object> p = (Map<String, Object>) c.params();
return p.get("text").toString();
})
.filter(t -> !t.equals("\r")) // drop the standalone submit keystroke
.toList());
}
@@ -77,6 +77,7 @@ class BridgeMcpAuthzTest {
private static final Principal PRIMARY = Principal.primary(100);
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal ANON = Principal.anonymous();
private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400);
// --- the table, enforced on THIS path too ---------------------------------------------------
@@ -119,6 +120,33 @@ class BridgeMcpAuthzTest {
assertNotNull(m.denyFor(PRIMARY, Authz.Action.ASK, "term_a"));
}
// --- CB-548: the architect on this path ------------------------------------------------
@Test
void anArchitectMaySendAndReadButNotOrchestrateOverMcp() {
BridgeMcp m = mcp(true);
assertNull(m.denyFor(ARCH_DESIGN, Authz.Action.SEND, "term_a"),
"delegating a turn is the architect's job");
assertNull(m.denyFor(ARCH_DESIGN, Authz.Action.READ, null));
for (Authz.Action a : new Authz.Action[]{Authz.Action.SPAWN, Authz.Action.STOP,
Authz.Action.DRAIN}) {
McpSchema.CallToolResult denied = m.denyFor(ARCH_DESIGN, a, null);
assertNotNull(denied, a + " must be refused to an architect");
assertTrue(denied.isError(), "a refusal is returned as an MCP tool error");
}
}
@Test
void anArchitectMayReplyAndAskOnlyAsItsOwnPaneOverMcp() {
BridgeMcp m = mcp(true);
assertNull(m.denyFor(ARCH_DESIGN, Authz.Action.REPLY, "term_design"));
assertNull(m.denyFor(ARCH_DESIGN, Authz.Action.ASK, "term_design"));
assertNotNull(m.denyFor(ARCH_DESIGN, Authz.Action.REPLY, "term_a"),
"architect 'lead-designer' must not reply on worker term_a's session");
}
@Test
void anonymousIsRefusedEverythingAndCountedAsUnauthenticated() {
BridgeMcp m = mcp(true);
@@ -156,6 +184,11 @@ class BridgeMcpAuthzTest {
assertEquals("term_a", BridgeMcp.principalFrom("WORKER", "term_a", 7).terminal());
assertEquals(Role.PRIMARY, BridgeMcp.principalFrom("PRIMARY", null, 7).role());
assertEquals(Role.ANONYMOUS, BridgeMcp.principalFrom("ANONYMOUS", null, -1).role());
// CB-548: an architect round-trips through the same stash, carrying its slot name.
Principal arch = BridgeMcp.principalFrom("ARCHITECT", "term_design", 7, "lead-designer");
assertEquals(Role.ARCHITECT, arch.role());
assertEquals("lead-designer", arch.name());
assertEquals("term_design", arch.terminal());
}
@Test
@@ -15,6 +15,8 @@ import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.modelcontextprotocol.spec.McpSchema;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.Map;
@@ -31,10 +33,19 @@ import static org.junit.jupiter.api.Assertions.*;
*/
class BridgeMcpTest {
private static final String T = "term_a";
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final MessageService messages = new MessageService(agents, new Injector(agents), rendezvous);
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
private final MessageService messages = new MessageService(agents, new Injector(agents), rendezvous, inbox);
@BeforeEach
void setUp() {
// CB-520: the inbox only peeks/acks targets it owns.
inbox.own(T);
}
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
@@ -216,12 +227,15 @@ class BridgeMcpTest {
@Test
void spawnReturnsTheNewWorkersSessionAndPane() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null);
SessionManager sm = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
McpSchema.CallToolResult res = BridgeMcp.spawn(sm, null);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"sessionId\":\"term_new_1\""), out);
assertTrue(out.contains("\"paneId\":\"w9:pW_1\""), out);
// CB-519: the "paneId" wire field now carries the host-unique opaque id, not the herdr pane.
WorkerSession s = sm.roster().getFirst();
assertTrue(out.contains("\"paneId\":\"" + s.paneId() + "\""), out);
assertNotEquals("w9:pRoot_1", s.paneId(), "the id is decoupled from the herdr pane coordinate");
assertTrue(out.contains("\"status\":\"spawning\""), out);
}
@@ -250,9 +264,10 @@ class BridgeMcpTest {
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "/req/dir", null, null, null);
assertNotEquals(Boolean.TRUE, res.isError());
// Protocol 19: the requested cwd roots the worker's pane at creation (tab.create).
@SuppressWarnings("unchecked")
Map<String, Object> start = (Map<String, Object>) h.lastCall("agent.start").params();
assertEquals("/req/dir", start.get("cwd"));
Map<String, Object> create = (Map<String, Object>) h.lastCall("tab.create").params();
assertEquals("/req/dir", create.get("cwd"));
}
@Test
@@ -273,7 +288,8 @@ class BridgeMcpTest {
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-304", null));
McpSchema.CallToolResult res = BridgeMcp.listWorkers(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions);
McpSchema.CallToolResult res = BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), "");
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
@@ -287,6 +303,58 @@ class BridgeMcpTest {
assertTrue(out.contains("\"liveStatus\":\"unknown\""), out);
}
@Test
void listReportsLeadsAndFlagsTheCallersOwnRow() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
McpSchema.CallToolResult res = BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions,
Map.of("term_me", "opus-5.0", "term_peer", "gpt-sol-5.6"), "term_me");
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"name\":\"opus-5.0\""), out);
assertTrue(out.contains("\"name\":\"gpt-sol-5.6\""), out);
assertTrue(out.contains("\"sessionId\":\"term_peer\""), out);
// The caller's own row is flagged, and only the caller's — a peer must be distinguishable
// from self without a second bridge_whoami call.
assertEquals(1, out.split("\"self\":true", -1).length - 1, out);
assertTrue(out.indexOf("term_me") < out.indexOf("\"self\":true"), out);
}
@Test
void listReportsBothHalvesEvenWhenEmpty() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
McpSchema.CallToolResult res = BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), "");
// An absent "leads" key is what made an empty worker roster read as "no peers" (CB-535).
String out = textOf(res);
assertTrue(out.contains("\"leads\":[]"), out);
assertTrue(out.contains("\"workers\":[]"), out);
}
@Test
void listReportsALeadHerdrCannotSeeAsUnknown() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
McpSchema.CallToolResult res = BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions,
Map.of("term_ghost", "gone-away"), "term_me");
// Reported, not hidden: an unreachable peer is exactly what a would-be sender needs to see.
String out = textOf(res);
assertTrue(out.contains("\"name\":\"gone-away\""), out);
assertTrue(out.contains("\"status\":\"unknown\""), out);
}
@Test
void stopTearsDownAWorkerByPane() {
FakeHerdr h = new FakeHerdr();
@@ -398,4 +466,55 @@ class BridgeMcpTest {
assertTrue(out.contains("\"sessionId\":\"term_orphan\""), out);
assertFalse(out.contains("profile"), out); // nothing invented for a session we don't track
}
/**
* CB-548: an architect reports its role and which gateway-local slot its pane is bound to —
* the same shape as a lead, under the architect key, so it can tell a peer where to reach it.
*/
@Test
void whoamiReportsAnArchitectWithItsSlotAndPane() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.whoami(
Principal.architect("lead-designer", "term_design", 400),
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"architect\""), out);
assertTrue(out.contains("\"architect\":\"lead-designer\""), out);
assertTrue(out.contains("\"sessionId\":\"term_design\""), out);
}
/**
* CB-548: an architect SEND delegates as its own pane (recording the per-target delegation) but
* must NEVER become the legacy singleton "primary" fallback — the per-target map does not cure
* the singleton, so an architect left there would draw no-delegation inbox nudges meant for a
* primary. Only PRIMARY callers (the unnamed primary and named leads alike) may claim it, and
* the decision keys on the resolved role, not name/kind sniffing.
*/
@Test
void architectSendDoesNotClaimThePrimarySingletonButALeadSendStillCan() {
// Architect SEND: does not change the legacy primary fallback.
PrimaryRegistry reg = new PrimaryRegistry(null);
BridgeMcp.recordPrimarySingleton(reg, "term_design", Principal.architect("design", "term_design", 400));
assertTrue(reg.primaryTerminal().isEmpty(),
"an architect must never become the legacy primary fallback");
// Lead SEND (a named PRIMARY) still claims it — preserved from CB-530/CB-532.
PrimaryRegistry leadReg = new PrimaryRegistry(null);
BridgeMcp.recordPrimarySingleton(leadReg, "term_lead_opus", Principal.leader("opus", "term_lead_opus", 100));
assertEquals("term_lead_opus", leadReg.primaryTerminal().orElseThrow(),
"a named lead is a primary and may claim the fallback");
// Unnamed primary likewise.
PrimaryRegistry primaryReg = new PrimaryRegistry(null);
BridgeMcp.recordPrimarySingleton(primaryReg, "term_p", Principal.primary(50));
assertEquals("term_p", primaryReg.primaryTerminal().orElseThrow(),
"an unnamed primary may claim the fallback");
// A null caller (legacy/no-auth path) records nothing.
PrimaryRegistry legacy = new PrimaryRegistry(null);
BridgeMcp.recordPrimarySingleton(legacy, "term_x", null);
assertTrue(legacy.primaryTerminal().isEmpty(), "no caller means nothing is recorded");
}
}
@@ -105,4 +105,68 @@ class PrimaryRegistryTest {
reg.record("term_found");
assertEquals("term_found", reg.primaryTerminal().get());
}
// ── CB-532: nudges follow the delegating lead, not "the primary" ────────────────────────────
/**
* The bug that made `primary.terminal` unretirable: with one slot, whichever lead called
* bridge_send first captured every nudge — including nudges for the other lead's delegations.
*/
@Test
void aNudgeGoesToTheLeadThatDelegatedToThatWorker() {
var reg = new PrimaryRegistry(null);
reg.recordDelegation("term_worker_a", "term_lead_opus");
reg.recordDelegation("term_worker_b", "term_lead_sol");
assertEquals("term_lead_opus", reg.nudgeTargetFor("term_worker_a").orElseThrow());
assertEquals("term_lead_sol", reg.nudgeTargetFor("term_worker_b").orElseThrow());
}
@Test
void theMostRecentDelegatorWinsWhenAWorkerChangesHands() {
var reg = new PrimaryRegistry(null);
reg.recordDelegation("term_worker", "term_lead_opus");
reg.recordDelegation("term_worker", "term_lead_sol");
assertEquals("term_lead_sol", reg.nudgeTargetFor("term_worker").orElseThrow(),
"the lead waiting on the reply is the one that sent the work most recently");
}
/** After a restart the map is empty while the durable inbox still holds the reply. */
@Test
void anUnknownDelegationFallsBackToThePinnedPrimary() {
var reg = new PrimaryRegistry("term_pinned");
assertEquals("term_pinned", reg.nudgeTargetFor("term_never_seen").orElseThrow());
}
@Test
void withNoPinAndNoDelegationNobodyIsNudged() {
var reg = new PrimaryRegistry(null);
assertTrue(reg.nudgeTargetFor("term_worker").isEmpty(),
"guessing would interrupt the wrong lead with someone else's result; the durable "
+ "inbox makes pull the correct degradation");
}
@Test
void releasingASessionForgetsItsLead() {
var reg = new PrimaryRegistry(null);
reg.recordDelegation("term_worker", "term_lead");
reg.forgetDelegation("term_worker");
assertTrue(reg.nudgeTargetFor("term_worker").isEmpty());
}
@Test
void aBlankOrNullDelegationIsIgnoredRatherThanStored() {
var reg = new PrimaryRegistry(null);
reg.recordDelegation("term_worker", " ");
reg.recordDelegation(null, "term_lead");
reg.recordDelegation(" ", "term_lead");
assertTrue(reg.nudgeTargetFor("term_worker").isEmpty());
assertTrue(reg.nudgeTargetFor(null).isEmpty());
}
}
@@ -1,5 +1,10 @@
package dev.ltms.bridged.msg;
import com.rabbitmq.client.AMQP;
import com.rabbitmq.client.Channel;
import com.rabbitmq.client.Connection;
import com.rabbitmq.client.ConnectionFactory;
import org.junit.jupiter.api.BeforeAll;
import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import org.testcontainers.containers.RabbitMQContainer;
@@ -19,19 +24,46 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* excluded from {@code mvn test}/{@code mvn clean install} (which stay hermetic and need no Docker);
* run it with Docker present via {@code mvn test -Pcontract}.
*
* <p>Two broker modes:
* <ul>
* <li><b>Locally</b> ({@code AMQP_URI} unset): Testcontainers spins a RabbitMQ container. Requires
* a working Docker engine; see the "Running the contract tests" note in
* {@code docs/CB-307-Reliable-Delivery.md} for the {@code api.version} engine-compat pin.</li>
* <li><b>In CI</b> ({@code AMQP_URI} set): a RabbitMQ service container provisions the broker and
* {@code AMQP_URI} points at it, so the contract job needs <em>no</em> Docker on the runner.</li>
* </ul>
*
* <p>It proves the port contract on genuine infrastructure: eventual visibility of a published reply,
* ack removal, msgId dedup, and — the reason Stage 2 exists — cross-restart durability: an unacked
* reply survives closing the inbox and is redelivered to a fresh connection.
* ack removal, msgId dedup, explicit ownership, and — the reason Stage 2 exists — cross-restart
* durability: an unacked reply survives closing the inbox and is redelivered to a fresh connection.
*/
@Tag("contract")
@Testcontainers
// disabledWithoutDocker=false on purpose: on the CI path (AMQP_URI set) no container is started and
// the class must still RUN against the external broker even though the runner has no Docker — a
// disabled-without-docker check would silently skip the whole contract suite there.
@Testcontainers(disabledWithoutDocker = false)
class AmqpReplyInboxContractTest {
@Container
static final RabbitMQContainer BROKER =
// When a broker is provisioned out-of-band (CI service container), AMQP_URI takes us straight to
// it and we never touch Testcontainers. Unset locally → Testcontainers starts the container below.
private static final String EXTERNAL_URI = System.getenv("AMQP_URI");
private static final RabbitMQContainer BROKER =
new RabbitMQContainer(DockerImageName.parse("rabbitmq:3.13-management"));
// No @Container on BROKER: the JUnit 5 extension would force-start it even when AMQP_URI is set.
// Start it manually only on the local (no-external-broker) path; Ryuk reaps it on JVM exit.
@BeforeAll
static void startBrokerUnlessExternal() {
if (EXTERNAL_URI == null) {
BROKER.start();
}
}
private static String uri() {
if (EXTERNAL_URI != null) {
return EXTERNAL_URI;
}
// guest/guest against the mapped AMQP port. No trailing slash: an empty path is vhost "",
// which does not exist — omitting it selects the default vhost "/".
return "amqp://guest:guest@" + BROKER.getHost() + ":" + BROKER.getAmqpPort();
@@ -41,6 +73,7 @@ class AmqpReplyInboxContractTest {
void publishThenPeekThenAck() throws Exception {
String target = "worker-pub-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.own(target);
inbox.publish(target, "m1", "hello primary");
List<ReplyInbox.InboxMessage> got = awaitPeek(inbox, target);
@@ -58,6 +91,7 @@ class AmqpReplyInboxContractTest {
void duplicateMsgIdIsNotDoubleQueued() throws Exception {
String target = "worker-dedup-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.own(target);
inbox.publish(target, "dup", "first");
awaitPeek(inbox, target);
inbox.publish(target, "dup", "second"); // same msgId — must be a no-op
@@ -74,8 +108,9 @@ class AmqpReplyInboxContractTest {
void unackedReplySurvivesRestartAndIsRedelivered() throws Exception {
String target = "worker-durable-" + System.nanoTime();
// First "process life": publish, see it held, but crash before acking.
// First "process life": own, publish, see it held, but crash before acking.
try (AmqpReplyInbox first = AmqpReplyInbox.open(uri())) {
first.own(target);
first.publish(target, "persist-1", "survive me");
assertEquals(1, awaitPeek(first, target).size());
// no ack — simulate a java -jar bounce with the reply still pending
@@ -83,6 +118,7 @@ class AmqpReplyInboxContractTest {
// Second "process life": a fresh connection to the same broker must be redelivered the reply.
try (AmqpReplyInbox second = AmqpReplyInbox.open(uri())) {
second.own(target);
List<ReplyInbox.InboxMessage> got = awaitPeek(second, target);
assertEquals(1, got.size(), "an unacked persistent reply is redelivered after restart");
assertEquals("persist-1", got.getFirst().msgId());
@@ -93,11 +129,43 @@ class AmqpReplyInboxContractTest {
// Third life: once acked, it is gone for good — durability is not endless replay.
try (AmqpReplyInbox third = AmqpReplyInbox.open(uri())) {
third.own(target);
Thread.sleep(500);
assertTrue(third.peek(target).isEmpty(), "an acked reply does not come back on the next restart");
}
}
@Test
void publishDoesNotAttachAConsumer() throws Exception {
String target = "worker-pub-no-consumer-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri());
Connection inspect = newConnection()) {
inbox.own(target);
inbox.publish(target, "m1", "published");
awaitPeek(inbox, target); // ensure the owner's consumer received it
try (Channel ch = inspect.createChannel()) {
AMQP.Queue.DeclareOk ok = ch.queueDeclare(queueName(target), true, false, false, null);
assertEquals(1, ok.getConsumerCount(),
"publish must not attach a consumer; only the owner's consumer should exist");
}
}
}
@Test
void releaseCancelsConsumerAndClearsHeld() throws Exception {
String target = "worker-release-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.own(target);
inbox.publish(target, "m1", "release me");
assertEquals(1, awaitPeek(inbox, target).size(), "owned target holds the reply");
inbox.release(target);
assertTrue(inbox.peek(target).isEmpty(),
"release clears the local held snapshot");
}
}
/** Poll peek (broker delivery is async) until a reply for {@code target} appears or ~10s elapse. */
@SuppressWarnings("BusyWait") // deliberate poll for async broker delivery, bounded by the deadline
private static List<ReplyInbox.InboxMessage> awaitPeek(AmqpReplyInbox inbox, String target)
@@ -110,4 +178,15 @@ class AmqpReplyInboxContractTest {
}
return msgs;
}
/** A separate broker connection for inspecting queue state without disturbing the inbox. */
private static Connection newConnection() throws Exception {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri());
return factory.newConnection("contract-inspector");
}
private static String queueName(String target) {
return "agent." + target + ".inbox";
}
}
@@ -1,5 +1,6 @@
package dev.ltms.bridged.msg;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CountDownLatch;
@@ -10,13 +11,18 @@ import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for {@link InMemoryReplyInbox}: publish, peek, ack, dedup, FIFO ordering, and thread
* safety under concurrent publish vs. drain.
* Unit tests for {@link InMemoryReplyInbox}: publish, peek, ack, dedup, FIFO ordering, thread
* safety under concurrent publish vs. drain, and explicit ownership.
*/
class InMemoryReplyInboxTest {
private final ReplyInbox inbox = new InMemoryReplyInbox();
@BeforeEach
void setUp() {
inbox.own("term_a");
}
@Test
void publishThenPeekReturnsTheMessage() {
inbox.publish("term_a", "m1", "hello");
@@ -72,6 +78,7 @@ class InMemoryReplyInboxTest {
@Test
void perTargetIsolation() {
inbox.own("term_b");
inbox.publish("term_a", "m1", "for-a");
inbox.publish("term_b", "m2", "for-b");
assertEquals(1, inbox.peek("term_a").size());
@@ -142,4 +149,40 @@ class InMemoryReplyInboxTest {
exec.shutdown();
}
}
@Test
void publishWithoutOwnDoesNotClaimOwnership() {
String unowned = "term_unowned";
inbox.publish(unowned, "m1", "hello");
// Without an owner, peek returns nothing — publish did not imply consume.
assertTrue(inbox.peek(unowned).isEmpty(),
"publishing to an unowned target must not make it peekable");
// Owning afterwards makes the already-published message visible.
inbox.own(unowned);
var msgs = inbox.peek(unowned);
assertEquals(1, msgs.size());
assertEquals("m1", msgs.getFirst().msgId());
}
@Test
void peekAndAckAreNoOpsForUnownedTarget() {
assertTrue(inbox.peek("term_not_owned").isEmpty());
inbox.ack("term_not_owned", "m1"); // no-op, should not throw
}
@Test
void releaseStopsConsumingAndClearsHeld() {
inbox.publish("term_a", "m1", "hello");
assertEquals(1, inbox.peek("term_a").size());
inbox.release("term_a");
assertTrue(inbox.peek("term_a").isEmpty(),
"after release, the local snapshot is cleared");
}
@Test
void ownIsIdempotent() {
inbox.own("term_a");
inbox.publish("term_a", "m1", "hello");
assertEquals(1, inbox.peek("term_a").size());
}
}
@@ -5,7 +5,9 @@ import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.inject.Injector;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CompletableFuture;
@@ -15,6 +17,7 @@ import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
@@ -30,7 +33,14 @@ class MessageServiceTest {
private final Rendezvous rendezvous = new Rendezvous();
private final CompletionResolver completion = new CompletionResolver(agents, rendezvous);
private final Injector injector = new Injector(agents, completion);
private final MessageService messages = new MessageService(agents, injector, rendezvous);
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
private final MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
@BeforeEach
void setUp() {
// CB-520: the inbox only peeks/acks targets it owns.
inbox.own(T);
}
/** Run {@code send} on a background thread; the current thread drives the worker's turn. */
private CompletableFuture<MessageService.Reply> sendAsync() {
@@ -313,6 +323,141 @@ class MessageServiceTest {
assertEquals("first done", firstReply.text());
}
// --- CB-548: delegator ownership is recorded only on an ACCEPTED send ----------------------
private static final String LEAD_L = "term_lead_l";
private static final String LEAD_A = "term_lead_a";
/**
* The bug CB-548 fixes: L holds worker W, then architect A attempts W and times out BUSY. With
* delegator ownership recorded at {@code bridge_send} <em>request</em> time, A's rejected call
* would overwrite L — and W's late no-waiter reply would be pushed to A, who never owned the
* turn. The accepted-delivery hook must not fire for a BUSY send, so L stays the delegator.
*/
@Test
void busySenderDoesNotBecomeTheDelegatingOwner() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null);
// L accepts a delegation to W: the send wins the lock and queues delivery → L is recorded.
CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync(
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L)));
awaitWaiting();
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"an accepted send owns the delegation");
// A attempts W while L holds it → BUSY (lock never taken) → its hook never fires.
MessageService.Reply busy = messages.send(T, "second", 100, () -> reg.recordDelegation(T, LEAD_A));
assertEquals(MessageService.Outcome.BUSY, busy.outcome());
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"a BUSY send must not steal the delegator ownership it never earned");
// L completes so the test thread is not left pinned.
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
assertTrue(rendezvous.resolve(T, "first done"));
MessageService.Reply firstReply = first.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, firstReply.outcome());
}
/**
* Once L's accepted delegation is fully done, a later <em>accepted</em> send from A may
* legitimately become the new delegator — ownership follows the turn, not the first caller.
*/
@Test
void anAcceptedSendAfterThePriorOwnerFinishesBecomesTheNewOwner() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null);
CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync(
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L)));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
assertTrue(rendezvous.resolve(T, "first done"));
MessageService.Reply firstReply = first.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, firstReply.outcome());
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(), "L owned the first turn");
// L finished; A's later accepted send takes the delegation over.
CompletableFuture<MessageService.Reply> second = CompletableFuture.supplyAsync(
() -> messages.send(T, "second", 5000, () -> reg.recordDelegation(T, LEAD_A)));
awaitWaiting();
assertEquals(LEAD_A, reg.nudgeTargetFor(T).orElseThrow(),
"an accepted send after the owner finished becomes the new delegator");
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
assertTrue(rendezvous.resolve(T, "second done"));
MessageService.Reply secondReply = second.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, secondReply.outcome());
}
/**
* CB-548 requirement: answering an existing {@code bridge_ask} is the SAME delegation, so it must
* not rewrite ownership. L accepted the send (owned), the worker paused to ask, and L answers via
* turnId — ownership stays L throughout; the answer path never touches the registry.
*/
@Test
void answeringAnAskDoesNotRewriteDelegatorOwnership() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null);
CompletableFuture<MessageService.Reply> send = CompletableFuture.supplyAsync(
() -> messages.send(T, "do X", 5000, () -> reg.recordDelegation(T, LEAD_L)));
awaitWaiting();
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(), "L owns the delegation");
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up, then pauses to ask
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertNotNull(q.turnId());
// L answers the ask on the same turn; the answer path must not touch ownership.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting(); // the answering send reopened its forward waiter
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"answering an ask keeps L as the delegator — ownership is not rewritten");
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
/**
* CB-548: {@code onAccepted} is a public callback, so a throwing one must not orphan the turn.
* The waiter is opened first, then the hook runs BEFORE delivery is queued — so a throw fails
* the send loudly, closes its waiter, and never enqueues a message the worker would pick up and
* reply into the void.
*/
@Test
void aThrowingAcceptedHookLeavesNoStaleWaiterOrQueuedOrphan() {
assertThrows(IllegalStateException.class,
() -> messages.send(T, "doomed", 500,
() -> { throw new IllegalStateException("ownership hook failed"); }),
"a throwing ownership hook fails the send loudly");
assertFalse(rendezvous.isWaiting(T), "the failed send must not leave a stale rendezvous waiter");
// Give the injector a delivery window: with nothing enqueued, nothing may reach the worker.
injector.onStatus(T, AgentStatus.IDLE);
boolean doomedQueued = herdr.calls.stream()
.anyMatch(c -> c.method().equals("agent.prompt")
&& String.valueOf(c.params()).contains("doomed"));
assertFalse(doomedQueued, "a throwing ownership hook must not leave a queued, orphanable message");
}
/**
* The async (fire-and-poll) path runs the same {@code send} on a background thread, so the
* accepted-delivery hook must thread through it — ownership is recorded exactly as blocking sends.
*/
@Test
void asyncSendRecordsOwnershipOnAcceptance() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null);
messages.sendAsync(T, "async task", () -> reg.recordDelegation(T, LEAD_L));
awaitWaiting(); // the background send won the lock, queued, and opened its waiter
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"the async path records delegator ownership on acceptance, like the blocking path");
}
// --- CB-307 reply inbox ----------------------------------------------------------------
@Test
@@ -8,6 +8,8 @@ import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
@@ -20,6 +22,26 @@ class RendezvousTest {
private final Rendezvous rendezvous = new Rendezvous();
/**
* CB-548: an {@code open} is atomic fail-if-present so a double open can never replace the first
* waiter another send is blocked on. {@code MessageService} serializes sends per session, so in
* correct code this cannot happen — the rejection is a loud tripwire for an invariant violation,
* and the first waiter must survive it and stay resolvable.
*/
@Test
void openRejectsADoubleOpenAndKeepsTheFirstWaiterRegisteredAndResolvable() {
CompletableFuture<Rendezvous.Resolution> first = rendezvous.open(W);
assertThrows(IllegalStateException.class, () -> rendezvous.open(W),
"a second open while one is registered is rejected loudly, not a silent replace");
assertSame(first, rendezvous.currentWaiter(W), "the first waiter remains the registered one");
assertTrue(rendezvous.resolve(W, "first wins"), "the first waiter is still resolvable");
Rendezvous.Resolution r = first.getNow(null);
assertEquals(Rendezvous.Kind.REPLY, r.kind());
assertEquals("first wins", r.text(), "the resolution lands on the first waiter, not the rejected one");
}
@Test
void openAskMintsAUniqueTurnScopedToItsSessionAndCoalescesDuplicates() {
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
@@ -45,6 +45,7 @@ class ReplyPushLoopTest {
void setUp() {
registry = new PrimaryRegistry(PRIMARY);
inbox = new InMemoryReplyInbox();
inbox.own(WORKER); // CB-520: the inbox only peeks/acks targets it owns
scheduler = Executors.newSingleThreadScheduledExecutor();
}
@@ -133,10 +134,10 @@ class ReplyPushLoopTest {
loop(1, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"one nudge (2 agent.send calls) should have been sent");
"one nudge (1 agent.prompt call) should have been sent");
// Exactly one nudge = exactly 2 agent.send calls (text + submit)
assertEquals(2, rec.sendCount());
// Exactly one nudge = exactly 1 agent.prompt call (it submits itself)
assertEquals(1, rec.sendCount());
assertTrue(rec.sentParams().stream()
.anyMatch(e -> e.getValue().toString().contains("bridge_poll")),
"nudge text should contain bridge_poll");
@@ -153,9 +154,9 @@ class ReplyPushLoopTest {
loop.onReplyQueued(WORKER); // second call — should be a no-op
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"expected exactly one nudge (2 sends)");
"expected exactly one nudge (1 prompt)");
Thread.sleep(200);
assertEquals(2, rec.sendCount(),
assertEquals(1, rec.sendCount(),
"second onReplyQueued must not trigger another nudge");
}
@@ -165,15 +166,15 @@ class ReplyPushLoopTest {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
rec.sendLatch = new CountDownLatch(cap * 2);
rec.sendLatch = new CountDownLatch(cap);
loop(cap, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(5, TimeUnit.SECONDS),
cap + " nudges (" + (cap * 2) + " sends) should have fired");
cap + " nudges (" + cap + " prompts) should have fired");
Thread.sleep(300);
assertEquals(cap * 2, rec.sendCount(),
"exactly " + (cap * 2) + " agent.send calls (cap=" + cap + ")");
assertEquals(cap, rec.sendCount(),
"exactly " + cap + " agent.prompt calls (cap=" + cap + ")");
}
// --- nudge format --------------------------------------------------------------------------
@@ -197,7 +198,7 @@ class ReplyPushLoopTest {
loop(1, 50, metrics).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"one nudge (2 agent.send calls) should have been sent");
"one nudge (1 agent.prompt call) should have been sent");
// The delivered count is bumped on the scheduler thread right after the send that releases
// the latch — settle briefly so the counter is published before we read it.
Thread.sleep(200);
@@ -261,13 +262,13 @@ class ReplyPushLoopTest {
}
/**
* Thread-safe recording fake that counts agent.send calls. Uses synchronized access
* so the scheduler thread and test thread never race.
* Thread-safe recording fake that counts agent.prompt calls (protocol 19: one nudge = one
* prompt). Uses synchronized access so the scheduler thread and test thread never race.
*/
private static final class RecordingHerdrClient implements HerdrClient {
private final List<Map.Entry<String, Object>> calls =
Collections.synchronizedList(new ArrayList<>());
volatile CountDownLatch sendLatch = new CountDownLatch(2);
volatile CountDownLatch sendLatch = new CountDownLatch(1);
@Override
public JsonNode call(String method, Object params) {
@@ -277,7 +278,7 @@ class ReplyPushLoopTest {
.put("terminal_id", PRIMARY)
.put("agent_status", "idle")); // recording double is always injectable
}
if ("agent.send".equals(method)) {
if ("agent.prompt".equals(method)) {
calls.add(Map.entry(method, params));
sendLatch.countDown();
}
@@ -0,0 +1,183 @@
package dev.ltms.bridged.placement;
import org.junit.jupiter.api.Test;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for the placement policies. They run with no herdr and no launcher — pure selection
* logic exercised through the descriptor type so CB-308 host expansion will not need to rewrite
* these assertions.
*/
class PlacementPolicyTest {
private static Function<String, Integer> noSessions() {
return name -> 0;
}
private static PlacementContext ctx(List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable) {
return new PlacementContext("b", candidates, liveCount, unreachable);
}
private static PlacementContext ctx(List<PlacementCandidate> candidates,
Function<String, Integer> liveCount) {
return ctx(candidates, liveCount, Set.of());
}
@Test
void fixedReturnsDefaultEvenIfOtherProfilesExist() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = ctx(List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b")), noSessions());
assertEquals("b", policy.select(ctx).profile());
}
@Test
void fixedFallsBackToFirstCandidateWhenNoDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null,
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of());
assertEquals("a", policy.select(ctx).profile());
}
@Test
void fixedThrowsWhenNoProfilesAndNoDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null, List.of(), noSessions(), Set.of());
assertThrows(PlacementException.class, () -> policy.select(ctx));
}
@Test
void roundRobinCyclesThroughAvailableProfiles() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b"),
PlacementCandidate.profile("c"));
assertEquals("a", policy.select(ctx(candidates, noSessions())).profile());
assertEquals("b", policy.select(ctx(candidates, noSessions())).profile());
assertEquals("c", policy.select(ctx(candidates, noSessions())).profile());
assertEquals("a", policy.select(ctx(candidates, noSessions())).profile());
}
@Test
void roundRobinSkipsProfilesAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 2),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = Map.of("a", 2)::get;
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, liveCount)).profile());
}
}
@Test
void roundRobinThrowsWhenAllAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, 1));
Function<String, Integer> liveCount = Map.of("a", 1, "b", 1)::get;
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
@Test
void weightedAlternatesEvenlyWithEqualWeights() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.5f, null),
PlacementCandidate.profile("b", 0.5f, null));
int a = 0, b = 0;
for (int i = 0; i < 100; i++) {
String p = policy.select(ctx(candidates, noSessions())).profile();
if ("a".equals(p)) a++;
else if ("b".equals(p)) b++;
}
assertEquals(50, a, "equal weights should split 50/50");
assertEquals(50, b);
}
@Test
void weightedHoldsThreeToOneRatio() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.75f, null),
PlacementCandidate.profile("b", 0.25f, null));
int a = 0, b = 0;
for (int i = 0; i < 40; i++) {
String p = policy.select(ctx(candidates, noSessions())).profile();
if ("a".equals(p)) a++;
else if ("b".equals(p)) b++;
}
assertEquals(30, a, "0.75/0.25 should yield a 3:1 ratio over a multiple of 4");
assertEquals(10, b);
}
@Test
void weightedSkipsProfileAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = name -> "a".equals(name) ? 1 : 0;
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, liveCount)).profile());
}
}
@Test
void weightedThrowsWhenAllAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, 1));
Function<String, Integer> liveCount = name -> 1;
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
@Test
void weightedThrowsWhenAllUnreachable() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b"));
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, noSessions(), Set.of("a", "b"))));
assertTrue(e.getMessage().contains("unreachable"), e.getMessage());
}
@Test
void mixedExclusionMessageNamesBothReasons() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = name -> "a".equals(name) ? 1 : 0;
Set<String> unreachable = new HashSet<>();
unreachable.add("b");
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount, unreachable)));
assertTrue(e.getMessage().contains("1 at maxLoad"), e.getMessage());
assertTrue(e.getMessage().contains("1 unreachable"), e.getMessage());
}
@Test
void unknownPolicyNameThrows() {
assertThrows(IllegalArgumentException.class, () -> PlacementPolicies.fromName("random"));
}
}
@@ -10,6 +10,7 @@ import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
@@ -28,6 +29,7 @@ import java.net.http.HttpResponse;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import static org.junit.jupiter.api.Assertions.*;
@@ -74,7 +76,12 @@ class BridgedAppTest {
poller = new StatusPoller(agents, injector, 5); // delivers when the fake reports idle
poller.start();
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
sessions.onAcquire(inbox::own);
// The static REST tests address "term_a" without acquiring it through SessionManager, so own
// it directly so the inbox contract holds for those endpoints.
inbox.own("term_a");
MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
app = new BridgedApp(herdr, workers, sessions, messages, this.presence, null)
.build().start("127.0.0.1", 0);
return app.port();
@@ -118,7 +125,7 @@ class BridgedAppTest {
assertEquals(200, res.statusCode());
JsonNode body = mapper.readTree(res.body());
assertEquals("ok", body.get("status").asText());
assertEquals(14, body.get("herdr").get("protocol").asInt());
assertEquals(19, body.get("herdr").get("protocol").asInt());
}
@Test
@@ -156,22 +163,27 @@ class BridgedAppTest {
HttpResponse<String> res = req(port, "POST", "/workers");
assertEquals(201, res.statusCode());
JsonNode body = mapper.readTree(res.body());
assertEquals("w9:pW_1", body.get("paneId").asText());
// CB-519: the responded paneId is a host-unique opaque UUID, not the herdr pane coordinate.
String id = body.get("paneId").asText();
assertNotEquals("w9:pRoot_1", id, "paneId is the host-unique id, not the herdr pane");
assertDoesNotThrow(() -> UUID.fromString(id), "paneId must be a UUID: " + id);
assertEquals("spawning", body.get("state").asText());
// Subscription boundary: agent.start carried base_url + token in its env map.
Map<String, Object> start = params(herdr, "agent.start");
// Subscription boundary (protocol 19): tab.create carried base_url + token in its env
// map — the seed shell the agent starts into is what inherits them.
Map<String, Object> create = params(herdr, "tab.create");
@SuppressWarnings("unchecked")
Map<String, String> env = (Map<String, String>) start.get("env");
Map<String, String> env = (Map<String, String>) create.get("env");
assertEquals("http://gx00.gw:8000", env.get("ANTHROPIC_BASE_URL"));
assertEquals("tok-abc", env.get("ANTHROPIC_AUTH_TOKEN"));
assertEquals(List.of("claude"), start.get("argv"));
// Placement: worker space ensured, worker started INTO its own tab, seed shell
// dropped, and the tab given a friendly label.
// Placement: worker space ensured, worker started INTO its tab's seed pane (which
// becomes the worker pane — nothing is dropped), and the tab given a friendly label.
Map<String, Object> start = params(herdr, "agent.start");
assertTrue(herdr.called("workspace.create"), "worker space must be found-or-created");
assertEquals("w9:t2", start.get("tab_id"), "worker must start into its dedicated tab");
assertEquals("w9:pRoot", params(herdr, "pane.close").get("pane_id"), "seed shell pane dropped");
assertEquals("claude", start.get("kind"), "herdr resolves the executable from kind");
assertEquals("w9:pRoot_1", start.get("pane_id"), "worker must start into its tab's seed pane");
assertFalse(herdr.called("pane.close"), "the seed pane IS the worker pane — never dropped");
assertEquals("worker: ltms-local #1", params(herdr, "tab.rename").get("label"),
"tab label carries the worker number so siblings stay distinct");
}
@@ -215,7 +227,7 @@ class BridgedAppTest {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(201, req(port, "POST", "/workers?cwd=/tmp/proj").statusCode());
assertEquals("/tmp/proj", params(herdr, "agent.start").get("cwd"), "the worker starts in cwd");
assertEquals("/tmp/proj", params(herdr, "tab.create").get("cwd"), "the worker starts in cwd");
}
@Test
@@ -322,7 +334,7 @@ class BridgedAppTest {
HttpResponse<String> res = send.get(6, java.util.concurrent.TimeUnit.SECONDS);
assertEquals(200, res.statusCode());
assertEquals("LGTM ship it", mapper.readTree(res.body()).get("reply").asText());
// (injection via agent.send is covered deterministically by the timeout-working test)
// (injection via agent.prompt is covered deterministically by the timeout-working test)
}
@Test
@@ -394,7 +406,7 @@ class BridgedAppTest {
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":150}");
assertEquals(202, res.statusCode());
assertEquals("queued", mapper.readTree(res.body()).get("status").asText());
assertFalse(herdr.called("agent.send"), "no injection while the worker is mid-turn");
assertFalse(herdr.called("agent.prompt"), "no injection while the worker is mid-turn");
}
@Test
@@ -405,7 +417,7 @@ class BridgedAppTest {
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":250}");
assertEquals(202, res.statusCode());
assertEquals("working", mapper.readTree(res.body()).get("status").asText());
assertTrue(herdr.called("agent.send"), "message was injected");
assertTrue(herdr.called("agent.prompt"), "message was injected");
}
@Test
@@ -0,0 +1,208 @@
package dev.ltms.bridged.session;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-525 acceptance test for tool-surface isolation. This is one of the few tests that drives real
* {@code git} — the behaviour under test is precisely what {@link GitWorktrees} does to a checkout,
* so a fake would assert nothing. Everything happens inside a {@link TempDir} throwaway repo.
*/
class GitWorktreesTest {
/** A project MCP config with servers in it — what this repo actually commits. */
private static final String WITH_SERVERS = """
{
"mcpServers": {
"jetbrains": { "type": "sse", "url": "http://localhost:64342/sse" }
}
}
""";
/** An opencode config carrying a {@code {file:.secrets/...}} reference — CB-543's crash repro. */
private static final String OPENCODE_WITH_FILE_REF = """
{
"env": {
"CONTEXT7_TOKEN": "{file:.secrets/context7-token}"
}
}
""";
/** A non-empty autoenv file — the form that would prompt for authorization in a worktree. */
private static final String AUTOENV_WITH_DIRECTIVE = "export HELLO=world\n";
private static Path initRepo(Path dir) throws Exception {
Files.createDirectories(dir);
git(dir, "init", "-q", "-b", "main");
git(dir, "config", "user.email", "test@example.invalid");
git(dir, "config", "user.name", "Test");
Files.writeString(dir.resolve(".mcp.json"), WITH_SERVERS);
Files.writeString(dir.resolve("README.md"), "seed\n");
git(dir, "add", ".mcp.json", "README.md");
git(dir, "commit", "-q", "-m", "seed");
return dir;
}
private static void git(Path cwd, String... args) throws Exception {
List<String> cmd = new java.util.ArrayList<>(List.of("git"));
cmd.addAll(List.of(args));
Process p = new ProcessBuilder(cmd).directory(cwd.toFile()).redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git timed out: " + String.join(" ", cmd));
assertEquals(0, p.exitValue(), "git " + String.join(" ", args) + " failed:\n" + out);
}
/** Pending changes to {@code file} in {@code cwd}, empty when git considers it unmodified. */
private static String status(Path cwd, String file) throws Exception {
Process p = new ProcessBuilder("git", "status", "--porcelain", "--", file)
.directory(cwd.toFile()).redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git status timed out");
return out;
}
/**
* The heart of CB-525: a provisioned worktree must not inherit the primary's MCP servers. Without
* the isolation step the checked-out {@code .mcp.json} carries them in, and a worker navigating
* through the primary's IDE servers edits the primary's tree while building its own.
*/
@Test
void aProvisionedWorktreeInheritsNoMcpServers(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-525-a", "HEAD");
Path mcp = Path.of(wt).resolve(".mcp.json");
assertTrue(Files.exists(mcp), ".mcp.json must still exist — present and explicitly empty");
String body = Files.readString(mcp);
assertFalse(body.contains("jetbrains"), "worktree inherited the primary's MCP servers:\n" + body);
assertTrue(body.replaceAll("\\s+", "").contains("\"mcpServers\":{}"),
"expected an explicitly empty server map, got:\n" + body);
}
/** Neutralizing must not look like work in progress, or a worker would commit it into its PR. */
@Test
void theNeutralizedConfigIsNotAPendingLocalModification(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-525-b", "HEAD");
assertEquals("", status(Path.of(wt), ".mcp.json"),
"the neutralized .mcp.json shows as modified — --skip-worktree did not take");
}
/** Isolation is the worktree's business only; the primary's own checkout must be untouched. */
@Test
void thePrimaryCheckoutIsLeftAlone(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-525-c", "HEAD");
assertEquals(WITH_SERVERS, Files.readString(repo.resolve(".mcp.json")),
"the primary's .mcp.json was rewritten — isolation reached out of the worktree");
}
/** A repo that commits no {@code .mcp.json} still gets one, so nothing can be inherited later. */
@Test
void aRepoWithoutAnMcpConfigStillGetsANeutralOne(@TempDir Path tmp) throws Exception {
Path repo = tmp.resolve("repo");
Files.createDirectories(repo);
git(repo, "init", "-q", "-b", "main");
git(repo, "config", "user.email", "test@example.invalid");
git(repo, "config", "user.name", "Test");
Files.writeString(repo.resolve("README.md"), "seed\n");
git(repo, "add", "README.md");
git(repo, "commit", "-q", "-m", "seed");
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-525-d", "HEAD");
// Untracked is the normal case here, so the --skip-worktree branch must be skipped rather
// than run and fail: `update-index --skip-worktree` on an unknown path exits non-zero.
String body = Files.readString(Path.of(wt).resolve(".mcp.json"));
assertTrue(body.replaceAll("\\s+", "").contains("\"mcpServers\":{}"), body);
}
/**
* CB-543's repro: a tracked {@code opencode.json} carries a {@code {file:.secrets/...}} reference
* to a gitignored secret that never reaches a worktree, and opencode refuses to start on it. The
* worktree's copy must be neutralized and hidden like {@code .mcp.json}.
*/
@Test
void aTrackedOpencodeConfigIsNeutralizedAndHidden(@TempDir Path tmp) throws Exception {
Path repo = tmp.resolve("repo");
Files.createDirectories(repo);
git(repo, "init", "-q", "-b", "main");
git(repo, "config", "user.email", "test@example.invalid");
git(repo, "config", "user.name", "Test");
Files.writeString(repo.resolve(".mcp.json"), WITH_SERVERS);
Files.writeString(repo.resolve("opencode.json"), OPENCODE_WITH_FILE_REF);
Files.writeString(repo.resolve("README.md"), "seed\n");
git(repo, "add", ".mcp.json", "opencode.json", "README.md");
git(repo, "commit", "-q", "-m", "seed");
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-543-a", "HEAD");
String body = Files.readString(Path.of(wt).resolve("opencode.json"));
assertFalse(body.contains(".secrets"),
"worktree kept a dangling {file:...} secret reference:\n" + body);
assertEquals("{}", body.replaceAll("\\s+", ""),
"expected an empty JSON object stub, got:\n" + body);
assertEquals("", status(Path.of(wt), "opencode.json"),
"the neutralized opencode.json shows as modified — --skip-worktree did not take");
}
/** A config the repo does not carry must be skipped — no stub invented, provisioning still succeeds. */
@Test
void anAbsentConfigIsSkippedWithoutError(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo")); // only .mcp.json + README are committed
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-543-b", "HEAD");
assertFalse(Files.exists(Path.of(wt).resolve("opencode.json")),
"a stub was invented for a config the repo does not carry");
assertFalse(Files.exists(Path.of(wt).resolve(".autoenv")),
"a stub was invented for a config the repo does not carry");
// .mcp.json's long-standing create-always behaviour must be unchanged.
assertTrue(Files.exists(Path.of(wt).resolve(".mcp.json")), ".mcp.json stub was dropped");
}
/** All three protected configs are covered: each one present in a worktree is neutralized and hidden. */
@Test
void allThreeConfigsAreNeutralizedWhenPresent(@TempDir Path tmp) throws Exception {
Path repo = tmp.resolve("repo");
Files.createDirectories(repo);
git(repo, "init", "-q", "-b", "main");
git(repo, "config", "user.email", "test@example.invalid");
git(repo, "config", "user.name", "Test");
Files.writeString(repo.resolve(".mcp.json"), WITH_SERVERS);
Files.writeString(repo.resolve("opencode.json"), OPENCODE_WITH_FILE_REF);
Files.writeString(repo.resolve(".autoenv"), AUTOENV_WITH_DIRECTIVE);
Files.writeString(repo.resolve("README.md"), "seed\n");
git(repo, "add", ".mcp.json", "opencode.json", ".autoenv", "README.md");
git(repo, "commit", "-q", "-m", "seed");
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-543-c", "HEAD");
assertTrue(Files.readString(Path.of(wt).resolve(".mcp.json"))
.replaceAll("\\s+", "").contains("\"mcpServers\":{}"),
".mcp.json was not neutralized");
assertEquals("{}", Files.readString(Path.of(wt).resolve("opencode.json")).replaceAll("\\s+", ""),
"opencode.json was not neutralized");
assertEquals("", Files.readString(Path.of(wt).resolve(".autoenv")),
".autoenv was not neutralized");
assertEquals("", status(Path.of(wt), ".mcp.json"), ".mcp.json still shows as modified");
assertEquals("", status(Path.of(wt), "opencode.json"), "opencode.json still shows as modified");
assertEquals("", status(Path.of(wt), ".autoenv"), ".autoenv still shows as modified");
}
}
@@ -39,13 +39,29 @@ class SessionManagerTest {
}
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap) {
return sessionManager(herdr, clock, contextCap, false);
}
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap,
boolean clearAfterTurn) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
return new SessionManager(workers, new GitWorktrees(), clock, contextCap);
return new SessionManager(workers, new GitWorktrees(), clock, contextCap, clearAfterTurn);
}
@Test
void primaryContactWithNoTerminalIsNotAReadinessSignal() {
// The MCP context extractor calls presence.markPresent(p.terminal()) on EVERY request,
// and the primary's terminal is null — the presence bridge must treat that as a no-op,
// not feed it into the READY transition (which NPEd on the first real primary contact).
SessionManager sessions = sessionManager(new FakeHerdr());
assertDoesNotThrow(() -> sessions.asPresence().markPresent(null));
assertDoesNotThrow(() -> sessions.asPresence().markPresent(" "));
}
@Test
@@ -69,6 +85,25 @@ class SessionManagerTest {
assertEquals(2, sessions.roster().size(), "both sessions are registered");
}
@Test
void aNullTerminalFromThePrimaryIsANoOpEvenWithSessionsRegistered() {
// The primary resolves to a Principal with no terminal, and BridgeMcp's context extractor
// forwards that null into markPresent on EVERY MCP call. It only reached the registry scan
// once a session existed, so this NPE'd the primary's second spawn while the first passed.
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
assertDoesNotThrow(() -> sessions.asPresence().markPresent(null),
"the primary's null terminal must not blow up an unrelated tool call");
assertDoesNotThrow(() -> sessions.onDelivered(null));
assertDoesNotThrow(() -> sessions.onTurnComplete(null));
assertDoesNotThrow(() -> sessions.onTurnFailed(null));
assertEquals(WorkerSession.State.SPAWNING, sessions.get(session.paneId()).orElseThrow().state(),
"and must not transition any registered session");
}
@Test
void presenceMovesSpawningToReadyAndDeliveredTurnMovesToDone() {
FakeHerdr herdr = new FakeHerdr();
@@ -142,9 +177,11 @@ class SessionManagerTest {
assertEquals(1, sessions.roster().size(), "only the fresh session remains");
assertEquals(fresh.paneId(), sessions.roster().getFirst().paneId());
// The old session was the first spawn → pane w9:pRoot_1 (CB-519: the registry key is the
// uuid id, so teardown is asserted on the real pane coordinate).
long paneCloseCount = herdr.calls.stream()
.filter(c -> "pane.close".equals(c.method()))
.filter(c -> oldPane.equals(((Map<?, ?>) c.params()).get("pane_id")))
.filter(c -> "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
assertEquals(1, paneCloseCount, "the old worker was torn down");
}
@@ -277,7 +314,7 @@ class SessionManagerTest {
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(WorkerSession.State.DONE, updated.state(), "session finishes second turn");
assertEquals(2, updated.turnCount(), "turn count tracks both deliveries");
long releaseCloseCount = paneCloseCallsFor(herdr, session.paneId());
long releaseCloseCount = paneCloseCallsFor(herdr, "w9:pRoot_1"); // the real pane coordinate
assertEquals(0, releaseCloseCount, "cap disabled — no forced release of the worker pane");
}
@@ -300,10 +337,54 @@ class SessionManagerTest {
assertTrue(sessions.get(session.paneId()).isEmpty(), "session released after cap reached");
assertTrue(sessions.roster().isEmpty(), "released session leaves roster");
assertEquals(1, paneCloseCallsFor(herdr, session.paneId()),
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_1"),
"forced release tears the worker pane down exactly once");
}
@Test
void clearAfterTurnResetsContextWithoutDoubleCountingTheTurn() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 0, true);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
assertTrue(sessions.onTurnCompleteWithPostAction(session.terminalId()));
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(1, updated.turnCount(), "the reset is housekeeping, not a second delegation");
assertEquals(List.of("/clear"), promptTexts(herdr));
}
@Test
void contextCapReleaseWinsOverClearAfterTurn() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 1, true);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
assertFalse(sessions.hasPostTurnAction(session.terminalId()),
"a session at its cap will be released, not reset for reuse");
assertFalse(sessions.onTurnCompleteWithPostAction(session.terminalId()));
assertTrue(sessions.get(session.paneId()).isEmpty());
assertTrue(promptTexts(herdr).isEmpty(), "never send /clear into a worker being torn down");
}
@Test
void clearAfterTurnFalsePreservesCompletionWithoutAControlPrompt() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 0, false);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
sessions.onTurnComplete(session.terminalId());
assertEquals(WorkerSession.State.DONE, sessions.get(session.paneId()).orElseThrow().state());
assertTrue(promptTexts(herdr).isEmpty());
}
@Test
void drainAllReleasesBusyAndReadySessionsAndWaitsForBusy() {
long[] clock = {0};
@@ -321,9 +402,10 @@ class SessionManagerTest {
assertTrue(sessions.roster().isEmpty(), "drain clears the roster");
assertTrue(sessions.get(ready.paneId()).isEmpty(), "ready session is released");
assertTrue(sessions.get(busy.paneId()).isEmpty(), "busy session is released after timeout");
assertEquals(1, paneCloseCallsFor(herdr, ready.paneId()),
// ready is the first spawn → pane w9:pRoot_1, busy the second → w9:pRoot_2 (FakeHerdr order).
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_1"),
"ready worker pane is torn down");
assertEquals(1, paneCloseCallsFor(herdr, busy.paneId()),
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_2"),
"busy worker pane is torn down");
}
@@ -334,6 +416,13 @@ class SessionManagerTest {
.count();
}
private static List<String> promptTexts(FakeHerdr herdr) {
return herdr.calls.stream()
.filter(c -> "agent.prompt".equals(c.method()))
.map(c -> String.valueOf(((Map<?, ?>) c.params()).get("text")))
.toList();
}
// --- CB-306 spawn-readiness gate: no half-registered session on timeout ----------------
@Test
@@ -32,7 +32,8 @@ class WorktreeSessionManagerTest {
private static String startCwd(FakeHerdr herdr) {
@SuppressWarnings("unchecked")
Map<String, Object> start = (Map<String, Object>) herdr.lastCall("agent.start").params();
// Protocol 19: the worker's cwd rides on pane creation (tab.create), not agent.start.
Map<String, Object> start = (Map<String, Object>) herdr.lastCall("tab.create").params();
Object cwd = start.get("cwd");
return cwd == null ? null : cwd.toString();
}
@@ -81,7 +82,7 @@ class WorktreeSessionManagerTest {
void worktreeAcquireRunsParityOverlayWithProfileDefaults() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.track(".mcp.json")
.track(".envrc")
.exists(".claude/settings.local.json");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
@@ -92,11 +93,14 @@ class WorktreeSessionManagerTest {
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
assertNotNull(overlay);
assertEquals("/repo", overlay.repoRoot());
assertEquals(List.of(".mcp.json", ".claude/settings.local.json", ".env", ".envrc"),
assertEquals(List.of(".claude/settings.local.json", ".env", ".envrc"),
overlay.requested(), "default parity overlay is used when unset");
assertEquals(List.of(".mcp.json", ".claude/settings.local.json"), overlay.copied(),
assertFalse(overlay.requested().contains(".mcp.json"),
"CB-525: replicating the primary's MCP config gives a worker the primary's IDE "
+ "servers, which navigate its edits out of its own worktree");
assertEquals(List.of(".claude/settings.local.json", ".envrc"), overlay.copied(),
"existing paths are copied; missing paths are skipped");
assertEquals(List.of(".mcp.json"), overlay.skipWorktree(),
assertEquals(List.of(".envrc"), overlay.skipWorktree(),
"tracked copied paths are --skip-worktree'd");
}
@@ -1,6 +1,7 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
@@ -14,6 +15,7 @@ import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.*;
@@ -29,52 +31,91 @@ class ClaudeCodeLauncherTest {
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
}
/** The {@code args} of the last agent.start — protocol 19: everything after the executable. */
@SuppressWarnings("unchecked")
private List<String> spawnedArgv(FakeHerdr herdr) {
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("argv");
private List<String> spawnedArgs(FakeHerdr herdr) {
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("args");
}
@Test
void appendsBridgeMcpAndReplyCharterWhenMcpUrlSet() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("ccs", "ltms-local"), "http://127.0.0.1:8765/mcp").spawn();
service(herdr, List.of("claude"), "http://127.0.0.1:8765/mcp").spawn();
List<String> argv = spawnedArgv(herdr);
assertEquals(List.of("ccs", "ltms-local"), argv.subList(0, 2), "base command preserved first");
assertTrue(argv.contains("--mcp-config"));
assertTrue(argv.stream().anyMatch(a -> a.contains("\"bridge\"") && a.contains("http://127.0.0.1:8765/mcp")),
List<String> args = spawnedArgs(herdr);
assertTrue(args.contains("--mcp-config"));
assertTrue(args.stream().anyMatch(a -> a.contains("\"bridge\"") && a.contains("http://127.0.0.1:8765/mcp")),
"inline bridge MCP config present");
assertTrue(argv.contains("--append-system-prompt"));
assertTrue(argv.stream().anyMatch(a -> a.contains("bridge_reply")), "reply charter present");
assertTrue(args.contains("--append-system-prompt"));
assertTrue(args.stream().anyMatch(a -> a.contains("bridge_reply")), "reply charter present");
}
@Test
void startRetriesWhileTheSeedShellBoots() {
// tab.create returns before the seed shell reaches its prompt; herdr refuses agent.start
// into a not-ready pane with agent_pane_busy. The launcher must wait it out, not fail.
FakeHerdr herdr = new FakeHerdr().agentPaneBusyTimes(2);
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
Map.of("ltms-local", new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null)),
"ltms-local", _ -> null,
0, () -> clock[0], () -> clock[0] += 50);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn succeeds once the shell is ready");
assertEquals(3, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"two busy rejections, then the successful start");
}
@Test
void startResolvesTheExecutableFromKindAndDropsArgvZero() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("claude"), null).spawn();
Map<?, ?> start = (Map<?, ?>) herdr.lastCall("agent.start").params();
assertEquals("claude", start.get("kind"), "herdr launches the canonical executable by kind");
// CB-533: the shared fixture pins model "coder", so the model flag is the whole args list.
// What this test guards is that argv[0] is NOT repeated — herdr supplies it from `kind`.
assertEquals(List.of("--model", "coder"), start.get("args"),
"the configured executable is not repeated in args");
}
@Test
void noBridgeFlagsWhenMcpUrlAbsent() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("bash", "-c", "sleep 1"), null).spawn();
assertEquals(List.of("bash", "-c", "sleep 1"), spawnedArgv(herdr), "argv untouched without mcpUrl");
service(herdr, List.of("claude", "--verbose"), null).spawn();
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("--mcp-config"), "no bridge mount without mcpUrl");
assertFalse(args.contains("--append-system-prompt"), "no reply charter without mcpUrl");
// CB-533: the model flag is independent of the MCP mount — pinning the model is not part of
// "mount the bridge", so an unmounted worker still runs the model its profile names.
assertEquals(List.of("--verbose", "--model", "coder"), args,
"the operator's own args are preserved, in order, ahead of the model flag");
}
private ClaudeCodeLauncher multiProfile(FakeHerdr herdr) {
BridgedConfig.Worker gx10 = new BridgedConfig.Worker("gx10", "http://gx10.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "gx10"), "tab", "bridged-workers", "w #{n}", null, null, null);
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
BridgedConfig.Worker ollama = new BridgedConfig.Worker("ollama", "http://ollama.ltms.dev", null,
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "ollama"), "tab", "bridged-workers", "w #{n}", null, null, null);
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx10.gw", "ollama.ltms.dev")),
Map.of("gx10", gx10, "ollama", ollama), "gx10", _ -> "tok");
}
@Test
@SuppressWarnings("unchecked")
void spawnPicksTheNamedProfilesBaseUrlAndArgv() {
void spawnPicksTheNamedProfilesBaseUrl() {
FakeHerdr herdr = new FakeHerdr();
multiProfile(herdr).spawn("ollama");
Map<String, Object> start = (Map<String, Object>) herdr.lastCall("agent.start").params();
Map<String, String> env = (Map<String, String>) start.get("env");
assertEquals("http://ollama.ltms.dev", env.get("ANTHROPIC_BASE_URL"), "the named profile's base_url");
assertEquals(List.of("ccs", "ollama"), start.get("argv"), "the named profile's launch command");
assertEquals("http://ollama.ltms.dev", startEnv(herdr).get("ANTHROPIC_BASE_URL"),
"the named profile's base_url");
}
@Test
@@ -86,8 +127,9 @@ class ClaudeCodeLauncherTest {
@SuppressWarnings("unchecked")
private static String startCwd(FakeHerdr herdr) {
// The worker's cwd is set on agent.start (an agent pane does not inherit the tab's cwd).
Object v = ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("cwd");
// Protocol 19: the worker's cwd is set at pane creation (tab.create), where the seed
// shell — which the agent starts into — is rooted.
Object v = ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("cwd");
return v == null ? null : v.toString();
}
@@ -119,9 +161,10 @@ class ClaudeCodeLauncherTest {
// --- CB-302 git-forge token injection (worker checkpoint grant) ------------
/** Protocol 19: the worker's env is injected at pane creation (tab.create), not agent.start. */
@SuppressWarnings("unchecked")
private static Map<String, String> startEnv(FakeHerdr herdr) {
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("env");
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
}
@Test
@@ -227,14 +270,18 @@ class ClaudeCodeLauncherTest {
// --- PeerHandle indirection ----------------------------------------------------------------
@Test
void spawnReturnsPeerHandleWithIdEqualToPaneId() {
void spawnReturnsPeerHandleWithHostUniqueOpaqueId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn must return a non-null handle");
assertEquals("w9:pW_1", handle.id(), "handle.id() must equal the agent's paneId");
// CB-519: id() is a host-unique opaque UUID, decoupled from the herdr pane coordinate.
assertNotEquals("w9:pRoot_1", handle.id(),
"handle.id() must NOT be the herdr pane id");
assertDoesNotThrow(() -> UUID.fromString(handle.id()),
"handle.id() must be a UUID: " + handle.id());
}
@Test
@@ -257,6 +304,20 @@ class ClaudeCodeLauncherTest {
assertTrue(caps.contains(Capability.MID_TURN_ASK), "every Claude Code peer supports mid-turn ask");
assertTrue(caps.contains(Capability.WORKTREE), "every CLI peer supports worktree cwd");
assertTrue(caps.contains(Capability.ORPHAN_REAP), "every herdr launcher supports orphan reap");
assertTrue(caps.contains(Capability.CONTEXT_RESET), "Claude Code supports /clear");
}
@Test
void clearContextUsesTheClaudeCommandThroughTheOwningHandle() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("claude"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertTrue(svc.clearContext(handle.id()));
Map<?, ?> prompt = (Map<?, ?>) herdr.lastCall("agent.prompt").params();
assertEquals("/clear", prompt.get("text"));
assertEquals("w9:pRoot_1", prompt.get("target"));
}
@Test
@@ -293,6 +354,65 @@ class ClaudeCodeLauncherTest {
assertEquals("/work/proj", cwd, "effectiveCwd via SpawnRequest must match the three-arg resolution");
}
// --- CB-547a: durable session identity (mint / resume / no-identity legacy) -----------------
@Test
void freshSpawnMintsASessionIdAndPassesTheName() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, null, "my-session", null));
List<String> args = spawnedArgs(herdr);
int flag = args.indexOf("--session-id");
assertTrue(flag >= 0 && flag + 1 < args.size(), "--session-id present: " + args);
String minted = args.get(flag + 1);
assertDoesNotThrow(() -> UUID.fromString(minted), "--session-id is a valid UUID: " + minted);
assertEquals("my-session", args.get(args.indexOf("-n") + 1), "the logical name rides as -n");
assertEquals(minted, handle.agentSessionId(),
"the resume handle is the minted id, known before the agent has written anything");
assertEquals("my-session", handle.sessionName(), "the handle carries the logical name");
}
@Test
void resumeSpawnPassesDashRAndNeverASessionId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, null, "my-session", "cb-resume-1"));
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("--session-id"), "--session-id must NOT be passed on a resume (conflicts with -r)");
assertEquals("cb-resume-1", args.get(args.indexOf("-r") + 1), "-r carries the prior session id");
assertEquals("cb-resume-1", handle.agentSessionId(), "a resume adopts the prior id as its own");
assertEquals("my-session", handle.sessionName(), "the logical name survives a resume");
}
@Test
void noIdentitySpawnKeepsTheLegacyArgvAndCarriesNoSessionHandle() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, null));
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("--session-id"), "no identity → no --session-id");
assertFalse(args.contains("-n"), "no identity → no -n");
assertFalse(args.contains("-r"), "no identity → no -r");
assertNull(handle.agentSessionId(), "no identity → no resume handle");
assertNull(handle.sessionName(), "no identity → no logical name");
}
@Test
void capabilitiesIncludeSessionNameAndSessionResume() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
Set<Capability> caps = svc.capabilities();
assertTrue(caps.contains(Capability.SESSION_NAME), "Claude Code surfaces the bridge's logical name (-n)");
assertTrue(caps.contains(Capability.SESSION_RESUME), "Claude Code can relaunch onto a prior conversation (-r)");
}
@Test
void profilesViaPeerLauncherMatchesExistingApi() {
FakeHerdr herdr = new FakeHerdr();
@@ -320,6 +440,40 @@ class ClaudeCodeLauncherTest {
assertTrue(herdr.called("pane.close"), "stop via handle.id() must close the pane");
}
// --- CB-519: host-unique id, decoupled from the pane coordinate ------------------------------
@Test
void twoSpawnsOnTheSamePaneNeverCollideOnHostUniqueId() {
// Two spawns may be placed on the same herdr pane coordinate (e.g. a pane that was reused
// or re-reported after a restart); the host-unique id must not collide even then.
FakeHerdr herdr = new FakeHerdr().pinNextStarts(2, "term_shared", "w9:pShared");
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle a = svc.spawn(new SpawnRequest(null, null, null));
PeerHandle b = svc.spawn(new SpawnRequest(null, null, null));
assertNotEquals(a.id(), b.id(),
"two spawns on the same pane coordinate get distinct host-unique ids");
assertNotEquals("w9:pShared", a.id(), "id is not the pane coordinate");
assertNotEquals("w9:pShared", b.id(), "id is not the pane coordinate");
}
@Test
void stopResolvesTheHostUniqueIdToThePaneThatSpawnedIt() {
// CB-519: id() != paneId, so stop(id) must tear down the exact pane the id names — and no
// other live peer's pane.
FakeHerdr herdr = new FakeHerdr(); // deterministic panes w9:pRoot_1, w9:pRoot_2 per spawn
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle a = svc.spawn(new SpawnRequest(null, null, null));
PeerHandle b = svc.spawn(new SpawnRequest(null, null, null));
svc.stop(b.id());
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_2"), "stop(b.id()) closes only b's pane");
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"), "a's pane is untouched");
}
// --- CB-306 spawn-readiness gate -----------------------------------------------------------
private static Map<String, BridgedConfig.Worker> workerConfigMap(String profile, String mcpUrl) {
@@ -355,8 +509,9 @@ class ClaudeCodeLauncherTest {
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn returns a handle when worker becomes injectable");
assertEquals("w9:pW_1", handle.id(), "handle id matches the started pane");
assertEquals(0, paneCloseCount(herdr, "w9:pW_1"),
assertNotEquals("w9:pRoot_1", handle.id(),
"handle id is a host-unique opaque id, not the started pane");
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"no pane.close when worker becomes injectable before timeout");
}
@@ -376,12 +531,12 @@ class ClaudeCodeLauncherTest {
PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(ex.getMessage().contains("w9:pW_1"),
assertTrue(ex.getMessage().contains("w9:pRoot_1"),
"exception message references the paneId: " + ex.getMessage());
assertTrue(ex.getMessage().contains("1000"),
"exception message references the timeout: " + ex.getMessage());
assertTrue(clock[0] >= 1000, "fake clock advanced past the timeout: " + clock[0]);
assertEquals(1, paneCloseCount(herdr, "w9:pW_1"),
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_1"),
"pane was closed on timeout (no orphan left behind)");
}
@@ -414,8 +569,9 @@ class ClaudeCodeLauncherTest {
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn still succeeds with zero timeout");
assertEquals(0, paneCloseCount(herdr, handle.id()),
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"no orphan pane close from the gate path");
assertDoesNotThrow(() -> UUID.fromString(handle.id()));
}
// --- CB-511: worker environment seeding -----------------------------------------------------
@@ -440,7 +596,7 @@ class ClaudeCodeLauncherTest {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("JAVA_HOME", "/opt/jdk", "PATH", "/profile/bin"));
null, Map.of("JAVA_HOME", "/opt/jdk", "PATH", "/profile/bin"), null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
k -> "PATH".equals(k) ? "/daemon/bin" : null).spawn();
@@ -461,7 +617,7 @@ class ClaudeCodeLauncherTest {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("ANTHROPIC_BASE_URL", "http://evil.example.com"));
null, Map.of("ANTHROPIC_BASE_URL", "http://evil.example.com"), null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null).spawn();
@@ -469,4 +625,161 @@ class ClaudeCodeLauncherTest {
assertEquals("http://gx00.gw:8000", startEnv(herdr).get("ANTHROPIC_BASE_URL"),
"the guard-checked baseUrl must win over any env: entry, or the boundary is bypassable");
}
// ── CB-533: the model is pinned on the command line, not only in the environment ────────────
/** A launcher for a profile identical but for its {@code model:} — the only variable here. */
private ClaudeCodeLauncher serviceWithModel(FakeHerdr herdr, String model) {
BridgedConfig.Worker cfg = profileWithModel(model);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null);
}
private static BridgedConfig.Worker profileWithModel(String model) {
return new BridgedConfig.Worker("sonnet", "http://gx00.gw:8000", model, null,
"BRIDGED_WORKER_TOKEN", List.of("ccs", "sonnet"), "tab", "bridged-workers",
"w #{n}", "http://127.0.0.1:8765/mcp", null, null);
}
@Test
void aConfiguredModelIsPassedAsAModelFlagAsWellAsTheEnvVar() {
// ANTHROPIC_MODEL alone loses to `ccs`, which exports its own model family over whatever it
// inherited — so a profile that set model: was silently overruled by its own launcher.
FakeHerdr herdr = new FakeHerdr();
serviceWithModel(herdr, "claude-sonnet-5").spawn("sonnet", null, null);
assertEquals("claude-sonnet-5", startEnv(herdr).get("ANTHROPIC_MODEL"));
List<String> args = spawnedArgs(herdr);
int flag = args.indexOf("--model");
assertTrue(flag >= 0, "the flag is what survives a wrapper argv like [ccs, sonnet]");
assertEquals("claude-sonnet-5", args.get(flag + 1));
}
@Test
void theModelFlagComesLastSoItOutranksTheOperatorsOwnArgv() {
FakeHerdr herdr = new FakeHerdr();
serviceWithModel(herdr, "claude-sonnet-5").spawn("sonnet", null, null);
List<String> args = spawnedArgs(herdr);
assertEquals(args.size() - 2, args.indexOf("--model"));
}
@Test
void aProfileWithNoModelGetsNoModelFlag() {
// gx10 deliberately leaves model: unset so ccs owns selection; adding a flag would make
// this file a second source of truth for exactly the thing it declines to decide.
FakeHerdr herdr = new FakeHerdr();
serviceWithModel(herdr, null).spawn("sonnet", null, null);
assertFalse(spawnedArgs(herdr).contains("--model"));
assertNull(startEnv(herdr).get("ANTHROPIC_MODEL"));
}
// --- CB-539: subscription-profile opt-in ----------------------------------------------------
/** A claude-code profile on the subscription: no baseUrl (by design), no off-sub endpoint. */
private static BridgedConfig.Worker subscriptionCfg(String profile, String baseUrl) {
return new BridgedConfig.Worker(
profile, baseUrl, "sonnet", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", profile), "tab", "bridged-workers", "w #{n}", null, null, null,
null, null, null, Map.of(), null, null, true);
}
@Test
void defaultRefusalIsPreservedForClaudeProfileWithNoBaseUrl() {
// Requirement 1: absent subscription:true ⇒ byte-identical refusal to today. A claude-code
// profile with no baseUrl and no subscription must still be refused (it would bill the sub).
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", null, "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers", "w #{n}", null, null, null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
GuardException ex = assertThrows(GuardException.class, () -> svc.spawn("ltms-local", null, null));
assertTrue(ex.getMessage().contains("no ANTHROPIC_BASE_URL"),
"the refusal names the missing baseUrl: " + ex.getMessage());
assertEquals(0, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"nothing was spawned before the refusal");
}
@Test
void subscriptionProfileSpawnsWithoutInjectedAnthropicVars() {
// Requirement on subscription:true: no baseUrl is required (or injected), and neither
// ANTHROPIC_BASE_URL nor ANTHROPIC_AUTH_TOKEN is injected even though the token env would
// resolve one if asked.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = subscriptionCfg("sonnet", null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "would-be-token").spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"), "no baseUrl injected for a subscription profile");
assertNull(env.get("ANTHROPIC_AUTH_TOKEN"), "no auth token injected for a subscription profile");
assertEquals("sonnet", env.get("ANTHROPIC_MODEL"),
"the model alias is still injected; only the subscription-boundary vars are dropped");
}
@Test
void subscriptionPlusBaseUrlIsRefused() {
// Requirement 2: subscription:true + a baseUrl state opposite intents — refuse at spawn,
// naming the profile, rather than silently picking a winner.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = subscriptionCfg("sonnet", "http://gx00.gw:8000");
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
IllegalStateException ex = assertThrows(IllegalStateException.class,
() -> svc.spawn("sonnet", null, null));
assertTrue(ex.getMessage().contains("sonnet"), "refusal names the profile: " + ex.getMessage());
assertTrue(ex.getMessage().contains("subscription"), "refusal explains the contradiction: " + ex.getMessage());
assertEquals(0, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"nothing was spawned before the contradiction was refused");
}
@Test
void nonSubscriptionProfilesAreStillAllowlistChecked() {
// Requirement 3: the guard keeps its teeth for every other profile — a base_url whose host is
// not on the allowlist is still refused, whether or not any subscription profile exists.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker rogue = new BridgedConfig.Worker(
"rogue", "http://evil.example.com:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "rogue"), "tab", "bridged-workers", "w #{n}", null, null, null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(rogue.profile(), rogue), rogue.profile(), _ -> null);
GuardException ex = assertThrows(GuardException.class, () -> svc.spawn("rogue", null, null));
assertTrue(ex.getMessage().contains("not on the"), "refusal cites the allowlist: " + ex.getMessage());
assertEquals(0, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"nothing was spawned before the allowlist refusal");
}
@Test
void aSubscriptionProfileHasAnEnvSuppliedAnthropicBindingStripped() {
// CB-542: even a subscription profile whose env: carries ANTHROPIC_BASE_URL (or AUTH_TOKEN)
// must not hand them to the worker — on the subscription path no guard would vet them. Config
// load refuses this loudly; this launcher-side strip is the belt-and-braces that makes the
// invariant hold for a profile built in code that never passed through that validation.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"sonnet", null, "sonnet", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "sonnet"), "tab", "bridged-workers", "w #{n}", null, null, null,
null, null, null,
Map.of("ANTHROPIC_BASE_URL", "http://evil.example.com",
"ANTHROPIC_AUTH_TOKEN", "sk-ant-bad", "JAVA_HOME", "/opt/jdk"),
null, null, true);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "would-be-token").spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"),
"the unguarded endpoint must not survive into the worker");
assertNull(env.get("ANTHROPIC_AUTH_TOKEN"),
"the unguarded token must not survive into the worker");
assertEquals("/opt/jdk", env.get("JAVA_HOME"),
"only the Anthropic binding keys are stripped; the rest of env: still applies");
}
}
@@ -1,18 +1,32 @@
package dev.ltms.bridged.worker;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.PlacementException;
import dev.ltms.bridged.placement.PlacementPolicies;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.EnumSet;
import java.util.HashMap;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.*;
@@ -46,6 +60,87 @@ class CompositePeerLauncherTest {
List.of(claudeAdapter(herdr), opencodeAdapter(herdr)), "claude");
}
/**
* A minimal concrete HerdrPeerLauncher for policy tests. It either returns a fake handle for the
* requested profile or throws, depending on {@code failProfiles}. buildLaunch is a stub; only
* spawn/stop/list/caps/reap are exercised by the composite.
*/
private static final class StubLauncher extends HerdrPeerLauncher {
private final Set<String> failProfiles;
private final Map<String, Integer> spawnCounts = new HashMap<>();
StubLauncher(String prefix, FakeHerdr herdr,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Set<String> failProfiles) {
super(prefix, new AgentControl(herdr), new WorkspaceControl(herdr),
profiles, defaultProfile, _ -> null, 0L, System::currentTimeMillis, () -> { });
this.failProfiles = Set.copyOf(failProfiles);
}
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
return new Launch(Map.of(), List.of());
}
@Override
public PeerHandle spawn(SpawnRequest req) {
String p = (req.profileName() == null || req.profileName().isBlank())
? defaultProfile() : req.profileName();
spawnCounts.merge(p, 1, Integer::sum);
if (failProfiles.contains(p)) {
throw new PeerUnreachableException(p + " is down");
}
return new PeerHandle() {
@Override public String id() { return "pane-" + p; }
@Override public String terminalId() { return "term-" + p; }
@Override public String profile() { return p; }
};
}
@Override
public void stop(String id) { }
@Override
public List<Agent> list() { return List.of(); }
@Override
public Set<Capability> capabilities() { return EnumSet.noneOf(Capability.class); }
@Override
public int reapOrphanWorkers() { return 0; }
int spawnCount(String profile) {
return spawnCounts.getOrDefault(profile, 0);
}
}
private static BridgedConfig.Worker stubWorker(String profile) {
return new BridgedConfig.Worker(profile, "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers",
"w #{n}", null, null, null, null, null, null, null, null, null);
}
private static BridgedConfig.Worker stubWorker(String profile, float weight, Integer maxLoad) {
return new BridgedConfig.Worker(profile, "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers",
"w #{n}", null, null, null, null, null, null, null,
weight, maxLoad);
}
/**
* An <em>order-preserving</em> profile map. Never {@code Map.of} here: its iteration order is
* salted per JVM run, and the weighted policy breaks an exact-weight tie on candidate order —
* so a {@code Map.of} would make "which profile is tried first" a coin flip per run and any
* assertion about the first attempt intermittently false.
*/
private static Map<String, BridgedConfig.Worker> ordered(String first, BridgedConfig.Worker a,
String second, BridgedConfig.Worker b) {
Map<String, BridgedConfig.Worker> m = new LinkedHashMap<>();
m.put(first, a);
m.put(second, b);
return m;
}
@SuppressWarnings("unchecked")
private static String startedName(FakeHerdr herdr) {
return (String) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("name");
@@ -127,15 +222,42 @@ class CompositePeerLauncherTest {
@Test
void stopTearsDownAPaneSpawnedThroughTheComposite() {
// CB-519: handle.id() is a host-unique opaque UUID, not the herdr pane — stop(id) must
// resolve it through the owning adapter down to the actual pane coordinate it spawned.
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
PeerHandle handle = composite.spawn(new SpawnRequest("gemini", null, null));
assertNotEquals("w9:pRoot_1", handle.id(), "the id is decoupled from the pane coordinate");
composite.stop(handle.id());
assertTrue(herdr.calls.stream()
.anyMatch(c -> c.method().equals("pane.close")
&& handle.id().equals(((Map<?, ?>) c.params()).get("pane_id"))),
"stop routes to the spawning adapter and closes that worker's pane");
&& "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id"))),
"stop routes to the spawning adapter and closes exactly that worker's pane");
}
@Test
void opencodeContextResetIsANoOpAndWarnsOnlyOnce() {
FakeHerdr herdr = new FakeHerdr();
CompositePeerLauncher composite = composite(herdr);
PeerHandle handle = composite.spawn(new SpawnRequest("gemini", null, null));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
assertFalse(composite.clearContext(handle.id()));
assertFalse(composite.clearContext(handle.id()));
} finally {
logger.detachAppender(appender);
}
assertFalse(opencodeAdapter(herdr).capabilities().contains(Capability.CONTEXT_RESET));
assertTrue(herdr.calls.stream().noneMatch(c -> "agent.prompt".equals(c.method())),
"never type Claude's /clear into an opencode prompt");
assertEquals(1, appender.list.stream()
.filter(e -> e.getFormattedMessage().contains("context reset is unsupported"))
.count(), "unsupported reset is logged once per adapter, not once per turn");
}
@Test
@@ -155,4 +277,167 @@ class CompositePeerLauncherTest {
() -> new CompositePeerLauncher(List.of(), "claude"),
"at least one adapter must be configured");
}
@Test
void fixedDefaultIsNoOpForUnqualifiedSpawns() {
FakeHerdr herdr = new FakeHerdr();
CompositePeerLauncher composite = composite(herdr);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertTrue(startedName(herdr).startsWith("claude-"),
"fixed placement still routes an unqualified spawn to the default profile");
assertEquals("claude", h.profile(), "the returned handle carries the resolved default profile");
}
@Test
void weightedPolicyGatesProfileAtMaxLoad() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"a", stubWorker("a", 1.0f, 1),
"b", stubWorker("b", 1.0f, null));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> "a".equals(name) ? 1 : 0);
for (int i = 0; i < 5; i++) {
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("b", h.profile(), "profile a is at maxLoad, so every spawn must land on b");
}
}
@Test
void weightedPolicyDistributesAccordingToWeightRatio() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"a", stubWorker("a", 0.75f, null),
"b", stubWorker("b", 0.25f, null));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> 0);
int a = 0, b = 0;
for (int i = 0; i < 40; i++) {
String p = composite.spawn(new SpawnRequest(null, null, null)).profile();
if ("a".equals(p)) a++;
else if ("b".equals(p)) b++;
}
assertEquals(30, a, "weighted distribution should hold the 3:1 ratio");
assertEquals(10, b);
}
@Test
void failoverRetriesNextCandidateWhenProfileIsUnreachable() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"a", stubWorker("a"),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of("a"));
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> 0);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("b", h.profile(), "the spawn must fail over from unreachable a to b");
assertEquals(1, adapter.spawnCount("a"), "a was tried once and failed");
assertEquals(1, adapter.spawnCount("b"), "b was tried once and succeeded");
}
/**
* Definition order — not hash order — decides an exact-weight tie. Paired with the test above
* (same two profiles, opposite declaration order, opposite expected first attempt) this pins the
* ordering contract from both sides: under a salted map one of the two must fail on every run.
*/
@Test
void reversingDefinitionOrderReversesWhichProfileIsTriedFirst() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"b", stubWorker("b"),
"a", stubWorker("a"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of("b"));
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 0);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("a", h.profile(), "b is declared first and unreachable, so the spawn lands on a");
assertEquals(1, adapter.spawnCount("b"), "b, declared first, is the one tried first");
}
@Test
void failoverBoundedByCandidateCount() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"a", stubWorker("a"),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of("a", "b"));
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> 0);
PeerUnreachableException e = assertThrows(PeerUnreachableException.class,
() -> composite.spawn(new SpawnRequest(null, null, null)));
assertTrue(e.getMessage().contains("no reachable worker profile"), e.getMessage());
assertEquals(1, adapter.spawnCount("a"));
assertEquals(1, adapter.spawnCount("b"));
}
@Test
void explicitSpawnAtMaxLoadThrowsPlacementExceptionNamingProfileLiveAndCap() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"a", stubWorker("a", 1.0f, 2),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), name -> "a".equals(name) ? 2 : 0);
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest("a", null, null)));
assertTrue(e.getMessage().contains("'a'"), "message names the profile: " + e.getMessage());
assertTrue(e.getMessage().contains("2 live"), "message names the live count: " + e.getMessage());
assertTrue(e.getMessage().contains("2 cap"), "message names the cap: " + e.getMessage());
assertEquals(0, adapter.spawnCount("a"), "at cap, the spawn is refused before any delegation");
}
@Test
void explicitSpawnUnderMaxLoadStillSucceeds() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"a", stubWorker("a", 1.0f, 2),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), name -> "a".equals(name) ? 1 : 0);
PeerHandle h = composite.spawn(new SpawnRequest("a", null, null));
assertEquals("a", h.profile(), "a profile under its cap accepts an explicit spawn");
assertEquals(1, adapter.spawnCount("a"), "the under-cap spawn is delegated");
}
@Test
void explicitSpawnWithNullMaxLoadIsNeverCapped() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"a", stubWorker("a", 1.0f, null),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
// A deliberately absurd live count: an unset maxLoad means unlimited, so it must never refuse.
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), name -> 1000);
PeerHandle h = composite.spawn(new SpawnRequest("a", null, null));
assertEquals("a", h.profile(), "a profile with no maxLoad is never capped, however many live workers");
}
@Test
void emptyCandidateSetThrowsClearException() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"a", stubWorker("a", 1.0f, 1),
"b", stubWorker("b", 1.0f, 1));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> 1);
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest(null, null, null)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
}
@@ -37,7 +37,7 @@ class OpenCodeLauncherTest {
private OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, BridgedConfig.Worker cfg) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), k -> "GITEA_ACCESS_TOKEN".equals(k) ? "tok" : null,
0, System::currentTimeMillis, () -> { }, configRoot);
0, System::currentTimeMillis, () -> { }, configRoot, configRoot);
}
@SuppressWarnings("unchecked")
@@ -45,14 +45,18 @@ class OpenCodeLauncherTest {
return (Map<String, Object>) herdr.lastCall("agent.start").params();
}
/** Protocol 19: the worker's env is injected at pane creation (tab.create), not agent.start. */
@SuppressWarnings("unchecked")
private static Map<String, String> startEnv(FakeHerdr herdr) {
return (Map<String, String>) lastStart(herdr).get("env");
Map<String, String> env =
(Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
return env == null ? Map.of() : env;
}
/** Protocol 19: agent.start carries only the args after the kind-resolved executable. */
@SuppressWarnings("unchecked")
private static List<String> startArgv(FakeHerdr herdr) {
return (List<String>) lastStart(herdr).get("argv");
private static List<String> startArgs(FakeHerdr herdr) {
return (List<String>) lastStart(herdr).get("args");
}
@Test
@@ -95,22 +99,24 @@ class OpenCodeLauncherTest {
}
@Test
void passesTheModelAsDashMFlag(@TempDir Path root) {
void passesTheModelAsDashMFlagAlongsideAutoApprove(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null)).spawn();
List<String> argv = startArgv(herdr);
assertEquals("opencode", argv.getFirst(), "base opencode command preserved first");
int m = argv.indexOf("-m");
List<String> args = startArgs(herdr);
assertTrue(args.contains("--auto"),
"--auto is present alongside -m so a spawned peer never blocks on approval");
int m = args.indexOf("-m");
assertTrue(m >= 0, "model is selected with -m");
assertEquals("google/gemini-2.5-pro", argv.get(m + 1), "the provider/model selector follows -m");
assertEquals("google/gemini-2.5-pro", args.get(m + 1), "the provider/model selector follows -m");
}
@Test
void noModelFlagWhenModelBlank(@TempDir Path root) {
void autoApproveIsUnconditionalWhenModelBlank(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg(null, null, null)).spawn();
assertEquals(List.of("opencode"), startArgv(herdr), "no model → argv is the bare opencode command");
assertEquals(List.of("--auto"), startArgs(herdr),
"--auto is unconditional: a model-less worker still must never block on approval");
}
@Test
@@ -124,14 +130,65 @@ class OpenCodeLauncherTest {
@Test
void capabilitiesDeclareOrphanReapAndMcpAskAndConditionalSelfPr(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
assertEquals(java.util.Set.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP),
assertEquals(java.util.Set.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP,
Capability.SESSION_RESUME),
service(herdr, root, opencodeCfg(null, null, null)).capabilities(),
"no git token → no SELF_PR");
"opencode can be resumed by its own session id, so SESSION_RESUME is always declared");
assertFalse(service(herdr, root, opencodeCfg(null, null, null))
.capabilities().contains(Capability.SESSION_NAME),
"opencode has no display-name flag, so SESSION_NAME must NOT be declared");
assertTrue(service(herdr, root, opencodeCfg(null, null, "GITEA_ACCESS_TOKEN"))
.capabilities().contains(Capability.SELF_PR),
"a git-token profile adds SELF_PR");
}
// --- CB-547: resume + post-hoc session discovery --------------------------------------------
@Test
void aResumeSpawnPassesTheSessionIdAsDashS(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null))
.spawn(new SpawnRequest(null, null, null, null, "ses_41b79fc90ffeI9E8uZv6VprUn2"));
List<String> args = startArgs(herdr);
int s = args.indexOf("-s");
assertTrue(s >= 0, "a resumed spawn carries opencode's -s flag");
assertEquals("ses_41b79fc90ffeI9E8uZv6VprUn2", args.get(s + 1),
"the resume target id follows -s");
}
@Test
void aFreshSpawnCarriesNoSessionFlag(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg(null, null, null))
.spawn(new SpawnRequest(null, null, null, null, null));
assertFalse(startArgs(herdr).contains("-s"),
"no resume target → a fresh session with no -s flag");
}
@Test
void theHandleDiscoversTheSessionIdForTheWorkersCwdOnlyAfterItAppears(@TempDir Path root,
@TempDir Path discRoot)
throws Exception {
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = new OpenCodeLauncher(new AgentControl(herdr),
new WorkspaceControl(herdr), Map.of("gemini", opencodeCfg(null, null, null)),
"gemini", _ -> null, 0, System::currentTimeMillis, () -> { }, root, discRoot);
PeerHandle handle = launcher.spawn(new SpawnRequest(null, "/work/dir", null));
// opencode writes the record only when the session is first persisted — the instant the
// pane is ready it does not exist, so agentSessionId() is null (never a spawn failure).
assertNull(handle.agentSessionId(), "no record yet → null, not a spawn-time block");
// Once the record appears (here: same cwd), lazy discovery resolves it — the handle's
// session id matches its own worktree, not another's.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "p1", "ses_a.json",
"ses_resolved", "/work/dir", 1000L);
assertEquals("ses_resolved", handle.agentSessionId(),
"agentSessionId() re-scans and picks up a record that has since been written");
}
@Test
void foreignWorkerMatchesOpencodePrefixButNotClaude() {
String nonce = "abc123";
@@ -163,14 +220,14 @@ class OpenCodeLauncherTest {
long[] clock = {0};
OpenCodeLauncher svc = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("gemini", opencodeCfg(null, null, null)), "gemini", _ -> null,
1000, () -> clock[0], () -> clock[0] += 50, root);
1000, () -> clock[0], () -> clock[0] += 50, root, root);
PeerUnreachableException ex = assertThrows(PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(clock[0] >= 1000, "the fake clock advanced past the timeout: " + clock[0]);
long closes = herdr.calls.stream()
.filter(c -> c.method().equals("pane.close"))
.filter(c -> "w9:pW_1".equals(((Map<?, ?>) c.params()).get("pane_id")))
.filter(c -> "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
assertEquals(1, closes, "the worker pane was reaped on timeout (no orphan)");
assertNotNull(ex.getMessage());
@@ -0,0 +1,90 @@
package dev.ltms.bridged.worker;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.FileTime;
import static org.junit.jupiter.api.Assertions.*;
/**
* {@link OpenCodeSessionDiscovery} matches an opencode session record by the worker's cwd (its
* {@code directory}) against opencode's on-disk storage. These tests populate a TEMP storage root
* themselves — never the operator's real {@code ~/.local/share/opencode}.
*/
class OpenCodeSessionDiscoveryTest {
/**
* Write a session record {@code {"id":..., "directory":...}} under
* {@code <root>/session/<projectID>/<fileName>} and stamp it with a known last-modified time,
* so "most recently modified wins" is deterministic. Static so the launcher test can reuse it.
*/
static void writeRecord(Path root, String projectId, String fileName, String id,
String directory, long lastModifiedEpochMillis) throws Exception {
Path dir = root.resolve("session").resolve(projectId);
Files.createDirectories(dir);
Path file = dir.resolve(fileName);
Files.writeString(file, "{\"id\":\"" + id + "\",\"directory\":\"" + directory
+ "\",\"projectID\":\"" + projectId + "\",\"version\":\"1.1.31\"}");
Files.setLastModifiedTime(file, FileTime.fromMillis(lastModifiedEpochMillis));
}
@Test
void findsTheRecordWhoseDirectoryEqualsTheCwd(@TempDir Path root) throws Exception {
writeRecord(root, "p1", "ses_a.json", "ses_aaa", "/w/a", 1000L);
writeRecord(root, "p2", "ses_b.json", "ses_bbb", "/w/b", 2000L);
assertEquals("ses_bbb", new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/b"),
"the record whose directory equals the cwd is the one found");
assertEquals("ses_aaa", new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"));
}
@Test
void aNonMatchingDirectoryYieldsNullRatherThanAMismatch(@TempDir Path root) throws Exception {
writeRecord(root, "p1", "ses_a.json", "ses_aaa", "/w/a", 1000L);
assertNull(new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/other"),
"no record for this cwd yet → null, not a wrong session");
}
@Test
void prefersTheMostRecentlyModifiedRecordWhenSeveralMatch(@TempDir Path root) throws Exception {
writeRecord(root, "p1", "old.json", "ses_old", "/w/a", 1000L);
writeRecord(root, "p2", "new.json", "ses_new", "/w/a", 5000L);
assertEquals("ses_new", new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"),
"the freshest record for the cwd wins");
}
@Test
void aMissingOrEmptyStorageRootYieldsNullWithoutThrowing(@TempDir Path root) throws Exception {
// Missing: no session dir at all under the root.
assertNull(new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"));
// Present but empty: a session dir with nothing in it produces no match, not a throw.
Path emptyRoot = root.resolve("empty");
Files.createDirectories(emptyRoot.resolve("session"));
assertNull(new OpenCodeSessionDiscovery(emptyRoot).sessionIdForDirectory("/w/a"));
}
@Test
void aBlankOrNullDirectoryYieldsNull(@TempDir Path root) {
OpenCodeSessionDiscovery discovery = new OpenCodeSessionDiscovery(root);
assertNull(discovery.sessionIdForDirectory(null));
assertNull(discovery.sessionIdForDirectory(" "));
}
@Test
void aMalformedRecordIsSkippedRatherThanFatal(@TempDir Path root) throws Exception {
// A record that fails to parse must not abort the scan of its siblings.
Path dir = root.resolve("session").resolve("p1");
Files.createDirectories(dir);
Files.writeString(dir.resolve("broken.json"), "{not valid json");
writeRecord(root, "p1", "good.json", "ses_good", "/w/a", 1000L);
assertEquals("ses_good", new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"),
"an unreadable record is skipped; a later valid one still matches");
}
}
+37
View File
@@ -162,3 +162,40 @@ Stage-2 AMQP adapter: absent → in-memory, present → AMQP).
Commit on your feature branch and reply with: the commit SHA, the surefire total (run/failures/errors), a
one-line note on the drain-surface decision (§2.4) you shipped, and confirmation that `.mcp.json`/`wiki/`
were untouched. The primary re-gates and integrates.
## 8. Running the contract tests (CB-521)
`AmqpReplyInboxContractTest` is the real-broker proof of the `ReplyInbox` port (eventual visibility, ack
removal, msgId dedup, cross-restart redelivery). It is `@Tag("contract")`, so the default
`mvn test` / `mvn clean install` **skip it** — that hermetic, Docker-free default is deliberate and
untouched. Run it explicitly when Docker (or a broker) is available:
```bash
cd bridged
mvn -Pcontract test -Dtest=AmqpReplyInboxContractTest # local: spins a RabbitMQ Testcontainers fixture
```
### Two broker modes
| Mode | Trigger | Broker | Needs Docker? |
|---|---|---|---|
| Local | `AMQP_URI` unset | Testcontainers starts `rabbitmq:3.13-management` | Yes |
| CI / external | `AMQP_URI` set | the broker at that URI (CI RabbitMQ service container) | **No** — binds straight to the URI, never touches Testcontainers |
In CI the broker is provided as a RabbitMQ **service container** and `AMQP_URI` points at it, so the
contract job runs the same assertions with no Docker on the runner and no skipped test
(see `.gitea/workflows/ci.yml` → `contract`). The `build` job stays hermetic and Docker-free — keep
that separation.
### Docker-engine discovery (why the contract profile pins `api.version`)
Out of the box, Testcontainers 1.20.4's docker-java client defaults to Docker API **1.32** when no
version is requested. Modern engines reject that as too old — on this host's OrbStack (`min API 1.40`)
testcontainers fails with *"Could not find a valid Docker environment … client version 1.32 is too
old"* even though the `docker` CLI works (the CLI negotiates a newer API).
The `contract` Maven profile sets `api.version=1.43` in surefire, which works on OrbStack and Docker
24+, and is overridable per host: `mvn -Pcontract -Dapi.version=1.54 test …`. It only applies under
`-Pcontract`, so the default build is unaffected. If your engine differs, set `-Dapi.version` to a
version ≥ your engine's minimum API (e.g. `docker version` shows `API version`).
+106 -8
View File
@@ -1,6 +1,6 @@
# CB-308 — Multi-Host Federation (Stage 5)
**Status:** design note (proposal)
**Status:** design note (proposal) — core design decisions resolved 2026-08-10 (§7)
**Depends on:** CB-307 (broker-based reliable delivery) — CB-308 is the multi-host layer built *on*
CB-307's broker fabric.
**Relates to:** CB-401 (`PeerHandle` opaque id), CB-304 (`rosterView`), CB-306 (spawn-readiness),
@@ -143,6 +143,8 @@ extending the same broker from "worker→primary reliability" to "gateway↔gate
5. **Trust** — the broker connection is now the security boundary. A gateway injects env/tokens at
daemon privilege (the CB-401 Stage-C concern), so a **remote-triggered spawn/send** needs
authn/authz: who may act on which host, and which control channels a gateway will honour.
*Authenticity* is resolved — signed messages, §7.1; *authorization* (who may do what) remains
open — §8.
## 5. The one thing the broker does NOT dissolve
@@ -191,15 +193,111 @@ useful), and build CB-308's items on top once the broker fabric exists. Choose C
channel naming **multi-host-ready** now (per-agent routing keys, a `roster.*` topic namespace) so
CB-308 doesn't have to repaint the topology.
## 7. Open questions
## 7. Resolved design decisions (2026-08-10)
Settled in a design review of this note + wiki chapter 10. The broker-level operational rules
(inbox caps, TLS + private broker, schema versioning, trace id, exclusive consumers, U8 broadcast)
are recorded in wiki 10 §10; the CB-308-side decisions are below. Entries 1–6 are the first-pass
decisions; 7–10 came out of the adversarial second-pass review (same day) and supersede 1–6 where
they overlap (notably: the envelope is no longer optional, and dedup is split by path).
1. **Sender authenticity — sign every message.** Each gateway holds its own signing key and signs
what it publishes (sender gid, `msgId`, timestamp). The receiving gateway verifies the
signature **and** checks against the roster that the claimed sender lives on the signing
gateway's host. This extends the single-host invariant — *identity comes from the connection,
never an argument* — across the broker: cross-host, identity comes from the key. Complements
(not replaces) per-gateway broker logins over TLS.
2. **Profiles are owned by the worker's host.** `bridge_spawn(profile, host)` resolves the name in
the *target* gateway's `bridged.yaml`. Gateways advertise their profile names in presence
heartbeats, so a leader sees what each host offers before spawning; an unknown name is a clear
error from the target. Secrets (base URLs, tokens) never leave the host that uses them.
3. **Repo provisioning — clone from the forge, pinned.** A cross-host spawn names the repo URL and
the exact commit. The target gateway clones from the forge into a local cache (first spawn
only), then cuts a per-worker worktree — the CB-301-ext flow with a clone step in front,
covered by the same repo-scoped forge token (CB-302). Git stays the only channel code moves
through.
4. **Asks are live-only, with expiry.** `ASK`/`ANSWER` (U2) traverse the broker as short-lived
(TTL'd) messages carrying the `turn_id`, and are never held durably — the single-host rule
kept. An answer arriving after its turn ended is **not** injected; it is dropped and the leader
gets a `TOO_LATE` notice, so the one failure case is loud rather than weird. Only terminal
replies are durable. Walkthrough: wiki 10 §7.4.
5. **Spawn dedup — a spawn id, remembered on the target.** The control queue redelivers like any
queue; a replayed `SpawnRequest` must not double-spawn. Requests carry a unique spawn id; the
target gateway keeps a short memory of handled ids and answers a redelivery with the existing
`PeerHandle`. CB-117's orphan reap stays as the backstop.
6. **Broker down — local unaffected, remote fails fast.** The routing fork (§3.2) means same-host
traffic never touches the broker; that is now a written promise. A send to a remote agent while
the broker is unreachable **fails immediately** with a clear error — the gateway never buffers
on the broker's behalf (it stays soft-state, so a crash cannot lose messages it claimed to
deliver). Gateways auto-reconnect; remote hosts read as unknown in the roster meanwhile. Broker
HA is a later ops choice, not a design requirement.
7. **Turn state — split by where the signals are.** The *worker's* gateway owns the turn record
(turnId minting, ask coalescing, STALE_TURN, the completion/failure fallbacks, CB-516 abandon):
every input to those decisions — pane status, injection, teardown — is local to it. The
*sender's* gateway owns only the waiter. The two are stitched by terminal-outcome envelope
kinds (`REPLY` / `FAILED` / `ABANDONED`) published to the sender's inbox: a worker dying on B
fails A's waiter fast because gateway B sees the death synchronously and says so.
**`ABANDONED` is belt-and-braces over the waiter's own timeout and roster expiry, never a
replacement** — the case where the waiter hangs longest is gateway B itself dying, which is
exactly when B can publish nothing.
8. **Dual ack model + spawn idempotence by construction.** Forward path (a brief into a worker):
ack **before** the inject — at-most-once, duplicates structurally impossible; the loss window
is closed by an `INJECTED` confirmation published after the inject lands (no `INJECTED` within
a bound = loud fast failure at the sender, not a silent send-timeout). Reply/pull path keeps
ack-after-drain — a duplicate reply is benign, deduped by `msgId`. Spawn: the requester mints
**spawn id = the new worker's gid**; the target checks it against the **live pane registry**,
and the gid is **stored in the herdr pane itself** (label/env, readable back), so a restarted
gateway rebuilds gid↔pane from herdr and the check survives restarts with *no persisted
ledger* — this storage point is the load-bearing detail of the no-ledger position. An
**in-flight reservation set**, entered before the launcher call, absorbs a redelivery arriving
while the first spawn is still inside CB-306's readiness gate; a crash mid-spawn leaves a
half-built pane, which is exactly what CB-117 reaps.
9. **Publish is enforced, not fire-and-forget.** Publisher confirms + the `mandatory` flag + a
return listener, on a **publish channel separate from the consume/ack channel** — synchronous
confirms on the single shared channel would hold its lock across a broker round trip and
serialize acks fleet-wide. Ordering caveat: a *return* (unroutable) arrives **before** the
confirm, so "confirmed" ≠ "routed"; the sender checks the returned-set at confirm time.
`mandatory` is false only for `BROADCAST`, where an empty group is legal silence.
10. **Queue lifecycle is session lifecycle.** `bridge_stop`/reap deletes the worker's inbox queue
(its `broadcast.*` bindings die with it — no broadcasts to the dead); `x-expires` collects
queues orphaned by a crashed gateway (long for main/orchestrator inboxes, short for workers).
Queue names carry a version suffix (`.v2`): AMQP refuses to redeclare an existing durable
queue with new arguments (`PRECONDITION_FAILED` — a crash loop on an in-place upgrade from
v1.0.0), and the suffix keeps old sender-keyed and new recipient-keyed queues apart during
the keying migration (wiki 10 §3 footnote).
## 8. Still open
- **Directory ground-truth:** pure soft-state presence (heartbeats) vs. also treating broker queue
existence as authoritative. Lean soft-state to preserve the persistence boundary; revisit if
split-brain roster views cause mis-routing.
- **Global id scheme:** `<host>/<paneId>` (human-legible, leaks host) vs. opaque UUID (clean, needs
the directory to resolve host). Probably UUID in the protocol, host as directory metadata.
- **Gateway discovery:** how gateways find the broker and each other (static config vs. discovery).
- **Trust model shape:** per-host shared secret vs. mTLS on the broker vs. a capability token per
control action — ties into CB-401 Stage-C.
- **Failure semantics:** a host/gateway dies mid-turn — how the federated roster reaps it (missed
heartbeat) and whether in-flight primary-bound messages survive (broker durability = yes).
- **Control authorization — THE GATE ON U4.** Signing (§7.1) settles *who sent it*; authorization
is *who may do what*. **Cross-host spawn must not land before the minimal version exists**: a
per-host allowlist in `bridged.yaml` — beside the peer public keys — of gateway ids permitted to
publish control to this host, checked against the verified signature. A few lines of config and
check; without them, any principal holding broker credentials can start processes on every host
in the fleet.
- **Key distribution & rotation:** static config (host → public key in each `bridged.yaml`) is
fine at the current 2–3 host scale; rotation is manual. A refinement, not a blocker.
- **Gateway death mid-turn:** the roster reaps it by missed heartbeat, and in-flight primary-bound
messages survive by broker durability; still open is reconciling *worker* state when the dead
gateway's host comes back (orphaned panes vs. still-valid sessions).
*(Resolved and moved up: the global id scheme — an opaque UUID minted by the spawn requester as
the spawn id, host carried as roster metadata; §7.8.)*
## 9. Implementation order (each step verifiable single-host)
1. **As-built fixes, independent of CB-308** (v1.0.x tickets): `basicQos` prefetch on the AMQP
consumer (today the queue drains into gateway heap, so any cap would guard an empty queue);
publisher confirms + `mandatory` (§7.9); the `drainReplies` javadoc that claims "the ack is
local" — false for the AMQP adapter.
2. Envelope + signing (wiki 10 §2.1) — testable against the single-host broker.
3. Recipient-keyed queue migration (`.v2` names, drain-by-`from`, `ReplyPushLoop` rekeyed).
4. Global id + queue lifecycle (§7.8, §7.10).
5. Roster: host-level heartbeat + signed presence; then the routing fork (§3.2).
6. U2 cross-host with the terminal-outcome kinds (§7.4, §7.7).
7. U4 cross-host spawn — **gated on the control allowlist (§8)**.
8. U8 broadcast **last** — it is the feature that punishes an unfinished queue lifecycle.
+119
View File
@@ -0,0 +1,119 @@
# v1.0.0 — One leader, one host, complete
This is the first release of **`bridged`**.
`bridged` lets one main Claude Code session (the **leader**, on your Pro/Max subscription)
run a team of **workers** — extra Claude Code sessions on a cheaper or local model, and
non-Claude agents too. The leader's own session is never touched: it stays on subscription,
with a clean environment.
**This release finishes a full scope — it does not stop halfway.** The scope is *one leader
on one machine*, running many workers. Everything that setup needs is now built, tested, and
used daily: who-is-who, the subscription line, messages in both directions, worker start and
stop, security, and monitoring. Nothing on the single-machine path is left as a known gap.
This is also how the code is shaped: `PrimaryRegistry` holds exactly **one** leader. Running
across many machines is the next big step (see *What comes next* below) — not a missing piece
of this one.
## One gateway for all messages
- **Everyone talks through the same door.** The leader and every worker connect to the same
MCP server and use only its tools: `bridge_whoami` · `bridge_profiles` · `bridge_spawn` ·
`bridge_list` · `bridge_status` · `bridge_send` · `bridge_reply` · `bridge_ask` ·
`bridge_poll` · `bridge_ack` · `bridge_stop`.
- **You are who your connection says you are.** The bridge finds out who is calling from the
connection itself, never from a name the caller sends. So a worker cannot pretend to be
someone else, and `bridge_whoami` tells each agent its own role — no guessing.
- **The subscription line cannot be crossed.** Only a spawned worker gets
`ANTHROPIC_BASE_URL`; the leader never does. Each worker profile has a list of allowed
model hosts, checked before anything starts.
- **Messages wait for the right moment.** The bridge sends one message per turn, only when
the other side is ready — no hammering a busy agent.
## Worker lifecycle
- **Start → work → stop.** Each worker gets its own git worktree (its own copy of the repo)
with the same setup as the leader — `CLAUDE.md`, skills, hooks — so it commits on its own
branch and opens its own PR. Ready-made playbooks ship in the repo:
`.claude/skills/implementer` and `.claude/skills/reviewer`.
- **Fast failure, not a silent hang.** Starting a worker waits until it is really connected.
If it never connects, you get a clear error (`PeerUnreachableException`) instead of a stuck
send.
- **Workers don't live forever.** Idle workers are cleaned up (`idle_ttl`), long sessions have
a turn limit (`context_cap`), shutdown drains work first, and workers left behind by an old
daemon are found and removed at startup.
- **Predictable placement.** Each worker gets its own tab in a shared worker space, in the
same order every time.
## No reply gets lost
MCP only lets the client call the server, so the bridge could push to a worker but the leader
had to ask for its replies. That gap is now closed on a single machine:
- **Replies are kept, never dropped.** If a reply arrives and nobody is waiting, the bridge
holds it until the leader picks it up.
- **Replies can survive a restart.** With a broker (LavinMQ or RabbitMQ) set up, held replies
live on the broker, so a daemon restart does not lose them — they come back, and repeats are
filtered out by `msgId`. No `broker:` in the config → replies are held in memory instead.
- **The leader gets a tap on the shoulder.** When a reply lands, the bridge nudges the
leader's own pane — only when the leader is free, and only a few times. If the leader is on
another machine, this quietly falls back to pick-up mode; the reply still waits.
- **Workers can ask questions.** With `bridge_ask`, a worker can pause mid-task, ask the
leader something, and continue the *same* task with the answer.
## More than one kind of worker
Workers are started through a small plug-in interface (`PeerLauncher`). Two plug-ins ship:
one for Claude Code and one for **opencode** (tested live against opencode 1.18.5). The
opencode one proves the interface is neutral — it uses nothing Claude-specific. Each worker
only sees the tools its own launcher gives it.
## Security & operations
- **Auth.** Default is `loopback-trust`: only same-machine callers are trusted. Or set a
bearer `token`. Unknown callers count as `ANONYMOUS` — nobody is trusted by accident. If
the config would expose the daemon to the network without a token, it **refuses to start**.
Workers never need the token, so turning auth on cannot lock them out.
- **Rules + audit log.** The role rules are checked on both doors (REST and MCP). A worker
may only act as itself. The audit log is JSON and never contains message text.
- **Monitoring.** `/healthz` for liveness, `/metrics` for Prometheus — no extra libraries.
- **Runs as a service.** launchd (macOS) and systemd (Linux) files are included. If herdr
isn't up yet at boot, the daemon waits up to 30 seconds and then runs in a reduced mode
instead of crash-looping.
- **CI.** Every push builds and tests on the Gitea runner, including the broker test against
a real broker. Only the live-herdr test stays local (`-Pcontract`).
## What you need
Java 25 · Maven · **herdr 0.8.0 (protocol 19)** · optionally LavinMQ or RabbitMQ for
restart-proof replies · macOS (launchd) or Linux (systemd). For TLS, put a reverse proxy in
front — the daemon does not do TLS itself, by design.
## Tested
`mvn clean install` is green at `84081b2`: **399 tests**, coverage **75.6%** of instructions /
**64.5%** of branches. Live end-to-end runs under `e2e/`: a worker asking the leader a
question, one leader running several workers at once on a bug hunt, and a 30-turn
back-and-forth conversation. The bridge is used on itself — worker-run code reviews have led
to real committed fixes in this repo.
## What comes next
The single-machine story is done. The next big step stretches the same rules across machines:
- **Many machines (CB-308).** A leader on machine A with workers on machines B and C — built
on this release's broker layer, with one gateway per machine. The design is written
(`docs/CB-308-Multi-Host-Federation.md`); the open question is trust between machines.
- **Many leaders.** Today the bridge holds exactly one leader. The next idea is a small
council — for example one Claude and one Codex — that can discuss a problem together,
compare answers, and agree on a decision before the work is handed to workers. This needs
leader-to-leader messages and a simple way to settle disagreement, neither of which exists
yet.
- **Third-party launcher plug-ins.** Loading launcher plug-ins from outside the project needs
a trust model first, because a launcher runs with daemon rights and can hand secrets to
workers.
One thing we chose **not** to build, so it isn't read as a gap: the originally planned heavy
message envelope (CB-201). Connection identity already routes every reply to the right place,
so only a small `QUESTION` message kind and a `turn_id` were added.
+34
View File
@@ -0,0 +1,34 @@
{
"$schema": "https://opencode.ai/config.json",
"instructions": [
"CLAUDE.md"
],
"mcp": {
"bridged": {
"type": "remote",
"url": "http://127.0.0.1:8765/mcp",
"enabled": true
},
"context7": {
"type": "remote",
"url": "https://ct7.ltms.dev/mcp",
"enabled": true,
"headers": {
"Authorization": "Bearer {file:.secrets/context7-token}"
}
},
"gitea": {
"type": "local",
"command": [
"gitea-mcp",
"-t",
"stdio"
],
"enabled": true,
"environment": {
"GITEA_ACCESS_TOKEN": "{file:.secrets/gitea-token}",
"GITEA_HOST": "{file:.secrets/gitea-host}"
}
}
}
}
+18
View File
@@ -0,0 +1,18 @@
{
"name": "claude-bridge",
"description": "Make a project bridge-ready: mount the bridged MCP gateway and set up standard Claude Code settings so this session can orchestrate a fleet of delegated workers. Ships no credentials.",
"version": "0.1.0",
"author": {
"name": "LTMS"
},
"homepage": "https://git.ltms.dev/lms/claude-bridge",
"repository": "https://git.ltms.dev/lms/claude-bridge",
"license": "MIT",
"keywords": [
"mcp",
"orchestration",
"multi-agent",
"delegation",
"codex"
]
}
+8
View File
@@ -0,0 +1,8 @@
{
"mcpServers": {
"bridged": {
"type": "http",
"url": "http://127.0.0.1:8765/mcp"
}
}
}
+54
View File
@@ -0,0 +1,54 @@
# claude-bridge (Claude Code plugin)
Makes a project **bridge-ready**: mounts the `bridged` MCP gateway and applies standard Claude Code
settings, so the session can orchestrate a fleet of delegated workers.
**This plugin ships no credentials.** Every secret is referenced by environment-variable *name*;
the values stay with the user. Nothing the plugin writes is unsafe to commit.
## What it is not
The plugin is the **client-side setup**, not the bridge. `bridged` is a separate daemon and `herdr`
is a separate PTY multiplexer, each with its own lifecycle and install. The plugin mounts an
already-running daemon and tells you what is missing when one isn't there — it deliberately does
not try to install system services on your behalf.
## Install
```shell
/plugin marketplace add ltms/claude-bridge
/plugin install claude-bridge@claude-bridge
```
Then, in the project you want to onboard:
```shell
/claude-bridge:setup
```
## What you get
| Component | Effect |
|---|---|
| `.mcp.json` | mounts `bridged` at `http://127.0.0.1:8765/mcp` for any session with the plugin enabled |
| `skills/setup` | `/claude-bridge:setup` — preflight, project settings, credential guidance, and verification |
Because the plugin carries its own `.mcp.json`, an installed plugin needs no project-level MCP
file at all. The setup skill writes one only when you want the mount to work *without* the plugin —
for teammates who haven't installed it, or for CI.
## Verifying a setup
The setup skill ends by requiring a **real spawn**, not a health check. `/healthz` only reports
that the daemon can reach herdr; a protocol mismatch between the daemon's adapter and the herdr
binary leaves health green while every spawn fails. Only a spawn that reaches `ready` proves the
fleet.
## Local development
```shell
claude --plugin-dir ./plugin
claude plugin validate ./plugin
```
`/reload-plugins` picks up edits without restarting the session.
+227
View File
@@ -0,0 +1,227 @@
---
name: setup
description: Make the current project bridge-ready — check the prerequisites, mount the bridged MCP gateway into the project's .mcp.json, apply standard Claude Code settings, and verify this session resolves as the primary. Writes no credentials. Load this when asked to set up, install, configure, or onboard a project onto claude-bridge, or when bridge_* tools are expected but absent.
---
# Bridge setup — make this project bridge-ready
This skill configures **the project you are currently in** so that this Claude Code session can
orchestrate a fleet of delegated workers through `bridged`.
**It writes no credentials, ever.** Every secret is referenced by environment-variable *name*, and
the user exports the value themselves. Nothing this skill creates is unsafe to commit. If you are
ever about to write a token, key, or password into a file, you have misread this skill — stop.
Work through the steps in order. Each one has a check; **report what actually happened**, including
failures. A setup that half-worked and was reported as done is worse than one that failed loudly.
## 0. Establish where you are
```bash
pwd
git rev-parse --show-toplevel 2>/dev/null || echo "(not a git repo)"
ls -a | head -30
```
Everything below is written into **this** project root. If the user meant a different directory,
confirm before writing anything.
## 1. Preflight — what must already exist
The bridge is three moving parts, and the plugin is only one of them. Check all of it before
changing any file, so you can tell the user the whole story at once instead of failing one step at
a time.
```bash
command -v herdr && herdr --version 2>&1 | head -1 || echo "MISSING: herdr"
command -v ccs && ccs version 2>&1 | head -1 || echo "MISSING: ccs (needed for worker profiles)"
command -v codex && codex --version 2>&1 | head -1 || echo "absent: codex (optional)"
curl -s -m 5 http://127.0.0.1:8765/healthz || echo "MISSING: bridged daemon is not reachable"
```
A healthy daemon answers with its status **and the herdr protocol it negotiated**:
```json
{"status":"ok","herdr":{"version":"0.8.0","protocol":19}}
```
| Missing | What to tell the user |
|---|---|
| `herdr` | The PTY multiplexer that owns worker terminals. Install it first; nothing else works without it. |
| `bridged` | The daemon. It is a separate service, not part of this plugin — the plugin only *mounts* it. Point the user at the project's own install instructions. |
| `ccs` | Only needed to launch worker profiles. The bridge itself will still start. |
| `codex` | Optional. Only needed if this fleet will run Codex peers. |
**Do not attempt to install these yourself.** They are system services with their own lifecycles;
guessing at an install is how you end up with two daemons on one socket. Report what is missing and
let the user install it.
## 2. Mount the bridge MCP — merge, never overwrite
The project's `.mcp.json` may already declare servers. **Read it first and merge**; clobbering
someone's existing MCP config is not a recoverable mistake.
```bash
cat .mcp.json 2>/dev/null || echo "(no .mcp.json yet)"
```
The entry to add, exactly:
```json
{
"mcpServers": {
"bridged": {
"type": "http",
"url": "http://127.0.0.1:8765/mcp"
}
}
}
```
If `.mcp.json` already exists, add only the `bridged` key and leave every other server untouched.
If a `bridged` entry is already there with a different URL, **ask** rather than assuming yours is
right — a non-default port usually means a deliberate second daemon.
> **If this plugin is installed, you can skip this step entirely.** The plugin ships its own
> `.mcp.json`, so `bridged` is already mounted for any session with the plugin enabled. Write the
> project-level file only when the user wants the mount to work *without* the plugin — for
> teammates who have not installed it, or for CI.
**Before writing it, settle whether `.mcp.json` is committed here:**
```bash
git ls-files --error-unmatch .mcp.json 2>/dev/null && echo "TRACKED" || echo "untracked"
```
A tracked `.mcp.json` is inherited by every checkout of this repo — including git worktrees the
bridge provisions for workers. Servers bound to *your* machine (an IDE index, a local language
server) will then be mounted by workers too, and every path they return points into **your**
checkout rather than the worker's. That failure is silent and expensive: it has produced a worker
that made all of its edits in the wrong tree while its builds passed, because it was building the
tree it was not editing. Keep machine-local servers out of a tracked `.mcp.json`, or keep the file
untracked.
## 3. Standard project settings
Create or merge `.claude/settings.json`. These are defaults, not requirements — keep anything the
project already set.
```json
{
"$schema": "https://json.schemastore.org/claude-code-settings.json",
"permissions": {
"allow": [
"mcp__bridged__bridge_whoami",
"mcp__bridged__bridge_list",
"mcp__bridged__bridge_status",
"mcp__bridged__bridge_profiles",
"mcp__bridged__bridge_poll"
]
}
}
```
Only the **read-only** bridge verbs are pre-allowed. `bridge_spawn`, `bridge_send`, and
`bridge_stop` start processes, deliver work, and tear down terminals — those stay behind a prompt
on purpose. Do not "helpfully" add them.
Never write `settings.local.json` on the user's behalf; that file is personal and usually
gitignored.
## 4. Credentials — by reference only
The bridge takes every secret from the **environment**, and the daemon's config names the variable
rather than holding the value. Your job is to tell the user which variables to export, not to
collect or store them.
| Variable | Needed for | Notes |
|---|---|---|
| `BRIDGED_WORKER_TOKEN` | authenticating a worker to the daemon | only when the daemon is configured with `tokenEnv` |
| `GITEA_TOKEN` / equivalent | letting a worker open its own PR | **minimal `write:repository` scope** — see below |
| `GITEA_HOST` | the forge base URL | no secret; safe anywhere |
Two rules to state plainly to the user:
- **The PR token must not be able to merge.** A worker opens a PR; the primary is the gate. A token
that can merge makes the gate decorative. Mint a narrow, repo-scoped token — never reuse a
personal admin token.
- **Never set `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN`** in this project, this shell, or any
settings file. The primary stays on subscription; only the daemon moves a *worker* off it, at
spawn. Mounting the bridge must never move a session across that boundary — if setup appears to
need this, something is wrong and you should stop and say so.
Write none of these into any file. Show the user the `export` lines to run themselves.
## 5. Verify — and do not trust a green health check
Reconnect MCP if needed (`/mcp`), then confirm the tools are live and this session is the primary:
```
bridge_whoami
```
- `{"role":"primary"}` — correct, you are done with this step.
- `{"role":"worker", …}` — **this is the trap.** If the primary runs inside a herdr pane, the
daemon resolves it to a terminal and classifies it as a worker, refusing `spawn`/`send`/`stop`:
every verb an orchestrator exists to call. It is **self-locking**, because the daemon can only
*learn* the primary's terminal from those same refused calls. The only way out is an
operator-set pin in the daemon's config:
```yaml
primary:
terminal: term_xxxxxxxxxxxx # the terminalId bridge_whoami just reported
```
The daemon reads this **at boot**, so it needs a restart. Re-pin whenever the primary moves
panes — a stale pin fails exactly as silently as no pin.
Then prove the fleet actually works, with a real spawn:
```
bridge_profiles → the configured backends
bridge_spawn{profile: "<one of them>"} → must reach state "ready"
bridge_stop{paneId: "<from spawn>"}
```
**`/healthz` reporting `ok` is not evidence that spawning works.** It reports that the daemon can
reach herdr — nothing more. A version mismatch between the daemon's adapter and the herdr binary
leaves health green while every single spawn fails. Only a real spawn proves the fleet. Do this
even when everything above looked fine.
## 6. Optional — Codex parity
Only if the user wants Codex and Claude Code to share instructions, skills, and MCP config:
```bash
npm install -g ai-config-sync-manager
ai-config-sync connect
ai-config-sync status # compare both hosts
ai-config-sync sync --dry-run # preview — always look before applying
ai-config-sync sync --apply
```
It maps `~/.claude/CLAUDE.md` ↔ `~/.codex/AGENTS.md`, `~/.claude/skills/` ↔ `~/.codex/skills/`, and
Claude's MCP servers ↔ `[mcp_servers.*]` in `~/.codex/config.toml`.
**Raise the boundary before running it.** That sync is *user-level* and bidirectional, while the
bridge deliberately isolates each worker's tool surface (§2). Syncing your MCP servers into
`~/.codex/config.toml` gives every Codex session your machine-local servers — the same
wrong-tree failure as §2, in a different runtime. Use the sync for the two CLIs *you* drive
interactively; leave anything the bridge spawns isolated. Always `--dry-run` first.
## 7. Report
State plainly:
```
prereqs: herdr <version> · bridged <protocol> · ccs <version> · codex <version|absent>
written: <files created or merged, or "none">
role: <bridge_whoami result — and the pin, if one was needed>
spawn: <real spawn result: profile, state reached, torn down>
env: <variables the USER still needs to export — names only, never values>
skipped: <anything not done, and why>
```
Never report a step as done that you did not verify. If the daemon was unreachable, say so and stop
— the remaining steps cannot be checked, and guessing at them is how a broken setup gets called
finished.
+1 -1
Submodule wiki updated: 0c896eb49b...8c63db5da6