Fix two handler-level bugs found in PR #21:
- Only PRIMARY callers may update PrimaryRegistry.record (the legacy singleton 'primary'
fallback for no-delegation inbox nudges). An architect SEND previously recorded its terminal
as the fallback; the per-target delegation map does not cure the singleton. New
BridgeMcp.recordPrimarySingleton uses the resolved role (caller.isPrimary()) — named leads
(PRIMARY) still record, architects never do.
- MessageService.send now opens the rendezvous waiter BEFORE queueing delivery, fixing both the
enqueue-before-open fast-reply race (a fast reply no longer orphans into the inbox) and
callback-failure ordering: a throwing onAccepted (public callback) fails the send cleanly with
no stale waiter and no queued, orphanable message.
Tests: architect SEND vs lead SEND primary-singleton regression; throwing onAccepted leaves no
stale waiter or queued orphan.
Record PrimaryRegistry delegator ownership via a MessageService accepted-delivery
hook (won the session lock + queued delivery), never at bridge_send request time, so
a concurrent sender that times out BUSY cannot steal a live turn's reply routing.
Make Rendezvous.open atomic fail-if-present so a double open trips loudly instead of
replacing the waiter another send is blocked on. Answering a bridge_ask keeps the same
ownership (no rewrite). Adds ownership/rendezvous regression tests.
Add a per-profile subscription: true opt-in that lets a claude-code worker run on
the operator's Claude subscription when there is no off-subscription endpoint for it
(e.g. sonnet on ccs). When set, the launcher injects neither ANTHROPIC_BASE_URL nor
ANTHROPIC_AUTH_TOKEN and skips SubscriptionGuard's base_url requirement for that
profile only, logging a WARN naming the profile. subscription: true alongside a
baseUrl is refused as contradictory. The default (absent/false) keeps today's hard
refusal unchanged; every other profile stays allowlist-checked and SubscriptionGuard
is untouched.
A spawned peer has no human at its pane, so an approval prompt is not a pause —
it is a wedge. The agent stops, looks identical to a legitimate mid-turn wait,
and can never reach its bridge_reply, so the delegation dies silently and the
lead learns nothing until the timeout.
Unconditional rather than a per-profile knob, which is the right call: there is
no configuration under which a bridge-spawned opencode worker WANTS to block on
an approval it has no way to answer. opencode's help calls --auto 'dangerous!',
and that warning is written for a human at a terminal; the blast radius here is
already bounded by the layer above — a worker runs in its own git worktree, on
its own branch, off-subscription, and cannot merge. The lead is the gate.
Reviewed by me rather than fanned out: 24 lines across one method and its two
tests, below the threshold where a reviewer pass pays for itself.
argvWithModel is re-signatured to take composed argv instead of building it,
so the two flag-appenders compose rather than each owning construction.
Two leads now work as peers rather than one primary plus workers. The arc:
CB-530/531 lead identity: `leaders:` names panes, `leadScan:` discovers them by
tab label (LeadTabScanner, TTL-cached, worker spaces excluded).
CB-532 leads can message each other AND be answered. Principal.leader now
carries its terminal, so ownsSession() can be true for a lead; the
"and you must be a worker" conjunct beside it protected nothing.
Retires `primary:` — reply nudges follow the delegating lead, a
binding recorded at bridge_send where both halves are known.
CB-533 ClaudeCodeLauncher passes --model. argv is usually a wrapper
(`ccs <profile>`) that re-exports its own model family, so
ANTHROPIC_MODEL alone was silently overruled.
CB-534 a lead is deliverable. The CB-113 readiness gate only opened for
terminals in WorkerPresence, which only workers ever enter, so every
lead->lead send waited out the ~60s grace and failed having never
been typed. The gate guards a *spawned* peer's boot window; a lead
is never spawned.
CB-535 bridge_list returns `leads` alongside `workers`, with `self` on the
caller's row. An empty worker roster no longer reads as "no peers".
CB-536 CLAUDE.md: lead<->lead is coordinate-only, never sideways delegation.
Propagated byte-identically to wiki/7-Use-Cases.md.
MIXED PROVENANCE — recorded deliberately rather than hidden. This tree also carries
in-progress CB-537 (context separation) authored by the peer lead gpt-sol-5.6 and
its worker: Capability.CONTEXT_RESET, SessionManager.clearAfterTurn, and the
Injector/TurnListener/CompletionResolver/launcher changes around it. That work was
done in this shared working tree rather than a worktree, and is entangled with the
above in BridgedConfig.java, Bridged.java and ClaudeCodeLauncher.java, so neither
lead could stage its own half without sweeping in the other's. Committing the whole
green state is the honest resolution; the peer branches from here.
Note for whoever picks CB-537 up: the design in this commit is SUPERSEDED. Both
leads agreed to replace the global `clearAfterTurn` boolean with per-delivery
policy (inherit|fresh|thread) applied PRE-delivery, because a post-turn reset races
by construction — Injector.onStatus clears awaitingCompletion and dequeues the next
message in the same tick. `fresh` is also a correctness guarantee, so an adapter
without a reset capability must refuse it rather than log a no-op.
mvn clean install: Tests run: 464, Failures: 0, Errors: 0, Skipped: 0. BUILD SUCCESS.
A worker in a provisioned worktree was inheriting the primary's MCP servers by
two independent routes: the repo commits a .mcp.json declaring the IDE servers,
so a fresh checkout mounts them, and the default parity overlay then copied the
primary's own copy over the top.
Those servers are bound to the primary's IntelliJ project, so every path they
hand back points into the primary's checkout. A CB-523 worker made all 59 of its
edits there while running `mvn -f bridged/pom.xml` against its worktree — every
build it ran was of code that did not contain its changes, and it passed. The
worker's own `ls` of the file it had "edited" returned "No such file".
GitWorktrees now neutralizes .mcp.json at provisioning: an explicitly empty
server map, --skip-worktree'd when tracked so it never reads as pending work a
worker might commit. Unconditional, because the overlay was only half the leak.
The bridge itself is unaffected — it reaches a worker through the launcher's
--mcp-config flag, not the project file, so bridge_reply still works.
- BridgedConfig: .mcp.json out of the default parity overlay
- GitWorktrees: isolateToolSurface() on add(), with the rationale in javadoc
- GitWorktreesTest: 4 real-git acceptance tests (2 fail if the call is removed)
- implementer skill: work from $PWD, and quote a green unpiped `mvn clean
install` from the worktree as the acceptance criterion
mvn clean install: Tests run: 392, Failures: 0, Errors: 0 — BUILD SUCCESS
The weighted policy breaks an exact-weight tie on candidate list order
(WeightedRoundRobinPolicy picks the first candidate with a strictly greater
score), and that list comes from CompositePeerLauncher.candidates(), which
iterates profileConfigs. Both that map and BridgedConfig.workerProfiles() were
built with Map.copyOf, whose iteration order is salted per JVM run — so the
"in definition order" contract candidates() documents was not held.
Two consequences. In production, a config with equal weights (ollama 0.5 /
gx10 0.5) placed its first worker on a profile chosen at random on every daemon
restart. In the suite, CompositePeerLauncherTest.failoverRetriesNextCandidate-
WhenProfileIsUnreachable failed roughly one run in four, because whether
profile "a" was tried first depended on the salt.
Preserve definition order at every layer: unmodifiable LinkedHashMap for
workerProfiles(), profileConfigs, and byProfile (which also feeds the
user-visible bridge_profiles listing). The tests build profile maps with an
ordered helper rather than Map.of, which is salted for the same reason.
Guarded by a pair of tests declaring the same two profiles in opposite order
and asserting opposite first attempts, so any order-scrambling implementation
must fail one of them. Verified by mutation: reverting profileConfigs to
Map.copyOf fails 8/8 runs (6 caught by the original test, 2 only by the new
reversed-order one); with the fix, 10/10 fresh JVMs pass, 388 tests green.
Caller identity resolved any loopback PID that mapped to a herdr pane as a
WORKER, and PaneLocator scans every pane -- not just bridged-spawned ones. A
primary running inside a herdr pane therefore classified itself as a worker and
was refused SPAWN/SEND/STOP, i.e. every orchestration verb it exists to call.
The failure is self-locking: PrimaryRegistry only learns the primary's terminal
from bridge_send/bridge_spawn, the exact calls being refused, so the learned
value can never bootstrap. Only an operator-set pin breaks the cycle.
CallerResolver now consults primary.terminal from config *before* the pane
lookup. Deliberately the pinned value only, never the learned one -- the learned
terminal is populated by the callers this method is itself classifying, so
trusting it would be circular. Config is operator input, never network input, so
this widens no attack surface; bridge_whoami and the authz gate still share one
resolution.
Fixing that exposed a second, older bug. BridgeMcp's context extractor forwards
the caller's terminal into markPresent on every MCP call, documented as "no-op
for the primary (null terminal)". WorkerPresence.markPresent honours that, but
PresenceBridge overrides it and forwards the same null into SessionManager.
onReady -> transitionByTerminal -> findByTerminal, which called
terminalId.equals(...) unguarded. It only reached the scan once the registry was
non-empty, so the primary's first spawn succeeded and every later call NPE'd
with an HTTP 500 -- and it would have fired for ANY primary not living in a
herdr pane, pinned or not.
findByTerminal is now total. That covers onReady, onDelivered, onTurnComplete
and onTurnFailed at once; a null id could never match a registered session
anyway, so "no match" is the honest answer rather than taking down an unrelated
tool call.
Also drops two dead pass-throughs on CallerResolver (cwdForPid, tokenMode) that
IDE inspections flagged -- callers use ConnectionIdentity and BridgedConfig.Auth
directly.
The example config now states that primary.terminal is REQUIRED, not just a
push-loop optimisation, when the primary shares a herdr pane.
mvn clean install: 360 tests, 0 failures. Verified live: daemon restarted on
this jar, bridge_whoami reports primary, and four concurrent worktree spawns --
the exact shape that NPE'd -- now all succeed.
A primary running INSIDE a herdr pane was resolved as a worker by the
pane-match rule and refused every orchestration tool — the exact lockout
bridge_whoami surfaced on this deployment. The CB-307 primary.terminal pin
always claimed to replace connection-derived identity but only fed the push
loop; it now short-circuits CallerResolver ahead of the pane→worker rule
(the pane mapping is as unforgeable as a worker's, so no credential needed,
even in token mode). bridged.example.yaml documents the block.
Also guard the presence bridge against the primary's null terminal: the MCP
context extractor marks presence on every request, and the first genuine
primary contact NPEd into the SPAWNING→READY transition.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
herdr 0.8.0 redesigned the agent API out from under the daemon: agent.start
now launches a supported kind INTO an existing pane, env/cwd move to pane
creation (tab.create / pane.split — the subscription-boundary seam now),
agent.send is replaced by agent.prompt (self-submitting) plus agent.send_keys
for the Enter nudge, and terminal ids are no longer valid agent.* targets.
- AgentControl: start(name, kind, args, paneId); prompt/send_keys delivery;
cached terminal→pane target translation (invalidated on agent_not_found).
- WorkspaceControl: tab.create carries cwd+env; pane.split for legacy placement.
- HerdrPeerLauncher: the seed pane IS the worker pane (no drop step); retry
agent.start while the seed shell boots (agent_pane_busy).
- FakeHerdr and the test suite model protocol 19 (unique seed panes, required
kind/pane_id, prompt-based delivery); contract tests probe the seed shell
instead of arbitrary-command agents, which protocol 19 removed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
The communication rules lived only in two opt-in skills, so nothing
always-on told the primary how to orchestrate and nothing guaranteed a
worker loaded its playbook. Move protocol and policy into CLAUDE.md,
which a worker inherits for free (its worktree is a checkout of this
repo), and leave the skills as pure per-job procedure.
bridge_whoami closes the load-bearing gap: every tool already consumed
the caller identity ConnectionIdentity resolves from the connection, but
none reported it, so an agent had to infer its own role from side
channels the daemon does not control. Guessing fails asymmetrically — a
primary acting as a worker is refused by the authz gate and learns at
once, while a worker acting as the primary ends its turn without
bridge_reply and the sender silently receives nothing. The tool reuses
the same Principal the gate is built on, so the two cannot disagree; the
primary gets role only (handing it a sessionId it does not own would
invite the forged reply Authz refuses), and a worker missing from the
registry still gets role + sessionId rather than 'unknown'.
The CLAUDE.md block is written to be copied as-is into any project that
mounts the bridge: repo-local details (Authz paths, the .mcp.json/wiki
exclusions, the skill names) moved below it into a project addendum, and
every role-inference fallback is stated one-way — the mount-name signal
only holds for mcp__bridge__* (the launcher fixes it), not for the
primary's mount, which each project names itself. The wiki carries the
block verbatim as the template, with a sync check.
Because this repo IS the bridge, that block is shipped surface, not
documentation: the addendum adds a mandatory checklist mapping each part
of the code to the part of the prompt it can invalidate.
Also: delegate-by-default policy for the primary — the test is not 'could
I do this faster myself' but 'can I write a brief good enough for a
worker'.
mvn clean install: 356 tests green (353 + 3 for whoami); ide_diagnostics
clean on both changed files.
Tearing a worker down left the send that was waiting on it stranded. The
rendezvous waiter stayed open, so a blocking bridge_send kept blocking and an
async one kept reporting PENDING until ASYNC_TIMEOUT_MS — thirty minutes —
even though the worker provably no longer existed and the delegation could
never complete.
Observed repeatedly this session: a task sitting at
{"phase":"pending","detail":"worker unknown"} long after I had personally
deleted the pane. The maddening part is that poll() already HAD the evidence —
it calls liveStatus() to build that detail string, gets back "unknown", and
reports PENDING anyway. It also never reached /metrics under any outcome, so a
stalled delegation was invisible to both the task view and the dashboard. The
only way I ever diagnosed one was reading the worker's pane over the herdr
socket by hand.
Fixed at the choke point rather than by guessing from status strings:
SessionManager.release() is the single path every teardown funnels through
(REST stop, MCP stop, idle-TTL reaper, recycle, shutdown drain), so it now
notifies a release listener with the terminalId, and Bridged.main wires that to
MessageService.abandon(). Abandon fails the open waiter with a real reason.
Deliberately NOT done by inferring "gone" from liveStatus(): that method
collapses a vanished pane, a wedged worker and a herdr hiccup into the same
"unknown" string, so acting on it would fail live delegations during a
transient blip. An explicit lifecycle signal cannot be ambiguous.
Notified before launcher.stop() so a blocked caller fails fast, and wrapped so
a listener failure can never prevent the teardown it is reacting to.
Resolving as a failure rather than letting it time out also means the outcome
is counted — a torn-down delegation now shows up as
sends_total{outcome="failed"} instead of nothing at all.
353 tests (was 346): abandon fails a waiting send / is a no-op with no waiter /
never clobbers a send the worker already answered, an abandoned async task
polls FAILED rather than PENDING, release notifies with the right terminal,
releasing an unknown pane notifies nobody, and a throwing listener does not
block teardown.
Verified live on the running daemon, reproducing the original scenario:
async send -> {"phase":"pending","detail":"worker working"}
DELETE the worker -> 204
poll -> {"phase":"failed","detail":"the worker session was
released before it replied"}
/metrics -> bridged_sends_total{outcome="failed"} 1
Previously that poll returned PENDING for thirty minutes and the counter stayed
empty.
NOT addressed here, and worth deciding separately: a worker that simply never
replies (rather than being released) still rides out the full 30-minute
ASYNC_TIMEOUT_MS. That is a policy question about long-running delegations, not
a correctness bug.
CB-505 claimed authorization is "enforced on both entry paths". It is — but
only REST was ever tested. Coverage showed BridgeMcp.deny(), principal(),
callerTerminal(), worktreeRequest() and every tool-registration lambda at ZERO
executed lines: no test had ever constructed a BridgeMcp, because the existing
BridgeMcpTest calls only the static handler methods. So the MCP half of the
security control had ten REST tests' worth of nothing behind it.
An unexercised security control is a claim, not a control.
Made testable by separating policy from plumbing rather than by reaching for a
mocking library the project does not use:
- denyFor(Principal, Action, target) is the decision — testable directly.
- deny(exchange, ...) shrinks to pulling the caller out of the SDK exchange.
- principalFrom(role, terminal, pid) extracts identity reconstruction from
McpSyncServerExchange, an SDK type with no fake available.
Moved the `authz == null` enforcement switch OUT of the exchange-facing wrapper
and INTO denyFor. Found by a failing test: as written, any future tool calling
denyFor directly would have silently skipped the gate. The switch now lives with
the decision it governs.
New BridgeMcpAuthzTest constructs a real BridgeMcp — which is why coverage moved
so far, since that also runs the constructor and all the tool wiring — and pins
the table on this path: primary orchestrates, worker cannot; worker replies only
as itself; the primary cannot forge a worker reply; anonymous gets nothing; and
401-shaped vs 403-shaped refusals are counted apart.
Verified as real controls, not decoration: with the gate forced open, 5 of the 9
fail. 335 tests (was 326).
Workers could not run `mvn` or `java`. Every delegated task that asked for a
build came back "mvn is not on PATH", and the worker was right.
Root cause: HerdrPeerLauncher seeded the worker environment with an EMPTY map,
so bridged passed only the vars it explicitly set (OPENCODE_CONFIG, GITEA_TOKEN,
ANTHROPIC_*) and never PATH. herdr merges that map into its own process env, so
a worker inherited whatever PATH the herdr SERVER was started with. On this host
that server (pid 79870, PPID 1) had been up since Jul 4 with a PATH containing
neither the JDK nor Maven. Confirmed on a live worker: its PATH was byte-identical
to herdr's, and the only var bridged had contributed was OPENCODE_CONFIG.
The failure was invisible and non-deterministic: the fleet's capabilities
depended on how a long-lived daemon happened to be launched weeks earlier. There
are three herdr processes on this box with three different PATHs; the one owning
the socket is the one without a toolchain. bridged itself HAD Maven on PATH the
whole time — it just never passed it on.
It also quietly contradicted the project's own principle that "a worker is a
full peer of the primary", and the implementer skill's instruction to build,
commit and open a PR. Every delegation so far has depended on the primary
running the build gate.
Fix: baseEnv(cfg) seeds each worker with the daemon's own PATH, then applies the
profile's new optional env: map. Adapter-specific vars are layered on top and
therefore win — that ordering is load-bearing, not incidental: it stops an env:
entry from overwriting ANTHROPIC_BASE_URL and slipping past SubscriptionGuard,
which is checked against the profile's baseUrl alone. Pinned by a test.
Because the default is now the daemon's PATH, both supervision units set PATH
explicitly — launchd and systemd do not source a login shell, so under CB-504
the daemon (and every worker) would otherwise get a bare /usr/bin:/bin and this
bug would silently return in production.
324 tests (was 321): daemon-PATH propagation, profile env: passthrough including
an explicit PATH override, and the guard-bypass ordering.
Verified live: daemon restarted, worker spawned, and asked to run the tools —
"Apache Maven 3.9.16", "java version 25.0.2". Previously both were absent.
Points an opencode worker at a local vLLM (or llama.cpp / LM Studio / TGI)
instead of opencode's own gateway. opencode has no ANTHROPIC_BASE_URL seam, so
this could not be a config-only change: setting baseUrl on a kind: opencode
profile now makes the launcher emit a custom `provider` block into the
generated opencode.json, using @ai-sdk/openai-compatible.
The provider id comes from the provider half of the model: selector, so one
field drives both the generated declaration and the -m flag and the two cannot
drift apart. A bare model name with a baseUrl set is rejected at spawn with a
message saying how to fix it — silently falling back to the default gateway
would leave a worker talking to the wrong LLM while looking perfectly healthy.
A bare host:port gets /v1 appended (where these servers mount the API); a URL
that already carries a path is used verbatim. tokenEnv, when set, becomes the
provider apiKey; local servers generally ignore it but the AI SDK requires a
non-empty value, so a placeholder is used otherwise.
Two supporting changes:
- writeConfig previously ran only when a bridge MCP url was set. A pinned
endpoint needs the config file too, so it now runs when either applies, and
the mcp/instructions half is emitted conditionally.
- The config is now built with Jackson instead of string concatenation. The
provider block is nested and interpolates operator-supplied values (URL,
model id, api key), so escaping has to be real rather than a hand-rolled
two-character replace.
No guard entry is required even with baseUrl set. SubscriptionGuard exists to
stop a worker borrowing the primary's Anthropic subscription, and an opencode
process has no Anthropic credential path at all — the asymmetry with the Claude
adapter reusing the same field is deliberate and documented at the call site.
Also fixes a brittle assertion in the existing MCP-mount test, which matched
the substring "\"type\": \"remote\"" and broke on Jackson's spacing. It now
parses the generated JSON and asserts on structure; whitespace is the
formatter's business, not the contract's.
318 tests (was 311): 5 new covering provider generation, /v1 normalisation,
path-preserving URLs, the missing-prefix rejection, MCP+provider coexistence,
and that no baseUrl still means no provider block.
Verified live end to end: daemon restarted on this build, worker spawned on the
opencode-local profile, generated config carries baseURL
http://127.0.0.1:8000/v1, and a blocking bridge_send returned
{"reply":"LOCAL-OK","replySource":"reply"} — a structured reply, not the
completion fallback. The worker pane reports
"Build · deepseek-v4-flash local-vllm (bridged)", confirming traffic reached
the local server rather than silently falling back.
POST /workers?worktree=true returned HTTP 500 with a NullPointerException out
of ProcessBuilder.start(): acquireWithWorktree resolved the repo root from
firstNonBlank(requestedCwd, callerCwd), and a plain REST spawn supplies
neither (BridgedApp hardcodes callerCwd=null, "no MCP caller over REST"). Both
null yielded null, putting `git -C null rev-parse --show-toplevel` on the
command line.
Now resolved through launcher.effectiveCwd, the CB-112 chain used everywhere
else (requested -> profile cwd -> caller -> daemon cwd -> "."), which is
documented never to return null. The non-worktree path in this same class
already went through it; only the worktree branch was missed.
Also fixes a second latent bug in the same line: firstNonBlank never consulted
the profile's configured cwd:, so a worktree spawn silently ignored a pinned
per-profile working directory. effectiveCwd honours it.
Removes firstNonBlank, now dead (this was its only call site) — javac ignores
an unused private method but IDE inspections flag it, and CLAUDE.md requires a
clean bill.
Why 311 tests missed it: the null/null case only arises over REST, and
WorktreeSessionManagerTest always passes an explicit cwd. Over MCP callerCwd is
populated from the caller PID, so the feature worked there. This is the third
REST-vs-MCP divergence found this month, after CB-505's path-trusted session id.
The one-line change was implemented by an opencode-free worker over the bridge
in an isolated worktree (branch worker/cb-507-worktree-cwd-npe-3e9c3b-3); the
dead-helper cleanup and the explanatory comment were added on integration.
A regression test is still outstanding and is being delegated separately.
The first cut spliced the timestamp on via a logback pattern:
{"ts":"%d{...}",%replace(%msg){'^\{',''}%n
Logback's variable substitution chokes on the literal braces
("All tokens consumed but was expecting }"), so the encoder failed to
configure. Caught by running the real jar and noticing logback had dumped its
internal status — which it only does when something failed to parse. The build
was green throughout: nothing asserted the audit trail was machine-readable.
AuditLog now emits the complete object including its own ISO-8601 "ts", and the
appender pattern is a bare %msg. Adds AuditLogTest, which parses each emitted
line with Jackson (so a malformed record fails the build) and pins that hostile
ids cannot escape their field to forge a second record.
311 tests green; logback now configures with zero internal errors.
Closes out single-host before the cross-host work. Sequenced BEFORE CB-308
deliberately: federation's own gating concern is the trust model, and it
inherits whatever identity shape lands here.
The finding this stage is built around: bridged had exactly ONE security
control, the loopback bind. ConnectionIdentity resolves a worker from its
connection (unforgeable), but every caller that was not a recognised worker
pane fell through to being treated as the PRIMARY -- the most privileged role
on the bus. Latent today; load-bearing the moment a bind widens.
CB-501 auth:
- Role/Principal/CallerResolver: connection identity first, bearer token
second, ANONYMOUS third. Inverts the old default so absence of identity
means nothing, not everything.
- Worker identity is never token-gated, so enabling auth cannot lock the
fleet out of bridge_reply.
- Constant-time token compare (MessageDigest.isEqual).
- validateAuthExposure(): a non-loopback bind under loopback-trust now
REFUSES TO START. Makes the dangerous config unrepresentable rather than
merely documented.
- TLS terminates at a reverse proxy by design (D3), not in the JVM.
CB-505 authz + audit, enforced on BOTH entry paths:
- The docs describe MCP as "a thin adapter over the REST core"; at code level
it is not. BridgeMcp calls MessageService directly, and /mcp is a raw
servlet on Jetty's context handler that never traverses Javalin's before
filter. Enforcing only at REST would have left /mcp open.
- Load-bearing rule is own-session-only: a worker may reply/ask only as
itself. Structurally true over MCP already; over REST the session id in the
URL path had simply been trusted.
- Audit: JSON lines to a dedicated appender, additivity=false. Never records
message content -- this bus carries source and prompts.
CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer
instead of the specced Micrometer, because this pom already hand-pins
jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against
skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated
dependency CVE gate could not be run (no JetBrains MCP server connected).
Instrumented at MessageService, the single funnel both surfaces share.
CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner.
Needs no contract-exclusion flag -- the pom's default-excludes profile
already sets excludedGroups=contract, so plain `mvn clean install` IS the
mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older).
CB-504 supervision: launchd agent (the real target -- this host is macOS,
there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds.
Ordering directives are advisory, so the actual fix is that startup now waits
up to 30s for the herdr socket and then serves degraded, instead of crashing
into a restart loop on a boot-order race.
Also fixes drift found while surveying:
- bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config
binds via plain Jackson with ignoreUnknown, so uncommenting it would have
been silently dropped and the default kept. Now camelCase, with a test that
loads the shipped example and one that pins every documented knob's
spelling -- no test had ever loaded that file.
- Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay,
gitTokenEnv, gitHostEnv, configDir, primary:).
- README "Next" listed bridge_ask and session lifecycle as upcoming; both
shipped long ago.
- docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre-
implementation" for work already merged.
307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS.
Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check
could not be run -- no JetBrains/intellij-index MCP server is connected this
session. mvn clean install is the only gate that ran.
Introduce the router the core holds when more than one adapter is
configured: one HerdrPeerLauncher per peer kind, dispatched by profile
(spawn/effectiveCwd/parityOverlay), by pane id (stop, via a spawn-time
owner map), and fanned out + combined for the fleet-wide queries
(list dedup by pane id, reap/caps union, profiles union). The ctor
rejects an empty adapter list and a profile two adapters both claim.
Wire it in Bridged.main: partition workerProfiles() by kind (claude-code
is the always-present default adapter; opencode is added when any profile
opts in) and front both with the composite. This lets BridgeMcp and
BridgedApp finally take the PeerLauncher SPI instead of a concrete
ClaudeCodeLauncher — the two (ClaudeCodeLauncher) casts in Bridged are
gone. list() elements are cast to herdr Agent at the point of the
herdr-specific roster view, where that assumption actually lives.
Add BridgedConfig.Worker.isClaudeCode()/isOpenCode() kind predicates
(the wiring uses isOpenCode; both are unit-tested). Drop the long-dead
'rendezvous' constructor param threaded into BridgeMcp and BridgedApp.
10 CompositePeerLauncherTest cases over two real adapters on one
FakeHerdr: profile routing (observed via the started agent's
claude-/opencode- name prefix), default resolution, unknown-profile
and duplicate-profile rejection, caps union, list dedup, reap sum,
and stop teardown. 266 tests green.
A HerdrPeerLauncher subclass for opencode, a provider-agnostic terminal
coding agent. It reuses every line of shared base transport (tab/pane
placement, CB-306 readiness gate, unique naming + CB-117 reap, teardown,
listing, cwd) and diverges only in buildLaunch:
- No subscription boundary: no ANTHROPIC_BASE_URL, no SubscriptionGuard
(the guard is a Claude-private concern, not part of the SPI).
- File-based MCP mount + instructions: writes an ephemeral opencode.json
declaring the bridge as a remote MCP server + a reply-charter file under
instructions, pointed at via OPENCODE_CONFIG (opencode has no inline
--mcp-config / --append-system-prompt).
- Model selected with -m provider/model, not an env var.
- 'opencode' name prefix so reap matches opencode-* panes only.
configRoot is injectable so tests inspect the generated config/charter under
a @TempDir. 10 tests cover config content, model flag, git-token grant,
capabilities, reap predicate, both production ctors, and the readiness gate
(throws PeerUnreachable on timeout, reaps only the worker pane).
Add a `kind` field to BridgedConfig.Worker — "claude-code" (default) or
"opencode" — the discriminator the CompositePeerLauncher will route spawn/reap
by so each adapter drives only its own peer kind. Normalised to lower-case;
blank/absent ⇒ claude-code, so every existing config and call site is
unchanged. argv now defaults to the kind's own binary (claude vs opencode)
rather than always `claude`, so an opencode profile never inherits the Claude
command.
kind is appended at the record tail; a new 14-arg back-compat constructor
(git fields, no kind) keeps the CB-302 call sites working, and the existing
12-arg constructor is untouched. Also drop the never-used Primary(String)
legacy constructor to keep the file warning-clean.
example.yaml documents the key and carries a commented opencode-gemini
profile. Tests cover default/normalisation/argv-defaulting. 245 tests green,
config files 0 IDE problems.
Behaviour-preserving refactor ahead of the second-adapter work. All herdr
transport shared by any peer kind — tab/pane placement, the CB-306
spawn-readiness gate, unique naming + CB-117 orphan reap, teardown, listing,
cwd resolution, and the peer-neutral git-forge grant — moves into a new
abstract HerdrPeerLauncher (Template-Method base). ClaudeCodeLauncher becomes
a final subclass supplying only the two Claude-specific seams: the `claude`
name prefix and buildLaunch(), which encodes the subscription boundary
(ANTHROPIC_BASE_URL + SubscriptionGuard assert, inline --mcp-config and
--append-system-prompt reply charter).
The base owns the injectable clock + sleeper for the readiness gate; the poll
interval is baked into the sleeper, so the vestigial spawnReadyPollMs field is
dropped from the base and from the full testability constructor (the explicit
sleeper already encodes it). The 6-arg and 8-arg production constructors keep
their signatures; three full-ctor test sites drop the now-unused poll argument.
No behaviour change: 242 tests green, both refactored files 0 IDE problems.
Primary verification pass over the delegated push-loop delivery:
- Remove the unused LongSupplier clock threaded into ReplyPushLoop
(timing is the scheduler's; the field was never read) from the
component, Bridged wiring, and both test call sites.
- Collapse the single-statement WAIT_BUSY switch arm (redundant block).
- Drop now-dead test scaffolding: the always-"idle" recordingClient
param and unused AgentStatus/AtomicReference imports.
IDE diagnostics 0/0 on all changed files; mvn clean install green
(242 tests, 0 failures).
Behind the existing ReplyInbox port, add an AMQP-backed adapter selected by a
`broker:` block in config (absent → the in-memory soft-state inbox; present →
AMQP). Mapping is consume-and-hold with deferred manual ack: each target owns a
durable queue `agent.<target>.inbox`; a manual-ack consumer pulls persistent
messages into an in-memory held map (dedup by msgId) but does not ack; peek
returns the snapshot; ack acks the broker delivery-tag and drops it. A crash
before caller-ack leaves messages unacked, so the broker redelivers on
reconnect — genuine durability with the port contract preserved. bridged still
owns no persistence; the broker does.
- msg/AmqpReplyInbox: the adapter (single synchronized channel; recovery
listener clears held on reconnect so fresh delivery-tags repopulate).
- config/BridgedConfig: nullable Broker(uri) record; isConfigured() gates it.
- Bridged.main: select adapter; close the AMQP connection in the ordered
shutdown hook (no-op for the in-memory inbox).
- deps: com.rabbitmq:amqp-client (main); testcontainers rabbitmq/junit-jupiter
(test). Pinned commons-compress 1.27.1 + commons-lang3 3.18.0 to clear the
test-scope CVEs those pull. Production default LavinMQ; RabbitMQ URI-swap.
- tests: BridgedConfigTest broker-selection cases (hermetic); AmqpReplyInbox
contract test (@Tag("contract"), Testcontainers RabbitMQ) proving
publish/peek/ack, msgId dedup, and cross-restart redelivery. Excluded from
the default build so `mvn clean install` stays hermetic (210 green).
Problem: the reverse (worker->primary) path was Rendezvous, a map of LIVE blocking
waiters only. A bridge_reply arriving with no open send hit Rendezvous.complete()
-> no waiter -> returned false -> the reply was silently DISCARDED (worker saw an
error / REST 409). No message-id/dedup/ack existed anywhere.
Stage 1 (no broker, soft-state) behind one port:
- ReplyInbox port + InboxMessage record; InMemoryReplyInbox adapter (per-target
FIFO via LinkedHashMap, dedup by msgId, thread-safe). Soft-state, not persistence.
- MessageService.reply(session, content): resolve an open send, else publish to the
inbox with a minted UUID (was a silent drop). drainReplies(target) = peek + ack.
- BridgeMcp.reply / BridgedApp.replyMessage repointed off bare Rendezvous.resolve
onto messages.reply -> no-waiter is now SUCCESS (queued), not error / 409.
- Drain surface: bridge_poll gains optional target; REST GET /sessions/{id}/replies.
- Rendezvous left untouched. QUESTION path (bridge_ask) NOT queued (interactive,
keeps NO_WAITER); completion/failure fallbacks NOT queued (captured-waiter).
- Bridged.main wires new InMemoryReplyInbox(); no broker: config yet (Stage 2 = AMQP).
Tests: +19 (188 -> 207), 0 failures/0 errors. New InMemoryReplyInboxTest (12) +
MessageService/BridgeMcp/BridgedApp coverage incl. guards proving a QUESTION and a
completion fallback are never queued.
Implemented via delegation to an off-sub gx10 worker in a pre-trusted worktree;
primary-verified (mvn clean install green, 207 tests) and committed by the primary
because the worker's completion replies were lost to the very bug this fixes.
Refs CB-307 (gitea #5), Stage 1 of 2.
The launcher now polls AgentControl.status(paneId) after starting the pane.
It returns the handle only once the worker reports an injectable state
(IDLE/BLOCKED/DONE). If the timeout elapses while still UNKNOWN, the
pane is self-reaped and a PeerUnreachableException is thrown — no orphan
left behind. The gate is disabled when spawnReadyTimeoutMs == 0 (legacy
non-blocking spawn, the default for the 6-arg constructor).
Key changes:
- PeerUnreachableException (new) in dev.ltms.bridged.peer
- BridgedConfig: spawnReadyTimeoutMs (default 20000), spawnReadyPollMs (default 300)
- ClaudeCodeLauncher: 3 constructor overloads:
(a) 6-arg backward-compat: gate disabled (timeout=0)
(b) 8-arg production: gate with config knobs + real clock/sleep
(c) 10-arg testability: full seam (LongSupplier clock + Runnable sleeper)
- waitUntilInjectableOrThrow() loop in spawn(SpawnRequest)
- sleepUninterruptibly() helper for the production sleeper
- BridgeMcp.spawn + BridgedApp.spawnWorker catch PeerUnreachableException
→ clean tool error / 502 response (not an uncaught 500)
- SessionManager.acquire inherently registers nothing on throw (both
worktree and non-worktree paths) — confirmed by new test
Tests:
- ClaudeCodeLauncherTest: 4 new tests
- unknown→idle: returns handle, no pane.close
- always-unknown: throws PeerUnreachableException, pane closed,
clock advanced past timeout
- timeout=0 (6-arg ctor): no agent.get calls, returns handle
- timeout=0 (10-arg ctor): no orphan pane close
- SessionManagerTest: 1 new test
- acquire → PeerUnreachableException: roster remains empty
Total: 188 tests, all pass (no existing test changed semantics)
Name the first-class Claude Code adapter explicitly, per the Peer Launcher SPI:
WorkerService was the de-facto Claude-Code launcher; as an in-tree PeerLauncher impl
it should say so. Pure IDE rename (class + file + WorkerServiceTest) plus stale
Javadoc/comment mentions swept to the new name. No behaviour change.
Gate: IDE diagnostics 0/0 on touched files; mvn clean install BUILD SUCCESS,
MVN_EXIT=0, 183 tests pass. Deferral #1 from issue #3 cleared.
The worker "checkpoint" is commit → push → open its own PR. Push is free over SSH
(same user, same keys); the only incremental grant is PR-create, so the daemon injects
a repo-scoped gitea token into the worker env — opt-in per profile, never mutating
bridged's own environment.
- BridgedConfig.Worker: gitTokenEnv/gitHostEnv fields (opt-in; gitHostEnv defaults to
GITEA_HOST). Backward-compat 12-arg constructor keeps pre-CB-302 call sites + YAML
working. hasGitToken() gates injection.
- WorkerService.spawn: inject GITEA_TOKEN (and paired GITEA_HOST) only when the profile
grants a token AND the host env resolves one. resolveEnv() tolerates unset var names.
- WorkerServiceTest: injection present for a granting profile; absent when not (proving
the gate is config, not a missing env var).
- .claude/skills/implementer/SKILL.md: worktree-aware playbook — confirm the worktree/
branch, implement, commit (never .mcp.json/wiki), push, open PR via gitea REST with
GITEA_TOKEN, hand off the PR URL via bridge_reply. Never merge; workers can't run IDE
diagnostics so never claim IDE-clean.
Whole-project gate: mvn clean install green, 164 tests, 0 failures.