Fix two handler-level bugs found in PR #21:
- Only PRIMARY callers may update PrimaryRegistry.record (the legacy singleton 'primary'
fallback for no-delegation inbox nudges). An architect SEND previously recorded its terminal
as the fallback; the per-target delegation map does not cure the singleton. New
BridgeMcp.recordPrimarySingleton uses the resolved role (caller.isPrimary()) — named leads
(PRIMARY) still record, architects never do.
- MessageService.send now opens the rendezvous waiter BEFORE queueing delivery, fixing both the
enqueue-before-open fast-reply race (a fast reply no longer orphans into the inbox) and
callback-failure ordering: a throwing onAccepted (public callback) fails the send cleanly with
no stale waiter and no queued, orphanable message.
Tests: architect SEND vs lead SEND primary-singleton regression; throwing onAccepted leaves no
stale waiter or queued orphan.
Record PrimaryRegistry delegator ownership via a MessageService accepted-delivery
hook (won the session lock + queued delivery), never at bridge_send request time, so
a concurrent sender that times out BUSY cannot steal a live turn's reply routing.
Make Rendezvous.open atomic fail-if-present so a double open trips loudly instead of
replacing the waiter another send is blocked on. Answering a bridge_ask keeps the same
ownership (no rewrite). Adds ownership/rendezvous regression tests.
Pairs with 244fbd9. The docs half lived on the submodule's own remote, so the
pointer bump is separate by necessity, not by preference.
The substantive part is not the new 'Run a worker on the subscription' entry but
the correction beside it: 'Give workers a toolchain' asserted flatly that an env:
entry cannot repoint a worker past the SubscriptionGuard. CB-539 made that false,
and the wiki went on claiming it until CB-542 closed the hole. The entry now
states the rule and its one exception together.
Verified by the lead in a clean worktree at 5afe8e1 rather than on the worker's
report: Tests run: 474, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS (main
was at 464).
The invariant this lands: there is no configuration in which a worker reaches an
Anthropic endpoint that no guard vetted. Closed at two layers — a fatal, profile-
naming refusal at config load, and a launcher-side strip so it holds for profiles
built in code that never passed validation.
The wiki half is a separate branch on the submodule's own remote; the pointer bump
follows as its own commit on main.
Add a per-profile subscription: true opt-in that lets a claude-code worker run on
the operator's Claude subscription when there is no off-subscription endpoint for it
(e.g. sonnet on ccs). When set, the launcher injects neither ANTHROPIC_BASE_URL nor
ANTHROPIC_AUTH_TOKEN and skips SubscriptionGuard's base_url requirement for that
profile only, logging a WARN naming the profile. subscription: true alongside a
baseUrl is refused as contradictory. The default (absent/false) keeps today's hard
refusal unchanged; every other profile stays allowlist-checked and SubscriptionGuard
is untouched.
Regression I introduced one commit ago, caught by two consecutive spawn failures
and reproduced end-to-end.
Tracking `.autoenv` means git checks it out into every provisioned worktree. A
worktree is a new path, and autoenv authorizes by path, so the file is always
unauthorized there — autoenv prints "[autoenv] Authorize this file? (y/n/d)" and
blocks on `read` (activate.sh:211-222). The pane's shell sits at that prompt, so
`ccs <profile>` never runs, the peer never becomes injectable, and CB-306's
readiness gate closes the pane after 20s. Symptom is a bare "spawn timed out";
nothing names autoenv, which is what made it worth writing down.
Confirmed rather than inferred: the peer lead's cb-537 worker, spawned BEFORE
f8182e4, has no `.autoenv` in its worktree and is still alive; a throwaway
worktree at HEAD reproduces the prompt on entry.
The file is still worth committing — the reasoning in f8182e4 stands, and it is
recoverable from there. What is missing is the other half: GitWorktrees already
neutralizes `.mcp.json` in a provisioned worktree ("worker tool surface is
launcher-mounted only"), and `.autoenv` needs exactly the same treatment for
exactly the same reason — a worker's environment is launcher-supplied, never
repo-supplied. Re-land it with that, tracked as CB-543.
Rejected the quicker fix of exporting AUTOENV_ASSUME_YES: it auto-approves
arbitrary repo-controlled shell code in every spawned worker, which is a worse
trade than one unset variable.
A spawned peer has no human at its pane, so an approval prompt is not a pause —
it is a wedge. The agent stops, looks identical to a legitimate mid-turn wait,
and can never reach its bridge_reply, so the delegation dies silently and the
lead learns nothing until the timeout.
Unconditional rather than a per-profile knob, which is the right call: there is
no configuration under which a bridge-spawned opencode worker WANTS to block on
an approval it has no way to answer. opencode's help calls --auto 'dangerous!',
and that warning is written for a human at a terminal; the blast radius here is
already bounded by the layer above — a worker runs in its own git worktree, on
its own branch, off-subscription, and cannot merge. The lead is the gate.
Reviewed by me rather than fanned out: 24 lines across one method and its two
tests, below the threshold where a reviewer pass pays for itself.
argvWithModel is re-signatured to take composed argv instead of building it,
so the two flag-appenders compose rather than each owning construction.
Moves the submodule pointer 4320c1c -> 0a21b49, matching the docs to the code in
bc13b8e rather than leaving chapter 11 describing a fleet with one lead in it.
7b5381b 11, 7: CB-530..536 — leaders:, leadScan:, lead-to-lead messaging,
lead deliverability, leads in bridge_list; 7's copy of the portable
block re-synced byte-identically
0a21b49 12: where opencode's secrets actually come from, and why {env:...}
is green under Claude Code and broken in a terminal
Both are already pushed to the wiki's own remote, so this pointer resolves for
anyone who clones — the ordering that matters, and the reason the wiki went
first.
This is the pointer only. `wiki/` is a submodule with its own remote and its
content is never committed here; bumping the recorded SHA in its own commit is
how this repo has always recorded a docs update (see "bump wiki to 4320c1c").
This file has always been designed to be committed and says so in its own header;
it simply never was, so every clone of this workspace has been reconstructing it
by hand or duplicating tokens instead.
It holds no secret. It reads `.secrets/` (gitignored, 0600) and exports three
variables, because Claude Code expands `${VAR}` in `.mcp.json` from the *process*
environment and cannot read a file — so without it, CONTEXT7_TOKEN and the gitea
pair must be duplicated as literals in `.claude/settings.local.json`. opencode
needs none of this: it reads `.secrets/` directly via `{file:...}`.
Verified before committing that no value appears in it, only the three names and
the paths they are read from.
Safe in a worktree by construction: a worktree receives tracked files only, so
`.secrets/` is absent there and the whole block is skipped rather than failing.
Workers are fed by the launcher's env instead — which is where the name mismatch
documented in wiki chapter 12 (GITEA_TOKEN vs GITEA_ACCESS_TOKEN) has to be
reconciled.
Two leads now work as peers rather than one primary plus workers. The arc:
CB-530/531 lead identity: `leaders:` names panes, `leadScan:` discovers them by
tab label (LeadTabScanner, TTL-cached, worker spaces excluded).
CB-532 leads can message each other AND be answered. Principal.leader now
carries its terminal, so ownsSession() can be true for a lead; the
"and you must be a worker" conjunct beside it protected nothing.
Retires `primary:` — reply nudges follow the delegating lead, a
binding recorded at bridge_send where both halves are known.
CB-533 ClaudeCodeLauncher passes --model. argv is usually a wrapper
(`ccs <profile>`) that re-exports its own model family, so
ANTHROPIC_MODEL alone was silently overruled.
CB-534 a lead is deliverable. The CB-113 readiness gate only opened for
terminals in WorkerPresence, which only workers ever enter, so every
lead->lead send waited out the ~60s grace and failed having never
been typed. The gate guards a *spawned* peer's boot window; a lead
is never spawned.
CB-535 bridge_list returns `leads` alongside `workers`, with `self` on the
caller's row. An empty worker roster no longer reads as "no peers".
CB-536 CLAUDE.md: lead<->lead is coordinate-only, never sideways delegation.
Propagated byte-identically to wiki/7-Use-Cases.md.
MIXED PROVENANCE — recorded deliberately rather than hidden. This tree also carries
in-progress CB-537 (context separation) authored by the peer lead gpt-sol-5.6 and
its worker: Capability.CONTEXT_RESET, SessionManager.clearAfterTurn, and the
Injector/TurnListener/CompletionResolver/launcher changes around it. That work was
done in this shared working tree rather than a worktree, and is entangled with the
above in BridgedConfig.java, Bridged.java and ClaudeCodeLauncher.java, so neither
lead could stage its own half without sweeping in the other's. Committing the whole
green state is the honest resolution; the peer branches from here.
Note for whoever picks CB-537 up: the design in this commit is SUPERSEDED. Both
leads agreed to replace the global `clearAfterTurn` boolean with per-delivery
policy (inherit|fresh|thread) applied PRE-delivery, because a post-turn reset races
by construction — Injector.onStatus clears awaitingCompletion and dequeues the next
message in the same tick. `fresh` is also a correctness guarantee, so an adapter
without a reset capability must refuse it rather than log a no-op.
mvn clean install: Tests run: 464, Failures: 0, Errors: 0, Skipped: 0. BUILD SUCCESS.
The orchestration contract had no distributable form. Every consuming project
hand-copied a block of CLAUDE.md and hand-wrote an .mcp.json, and we keep a
script whose only job is to notice those copies drifting apart. A plugin is
versioned, installed once, and updates in place.
Ships no credentials, deliberately: every secret is referenced by environment
variable NAME and the value never enters a file, which is what makes the
artifact safe to publish. The setup skill states the two rules that are easy to
get wrong for the right-sounding reasons — the PR token must not be able to
merge (a worker opens, the primary gates), and ANTHROPIC_BASE_URL must never be
set by setup, because mounting the bridge must not move a session off
subscription.
The plugin root is plugin/, not the repo root. An installed plugin's .mcp.json
is a committed file, while this repo's root .mcp.json is local-only and
--skip-worktree; rooting the plugin at the repo would commit the primary's IDE
servers and hand them to every worker — the exact failure CB-525 exists to
prevent.
Scope is client-side setup only. herdr and bridged stay separate services with
their own lifecycles, and the skill refuses to install them rather than guess.
It also refuses to accept /healthz as proof: health reports only that the daemon
can reach herdr, and CB-521 showed it staying green while every spawn failed, so
verification ends with a real spawn.
Both manifests pass `claude plugin validate --strict`.
The ollama backend is decommissioned, so the example config stops pointing
readers at a dead host and the guard allowlist stops carrying an entry with
no profile behind it — a stale entry there is dead permission, and that list
is the only thing keeping a worker off the primary's subscription.
The second illustrative profile survives as gx11: the example exists to show
`placement: weighted` having something to choose between, and a one-profile
example would quietly stop demonstrating that.
It also moves the CB-523 auto-compact override onto the surviving profile.
That guard had been attached to `ollama` alone, so retiring the profile would
have removed the fleet's only protection against the failure it was written
for — a worker whose prompt is rejected before auto-compact ever fires. The
window belongs on every profile, not on whichever one happened to hit it.
Four new resolved decisions (§7.7-7.10): turn state split by where the signals
are, with ABANDONED explicitly belt-and-braces over the waiter timeout; the
dual ack model with spawn idempotence by construction (gid stored IN the herdr
pane — the load-bearing detail of the no-ledger position — plus an in-flight
reservation for redelivery during a slow spawn); enforced publish semantics
(confirms + mandatory on a separate channel, return-before-confirm caveat);
queue lifecycle = session lifecycle with .v2 names for the redeclare hazard.
§8 reworked: global id scheme resolved and moved up; control authorization
sharpened into the hard gate on U4 (per-host allowlist beside the peer keys);
key distribution/rotation added. New §9: implementation order, each step
verifiable single-host, U4 gated, U8 last.
Wiki pointer bumped to 4320c1c (chapter 10 same-pass changes).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
Six decisions recorded in §7, replacing the matching open questions: signed
messages (identity from the key, extending identity-from-connection across the
broker), target-host-owned profiles advertised via presence, repo provisioning
by pinned forge clone, live-only asks with TTL + TOO_LATE notice, spawn-id
dedup on the target, and broker-outage semantics (local unaffected, remote
fails fast, gateway stays soft-state). §8 keeps what is genuinely still open,
with control *authorization* now separated from the resolved *authenticity*.
The broker-level half (U8 broadcast, exclusive consumers, inbox caps, TLS,
schema version, trace id) lands in wiki chapter 10 §10 — pointer bumped
(also picks up 1710a77, chapter 11 Features).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
Bump bridged to 1.0.0 and add the release notes: the single-leader,
single-host scope is closed — gateway, lifecycle, two-way delivery,
pluggable peers, auth/authz/audit, supervision, CI. Cross-host
federation (CB-308) is the next major line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
CLAUDE.md already carries a mandatory before-done checklist ("the prompt is part
of the product"). It covered the instruction surface but not the operator-facing
one, and the result was measurable: CB-506 through CB-525 shipped without a
single wiki mention, while the Roadmap went on claiming Stage 5 was finished.
Discipline is what already failed, so this rides the existing gate rather than
adding a new habit to remember: one more row, firing when a change touches
anything an operator can use, configure, or observe. The row names where the
other two kinds of change go too (contracts to Implementation, coverage to the
Roadmap), so "nothing to document" is a decision the table makes rather than a
default you fall into.
Addendum-only — the canonical block is untouched and still byte-identical to the
wiki template (verified).
Rebase integration only, no behaviour change. CB-519's tests named the pane the
pre-protocol-19 fake produced (w9:pW_n); upstream's herdr 0.8.0 port creates the
pane through tab.create and starts the agent into it, so the fake now reports
w9:pRoot_n. Five assertions were therefore counting closes of a pane that never
existed and reading 0.
mvn clean install: Tests run: 399, Failures: 0, Errors: 0 — BUILD SUCCESS
A worker in a provisioned worktree was inheriting the primary's MCP servers by
two independent routes: the repo commits a .mcp.json declaring the IDE servers,
so a fresh checkout mounts them, and the default parity overlay then copied the
primary's own copy over the top.
Those servers are bound to the primary's IntelliJ project, so every path they
hand back points into the primary's checkout. A CB-523 worker made all 59 of its
edits there while running `mvn -f bridged/pom.xml` against its worktree — every
build it ran was of code that did not contain its changes, and it passed. The
worker's own `ls` of the file it had "edited" returned "No such file".
GitWorktrees now neutralizes .mcp.json at provisioning: an explicitly empty
server map, --skip-worktree'd when tracked so it never reads as pending work a
worker might commit. Unconditional, because the overlay was only half the leak.
The bridge itself is unaffected — it reaches a worker through the launcher's
--mcp-config flag, not the project file, so bridge_reply still works.
- BridgedConfig: .mcp.json out of the default parity overlay
- GitWorktrees: isolateToolSurface() on add(), with the rationale in javadoc
- GitWorktreesTest: 4 real-git acceptance tests (2 fail if the call is removed)
- implementer skill: work from $PWD, and quote a green unpiped `mvn clean
install` from the worktree as the acceptance criterion
mvn clean install: Tests run: 392, Failures: 0, Errors: 0 — BUILD SUCCESS
The weighted policy breaks an exact-weight tie on candidate list order
(WeightedRoundRobinPolicy picks the first candidate with a strictly greater
score), and that list comes from CompositePeerLauncher.candidates(), which
iterates profileConfigs. Both that map and BridgedConfig.workerProfiles() were
built with Map.copyOf, whose iteration order is salted per JVM run — so the
"in definition order" contract candidates() documents was not held.
Two consequences. In production, a config with equal weights (ollama 0.5 /
gx10 0.5) placed its first worker on a profile chosen at random on every daemon
restart. In the suite, CompositePeerLauncherTest.failoverRetriesNextCandidate-
WhenProfileIsUnreachable failed roughly one run in four, because whether
profile "a" was tried first depended on the salt.
Preserve definition order at every layer: unmodifiable LinkedHashMap for
workerProfiles(), profileConfigs, and byProfile (which also feeds the
user-visible bridge_profiles listing). The tests build profile maps with an
ordered helper rather than Map.of, which is salted for the same reason.
Guarded by a pair of tests declaring the same two profiles in opposite order
and asserting opposite first attempts, so any order-scrambling implementation
must fail one of them. Verified by mutation: reverting profileConfigs to
Map.copyOf fails 8/8 runs (6 caught by the original test, 2 only by the new
reversed-order one); with the fix, 10/10 fresh JVMs pass, 388 tests green.
Caller identity resolved any loopback PID that mapped to a herdr pane as a
WORKER, and PaneLocator scans every pane -- not just bridged-spawned ones. A
primary running inside a herdr pane therefore classified itself as a worker and
was refused SPAWN/SEND/STOP, i.e. every orchestration verb it exists to call.
The failure is self-locking: PrimaryRegistry only learns the primary's terminal
from bridge_send/bridge_spawn, the exact calls being refused, so the learned
value can never bootstrap. Only an operator-set pin breaks the cycle.
CallerResolver now consults primary.terminal from config *before* the pane
lookup. Deliberately the pinned value only, never the learned one -- the learned
terminal is populated by the callers this method is itself classifying, so
trusting it would be circular. Config is operator input, never network input, so
this widens no attack surface; bridge_whoami and the authz gate still share one
resolution.
Fixing that exposed a second, older bug. BridgeMcp's context extractor forwards
the caller's terminal into markPresent on every MCP call, documented as "no-op
for the primary (null terminal)". WorkerPresence.markPresent honours that, but
PresenceBridge overrides it and forwards the same null into SessionManager.
onReady -> transitionByTerminal -> findByTerminal, which called
terminalId.equals(...) unguarded. It only reached the scan once the registry was
non-empty, so the primary's first spawn succeeded and every later call NPE'd
with an HTTP 500 -- and it would have fired for ANY primary not living in a
herdr pane, pinned or not.
findByTerminal is now total. That covers onReady, onDelivered, onTurnComplete
and onTurnFailed at once; a null id could never match a registered session
anyway, so "no match" is the honest answer rather than taking down an unrelated
tool call.
Also drops two dead pass-throughs on CallerResolver (cwdForPid, tokenMode) that
IDE inspections flagged -- callers use ConnectionIdentity and BridgedConfig.Auth
directly.
The example config now states that primary.terminal is REQUIRED, not just a
push-loop optimisation, when the primary shares a herdr pane.
mvn clean install: 360 tests, 0 failures. Verified live: daemon restarted on
this jar, bridge_whoami reports primary, and four concurrent worktree spawns --
the exact shape that NPE'd -- now all succeed.
A primary running INSIDE a herdr pane was resolved as a worker by the
pane-match rule and refused every orchestration tool — the exact lockout
bridge_whoami surfaced on this deployment. The CB-307 primary.terminal pin
always claimed to replace connection-derived identity but only fed the push
loop; it now short-circuits CallerResolver ahead of the pane→worker rule
(the pane mapping is as unforgeable as a worker's, so no credential needed,
even in token mode). bridged.example.yaml documents the block.
Also guard the presence bridge against the primary's null terminal: the MCP
context extractor marks presence on every request, and the first genuine
primary contact NPEd into the SPAWNING→READY transition.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
herdr 0.8.0 redesigned the agent API out from under the daemon: agent.start
now launches a supported kind INTO an existing pane, env/cwd move to pane
creation (tab.create / pane.split — the subscription-boundary seam now),
agent.send is replaced by agent.prompt (self-submitting) plus agent.send_keys
for the Enter nudge, and terminal ids are no longer valid agent.* targets.
- AgentControl: start(name, kind, args, paneId); prompt/send_keys delivery;
cached terminal→pane target translation (invalidated on agent_not_found).
- WorkspaceControl: tab.create carries cwd+env; pane.split for legacy placement.
- HerdrPeerLauncher: the seed pane IS the worker pane (no drop step); retry
agent.start while the seed shell boots (agent_pane_busy).
- FakeHerdr and the test suite model protocol 19 (unique seed panes, required
kind/pane_id, prompt-based delivery); contract tests probe the seed shell
instead of arbitrary-command agents, which protocol 19 removed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
Turns the primary's half of the bridge charter from a bullet list of
policies into a numbered 0-8 procedure, and splits delegated review out
of the merge step it used to sit beside. Wiki template kept byte-identical
by splicing; pointer bumped to 0c896eb in the same commit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw1FosEt3Noix5GG9wJ2r
CB-517 moved orchestration policy into CLAUDE.md but left it as a bullet
list, so the procedure was implicit: the order of operations had to be
reconstructed from a parallelisation bullet, and nothing said when to
review or when to tear down. A policy you have to reassemble on each
task is one you will reassemble differently each task. Restate the
primary's half as a numbered 0-8 flow — role check, split, gate, spawn
all, send all, collect, verify, review, adjudicate — so that following
it is checkable against the tool calls rather than a matter of recall.
Two steps carry the load. Spawn and send are separate on purpose:
folding them into one loop is what silently serialises work that was
meant to fan out. And review is now its own step ahead of the merge
rather than a clause inside it, because the two have opposite owners —
reviewers fan out over the diff (never the implementer of the scope
they review, and briefed from the diff rather than the author's
rationale, which carries the same blind spot), while adjudication, the
merge and teardown stay with the primary. Merging on a reviewer's word
is delegating the gate by proxy, so the step says so outright.
Nothing is dropped. The six bullets that trailed the tool table are
relocated into the step that owns each — profile explicitness into
spawn, playbook naming and self-containment into send, claim
verification into its own step, the ~60s blocking-send cap into a note
beneath the flow — and the table stays as the intent→tool lookup.
The wiki pointer moves with it. The block is canonical only if its
template matches byte for byte, so the template was produced by
splicing the block out of CLAUDE.md rather than by editing it in
parallel, and the sync check the repo documents passes. Bumping the
pointer in the same commit keeps charter and template versioned
together, as CB-517 did.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw1FosEt3Noix5GG9wJ2r
The CB-307 durable ReplyInbox needs an AMQP broker, but the one behind it
was run ad hoc and had simply vanished from the host — which takes the
whole daemon with it, since AmqpReplyInbox.open throws and Bridged.java:187
does not guard it. A missing broker is a hard startup failure, not a
degraded mode, so 'how the broker runs' is part of the system, not a local
detail.
Pinned to 2.9.1 (:latest would move the broker under a running daemon),
data on a named volume (held-but-unacked replies are the entire point of
Stage 2 — a plain 'compose down' would discard exactly what durability
protects), and restart: unless-stopped so it comes back after a reboot
instead of disappearing again.
Ports are bound to 127.0.0.1 deliberately: LavinMQ ships a default
guest/guest account, which is only acceptable while nothing off-host can
reach it.
Verified by driving the production AmqpReplyInbox against this deployment
(publish/peek/dedup/FIFO/ack, then reconnect): 8/8 including redelivery of
the unacked message. That pairing had never been exercised — the
@Tag("contract") test runs against a RabbitMQ container, and is excluded
from the default build, so mvn clean install covers the broker path zero
times.
The communication rules lived only in two opt-in skills, so nothing
always-on told the primary how to orchestrate and nothing guaranteed a
worker loaded its playbook. Move protocol and policy into CLAUDE.md,
which a worker inherits for free (its worktree is a checkout of this
repo), and leave the skills as pure per-job procedure.
bridge_whoami closes the load-bearing gap: every tool already consumed
the caller identity ConnectionIdentity resolves from the connection, but
none reported it, so an agent had to infer its own role from side
channels the daemon does not control. Guessing fails asymmetrically — a
primary acting as a worker is refused by the authz gate and learns at
once, while a worker acting as the primary ends its turn without
bridge_reply and the sender silently receives nothing. The tool reuses
the same Principal the gate is built on, so the two cannot disagree; the
primary gets role only (handing it a sessionId it does not own would
invite the forged reply Authz refuses), and a worker missing from the
registry still gets role + sessionId rather than 'unknown'.
The CLAUDE.md block is written to be copied as-is into any project that
mounts the bridge: repo-local details (Authz paths, the .mcp.json/wiki
exclusions, the skill names) moved below it into a project addendum, and
every role-inference fallback is stated one-way — the mount-name signal
only holds for mcp__bridge__* (the launcher fixes it), not for the
primary's mount, which each project names itself. The wiki carries the
block verbatim as the template, with a sync check.
Because this repo IS the bridge, that block is shipped surface, not
documentation: the addendum adds a mandatory checklist mapping each part
of the code to the part of the prompt it can invalidate.
Also: delegate-by-default policy for the primary — the test is not 'could
I do this faster myself' but 'can I write a brief good enough for a
worker'.
mvn clean install: 356 tests green (353 + 3 for whoami); ide_diagnostics
clean on both changed files.
Tearing a worker down left the send that was waiting on it stranded. The
rendezvous waiter stayed open, so a blocking bridge_send kept blocking and an
async one kept reporting PENDING until ASYNC_TIMEOUT_MS — thirty minutes —
even though the worker provably no longer existed and the delegation could
never complete.
Observed repeatedly this session: a task sitting at
{"phase":"pending","detail":"worker unknown"} long after I had personally
deleted the pane. The maddening part is that poll() already HAD the evidence —
it calls liveStatus() to build that detail string, gets back "unknown", and
reports PENDING anyway. It also never reached /metrics under any outcome, so a
stalled delegation was invisible to both the task view and the dashboard. The
only way I ever diagnosed one was reading the worker's pane over the herdr
socket by hand.
Fixed at the choke point rather than by guessing from status strings:
SessionManager.release() is the single path every teardown funnels through
(REST stop, MCP stop, idle-TTL reaper, recycle, shutdown drain), so it now
notifies a release listener with the terminalId, and Bridged.main wires that to
MessageService.abandon(). Abandon fails the open waiter with a real reason.
Deliberately NOT done by inferring "gone" from liveStatus(): that method
collapses a vanished pane, a wedged worker and a herdr hiccup into the same
"unknown" string, so acting on it would fail live delegations during a
transient blip. An explicit lifecycle signal cannot be ambiguous.
Notified before launcher.stop() so a blocked caller fails fast, and wrapped so
a listener failure can never prevent the teardown it is reacting to.
Resolving as a failure rather than letting it time out also means the outcome
is counted — a torn-down delegation now shows up as
sends_total{outcome="failed"} instead of nothing at all.
353 tests (was 346): abandon fails a waiting send / is a no-op with no waiter /
never clobbers a send the worker already answered, an abandoned async task
polls FAILED rather than PENDING, release notifies with the right terminal,
releasing an unknown pane notifies nobody, and a throwing listener does not
block teardown.
Verified live on the running daemon, reproducing the original scenario:
async send -> {"phase":"pending","detail":"worker working"}
DELETE the worker -> 204
poll -> {"phase":"failed","detail":"the worker session was
released before it replied"}
/metrics -> bridged_sends_total{outcome="failed"} 1
Previously that poll returned PENDING for thirty minutes and the counter stayed
empty.
NOT addressed here, and worth deciding separately: a worker that simply never
replies (rather than being released) still rides out the full 30-minute
ASYNC_TIMEOUT_MS. That is a policy question about long-running delegations, not
a correctness bug.
Five tests pinning the invariants that decide WHICH turn a reply belongs to.
These protect against a silent correctness bug — a reply attributed to the
wrong turn — not against a crash, which is why they were worth picking over
higher-percentage coverage gaps.
Chosen by blast radius, not by uncovered-line count. Both guards are compound
conditions with a side that never executed, i.e. exactly the shape where a
clause can be deleted as "redundant" and every existing test still passes.
CompletionResolver:
- The CB-115 misattribution guard suppresses a completion when the scrape is
byte-identical to the pane at delivery. Its !scrapeFailed clause was
unexercised: delete it and a FAILED read is misread as "no output change",
so the send is suppressed and hangs to the caller's timeout instead of
resolving. The new test sets the baseline to "" so the empty tail from a
failed read would byte-match and wrongly suppress — built to die precisely
when that clause dies.
- The fail() guard leaves an already-resolved waiter alone. The new test also
asserts agent.read is never called, so the worker is not scraped for a send
nobody is waiting on.
Rendezvous: a second resolution of an already-completed waiter returns false
and does not overwrite the first value, for both resolveCompletion and
resolveFailure.
Verified by sabotage, one guard at a time: removing !scrapeFailed reds
resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent; removing the
isDone() clause reds failLeavesAnAlreadyResolvedWaiterUntouchedAndSkipsTheScrape.
(The first attempt at the second sabotage reported a false pass — the patch hit
an identically-worded guard earlier in the file. Line-targeted and re-run.)
346 tests, was 341. Worker-implemented on the local-vLLM profile; it noticed
three of the eight cases I asked for already existed and said so with names
rather than duplicating them.
Also of note: the first delegation of this ticket wedged the worker — the pane
showed a zsh parse error and it went idle with an untouched worktree, task stuck
pending. The retry differed only in phrasing the same requirements as prose
instead of quoting Java boolean expressions. Filed as a bridge robustness
concern: injected content shares a channel with control, and a wedged turn is
invisible in both the task view and /metrics.
Six tests on the delivery core, the most load-bearing class in the project,
which sat at 74% with its hardest paths unexercised: answer() (the CB-205
ask-answer resolution) had 11 of 23 lines uncovered, send()'s timeout branches
7 of 21, poll() 7 of 18, and tryLock 3 of 4.
Covered: TIMED_OUT_QUEUED vs TIMED_OUT_WORKING (undelivered vs delivered-but-
silent), an answered worker that never sends its follow-up reply, an unknown
ticket, a completed async ticket reporting its reply and replySource, and a
second concurrent send to the same session returning BUSY rather than hanging.
Worker-implemented on the local-vLLM profile, self-verified: it reported
'Tests run: 341, Failures: 0' and an independent run of its branch agrees
exactly. Additive only — 101 lines in one test file, no main/ source touched.
Quality is good on its own terms, not just green: the async-completion test
polls to a deadline instead of sleeping and hoping, the contention test
releases the blocked send so the test thread is not left pinned, and every
assertion is on a specific Outcome rather than 'nothing threw' — the failure
mode an earlier worker produced in CB-510.
Branch pushed by the worker; PR left to the primary since GITEA_TOKEN is not
granted to that profile by design.
CB-505 claimed authorization is "enforced on both entry paths". It is — but
only REST was ever tested. Coverage showed BridgeMcp.deny(), principal(),
callerTerminal(), worktreeRequest() and every tool-registration lambda at ZERO
executed lines: no test had ever constructed a BridgeMcp, because the existing
BridgeMcpTest calls only the static handler methods. So the MCP half of the
security control had ten REST tests' worth of nothing behind it.
An unexercised security control is a claim, not a control.
Made testable by separating policy from plumbing rather than by reaching for a
mocking library the project does not use:
- denyFor(Principal, Action, target) is the decision — testable directly.
- deny(exchange, ...) shrinks to pulling the caller out of the SDK exchange.
- principalFrom(role, terminal, pid) extracts identity reconstruction from
McpSyncServerExchange, an SDK type with no fake available.
Moved the `authz == null` enforcement switch OUT of the exchange-facing wrapper
and INTO denyFor. Found by a failing test: as written, any future tool calling
denyFor directly would have silently skipped the gate. The switch now lives with
the decision it governs.
New BridgeMcpAuthzTest constructs a real BridgeMcp — which is why coverage moved
so far, since that also runs the constructor and all the tool wiring — and pins
the table on this path: primary orchestrates, worker cannot; worker replies only
as itself; the primary cannot forge a worker reply; anonymous gets nothing; and
401-shaped vs 403-shaped refusals are counted apart.
Verified as real controls, not decoration: with the gate forced open, 5 of the 9
fail. 335 tests (was 326).
Fixes one of the three counters declared in BridgedMetrics but never
incremented, so bridged_push_nudges_total{outcome=delivered|exhausted} now
actually appears on /metrics. An absent series reads as 'no push failures ever'
rather than 'not measured', which is the misleading case.
Implemented end to end by an opencode worker on the local-vLLM profile in an
isolated worktree, and this is the first delegation where the worker verified
its own work: it ran mvn, hit a real test failure, iterated, and reported
'Tests run: 326, Failures: 0' — which matches an independent run of its branch
exactly. Every earlier delegation reported results it had no way to check,
because CB-511 had not yet given workers a PATH with a toolchain.
It also committed and pushed its own branch unprompted, and was straight about
the one thing it could not do: opening the PR, since GITEA_HOST/GITEA_TOKEN are
not granted to that profile ('URL rejected: No host part'). That grant is opt-in
per profile by design, so the merge is primary-side as intended.
Diff reviewed and correct on every constraint, including the subtle ones:
delivered counted inside the try after a successful send (not the catch),
exhausted only on the reminder-cap branch and not the other two STOP paths, and
a single Metrics instance moved above pushLoop and shared with MessageService.