The repo carried a gitignored .secrets/ directory with four files. Two of them
(context7-token, gitea-token) were byte-identical copies of variables the login
shell already exported. One (gitea-host) is not a secret. The fourth
(worker-gitea-token) was the only copy anywhere, and nothing exported it, so
bridged read gitTokenEnv from an environment that never had it and every worker
push got an empty token.
All four values now live in the operator's single sourced secrets file, verified by
sha256 before the copies were removed. opencode.json reads them as {env:...}, which
.mcp.json already did. A second copy of a secret is the problem: the copy you forget
is the one that leaks or goes stale.
This makes worktree isolation load-bearing rather than a workaround. opencode.json is
tracked, so it lands in every worktree. It used to fail there, because {file:.secrets/}
pointed at files a worktree never receives and OpenCode refuses to start on a dangling
reference. With {env:...} the reference resolves, and a member would silently inherit
the primary's admin-scoped GITEA_ACCESS_TOKEN. GitWorktrees already neutralizes the
file; only its stated reason changes, and it is now a confidentiality boundary.
The port-to-opencode skill taught {file:.secrets/} as the preferred pattern, so it is
rewritten to teach the central store and to say why we moved. .gitignore keeps the
.secrets/ line as a backstop against habit.
Includes the wiki pointer, which also carries the CB-559 config-reload correction.
The shipped docs and javadoc said a profile's `model` and `tabLabel` take effect on
the next spawn. They do not, and ConfigRef did not detect the change either, so a
reload logged a clean "config reloaded" and silently did nothing. That is the worst
outcome a reload can produce: the operator has no reason to doubt it.
What makes a key hot is who reads it and when, not that it is config. Placement
reads weight and maxLoad through a supplier on CompositePeerLauncher, so those are
genuinely hot. HerdrPeerLauncher takes Map.copyOf(profiles) at construction and
resolves every spawn out of that copy, so model, baseUrl, argv, env and the rest
cannot move until the daemon restarts.
changedDeferredKeys now compares every launch component of an existing profile,
excluding weight and maxLoad, and names the profiles that need a restart. The
javadoc and bridged.example.yaml say the same thing. Two tests pin the pair:
weight/maxLoad reports nothing deferred, a changed model reports the profile by name.
Tuning a fleet meant restarting bridged, and a restart tears down every lead
and worker it owns. Changing one pool's weight cost the whole fleet's state,
so in practice nobody changed it.
ConfigRef holds the live BridgedConfig in an AtomicReference. Consumers read
it at the point of use, so a change reaches the next spawn with nothing
rebuilt. The launchers that used to capture config into fields now take
suppliers: the fleet tabLabel template, the profile map, the placement policy
and the fleet block.
Keys fall into three classes, and the difference is what already exists when
the reload happens:
hot fleet: (pools + tabLabel), placement:, and an existing profile's
weight / maxLoad / model / tabLabel — live on the next spawn.
deferred lifecycle:, leadHeartbeat:, guard:, worktreeRoot:, spawnReady*,
and adding/removing a profile — accepted, but the startup wiring
keeps the old value. The reload logs these by name.
cold bind:, herdrSocket:, broker:, auth: — refuses the WHOLE reload.
A cold change refuses everything rather than applying the hot half. A
half-applied reload leaves the daemon matching no file on disk, which is the
worst thing a reload can do to an operator reading that file to work out what
the daemon is doing. Refusing keeps the invariant that the live config is
always some version of the file.
A parse failure or a failed startup validator is refused the same way, and
the running config stays live: a file being saved is sometimes read
mid-write, and degrading a working daemon over a half-written file is a bad
trade. The same four validators startup runs are re-run, so a config that
could not have booted cannot slip in through a reload.
ConfigWatcher polls the modified time on a daemon thread, opt-in through
configReload.enabled (default off, so an upgraded daemon is unchanged). It
stamps the timestamp BEFORE reloading, so a refused file is not retried every
tick — the next save earns a fresh attempt. A missing file is skipped
silently, because editors unlink briefly mid-save.
MicroProfile Config was the first idea and does not fit: @ConfigMapping needs
interfaces, resolves once at bootstrap, and reload would still mean rebuild
and swap. The port would also lose the raw-YAML duplicate-key detection,
since duplicates have already collapsed once the tree is flattened to
properties.
634 tests.
`fleet.leaders.<name>.instances` was descriptive. Now the daemon reads it: a
lead that names a `profile:` is started when fewer than `instances` are running.
A lead with only a `terminal:` stays recognise-only, as before.
A lead is not a member, and LeadLauncher exists to keep it that way. Every other
spawn path goes through HerdrPeerLauncher, which does three things a lead must
never get: it appends the worker reply charter ("you are an off-subscription
worker … end every turn with bridge_reply" — the opposite of an orchestrator);
it registers the session with SessionManager, whose idle reaper would kill a
lead for being idle, which is a lead's normal state; and it can move a peer off
the subscription. So this launcher talks to AgentControl/WorkspaceControl
directly. The duplicated argv/env assembly is the cheaper half of that trade.
Not double-spawning is the safety property, so liveness needs two pieces of
evidence. A running agent in a tab labelled `lead: <name>` finds an
auto-launched lead. A running agent on a pinned `terminal:` finds one the
operator opened by hand — without it, a pinned lead whose tab carries no
matching label would be relaunched on every boot. Member workspaces are
excluded, so a member in a matching tab is never counted. If herdr cannot be
reached, nothing is started: a second orchestrator is worse than none.
Liveness deliberately requires the AGENT, not just the label. LeadTabScanner
used to promise that bridged never writes a lead label, so there was no
round-trip from the daemon's own rename back into its next decision. That is no
longer true, and its javadoc now says so. The trust direction is unaffected — a
label is a name, not a capability — but staleness becomes real: a label left by
a crashed session would otherwise read as a live lead forever and disable
auto-launch permanently.
Two new knobs. `workspace:` (default "leads") is where a launched lead's tab
goes; it must not be a member workspace, because those are excluded from the
scan and a lead placed in one would never be found again. `cwd:` defaults to
bridged's own working directory.
Also: WorkspaceControl.listTabs, and a FakeHerdr tab seeder that leaves the
canned response byte-identical when no tab is seeded.
617 tests pass (16 new), IDE-clean.
The pools were config-only until now: the launchers still received one global
effectiveDefaultProfile and placement still ranged over every configured
profile, so a reviewer could be placed on an architect-only backend.
Three parts:
SessionManager computed the role, stored it on the MemberSession, and never put
it on the SpawnRequest. So the role reached the record that describes the spawn
but not the call that performs it — every launcher saw DEV. Both spawn paths
(plain and worktree) now carry it.
CompositePeerLauncher takes the Fleet and draws its candidates from
fleet.<role> instead of from all profiles. An absent or empty pool means
unconstrained, not blocked: a config that declares pools for some roles must
keep spawning the rest, so it falls back to every profile. A null Fleet is the
pre-CB-557 wiring and behaves exactly as before.
Bridged passes cfg.fleet() to the composite and cfg.fleet().tabLabel() to both
launchers. The tab-label knob was accepted by HerdrPeerLauncher but passed by
nobody, so it was inert — the label only looked right because the fallback
happened to match the configured template. Four tests now pin the wiring
instead of the coincidence.
An EXPLICIT profile stays exempt from the pool. `bridge_spawn{profile:"opus"}`
carries no role, so it defaults to DEV; judging it against the dev pool would
refuse a spawn the operator asked for by name. maxLoad still applies to it.
Also cleared the IDE warnings in the touched files: an immediately-rethrown
catch (the comment stays, the redundant block goes), unused lambda params, a
javadoc link to a package-private class, two unused imports.
601 tests pass.
Four top-level keys (leaders:, members:, leadScan:, defaultProfile:) become one
`fleet:` block, and a member's role becomes the map key that contains it rather
than a `role:` field inside it.
Why the key and not a field: a misspelled `role: architct` used to produce a
member with no contract, which nothing rejected. A misspelled pool name declares
nothing, which is a shape the loader can see.
`fleet.architects/developers/reviewers` are pools of profiles a role MAY run on.
That replaces the single global `defaultProfile:`, so an unqualified spawn now
resolves its profile from the pool of the role it asked for. Role and profile
stay orthogonal: a reviewer may run on the same profile as the dev it reviews,
and one profile may appear in several pools.
Tab labels are role-first — `dev: sonnet #4`. The template lives on `fleet:`
because a profile cannot know the role of the member launched on it; a profile
may still override it. The `{n}` counter is scoped per role+profile, so a dev
and a reviewer on one profile each start at #1. Making {role} the first field
also turns the lead/member namespace check into a structural guarantee: roles
are a closed enum, so only hand-written templates can still collide with a lead
tabPrefix.
Removed keys are hard errors that name their successor. `defaultProfile:` has no
single successor key, so its message explains the new model instead of pointing
at a key that does not exist.
Map order is kept with LinkedHashMap, deliberately not Map.copyOf — the latter
salts iteration order per JVM run, which would destroy the YAML definition order
that `placement: fixed` selects on.
Not yet wired: SessionManager still hands the launchers one effectiveDefault-
Profile, so pools are not enforced at spawn time yet, and placement still ranges
over all profiles.
595 tests pass.
Every spawned peer is now a member with a role, and the role travels with it
from the spawn call to the roster.
MCP:
bridge_spawn gains role: architect | dev | reviewer (default dev). An
unknown role is refused with the valid spellings in the message.
bridge_list returns "members" instead of "workers"; each row carries both
role (what it is for) and profile (which backend it runs on).
The spawn result echoes the role back, so a spawn that fell back to dev is
visible rather than silent.
REST:
GET/POST /members and DELETE /members/{paneId} replace /workers.
POST accepts role= as a query param or a body field; an unknown role is 400.
Code:
dev.ltms.bridged.worker package -> dev.ltms.bridged.member
WorkerSession -> MemberSession, plus a MemberRole role component
WorkerPresence -> MemberPresence
SessionManager.acquire gains a role parameter; the existing overloads keep
working and default to DEV, which is exactly what "worker" used to mean.
ClaudeCodeLauncher and OpenCodeLauncher keep their names on purpose — they
are named after the backend, not the role.
Not done here: the launch charter is still one string for every role, so a
member is told its role by nobody yet. That is the next ticket.
mvn clean install: 583 tests, 0 failures, 0 errors, BUILD SUCCESS.
A member is anything a lead spawns. Every member carries two independent
attributes:
role — which contract: architect, dev or reviewer. It picks the launch
charter, the role file, the playbook skill and the authz row.
profile — which backend: model, CLI adapter, credentials, cost.
They vary on their own. A reviewer may run on the same profile as the dev
whose diff it reads, which is the case that proves the two cannot be one
field.
Config changes (breaking — we are in active development, so no aliases):
workers: -> profiles: it was never a list of workers; it is a
catalogue of backends
defaultWorker: -> defaultProfile:
architects: -> members: each slot now names its role
An old config is rejected at load with the new key named, rather than being
warned about once and then running with zero profiles — that failure would
surface much later, at the first spawn, pointing nowhere near the cause.
Also:
- BridgedConfig.Worker -> BridgedConfig.Profile
- ArchitectRegistry -> MemberRegistry
- new peer.MemberRole enum, validated at startup
- profiles map is normalized once in the compact constructor, so the raw
map and the derived one can no longer disagree
- the legacy singular worker: block is dropped
- workerProfiles() -> profiles(); defaultProfile() -> effectiveDefaultProfile()
(the record component now owns the plain name)
mvn clean install: 578 tests, 0 failures, 0 errors, BUILD SUCCESS.
A worker that exhausts its context dies mid-turn and its bridge_reply — the
whole point of the turn — dies with it. Auto-compaction is a condition of the
turn contract for a spawned peer, not an operator preference.
OPENCODE_CONFIG is merged over ~/.config/opencode/config.json rather than
replacing it, so workers already inherited auto:true from the home file. That
inheritance is exactly what this removes as a dependency: the home file is
outside this repo, differs per machine, and is not ours to rely on.
The trade is recorded in the code: an OPENCODE_CONFIG value overrides the home
value, so an operator cannot disable compaction for bridged workers from home.
Deliberate for peers we spawn and whose turns we must land; per-profile control
would be a profile knob, not the removal of this line.
Origin: fleet-wide auto-compaction audit by peer lead gpt-sol-5.6, which found
autoCompactEnabled:false alongside a 300k window across the Claude instance
configs. Handed over uncommitted; taken deliberately, with rationale added.
Fix two handler-level bugs found in PR #21:
- Only PRIMARY callers may update PrimaryRegistry.record (the legacy singleton 'primary'
fallback for no-delegation inbox nudges). An architect SEND previously recorded its terminal
as the fallback; the per-target delegation map does not cure the singleton. New
BridgeMcp.recordPrimarySingleton uses the resolved role (caller.isPrimary()) — named leads
(PRIMARY) still record, architects never do.
- MessageService.send now opens the rendezvous waiter BEFORE queueing delivery, fixing both the
enqueue-before-open fast-reply race (a fast reply no longer orphans into the inbox) and
callback-failure ordering: a throwing onAccepted (public callback) fails the send cleanly with
no stale waiter and no queued, orphanable message.
Tests: architect SEND vs lead SEND primary-singleton regression; throwing onAccepted leaves no
stale waiter or queued orphan.
Record PrimaryRegistry delegator ownership via a MessageService accepted-delivery
hook (won the session lock + queued delivery), never at bridge_send request time, so
a concurrent sender that times out BUSY cannot steal a live turn's reply routing.
Make Rendezvous.open atomic fail-if-present so a double open trips loudly instead of
replacing the waiter another send is blocked on. Answering a bridge_ask keeps the same
ownership (no rewrite). Adds ownership/rendezvous regression tests.
Add a per-profile subscription: true opt-in that lets a claude-code worker run on
the operator's Claude subscription when there is no off-subscription endpoint for it
(e.g. sonnet on ccs). When set, the launcher injects neither ANTHROPIC_BASE_URL nor
ANTHROPIC_AUTH_TOKEN and skips SubscriptionGuard's base_url requirement for that
profile only, logging a WARN naming the profile. subscription: true alongside a
baseUrl is refused as contradictory. The default (absent/false) keeps today's hard
refusal unchanged; every other profile stays allowlist-checked and SubscriptionGuard
is untouched.
A spawned peer has no human at its pane, so an approval prompt is not a pause —
it is a wedge. The agent stops, looks identical to a legitimate mid-turn wait,
and can never reach its bridge_reply, so the delegation dies silently and the
lead learns nothing until the timeout.
Unconditional rather than a per-profile knob, which is the right call: there is
no configuration under which a bridge-spawned opencode worker WANTS to block on
an approval it has no way to answer. opencode's help calls --auto 'dangerous!',
and that warning is written for a human at a terminal; the blast radius here is
already bounded by the layer above — a worker runs in its own git worktree, on
its own branch, off-subscription, and cannot merge. The lead is the gate.
Reviewed by me rather than fanned out: 24 lines across one method and its two
tests, below the threshold where a reviewer pass pays for itself.
argvWithModel is re-signatured to take composed argv instead of building it,
so the two flag-appenders compose rather than each owning construction.
Two leads now work as peers rather than one primary plus workers. The arc:
CB-530/531 lead identity: `leaders:` names panes, `leadScan:` discovers them by
tab label (LeadTabScanner, TTL-cached, worker spaces excluded).
CB-532 leads can message each other AND be answered. Principal.leader now
carries its terminal, so ownsSession() can be true for a lead; the
"and you must be a worker" conjunct beside it protected nothing.
Retires `primary:` — reply nudges follow the delegating lead, a
binding recorded at bridge_send where both halves are known.
CB-533 ClaudeCodeLauncher passes --model. argv is usually a wrapper
(`ccs <profile>`) that re-exports its own model family, so
ANTHROPIC_MODEL alone was silently overruled.
CB-534 a lead is deliverable. The CB-113 readiness gate only opened for
terminals in WorkerPresence, which only workers ever enter, so every
lead->lead send waited out the ~60s grace and failed having never
been typed. The gate guards a *spawned* peer's boot window; a lead
is never spawned.
CB-535 bridge_list returns `leads` alongside `workers`, with `self` on the
caller's row. An empty worker roster no longer reads as "no peers".
CB-536 CLAUDE.md: lead<->lead is coordinate-only, never sideways delegation.
Propagated byte-identically to wiki/7-Use-Cases.md.
MIXED PROVENANCE — recorded deliberately rather than hidden. This tree also carries
in-progress CB-537 (context separation) authored by the peer lead gpt-sol-5.6 and
its worker: Capability.CONTEXT_RESET, SessionManager.clearAfterTurn, and the
Injector/TurnListener/CompletionResolver/launcher changes around it. That work was
done in this shared working tree rather than a worktree, and is entangled with the
above in BridgedConfig.java, Bridged.java and ClaudeCodeLauncher.java, so neither
lead could stage its own half without sweeping in the other's. Committing the whole
green state is the honest resolution; the peer branches from here.
Note for whoever picks CB-537 up: the design in this commit is SUPERSEDED. Both
leads agreed to replace the global `clearAfterTurn` boolean with per-delivery
policy (inherit|fresh|thread) applied PRE-delivery, because a post-turn reset races
by construction — Injector.onStatus clears awaitingCompletion and dequeues the next
message in the same tick. `fresh` is also a correctness guarantee, so an adapter
without a reset capability must refuse it rather than log a no-op.
mvn clean install: Tests run: 464, Failures: 0, Errors: 0, Skipped: 0. BUILD SUCCESS.
The ollama backend is decommissioned, so the example config stops pointing
readers at a dead host and the guard allowlist stops carrying an entry with
no profile behind it — a stale entry there is dead permission, and that list
is the only thing keeping a worker off the primary's subscription.
The second illustrative profile survives as gx11: the example exists to show
`placement: weighted` having something to choose between, and a one-profile
example would quietly stop demonstrating that.
It also moves the CB-523 auto-compact override onto the surviving profile.
That guard had been attached to `ollama` alone, so retiring the profile would
have removed the fleet's only protection against the failure it was written
for — a worker whose prompt is rejected before auto-compact ever fires. The
window belongs on every profile, not on whichever one happened to hit it.
Bump bridged to 1.0.0 and add the release notes: the single-leader,
single-host scope is closed — gateway, lifecycle, two-way delivery,
pluggable peers, auth/authz/audit, supervision, CI. Cross-host
federation (CB-308) is the next major line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
Rebase integration only, no behaviour change. CB-519's tests named the pane the
pre-protocol-19 fake produced (w9:pW_n); upstream's herdr 0.8.0 port creates the
pane through tab.create and starts the agent into it, so the fake now reports
w9:pRoot_n. Five assertions were therefore counting closes of a pane that never
existed and reading 0.
mvn clean install: Tests run: 399, Failures: 0, Errors: 0 — BUILD SUCCESS
A worker in a provisioned worktree was inheriting the primary's MCP servers by
two independent routes: the repo commits a .mcp.json declaring the IDE servers,
so a fresh checkout mounts them, and the default parity overlay then copied the
primary's own copy over the top.
Those servers are bound to the primary's IntelliJ project, so every path they
hand back points into the primary's checkout. A CB-523 worker made all 59 of its
edits there while running `mvn -f bridged/pom.xml` against its worktree — every
build it ran was of code that did not contain its changes, and it passed. The
worker's own `ls` of the file it had "edited" returned "No such file".
GitWorktrees now neutralizes .mcp.json at provisioning: an explicitly empty
server map, --skip-worktree'd when tracked so it never reads as pending work a
worker might commit. Unconditional, because the overlay was only half the leak.
The bridge itself is unaffected — it reaches a worker through the launcher's
--mcp-config flag, not the project file, so bridge_reply still works.
- BridgedConfig: .mcp.json out of the default parity overlay
- GitWorktrees: isolateToolSurface() on add(), with the rationale in javadoc
- GitWorktreesTest: 4 real-git acceptance tests (2 fail if the call is removed)
- implementer skill: work from $PWD, and quote a green unpiped `mvn clean
install` from the worktree as the acceptance criterion
mvn clean install: Tests run: 392, Failures: 0, Errors: 0 — BUILD SUCCESS
The weighted policy breaks an exact-weight tie on candidate list order
(WeightedRoundRobinPolicy picks the first candidate with a strictly greater
score), and that list comes from CompositePeerLauncher.candidates(), which
iterates profileConfigs. Both that map and BridgedConfig.workerProfiles() were
built with Map.copyOf, whose iteration order is salted per JVM run — so the
"in definition order" contract candidates() documents was not held.
Two consequences. In production, a config with equal weights (ollama 0.5 /
gx10 0.5) placed its first worker on a profile chosen at random on every daemon
restart. In the suite, CompositePeerLauncherTest.failoverRetriesNextCandidate-
WhenProfileIsUnreachable failed roughly one run in four, because whether
profile "a" was tried first depended on the salt.
Preserve definition order at every layer: unmodifiable LinkedHashMap for
workerProfiles(), profileConfigs, and byProfile (which also feeds the
user-visible bridge_profiles listing). The tests build profile maps with an
ordered helper rather than Map.of, which is salted for the same reason.
Guarded by a pair of tests declaring the same two profiles in opposite order
and asserting opposite first attempts, so any order-scrambling implementation
must fail one of them. Verified by mutation: reverting profileConfigs to
Map.copyOf fails 8/8 runs (6 caught by the original test, 2 only by the new
reversed-order one); with the fix, 10/10 fresh JVMs pass, 388 tests green.
Caller identity resolved any loopback PID that mapped to a herdr pane as a
WORKER, and PaneLocator scans every pane -- not just bridged-spawned ones. A
primary running inside a herdr pane therefore classified itself as a worker and
was refused SPAWN/SEND/STOP, i.e. every orchestration verb it exists to call.
The failure is self-locking: PrimaryRegistry only learns the primary's terminal
from bridge_send/bridge_spawn, the exact calls being refused, so the learned
value can never bootstrap. Only an operator-set pin breaks the cycle.
CallerResolver now consults primary.terminal from config *before* the pane
lookup. Deliberately the pinned value only, never the learned one -- the learned
terminal is populated by the callers this method is itself classifying, so
trusting it would be circular. Config is operator input, never network input, so
this widens no attack surface; bridge_whoami and the authz gate still share one
resolution.
Fixing that exposed a second, older bug. BridgeMcp's context extractor forwards
the caller's terminal into markPresent on every MCP call, documented as "no-op
for the primary (null terminal)". WorkerPresence.markPresent honours that, but
PresenceBridge overrides it and forwards the same null into SessionManager.
onReady -> transitionByTerminal -> findByTerminal, which called
terminalId.equals(...) unguarded. It only reached the scan once the registry was
non-empty, so the primary's first spawn succeeded and every later call NPE'd
with an HTTP 500 -- and it would have fired for ANY primary not living in a
herdr pane, pinned or not.
findByTerminal is now total. That covers onReady, onDelivered, onTurnComplete
and onTurnFailed at once; a null id could never match a registered session
anyway, so "no match" is the honest answer rather than taking down an unrelated
tool call.
Also drops two dead pass-throughs on CallerResolver (cwdForPid, tokenMode) that
IDE inspections flagged -- callers use ConnectionIdentity and BridgedConfig.Auth
directly.
The example config now states that primary.terminal is REQUIRED, not just a
push-loop optimisation, when the primary shares a herdr pane.
mvn clean install: 360 tests, 0 failures. Verified live: daemon restarted on
this jar, bridge_whoami reports primary, and four concurrent worktree spawns --
the exact shape that NPE'd -- now all succeed.
A primary running INSIDE a herdr pane was resolved as a worker by the
pane-match rule and refused every orchestration tool — the exact lockout
bridge_whoami surfaced on this deployment. The CB-307 primary.terminal pin
always claimed to replace connection-derived identity but only fed the push
loop; it now short-circuits CallerResolver ahead of the pane→worker rule
(the pane mapping is as unforgeable as a worker's, so no credential needed,
even in token mode). bridged.example.yaml documents the block.
Also guard the presence bridge against the primary's null terminal: the MCP
context extractor marks presence on every request, and the first genuine
primary contact NPEd into the SPAWNING→READY transition.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
herdr 0.8.0 redesigned the agent API out from under the daemon: agent.start
now launches a supported kind INTO an existing pane, env/cwd move to pane
creation (tab.create / pane.split — the subscription-boundary seam now),
agent.send is replaced by agent.prompt (self-submitting) plus agent.send_keys
for the Enter nudge, and terminal ids are no longer valid agent.* targets.
- AgentControl: start(name, kind, args, paneId); prompt/send_keys delivery;
cached terminal→pane target translation (invalidated on agent_not_found).
- WorkspaceControl: tab.create carries cwd+env; pane.split for legacy placement.
- HerdrPeerLauncher: the seed pane IS the worker pane (no drop step); retry
agent.start while the seed shell boots (agent_pane_busy).
- FakeHerdr and the test suite model protocol 19 (unique seed panes, required
kind/pane_id, prompt-based delivery); contract tests probe the seed shell
instead of arbitrary-command agents, which protocol 19 removed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
The communication rules lived only in two opt-in skills, so nothing
always-on told the primary how to orchestrate and nothing guaranteed a
worker loaded its playbook. Move protocol and policy into CLAUDE.md,
which a worker inherits for free (its worktree is a checkout of this
repo), and leave the skills as pure per-job procedure.
bridge_whoami closes the load-bearing gap: every tool already consumed
the caller identity ConnectionIdentity resolves from the connection, but
none reported it, so an agent had to infer its own role from side
channels the daemon does not control. Guessing fails asymmetrically — a
primary acting as a worker is refused by the authz gate and learns at
once, while a worker acting as the primary ends its turn without
bridge_reply and the sender silently receives nothing. The tool reuses
the same Principal the gate is built on, so the two cannot disagree; the
primary gets role only (handing it a sessionId it does not own would
invite the forged reply Authz refuses), and a worker missing from the
registry still gets role + sessionId rather than 'unknown'.
The CLAUDE.md block is written to be copied as-is into any project that
mounts the bridge: repo-local details (Authz paths, the .mcp.json/wiki
exclusions, the skill names) moved below it into a project addendum, and
every role-inference fallback is stated one-way — the mount-name signal
only holds for mcp__bridge__* (the launcher fixes it), not for the
primary's mount, which each project names itself. The wiki carries the
block verbatim as the template, with a sync check.
Because this repo IS the bridge, that block is shipped surface, not
documentation: the addendum adds a mandatory checklist mapping each part
of the code to the part of the prompt it can invalidate.
Also: delegate-by-default policy for the primary — the test is not 'could
I do this faster myself' but 'can I write a brief good enough for a
worker'.
mvn clean install: 356 tests green (353 + 3 for whoami); ide_diagnostics
clean on both changed files.
Tearing a worker down left the send that was waiting on it stranded. The
rendezvous waiter stayed open, so a blocking bridge_send kept blocking and an
async one kept reporting PENDING until ASYNC_TIMEOUT_MS — thirty minutes —
even though the worker provably no longer existed and the delegation could
never complete.
Observed repeatedly this session: a task sitting at
{"phase":"pending","detail":"worker unknown"} long after I had personally
deleted the pane. The maddening part is that poll() already HAD the evidence —
it calls liveStatus() to build that detail string, gets back "unknown", and
reports PENDING anyway. It also never reached /metrics under any outcome, so a
stalled delegation was invisible to both the task view and the dashboard. The
only way I ever diagnosed one was reading the worker's pane over the herdr
socket by hand.
Fixed at the choke point rather than by guessing from status strings:
SessionManager.release() is the single path every teardown funnels through
(REST stop, MCP stop, idle-TTL reaper, recycle, shutdown drain), so it now
notifies a release listener with the terminalId, and Bridged.main wires that to
MessageService.abandon(). Abandon fails the open waiter with a real reason.
Deliberately NOT done by inferring "gone" from liveStatus(): that method
collapses a vanished pane, a wedged worker and a herdr hiccup into the same
"unknown" string, so acting on it would fail live delegations during a
transient blip. An explicit lifecycle signal cannot be ambiguous.
Notified before launcher.stop() so a blocked caller fails fast, and wrapped so
a listener failure can never prevent the teardown it is reacting to.
Resolving as a failure rather than letting it time out also means the outcome
is counted — a torn-down delegation now shows up as
sends_total{outcome="failed"} instead of nothing at all.
353 tests (was 346): abandon fails a waiting send / is a no-op with no waiter /
never clobbers a send the worker already answered, an abandoned async task
polls FAILED rather than PENDING, release notifies with the right terminal,
releasing an unknown pane notifies nobody, and a throwing listener does not
block teardown.
Verified live on the running daemon, reproducing the original scenario:
async send -> {"phase":"pending","detail":"worker working"}
DELETE the worker -> 204
poll -> {"phase":"failed","detail":"the worker session was
released before it replied"}
/metrics -> bridged_sends_total{outcome="failed"} 1
Previously that poll returned PENDING for thirty minutes and the counter stayed
empty.
NOT addressed here, and worth deciding separately: a worker that simply never
replies (rather than being released) still rides out the full 30-minute
ASYNC_TIMEOUT_MS. That is a policy question about long-running delegations, not
a correctness bug.
Five tests pinning the invariants that decide WHICH turn a reply belongs to.
These protect against a silent correctness bug — a reply attributed to the
wrong turn — not against a crash, which is why they were worth picking over
higher-percentage coverage gaps.
Chosen by blast radius, not by uncovered-line count. Both guards are compound
conditions with a side that never executed, i.e. exactly the shape where a
clause can be deleted as "redundant" and every existing test still passes.
CompletionResolver:
- The CB-115 misattribution guard suppresses a completion when the scrape is
byte-identical to the pane at delivery. Its !scrapeFailed clause was
unexercised: delete it and a FAILED read is misread as "no output change",
so the send is suppressed and hangs to the caller's timeout instead of
resolving. The new test sets the baseline to "" so the empty tail from a
failed read would byte-match and wrongly suppress — built to die precisely
when that clause dies.
- The fail() guard leaves an already-resolved waiter alone. The new test also
asserts agent.read is never called, so the worker is not scraped for a send
nobody is waiting on.
Rendezvous: a second resolution of an already-completed waiter returns false
and does not overwrite the first value, for both resolveCompletion and
resolveFailure.
Verified by sabotage, one guard at a time: removing !scrapeFailed reds
resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent; removing the
isDone() clause reds failLeavesAnAlreadyResolvedWaiterUntouchedAndSkipsTheScrape.
(The first attempt at the second sabotage reported a false pass — the patch hit
an identically-worded guard earlier in the file. Line-targeted and re-run.)
346 tests, was 341. Worker-implemented on the local-vLLM profile; it noticed
three of the eight cases I asked for already existed and said so with names
rather than duplicating them.
Also of note: the first delegation of this ticket wedged the worker — the pane
showed a zsh parse error and it went idle with an untouched worktree, task stuck
pending. The retry differed only in phrasing the same requirements as prose
instead of quoting Java boolean expressions. Filed as a bridge robustness
concern: injected content shares a channel with control, and a wedged turn is
invisible in both the task view and /metrics.
CB-505 claimed authorization is "enforced on both entry paths". It is — but
only REST was ever tested. Coverage showed BridgeMcp.deny(), principal(),
callerTerminal(), worktreeRequest() and every tool-registration lambda at ZERO
executed lines: no test had ever constructed a BridgeMcp, because the existing
BridgeMcpTest calls only the static handler methods. So the MCP half of the
security control had ten REST tests' worth of nothing behind it.
An unexercised security control is a claim, not a control.
Made testable by separating policy from plumbing rather than by reaching for a
mocking library the project does not use:
- denyFor(Principal, Action, target) is the decision — testable directly.
- deny(exchange, ...) shrinks to pulling the caller out of the SDK exchange.
- principalFrom(role, terminal, pid) extracts identity reconstruction from
McpSyncServerExchange, an SDK type with no fake available.
Moved the `authz == null` enforcement switch OUT of the exchange-facing wrapper
and INTO denyFor. Found by a failing test: as written, any future tool calling
denyFor directly would have silently skipped the gate. The switch now lives with
the decision it governs.
New BridgeMcpAuthzTest constructs a real BridgeMcp — which is why coverage moved
so far, since that also runs the constructor and all the tool wiring — and pins
the table on this path: primary orchestrates, worker cannot; worker replies only
as itself; the primary cannot forge a worker reply; anonymous gets nothing; and
401-shaped vs 403-shaped refusals are counted apart.
Verified as real controls, not decoration: with the gate forced open, 5 of the 9
fail. 335 tests (was 326).
Workers could not run `mvn` or `java`. Every delegated task that asked for a
build came back "mvn is not on PATH", and the worker was right.
Root cause: HerdrPeerLauncher seeded the worker environment with an EMPTY map,
so bridged passed only the vars it explicitly set (OPENCODE_CONFIG, GITEA_TOKEN,
ANTHROPIC_*) and never PATH. herdr merges that map into its own process env, so
a worker inherited whatever PATH the herdr SERVER was started with. On this host
that server (pid 79870, PPID 1) had been up since Jul 4 with a PATH containing
neither the JDK nor Maven. Confirmed on a live worker: its PATH was byte-identical
to herdr's, and the only var bridged had contributed was OPENCODE_CONFIG.
The failure was invisible and non-deterministic: the fleet's capabilities
depended on how a long-lived daemon happened to be launched weeks earlier. There
are three herdr processes on this box with three different PATHs; the one owning
the socket is the one without a toolchain. bridged itself HAD Maven on PATH the
whole time — it just never passed it on.
It also quietly contradicted the project's own principle that "a worker is a
full peer of the primary", and the implementer skill's instruction to build,
commit and open a PR. Every delegation so far has depended on the primary
running the build gate.
Fix: baseEnv(cfg) seeds each worker with the daemon's own PATH, then applies the
profile's new optional env: map. Adapter-specific vars are layered on top and
therefore win — that ordering is load-bearing, not incidental: it stops an env:
entry from overwriting ANTHROPIC_BASE_URL and slipping past SubscriptionGuard,
which is checked against the profile's baseUrl alone. Pinned by a test.
Because the default is now the daemon's PATH, both supervision units set PATH
explicitly — launchd and systemd do not source a login shell, so under CB-504
the daemon (and every worker) would otherwise get a bare /usr/bin:/bin and this
bug would silently return in production.
324 tests (was 321): daemon-PATH propagation, profile env: passthrough including
an explicit PATH override, and the guard-bypass ordering.
Verified live: daemon restarted, worker spawned, and asked to run the tools —
"Apache Maven 3.9.16", "java version 25.0.2". Previously both were absent.
SessionReaper had no tests at all. Its TTL *policy* was already well covered
(SessionManager.reapIdle, 6 cases in SessionManagerTest); what was untested was
the thread wrapper around it — idempotent start/stop and whether the loop
actually runs and actually stops.
Observed through an injected clock rather than by sleeping and hoping: reapIdle
reads nowNanos exactly once per call, so the tick count IS the iteration count.
Waits are bounded polls, not fixed sleeps, and nothing asserts an exact
timing-derived number — flaky counts would be worse than no test.
321 tests (was 318); line coverage 66.9% -> 67.9%.
Drafted by an opencode worker on the new local-vLLM profile (branch
worker/cb-510-session-reaper-test-cd1793-1). Its structure and setup were good
and it was honest that it could not run mvn. But its third test asserted
NOTHING — it started the reaper, slept, stopped it, and relied on "no throw",
with a comment claiming that proved the loop had run. It did not: verified by
sabotage, all three of its tests passed against a start() replaced with an
immediate return.
Rewritten so the assertions can fail for the right reason. Same sabotage now
fails 2 of 3 (the third only pins stop()-before-start(), where "does not throw"
genuinely is the contract). Uncomfortably on the nose given this task began as
a hunt for tests that do not mean anything.
Build-time tooling only — never a compile or runtime dependency, so it adds
nothing to the shipped jar and no new transitive surface to the artifact.
(Noting per CLAUDE.md that the pom CVE gate could not be run: no JetBrains MCP
server is connected this session.)
Report at target/site/jacoco/index.html, machine-readable at jacoco.csv.
Deliberately NO check rule or threshold. A coverage gate rewards writing tests
that merely execute lines, which is the exact failure mode this codebase has
already been bitten by — CB-507 shipped a null-argument NPE with 311 green
tests because FakeWorktrees.repoRoot records its argument instead of shelling
out, so the broken line was covered and still wrong. Coverage is a map of where
to look, not a target to hit.
Baseline: 66.9% line, 60.6% branch, 76.4% method.