MemberRegistry.bind() had 18 call sites, every one in its own unit test.
Nothing in production ever bound a terminal to a slot, so CallerResolver's
architect branch was unreachable, every spawned architect resolved as WORKER,
and Authz's architect row was dead code. Since SEND is granted to
isPrimary() || isArchitect(), no architect could message anyone.
The spawn lifecycle now binds on both acquire paths and unbinds on release,
through a MemberLifecycle seam with a no-op default, so all six SessionManager
constructors are unchanged.
The escalation guard is deliberately double-sided. MemberRegistry flattens
every pool, so dev:* and reviewer:* slots sit in the same map the architect
resolver reads. The lifecycle binds only ARCHITECT roles, AND CallerResolver
independently checks roleForSlot(slot) == ARCHITECT before granting. Either
alone would be enough today; together, a regression in one cannot escalate a
worker. A reviewer confirmed a third barrier already existed: Authz gates
SPAWN on isPrimary(), so only a lead can request role=ARCHITECT at all.
Verified by the primary: mvn clean install, 640 tests, 0 failures.
Two findings recorded, neither blocking:
1. Known race, low severity. registry.put() and memberLifecycle.acquired()
are not atomic. A release landing between them unbinds nothing (no binding
exists yet), then acquire binds a dead terminal that no later release will
ever clear - the slot leaks. Only the shutdown drain can reach a SPAWNING
session (the idle reaper skips it, and an explicit stop needs a paneId the
spawn has not returned), so the leak dies with the process. Note that
binding before put does NOT fix it: the drain then iterates a roster the
session is not in yet. A real fix needs atomicity.
2. The legacy Supplier-based withLeadsAndMembers overload and the Map-form
constructors now silently disable architect resolution: memberSlotRoles
defaults to _ -> null, so the architect branch can never be taken there.
Production uses the registry form, and the tests were migrated, so nothing
fails - but a future caller gets workers with no error.
A member no longer receives the primary's pre-approved tool grants by default.
The grants were inert — GitWorktrees.isolateToolSurface already strips each
worktree's .mcp.json to an empty server map — so this is defence in depth: two
independent guards instead of one. Same reasoning as CB-525, which removed the
sibling .mcp.json and left this file behind.
An explicit parityOverlay: in config is unaffected; only the default changes.
Verified by the primary: mvn clean install, 637 tests, 0 failures.
The repo carried a gitignored .secrets/ directory with four files. Two of them
(context7-token, gitea-token) were byte-identical copies of variables the login
shell already exported. One (gitea-host) is not a secret. The fourth
(worker-gitea-token) was the only copy anywhere, and nothing exported it, so
bridged read gitTokenEnv from an environment that never had it and every worker
push got an empty token.
All four values now live in the operator's single sourced secrets file, verified by
sha256 before the copies were removed. opencode.json reads them as {env:...}, which
.mcp.json already did. A second copy of a secret is the problem: the copy you forget
is the one that leaks or goes stale.
This makes worktree isolation load-bearing rather than a workaround. opencode.json is
tracked, so it lands in every worktree. It used to fail there, because {file:.secrets/}
pointed at files a worktree never receives and OpenCode refuses to start on a dangling
reference. With {env:...} the reference resolves, and a member would silently inherit
the primary's admin-scoped GITEA_ACCESS_TOKEN. GitWorktrees already neutralizes the
file; only its stated reason changes, and it is now a confidentiality boundary.
The port-to-opencode skill taught {file:.secrets/} as the preferred pattern, so it is
rewritten to teach the central store and to say why we moved. .gitignore keeps the
.secrets/ line as a backstop against habit.
Includes the wiki pointer, which also carries the CB-559 config-reload correction.
The shipped docs and javadoc said a profile's `model` and `tabLabel` take effect on
the next spawn. They do not, and ConfigRef did not detect the change either, so a
reload logged a clean "config reloaded" and silently did nothing. That is the worst
outcome a reload can produce: the operator has no reason to doubt it.
What makes a key hot is who reads it and when, not that it is config. Placement
reads weight and maxLoad through a supplier on CompositePeerLauncher, so those are
genuinely hot. HerdrPeerLauncher takes Map.copyOf(profiles) at construction and
resolves every spawn out of that copy, so model, baseUrl, argv, env and the rest
cannot move until the daemon restarts.
changedDeferredKeys now compares every launch component of an existing profile,
excluding weight and maxLoad, and names the profiles that need a restart. The
javadoc and bridged.example.yaml say the same thing. Two tests pin the pair:
weight/maxLoad reports nothing deferred, a changed model reports the profile by name.
Four commits that together reshape how the daemon is told who it may run.
CB-557 (schema) — one `fleet:` block replaces `leaders:`, `members:`, `leadScan:`
and `defaultProfile:`. Role is the containing map key, not a `role:` field, so a
misspelled pool name declares nothing instead of producing a member with no
contract. Tab labels are role-first and counted per role+profile.
CB-557 (routing) — an unqualified spawn picks its profile from its role's pool.
Role and profile stay orthogonal: a reviewer may run on the same profile as the
dev whose diff it reads, so the two cannot be one field. An explicit
`bridge_spawn{profile:…}` stays exempt, because it carries no role and would
otherwise be refused for a profile the operator named.
CB-558 — the daemon launches a declared lead when fewer than `instances` are
running. An auto-launched lead is NOT a member: no reply charter, never
registered with SessionManager (the idle reaper would kill it), and no Anthropic
binding in its env. A lead counts as live only when herdr reports a running
agent, so one crash does not disable auto-launch forever, and an unreachable
herdr starts nothing at all.
CB-559 — `bridged.yaml` can be re-read without a restart, opt-in through
`configReload:`. Keys are hot (`fleet:`, `placement:`, an existing profile's
fields), deferred (lifecycle, guard, adding a profile) or cold (bind,
herdrSocket, broker, auth). A cold change refuses the WHOLE reload rather than
half-applying it, because a daemon matching no file on disk is worse for an
operator than no reload at all.
634 tests, IDE diagnostics clean. Verified live against the running daemon: the
lead launcher recognised the existing lead and started nothing; the reload
applied a hot change, refused a `bind:` change exactly once (not once per tick),
and reverted cleanly.
Tuning a fleet meant restarting bridged, and a restart tears down every lead
and worker it owns. Changing one pool's weight cost the whole fleet's state,
so in practice nobody changed it.
ConfigRef holds the live BridgedConfig in an AtomicReference. Consumers read
it at the point of use, so a change reaches the next spawn with nothing
rebuilt. The launchers that used to capture config into fields now take
suppliers: the fleet tabLabel template, the profile map, the placement policy
and the fleet block.
Keys fall into three classes, and the difference is what already exists when
the reload happens:
hot fleet: (pools + tabLabel), placement:, and an existing profile's
weight / maxLoad / model / tabLabel — live on the next spawn.
deferred lifecycle:, leadHeartbeat:, guard:, worktreeRoot:, spawnReady*,
and adding/removing a profile — accepted, but the startup wiring
keeps the old value. The reload logs these by name.
cold bind:, herdrSocket:, broker:, auth: — refuses the WHOLE reload.
A cold change refuses everything rather than applying the hot half. A
half-applied reload leaves the daemon matching no file on disk, which is the
worst thing a reload can do to an operator reading that file to work out what
the daemon is doing. Refusing keeps the invariant that the live config is
always some version of the file.
A parse failure or a failed startup validator is refused the same way, and
the running config stays live: a file being saved is sometimes read
mid-write, and degrading a working daemon over a half-written file is a bad
trade. The same four validators startup runs are re-run, so a config that
could not have booted cannot slip in through a reload.
ConfigWatcher polls the modified time on a daemon thread, opt-in through
configReload.enabled (default off, so an upgraded daemon is unchanged). It
stamps the timestamp BEFORE reloading, so a refused file is not retried every
tick — the next save earns a fresh attempt. A missing file is skipped
silently, because editors unlink briefly mid-save.
MicroProfile Config was the first idea and does not fit: @ConfigMapping needs
interfaces, resolves once at bootstrap, and reload would still mean rebuild
and swap. The port would also lose the raw-YAML duplicate-key detection,
since duplicates have already collapsed once the tree is flattened to
properties.
634 tests.
`fleet.leaders.<name>.instances` was descriptive. Now the daemon reads it: a
lead that names a `profile:` is started when fewer than `instances` are running.
A lead with only a `terminal:` stays recognise-only, as before.
A lead is not a member, and LeadLauncher exists to keep it that way. Every other
spawn path goes through HerdrPeerLauncher, which does three things a lead must
never get: it appends the worker reply charter ("you are an off-subscription
worker … end every turn with bridge_reply" — the opposite of an orchestrator);
it registers the session with SessionManager, whose idle reaper would kill a
lead for being idle, which is a lead's normal state; and it can move a peer off
the subscription. So this launcher talks to AgentControl/WorkspaceControl
directly. The duplicated argv/env assembly is the cheaper half of that trade.
Not double-spawning is the safety property, so liveness needs two pieces of
evidence. A running agent in a tab labelled `lead: <name>` finds an
auto-launched lead. A running agent on a pinned `terminal:` finds one the
operator opened by hand — without it, a pinned lead whose tab carries no
matching label would be relaunched on every boot. Member workspaces are
excluded, so a member in a matching tab is never counted. If herdr cannot be
reached, nothing is started: a second orchestrator is worse than none.
Liveness deliberately requires the AGENT, not just the label. LeadTabScanner
used to promise that bridged never writes a lead label, so there was no
round-trip from the daemon's own rename back into its next decision. That is no
longer true, and its javadoc now says so. The trust direction is unaffected — a
label is a name, not a capability — but staleness becomes real: a label left by
a crashed session would otherwise read as a live lead forever and disable
auto-launch permanently.
Two new knobs. `workspace:` (default "leads") is where a launched lead's tab
goes; it must not be a member workspace, because those are excluded from the
scan and a lead placed in one would never be found again. `cwd:` defaults to
bridged's own working directory.
Also: WorkspaceControl.listTabs, and a FakeHerdr tab seeder that leaves the
canned response byte-identical when no tab is seeded.
617 tests pass (16 new), IDE-clean.
The pools were config-only until now: the launchers still received one global
effectiveDefaultProfile and placement still ranged over every configured
profile, so a reviewer could be placed on an architect-only backend.
Three parts:
SessionManager computed the role, stored it on the MemberSession, and never put
it on the SpawnRequest. So the role reached the record that describes the spawn
but not the call that performs it — every launcher saw DEV. Both spawn paths
(plain and worktree) now carry it.
CompositePeerLauncher takes the Fleet and draws its candidates from
fleet.<role> instead of from all profiles. An absent or empty pool means
unconstrained, not blocked: a config that declares pools for some roles must
keep spawning the rest, so it falls back to every profile. A null Fleet is the
pre-CB-557 wiring and behaves exactly as before.
Bridged passes cfg.fleet() to the composite and cfg.fleet().tabLabel() to both
launchers. The tab-label knob was accepted by HerdrPeerLauncher but passed by
nobody, so it was inert — the label only looked right because the fallback
happened to match the configured template. Four tests now pin the wiring
instead of the coincidence.
An EXPLICIT profile stays exempt from the pool. `bridge_spawn{profile:"opus"}`
carries no role, so it defaults to DEV; judging it against the dev pool would
refuse a spawn the operator asked for by name. maxLoad still applies to it.
Also cleared the IDE warnings in the touched files: an immediately-rethrown
catch (the comment stays, the redundant block goes), unused lambda params, a
javadoc link to a package-private class, two unused imports.
601 tests pass.
Four top-level keys (leaders:, members:, leadScan:, defaultProfile:) become one
`fleet:` block, and a member's role becomes the map key that contains it rather
than a `role:` field inside it.
Why the key and not a field: a misspelled `role: architct` used to produce a
member with no contract, which nothing rejected. A misspelled pool name declares
nothing, which is a shape the loader can see.
`fleet.architects/developers/reviewers` are pools of profiles a role MAY run on.
That replaces the single global `defaultProfile:`, so an unqualified spawn now
resolves its profile from the pool of the role it asked for. Role and profile
stay orthogonal: a reviewer may run on the same profile as the dev it reviews,
and one profile may appear in several pools.
Tab labels are role-first — `dev: sonnet #4`. The template lives on `fleet:`
because a profile cannot know the role of the member launched on it; a profile
may still override it. The `{n}` counter is scoped per role+profile, so a dev
and a reviewer on one profile each start at #1. Making {role} the first field
also turns the lead/member namespace check into a structural guarantee: roles
are a closed enum, so only hand-written templates can still collide with a lead
tabPrefix.
Removed keys are hard errors that name their successor. `defaultProfile:` has no
single successor key, so its message explains the new model instead of pointing
at a key that does not exist.
Map order is kept with LinkedHashMap, deliberately not Map.copyOf — the latter
salts iteration order per JVM run, which would destroy the YAML definition order
that `placement: fixed` selects on.
Not yet wired: SessionManager still hands the launchers one effectiveDefault-
Profile, so pools are not enforced at spawn time yet, and placement still ranges
over all profiles.
595 tests pass.
Every spawned peer is now a member with a role, and the role travels with it
from the spawn call to the roster.
MCP:
bridge_spawn gains role: architect | dev | reviewer (default dev). An
unknown role is refused with the valid spellings in the message.
bridge_list returns "members" instead of "workers"; each row carries both
role (what it is for) and profile (which backend it runs on).
The spawn result echoes the role back, so a spawn that fell back to dev is
visible rather than silent.
REST:
GET/POST /members and DELETE /members/{paneId} replace /workers.
POST accepts role= as a query param or a body field; an unknown role is 400.
Code:
dev.ltms.bridged.worker package -> dev.ltms.bridged.member
WorkerSession -> MemberSession, plus a MemberRole role component
WorkerPresence -> MemberPresence
SessionManager.acquire gains a role parameter; the existing overloads keep
working and default to DEV, which is exactly what "worker" used to mean.
ClaudeCodeLauncher and OpenCodeLauncher keep their names on purpose — they
are named after the backend, not the role.
Not done here: the launch charter is still one string for every role, so a
member is told its role by nobody yet. That is the next ticket.
mvn clean install: 583 tests, 0 failures, 0 errors, BUILD SUCCESS.
A member is anything a lead spawns. Every member carries two independent
attributes:
role — which contract: architect, dev or reviewer. It picks the launch
charter, the role file, the playbook skill and the authz row.
profile — which backend: model, CLI adapter, credentials, cost.
They vary on their own. A reviewer may run on the same profile as the dev
whose diff it reads, which is the case that proves the two cannot be one
field.
Config changes (breaking — we are in active development, so no aliases):
workers: -> profiles: it was never a list of workers; it is a
catalogue of backends
defaultWorker: -> defaultProfile:
architects: -> members: each slot now names its role
An old config is rejected at load with the new key named, rather than being
warned about once and then running with zero profiles — that failure would
surface much later, at the first spawn, pointing nowhere near the cause.
Also:
- BridgedConfig.Worker -> BridgedConfig.Profile
- ArchitectRegistry -> MemberRegistry
- new peer.MemberRole enum, validated at startup
- profiles map is normalized once in the compact constructor, so the raw
map and the derived one can no longer disagree
- the legacy singular worker: block is dropped
- workerProfiles() -> profiles(); defaultProfile() -> effectiveDefaultProfile()
(the record component now owns the plain name)
mvn clean install: 578 tests, 0 failures, 0 errors, BUILD SUCCESS.
A worker that exhausts its context dies mid-turn and its bridge_reply — the
whole point of the turn — dies with it. Auto-compaction is a condition of the
turn contract for a spawned peer, not an operator preference.
OPENCODE_CONFIG is merged over ~/.config/opencode/config.json rather than
replacing it, so workers already inherited auto:true from the home file. That
inheritance is exactly what this removes as a dependency: the home file is
outside this repo, differs per machine, and is not ours to rely on.
The trade is recorded in the code: an OPENCODE_CONFIG value overrides the home
value, so an operator cannot disable compaction for bridged workers from home.
Deliberate for peers we spawn and whose turns we must land; per-profile control
would be a profile knob, not the removal of this line.
Origin: fleet-wide auto-compaction audit by peer lead gpt-sol-5.6, which found
autoCompactEnabled:false alongside a 300k window across the Claude instance
configs. Handed over uncommitted; taken deliberately, with rationale added.
Fix two handler-level bugs found in PR #21:
- Only PRIMARY callers may update PrimaryRegistry.record (the legacy singleton 'primary'
fallback for no-delegation inbox nudges). An architect SEND previously recorded its terminal
as the fallback; the per-target delegation map does not cure the singleton. New
BridgeMcp.recordPrimarySingleton uses the resolved role (caller.isPrimary()) — named leads
(PRIMARY) still record, architects never do.
- MessageService.send now opens the rendezvous waiter BEFORE queueing delivery, fixing both the
enqueue-before-open fast-reply race (a fast reply no longer orphans into the inbox) and
callback-failure ordering: a throwing onAccepted (public callback) fails the send cleanly with
no stale waiter and no queued, orphanable message.
Tests: architect SEND vs lead SEND primary-singleton regression; throwing onAccepted leaves no
stale waiter or queued orphan.
Record PrimaryRegistry delegator ownership via a MessageService accepted-delivery
hook (won the session lock + queued delivery), never at bridge_send request time, so
a concurrent sender that times out BUSY cannot steal a live turn's reply routing.
Make Rendezvous.open atomic fail-if-present so a double open trips loudly instead of
replacing the waiter another send is blocked on. Answering a bridge_ask keeps the same
ownership (no rewrite). Adds ownership/rendezvous regression tests.
Pairs with 244fbd9. The docs half lived on the submodule's own remote, so the
pointer bump is separate by necessity, not by preference.
The substantive part is not the new 'Run a worker on the subscription' entry but
the correction beside it: 'Give workers a toolchain' asserted flatly that an env:
entry cannot repoint a worker past the SubscriptionGuard. CB-539 made that false,
and the wiki went on claiming it until CB-542 closed the hole. The entry now
states the rule and its one exception together.
Verified by the lead in a clean worktree at 5afe8e1 rather than on the worker's
report: Tests run: 474, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS (main
was at 464).
The invariant this lands: there is no configuration in which a worker reaches an
Anthropic endpoint that no guard vetted. Closed at two layers — a fatal, profile-
naming refusal at config load, and a launcher-side strip so it holds for profiles
built in code that never passed validation.
The wiki half is a separate branch on the submodule's own remote; the pointer bump
follows as its own commit on main.
Add a per-profile subscription: true opt-in that lets a claude-code worker run on
the operator's Claude subscription when there is no off-subscription endpoint for it
(e.g. sonnet on ccs). When set, the launcher injects neither ANTHROPIC_BASE_URL nor
ANTHROPIC_AUTH_TOKEN and skips SubscriptionGuard's base_url requirement for that
profile only, logging a WARN naming the profile. subscription: true alongside a
baseUrl is refused as contradictory. The default (absent/false) keeps today's hard
refusal unchanged; every other profile stays allowlist-checked and SubscriptionGuard
is untouched.
Regression I introduced one commit ago, caught by two consecutive spawn failures
and reproduced end-to-end.
Tracking `.autoenv` means git checks it out into every provisioned worktree. A
worktree is a new path, and autoenv authorizes by path, so the file is always
unauthorized there — autoenv prints "[autoenv] Authorize this file? (y/n/d)" and
blocks on `read` (activate.sh:211-222). The pane's shell sits at that prompt, so
`ccs <profile>` never runs, the peer never becomes injectable, and CB-306's
readiness gate closes the pane after 20s. Symptom is a bare "spawn timed out";
nothing names autoenv, which is what made it worth writing down.
Confirmed rather than inferred: the peer lead's cb-537 worker, spawned BEFORE
f8182e4, has no `.autoenv` in its worktree and is still alive; a throwaway
worktree at HEAD reproduces the prompt on entry.
The file is still worth committing — the reasoning in f8182e4 stands, and it is
recoverable from there. What is missing is the other half: GitWorktrees already
neutralizes `.mcp.json` in a provisioned worktree ("worker tool surface is
launcher-mounted only"), and `.autoenv` needs exactly the same treatment for
exactly the same reason — a worker's environment is launcher-supplied, never
repo-supplied. Re-land it with that, tracked as CB-543.
Rejected the quicker fix of exporting AUTOENV_ASSUME_YES: it auto-approves
arbitrary repo-controlled shell code in every spawned worker, which is a worse
trade than one unset variable.
A spawned peer has no human at its pane, so an approval prompt is not a pause —
it is a wedge. The agent stops, looks identical to a legitimate mid-turn wait,
and can never reach its bridge_reply, so the delegation dies silently and the
lead learns nothing until the timeout.
Unconditional rather than a per-profile knob, which is the right call: there is
no configuration under which a bridge-spawned opencode worker WANTS to block on
an approval it has no way to answer. opencode's help calls --auto 'dangerous!',
and that warning is written for a human at a terminal; the blast radius here is
already bounded by the layer above — a worker runs in its own git worktree, on
its own branch, off-subscription, and cannot merge. The lead is the gate.
Reviewed by me rather than fanned out: 24 lines across one method and its two
tests, below the threshold where a reviewer pass pays for itself.
argvWithModel is re-signatured to take composed argv instead of building it,
so the two flag-appenders compose rather than each owning construction.
Moves the submodule pointer 4320c1c -> 0a21b49, matching the docs to the code in
bc13b8e rather than leaving chapter 11 describing a fleet with one lead in it.
7b5381b 11, 7: CB-530..536 — leaders:, leadScan:, lead-to-lead messaging,
lead deliverability, leads in bridge_list; 7's copy of the portable
block re-synced byte-identically
0a21b49 12: where opencode's secrets actually come from, and why {env:...}
is green under Claude Code and broken in a terminal
Both are already pushed to the wiki's own remote, so this pointer resolves for
anyone who clones — the ordering that matters, and the reason the wiki went
first.
This is the pointer only. `wiki/` is a submodule with its own remote and its
content is never committed here; bumping the recorded SHA in its own commit is
how this repo has always recorded a docs update (see "bump wiki to 4320c1c").
This file has always been designed to be committed and says so in its own header;
it simply never was, so every clone of this workspace has been reconstructing it
by hand or duplicating tokens instead.
It holds no secret. It reads `.secrets/` (gitignored, 0600) and exports three
variables, because Claude Code expands `${VAR}` in `.mcp.json` from the *process*
environment and cannot read a file — so without it, CONTEXT7_TOKEN and the gitea
pair must be duplicated as literals in `.claude/settings.local.json`. opencode
needs none of this: it reads `.secrets/` directly via `{file:...}`.
Verified before committing that no value appears in it, only the three names and
the paths they are read from.
Safe in a worktree by construction: a worktree receives tracked files only, so
`.secrets/` is absent there and the whole block is skipped rather than failing.
Workers are fed by the launcher's env instead — which is where the name mismatch
documented in wiki chapter 12 (GITEA_TOKEN vs GITEA_ACCESS_TOKEN) has to be
reconciled.