Closes the one known-unverified item before cross-host. CB-402 merged in
ded226a with increment 5 (the §5 live checklist) deferred; it has now run.
Provider question (§7 Q1) resolved with no credentials needed: opencode's own
gateway serves free-tier models. `opencode auth list` reports 0 credentials,
yet `opencode run -m opencode/north-mini-code-free` answers. Distinct from the
primary's subscription by construction, and needs no guard entry — opencode
carries no ANTHROPIC_BASE_URL, so SubscriptionGuard never applies to it.
The schema-drift risk was the real one and it did not bite. The adapter was
designed against opencode 1.1.31; installed is 1.18.5. The generated config
still validates unchanged (type:"remote" + instructions:[path]), and
`OPENCODE_CONFIG=… opencode mcp list` reports the bridge connected. Pinned as
a verified fact for 1.18.5.
Full lifecycle through REST: spawn (201, kind-routed to OpenCodeLauncher) ->
CB-306 gate passed ~0.6s -> ready -> send -> {"replySource":"reply"} (a
STRUCTURED bridge_reply, not the CB-115 completion fallback) -> delete (204,
tolerant teardown).
Unplanned cross-validation with CB-501: the audit trail recorded the reply as
role=WORKER actor=worker:term_657c… — connection-based identity classified an
opencode process as a worker with no opencode-specific handling. The identity
model is peer-kind-agnostic, which is what CB-308 needs when the roster
stretches across hosts.
Stage 5 verified live on the same run: /workers (CB-304) answers where the
13-day-old daemon 404'd, /metrics counted the delegation
(sends_total{outcome=replied} 1, replies_total{path=rendezvous} 1,
inbox_depth 0), and the audit log captured SPAWN/SEND/REPLY with correct roles.
Adds the opencode-free dogfood profile to the local bridged.yaml (gitignored;
recorded here for reproducibility) and docs/CB-402 §8 as-built.
Closes out single-host before the cross-host work. Sequenced BEFORE CB-308
deliberately: federation's own gating concern is the trust model, and it
inherits whatever identity shape lands here.
The finding this stage is built around: bridged had exactly ONE security
control, the loopback bind. ConnectionIdentity resolves a worker from its
connection (unforgeable), but every caller that was not a recognised worker
pane fell through to being treated as the PRIMARY -- the most privileged role
on the bus. Latent today; load-bearing the moment a bind widens.
CB-501 auth:
- Role/Principal/CallerResolver: connection identity first, bearer token
second, ANONYMOUS third. Inverts the old default so absence of identity
means nothing, not everything.
- Worker identity is never token-gated, so enabling auth cannot lock the
fleet out of bridge_reply.
- Constant-time token compare (MessageDigest.isEqual).
- validateAuthExposure(): a non-loopback bind under loopback-trust now
REFUSES TO START. Makes the dangerous config unrepresentable rather than
merely documented.
- TLS terminates at a reverse proxy by design (D3), not in the JVM.
CB-505 authz + audit, enforced on BOTH entry paths:
- The docs describe MCP as "a thin adapter over the REST core"; at code level
it is not. BridgeMcp calls MessageService directly, and /mcp is a raw
servlet on Jetty's context handler that never traverses Javalin's before
filter. Enforcing only at REST would have left /mcp open.
- Load-bearing rule is own-session-only: a worker may reply/ask only as
itself. Structurally true over MCP already; over REST the session id in the
URL path had simply been trusted.
- Audit: JSON lines to a dedicated appender, additivity=false. Never records
message content -- this bus carries source and prompts.
CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer
instead of the specced Micrometer, because this pom already hand-pins
jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against
skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated
dependency CVE gate could not be run (no JetBrains MCP server connected).
Instrumented at MessageService, the single funnel both surfaces share.
CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner.
Needs no contract-exclusion flag -- the pom's default-excludes profile
already sets excludedGroups=contract, so plain `mvn clean install` IS the
mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older).
CB-504 supervision: launchd agent (the real target -- this host is macOS,
there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds.
Ordering directives are advisory, so the actual fix is that startup now waits
up to 30s for the herdr socket and then serves degraded, instead of crashing
into a restart loop on a boot-order race.
Also fixes drift found while surveying:
- bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config
binds via plain Jackson with ignoreUnknown, so uncommenting it would have
been silently dropped and the default kept. Now camelCase, with a test that
loads the shipped example and one that pins every documented knob's
spelling -- no test had ever loaded that file.
- Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay,
gitTokenEnv, gitHostEnv, configDir, primary:).
- README "Next" listed bridge_ask and session lifecycle as upcoming; both
shipped long ago.
- docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre-
implementation" for work already merged.
307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS.
Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check
could not be run -- no JetBrains/intellij-index MCP server is connected this
session. mvn clean install is the only gate that ran.
Clarifies the open fork from §4/§9: Development A (sandbox launcher) and CB-308
(per-host federation) COMPOSE — each host runs a bridged gateway whose launcher
spawns agents into that host's LOCAL sandboxes; the broker moves messages +
presence, never keystrokes.
The forcing fact: delivery is herdr keystroke-injection into a locally-owned PTY,
so "a sandboxed agent on another host" ≡ "a sandbox spawned by that host's
gateway" (a remote container with no local herdr can't be injected into). Rules
out a central daemon reaching remote PTYs.
Adds Figure 11 (composed topology) + Figure 12 (remote-delegation sequence: the
local?inject:publish fork with a sandboxed far side — both injection points stay
local, only the middle hop crosses the broker), the two forced reachability
changes (host-routable mcpUrl; PTY in the local gateway's herdr), and a
maps-to-existing-seams table (CB-308 gateway × Dev-A launcher, CB-117 reap,
CB-303 container lifecycle, CB-308 #5 trust). No new pillars. Both diagrams
mmdc-validated; §4/§9 updated to point at §11.
Scopes the lead's next direction on top of the peer-launcher arc, as a
proposal (ticket split deferred):
- A · sandboxed, role-specific workers — a sandbox is a *placement*, so it
slots into the CB-401 SPI as a new kind: sandbox adapter exactly the way
CB-402's opencode slotted in as a new provider (SPI proven placement-neutral,
not just provider-neutral). Ownership line held: bridge launches INTO a
peer-owned image, never provisions the IDE/dev-tools inside it.
- B · main-agent pairs (Opus + cloud) — both mains are MCP clients so neither
can be called into; each needs a pull inbox → depends on CB-308's per-agent
channels. PrimaryRegistry single-slot → multi-slot.
- C · orchestrator tier — SessionManager recursed one tier up
(orchestrator:mains :: main:workers) + context scoping; re-roots the human
from a live primary to the orchestrator.
10 mermaid diagrams (component + sequence per development, ownership guardrail,
staging graph), all mmdc-validated and theme-safe. §7 pins the bus-vs-env-manager
boundary as an acceptance criterion on A; §8 stages A → CB-308 substrate → B → C.
Design-only. Audits the as-built Claude/herdr coupling (concentrated in
WorkerService), defines a PeerLauncher SPI + opaque PeerHandle + capability
model so the bus delegates peer materialization to a config-selected adapter.
ClaudeCodeLauncher = adapted WorkerService. Stages A/B/C with the Stage-C
plugin-loading security gate called out. No production code touched.
Worktree is code-only isolation; SessionManager hydrates it to full config
parity (overlay untracked local settings/.mcp.json/.env) so a worker is a
full peer of the primary, differing only in the LLM provider. Worker commits,
pushes over SSH, and opens its own PR (gitea REST + repo-scoped token).
The as-built audit surfaced that WorkerService keeps no registry of what it
spawned (its own Javadoc: 'there is no registry; list() only asks herdr').
CB-301 adds a SessionManager wrapping WorkerService: an authoritative in-daemon
roster with a per-session lifecycle FSM (SPAWNING/READY/BUSY/DONE/RELEASED/
FAILED), deterministic release, and recycle (= release + fresh acquire, no
reuse). Leaves clean seams for CB-302 (checkpoint on release), CB-303 (idle_ttl/
context_cap/drain policy over roster), CB-304 (bridge_list reads roster).
A worker now opens the same directory the primary is in, unless told otherwise.
Resolution: explicit spawn cwd → per-profile config cwd → the primary's cwd
(auto-detected from the bridge_spawn caller's PID via lsof -d cwd) → the daemon's
cwd. Never $HOME.
Mechanism (found by live probe, corrects the earlier assumption): an agent.start
pane does NOT inherit its tab's or workspace's cwd — it starts in $HOME. herdr's
agent.start honours an (undocumented) cwd param, so the resolved cwd is threaded
onto agent.start {cwd} (both tab and pane placement), not tab.create.
Surfaces: bridge_spawn {cwd?} + auto-detect via ConnectionIdentity.resolve (peer
PID) + ProcessCwdLookup (lsof); REST POST /workers ?cwd= / body cwd; per-profile
'cwd:' config. Validated live: explicit cwd → worker rooted there; no cwd over
REST → daemon cwd, not $HOME. Also clears the ccs folder-trust prompt when the
project dir is already trusted (see docs/Worker-Startup-and-Trust.md).
Documents (a) the rule that a worker inherits the PRIMARY's working directory,
never $HOME — resolved from an explicit cwd, else the bridge_spawn caller's PID
(lsof -d cwd), else the daemon cwd; herdr's seam is workspace.create {cwd}, since
agent.start has no cwd; (b) ccs (Claude Code) as the assumed launcher and its
folder-trust model (hasTrustDialogAccepted per project in each instance's
.claude.json), so trusting the project dir once per profile clears the prompt;
(c) that other CLIs have their own startup gates, documented per launcher.
Marks the cwd-inheritance as target design (follow-up), not yet wired.