wiki: add chapter 11 Features, and stop claiming Stage 5 is finished

Roadmap said Stage 5 was "✅ All landed" and listed CB-501–505. Twenty tickets
shipped after that line was written (CB-506…CB-525) and none of them appeared
anywhere in the wiki — so the page was not merely incomplete, it asserted
something false.

Two fixes, because there were two problems:

- The stage row is corrected and a "Stage 5 continued (as-built)" section records
  all twenty, split by who needs them: capabilities, contract changes, and the
  quality infrastructure that makes the rest trustworthy. It also carries the
  CB-521 version-coupling warning — /healthz can be green while every spawn
  fails, because health pings herdr without checking that the adapter and the
  binary agree on a protocol.

- Chapter 11 is new, and it is the structural fix. Chapters 1–10 are organised by
  design topic and every one answers "how is this built / why this way". None
  answers "what can it do and how do I turn it on", so a shipped knob like
  `placement: weighted` had no page that wanted it and landed nowhere. Eleven
  entries, each: what it does · the knob · why it exists · the gotcha. The `why`
  line is the load-bearing one — it is what stops a decision being re-litigated
  a month later.

Written from verified behaviour only; six entries that still need their config
surface confirmed against the code are listed as an explicit backfill list rather
than guessed at.

Mermaid block rendered with mermaid-cli to confirm it parses.
Dai Ha
2026-08-09 19:38:22 +02:00
parent 0c896eb49b
commit 1710a77827
3 changed files with 266 additions and 1 deletions
+219
@@ -0,0 +1,219 @@
# 11 — Features
**What this page is for.** The other chapters answer *how is this built* and *why this way*. This
one answers **"what can it do, and how do I turn it on"** — one entry per operator-facing
capability, so a feature that shipped six weeks ago is still findable without reading a design doc
or a commit log.
**What belongs here.** A capability an operator can *use, configure, or observe*: an MCP tool, a
`bridged.yaml` knob, an endpoint, or a behaviour visible from outside the daemon. Internal contract
changes go to [Implementation](9-Implementation); test and coverage work is a
[Roadmap](8-Roadmap) line. If a change adds none of those, it has no entry here — that is a normal
outcome, not an omission.
**What every entry states, in this order:** what it does · how you turn it on · why it exists ·
the gotcha. The *why* is the load-bearing line — it is what stops a decision being re-litigated in
six weeks, and the table alone will not carry it.
## Index
| Capability | Turn it on with | Since | Code |
|---|---|---|---|
| [Ask the bridge who you are](#ask-the-bridge-who-you-are) | `bridge_whoami` | CB-517 | `mcp/BridgeMcp` |
| [Primary inside a herdr pane](#primary-inside-a-herdr-pane) | `primary.terminal:` | CB-522 | `auth/CallerResolver` |
| [Weighted worker placement](#weighted-worker-placement) | `placement: weighted` + `weight` / `maxLoad` | CB-518 | `placement/` |
| [Give workers a toolchain](#give-workers-a-toolchain) | per-profile `env:` | CB-511 | `worker/HerdrPeerLauncher` |
| [Isolated worktree per worker](#isolated-worktree-per-worker) | `bridge_spawn{worktree, ticket}` | CB-301-ext | `session/GitWorktrees` |
| [Worker tool-surface isolation](#worker-tool-surface-isolation) | automatic | CB-525 | `session/GitWorktrees` |
| [Worker opens its own PR](#worker-opens-its-own-pr) | `gitTokenEnv:` / `gitHostEnv:` | CB-302 | `worker/HerdrPeerLauncher` |
| [Session lifecycle caps](#session-lifecycle-caps) | `lifecycle:` | CB-303 | `session/SessionManager` |
| [Durable reply inbox](#durable-reply-inbox) | `broker:` | CB-307 | `msg/AmqpReplyInbox` |
| [Pin an opencode endpoint](#pin-an-opencode-endpoint) | profile `baseUrl:` | CB-508 | `worker/OpenCodeLauncher` |
Nearly every knob above lives in one file, on one profile:
```mermaid
flowchart LR
Y["bridged.yaml"] --> G["bind / auth / primary.terminal"]
Y --> B["broker"]
Y --> W["workers:"]
W --> P1["profile: gx10"]
W --> P2["profile: ollama"]
P1 --> K["placement · weight · maxLoad<br/>env · gitTokenEnv · parityOverlay<br/>configDir · cwd · argv"]
P2 --> K
Y --> L["lifecycle"]
```
*The configuration surface: daemon-wide settings, then one block per worker profile. A profile is
a backend, not a host.*
---
## Ask the bridge who you are
**What.** `bridge_whoami` returns `{"role":"primary"}` or `{"role":"worker", sessionId, profile,
worktree, branch}`.
**On.** Always available; no configuration.
**Why.** Every rule in the bridge charter is role-conditional, and both roles read the same
`CLAUDE.md` — a worker runs in a worktree of the same repo, so it inherits the file verbatim. Before
this, a session had to *infer* its role from side-channels (the mount name, `ANTHROPIC_BASE_URL`,
the system prompt), each of which is one-way and some of which are absent for Claude-model workers.
The daemon already resolves the role from the connection for its authorization gate; `bridge_whoami`
just exposes that same answer, so guessing is never necessary.
**Gotcha.** The answer comes from the connection and cannot be forged or overridden by an argument.
If it disagrees with what you expect, the daemon is right and your assumption is wrong — check
`primary.terminal` next.
## Primary inside a herdr pane
**What.** Lets the orchestrating session run inside a herdr pane instead of an outside terminal.
**On.** `primary: terminal: term_<id>` — read the id from `bridge_whoami`, re-pin whenever the
primary moves panes.
**Why.** Caller identity resolves a loopback PID to its herdr pane, and the pane scan covers *every*
pane, not just bridged-spawned ones. So a primary living in a pane classified itself as a worker and
was refused spawn/send/stop — every verb it exists to call. The failure is **self-locking**: the
daemon can also *learn* the primary's terminal, but only from `bridge_send`/`bridge_spawn`, the
exact calls being refused. Only an operator-set pin breaks the cycle, which is why the pinned value
is consulted and the learned one deliberately is not.
**Gotcha.** A stale pin is silent. You are simply demoted to worker and every orchestration call is
refused. Re-pin after the primary changes panes, and note the daemon reads this at boot — a change
needs a restart.
## Weighted worker placement
**What.** Spreads unqualified spawns across profiles by weight, with a concurrency cap per profile
and failover to the next candidate when one is unreachable.
**On.** `placement: weighted` plus per-profile `weight:` and `maxLoad:`. The default `fixed` policy
reproduces the historical always-the-default-profile behaviour.
**Why.** Profiles differ in model and cost, not tier. Without weights, every unqualified spawn piles
onto one backend regardless of what it costs or how loaded it is.
**Gotcha.** An explicit `profile:` on `bridge_spawn` bypasses the policy entirely — placement only
governs *unqualified* spawns. Equal-weight candidates tie-break on **YAML definition order**, so
that order is load-bearing config, not cosmetics (CB-524).
## Give workers a toolchain
**What.** Propagates the daemon's own `PATH` to every worker, plus a literal per-profile `env:` map.
**On.** Automatic for `PATH`; add `env: {JAVA_HOME: ..., ...}` on a profile to extend or override.
**Why.** A worker's environment does **not** come from your shell. `bridged` hands herdr an explicit
env map and herdr merges it into *its own* process env — so before this, a worker inherited whatever
`PATH` the herdr server happened to be started with. On a long-lived herdr that can predate your
toolchain entirely, leaving workers unable to run `mvn` or `java` at all.
**Gotcha.** Adapter-owned variables win over `env:` — the `ANTHROPIC_*`/`CLAUDE_*` wiring is applied
after it, so an `env:` entry cannot repoint a worker past the `SubscriptionGuard`. Since the default
is the *daemon's* `PATH`, start the daemon with a good one (see the `PATH` lines in
`deploy/dev.ltms.bridged.plist` and `deploy/bridged.service`).
## Isolated worktree per worker
**What.** Provisions a git worktree on its own branch per worker, and copies a configurable set of
local config files in ("parity overlay") so the worker sees the same local setup.
**On.** `bridge_spawn{worktree: true, ticket: "cb-123"}`; overlay list via `parityOverlay:`
(default `[.claude/settings.local.json, .env, .envrc]`).
**Why.** Parallel workers editing one checkout collide. A worktree gives each its own branch and
files at the cost of a checkout.
**Gotcha.** Release removes the checkout but **keeps the branch** — unmerged work survives a
teardown. Overlay-copied files that are tracked get `--skip-worktree` so they never read as pending
changes. Never add `.mcp.json` to the overlay; see the next entry for why.
## Worker tool-surface isolation
**What.** A provisioned worktree's project `.mcp.json` is neutralized to an empty server map, so a
worker's tools are only what its launcher mounts (the bridge).
**On.** Automatic at provisioning. Nothing to configure.
**Why.** The repo commits a `.mcp.json` declaring the primary's IDE servers, so a fresh checkout
mounted them *and* the parity overlay copied the primary's own copy on top. Those servers are bound
to the primary's IDE project, so every path they return points into the **primary's checkout**. This
is not hypothetical: a worker made all 59 of its edits in the primary's tree while running `mvn`
against its worktree — so every build it ran was of code that did not contain its changes, and every
build passed.
**Gotcha.** This is why a worker cannot run IDE diagnostics and `mvn` is its only verification. That
is deliberate — the bridge is a message bus, and the primary is the gate. Never accept a worker's
claim about a check it had no way to run.
## Worker opens its own PR
**What.** Injects a repo-scoped forge token so a worker can commit, push over SSH, and open its own
pull request at checkpoint.
**On.** `gitTokenEnv:` (host env var holding the token) and optionally `gitHostEnv:` on a profile.
**Why.** Opt-in by design: omit it and the worker gets no PR-create grant, while push over SSH still
works.
**Gotcha.** Use a **minimal `write:repository` token**, never an admin one. The token can create a
PR but must not be able to merge — the primary is the gate, and a worker that can merge is not
gated.
## Session lifecycle caps
**What.** Reaps idle sessions, caps turns per session, and drains cleanly on shutdown.
**On.** `lifecycle: { idleTtlSeconds, contextCap, drainTimeoutSeconds, clearAfterTurn }`.
**Why.** Workers are disposable but not free; without caps an abandoned session holds a pane and a
context indefinitely.
**Gotcha.** Read at boot — changes need a daemon restart.
## Durable reply inbox
**What.** A worker's reply survives with no waiter attached: it is queued and collected later by
`bridge_poll` / `bridge_ack`, over a real AMQP broker when one is configured.
**On.** `broker:` pointing at an AMQP URI; omit it for an in-memory inbox.
**Why.** A blocking `bridge_send` is capped by the *caller's* MCP client timeout (~60s), far below a
real task's runtime. Without a durable inbox, a reply arriving after that window lands nowhere.
**Gotcha.** The broker is **LavinMQ**, not RabbitMQ. Point it at a stray local RabbitMQ and you are
writing into someone else's broker. Also see the known hole: an async ticket that times out while
the session is still BUSY currently discards the later completion rather than parking it.
## Pin an opencode endpoint
**What.** An opencode profile can target its own OpenAI-compatible endpoint.
**On.** `baseUrl:` on a `kind: opencode` profile.
**Why.** opencode is provider-agnostic and shares none of Claude's private seams — no
`ANTHROPIC_BASE_URL`, no `SubscriptionGuard`, no `--mcp-config`. It is the adapter that proves the
`PeerLauncher` SPI is genuinely provider-neutral rather than Claude-shaped.
**Gotcha.** Because it bypasses `SubscriptionGuard`, the guard's allowlist does not protect this
path — the endpoint you name is the endpoint it uses.
---
## Backfill status
This page was started after the fact, so it is **not yet complete**. Entries above are written from
verified behaviour. Still to catalogue — each needs its config surface and gotcha confirmed against
the code before it earns an entry:
- `/metrics` and `/healthz`, and what each does *not* tell you (CB-502; `/healthz` reports herdr
reachability only — it went green while every spawn failed, see the CB-521 version-coupling note)
- bearer-token auth and the non-loopback-bind fail-fast (CB-501)
- the per-session authz table and the audit log (CB-505)
- launchd / systemd supervision (CB-504)
- multi-profile routing and `kind:` adapter selection (CB-305, CB-401/402)
- the reply push loop and its nudge budget (CB-307 step 2)
+46 -1
@@ -117,7 +117,7 @@ Compact scope; expand into detailed tickets when a stage starts (as Stage 1 is b
| **2** | `CB-201` envelope schema + codec · `CB-202` worker `bridge_reply` tool + reviewer skill · `CB-203` reply rendezvous (corr match; resolve on reply *or* the `working→idle` edge) · `CB-204` subscription guard via `ccs env` + allowlist · `CB-205` blocked-worker path (`bridge_ask`) |
| **3** | `CB-301` ✅ session manager (spawn/reuse/recycle) · `CB-301-ext` ✅ per-worker git worktree + config-parity overlay (`97ecc71`) · `CB-302` ✅ worker checkpoint — **shipped as commit→push→**_**worker-opened PR**_ (`64e70ef`: repo-scoped forge-token injection + the implementer skill), which **supersedes** this row's original `STATE.md`-file framing; see `docs/Worker-Git-Workflow.md` · `CB-303` ✅ `idle_ttl`/`context_cap`/drain · `CB-304` ✅ `bridge_list` roster+live (`9fe04bf`) · `CB-305` ✅ multi-profile routing |
| **4** | `CB-401` ✅ PeerLauncher SPI Stage A — extracted in-tree, one adapter (`ClaudeCodeLauncher`), core uses the `PeerLauncher` interface, main @ `3aa69a9`. `CB-402` ✅ **Stage B complete** (`ded226a`, dogfooded 2026-07-29) — `OpenCodeLauncher` as the SPI-proving second adapter: shares none of Claude's private seams (no `ANTHROPIC_BASE_URL`, no `SubscriptionGuard`), mounts the bridge MCP via a generated `OPENCODE_CONFIG`, and is routed by `kind:` through `CompositePeerLauncher`. Live spawn→send→`bridge_reply`→teardown verified against opencode 1.18.5. Stage C — dynamic external plugin loading, future, gated by trust/capability model. |
| **5** | ✅ **All landed.** `CB-501` bearer auth + non-loopback-bind fail-fast (TLS at a proxy, not in-JVM — see `docs/CB-5xx-Hardening.md` D3) · `CB-502` `/metrics` (zero-dependency Prometheus renderer, D4) + `/healthz` · `CB-503` mock-socket CI (`.gitea/workflows/ci.yml`) · `CB-504` launchd agent + systemd unit + herdr-socket startup wait · `CB-505` per-session authz table + audit log |
| **5** | ✅ **CB-501–505 landed** (the stage as originally scoped): `CB-501` bearer auth + non-loopback-bind fail-fast (TLS at a proxy, not in-JVM — see `docs/CB-5xx-Hardening.md` D3) · `CB-502` `/metrics` (zero-dependency Prometheus renderer, D4) + `/healthz` · `CB-503` mock-socket CI (`.gitea/workflows/ci.yml`) · `CB-504` launchd agent + systemd unit + herdr-socket startup wait · `CB-505` per-session authz table + audit log. **The 5xx line did not stop there** — `CB-506`…`CB-525` shipped after this row was written; see [Stage 5 continued](#cb-506525--stage-5-continued-as-built) below and [Features](11-Features) for the operator-facing ones. |
## CB-401 — Peer Launcher SPI (Stage 4)
@@ -180,6 +180,51 @@ default: `ANONYMOUS` is now the fallback and `PRIMARY` must be established.
the session id in the URL path). Audit lines are JSON to a dedicated appender and **never carry
message content**.
## CB-506–525 — Stage 5 continued (as-built)
The 5xx line kept running after the stage's original five tickets. Two things drove it: **dogfooding
the bridge on its own development** surfaced defects the test suite could not (a primary that
classified itself as a worker, a worker that edited the wrong checkout), and a **coverage push**
turned "it works when I try it" into guarded behaviour. Operator-facing entries are catalogued in
[Features](11-Features); this section is the ticket-level record.
**Capabilities** — what an operator gained.
| | |
|---|---|
| `CB-508` | an opencode profile can pin its own OpenAI-compatible endpoint |
| `CB-511` | workers get a real toolchain — the daemon's `PATH` is propagated, plus a per-profile `env:` map. Before this a worker inherited whatever `PATH` the herdr *server* was started with, which on a long-lived herdr can predate your toolchain entirely and leave workers unable to run `mvn` at all |
| `CB-517` | `bridge_whoami` — a session asks the daemon for its own role instead of inferring it; and the `CLAUDE.md` bridge block became a portable charter copied verbatim into every project that mounts the bridge. LavinMQ pinned as a durable, self-restarting broker |
| `CB-518` | weighted placement (`placement: weighted`, per-profile `weight` / `maxLoad`), smooth weighted round-robin with failover to the next candidate |
| `CB-521` | herdr adapter ported to **protocol 19** (herdr 0.8.0) — a hard version coupling, see the note below |
| `CB-522` | `primary.terminal` — the primary may run *inside* a herdr pane. Without the pin the pane lookup reads it as a worker and refuses every orchestration verb, and the failure is self-locking: the learned terminal is populated by the very calls being refused |
| `CB-525` | a provisioned worktree's tool surface is isolated to what its launcher mounts |
**Contracts** — internal shape changes a maintainer needs to know.
| | |
|---|---|
| `CB-507` | fixed an NPE (HTTP 500) on a worktree spawn with no cwd, plus regression tests |
| `CB-516` | a delegation now *fails* when its worker session is released, rather than hanging |
| `CB-519` | `PeerHandle.id()` is a host-unique opaque UUID, deliberately decoupled from the herdr pane id: the id is the registry/routing key and must never collide across daemon processes on one host, while the pane id stays a launcher-private placement/teardown coordinate |
| `CB-520` | `ReplyInbox` split into explicit own/release and publish halves |
| `CB-524` | placement is reproducible across restarts — `Map.of`/`Map.copyOf` iteration order is salted per JVM run, which silently discarded YAML definition order and made equal-weight placement differ run to run |
**Quality infrastructure** — invisible, but it is why the above is trustworthy.
`CB-505 fix` audit lines are valid JSON · `CB-506` the test suite stays out of the production audit
log · `CB-509` JaCoCo coverage reporting · `CB-510` SessionReaper 0 → 86.7% · `CB-512`
`bridged_push_nudges_total` increments wired · `CB-513` the MCP-side authorization gate, 27.4 →
57.4% · `CB-514` MessageService timeout/answer/poll/lock edges · `CB-515` the turn-attribution
guards regression-protected · `CB-521` the AMQP contract test runnable both locally and in CI.
> **Version coupling (CB-521).** `bridged` speaks one herdr wire protocol; the adapter is ported
> wholesale on a bump, with no negotiation or compat shim. The daemon can therefore be perfectly
> healthy — `/healthz` ok, `bridge_whoami` resolving, profiles listed — while **every spawn fails**,
> because health only pings herdr and never checks that the adapter and the binary agree. Symptom:
> `invalid_request: missing field <x>`. After any restart onto a jar carrying an adapter change,
> verify with a real `bridge_spawn`, not with `/healthz`.
## Delivery reliability & multi-host (CB-306 – CB-308)
A cross-cutting track that came out of a **"communication break" review** of the reverse
+1
@@ -14,6 +14,7 @@
8. [Roadmap](8-Roadmap) — stages, tech stack, tickets
9. [Implementation](9-Implementation) — as-built code map · classes · flows · state machines
10. [Cross-Host Messaging](10-Cross-Host-Messaging) — broker topology · exchanges · queues per entity
11. [Features](11-Features) — what it can do · the knob that turns it on · why · the gotcha
---
🟢 herdr-centric `bridged` · AgentAPI = fallback