docs: add chapter 13, the operator user guide, for release 1.1

The wiki had twelve chapters and none of them told an operator how to run
the thing. Chapters 1-3 explain why the design is what it is, 9-12 explain
how the code is put together, 11 lists capabilities. The two pages that were
meant to cover bring-up and day-2 - 4-Setup and 5-Operations - were never
written past their scope note, and every technical detail in them had gone
wrong: a decommissioned model host, port 8080, herdr protocol 14, a systemd
unit that does not exist, Redis Streams and NATS that were never built,
"no per-session authz yet" after Authz shipped, and recycle() events that
have no code behind them.

So this adds 13-User-Guide.md, written against the running system on
2026-08-17, and points the two stubs at it rather than leaving wrong claims
in place.

The guide covers:

  1. what this is and, more usefully, the five things it is NOT, each with
     the reason it is not that;
  2. install - herdr (check the PROTOCOL number, not the version), the
     daemon, the login-shell rule for secrets.sh, and the lead's tab label;
  3. configure - the four knobs that cost money, the live profile table with
     who pays for each, the gateway paths, and memberCredentials' two halves;
  4. run - the redeploy script, and the four checks that go beyond /healthz,
     because health is green while every spawn fails;
  5. delegate - the eleven tools, the spawn-all-then-send-all rule, the ~60s
     client cap on a blocking send, and the authz table;
  6. when it breaks - twelve traps hit for real this year, grouped by
     bring-up, losing a member's work, and merging a member's work;
  7. where to look next.

Home.md is corrected too: it claimed members launch against ollama.ltms.dev,
a host that no longer exists (the gateway is llm.ltms.dev), it framed the
system as Claude-only with one worker, it listed 8 of the 13 pages, and its
status still said "Design".
Dai Ha
2026-08-17 16:17:22 +02:00
parent 4e71c5ca5c
commit f47978cc5d
5 changed files with 520 additions and 80 deletions
+444
@@ -0,0 +1,444 @@
# 13 — User Guide
> **Status: 🟢 Written against the running system, 2026-08-17 (release 1.1).**
> Every command here was run on this host. Where a claim is not measured, it says so.
>
> This is the operator page. Chapters 1–3 say *why the system is built this way*, 9–12 say *how the
> code is put together*, 11 lists *what it can do*. This page is the one you read when you have to
> bring it up, run it, or fix it. It assumes you have done this before and forgotten the details.
---
## 1. What this is, and what it is not
`bridged` is a **message bus between AI agent sessions**. One session orchestrates (the **lead**),
and it delegates work to **members** running in other terminal panes, often on other models and
other vendors. All traffic goes through `bridged`'s MCP tools. No session talks to another session,
to a broker, or to the network directly.
```mermaid
flowchart LR
LEAD["lead session<br/>(Claude Code, subscription)"]
subgraph BD["bridged — a plain Java daemon"]
SRV["MCP + REST<br/>policy, authz, lifecycle"]
INJ["injector<br/>status-gated"]
SRV --> INJ
end
HERDR["herdr<br/>owns panes and PTYs"]
M1["member pane<br/>claude-code"]
M2["member pane<br/>opencode"]
GW["llm.ltms.dev<br/>the one gateway"]
LEAD -->|"bridge_send"| SRV
M1 -.->|"bridge_reply"| SRV
M2 -.->|"bridge_reply"| SRV
INJ -->|"unix socket"| HERDR
HERDR --> M1
HERDR --> M2
M1 --> GW
M2 --> GW
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
class LEAD,GW ext
class SRV,INJ,HERDR core
```
*Figure 1 — who talks to whom. The lead never touches herdr; the members never touch each other.*
**What it is not**, and each "not" is a decision, not a gap:
| Not | Why |
|---|---|
| Not a build system | It does not compile, lint, or test your code. Members run whatever the repo gives them. The lead verifies. |
| Not an environment manager | It does not install toolchains or set up a member's language runtime. It only sets a small, named set of environment variables at spawn. |
| Not a Claude-only tool | `kind: claude-code` and `kind: opencode` are both first-class. A member can be GPT, DeepSeek, or Claude. |
| Not a scheduler with a budget | There is **no cost metering and no refusal on spend**. `maxLoad` per profile is the only throttle you have. See §3. |
| Not a security boundary against your own members | A member runs as you, on your machine. `memberCredentials` limits which secrets it sees. It does not sandbox anything else. |
**The two invariants.** Break either and nothing else matters.
1. The lead **never** sets `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN`. It stays on the
subscription. Only the daemon moves a member off it, at spawn.
2. The bridge is the **only** channel. Text printed in a pane reaches nobody. An answer that is not
in a `bridge_*` call is discarded silently.
---
## 2. Install
Four things must be true before anything works. In order, because each one depends on the last.
### 2.1 herdr
`bridged` does not own terminals. `herdr` does. `bridged` drives it over a Unix socket.
```bash
herdr --version # 0.8.0 on this host
curl -s http://127.0.0.1:8765/healthz
# {"status":"ok","herdr":{"version":"0.8.0","protocol":19}}
```
**The protocol number is the thing to check, not the version.** `bridged`'s adapter speaks one
herdr wire protocol. If herdr is upgraded and the protocol moves, `/healthz` still says `ok` —
because the socket connects — and **every spawn fails**. Green health with a broken fleet is the
normal way this breaks. See trap 2 in §6.
### 2.2 The daemon
Build and run from this repo. There is no POM at the repo root:
```bash
mvn -f bridged/pom.xml clean install # never pipe this — a pipe hides BUILD FAILURE
java -jar bridged/target/bridged.jar bridged/bridged.yaml
```
In practice you never run those two by hand. Use the script (§4).
### 2.3 Secrets, and the login shell
All credentials live in one file: `${SHARED_ENV}/tools/secrets.sh`. It is sourced only by a **login
shell**. This one fact causes more lost hours than anything else in the system.
- The daemon inherits `AI_GATEWAY_TOKEN` and `WORKER_GITEA_TOKEN` **from the shell that started it**.
- Start it from a non-login shell and both are empty. The daemon boots, `/healthz` is green, and
nothing complains.
- The failure surfaces hours later: a gateway profile gets a 401, or a member cannot open a pull
request.
`launchd` does not run a login shell either. That is the only reason
`scripts/bridged-launchd-wrapper.sh` exists — it `exec`s `zsh -lc` so the store gets sourced, while
keeping one process so launchd's PID tracking still works. Read its header; it explains the trap
better than this paragraph.
> On this host the daemon is **not** under launchd. It runs as a plain `java -jar` started by the
> redeploy script from a login shell. `launchctl list | grep bridg` returns nothing.
The daemon logs which required secret names resolved at startup
(`Bridged.reportRequiredSecrets`). Read those lines. But note the gap: it **skips profiles marked
`subscription: true`**, on purpose, because they need no token. So a green secret report says
nothing about your subscription profiles.
### 2.4 The lead's tab
The daemon finds the lead by its **herdr tab label**, and by nothing else:
```yaml
fleet:
leaders:
opus:
tab: "lead: opus" # must match the real tab label
```
Matching is case-insensitive and both ends are stripped. Internal spacing is **not** forgiving:
`"lead: opus"` and `"lead:opus"` are different. A session whose tab does not match is resolved as a
**worker**, and every orchestration call it makes is refused.
The old `terminal:` key is gone. A daemon at CB-579 or later **refuses to start** if `terminal:` is
still in the config, and also if `tab:` is missing. That refusal is deliberate — a silently demoted
lead was worse than a daemon that will not boot.
---
## 3. Configure
`bridged/bridged.yaml` is the live config. It is **gitignored**. `bridged.example.yaml` is the
tracked, documented copy. Two consequences you will meet:
- Members cannot see the live config. They work in worktrees of the tracked repo. So a change to
what a config **key means** can pass review, merge green, and break the running fleet — because
nobody who reviewed it could see the file it breaks. Check and fix the live config yourself at
merge time.
- The example file is the documentation. A key that is read by code but missing from the example is
a real defect (that is how `subscription:` stayed undocumented for months — CB-610).
### The knobs that cost money
This is the section to reread before you change anything.
| Knob | What it does | The cost |
|---|---|---|
| `subscription: true` | The member runs on **your own Claude plan**. No `ANTHROPIC_BASE_URL`, no token — it inherits your Claude Code auth. | Every spawn bills your plan and eats your usage limit. Off-subscription is the whole point of the daemon, so treat `true` as a deliberate exception. Mutually exclusive with `baseUrl`; setting both is refused at load. |
| `maxLoad` | Members allowed at once on this profile. | **The only throttle that exists.** There is no metering, no budget, no refusal on spend. When `maxLoad` is reached, a spawn is refused with `no capacity … no fallback` — it does not fall back to a cheaper profile. |
| `weight` | Share of automatic placement. | `weighted` placement is **not cheapest-first**. It spreads by ratio across every profile that has a free slot, so paid spawns happen while the free local profile is idle. The workaround here is `local.weight: 100`; the real fix is open as CB-589. |
| `fleet.reviewers` | Which profiles an unqualified reviewer spawn may land on. | Review is the fan-out step — several members per pull request. Leaving the free profile out of this pool was the single largest avoidable cost in the fleet. |
Live profiles on this host:
| Profile | Kind | Model | maxLoad | Pays |
|---|---|---|---|---|
| `local` | claude-code | deepseek-v4-flash via `llm.ltms.dev/anthropic` | 2 | gateway |
| `local-direct` | claude-code | deepseek-v4-flash direct to `gx00.gw:8000` | 2 | nothing, LAN only, `weight: 0` |
| `gx` | opencode | `gx/deepseek-v4-flash` via `llm.ltms.dev/v1` | 2 | gateway |
| `opus` | claude-code | `claude-opus-5` | 1 | **your Claude plan** |
| `sonnet` | claude-code | `claude-sonnet-5` | 3 | **your Claude plan** |
| `sol` | opencode | `openai/gpt-5.6-sol` | 1 | OpenAI |
| `terra` | opencode | `openai/gpt-5.6-terra` | 2 | OpenAI |
`sol` and `terra` share **one** OpenAI credential. When it hits its limit, both die at once, and an
opencode member dying looks like a short clipped reply rather than an error.
### The gateway
`llm.ltms.dev` is the one front door, and `AI_GATEWAY_TOKEN` is the single key.
- `kind: claude-code` → `https://llm.ltms.dev/anthropic`. **Never `/v1`** — on `/v1` the model's
thinking output silently disappears.
- `kind: opencode` → `https://llm.ltms.dev/v1`. The path is taken as-is.
Both known gateway ceilings were fixed on 2026-08-15: the body limit went 32 KiB → 32 MiB, and the
route timeout 60 s → 86400 s. The remaining risk is an Envoy HTTP/1.1 bug that can truncate a
stream with a `200` and no terminator. Our members cannot detect that; it looks like a short answer.
### Credentials a member can see
`memberCredentials` is deny-by-default. You name what is allowed; everything else known is replaced
with a sentinel string. Current state: **34 known, 7 allowed, 29 blocked**, measured inside a live
member pane on 2026-08-17.
It has **two halves**, and only both together work:
1. The launcher writes the member's environment at spawn.
2. A `BRIDGED_MEMBER`-guarded block at the **end** of `${SHARED_ENV}/tools/secrets.sh` re-applies it.
Half 2 is not optional. The member's pane runs a **login shell**, which re-sources the whole secret
store and overwrites whatever the launcher set. The guarded block must be the last thing in the
file, or the store overwrites it in turn.
The policy is re-read **on every spawn**, so an allow-list edit needs no restart. That is measured,
not assumed.
> Rule in force: only the lead and architects may use `GITEA_ACCESS_TOKEN`. Members get
> `WORKER_GITEA_TOKEN`, which is a minimal `write:repository` token, never the admin one.
---
## 4. Run, and prove it runs
### Start and restart
Use the script. Do not hand-roll the steps.
```bash
scripts/redeploy-bridged.sh --check # read-only: reports state, changes nothing
scripts/redeploy-bridged.sh # build, confirm drain, restart, verify
scripts/redeploy-bridged.sh --yes # skip the drain prompt (fleet already checked)
scripts/redeploy-bridged.sh --no-build # restart the jar you already have
```
Run `--check` first, always. It is the only thing that reports whether the forge token resolves,
and it reports that **without printing the value**.
The script builds **before** it stops anything, so a failed build never leaves the fleet down. It
waits for the old process to exit instead of assuming. It polls `/healthz`. And it anchors its log
checks to a marker taken before the restart, so old errors cannot be misread as new ones.
**A merge is not a deployment.** The running daemon holds the jar it started with. Code merged to
`main` does nothing until you rebuild and restart. Saying "shipped" about code the live daemon has
never loaded is a false report.
**Drain first.** `bridge_list`, collect anything you still want with `bridge_poll`, then
`bridge_stop` each member. A restart drops in-flight tickets and rendezvous. A member's report is
not recoverable once its ticket is gone.
### Verify — `/healthz` is not enough
`/healthz` proves the socket connects. It does not prove a spawn works, that identity resolves, or
that the new jar is the one running. Four checks, in order:
1. **A fresh boot line.** Confirm a new `bridged listening` line at the end of `bridged/bridged.out`,
dated after the restart. An old daemon that never died looks identical from outside.
2. **Deferred keys.** The startup log names which config keys it accepted and which it deferred. A
deferred key needing a restart is usually the whole reason you restarted. Read those lines rather
than assuming.
3. **Identity.** `bridge_whoami` must still answer `primary`. If the tab label changed, the lead is
now a worker and every orchestration call is refused.
4. **A real spawn.** Spawn one cheap member and stop it. This is the only check that catches a herdr
protocol mismatch.
> Restarting the daemon **cuts your own MCP mount**, and it does not reconnect. So you cannot run
> `bridge_whoami` from the session that restarted it. Ask the operator to run `/mcp` to reconnect.
> This is why the restart is done from the lead but verified after a reconnect.
---
## 5. Delegate
Eleven tools. This is the whole surface.
| Intent | Tool |
|---|---|
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` |
| See the fleet | `bridge_list` → `leads` + `members` |
| One member's state | `bridge_status{sessionId}` |
| Delegate, blocking | `bridge_send{sessionId, content}` |
| Delegate, long task | `bridge_send{sessionId, content, wait:false}` → ticket |
| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Collect a reply | `bridge_poll{ticket}` or `bridge_poll{target}` |
| Clear a reply from the inbox | `bridge_ack{target, msgId}` — **`target`, not `ticket`** |
| Member ends its turn | `bridge_reply{content}` |
| Member asks the lead | `bridge_ask{question}` |
| Tear down | `bridge_stop{paneId}` |
### The loop
```mermaid
sequenceDiagram
participant L as lead
participant B as bridged
participant M as member
L->>B: bridge_spawn — all units first
L->>B: bridge_send with wait false — then all sends
B-->>L: ticket
B->>M: injected when the member is idle
M->>B: bridge_reply
B-->>L: nudge into the lead's own pane
L->>B: bridge_poll by ticket
L->>B: bridge_ack by target and msgId
L->>B: bridge_stop by paneId
```
*Figure 2 — the delegation loop. Spawn is separate from send on purpose.*
**Spawn every unit first, then send them all.** Spawning and sending in one loop is how parallel
work silently becomes serial. It is the most common way this whole layer is wasted.
**Prefer `wait:false`.** A blocking `bridge_send` is capped by **your own MCP client timeout**,
about 60 seconds — far below any real task's runtime. The cap is in your client, not in the daemon,
so no server setting fixes it.
**Line 1 of every brief names a skill.** `Load the implementer skill.` or `Load the reviewer skill.`
Those skills are opt-in, and that line is what makes them reliable. The brief must be
self-contained: the member sees your message and the repo, and nothing of your context, your plan,
or your screen.
**Since CB-588, a `wait:false` ticket that finishes nudges your pane by itself.** You do not have to
poll on a timer. It needs an injectable pane and is capped at 5 pending nudges, so still watch
anything you cannot afford to lose.
### Who may call what
| Call | Allowed |
|---|---|
| spawn / stop / drain | lead only |
| send | lead or architect |
| reply / ask | any member, **only as itself** |
| read / status | any authenticated caller |
Identity comes from the connection, never from an argument. You cannot act as another session. A
call outside your role is refused, not queued.
### The lead's own job
Delegation does not delegate responsibility. These stay with you and are not delegable: the
conversation with the operator, decomposition, the final judgment, **verification**, and the
**merge**. Merging on a reviewer's word is delegating the merge by proxy.
A member cannot run your IDE tooling. Any forge tools it appears to have hold a deliberately blocked
credential and fail every call. And a piped command hides failures behind a zero exit code. So never
promote a member's "clean" to a fact — re-run the build yourself.
---
## 6. When it breaks
Twelve traps, all hit for real this year. Grouped by where they bite.
### Bring-up
**1. The daemon started from a non-login shell.**
Symptom: everything green for hours, then a member cannot open a pull request, or a gateway profile
gets a 401. Nothing logs it at the time. Fix: `scripts/redeploy-bridged.sh --check` is the only
thing that reports it. Restart from a login shell, or via the launchd wrapper.
**2. `/healthz` is green and every spawn fails.**
Cause: herdr was upgraded and its wire protocol moved past what the adapter speaks. The socket still
connects, so health is `ok`. Fix: check `protocol` in the `/healthz` body, and verify any restart
with one real spawn — not with health.
**3. The lead is resolved as a worker.**
Cause: the herdr tab label no longer matches `fleet.leaders.*.tab`. Every orchestration call is
refused. Fix: match the tab exactly — internal spacing counts. Note the old `terminal:` key is
removed; a current daemon refuses to start if it is still present.
**4. A merge is not a deployment.**
The running daemon holds its original jar. Rebuild and restart, then prove it with a **fresh**
`bridged listening` line. Also: the restart cuts your own MCP mount and it never reconnects, so you
cannot verify identity from that session — ask for `/mcp`.
### Losing a member's work
**5. The ticket expired.**
Ticket time-to-live is about 10 minutes. After that `bridge_poll{ticket}` returns
`timed_out_working`. The member is usually fine and its real answer arrives later — in the member's
**inbox**. Drain it with `bridge_poll{target}`, then `bridge_ack{target, msgId}`.
**6. Reading the reply inbox is destructive.**
`GET /sessions/{id}/replies` drains on first read. If you `curl` it through `head` or a parser that
dies, the reply is gone. Always redirect it to a **file** first.
**7. The idle reaper ate the report.**
A member that is `done` holds its report only until `lifecycle.idleTtlSeconds` (1800 s here). Then
the ticket, the pane, and the report are all gone, with no scrape fallback. Poll before the reaper.
**8. A second send to a busy member never lands.**
`timed_out_queued` means **never delivered** — the member is fine. Worse, the stale brief stays
queued and restarts the member when it next goes idle. Back up the member's commit first:
`git fetch <worktree> <branch>:refs/backup/x`.
**9. `liveStatus: working` is not progress.**
A member can report busy for hours doing nothing. Check file modification times in its worktree and
snapshot `git diff` before and after. Do not read its terminal — use `bridge_status`.
**10. Never brief a member to "ask me".**
`bridge_ask` blocks for about 55 seconds and no nudge extends it. It is also invisible to
`bridge_poll`. If you are running async, you will not see the question in time. Decide before you
delegate, or give the member an explicit default.
### Merging a member's work
**11. The reported branch is not the branch it committed to.**
`bridge_list` reports the branch **spawn provisioned**, not the one the member actually used. Check
`git -C <worktree> branch --show-current`, or you push an empty ref and the merge says
"Already up to date".
**12. Members see a months-old wiki, and cannot see the live config.**
`wiki/` is a submodule whose pointer is never advanced, so a member's checkout is a stale snapshot.
Never brief "read `wiki/…`" — paste the text, and commit the member's wiki entry yourself. Same
shape for `bridged.yaml`: it is gitignored, so a member cannot see the file its change may break.
### The general shape behind several of these
A checker that is **narrower than what everyone believes it checks**, with nothing saying so. The
credential probe hardcoded 31 names against a policy of 29 and exited `0`. A config test checked
only column-0 keys, so a nested key that bills your Claude plan stayed undocumented. A scheduled
sweep never ran once because its gate used a sentinel in a numeric field.
Two habits fix most of it: **make a checker print its own denominator** ("checked 26 of 29", not
"26 blocked"), and when a green check is the evidence for a claim, **ask what it does not cover
before believing it**.
---
## 7. Where to look next
| You want | Go to |
|---|---|
| Why the design is herdr-centric at all | [1 Architecture](1-Architecture), [3 Approaches](3-Approaches) |
| The daemon's own design | [2 Message Server](2-Message-Server) |
| The as-built code map, classes, state machines | [9 Implementation](9-Implementation) |
| What a capability does, the knob, why it exists, the gotcha | [11 Features](11-Features) |
| Orchestrating a mixed-vendor fleet | [6 Team](6-Team), [7 Use Cases](7-Use-Cases) |
| Adding a second host or a non-Claude peer | [10 Cross-Host Messaging](10-Cross-Host-Messaging), [12 Claude → OpenCode](12-Claude-to-OpenCode) |
| What is left to build | [8 Roadmap](8-Roadmap) |
| The rules every session must follow | `CLAUDE.md` in the repo — it loads into every session |
| Rendezvous, `bridge_ask`, detached delivery, turn-done fallback | `docs/MCP-Contract.md` **§6 only** — the rest of that page is a pre-build design doc and its tool names never caught up with the code |
| Every config key, documented | `bridged/bridged.example.yaml` |
**The one thing that is easy to forget.** This repo *is* the bridge, so the charter block in
`CLAUDE.md` is not documentation about someone else's system — it is the instruction surface this
codebase ships. A code change that quietly makes it untrue is an incomplete change. The block in
`CLAUDE.md` and the template in [7 Use Cases](7-Use-Cases) must stay byte-identical; there is a
check script in `CLAUDE.md` that proves it.
+16 -31
@@ -1,40 +1,25 @@
# 4. Setup
> **Status:** 🟠 Stub — scope defined, procedure not yet written. `bridged` is at the design
> stage ([Message Server](2-Message-Server)); concrete install steps land with **M0–M1** of the build plan.
> **Status: ⚫ Superseded by [13 User Guide](13-User-Guide) §2 (2026-08-17).**
> This page was written in the design era as a scope note and the procedure was never filled in.
> Its content had gone wrong in every detail — a decommissioned model host, port 8080, herdr
> protocol 14, a systemd unit that does not exist on this host, Redis Streams / NATS that were never
> built. Rather than leave those claims in place, the page now points at the written procedure.
This page will cover standing up the bridge on an off-subscription worker host.
**Go to [13 User Guide](13-User-Guide) → §2 Install.** It covers, against the running system:
## Scope (what this page will contain)
- herdr, and why you check the **protocol number** rather than the version;
- building and starting the daemon;
- the **login shell** rule for `${SHARED_ENV}/tools/secrets.sh` — the single most expensive trap in
bring-up, and why `scripts/bridged-launchd-wrapper.sh` exists;
- the lead's **tab label**, which is how the daemon resolves who the lead is.
1. **Prerequisites** — [herdr](https://herdr.dev) installed and its server running; a worker
`claude` that inherits your `CLAUDE.md`/hooks/skills/MCP; a reachable worker model
(`ollama.ltms.dev` or GX10 vLLM) with a bearer token.
2. **herdr** — start the headless server; confirm the socket at
`~/.config/herdr/herdr.sock` (or `HERDR_SOCKET_PATH`); verify with `ping` (expect
`protocol: 14` on herdr 0.7.0) and `workspace.list`.
3. **`bridged`** — deploy the binary, `bridged.yaml` (worker model, `base_url` allowlist,
auth token, bind address), and a **systemd** unit ordered *after* herdr.
4. **Worker session** — create the first worker pane with the env-prefixed launch line
(`ANTHROPIC_BASE_URL=https://ollama.ltms.dev ANTHROPIC_AUTH_TOKEN=… claude`); confirm the
subscription guard accepts it and `pane.process_info` shows the expected egress host.
5. **MCP mount (unified, both sides)** — register `bridged` as an MCP server on the primary
**and** each worker with one line —
`claude mcp add --transport http bridge http://127.0.0.1:8080/mcp` (or a shared
`.mcp.json` / `CLAUDE.md` entry every session inherits). The primary then delegates via the
`bridge_send` tool (single blocking call per delegation) and workers reply via
`bridge_reply`. Same-host needs nothing more — `bridged` delivers async by injecting an idle
pane. Only for a **split-host** primary (not a herdr pane) add a `Stop`-hook that long-polls
**`bridged`** (never a broker) for detached wake-ups.
6. **Topology choice** — single-host vs split-host (see [Message Server](2-Message-Server) → *Deployment
model*). If durability or cross-host async is needed, configure `bridged`'s **internal** queue
(Redis Streams / NATS); it stays behind the gateway — no Claude session connects to it.
## The one rule that was already right on this page
## Non-negotiable during setup
The **primary** host/process must **never** be given `ANTHROPIC_BASE_URL`. Only worker panes
carry it. See [Architecture](1-Architecture) → *Subscription boundary*.
The **lead** host and process must **never** be given `ANTHROPIC_BASE_URL` or
`ANTHROPIC_AUTH_TOKEN`. Only member panes carry them, and only the daemon sets them, at spawn. See
[Architecture](1-Architecture) → *Subscription boundary*.
## Related
- [Message Server](2-Message-Server) · [Architecture](1-Architecture) · [Operations](5-Operations) · [Approaches](3-Approaches)
- [13 User Guide](13-User-Guide) · [2 Message Server](2-Message-Server) · [1 Architecture](1-Architecture) · [11 Features](11-Features)
+25 -29
@@ -1,39 +1,35 @@
# 5. Operations
> **Status:** 🟠 Stub — scope defined, runbook not yet written. Fills in as `bridged` reaches
> **M4 — Harden** (auth/TLS, metrics, systemd) in the [Message Server](2-Message-Server) build plan.
> **Status: ⚫ Superseded by [13 User Guide](13-User-Guide) §4 and §6 (2026-08-17).**
> This page was a design-era scope note; the runbook was never written. Several of its claims were
> overtaken by the build — `recycle()` events do not exist, per-session authorization shipped
> (`auth/Authz`), and the internal broker is optional rather than required.
Day-2 runbook for a running bridge.
**Go to [13 User Guide](13-User-Guide):**
## Scope (what this page will contain)
- **§4 Run, and prove it runs** — `scripts/redeploy-bridged.sh`, why a merge is not a deployment,
draining before a restart, and the four checks that go beyond `/healthz`.
- **§6 When it breaks** — twelve traps hit for real this year, grouped by bring-up, losing a
member's work, and merging a member's work.
- **Health** — `GET /healthz` liveness, `GET /metrics` (Prometheus), reading live
`agent_status` per session via `GET /sessions`, and confirming the **MCP endpoint** is
reachable from both the primary and the workers (`claude mcp list` shows `bridge` connected).
- **Restart & recovery** — ordered restart (herdr before `bridged`); how `bridged`
re-attaches to existing panes via `workspace.list`/`pane.list`; internal broker/queue replay of unacked items. See
[Architecture](1-Architecture) → *Failure modes & single points of failure* for what each outage costs.
- **Model swaps** — repoint a worker to a different `base_url`/model by recycling its pane
(Ralph loop); the subscription guard re-validates the new host against the allowlist.
- **Lifecycle / context ceilings** — observing recycle events; confirming workers externalize
state (git + `STATE.md`) before a recycle so continuity survives (see [Message Server](2-Message-Server) →
*Worker session lifecycle*).
- **Troubleshooting** — stuck `working` (model endpoint down), `blocked` awaiting input,
injection collisions on a hand-driven pane, envelope-vs-scrape reply mismatches.
- **Security ops** — token rotation, keeping the port off public interfaces; note that one
`bridged` is currently **one trust domain** (no per-session authz yet — [Message Server](2-Message-Server)
→ *Security*).
## What actually shipped, against what this page predicted
## Guardrails to watch (from [Architecture](1-Architecture))
| This page said | What is true now |
|---|---|
| "no per-session authz yet" | Per-session authorization shipped. `auth/Authz` holds the role table; identity comes from the connection, never from an argument. |
| "observing recycle events" | There is no `recycle()`. Lifecycle is an idle reaper plus a context cap — `lifecycle.idleTtlSeconds`, `lifecycle.contextCap`, `lifecycle.drainTimeoutSeconds`. |
| "internal broker/queue replay" | The broker is optional. With `broker:` commented out the daemon uses an in-memory inbox. When it is on, it is **LavinMQ**, never RabbitMQ. |
| "`GET /healthz` liveness" | Still true, and still not sufficient. Health can be green while every spawn fails, because the socket connects but the herdr wire protocol has moved. Check `protocol` in the body and verify with a real spawn. |
| "recycling a pane to swap models" | Spawn a member on a different profile instead. Profiles are the unit of model choice. |
- The **primary must never perpetual-poll** — quota burn. Async wake-ups are `bridged`
inject-on-idle by default; the `Stop`-hook is only the split-host-primary exception, and it
polls `bridged` (never a broker).
- Cross-agent ping-pong needs a round/turn budget — enforced centrally in `bridged` (sole
gateway), not per-session sentinels.
- The **internal** broker/queue must run with **ack + visibility timeout + consumer groups**
so a mid-turn crash re-delivers instead of dropping.
## Guardrails that still hold
- The lead must **never** busy-poll a member's pane — it burns your subscription for nothing. Async
wake-ups are injected into the lead's own pane when a ticket goes terminal (CB-588).
- Keep the port off public interfaces. The daemon binds loopback and trusts it; the exposure check
only fires on a non-loopback bind.
- Rotate tokens. Members get a minimal `write:repository` forge token, never the admin one.
## Related
- [Message Server](2-Message-Server) · [Architecture](1-Architecture) · [Setup](4-Setup) · [Approaches](3-Approaches)
- [13 User Guide](13-User-Guide) · [2 Message Server](2-Message-Server) · [1 Architecture](1-Architecture) · [11 Features](11-Features)
+32 -18
@@ -1,12 +1,16 @@
# claude-bridge
A **subscription-safe bridge** that lets a primary **Claude Code (Opus 4.8, on Pro/Max)**
session drive a **secondary Claude agent running a different model** via its own
`ANTHROPIC_BASE_URL` — without ever putting a proxy on the primary.
A **subscription-safe bridge** that lets a lead **Claude Code** session on Pro/Max delegate work to
**members running on other models and other vendors** — without ever putting a proxy on the lead.
> **New here, or here to operate it? Start at [13 User Guide](13-User-Guide).** That page is written
> against the running system: install, configure, run, delegate, and the traps. This page explains
> the shape of the design and why it was chosen.
> Sibling of [`crush-bridge`](https://git.ltms.dev/systems/vms) (Tier 2 → headless Crush
> on GX10 DeepSeek). `claude-bridge` keeps the worker a *real Claude Code process* so it
> inherits `CLAUDE.md`, hooks, skills, and MCP — just pointed at a cheaper/local model.
> on GX10 DeepSeek). `claude-bridge` keeps a Claude member a *real Claude Code process* so it
> inherits `CLAUDE.md`, hooks, skills, and MCP — just pointed at a cheaper model. It also runs
> non-Claude members (`kind: opencode`) as first-class peers.
## Leading approach — herdr-centric message server (`bridged`)
@@ -17,9 +21,9 @@ workers mount** — one unified Claude setup and the **sole communication gatewa
stays for non-Claude clients; any broker is `bridged`-internal, below the gateway). herdr owns
the PTYs, multiplexing, persistence, and **agent-status
events**; `bridged` owns policy (subscription boundary, session lifecycle, status-gated
delivery) and the client contract. The worker `claude` launches with
`ANTHROPIC_BASE_URL=https://ollama.ltms.dev` + a bearer token; the primary Opus stays
env-clean and calls `bridged`'s MCP tools.
delivery) and the client contract. A Claude member launches with `ANTHROPIC_BASE_URL` pointed at
the gateway, `https://llm.ltms.dev/anthropic`, plus a bearer token; the lead stays env-clean and
calls `bridged`'s MCP tools.
```mermaid
flowchart LR
@@ -30,8 +34,8 @@ flowchart LR
SRV --> CLI
end
HERDR["herdr<br/>panes · agent-status"]
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
M["ollama.ltms.dev<br/>(worker model)"]
W["member pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
M["llm.ltms.dev<br/>(the one gateway)"]
OPUS -->|"MCP bridge_send (blocks)"| SRV
W -.->|"MCP bridge_reply"| SRV
@@ -72,14 +76,24 @@ Read in order (the sidebar mirrors this):
1. **[Architecture](1-Architecture)** — process model, the two invariants, two traffic modes
2. **[Message Server](2-Message-Server)** — 🟢 **`bridged`**, the herdr-centric message server (primary approach)
3. **[Approaches](3-Approaches)** — herdr-centric vs AgentAPI vs Agent SDK vs bus/tmux (research matrix)
4. **[Setup](4-Setup)** — running herdr + `bridged` + a worker pointed at `ollama.ltms.dev`
5. **[Operations](5-Operations)** — health, restart, model swaps, troubleshooting
6. **[Team](6-Team)** — team-lead orchestrating a mixed Claude + local-LLM worker fleet
7. **[Use Cases](7-Use-Cases)** — flagship code-review conversation (Opus ↔ gx00 worker) + the five mechanisms
8. **[Roadmap](8-Roadmap)** — walking-skeleton-first stages, tech stack, and tickets (Stage 1 detailed)
4. **[Setup](4-Setup)** — 🟠 stub, never written. Use **[13 User Guide](13-User-Guide) §2** instead.
5. **[Operations](5-Operations)** — 🟠 stub, never written. Use **[13 User Guide](13-User-Guide) §4 and §6** instead.
6. **[Team](6-Team)** — a lead orchestrating a mixed-vendor fleet
7. **[Use Cases](7-Use-Cases)** — the code-review scenario, the mechanisms, and the portable `CLAUDE.md` block
8. **[Roadmap](8-Roadmap)** — stages, tech stack, and tickets
9. **[Implementation](9-Implementation)** — as-built code map, classes, flows, state machines
10. **[Cross-Host Messaging](10-Cross-Host-Messaging)** — broker topology, exchanges, queues per entity
11. **[Features](11-Features)** — what it can do, the knob that turns it on, why it exists, the gotcha
12. **[Claude → OpenCode](12-Claude-to-OpenCode)** — porting a workspace to a second host
13. **[User Guide](13-User-Guide)** — 🟢 **the operator page.** Install, configure, run, delegate, and the traps.
## Status
🟢 Design — **herdr-centric `bridged` message server** selected as the primary approach
(2026-07-11), superseding the AgentAPI plan (2026-07-08). AgentAPI is retained as a fallback
injector. See **[Message Server](2-Message-Server)**.
🟢 **Running.** Release 1.1 is code-complete (2026-08-16): 20 of 20 tickets closed, 870 tests green,
the daemon live on this host. Chapters 4 and 5 were never written past their scope note; chapter 13
replaced them.
The **herdr-centric `bridged` message server** was selected on 2026-07-11, superseding the AgentAPI
plan of 2026-07-08. AgentAPI is retained as a fallback injector and has not been needed. See
**[Message Server](2-Message-Server)** for the design and **[13 User Guide](13-User-Guide)** for how
to run it.
+3 -2
@@ -7,8 +7,8 @@
1. [Architecture](1-Architecture) — system · 2 invariants · 2 modes
2. [Message Server](2-Message-Server) — the `bridged` design
3. [Approaches](3-Approaches) — transports compared, why herdr
4. [Setup](4-Setup) — bring-up
5. [Operations](5-Operations) — day-2 runbook
4. [Setup](4-Setup) — ⚫ superseded by 13
5. [Operations](5-Operations) — ⚫ superseded by 13
6. [Team](6-Team) — orchestrating a mixed fleet
7. [Use Cases](7-Use-Cases) — the review scenario + mechanisms
8. [Roadmap](8-Roadmap) — stages, tech stack, tickets
@@ -16,6 +16,7 @@
10. [Cross-Host Messaging](10-Cross-Host-Messaging) — broker topology · exchanges · queues per entity
11. [Features](11-Features) — what it can do · the knob that turns it on · why · the gotcha
12. [Claude → OpenCode](12-Claude-to-OpenCode) — porting a workspace to a second host
13. **[User Guide](13-User-Guide)** — 🟢 install · configure · run · delegate · the traps
---
🟢 herdr-centric `bridged` · AgentAPI = fallback