docs: add chapter 13, the operator user guide, for release 1.1
The wiki had twelve chapters and none of them told an operator how to run
the thing. Chapters 1-3 explain why the design is what it is, 9-12 explain
how the code is put together, 11 lists capabilities. The two pages that were
meant to cover bring-up and day-2 - 4-Setup and 5-Operations - were never
written past their scope note, and every technical detail in them had gone
wrong: a decommissioned model host, port 8080, herdr protocol 14, a systemd
unit that does not exist, Redis Streams and NATS that were never built,
"no per-session authz yet" after Authz shipped, and recycle() events that
have no code behind them.
So this adds 13-User-Guide.md, written against the running system on
2026-08-17, and points the two stubs at it rather than leaving wrong claims
in place.
The guide covers:
1. what this is and, more usefully, the five things it is NOT, each with
the reason it is not that;
2. install - herdr (check the PROTOCOL number, not the version), the
daemon, the login-shell rule for secrets.sh, and the lead's tab label;
3. configure - the four knobs that cost money, the live profile table with
who pays for each, the gateway paths, and memberCredentials' two halves;
4. run - the redeploy script, and the four checks that go beyond /healthz,
because health is green while every spawn fails;
5. delegate - the eleven tools, the spawn-all-then-send-all rule, the ~60s
client cap on a blocking send, and the authz table;
6. when it breaks - twelve traps hit for real this year, grouped by
bring-up, losing a member's work, and merging a member's work;
7. where to look next.
Home.md is corrected too: it claimed members launch against ollama.ltms.dev,
a host that no longer exists (the gateway is llm.ltms.dev), it framed the
system as Claude-only with one worker, it listed 8 of the 13 pages, and its
status still said "Design".
+444
@@ -0,0 +1,444 @@
|
||||
# 13 — User Guide
|
||||
|
||||
> **Status: 🟢 Written against the running system, 2026-08-17 (release 1.1).**
|
||||
> Every command here was run on this host. Where a claim is not measured, it says so.
|
||||
>
|
||||
> This is the operator page. Chapters 1–3 say *why the system is built this way*, 9–12 say *how the
|
||||
> code is put together*, 11 lists *what it can do*. This page is the one you read when you have to
|
||||
> bring it up, run it, or fix it. It assumes you have done this before and forgotten the details.
|
||||
|
||||
---
|
||||
|
||||
## 1. What this is, and what it is not
|
||||
|
||||
`bridged` is a **message bus between AI agent sessions**. One session orchestrates (the **lead**),
|
||||
and it delegates work to **members** running in other terminal panes, often on other models and
|
||||
other vendors. All traffic goes through `bridged`'s MCP tools. No session talks to another session,
|
||||
to a broker, or to the network directly.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
LEAD["lead session<br/>(Claude Code, subscription)"]
|
||||
subgraph BD["bridged — a plain Java daemon"]
|
||||
SRV["MCP + REST<br/>policy, authz, lifecycle"]
|
||||
INJ["injector<br/>status-gated"]
|
||||
SRV --> INJ
|
||||
end
|
||||
HERDR["herdr<br/>owns panes and PTYs"]
|
||||
M1["member pane<br/>claude-code"]
|
||||
M2["member pane<br/>opencode"]
|
||||
GW["llm.ltms.dev<br/>the one gateway"]
|
||||
|
||||
LEAD -->|"bridge_send"| SRV
|
||||
M1 -.->|"bridge_reply"| SRV
|
||||
M2 -.->|"bridge_reply"| SRV
|
||||
INJ -->|"unix socket"| HERDR
|
||||
HERDR --> M1
|
||||
HERDR --> M2
|
||||
M1 --> GW
|
||||
M2 --> GW
|
||||
|
||||
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
|
||||
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
|
||||
class LEAD,GW ext
|
||||
class SRV,INJ,HERDR core
|
||||
```
|
||||
|
||||
*Figure 1 — who talks to whom. The lead never touches herdr; the members never touch each other.*
|
||||
|
||||
**What it is not**, and each "not" is a decision, not a gap:
|
||||
|
||||
| Not | Why |
|
||||
|---|---|
|
||||
| Not a build system | It does not compile, lint, or test your code. Members run whatever the repo gives them. The lead verifies. |
|
||||
| Not an environment manager | It does not install toolchains or set up a member's language runtime. It only sets a small, named set of environment variables at spawn. |
|
||||
| Not a Claude-only tool | `kind: claude-code` and `kind: opencode` are both first-class. A member can be GPT, DeepSeek, or Claude. |
|
||||
| Not a scheduler with a budget | There is **no cost metering and no refusal on spend**. `maxLoad` per profile is the only throttle you have. See §3. |
|
||||
| Not a security boundary against your own members | A member runs as you, on your machine. `memberCredentials` limits which secrets it sees. It does not sandbox anything else. |
|
||||
|
||||
**The two invariants.** Break either and nothing else matters.
|
||||
|
||||
1. The lead **never** sets `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN`. It stays on the
|
||||
subscription. Only the daemon moves a member off it, at spawn.
|
||||
2. The bridge is the **only** channel. Text printed in a pane reaches nobody. An answer that is not
|
||||
in a `bridge_*` call is discarded silently.
|
||||
|
||||
---
|
||||
|
||||
## 2. Install
|
||||
|
||||
Four things must be true before anything works. In order, because each one depends on the last.
|
||||
|
||||
### 2.1 herdr
|
||||
|
||||
`bridged` does not own terminals. `herdr` does. `bridged` drives it over a Unix socket.
|
||||
|
||||
```bash
|
||||
herdr --version # 0.8.0 on this host
|
||||
curl -s http://127.0.0.1:8765/healthz
|
||||
# {"status":"ok","herdr":{"version":"0.8.0","protocol":19}}
|
||||
```
|
||||
|
||||
**The protocol number is the thing to check, not the version.** `bridged`'s adapter speaks one
|
||||
herdr wire protocol. If herdr is upgraded and the protocol moves, `/healthz` still says `ok` —
|
||||
because the socket connects — and **every spawn fails**. Green health with a broken fleet is the
|
||||
normal way this breaks. See trap 2 in §6.
|
||||
|
||||
### 2.2 The daemon
|
||||
|
||||
Build and run from this repo. There is no POM at the repo root:
|
||||
|
||||
```bash
|
||||
mvn -f bridged/pom.xml clean install # never pipe this — a pipe hides BUILD FAILURE
|
||||
java -jar bridged/target/bridged.jar bridged/bridged.yaml
|
||||
```
|
||||
|
||||
In practice you never run those two by hand. Use the script (§4).
|
||||
|
||||
### 2.3 Secrets, and the login shell
|
||||
|
||||
All credentials live in one file: `${SHARED_ENV}/tools/secrets.sh`. It is sourced only by a **login
|
||||
shell**. This one fact causes more lost hours than anything else in the system.
|
||||
|
||||
- The daemon inherits `AI_GATEWAY_TOKEN` and `WORKER_GITEA_TOKEN` **from the shell that started it**.
|
||||
- Start it from a non-login shell and both are empty. The daemon boots, `/healthz` is green, and
|
||||
nothing complains.
|
||||
- The failure surfaces hours later: a gateway profile gets a 401, or a member cannot open a pull
|
||||
request.
|
||||
|
||||
`launchd` does not run a login shell either. That is the only reason
|
||||
`scripts/bridged-launchd-wrapper.sh` exists — it `exec`s `zsh -lc` so the store gets sourced, while
|
||||
keeping one process so launchd's PID tracking still works. Read its header; it explains the trap
|
||||
better than this paragraph.
|
||||
|
||||
> On this host the daemon is **not** under launchd. It runs as a plain `java -jar` started by the
|
||||
> redeploy script from a login shell. `launchctl list | grep bridg` returns nothing.
|
||||
|
||||
The daemon logs which required secret names resolved at startup
|
||||
(`Bridged.reportRequiredSecrets`). Read those lines. But note the gap: it **skips profiles marked
|
||||
`subscription: true`**, on purpose, because they need no token. So a green secret report says
|
||||
nothing about your subscription profiles.
|
||||
|
||||
### 2.4 The lead's tab
|
||||
|
||||
The daemon finds the lead by its **herdr tab label**, and by nothing else:
|
||||
|
||||
```yaml
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus" # must match the real tab label
|
||||
```
|
||||
|
||||
Matching is case-insensitive and both ends are stripped. Internal spacing is **not** forgiving:
|
||||
`"lead: opus"` and `"lead:opus"` are different. A session whose tab does not match is resolved as a
|
||||
**worker**, and every orchestration call it makes is refused.
|
||||
|
||||
The old `terminal:` key is gone. A daemon at CB-579 or later **refuses to start** if `terminal:` is
|
||||
still in the config, and also if `tab:` is missing. That refusal is deliberate — a silently demoted
|
||||
lead was worse than a daemon that will not boot.
|
||||
|
||||
---
|
||||
|
||||
## 3. Configure
|
||||
|
||||
`bridged/bridged.yaml` is the live config. It is **gitignored**. `bridged.example.yaml` is the
|
||||
tracked, documented copy. Two consequences you will meet:
|
||||
|
||||
- Members cannot see the live config. They work in worktrees of the tracked repo. So a change to
|
||||
what a config **key means** can pass review, merge green, and break the running fleet — because
|
||||
nobody who reviewed it could see the file it breaks. Check and fix the live config yourself at
|
||||
merge time.
|
||||
- The example file is the documentation. A key that is read by code but missing from the example is
|
||||
a real defect (that is how `subscription:` stayed undocumented for months — CB-610).
|
||||
|
||||
### The knobs that cost money
|
||||
|
||||
This is the section to reread before you change anything.
|
||||
|
||||
| Knob | What it does | The cost |
|
||||
|---|---|---|
|
||||
| `subscription: true` | The member runs on **your own Claude plan**. No `ANTHROPIC_BASE_URL`, no token — it inherits your Claude Code auth. | Every spawn bills your plan and eats your usage limit. Off-subscription is the whole point of the daemon, so treat `true` as a deliberate exception. Mutually exclusive with `baseUrl`; setting both is refused at load. |
|
||||
| `maxLoad` | Members allowed at once on this profile. | **The only throttle that exists.** There is no metering, no budget, no refusal on spend. When `maxLoad` is reached, a spawn is refused with `no capacity … no fallback` — it does not fall back to a cheaper profile. |
|
||||
| `weight` | Share of automatic placement. | `weighted` placement is **not cheapest-first**. It spreads by ratio across every profile that has a free slot, so paid spawns happen while the free local profile is idle. The workaround here is `local.weight: 100`; the real fix is open as CB-589. |
|
||||
| `fleet.reviewers` | Which profiles an unqualified reviewer spawn may land on. | Review is the fan-out step — several members per pull request. Leaving the free profile out of this pool was the single largest avoidable cost in the fleet. |
|
||||
|
||||
Live profiles on this host:
|
||||
|
||||
| Profile | Kind | Model | maxLoad | Pays |
|
||||
|---|---|---|---|---|
|
||||
| `local` | claude-code | deepseek-v4-flash via `llm.ltms.dev/anthropic` | 2 | gateway |
|
||||
| `local-direct` | claude-code | deepseek-v4-flash direct to `gx00.gw:8000` | 2 | nothing, LAN only, `weight: 0` |
|
||||
| `gx` | opencode | `gx/deepseek-v4-flash` via `llm.ltms.dev/v1` | 2 | gateway |
|
||||
| `opus` | claude-code | `claude-opus-5` | 1 | **your Claude plan** |
|
||||
| `sonnet` | claude-code | `claude-sonnet-5` | 3 | **your Claude plan** |
|
||||
| `sol` | opencode | `openai/gpt-5.6-sol` | 1 | OpenAI |
|
||||
| `terra` | opencode | `openai/gpt-5.6-terra` | 2 | OpenAI |
|
||||
|
||||
`sol` and `terra` share **one** OpenAI credential. When it hits its limit, both die at once, and an
|
||||
opencode member dying looks like a short clipped reply rather than an error.
|
||||
|
||||
### The gateway
|
||||
|
||||
`llm.ltms.dev` is the one front door, and `AI_GATEWAY_TOKEN` is the single key.
|
||||
|
||||
- `kind: claude-code` → `https://llm.ltms.dev/anthropic`. **Never `/v1`** — on `/v1` the model's
|
||||
thinking output silently disappears.
|
||||
- `kind: opencode` → `https://llm.ltms.dev/v1`. The path is taken as-is.
|
||||
|
||||
Both known gateway ceilings were fixed on 2026-08-15: the body limit went 32 KiB → 32 MiB, and the
|
||||
route timeout 60 s → 86400 s. The remaining risk is an Envoy HTTP/1.1 bug that can truncate a
|
||||
stream with a `200` and no terminator. Our members cannot detect that; it looks like a short answer.
|
||||
|
||||
### Credentials a member can see
|
||||
|
||||
`memberCredentials` is deny-by-default. You name what is allowed; everything else known is replaced
|
||||
with a sentinel string. Current state: **34 known, 7 allowed, 29 blocked**, measured inside a live
|
||||
member pane on 2026-08-17.
|
||||
|
||||
It has **two halves**, and only both together work:
|
||||
|
||||
1. The launcher writes the member's environment at spawn.
|
||||
2. A `BRIDGED_MEMBER`-guarded block at the **end** of `${SHARED_ENV}/tools/secrets.sh` re-applies it.
|
||||
|
||||
Half 2 is not optional. The member's pane runs a **login shell**, which re-sources the whole secret
|
||||
store and overwrites whatever the launcher set. The guarded block must be the last thing in the
|
||||
file, or the store overwrites it in turn.
|
||||
|
||||
The policy is re-read **on every spawn**, so an allow-list edit needs no restart. That is measured,
|
||||
not assumed.
|
||||
|
||||
> Rule in force: only the lead and architects may use `GITEA_ACCESS_TOKEN`. Members get
|
||||
> `WORKER_GITEA_TOKEN`, which is a minimal `write:repository` token, never the admin one.
|
||||
|
||||
---
|
||||
|
||||
## 4. Run, and prove it runs
|
||||
|
||||
### Start and restart
|
||||
|
||||
Use the script. Do not hand-roll the steps.
|
||||
|
||||
```bash
|
||||
scripts/redeploy-bridged.sh --check # read-only: reports state, changes nothing
|
||||
scripts/redeploy-bridged.sh # build, confirm drain, restart, verify
|
||||
scripts/redeploy-bridged.sh --yes # skip the drain prompt (fleet already checked)
|
||||
scripts/redeploy-bridged.sh --no-build # restart the jar you already have
|
||||
```
|
||||
|
||||
Run `--check` first, always. It is the only thing that reports whether the forge token resolves,
|
||||
and it reports that **without printing the value**.
|
||||
|
||||
The script builds **before** it stops anything, so a failed build never leaves the fleet down. It
|
||||
waits for the old process to exit instead of assuming. It polls `/healthz`. And it anchors its log
|
||||
checks to a marker taken before the restart, so old errors cannot be misread as new ones.
|
||||
|
||||
**A merge is not a deployment.** The running daemon holds the jar it started with. Code merged to
|
||||
`main` does nothing until you rebuild and restart. Saying "shipped" about code the live daemon has
|
||||
never loaded is a false report.
|
||||
|
||||
**Drain first.** `bridge_list`, collect anything you still want with `bridge_poll`, then
|
||||
`bridge_stop` each member. A restart drops in-flight tickets and rendezvous. A member's report is
|
||||
not recoverable once its ticket is gone.
|
||||
|
||||
### Verify — `/healthz` is not enough
|
||||
|
||||
`/healthz` proves the socket connects. It does not prove a spawn works, that identity resolves, or
|
||||
that the new jar is the one running. Four checks, in order:
|
||||
|
||||
1. **A fresh boot line.** Confirm a new `bridged listening` line at the end of `bridged/bridged.out`,
|
||||
dated after the restart. An old daemon that never died looks identical from outside.
|
||||
2. **Deferred keys.** The startup log names which config keys it accepted and which it deferred. A
|
||||
deferred key needing a restart is usually the whole reason you restarted. Read those lines rather
|
||||
than assuming.
|
||||
3. **Identity.** `bridge_whoami` must still answer `primary`. If the tab label changed, the lead is
|
||||
now a worker and every orchestration call is refused.
|
||||
4. **A real spawn.** Spawn one cheap member and stop it. This is the only check that catches a herdr
|
||||
protocol mismatch.
|
||||
|
||||
> Restarting the daemon **cuts your own MCP mount**, and it does not reconnect. So you cannot run
|
||||
> `bridge_whoami` from the session that restarted it. Ask the operator to run `/mcp` to reconnect.
|
||||
> This is why the restart is done from the lead but verified after a reconnect.
|
||||
|
||||
---
|
||||
|
||||
## 5. Delegate
|
||||
|
||||
Eleven tools. This is the whole surface.
|
||||
|
||||
| Intent | Tool |
|
||||
|---|---|
|
||||
| Confirm your own role | `bridge_whoami` |
|
||||
| See backends available | `bridge_profiles` |
|
||||
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` |
|
||||
| See the fleet | `bridge_list` → `leads` + `members` |
|
||||
| One member's state | `bridge_status{sessionId}` |
|
||||
| Delegate, blocking | `bridge_send{sessionId, content}` |
|
||||
| Delegate, long task | `bridge_send{sessionId, content, wait:false}` → ticket |
|
||||
| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
|
||||
| Collect a reply | `bridge_poll{ticket}` or `bridge_poll{target}` |
|
||||
| Clear a reply from the inbox | `bridge_ack{target, msgId}` — **`target`, not `ticket`** |
|
||||
| Member ends its turn | `bridge_reply{content}` |
|
||||
| Member asks the lead | `bridge_ask{question}` |
|
||||
| Tear down | `bridge_stop{paneId}` |
|
||||
|
||||
### The loop
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant L as lead
|
||||
participant B as bridged
|
||||
participant M as member
|
||||
L->>B: bridge_spawn — all units first
|
||||
L->>B: bridge_send with wait false — then all sends
|
||||
B-->>L: ticket
|
||||
B->>M: injected when the member is idle
|
||||
M->>B: bridge_reply
|
||||
B-->>L: nudge into the lead's own pane
|
||||
L->>B: bridge_poll by ticket
|
||||
L->>B: bridge_ack by target and msgId
|
||||
L->>B: bridge_stop by paneId
|
||||
```
|
||||
|
||||
*Figure 2 — the delegation loop. Spawn is separate from send on purpose.*
|
||||
|
||||
**Spawn every unit first, then send them all.** Spawning and sending in one loop is how parallel
|
||||
work silently becomes serial. It is the most common way this whole layer is wasted.
|
||||
|
||||
**Prefer `wait:false`.** A blocking `bridge_send` is capped by **your own MCP client timeout**,
|
||||
about 60 seconds — far below any real task's runtime. The cap is in your client, not in the daemon,
|
||||
so no server setting fixes it.
|
||||
|
||||
**Line 1 of every brief names a skill.** `Load the implementer skill.` or `Load the reviewer skill.`
|
||||
Those skills are opt-in, and that line is what makes them reliable. The brief must be
|
||||
self-contained: the member sees your message and the repo, and nothing of your context, your plan,
|
||||
or your screen.
|
||||
|
||||
**Since CB-588, a `wait:false` ticket that finishes nudges your pane by itself.** You do not have to
|
||||
poll on a timer. It needs an injectable pane and is capped at 5 pending nudges, so still watch
|
||||
anything you cannot afford to lose.
|
||||
|
||||
### Who may call what
|
||||
|
||||
| Call | Allowed |
|
||||
|---|---|
|
||||
| spawn / stop / drain | lead only |
|
||||
| send | lead or architect |
|
||||
| reply / ask | any member, **only as itself** |
|
||||
| read / status | any authenticated caller |
|
||||
|
||||
Identity comes from the connection, never from an argument. You cannot act as another session. A
|
||||
call outside your role is refused, not queued.
|
||||
|
||||
### The lead's own job
|
||||
|
||||
Delegation does not delegate responsibility. These stay with you and are not delegable: the
|
||||
conversation with the operator, decomposition, the final judgment, **verification**, and the
|
||||
**merge**. Merging on a reviewer's word is delegating the merge by proxy.
|
||||
|
||||
A member cannot run your IDE tooling. Any forge tools it appears to have hold a deliberately blocked
|
||||
credential and fail every call. And a piped command hides failures behind a zero exit code. So never
|
||||
promote a member's "clean" to a fact — re-run the build yourself.
|
||||
|
||||
---
|
||||
|
||||
## 6. When it breaks
|
||||
|
||||
Twelve traps, all hit for real this year. Grouped by where they bite.
|
||||
|
||||
### Bring-up
|
||||
|
||||
**1. The daemon started from a non-login shell.**
|
||||
Symptom: everything green for hours, then a member cannot open a pull request, or a gateway profile
|
||||
gets a 401. Nothing logs it at the time. Fix: `scripts/redeploy-bridged.sh --check` is the only
|
||||
thing that reports it. Restart from a login shell, or via the launchd wrapper.
|
||||
|
||||
**2. `/healthz` is green and every spawn fails.**
|
||||
Cause: herdr was upgraded and its wire protocol moved past what the adapter speaks. The socket still
|
||||
connects, so health is `ok`. Fix: check `protocol` in the `/healthz` body, and verify any restart
|
||||
with one real spawn — not with health.
|
||||
|
||||
**3. The lead is resolved as a worker.**
|
||||
Cause: the herdr tab label no longer matches `fleet.leaders.*.tab`. Every orchestration call is
|
||||
refused. Fix: match the tab exactly — internal spacing counts. Note the old `terminal:` key is
|
||||
removed; a current daemon refuses to start if it is still present.
|
||||
|
||||
**4. A merge is not a deployment.**
|
||||
The running daemon holds its original jar. Rebuild and restart, then prove it with a **fresh**
|
||||
`bridged listening` line. Also: the restart cuts your own MCP mount and it never reconnects, so you
|
||||
cannot verify identity from that session — ask for `/mcp`.
|
||||
|
||||
### Losing a member's work
|
||||
|
||||
**5. The ticket expired.**
|
||||
Ticket time-to-live is about 10 minutes. After that `bridge_poll{ticket}` returns
|
||||
`timed_out_working`. The member is usually fine and its real answer arrives later — in the member's
|
||||
**inbox**. Drain it with `bridge_poll{target}`, then `bridge_ack{target, msgId}`.
|
||||
|
||||
**6. Reading the reply inbox is destructive.**
|
||||
`GET /sessions/{id}/replies` drains on first read. If you `curl` it through `head` or a parser that
|
||||
dies, the reply is gone. Always redirect it to a **file** first.
|
||||
|
||||
**7. The idle reaper ate the report.**
|
||||
A member that is `done` holds its report only until `lifecycle.idleTtlSeconds` (1800 s here). Then
|
||||
the ticket, the pane, and the report are all gone, with no scrape fallback. Poll before the reaper.
|
||||
|
||||
**8. A second send to a busy member never lands.**
|
||||
`timed_out_queued` means **never delivered** — the member is fine. Worse, the stale brief stays
|
||||
queued and restarts the member when it next goes idle. Back up the member's commit first:
|
||||
`git fetch <worktree> <branch>:refs/backup/x`.
|
||||
|
||||
**9. `liveStatus: working` is not progress.**
|
||||
A member can report busy for hours doing nothing. Check file modification times in its worktree and
|
||||
snapshot `git diff` before and after. Do not read its terminal — use `bridge_status`.
|
||||
|
||||
**10. Never brief a member to "ask me".**
|
||||
`bridge_ask` blocks for about 55 seconds and no nudge extends it. It is also invisible to
|
||||
`bridge_poll`. If you are running async, you will not see the question in time. Decide before you
|
||||
delegate, or give the member an explicit default.
|
||||
|
||||
### Merging a member's work
|
||||
|
||||
**11. The reported branch is not the branch it committed to.**
|
||||
`bridge_list` reports the branch **spawn provisioned**, not the one the member actually used. Check
|
||||
`git -C <worktree> branch --show-current`, or you push an empty ref and the merge says
|
||||
"Already up to date".
|
||||
|
||||
**12. Members see a months-old wiki, and cannot see the live config.**
|
||||
`wiki/` is a submodule whose pointer is never advanced, so a member's checkout is a stale snapshot.
|
||||
Never brief "read `wiki/…`" — paste the text, and commit the member's wiki entry yourself. Same
|
||||
shape for `bridged.yaml`: it is gitignored, so a member cannot see the file its change may break.
|
||||
|
||||
### The general shape behind several of these
|
||||
|
||||
A checker that is **narrower than what everyone believes it checks**, with nothing saying so. The
|
||||
credential probe hardcoded 31 names against a policy of 29 and exited `0`. A config test checked
|
||||
only column-0 keys, so a nested key that bills your Claude plan stayed undocumented. A scheduled
|
||||
sweep never ran once because its gate used a sentinel in a numeric field.
|
||||
|
||||
Two habits fix most of it: **make a checker print its own denominator** ("checked 26 of 29", not
|
||||
"26 blocked"), and when a green check is the evidence for a claim, **ask what it does not cover
|
||||
before believing it**.
|
||||
|
||||
---
|
||||
|
||||
## 7. Where to look next
|
||||
|
||||
| You want | Go to |
|
||||
|---|---|
|
||||
| Why the design is herdr-centric at all | [1 Architecture](1-Architecture), [3 Approaches](3-Approaches) |
|
||||
| The daemon's own design | [2 Message Server](2-Message-Server) |
|
||||
| The as-built code map, classes, state machines | [9 Implementation](9-Implementation) |
|
||||
| What a capability does, the knob, why it exists, the gotcha | [11 Features](11-Features) |
|
||||
| Orchestrating a mixed-vendor fleet | [6 Team](6-Team), [7 Use Cases](7-Use-Cases) |
|
||||
| Adding a second host or a non-Claude peer | [10 Cross-Host Messaging](10-Cross-Host-Messaging), [12 Claude → OpenCode](12-Claude-to-OpenCode) |
|
||||
| What is left to build | [8 Roadmap](8-Roadmap) |
|
||||
| The rules every session must follow | `CLAUDE.md` in the repo — it loads into every session |
|
||||
| Rendezvous, `bridge_ask`, detached delivery, turn-done fallback | `docs/MCP-Contract.md` **§6 only** — the rest of that page is a pre-build design doc and its tool names never caught up with the code |
|
||||
| Every config key, documented | `bridged/bridged.example.yaml` |
|
||||
|
||||
**The one thing that is easy to forget.** This repo *is* the bridge, so the charter block in
|
||||
`CLAUDE.md` is not documentation about someone else's system — it is the instruction surface this
|
||||
codebase ships. A code change that quietly makes it untrue is an incomplete change. The block in
|
||||
`CLAUDE.md` and the template in [7 Use Cases](7-Use-Cases) must stay byte-identical; there is a
|
||||
check script in `CLAUDE.md` that proves it.
|
||||
+16
-31
@@ -1,40 +1,25 @@
|
||||
# 4. Setup
|
||||
|
||||
> **Status:** 🟠 Stub — scope defined, procedure not yet written. `bridged` is at the design
|
||||
> stage ([Message Server](2-Message-Server)); concrete install steps land with **M0–M1** of the build plan.
|
||||
> **Status: ⚫ Superseded by [13 User Guide](13-User-Guide) §2 (2026-08-17).**
|
||||
> This page was written in the design era as a scope note and the procedure was never filled in.
|
||||
> Its content had gone wrong in every detail — a decommissioned model host, port 8080, herdr
|
||||
> protocol 14, a systemd unit that does not exist on this host, Redis Streams / NATS that were never
|
||||
> built. Rather than leave those claims in place, the page now points at the written procedure.
|
||||
|
||||
This page will cover standing up the bridge on an off-subscription worker host.
|
||||
**Go to [13 User Guide](13-User-Guide) → §2 Install.** It covers, against the running system:
|
||||
|
||||
## Scope (what this page will contain)
|
||||
- herdr, and why you check the **protocol number** rather than the version;
|
||||
- building and starting the daemon;
|
||||
- the **login shell** rule for `${SHARED_ENV}/tools/secrets.sh` — the single most expensive trap in
|
||||
bring-up, and why `scripts/bridged-launchd-wrapper.sh` exists;
|
||||
- the lead's **tab label**, which is how the daemon resolves who the lead is.
|
||||
|
||||
1. **Prerequisites** — [herdr](https://herdr.dev) installed and its server running; a worker
|
||||
`claude` that inherits your `CLAUDE.md`/hooks/skills/MCP; a reachable worker model
|
||||
(`ollama.ltms.dev` or GX10 vLLM) with a bearer token.
|
||||
2. **herdr** — start the headless server; confirm the socket at
|
||||
`~/.config/herdr/herdr.sock` (or `HERDR_SOCKET_PATH`); verify with `ping` (expect
|
||||
`protocol: 14` on herdr 0.7.0) and `workspace.list`.
|
||||
3. **`bridged`** — deploy the binary, `bridged.yaml` (worker model, `base_url` allowlist,
|
||||
auth token, bind address), and a **systemd** unit ordered *after* herdr.
|
||||
4. **Worker session** — create the first worker pane with the env-prefixed launch line
|
||||
(`ANTHROPIC_BASE_URL=https://ollama.ltms.dev ANTHROPIC_AUTH_TOKEN=… claude`); confirm the
|
||||
subscription guard accepts it and `pane.process_info` shows the expected egress host.
|
||||
5. **MCP mount (unified, both sides)** — register `bridged` as an MCP server on the primary
|
||||
**and** each worker with one line —
|
||||
`claude mcp add --transport http bridge http://127.0.0.1:8080/mcp` (or a shared
|
||||
`.mcp.json` / `CLAUDE.md` entry every session inherits). The primary then delegates via the
|
||||
`bridge_send` tool (single blocking call per delegation) and workers reply via
|
||||
`bridge_reply`. Same-host needs nothing more — `bridged` delivers async by injecting an idle
|
||||
pane. Only for a **split-host** primary (not a herdr pane) add a `Stop`-hook that long-polls
|
||||
**`bridged`** (never a broker) for detached wake-ups.
|
||||
6. **Topology choice** — single-host vs split-host (see [Message Server](2-Message-Server) → *Deployment
|
||||
model*). If durability or cross-host async is needed, configure `bridged`'s **internal** queue
|
||||
(Redis Streams / NATS); it stays behind the gateway — no Claude session connects to it.
|
||||
## The one rule that was already right on this page
|
||||
|
||||
## Non-negotiable during setup
|
||||
|
||||
The **primary** host/process must **never** be given `ANTHROPIC_BASE_URL`. Only worker panes
|
||||
carry it. See [Architecture](1-Architecture) → *Subscription boundary*.
|
||||
The **lead** host and process must **never** be given `ANTHROPIC_BASE_URL` or
|
||||
`ANTHROPIC_AUTH_TOKEN`. Only member panes carry them, and only the daemon sets them, at spawn. See
|
||||
[Architecture](1-Architecture) → *Subscription boundary*.
|
||||
|
||||
## Related
|
||||
|
||||
- [Message Server](2-Message-Server) · [Architecture](1-Architecture) · [Operations](5-Operations) · [Approaches](3-Approaches)
|
||||
- [13 User Guide](13-User-Guide) · [2 Message Server](2-Message-Server) · [1 Architecture](1-Architecture) · [11 Features](11-Features)
|
||||
|
||||
+25
-29
@@ -1,39 +1,35 @@
|
||||
# 5. Operations
|
||||
|
||||
> **Status:** 🟠 Stub — scope defined, runbook not yet written. Fills in as `bridged` reaches
|
||||
> **M4 — Harden** (auth/TLS, metrics, systemd) in the [Message Server](2-Message-Server) build plan.
|
||||
> **Status: ⚫ Superseded by [13 User Guide](13-User-Guide) §4 and §6 (2026-08-17).**
|
||||
> This page was a design-era scope note; the runbook was never written. Several of its claims were
|
||||
> overtaken by the build — `recycle()` events do not exist, per-session authorization shipped
|
||||
> (`auth/Authz`), and the internal broker is optional rather than required.
|
||||
|
||||
Day-2 runbook for a running bridge.
|
||||
**Go to [13 User Guide](13-User-Guide):**
|
||||
|
||||
## Scope (what this page will contain)
|
||||
- **§4 Run, and prove it runs** — `scripts/redeploy-bridged.sh`, why a merge is not a deployment,
|
||||
draining before a restart, and the four checks that go beyond `/healthz`.
|
||||
- **§6 When it breaks** — twelve traps hit for real this year, grouped by bring-up, losing a
|
||||
member's work, and merging a member's work.
|
||||
|
||||
- **Health** — `GET /healthz` liveness, `GET /metrics` (Prometheus), reading live
|
||||
`agent_status` per session via `GET /sessions`, and confirming the **MCP endpoint** is
|
||||
reachable from both the primary and the workers (`claude mcp list` shows `bridge` connected).
|
||||
- **Restart & recovery** — ordered restart (herdr before `bridged`); how `bridged`
|
||||
re-attaches to existing panes via `workspace.list`/`pane.list`; internal broker/queue replay of unacked items. See
|
||||
[Architecture](1-Architecture) → *Failure modes & single points of failure* for what each outage costs.
|
||||
- **Model swaps** — repoint a worker to a different `base_url`/model by recycling its pane
|
||||
(Ralph loop); the subscription guard re-validates the new host against the allowlist.
|
||||
- **Lifecycle / context ceilings** — observing recycle events; confirming workers externalize
|
||||
state (git + `STATE.md`) before a recycle so continuity survives (see [Message Server](2-Message-Server) →
|
||||
*Worker session lifecycle*).
|
||||
- **Troubleshooting** — stuck `working` (model endpoint down), `blocked` awaiting input,
|
||||
injection collisions on a hand-driven pane, envelope-vs-scrape reply mismatches.
|
||||
- **Security ops** — token rotation, keeping the port off public interfaces; note that one
|
||||
`bridged` is currently **one trust domain** (no per-session authz yet — [Message Server](2-Message-Server)
|
||||
→ *Security*).
|
||||
## What actually shipped, against what this page predicted
|
||||
|
||||
## Guardrails to watch (from [Architecture](1-Architecture))
|
||||
| This page said | What is true now |
|
||||
|---|---|
|
||||
| "no per-session authz yet" | Per-session authorization shipped. `auth/Authz` holds the role table; identity comes from the connection, never from an argument. |
|
||||
| "observing recycle events" | There is no `recycle()`. Lifecycle is an idle reaper plus a context cap — `lifecycle.idleTtlSeconds`, `lifecycle.contextCap`, `lifecycle.drainTimeoutSeconds`. |
|
||||
| "internal broker/queue replay" | The broker is optional. With `broker:` commented out the daemon uses an in-memory inbox. When it is on, it is **LavinMQ**, never RabbitMQ. |
|
||||
| "`GET /healthz` liveness" | Still true, and still not sufficient. Health can be green while every spawn fails, because the socket connects but the herdr wire protocol has moved. Check `protocol` in the body and verify with a real spawn. |
|
||||
| "recycling a pane to swap models" | Spawn a member on a different profile instead. Profiles are the unit of model choice. |
|
||||
|
||||
- The **primary must never perpetual-poll** — quota burn. Async wake-ups are `bridged`
|
||||
inject-on-idle by default; the `Stop`-hook is only the split-host-primary exception, and it
|
||||
polls `bridged` (never a broker).
|
||||
- Cross-agent ping-pong needs a round/turn budget — enforced centrally in `bridged` (sole
|
||||
gateway), not per-session sentinels.
|
||||
- The **internal** broker/queue must run with **ack + visibility timeout + consumer groups**
|
||||
so a mid-turn crash re-delivers instead of dropping.
|
||||
## Guardrails that still hold
|
||||
|
||||
- The lead must **never** busy-poll a member's pane — it burns your subscription for nothing. Async
|
||||
wake-ups are injected into the lead's own pane when a ticket goes terminal (CB-588).
|
||||
- Keep the port off public interfaces. The daemon binds loopback and trusts it; the exposure check
|
||||
only fires on a non-loopback bind.
|
||||
- Rotate tokens. Members get a minimal `write:repository` forge token, never the admin one.
|
||||
|
||||
## Related
|
||||
|
||||
- [Message Server](2-Message-Server) · [Architecture](1-Architecture) · [Setup](4-Setup) · [Approaches](3-Approaches)
|
||||
- [13 User Guide](13-User-Guide) · [2 Message Server](2-Message-Server) · [1 Architecture](1-Architecture) · [11 Features](11-Features)
|
||||
|
||||
+32
-18
@@ -1,12 +1,16 @@
|
||||
# claude-bridge
|
||||
|
||||
A **subscription-safe bridge** that lets a primary **Claude Code (Opus 4.8, on Pro/Max)**
|
||||
session drive a **secondary Claude agent running a different model** via its own
|
||||
`ANTHROPIC_BASE_URL` — without ever putting a proxy on the primary.
|
||||
A **subscription-safe bridge** that lets a lead **Claude Code** session on Pro/Max delegate work to
|
||||
**members running on other models and other vendors** — without ever putting a proxy on the lead.
|
||||
|
||||
> **New here, or here to operate it? Start at [13 User Guide](13-User-Guide).** That page is written
|
||||
> against the running system: install, configure, run, delegate, and the traps. This page explains
|
||||
> the shape of the design and why it was chosen.
|
||||
|
||||
> Sibling of [`crush-bridge`](https://git.ltms.dev/systems/vms) (Tier 2 → headless Crush
|
||||
> on GX10 DeepSeek). `claude-bridge` keeps the worker a *real Claude Code process* so it
|
||||
> inherits `CLAUDE.md`, hooks, skills, and MCP — just pointed at a cheaper/local model.
|
||||
> on GX10 DeepSeek). `claude-bridge` keeps a Claude member a *real Claude Code process* so it
|
||||
> inherits `CLAUDE.md`, hooks, skills, and MCP — just pointed at a cheaper model. It also runs
|
||||
> non-Claude members (`kind: opencode`) as first-class peers.
|
||||
|
||||
## Leading approach — herdr-centric message server (`bridged`)
|
||||
|
||||
@@ -17,9 +21,9 @@ workers mount** — one unified Claude setup and the **sole communication gatewa
|
||||
stays for non-Claude clients; any broker is `bridged`-internal, below the gateway). herdr owns
|
||||
the PTYs, multiplexing, persistence, and **agent-status
|
||||
events**; `bridged` owns policy (subscription boundary, session lifecycle, status-gated
|
||||
delivery) and the client contract. The worker `claude` launches with
|
||||
`ANTHROPIC_BASE_URL=https://ollama.ltms.dev` + a bearer token; the primary Opus stays
|
||||
env-clean and calls `bridged`'s MCP tools.
|
||||
delivery) and the client contract. A Claude member launches with `ANTHROPIC_BASE_URL` pointed at
|
||||
the gateway, `https://llm.ltms.dev/anthropic`, plus a bearer token; the lead stays env-clean and
|
||||
calls `bridged`'s MCP tools.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
@@ -30,8 +34,8 @@ flowchart LR
|
||||
SRV --> CLI
|
||||
end
|
||||
HERDR["herdr<br/>panes · agent-status"]
|
||||
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
|
||||
M["ollama.ltms.dev<br/>(worker model)"]
|
||||
W["member pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
|
||||
M["llm.ltms.dev<br/>(the one gateway)"]
|
||||
|
||||
OPUS -->|"MCP bridge_send (blocks)"| SRV
|
||||
W -.->|"MCP bridge_reply"| SRV
|
||||
@@ -72,14 +76,24 @@ Read in order (the sidebar mirrors this):
|
||||
1. **[Architecture](1-Architecture)** — process model, the two invariants, two traffic modes
|
||||
2. **[Message Server](2-Message-Server)** — 🟢 **`bridged`**, the herdr-centric message server (primary approach)
|
||||
3. **[Approaches](3-Approaches)** — herdr-centric vs AgentAPI vs Agent SDK vs bus/tmux (research matrix)
|
||||
4. **[Setup](4-Setup)** — running herdr + `bridged` + a worker pointed at `ollama.ltms.dev`
|
||||
5. **[Operations](5-Operations)** — health, restart, model swaps, troubleshooting
|
||||
6. **[Team](6-Team)** — team-lead orchestrating a mixed Claude + local-LLM worker fleet
|
||||
7. **[Use Cases](7-Use-Cases)** — flagship code-review conversation (Opus ↔ gx00 worker) + the five mechanisms
|
||||
8. **[Roadmap](8-Roadmap)** — walking-skeleton-first stages, tech stack, and tickets (Stage 1 detailed)
|
||||
4. **[Setup](4-Setup)** — 🟠 stub, never written. Use **[13 User Guide](13-User-Guide) §2** instead.
|
||||
5. **[Operations](5-Operations)** — 🟠 stub, never written. Use **[13 User Guide](13-User-Guide) §4 and §6** instead.
|
||||
6. **[Team](6-Team)** — a lead orchestrating a mixed-vendor fleet
|
||||
7. **[Use Cases](7-Use-Cases)** — the code-review scenario, the mechanisms, and the portable `CLAUDE.md` block
|
||||
8. **[Roadmap](8-Roadmap)** — stages, tech stack, and tickets
|
||||
9. **[Implementation](9-Implementation)** — as-built code map, classes, flows, state machines
|
||||
10. **[Cross-Host Messaging](10-Cross-Host-Messaging)** — broker topology, exchanges, queues per entity
|
||||
11. **[Features](11-Features)** — what it can do, the knob that turns it on, why it exists, the gotcha
|
||||
12. **[Claude → OpenCode](12-Claude-to-OpenCode)** — porting a workspace to a second host
|
||||
13. **[User Guide](13-User-Guide)** — 🟢 **the operator page.** Install, configure, run, delegate, and the traps.
|
||||
|
||||
## Status
|
||||
|
||||
🟢 Design — **herdr-centric `bridged` message server** selected as the primary approach
|
||||
(2026-07-11), superseding the AgentAPI plan (2026-07-08). AgentAPI is retained as a fallback
|
||||
injector. See **[Message Server](2-Message-Server)**.
|
||||
🟢 **Running.** Release 1.1 is code-complete (2026-08-16): 20 of 20 tickets closed, 870 tests green,
|
||||
the daemon live on this host. Chapters 4 and 5 were never written past their scope note; chapter 13
|
||||
replaced them.
|
||||
|
||||
The **herdr-centric `bridged` message server** was selected on 2026-07-11, superseding the AgentAPI
|
||||
plan of 2026-07-08. AgentAPI is retained as a fallback injector and has not been needed. See
|
||||
**[Message Server](2-Message-Server)** for the design and **[13 User Guide](13-User-Guide)** for how
|
||||
to run it.
|
||||
|
||||
+3
-2
@@ -7,8 +7,8 @@
|
||||
1. [Architecture](1-Architecture) — system · 2 invariants · 2 modes
|
||||
2. [Message Server](2-Message-Server) — the `bridged` design
|
||||
3. [Approaches](3-Approaches) — transports compared, why herdr
|
||||
4. [Setup](4-Setup) — bring-up
|
||||
5. [Operations](5-Operations) — day-2 runbook
|
||||
4. [Setup](4-Setup) — ⚫ superseded by 13
|
||||
5. [Operations](5-Operations) — ⚫ superseded by 13
|
||||
6. [Team](6-Team) — orchestrating a mixed fleet
|
||||
7. [Use Cases](7-Use-Cases) — the review scenario + mechanisms
|
||||
8. [Roadmap](8-Roadmap) — stages, tech stack, tickets
|
||||
@@ -16,6 +16,7 @@
|
||||
10. [Cross-Host Messaging](10-Cross-Host-Messaging) — broker topology · exchanges · queues per entity
|
||||
11. [Features](11-Features) — what it can do · the knob that turns it on · why · the gotcha
|
||||
12. [Claude → OpenCode](12-Claude-to-OpenCode) — porting a workspace to a second host
|
||||
13. **[User Guide](13-User-Guide)** — 🟢 install · configure · run · delegate · the traps
|
||||
|
||||
---
|
||||
🟢 herdr-centric `bridged` · AgentAPI = fallback
|
||||
|
||||
Reference in New Issue
Block a user