docs: correct Home.md from the two page audits

Two members audited Home.md + 4-Setup.md and 5-Operations.md against the
code. Three findings changed what Home.md says.

1. The AgentAPI fallback does not exist. The page said it was "retained as a
   swappable fallback injector"; `grep -ri agentapi bridged/src/main` returns
   nothing, and the only injection path in shipped code is the herdr one. It
   is a discarded option, not something you can switch to, and reading it as
   a fallback would send an operator looking for a lever that was never
   built.

2. "One `claude mcp add` line for both sides" is claude-code framing only.
   An opencode member mounts through a generated opencode.json and gets no
   ANTHROPIC_* variables at all (OpenCodeLauncher). Also added what the
   broker actually is - optional, and LavinMQ over AMQP when on, not the
   Redis Streams / NATS the design era assumed.

3. The subscription boundary now has a deliberate exception. A profile with
   `subscription: true` runs its members on the operator's own plan on
   purpose, and an opencode member sits outside the boundary entirely on its
   own provider credential. The old absolute wording hid the one setting in
   the system that spends money.

Also corrected the reply paragraph: the blocking bridge_send is capped by the
lead's own MCP client timeout at about 60 seconds, so real work uses
wait:false and a ticket. The page described only the blocking path.
Dai Ha
2026-08-17 16:19:05 +02:00
parent f47978cc5d
commit 92a6c6da3a
2 changed files with 54 additions and 18 deletions
+29
@@ -260,6 +260,35 @@ that the new jar is the one running. Four checks, in order:
> `bridge_whoami` from the session that restarted it. Ask the operator to run `/mcp` to reconnect.
> This is why the restart is done from the lead but verified after a reconnect.
### The REST surface
When the MCP mount is down — which is exactly when you need it most — this is the whole API. It is
loopback only.
| Route | Does |
|---|---|
| `GET /healthz` | `200` with `herdr.protocol` in the body; `503 degraded` when herdr does not answer ping |
| `GET /metrics` | Prometheus. Only present when metrics are enabled |
| `GET /sessions` | every session, each row carrying its live `agentStatus` |
| `GET /agents` | agents as herdr sees them |
| `GET /members` · `POST /members` · `DELETE /members/{paneId}` | list, spawn, tear down |
| `GET /profiles` | the backends configured |
| `GET /sessions/{id}/status` | one session — the same view as `bridge_status` |
| `POST /sessions/{id}/message` · `/reply` · `/ask` | the three message kinds |
| `GET /sessions/{id}/replies` | drain the reply inbox — **destructive, see trap 6** |
| `GET /tasks/{ticket}` | poll a detached ticket |
### Where it runs, and where the logs are
- Log file: **`bridged/bridged.out`**, in both supervised and unsupervised modes.
- Audit log: `bridged/logs/audit.log`, rotated daily, 30 days kept.
- Service units ship in `deploy/`: `bridged.service` for Linux systemd (ordered
`After=herdr.service`) and `dev.ltms.bridged.plist` for macOS launchd. `deploy/lavinmq` holds the
optional broker.
- **On this host neither is loaded.** The daemon runs as a plain `java -jar` started by
`scripts/redeploy-bridged.sh` from a login shell. Verified with `launchctl list | grep bridg`,
which returns nothing. If you expected launchd here, that expectation is the bug.
---
## 5. Delegate
+25 -18
@@ -49,25 +49,32 @@ flowchart LR
class SRV,CLI,HERDR core
```
- **Subscription boundary:** the *primary* never sets `ANTHROPIC_BASE_URL` (stays on
Pro/Max). Only the *secondary* process is off-subscription; `bridged` is a plain daemon
(no Anthropic quota), so it may poll/subscribe freely.
- **One gateway:** `bridged` is the **sole communication path** for every Claude session. Both
the primary and the workers mount it as an MCP server (one `claude mcp add` line) and talk
*only* to it — **no Claude session ever addresses a broker, a peer, or the network directly.**
Any broker/queue is `bridged`-internal, below the gateway.
- **How the primary gets a reply:** it makes **one blocking MCP tool call** (`bridge_send`)
that `bridged` holds open until the worker replies (`bridge_reply`) or its turn completes,
then returns the reply as the tool result. Worker → primary rides `bridged`'s state, so no
keystroke into the primary pane is needed even single-host. Long/detached work comes back the
same way — `bridged` **injects the primary's idle pane** when the result is ready (a
split-host primary's `Stop`-hook polls `bridged`, not a broker). The primary never busy-polls
and never touches a broker.
- **Different model** per worker process sidesteps Claude Code's lack of per-subagent
- **Subscription boundary:** the *lead* never sets `ANTHROPIC_BASE_URL` (stays on Pro/Max). Only a
member process is moved off-subscription, and only by the daemon, at spawn; `bridged` is a plain
daemon with no Anthropic quota, so it may poll and subscribe freely. **One exception, and it
costs money:** a profile marked `subscription: true` runs its members on *your* plan on purpose —
see [13 User Guide](13-User-Guide) → *The knobs that cost money*. An `opencode` member sits
outside this boundary entirely and uses its own provider credential.
- **One gateway:** `bridged` is the **sole communication path** for every session. The lead and
every member mount it as an MCP server and talk *only* to it — **no session ever addresses a
broker, a peer, or the network directly.** How the mount happens differs by backend: a
`claude-code` member gets one `claude mcp add` line or a shared `.mcp.json`, while an `opencode`
member gets a generated `opencode.json` and no `ANTHROPIC_*` variables at all. Any broker or
queue is `bridged`-internal, below the gateway — and it is optional. With `broker:` commented
out the daemon uses an in-memory inbox; when it is on, it is **LavinMQ** over AMQP.
- **How the lead gets a reply:** in the simple case, **one blocking MCP tool call** (`bridge_send`)
that `bridged` holds open until the member replies (`bridge_reply`) or its turn completes, then
returns the reply as the tool result. In practice that call is capped by the lead's own MCP client
timeout (about 60 seconds), so real work uses `wait:false` and a ticket instead. Either way the
lead never busy-polls and never touches a broker: when a detached ticket goes terminal, `bridged`
**injects the lead's own idle pane** to wake it.
- **Different model** per member process sidesteps Claude Code's lack of per-subagent
provider routing — the worker isn't a subagent, it's its own configured process.
- **AgentAPI** ([`coder/agentapi`](https://github.com/coder/agentapi)) is retained only as a
swappable *fallback injector* behind the same interface. See [Approaches](3-Approaches) for the
comparison and [Message Server](2-Message-Server) for the full design.
- **AgentAPI** ([`coder/agentapi`](https://github.com/coder/agentapi)) was kept on paper as a
swappable *fallback injector*. It was **never built** — `grep -ri agentapi bridged/src/main`
returns nothing, and the only injection path in the shipped code is the herdr one. Treat it as a
discarded option, not a fallback you can switch to. See [Approaches](3-Approaches) for why herdr
won and [Message Server](2-Message-Server) for the full design.
## Pages