wiki: rewrite 1-Architecture from scratch (clean-slate, final consolidated design)

Discard the incrementally-accreted Architecture page and compose a fresh,
coherent one around the two invariants (subscription boundary + sole gateway).
Tighter structure, 4 purposeful diagrams (system overview, blocking seq, async
seq, lifecycle state), no legacy framing. All validated with mmdc.
Dai Ha
2026-07-12 06:57:01 +02:00
parent 78bb89fcf0
commit 22c1e20e83
+190 -217
@@ -1,187 +1,161 @@
# 1. Architecture
`claude-bridge` connects a **primary** Claude Code session (Opus 4.8, on your Pro/Max
subscription) to one or more **secondary** Claude Code **workers** running a *different*
model.
`claude-bridge` lets a **primary** Claude Code session — Opus 4.8 on your Pro/Max
subscription — drive one or more **secondary worker** Claude Code sessions running a
*different, cheaper/local* model, **without ever putting a proxy on the primary**. A
standalone daemon, **`bridged`**, sits between them: it drives [herdr](https://herdr.dev) (an
agent multiplexer) to inject turns and read live agent-status, and exposes an **MCP server**
that every Claude session mounts.
> **The gateway invariant.** `bridged` is the **sole communication gateway**: every Claude
> session — primary and workers alike — talks *only* to `bridged`, over the MCP tools it
> mounts. **No Claude session ever addresses the broker, a peer session, or the network
> directly.** Anything else — the herdr socket, a durable queue, an event-bus, cross-host
> transport — lives *inside or south of* `bridged` and is invisible to the Claude sessions.
Two invariants define the whole design; everything else follows from them.
Two modes of traffic cross that gateway, but both are pure MCP from the Claude side:
> **1 — Subscription boundary.** Only a *worker* process ever sets `ANTHROPIC_BASE_URL`; the
> primary never does, so it stays on Pro/Max. `bridged` is not a `claude` process and holds
> zero Anthropic quota.
>
> **2 — Sole gateway.** Every Claude session talks *only* to `bridged`, over the MCP tools it
> mounts. No session addresses a broker, a peer session, or the network directly.
1. **Request/response (blocking)** — the primary delegates a task with **one blocking MCP
tool call** (`bridge_send`) that `bridged` (driving [herdr](https://herdr.dev)) holds open
until the worker's turn completes, then returns the reply as the tool result. This is the
main channel; its full design is in **[Message Server](2-Message-Server)**. "Blocking" here
means a single tool call parked on a result — *not* a busy-poll — so it costs the primary
no quota. Both the primary and the workers reach `bridged` by **mounting it as an MCP
server** — one unified Claude setup (see [Message Server](2-Message-Server)).
2. **Asynchronous / duplex** — for traffic with no caller waiting on a connection (a detached
progress report, an out-of-band question, a webhook injecting work), **`bridged` delivers
it by injecting the recipient's idle pane over herdr** — event-driven off the live
`agent_status`, not a poll loop. If `bridged` needs durability or a host hop, *it* owns a
broker for that, **below** the gateway line. A `Stop`-hook survives only as a split-host
escape hatch for a primary `bridged` cannot inject into — and even then it polls `bridged`,
never the broker (see [Channel 2](#channel-2--bridged-mediated-async-duplex)).
The engine of the sync channel is **herdr**, fronted by `bridged`: herdr owns the PTYs,
multiplexing, persistence, and **agent-status events**; `bridged` owns policy (the
subscription boundary, session lifecycle, status-gated delivery) and the client-facing
contract — an **MCP server** both Claude sessions mount, with REST/SSE for non-Claude
clients. (This supersedes the earlier plan, which used `coder/agentapi` as the
sync transport and treated herdr as an optional ops layer — see [Approaches](3-Approaches) for why the
positions swapped.)
## Components
## System overview
```mermaid
flowchart TB
subgraph prim["PRIMARY — subscription (env CLEAN) · MCP client"]
OPUS["Claude Code · Opus 4.8<br/>leads, reviews, merges"]
subgraph herd["herdr — agent multiplexer (Claude sessions run as panes)"]
PP["primary pane · Opus 4.8<br/>env CLEAN · MCP client"]
WP["worker pane(s) · claude<br/>ANTHROPIC_BASE_URL set · MCP client"]
end
subgraph work["SECONDARY worker host — off-subscription"]
subgraph BD["bridged — SOLE GATEWAY (standalone daemon, not a claude process)"]
SRV["SERVER face<br/>MCP server · REST/SSE · policy"]
CLI["CLIENT face<br/>herdr socket client"]
SRV --> CLI
end
HERDR["herdr<br/>panes · agent-status"]
WCC["worker pane · claude<br/>ANTHROPIC_BASE_URL set · MCP client"]
BROKER["broker / durable queue<br/>(bridged-owned · below the gateway)"]
CLI -->|"Unix socket<br/>send_text · events.subscribe"| HERDR
HERDR -->|"drives PTY"| WCC
SRV -.->|"durability · cross-host (internal)"| BROKER
subgraph BD["bridged — standalone daemon · THE gateway (no Anthropic quota)"]
SRV["SERVER face<br/>MCP server · REST/SSE · policy brain"]
CLI["CLIENT face<br/>status-gated injector · herdr socket client"]
SRV --> CLI
end
MODEL["ollama.ltms.dev /v1<br/>or GX10 vLLM<br/>(worker model)"]
Q["broker / queue<br/>(bridged-owned · below the gateway)"]
M["worker model<br/>ollama.ltms.dev · GX10 vLLM"]
OPUS -->|"MCP bridge_send → reply in tool result"| SRV
WCC -.->|"MCP bridge_reply / bridge_ask"| SRV
SRV -.->|"SSE status (observers)"| OPUS
WCC -->|"inference"| MODEL
PP -->|"MCP tools"| SRV
WP -->|"MCP tools"| SRV
CLI -->|"Unix socket · send_text · events.subscribe"| herd
SRV -.->|"durability · cross-host (internal)"| Q
WP -->|"inference"| M
classDef sub fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
classDef pick fill:#2f855a,stroke:#22543d,color:#ffffff;
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
classDef warn fill:#b7791f,stroke:#7b341e,color:#ffffff;
class OPUS sub
class SRV,CLI,HERDR,WCC pick
class BROKER warn
class PP ext
class SRV,CLI core
class Q,M warn
```
*Figure: every Claude session — primary and worker — connects **only** to `bridged`'s SERVER
face over MCP; async delivery is `bridged` injecting an idle pane through its CLIENT face over
herdr. The broker sits **below the gateway line**, owned by `bridged` for durability/cross-host
and never touched by a Claude session. The primary's env stays clean; only the worker sets
`ANTHROPIC_BASE_URL`, and `bridged` enforces that boundary in code.*
*Figure: the Claude sessions are **herdr panes**; they reach *up* into `bridged`'s SERVER face
over MCP, while `bridged`'s CLIENT face drives them *down* through herdr's socket. `bridged` is
the only thing any session connects to — the broker sits below the gateway line, owned by
`bridged`, touched by no session. Only worker panes carry `ANTHROPIC_BASE_URL`.*
## Subscription boundary (the non-negotiable)
## The two invariants
The whole design exists to keep the **primary** on Pro/Max while the **worker** runs a
cheaper/local model — without a policy-violating proxy on the primary.
### Subscription boundary — why the bridge exists
- The **primary** `claude` process **never** sets `ANTHROPIC_BASE_URL`. It reaches the worker
**only through `bridged`'s MCP tools** — never by re-pointing its own endpoint, and never by
addressing a broker or the worker directly.
- Only the **worker's** `claude` process launches with
`ANTHROPIC_BASE_URL=https://ollama.ltms.dev` (+ `ANTHROPIC_AUTH_TOKEN` bearer) or a GX10
vLLM URL. Because it is a *separate process*, model selection is just per-process env —
which also sidesteps Claude Code's lack of per-subagent provider routing (the worker is
not a subagent).
- **`bridged` is not a `claude` process** and consumes zero Anthropic quota, so it may
busy-poll the broker and hold a permanent herdr event subscription with no policy concern.
Its **subscription guard** refuses to spawn a *worker* pane without an off-subscription
`ANTHROPIC_BASE_URL` and refuses to ever set it on a pane tagged *primary*. The guard only
constrains panes `bridged` itself spawns; it **cannot** inspect a primary it does not host
(e.g. an Opus on your Mac, split-host) — that primary's env cleanliness is the operator's
responsibility, backed by a best-effort startup self-check (see [Message Server](2-Message-Server)).
- **Worker → primary rides `bridged`'s MCP rendezvous, not a keystroke.** A blocking
`bridge_send` is resolved by the worker's `bridge_reply` (or the `done` event) and the
primary reads the answer as an ordinary MCP **tool result** — no typing into the primary
pane, even single-host. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so it is
subscription-safe by construction. Two fallbacks remain: (a) *single-host, non-MCP primary*
— `bridged` can type into the primary pane (simulated typing, identical to the human at the
keyboard; the primary still authenticates to `api.anthropic.com` on Pro/Max); (b)
*split-host / detached* — a primary `bridged` cannot inject into wakes via its own
`Stop`-hook, which **long-polls `bridged`** (not the broker) for queued messages (Channel 2).
In both fallbacks the Claude side still speaks only to `bridged`. Read every page with that
topology split in mind.
The point of the bridge is to keep the **primary** on Pro/Max while a **worker** runs a
cheaper/local model, with no policy-violating proxy on the primary.
> **Rule:** anything that sets `ANTHROPIC_BASE_URL` is, by definition, the worker. If you
> ever feel tempted to set it on the primary, stop — that is the subscription line.
- The **primary** `claude` **never** sets `ANTHROPIC_BASE_URL`. It authenticates to
`api.anthropic.com` on your subscription and reaches the worker *only* through `bridged`'s
MCP tools — never by re-pointing its own endpoint.
- Only a **worker** `claude` launches with `ANTHROPIC_BASE_URL=https://ollama.ltms.dev`
(+ `ANTHROPIC_AUTH_TOKEN`) or a GX10 vLLM URL. Because model choice is just per-process env,
this also sidesteps Claude Code's lack of per-subagent provider routing — the worker is a
*separate process*, not a subagent.
- **`bridged` holds no quota**, so it may hold a permanent herdr event subscription and poll
its own internal queue with no policy concern. Its **subscription guard** refuses to spawn a
worker pane whose resolved `ANTHROPIC_BASE_URL` host isn't on an off-subscription allowlist,
and refuses to ever set that var on a pane tagged *primary*. (The guard validates the
resolved host and confirms egress post-spawn — a substring check on the launch string is not
enough; see [Message Server](2-Message-Server).)
## Channel 1 — `bridged` over herdr (blocking request/response, the main path)
> **Rule:** anything that sets `ANTHROPIC_BASE_URL` is, by definition, a worker. If you are
> ever tempted to set it on the primary, stop — that is the subscription line.
`bridged` exposes an **MCP server** that both the primary and the workers mount (`bridge_send`,
`bridge_reply`, `bridge_ask`, `bridge_status`), translates those tool calls into herdr socket
calls, and — unlike a screen-stability heuristic — gates every injection on herdr's live
`agent_status_changed` events. REST/SSE remains for non-Claude clients. One session per pane;
many panes per herdr. Full contract, components, and the Go interface sketch are in
**[Message Server](2-Message-Server)**.
### Sole gateway — how sessions communicate
The primary's `bridge_send` **blocks** until `bridged` resolves the reply — on the worker's
`bridge_reply` (structured) or the `agent_status = done` event, whichever lands first — and
returns it as the tool result. Because the reply rides `bridged`'s own state, worker → primary
needs **no** keystroke into the primary pane, even single-host. SSE (`GET /events`) is a
parallel channel for *observers* (a human, a dashboard) to watch `working`/`blocked`/`done`
transitions; the primary does not need to hold it. For a task that may outrun a reasonable
request timeout, prefer Channel 2 (fire-and-forget, reply comes back async).
Because every session mounts `bridged` over MCP, `bridged` is the single chokepoint for *all*
agent traffic. This is a deliberate simplification with real payoffs:
- **No session-side transport.** A Claude session calls MCP tools and nothing else — no
`curl`, no broker client, no hand-rolled hook polling a queue. There is no shell step that
could leak env, so mounting the bridge is subscription-safe by construction.
- **Async is a push, not a poll.** When a message arrives for an idle session, `bridged`
**injects it into that session's pane over herdr**, gated on the live `agent_status`. The
session is never asked to busy-poll anything; the old "primary must never perpetual-poll"
footgun disappears because there is nothing for it to poll.
- **Guardrails are centrally enforced.** Round/turn budgets (anti ping-pong), rate limits,
per-session authz, and audit all live in `bridged` — one place — instead of cooperative
sentinels each session must honour.
- **The broker is infrastructure, below the line.** If `bridged` needs durability or a host
hop it owns a broker/queue for that. No Claude session sees it.
## Components
| Component | Role | Notes |
|---|---|---|
| **`bridged` — SERVER face** | The gateway: **MCP server** (the contract every session mounts) + REST/SSE for non-Claude clients, over the **policy brain** — session tracker, subscription guard, reply rendezvous. | `bridge_send` · `bridge_reply` · `bridge_ask` · `bridge_status` · `bridge_poll` · `bridge_sessions`. |
| **`bridged` — CLIENT face** | Drives herdr: a **status-gated injector** (per-pane FIFO, delivers only when `agent_status ∈ {idle, blocked}`) over a **herdr socket client** (NDJSON, id-correlated, live event stream). | Single writer per pane → no injector-vs-injector races. |
| **herdr** | Agent multiplexer. Owns the PTYs, panes, persistence, and — crucially — **`agent_status_changed` events**. Claude sessions run here as panes. | Socket is **local-only**; young/single-dev → injector kept pluggable. |
| **Worker `claude`** | A *real* Claude Code process (inherits `CLAUDE.md`, hooks, skills, MCP), pointed at a different model. Recyclable, not immortal. | Only these carry `ANTHROPIC_BASE_URL`. |
| **Broker / queue** *(optional, internal)* | `bridged`-owned durability + cross-host transport, **below the gateway**. Enqueues async messages `bridged` will later inject. | Redis Streams / NATS JetStream, or an embedded queue for a single host. |
| **AgentAPI** *(fallback)* | Swappable injector behind the CLIENT face if herdr is unavailable. | Screen-stability heuristic instead of events — see [Approaches](3-Approaches). |
## Traffic: two modes across the gateway
Both modes are pure MCP from the Claude side; they differ only in whether a caller waits on
the connection.
### Mode 1 — blocking request/response (the main path)
The primary delegates with **one** `bridge_send` tool call. `bridged` holds it open (the
primary is idle-waiting, spending no quota) and **resolves it on whichever lands first**: the
worker's structured `bridge_reply`, or herdr's `agent_status = done` event. The reply comes
back as the tool result — so worker → primary rides `bridged`'s state and needs **no keystroke
into the primary pane**, even single-host.
```mermaid
sequenceDiagram
participant P as "Primary (Opus)"
participant S as "bridged"
participant H as "herdr"
participant W as "Worker claude (other model)"
P->>S: "bridge_send(task, full context inlined) — tool call PARKS"
S->>S: "await pane status = idle"
S->>H: "pane.send_text + send_keys enter"
H->>W: "inject into running session"
activate W
H-->>S: "event: agent_status_changed = working"
S-->>P: "SSE: status working (observers only)"
W->>S: "bridge_reply(result) — or agent_status=done (fallback)"
H-->>S: "event: agent_status_changed = done"
deactivate W
S-->>P: "tool result = assistant reply (unparks the call)"
Note over P: "review diff / result, merge"
alt worker is blocked
H-->>S: "event: agent_status_changed = blocked"
S-->>P: "tool result = blocked + question"
P->>S: "bridge_send(answer inlined) — new blocking call"
end
```
*Figure: the primary's request blocks until the reply is ready and reads it from the response
body; SSE carries status to observers, not the answer. `bridged` injects only when the pane is
`idle`, and a blocked worker surfaces on a real herdr event (not a guess). A block returns to
the primary as the call's result; the primary re-sends the answer with a fresh blocking call.*
Why herdr over the earlier `agentapi` plan: structured **agent-status events** (vs a
screen-stability heuristic), **symmetric** injection into either pane (single-host), native
**multiplexing** of a herd of workers, and **persistence/detach**. AgentAPI is kept as a
swappable *fallback injector* behind the same interface. See [Approaches](3-Approaches) for the full
transport comparison and [Message Server](2-Message-Server) for the design.
## Channel 2 — bridged-mediated async (duplex)
For traffic with **no caller waiting on a connection** — a long-running worker posting
progress, an out-of-band question, another agent or a webhook injecting work — the recipient
must be *woken*. Under the gateway invariant, **`bridged` does the waking by injecting the
recipient's idle pane over herdr** — driven by the live `agent_status`, so it delivers the
instant the pane goes idle rather than on a timed poll. No Claude session polls, writes to, or
even knows about a broker; if `bridged` needs durability or a host hop it enqueues internally
(below the gateway) and still delivers by injection.
```mermaid
sequenceDiagram
participant SRC as "Source (worker bridge_reply/ask · webhook · bus)"
participant P as "Primary (Opus) — MCP client"
participant S as "bridged (gateway)"
participant Q as "broker / queue (internal)"
participant H as "herdr"
participant W as "Worker (other model)"
P->>S: "bridge_send(worker, task) — tool call PARKS"
S->>S: "await pane agent_status = idle"
S->>H: "pane.send_text + send_keys (inject)"
H->>W: "new turn"
activate W
H-->>S: "event: agent_status = working"
W->>S: "bridge_reply(result) — structured (preferred)"
H-->>S: "event: agent_status = done"
deactivate W
S-->>P: "tool result = reply (unparks the call)"
Note over P,W: "resolves on bridge_reply or the done event — whichever lands first.<br/>a blocked worker returns as the tool result, then the primary re-answers"
```
*Figure: a single parked tool call, not a busy-poll. SSE (`GET /events`) carries status to
*observers* (a human, a dashboard) in parallel; the primary never has to hold it. For work that
may outrun a sane request timeout, use Mode 2.*
### Mode 2 — asynchronous delivery (bridged-mediated)
For traffic with no caller waiting — a detached progress report, an out-of-band question, a
webhook injecting work — the recipient must be *woken*. `bridged` does the waking by
**injecting the recipient's idle pane**, driven by the live `agent_status`, so delivery is
event-driven rather than a timed poll. If durability or a host hop is needed, `bridged`
enqueues internally first; the recipient still receives by injection.
```mermaid
sequenceDiagram
participant SRC as "Source: worker bridge_reply/ask · webhook · bus"
participant S as "bridged (gateway)"
participant Q as "queue (internal)"
participant H as "herdr"
participant R as "Recipient pane (idle Claude)"
@@ -192,87 +166,86 @@ sequenceDiagram
S->>H: "await recipient agent_status = idle"
H-->>S: "event: idle"
S->>H: "pane.send_text + send_keys (inject)"
H->>R: "new turn = the async message"
Note over S,R: "recipient polled nothing —<br/>bridged pushed on the idle edge"
H->>R: "new turn = the message"
Note over S,R: "recipient polled nothing — bridged pushed on the idle edge"
```
*Figure: `bridged` is the mediator for async too. It accepts the message over MCP (or REST for
non-Claude sources), optionally parks it on its **internal** queue, waits for the recipient's
idle event, and injects. The **only** exception is a primary `bridged` cannot inject into
(split-host, off-herdr): that primary runs a `Stop`-hook which long-polls **`bridged`'s**
inbox endpoint — still the gateway, still never the broker.*
*Figure: `bridged` mediates async in both directions. Sources reach it over the gateway (MCP
for Claude, REST for external), it optionally parks the message on its internal queue, waits
for the idle event, and injects.*
### Guardrails (enforced centrally in `bridged`)
> **The one exception.** A **split-host primary that is not a herdr pane** (e.g. Opus on your
> Mac) is the sole session `bridged` cannot inject into. There, the primary runs a `Stop`-hook
> that **long-polls `bridged`** for queued messages — still the gateway, **never** the broker.
> The gateway invariant holds in every topology.
Collapsing everything behind one gateway turns these from cooperative conventions into
`bridged`-enforced policy — a real win over per-hook envelope sentinels:
## Worker lifecycle — the Ralph loop
| Risk | Mitigation |
|------|------------|
| **Primary quota burn** | Structurally impossible now — the primary has no broker to poll and no perpetual loop. It either blocks on one MCP call (`bridged` holds it, idle-waiting) or is injected on its idle edge. The split-host `Stop`-hook fires only at a natural turn boundary, never in a spin. |
| **Cross-agent ping-pong** | A→B→A→B can loop forever. `bridged` sees every hop (sole gateway), so it enforces a **round/turn budget centrally** and drops on breach — no reliance on a `no-reply-needed` sentinel each side must honor. |
| **Lost / double-processed messages** | `bridged`'s internal queue uses **ack + visibility timeout + consumer groups** (Redis Streams `XACK`, NATS JetStream). A crash mid-turn re-delivers instead of dropping. |
| **Context growth** | A perpetual worker's context window fills up. `bridged` caps idle cycles / tokens, then **recycles the pane fresh with state on the filesystem** (Ralph loop — see [Message Server](2-Message-Server)). Perpetual != one infinite session. |
## Deployment shape (target)
`bridged` treats a worker as a **recyclable** resource, not one immortal session: a long-lived
pane fills its context window. On a context/idle cap it checkpoints and respawns fresh, with
continuity carried by **artifacts on disk** (git commits, a `STATE.md`/task file the worker is
told to keep current) — **not** `claude --resume`, which would reload the context you are
trying to shed.
```mermaid
flowchart LR
subgraph mac["Your machine (subscription)"]
OPUS["Primary Opus<br/>(Claude Code)"]
end
subgraph host["Worker host (off-subscription, near model)"]
direction TB
BD["bridged :8080<br/>MCP · REST/SSE"]
HS["herdr server"]
W2["worker claude pane(s)"]
BR["broker / queue<br/>(bridged-owned, internal)"]
BD -->|"Unix socket"| HS --> W2
BD -.->|"durability / cross-host"| BR
end
ML["ollama.ltms.dev / GX10 vLLM"]
OPUS -->|"MCP over HTTP — the only link (sync + async)"| BD
W2 --> ML
classDef sub fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
classDef pick fill:#2f855a,stroke:#22543d,color:#ffffff;
class OPUS sub
class BD,HS pick
stateDiagram-v2
[*] --> Spawning
Spawning --> Ready: "claude prompt detected"
Ready --> Working: "turn injected"
Working --> Blocked: "permission / question"
Blocked --> Working: "bridged answers (send_input)"
Working --> Ready: "agent_status = done"
Ready --> Recycling: "context / idle cap hit"
Recycling --> Spawning: "state persisted to disk"
Working --> Failed: "pane.exited (crash)"
Failed --> Spawning: "auto-restart + replay unacked"
Ready --> [*]: "drain / shutdown"
```
*Figure: `bridged` is the **only** thing the primary connects to. Sync replies come back on
the blocking MCP call; because this split-host primary isn't a herdr pane, async wake-ups
arrive via its `Stop`-hook polling **that same `bridged` endpoint** — never a broker. herdr's
socket and the broker are local to the worker host and `bridged`-owned; neither crosses to the
primary. Bind `bridged` to localhost + tunnel, or front it with a token; never expose the port
unauthenticated.*
*Figure: `Blocked`, `done`, and `exited` are real herdr events, not heuristics — the reason
herdr-centric beats screen scraping. Full lifecycle detail is in
[Message Server](2-Message-Server) → *Worker session lifecycle*.*
## Topologies
- **Same-host (default).** Primary, `bridged`, herdr, and workers on one off-subscription box.
Every session is a herdr pane, so `bridged` can inject *either* direction; a broker is
optional. Simplest to run and the focus of the design.
- **Split-host.** Primary Opus on your Mac; `bridged` + herdr + workers on the GPU host near
the model. herdr's socket stays local, so the Mac reaches the worker host **only over
`bridged`'s MCP/HTTP endpoint**. The primary isn't a herdr pane, so async wake-ups use the
`Stop`-hook-polls-`bridged` path above. Deployment diagrams: [Message Server](2-Message-Server)
→ *Deployment model*.
**Security:** `bridged` is an agent-control surface — `bridge_send` runs arbitrary prompts and
key-passthrough sends raw keystrokes into a live agent. Bind it to `localhost` + SSH tunnel, or
front it with a bearer token + TLS; never expose the port unauthenticated. One `bridged` is
today **one trust domain** (no per-session authz yet — an open item in
[Message Server](2-Message-Server)).
## Failure modes & single points of failure
Both channels have a distinct SPOF; neither is redundant in the target design, so degrade
Making `bridged` the sole gateway buys a clean model at the cost of a real SPOF. Degrade
deliberately:
| What dies | Effect | Degradation / recovery |
| What dies | Effect | Recovery |
|---|---|---|
| **`bridged`** | **The whole gateway is down** — no delegations, no async wake-ups, in-flight blocking calls error out (it is the sole gateway, so sync *and* async stop together) | herdr + workers keep running (state on disk / queue). systemd restarts `bridged`; it re-attaches to existing panes via `session.snapshot` and drains its queue. Nothing reaches a Claude session in the meantime — by design there is no side path. |
| **herdr** | No pane control at all; delivery (sync and async injection) dead | Workers' PTYs die with the herdr server (no detach survives a *server* crash, only client detach). Respawn from persisted worker state (Ralph loop); replay unacked queue items. |
| **Broker / queue** (internal) | Durability + cross-host async degrade; **same-host async still works** (direct idle-pane injection needs no queue) | `bridged` can deliver locally without it; only durable replay and host-hop traffic pause. Ack + visibility timeout re-deliver on recovery; nothing silently dropped. |
| **Model endpoint** (`ollama.ltms.dev` / vLLM) | Workers stall or error mid-turn | herdr status shows `working` stuck or `blocked`; `bridged` times out the blocking call and surfaces the error. Primary (subscription) is never affected. |
| **All three** | Full sync + async outage | Primary Opus remains fully usable on its own subscription — the bridge is additive, never on the primary's critical path. |
The primary is deliberately **not** downstream of any bridge component: a total bridge
outage costs you the workers, never your own session.
| **`bridged`** | **All** agent comms stop — sync *and* async — since it is the only gateway; in-flight blocking calls error out. | herdr + workers keep running (state on disk / queue). systemd restarts `bridged`; it re-attaches to existing panes via `session.snapshot` and drains its queue. This restart path is load-bearing — harden it. |
| **herdr** | No pane control; all delivery (sync + async injection) dead. | PTYs die with the *server* (only client detach survives). Respawn workers from persisted state (Ralph loop); replay unacked queue items. |
| **Broker / queue** (internal) | Durability + cross-host async degrade; **same-host async still works** (idle-injection needs no queue). | `bridged` delivers locally without it; ack + visibility timeout re-deliver on recovery. Nothing silently dropped. |
| **Model endpoint** | Workers stall or error mid-turn. | herdr status shows `working` stuck / `blocked`; `bridged` times out the blocking call and surfaces the error. |
| **All of the above** | Full bridge outage. | The **primary is never downstream** of any bridge component — it stays fully usable on its own subscription. The bridge is additive, never on the primary's critical path. |
## Related pages
- **[Message Server](2-Message-Server)** — the `bridged` design: herdr control contract, lifecycle, API, tech stack
- **[Approaches](3-Approaches)** — why herdr-centric, and the full transport comparison (AgentAPI, SDK, Stop-hook, tmux)
- **[Setup](4-Setup)** — running herdr + `bridged` + a worker pointed at `ollama.ltms.dev`
- **[Operations](5-Operations)** — health, restart, model swaps, troubleshooting
- **[Message Server](2-Message-Server)** — the `bridged` design in depth: MCP contract, herdr control, reply rendezvous, lifecycle, API, tech stack, milestones.
- **[Approaches](3-Approaches)** — why herdr-centric, and the full transport comparison (AgentAPI, Agent SDK, queue+Stop-hook, tmux).
- **[Team](6-Team)** — a team-lead orchestrating a mixed Claude + local-LLM worker fleet over the same gateway.
- **[Setup](4-Setup)** · **[Operations](5-Operations)** — bring-up and the day-2 runbook.
## Sources
- [herdr — socket API](https://herdr.dev/docs/socket-api/) · [agents / state detection](https://herdr.dev/docs/agents/) · [GitHub](https://github.com/ogulcancelik/herdr)
- [coder/agentapi](https://github.com/coder/agentapi) — fallback injector
- [Agent Room — Stop-hook async collaboration](https://dev.to/agent-room/how-a-claude-code-stop-hook-unlocks-async-multi-agent-collaboration-no-polling-required-2e0e)
- [Hooks reference — Claude Code Docs](https://code.claude.com/docs/en/hooks)
</content>
- [Model Context Protocol](https://modelcontextprotocol.io) — the client contract both sessions mount
- [Hooks reference — Claude Code Docs](https://code.claude.com/docs/en/hooks) — the split-host `Stop`-hook adapter