9daf1ec5ba
Closes out single-host before the cross-host work. Sequenced BEFORE CB-308 deliberately: federation's own gating concern is the trust model, and it inherits whatever identity shape lands here. The finding this stage is built around: bridged had exactly ONE security control, the loopback bind. ConnectionIdentity resolves a worker from its connection (unforgeable), but every caller that was not a recognised worker pane fell through to being treated as the PRIMARY -- the most privileged role on the bus. Latent today; load-bearing the moment a bind widens. CB-501 auth: - Role/Principal/CallerResolver: connection identity first, bearer token second, ANONYMOUS third. Inverts the old default so absence of identity means nothing, not everything. - Worker identity is never token-gated, so enabling auth cannot lock the fleet out of bridge_reply. - Constant-time token compare (MessageDigest.isEqual). - validateAuthExposure(): a non-loopback bind under loopback-trust now REFUSES TO START. Makes the dangerous config unrepresentable rather than merely documented. - TLS terminates at a reverse proxy by design (D3), not in the JVM. CB-505 authz + audit, enforced on BOTH entry paths: - The docs describe MCP as "a thin adapter over the REST core"; at code level it is not. BridgeMcp calls MessageService directly, and /mcp is a raw servlet on Jetty's context handler that never traverses Javalin's before filter. Enforcing only at REST would have left /mcp open. - Load-bearing rule is own-session-only: a worker may reply/ask only as itself. Structurally true over MCP already; over REST the session id in the URL path had simply been trusted. - Audit: JSON lines to a dedicated appender, additivity=false. Never records message content -- this bus carries source and prompts. CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer instead of the specced Micrometer, because this pom already hand-pins jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated dependency CVE gate could not be run (no JetBrains MCP server connected). Instrumented at MessageService, the single funnel both surfaces share. CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner. Needs no contract-exclusion flag -- the pom's default-excludes profile already sets excludedGroups=contract, so plain `mvn clean install` IS the mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older). CB-504 supervision: launchd agent (the real target -- this host is macOS, there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds. Ordering directives are advisory, so the actual fix is that startup now waits up to 30s for the herdr socket and then serves degraded, instead of crashing into a restart loop on a boot-order race. Also fixes drift found while surveying: - bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config binds via plain Jackson with ignoreUnknown, so uncommenting it would have been silently dropped and the default kept. Now camelCase, with a test that loads the shipped example and one that pins every documented knob's spelling -- no test had ever loaded that file. - Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay, gitTokenEnv, gitHostEnv, configDir, primary:). - README "Next" listed bridge_ask and session lifecycle as upcoming; both shipped long ago. - docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre- implementation" for work already merged. 307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS. Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check could not be run -- no JetBrains/intellij-index MCP server is connected this session. mvn clean install is the only gate that ran.
124 lines
7.3 KiB
Markdown
124 lines
7.3 KiB
Markdown
# claude-bridge
|
|
|
|
A **subscription-safe bridge** that lets a primary **Claude Code (Opus 4.8, on Pro/Max)**
|
|
session drive a **secondary Claude agent running a different model** via its own
|
|
`ANTHROPIC_BASE_URL` — without ever putting a proxy on the primary session.
|
|
|
|
Sibling of [`crush-bridge`](https://git.ltms.dev/systems/vms) (which drives a headless
|
|
**Crush** worker on GX10 DeepSeek). `claude-bridge` keeps the worker a *real Claude Code
|
|
process*, so it inherits `CLAUDE.md`, hooks, skills, and MCP — just pointed at a
|
|
cheaper/local model.
|
|
|
|
## Leading approach — herdr-centric message server (`bridged`)
|
|
|
|
A small always-on message server, **`bridged`**, controls
|
|
[herdr](https://herdr.dev) (an agent multiplexer) over its Unix-socket API and exposes a
|
|
clean 2-way messaging API as an **MCP server that both the primary and the workers mount** —
|
|
one unified Claude setup and the **sole communication gateway** (REST/SSE stays for non-Claude
|
|
clients; any broker is `bridged`-internal, below the gateway).
|
|
herdr owns the PTYs, multiplexing, persistence, and **agent-status events**; `bridged` owns
|
|
policy (subscription boundary, session lifecycle, status-gated delivery) and the client
|
|
contract. The worker `claude` launches with `ANTHROPIC_BASE_URL=https://ollama.ltms.dev` + a
|
|
bearer token; the primary Opus stays env-clean and calls `bridged`'s MCP tools.
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
|
|
subgraph BD["bridged — standalone daemon (not a claude process)"]
|
|
SRV["SERVER face<br/>MCP · REST/SSE · policy"]
|
|
CLI["CLIENT face<br/>status-gated injector · herdr socket"]
|
|
SRV --> CLI
|
|
end
|
|
HERDR["herdr<br/>panes · agent-status"]
|
|
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
|
|
M["ollama.ltms.dev<br/>(worker model)"]
|
|
|
|
OPUS -->|"MCP bridge_send (blocks)"| SRV
|
|
W -.->|"MCP bridge_reply"| SRV
|
|
CLI -->|"Unix socket<br/>send_text · events.subscribe"| HERDR
|
|
HERDR -->|"drives PTY"| W
|
|
W -->|"inference"| M
|
|
|
|
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
|
|
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
|
|
class OPUS ext
|
|
class SRV,CLI,HERDR core
|
|
```
|
|
|
|
- **Subscription boundary:** the *primary* never sets `ANTHROPIC_BASE_URL` (stays on
|
|
Pro/Max). Only the *secondary* process is off-subscription — and `bridged` itself is a
|
|
plain daemon (no Anthropic quota), so it may poll/subscribe freely.
|
|
- **One gateway (unified MCP setup):** `bridged` is the **sole communication path** for every
|
|
Claude session. Primary and workers each mount it as an MCP server (one `claude mcp add`
|
|
line, same on both) and talk over MCP tools — `bridge_send` / `bridge_reply` /
|
|
`bridge_status` (with `bridge_ask` planned for the blocked-worker path). **No Claude session
|
|
ever addresses a broker, a peer, or the network
|
|
directly**; any queue is `bridged`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so
|
|
mounting the bridge is subscription-safe by construction.
|
|
- **How the primary consumes a reply:** a single **blocking MCP call** (`bridge_send`);
|
|
`bridged` holds it open until the worker calls `bridge_reply` or its turn hits
|
|
`agent_status=done`, then returns the reply as the tool result. No cross-turn busy-poll, so
|
|
no quota burn. SSE is an optional side-channel for humans/dashboards watching status.
|
|
- **Worker → primary** rides `bridged`'s **MCP rendezvous** — the reply resolves the primary's
|
|
blocking call (or, for detached work, `bridged` **injects the primary's idle pane** when it's
|
|
ready), so *no keystroke-into-primary and no broker are involved, even single-host*. The one
|
|
exception: a split-host primary that isn't a herdr pane wakes via its own `Stop`-hook, which
|
|
polls **`bridged`** (never a broker). See the wiki for the two topologies.
|
|
- **Different model per process** sidesteps Claude Code's lack of per-subagent provider
|
|
routing — the worker isn't a subagent, it's its own configured process.
|
|
- **AgentAPI** ([`coder/agentapi`](https://github.com/coder/agentapi)) is retained only as a
|
|
swappable *fallback injector* behind the same interface. See the wiki for the full
|
|
design, comparison, and rationale.
|
|
|
|
## Docs
|
|
|
|
Full design, setup, and operations live in the **[wiki](https://git.ltms.dev/lms/claude-bridge/wiki)**,
|
|
vendored here as a submodule under [`wiki/`](./wiki):
|
|
|
|
```bash
|
|
git clone --recurse-submodules ssh://git@git.ltms.dev:2224/lms/claude-bridge.git
|
|
# or, after a plain clone:
|
|
git submodule update --init
|
|
```
|
|
|
|
Edit docs in `wiki/`, then `cd wiki && git commit && git push` to publish them to the
|
|
Gitea wiki.
|
|
|
|
## Status
|
|
|
|
🟢 **Implemented & dogfooded** — the herdr-centric **`bridged`** message server is built and in
|
|
real use: an Opus primary delegates tasks to off-subscription workers that reply through the
|
|
bridge (code reviews delegated this way have produced committed bug fixes). Selected as the
|
|
primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a
|
|
fallback injector.
|
|
|
|
**Shipped** (Java 25 · Maven · 266 unit/acceptance tests green; the live-herdr and broker contract
|
|
tests run separately via `mvn test -Pcontract`):
|
|
|
|
- **Core gateway** — herdr socket client (contract-tested vs live 0.7.0); guard-checked worker
|
|
spawn with `ANTHROPIC_BASE_URL` injected only into the worker's env; status-gated injector;
|
|
blocking `bridge_send` with reply rendezvous; MCP server as a thin adapter over the REST core.
|
|
- **MCP tools** — `bridge_send` / `bridge_reply` / `bridge_status` (messaging) and `bridge_spawn`
|
|
/ `bridge_list` / `bridge_stop` / `bridge_profiles` / `bridge_poll` (fleet). Caller identity is
|
|
connection-based (loopback peer PID → herdr pane), so the same mount serves primary and workers.
|
|
- **Delivery reliability** — completion fallback (a confirmed `working→idle` turn resolves a
|
|
send); async fire-and-poll (beats the caller's MCP call timeout for long tasks); and failure
|
|
detection for wedged (`unknown`), vanished, and never-ready workers so a send never hangs.
|
|
- **Fleet** — multiple worker profiles, each with an independent base_url guard check; workers
|
|
inherit the primary's working directory (never `$HOME`); a readiness gate holds delivery until
|
|
a worker's Claude has connected the bridge MCP (no paste lost into its boot window).
|
|
- **Blocked-worker path** — `bridge_ask` reverse rendezvous: a worker pauses its delegated turn to
|
|
ask the primary and resumes the *same* turn with the answer (CB-205).
|
|
- **Session lifecycle** — session manager with spawn/reuse/recycle, `idle_ttl` reaper, `context_cap`,
|
|
and graceful drain on shutdown (CB-301/CB-303); per-worker git worktrees on their own branch with
|
|
a config-parity overlay, so parallel implementers never stomp each other (CB-301-ext).
|
|
- **Reliable worker→primary delivery** — a durable `ReplyInbox` (in-memory by default, AMQP/LavinMQ
|
|
for cross-restart durability) holds a reply that arrives with no open send, and an active
|
|
status-gated push loop nudges the primary to drain it (CB-307).
|
|
- **Pluggable peers** — a `PeerLauncher` SPI with two in-tree adapters, `claude-code` and `opencode`,
|
|
routed by a `kind:` discriminator (CB-401/CB-402).
|
|
|
|
**Next** (see the [roadmap](wiki/8-Roadmap.md)) — Stage 5 hardening (auth/TLS, `/metrics`, CI,
|
|
service supervision, per-session authz + audit), then cross-host: CB-308 multi-host federation and
|
|
CB-500 multi-tier coordination.
|