kevin 5100f215cf
CI / build (push) Successful in 1m33s
CB-515: regression-protect the turn-attribution guards
Five tests pinning the invariants that decide WHICH turn a reply belongs to.
These protect against a silent correctness bug — a reply attributed to the
wrong turn — not against a crash, which is why they were worth picking over
higher-percentage coverage gaps.

Chosen by blast radius, not by uncovered-line count. Both guards are compound
conditions with a side that never executed, i.e. exactly the shape where a
clause can be deleted as "redundant" and every existing test still passes.

CompletionResolver:
- The CB-115 misattribution guard suppresses a completion when the scrape is
  byte-identical to the pane at delivery. Its !scrapeFailed clause was
  unexercised: delete it and a FAILED read is misread as "no output change",
  so the send is suppressed and hangs to the caller's timeout instead of
  resolving. The new test sets the baseline to "" so the empty tail from a
  failed read would byte-match and wrongly suppress — built to die precisely
  when that clause dies.
- The fail() guard leaves an already-resolved waiter alone. The new test also
  asserts agent.read is never called, so the worker is not scraped for a send
  nobody is waiting on.

Rendezvous: a second resolution of an already-completed waiter returns false
and does not overwrite the first value, for both resolveCompletion and
resolveFailure.

Verified by sabotage, one guard at a time: removing !scrapeFailed reds
resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent; removing the
isDone() clause reds failLeavesAnAlreadyResolvedWaiterUntouchedAndSkipsTheScrape.
(The first attempt at the second sabotage reported a false pass — the patch hit
an identically-worded guard earlier in the file. Line-targeted and re-run.)

346 tests, was 341. Worker-implemented on the local-vLLM profile; it noticed
three of the eight cases I asked for already existed and said so with names
rather than duplicating them.

Also of note: the first delegation of this ticket wedged the worker — the pane
showed a zsh parse error and it went idle with an untouched worktree, task stuck
pending. The retry differed only in phrasing the same requirements as prose
instead of quoting Java boolean expressions. Filed as a bridge robustness
concern: injected content shares a channel with control, and a wedged turn is
invisible in both the task view and /metrics.
2026-08-01 23:48:05 +07:00

claude-bridge

A subscription-safe bridge that lets a primary Claude Code (Opus 4.8, on Pro/Max) session drive a secondary Claude agent running a different model via its own ANTHROPIC_BASE_URL — without ever putting a proxy on the primary session.

Sibling of crush-bridge (which drives a headless Crush worker on GX10 DeepSeek). claude-bridge keeps the worker a real Claude Code process, so it inherits CLAUDE.md, hooks, skills, and MCP — just pointed at a cheaper/local model.

Leading approach — herdr-centric message server (bridged)

A small always-on message server, bridged, controls herdr (an agent multiplexer) over its Unix-socket API and exposes a clean 2-way messaging API as an MCP server that both the primary and the workers mount — one unified Claude setup and the sole communication gateway (REST/SSE stays for non-Claude clients; any broker is bridged-internal, below the gateway). herdr owns the PTYs, multiplexing, persistence, and agent-status events; bridged owns policy (subscription boundary, session lifecycle, status-gated delivery) and the client contract. The worker claude launches with ANTHROPIC_BASE_URL=https://ollama.ltms.dev + a bearer token; the primary Opus stays env-clean and calls bridged's MCP tools.

flowchart LR
    OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
    subgraph BD["bridged — standalone daemon (not a claude process)"]
        SRV["SERVER face<br/>MCP · REST/SSE · policy"]
        CLI["CLIENT face<br/>status-gated injector · herdr socket"]
        SRV --> CLI
    end
    HERDR["herdr<br/>panes · agent-status"]
    W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
    M["ollama.ltms.dev<br/>(worker model)"]

    OPUS -->|"MCP bridge_send (blocks)"| SRV
    W -.->|"MCP bridge_reply"| SRV
    CLI -->|"Unix socket<br/>send_text · events.subscribe"| HERDR
    HERDR -->|"drives PTY"| W
    W -->|"inference"| M

    classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
    classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
    class OPUS ext
    class SRV,CLI,HERDR core
  • Subscription boundary: the primary never sets ANTHROPIC_BASE_URL (stays on Pro/Max). Only the secondary process is off-subscription — and bridged itself is a plain daemon (no Anthropic quota), so it may poll/subscribe freely.
  • One gateway (unified MCP setup): bridged is the sole communication path for every Claude session. Primary and workers each mount it as an MCP server (one claude mcp add line, same on both) and talk over MCP tools — bridge_send / bridge_reply / bridge_status (with bridge_ask planned for the blocked-worker path). No Claude session ever addresses a broker, a peer, or the network directly; any queue is bridged-internal. MCP tool I/O never sets ANTHROPIC_BASE_URL, so mounting the bridge is subscription-safe by construction.
  • How the primary consumes a reply: a single blocking MCP call (bridge_send); bridged holds it open until the worker calls bridge_reply or its turn hits agent_status=done, then returns the reply as the tool result. No cross-turn busy-poll, so no quota burn. SSE is an optional side-channel for humans/dashboards watching status.
  • Worker → primary rides bridged's MCP rendezvous — the reply resolves the primary's blocking call (or, for detached work, bridged injects the primary's idle pane when it's ready), so no keystroke-into-primary and no broker are involved, even single-host. The one exception: a split-host primary that isn't a herdr pane wakes via its own Stop-hook, which polls bridged (never a broker). See the wiki for the two topologies.
  • Different model per process sidesteps Claude Code's lack of per-subagent provider routing — the worker isn't a subagent, it's its own configured process.
  • AgentAPI (coder/agentapi) is retained only as a swappable fallback injector behind the same interface. See the wiki for the full design, comparison, and rationale.

Docs

Full design, setup, and operations live in the wiki, vendored here as a submodule under wiki/:

git clone --recurse-submodules ssh://git@git.ltms.dev:2224/lms/claude-bridge.git
# or, after a plain clone:
git submodule update --init

Edit docs in wiki/, then cd wiki && git commit && git push to publish them to the Gitea wiki.

Status

🟢 Implemented & dogfooded — the herdr-centric bridged message server is built and in real use: an Opus primary delegates tasks to off-subscription workers that reply through the bridge (code reviews delegated this way have produced committed bug fixes). Selected as the primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a fallback injector.

Shipped (Java 25 · Maven · 266 unit/acceptance tests green; the live-herdr and broker contract tests run separately via mvn test -Pcontract):

  • Core gateway — herdr socket client (contract-tested vs live 0.7.0); guard-checked worker spawn with ANTHROPIC_BASE_URL injected only into the worker's env; status-gated injector; blocking bridge_send with reply rendezvous; MCP server as a thin adapter over the REST core.
  • MCP tools — bridge_send / bridge_reply / bridge_status (messaging) and bridge_spawn / bridge_list / bridge_stop / bridge_profiles / bridge_poll (fleet). Caller identity is connection-based (loopback peer PID → herdr pane), so the same mount serves primary and workers.
  • Delivery reliability — completion fallback (a confirmed working→idle turn resolves a send); async fire-and-poll (beats the caller's MCP call timeout for long tasks); and failure detection for wedged (unknown), vanished, and never-ready workers so a send never hangs.
  • Fleet — multiple worker profiles, each with an independent base_url guard check; workers inherit the primary's working directory (never $HOME); a readiness gate holds delivery until a worker's Claude has connected the bridge MCP (no paste lost into its boot window).
  • Blocked-worker path — bridge_ask reverse rendezvous: a worker pauses its delegated turn to ask the primary and resumes the same turn with the answer (CB-205).
  • Session lifecycle — session manager with spawn/reuse/recycle, idle_ttl reaper, context_cap, and graceful drain on shutdown (CB-301/CB-303); per-worker git worktrees on their own branch with a config-parity overlay, so parallel implementers never stomp each other (CB-301-ext).
  • Reliable worker→primary delivery — a durable ReplyInbox (in-memory by default, AMQP/LavinMQ for cross-restart durability) holds a reply that arrives with no open send, and an active status-gated push loop nudges the primary to drain it (CB-307).
  • Pluggable peers — a PeerLauncher SPI with two in-tree adapters, claude-code and opencode, routed by a kind: discriminator (CB-401/CB-402).

Next (see the roadmap) — Stage 5 hardening (auth/TLS, /metrics, CI, service supervision, per-session authz + audit), then cross-host: CB-308 multi-host federation and CB-500 multi-tier coordination.

S
Description
No description provided
Readme 16 MiB
2026-08-10 15:58:06 +02:00
Languages
Java 94%
Shell 5.1%
Python 0.9%