ltms 7d711942fe
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m52s
Merge #534: detect a died shutdown drain the ERROR count is blind to (fleetd #512 part 2)
Verified independently. The branch is based on 8335b12 while main is at bec87f9, so I merged locally first and tested the MERGED tree, not the branch — a clean auto-merge is not a working merge.

Merged tree checks:
- `bash -n` exit 0 on both scripts, under /bin/bash 3.2.57 and env bash 5.3.9.
- Suite exit 0, 0 lines matching `^FAIL:`, 255 bytes of output. 60 test functions defined, 60 invoked, and no defined-but-never-invoked orphan (checked with a comm against the invocation list, not by comparing two counts — two equal counts can both be wrong).
- My own comment fix from bec87f9 survived the merge and is still at :317.

I ran two mutations the worker did not, per "mutate the half the worker did not":

(A) The one that matters, because it is the defect this ticket exists to prevent: collapsed the `unknown` state into `complete`, so "cannot tell" reports as a pass. Result exit 1, one FAIL: `cannot-tell fixture must set REDEPLOY_DRAIN_STATE=unknown: expected unknown, got complete`. So the third state is genuinely load-bearing, not decoration.

(B) Broke the positive check: changed `find_drain_complete_line`'s pattern from `drain complete: released=` to `drain finished: released=`, one site. Result exit 1, one FAIL: `find_drain_complete_line did not capture the present line`.

Proof that (B) applied, against a pristine copy: the full grep line 1 -> 0, the mutant form 0 -> 1, and the bare phrase 2 -> 1 with the comment occurrence untouched. Both files restored byte-identical; `git diff --quiet` clean; green control re-run.

A note on my own proof cell for (B), because it was wrong the first time. I wrote the counts with escaped double quotes inside an already double-quoted command substitution, so the shell split the pattern on spaces and grep treated the words as filenames. It printed "2 and 2" alongside `ugrep: No such file or directory` warnings — a symmetric, plausible-looking pair that meant nothing. The kill itself was never in doubt, since the suite named the exact function, but the cell that was supposed to prove the mutation applied proved nothing. Re-done with single quotes. This is the same trap already written down for this repo, hit by me, in a cell whose only purpose was to guard against exactly this.

One thing I checked that no test covers: the main flow's `HAD_OLD_PID=0; [ -n "$OLD_PID" ] && HAD_OLD_PID=1` runs under `set -euo pipefail`, and on a cold start the test fails. Sourcing stops before the main flow, so no behavioural test reaches that line. If `set -e` fired there, every cold start would abort before the health checks. It does not: `set -e` exempts the left side of an `&&` list, confirmed by running it under both shells — `survived, HAD_OLD_PID=0` on 3.2.57 and on 5.3.9. Safe, but it is untested main-flow wiring, which is the same class as #528's item 1 and belongs on that list.

On the `n/a` fourth state, which the worker flagged for a reviewer's judgment rather than quietly keeping: accepted, and it is in scope. The ticket asked for a third state because a sentinel conflating "no" with "cannot tell" hides two causes needing opposite handling. "The question does not apply" is a third such cause, not a variant of "cannot tell". Without it, the new warning would fire on every clean cold start, and a warning that cries wolf on the most common path trains the operator to skip it — which destroys the absence signal just as surely as putting the completion line in a `finally` would. The worker also proved the gate is actually consulted, using a cold-start fixture whose content deliberately looks like a died drain, so the test would fail if the gate were skipped. That is the right way to test a gate.

Both remaining outcomes are correctly excluded from the "no ERROR lines since restart" summary: only `complete` and `n/a` let it print.
2026-09-12 08:04:15 +02:00

claude-bridge

A subscription-safe bridge that lets a primary Claude Code (Opus 4.8, on Pro/Max) session drive a secondary Claude agent running a different model via its own ANTHROPIC_BASE_URL — without ever putting a proxy on the primary session.

Sibling of crush-bridge (which drives a headless Crush worker on GX10 DeepSeek). claude-bridge keeps the worker a real Claude Code process, so it inherits CLAUDE.md, hooks, skills, and MCP — just pointed at a cheaper/local model.

Leading approach — herdr-centric message server (fleetd)

A small always-on message server, fleetd, controls herdr (an agent multiplexer) over its Unix-socket API and exposes a clean 2-way messaging API as an MCP server that both the primary and the workers mount — one unified Claude setup and the sole communication gateway (REST/SSE stays for non-Claude clients; any broker is fleetd-internal, below the gateway). herdr owns the PTYs, multiplexing, persistence, and agent-status events; fleetd owns policy (subscription boundary, session lifecycle, status-gated delivery) and the client contract. A Claude member launches with ANTHROPIC_BASE_URL pointed at the gateway, https://llm.ltms.dev/anthropic, plus a bearer token; the lead stays env-clean and calls fleetd's MCP tools. See the wiki's 13 User Guide to run it.

flowchart LR
    OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
    subgraph BD["fleetd — standalone daemon (not a claude process)"]
        SRV["SERVER face<br/>MCP · REST/SSE · policy"]
        CLI["CLIENT face<br/>status-gated injector · herdr socket"]
        SRV --> CLI
    end
    HERDR["herdr<br/>panes · agent-status"]
    W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
    M["llm.ltms.dev<br/>(the one gateway)"]

    OPUS -->|"MCP fleet_send (blocks)"| SRV
    W -.->|"MCP fleet_reply"| SRV
    CLI -->|"Unix socket<br/>send_text · events.subscribe"| HERDR
    HERDR -->|"drives PTY"| W
    W -->|"inference"| M

    classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
    classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
    class OPUS ext
    class SRV,CLI,HERDR core
  • Subscription boundary: the primary never sets ANTHROPIC_BASE_URL (stays on Pro/Max). Only the secondary process is off-subscription — and fleetd itself is a plain daemon (no Anthropic quota), so it may poll/subscribe freely.
  • One gateway (unified MCP setup): fleetd is the sole communication path for every Claude session. Primary and workers each mount it as an MCP server (one claude mcp add line, same on both) and talk over MCP tools — fleet_send / fleet_reply / fleet_status (with fleet_ask planned for the blocked-worker path). No Claude session ever addresses a broker, a peer, or the network directly; any queue is fleetd-internal. MCP tool I/O never sets ANTHROPIC_BASE_URL, so mounting the bridge is subscription-safe by construction. Tool naming: the tools are fleet_* (renamed from bridge_* in CB-622). The old bridge_* names were removed in CB-634 — only fleet_* answers now.
  • How the primary consumes a reply: a single blocking MCP call (fleet_send); fleetd holds it open until the worker calls fleet_reply or its turn hits agent_status=done, then returns the reply as the tool result. No cross-turn busy-poll, so no quota burn. SSE is an optional side-channel for humans/dashboards watching status.
  • Worker → primary rides fleetd's MCP rendezvous — the reply resolves the primary's blocking call (or, for detached work, fleetd injects the primary's idle pane when it's ready), so no keystroke-into-primary and no broker are involved, even single-host. The one exception: a split-host primary that isn't a herdr pane wakes via its own Stop-hook, which polls fleetd (never a broker). See the wiki for the two topologies.
  • Different model per process sidesteps Claude Code's lack of per-subagent provider routing — the worker isn't a subagent, it's its own configured process.
  • AgentAPI (coder/agentapi) is retained only as a swappable fallback injector behind the same interface. See the wiki for the full design, comparison, and rationale.

Docs

Full design, setup, and operations live in the wiki, vendored here as a submodule under wiki/:

git clone --recurse-submodules ssh://git@git.ltms.dev:2224/fleet/fleetd.git
# or, after a plain clone:
git submodule update --init

Edit docs in wiki/, then cd wiki && git commit && git push to publish them to the Gitea wiki.

Status

🟢 Implemented & dogfooded — the herdr-centric fleetd message server is built and in real use: an Opus primary delegates tasks to off-subscription workers that reply through the bridge (code reviews delegated this way have produced committed bug fixes). Selected as the primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a fallback injector.

Shipped (Java 25 · Maven · 266 unit/acceptance tests green; the live-herdr and broker contract tests run separately via mvn test -Pcontract):

  • Core gateway — herdr socket client (contract-tested vs live 0.7.0); guard-checked worker spawn with ANTHROPIC_BASE_URL injected only into the worker's env; status-gated injector; blocking fleet_send with reply rendezvous; MCP server as a thin adapter over the REST core.
  • MCP tools — fleet_send / fleet_reply / fleet_status (messaging) and fleet_spawn / fleet_list / fleet_stop / fleet_profiles / fleet_poll (fleet). Caller identity is connection-based (loopback peer PID → herdr pane), so the same mount serves primary and workers.
  • Delivery reliability — completion fallback (a confirmed working→idle turn resolves a send); async fire-and-poll (beats the caller's MCP call timeout for long tasks); and failure detection for wedged (unknown), vanished, and never-ready workers so a send never hangs.
  • Fleet — multiple worker profiles, each with an independent base_url guard check; workers inherit the primary's working directory (never $HOME); a readiness gate holds delivery until a worker's Claude has connected the bridge MCP (no paste lost into its boot window).
  • Blocked-worker path — fleet_ask reverse rendezvous: a worker pauses its delegated turn to ask the primary and resumes the same turn with the answer (CB-205).
  • Session lifecycle — session manager with spawn/reuse/recycle, idle_ttl reaper, context_cap, and graceful drain on shutdown (CB-301/CB-303); per-worker git worktrees on their own branch with a config-parity overlay, so parallel implementers never stomp each other (CB-301-ext).
  • Reliable worker→primary delivery — a durable ReplyInbox (in-memory by default, AMQP/LavinMQ for cross-restart durability) holds a reply that arrives with no open send, and an active status-gated push loop nudges the primary to drain it (CB-307).
  • Pluggable peers — a PeerLauncher SPI with two in-tree adapters, claude-code and opencode, routed by a kind: discriminator (CB-401/CB-402).

Next (see the roadmap) — Stage 5 hardening (auth/TLS, /metrics, CI, service supervision, per-session authz + audit), then cross-host: CB-308 multi-host federation and CB-500 multi-tier coordination.

S
Description
No description provided
Readme 16 MiB
2026-08-10 15:58:06 +02:00
Languages
Java 94%
Shell 5.1%
Python 0.9%