diff --git a/docs/CB-591-Gateway-Migration.md b/docs/CB-591-Gateway-Migration.md new file mode 100644 index 0000000..e82909c --- /dev/null +++ b/docs/CB-591-Gateway-Migration.md @@ -0,0 +1,299 @@ +# CB-591 — move the fleet onto the LLM and MCP gateway + +**Status:** plan, not started · **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway) +· **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31) + +The gateway went live on 2026-08-15 and replaced Bifrost. This plan says what that means for a +**member definition** in `bridged.yaml`, because that is the part of this repo the change actually +touches. + +--- + +## 1. What changed upstream + +One front door for every LLM and MCP client: `https://llm.ltms.dev`, one token per consumer. + +| Surface | URL | +|---|---| +| OpenAI chat | `https://llm.ltms.dev/v1/chat/completions` | +| OpenAI models | `https://llm.ltms.dev/v1/models` | +| **Anthropic messages** | `https://llm.ltms.dev/anthropic/v1/messages` | +| MCP, all servers multiplexed | `https://llm.ltms.dev/mcp` | + +Anything outside that list returns **404 before any token is checked**, on purpose — the gateway must +never become a blanket proxy. + +The model backend is unchanged: GX10 vLLM at `10.10.10.26:8000` (`gx00.gw`), model name exactly +`deepseek-v4-flash`. The direct LAN path stays open on purpose as an escape hatch. + +--- + +## 2. Where claude-bridge sits today + +We do **not** use the gateway. The `local` profile talks straight to the vLLM: + +```yaml +local: + kind: claude-code + baseUrl: http://gx00.gw:8000 # direct vLLM — no auth, LAN only + model: deepseek-v4-flash + configDir: /Users/dai.ha/.ccs/instances/gx10 +``` + +Three facts about our side that decide the shape of this work: + +1. **`baseUrl` becomes `ANTHROPIC_BASE_URL`** in the member's environment, and `tokenEnv` becomes + `ANTHROPIC_AUTH_TOKEN` (the value is read from a host env var and never stored in config). + `local` sets no `tokenEnv` today, because a direct vLLM needs no token. +2. **`SubscriptionGuard` refuses any host not on an allowlist**, and that allowlist is + `guard.offSubscriptionHosts: [gx00.gw]`. It is built once in `Bridged.java:93` and handed to the + launcher, so **it is a restart-required key**, not a hot one. Changing `baseUrl` without changing + this makes every `local` spawn throw. +3. **The wiki names us as a blocker.** Under *Not done yet*: retiring the shared `legacy` token is + blocked because "kb, brain, **claude-bridge** and the workstation still share it. Each needs its + own consumer first." + +Context7 is mounted twice today, both times straight at `https://ct7.ltms.dev/mcp` — once in +`.mcp.json` (the primary) and once in `opencode.json` (the `sol` and `terra` members). + +```mermaid +flowchart LR + subgraph now["Today"] + M1["local member
claude-code"] -->|"ANTHROPIC_BASE_URL"| V1["vLLM gx00.gw:8000
no auth, LAN only"] + M2["sol / terra
opencode"] --> CT1["ct7.ltms.dev/mcp"] + P1["primary"] --> CT1 + end + subgraph after["Proposed"] + M3["local member"] -->|"ANTHROPIC_BASE_URL
+ ANTHROPIC_AUTH_TOKEN"| G["llm.ltms.dev/anthropic
consumer: claude-bridge"] + G --> V2["vLLM gx00.gw:8000"] + M4["local-direct
weight 0, escape hatch"] --> V2 + end +``` + +*The member definition is the only thing that moves. The model behind it does not.* + +--- + +## 3. The member definition change + +The gateway serves an Anthropic surface *and* an OpenAI surface, so **both member kinds can point at +it**. That is the main opportunity here, and it is bigger than the `local` profile alone. + +### 3a. `local` — claude-code, on `/anthropic` + +| Key | Today | After | Note | +|---|---|---|---| +| `baseUrl` | `http://gx00.gw:8000` | `https://llm.ltms.dev/anthropic` | see the schema warning below | +| `tokenEnv` | *(unset)* | `AI_GATEWAY_TOKEN` | new consumer token, `llmk-claude-bridge-<32 hex>` | +| `model` | `deepseek-v4-flash` | unchanged | must stay **exact**; a regex match returns an empty `/v1/models` while completions keep working | +| `guard.offSubscriptionHosts` | `[gx00.gw]` | `[gx00.gw, llm.ltms.dev]` | **restart required** | + +### 3b. A new opencode profile on `/v1` — no code needed + +`OpenCodeLauncher` already supports a pinned OpenAI-compatible endpoint (CB-508). Given `baseUrl` it +writes a custom provider block into the worker's opencode config: + +- `baseUrl` → `options.baseURL`. `openAiBaseUrl` appends `/v1` to a bare host, and takes a URL that + already has a path **as-is** — so `https://llm.ltms.dev/v1` works unchanged. +- `tokenEnv` → `options.apiKey` (falls back to a placeholder when unset, since a local vLLM ignores it). +- `model:` **must** be `/` when `baseUrl` is set — a bare name is rejected loudly + rather than silently falling back to opencode's default gateway. + +So the profile is pure config: + +```yaml +gx: + kind: opencode + baseUrl: https://llm.ltms.dev/v1 + tokenEnv: AI_GATEWAY_TOKEN + model: gx/deepseek-v4-flash # provider id is ours to choose; the half after / is the model + argv: ["opencode"] + mcpUrl: http://127.0.0.1:8765/mcp + gitTokenEnv: WORKER_GITEA_TOKEN + weight: 100 # same tier as `local` — free + maxLoad: 2 + # deliberately NO credentialId — this is our own box, not the shared OpenAI account +``` + +**Why this matters more than it looks.** Today every opencode member is `sol` or `terra`, and those +are two models on **one** OpenAI account sharing `credentialId: openai-shared` — so an exhaustion on +either locks out both, and half the fleet's opencode capacity dies at once. A gateway-backed opencode +profile is free, is not on that credential, and therefore is not in that quarantine pair. It removes +a single point of failure rather than just adding capacity. + +**Note the asymmetry, it is deliberate:** `SubscriptionGuard` does not apply to opencode at all — the +guard exists to stop a *Claude* worker borrowing the operator's subscription, and opencode reads its +own provider credentials. So 3b needs **no allowlist change**; only 3a does. + +**Both still need a restart, for a different reason.** `tokenEnv` is resolved by +`HerdrPeerLauncher.resolveEnv` → `env.apply(name)`, which reads the **daemon's own process +environment**. The running `bridged` inherited its environment when it started, so a variable added to +`secrets.sh` afterwards is simply not there — the launcher would inject an empty token and the +gateway would answer 401. This is the same failure as trap 1 in `scripts/redeploy-bridged.sh` +(`WORKER_GITEA_TOKEN`), and it has the same fix: **restart from a login shell**, and use +`scripts/redeploy-bridged.sh --check` to confirm the name resolves before restarting anything. + +### 3c. What this does to `ccs` + +Once a profile carries `baseUrl`, `tokenEnv` and `model` itself, the ccs instance stops being what +routes a member. Be precise about what is left, though: `configDir` still supplies **folder trust** +and `settings.json`, and dropping it is what produced the trust dialog and the wrong-model error +recorded in `bridged.yaml`. So ccs goes from *deciding where the tokens go* to *holding client-side +state*. Less load-bearing, not removable. + +### Why `/anthropic` and never `/v1/chat/completions` + +The gateway declares its Anthropic backend as `schema.name: Anthropic`, which means **no +translation** — streaming, tool use and thinking blocks pass through exactly as they do against vLLM +directly. + +Declared as `OpenAI`, Envoy's translator looks for a `thinking_blocks` field that our vLLM does not +send (it sends `reasoning_content`), and **every thinking delta disappears silently**. Claude Code +speaks the Anthropic protocol, so `/anthropic` is both correct and the only safe choice. + +This is the exact failure shape this repo keeps hitting: it compiles, it answers, it looks healthy, +and a capability is quietly off. Treat it as a `silent-default` risk, not a config preference. + +**Open question for 3b, stated as open.** The opencode profile must use the OpenAI surface, because +that is the only thing an OpenAI-compatible provider can speak. vLLM serves that API natively, so I +*expect* it to be a passthrough as well — but the wiki documents the translation trap only for the +Anthropic path, and I have not checked how `reasoning_content` behaves through `/v1`. Verify it on +first spawn (§7) rather than assuming. If reasoning is dropped there, that is a limitation of the +opencode profile, not a reason to abandon it — opencode members do dev work, not deep reasoning. + +--- + +## 4. Decisions + +### D1 — switch, but keep the direct path as an explicit profile · **recommended** + +Switching buys four things we do not have: + +- **Free opencode capacity, off the shared credential.** The largest single win. See §3b — it retires + a real single point of failure, not just a cost line. +- **Per-consumer usage figures.** The cockpit counts requests per consumer. That is the first real + measurement of what the fleet consumes, and it feeds [CB-589](https://git.ltms.dev/lms/claude-bridge/issues/74) Gap 2 directly. +- **Our own revocable token.** One consumer to revoke if a worker ever leaks it, instead of a shared + `legacy` token used by four systems. +- **It works off-LAN.** `gx00.gw` resolves on the LAN only. + +The cost is honest and worth stating: we add a TLS edge, an auth proxy and a gateway to the path of +every member spawn. The wiki keeps the direct route open precisely because "if the gateway breaks, +nothing that matters is blocked." + +So keep it. Add a second profile `local-direct` pointing at `http://gx00.gw:8000` with **`weight: 0`** +— never auto-selected, still spawnable with an explicit `bridge_spawn{profile: "local-direct"}`. +That is exactly what CB-554 made `weight: 0` mean, and it turns the escape hatch into something the +lead can actually reach during an incident. + +### D2 — do members also mount the gateway's `/mcp`? · **OPEN, operator's call** + +Not a detail. `CLAUDE.md` states in two places that a member mounts **only** the bridge MCP, and a +worker's honesty rule leans on it ("never claim the result of a check you had no way to run"). + +- **Keep bridge-only.** The invariant stays true and simple. Workers stay cheap and narrow. +- **Add the gateway MCP.** Implementers get context7 documentation lookups, which is genuinely useful + for library work. But `mcpUrl` in `BridgedConfig.Profile` is a **single `String`**, so a + claude-code member can mount exactly one MCP — this needs a code change, not a config edit. + +Note the invariant is **already inaccurate**: `opencode.json` gives `sol` and `terra` both context7 +and gitea. So the choice is really "make the rule true" or "make the rule match reality". Either is +defensible; picking one is not mine to do. + +### D3 — token scope + +One consumer, `claude-bridge`, its token in `${SHARED_ENV}/tools/secrets.sh` as `AI_GATEWAY_TOKEN`, +referenced by name only. Never the literal value in `bridged.yaml` — `tokenEnv` exists for this. + +--- + +## 5. Units of work + +```mermaid +flowchart TB + U1["U1 · consumer token
issue via cockpit, add to secrets.sh"] + U2["U2 · profile + guard
bridged.yaml, restart"] + U3["U3 · verify live
spawn, prove thinking survives"] + U4["U4 · context7 via gateway
.mcp.json + opencode.json"] + U5["U5 · docs
CLAUDE.md, wiki 11-Features"] + U1 --> U2 --> U3 + U4 --> U5 + U3 --> U5 +``` + +| # | Scope | Who | Why | +|---|---|---|---| +| U1 | Issue the `claude-bridge` consumer at `auth.ltms.dev`; store as `AI_GATEWAY_TOKEN` | **operator** | touches secrets and a host we do not own | +| U2a | New `gx` opencode profile on `/v1` — pure config, no guard change | **lead** | `bridged.yaml` is gitignored, so a worker cannot see or edit it | +| U2b | `local` → `/anthropic`; add `local-direct` weight 0; add `llm.ltms.dev` to the guard allowlist | **lead** | same | +| U2c | One restart from a **login shell**, after U2a and U2b | **lead** | picks up `AI_GATEWAY_TOKEN` into the daemon env *and* the guard allowlist, in one stop | +| U3 | Live spawn on both new profiles; confirm reasoning survives on each surface | **lead** | needs real spawns and the running daemon | +| U4 | Point `.mcp.json` and `opencode.json` context7 at the gateway `/mcp`; rename pinned tools | delegatable | tracked files, self-contained | +| U5 | Fix the "members mount only the bridge" claim; add a `wiki/11-Features.md` entry | delegatable | writing, clear criteria | + +U1 blocks U2a, U2b and U3. U4 and U5 do not depend on it. + +**Write U2a and U2b, then restart once (U2c), then verify `gx` before `local`.** Since both profiles +need the same restart there is no reason to do two, but there is still a reason to *verify* in order: +`gx` exercises the token and the gateway with no guard involved, so if it fails the cause is upstream. +`local` adds the guard allowlist on top, so a failure there points at our config instead. Testing them +in that order separates the two causes instead of confusing them. + +> **U1 status, 2026-08-15:** the operator issued the consumer and exported it as `AI_GATEWAY_TOKEN` +> (one key for every agent and MCP client behind `llm.ltms.dev`). Confirmed: it resolves in a login +> shell, is 48 characters and carries the documented `llmk-` prefix. The value was never printed. + +--- + +## 6. Traps carried over from the wiki + +Each of these cost someone real debugging time upstream. They apply to us. + +1. **Rotating a token restarts the auth proxy, which drops in-flight streaming responses.** For us + that means rotating `AI_GATEWAY_TOKEN` kills every live member mid-turn, and an async ticket's + report goes with it. This is the same rule as a daemon redeploy: **drain the fleet first** + (`bridge_list` → `bridge_poll` anything wanted → `bridge_stop`), then rotate. +2. **The gateway's own `SecurityPolicy` fails open.** Standalone `aigw run` accepts it and silently + ignores it — an unauthenticated request returned **200**. Auth is the Caddy proxy in front, and + nothing else. Never reason as if the gateway authenticates. +3. **Exact model name.** A regex match routes fine but returns an **empty** `/v1/models` list while + completions keep working. A wrong name returns a bare 404 that reads exactly like a dead gateway. +4. **MCP tool names changed prefix separator.** Bifrost used one dash (`ct7-resolve-library-id`); the + gateway uses **two underscores** (`ct7__resolve-library-id`). Relevant only if U4 is done. +5. **`/v1/models` 404 vs empty list are different faults.** 404 means no route loaded at all; empty + means the model match is a regex. Do not conflate them when diagnosing. + +--- + +## 7. Verification — what would prove this works + +Merging config is not proving it. The checks, in order: + +1. `bridge_spawn{profile: "gx"}` succeeds and the member completes a real turn ending in + `bridge_reply`. This is the first proof of the token, the URL and the model name, and it risks + nothing the fleet depends on. +2. `bridge_spawn{profile: "local"}` succeeds. If the guard allowlist was missed, this **throws** — a + loud, self-correcting failure, which is the good kind. If the restart was missed, it also throws, + for the same reason. +3. A `local` member completes a turn. That exercises streaming through two TLS edges, the auth proxy + and the gateway. +4. **Reasoning survives, checked separately on each surface.** For `local` on `/anthropic` this is + the check that catches the `/v1` versus `/anthropic` mistake, and it is the only one that does — + nothing else distinguishes a working passthrough from a translator quietly dropping thinking + deltas. For `gx` on `/v1`, this answers the open question in §3 rather than assuming it. +5. The cockpit at `auth.ltms.dev` shows requests counted against the `claude-bridge` consumer, not + `legacy`. That is the whole point of taking our own token. +6. `bridge_spawn{profile: "local-direct"}` still works, so the escape hatch is real rather than + theoretical. +7. `bridge_list` shows `gx` carrying no `credentialId`, so a `sol`/`terra` exhaustion cannot + quarantine it. This is the single-point-of-failure claim in §3b, checked rather than asserted. + +--- + +## 8. Related + +- [CB-589 / #74](https://git.ltms.dev/lms/claude-bridge/issues/74) — cost-first placement and a + gateway that reports live capacity. The per-consumer figures this migration unlocks are the first + input that ticket actually needs. +- `docs/CB-500-Multi-Tier-Coordination.md` §11 — the distributed-sandbox topology this gateway is + part of. diff --git a/scripts/redeploy-bridged.sh b/scripts/redeploy-bridged.sh index 69564d8..b23b519 100755 --- a/scripts/redeploy-bridged.sh +++ b/scripts/redeploy-bridged.sh @@ -8,8 +8,10 @@ # This script exists to turn five remembered traps into one auditable command: # # 1. A piped `mvn` hides BUILD FAILURE behind a zero exit, so the build here is never piped. -# 2. The daemon must start from a LOGIN shell, or WORKER_GITEA_TOKEN is empty and workers cannot -# open a PR. Nothing in the daemon logs this, so the script checks it and says so out loud. +# 2. The daemon must start from a LOGIN shell, or the tokens it hands to members are empty: +# WORKER_GITEA_TOKEN (workers cannot open a PR) and AI_GATEWAY_TOKEN (401 at llm.ltms.dev). +# Both are read from the DAEMON's own environment at spawn time, so a value added to +# secrets.sh after startup is absent. Nothing logs this, so the script checks and says so. # 3. An old daemon that never actually died looks identical from the outside, so the script waits # for the process to exit and for the port to free before it starts a new one. # 4. "It started" is not "it works": the script polls /healthz until it answers, and reports the @@ -82,6 +84,18 @@ else warn "Fix \${SHARED_ENV}/tools/secrets.sh before relying on worker checkpoints." fi +# Same trap, second variable (CB-591). A profile's `tokenEnv:` is resolved from the DAEMON's own +# process environment by HerdrPeerLauncher.resolveEnv, so a token added to secrets.sh after the +# daemon started is simply absent. The launcher then injects an empty token and llm.ltms.dev answers +# 401 — long after the restart, and with nothing tying the two together. +if zsh -lc '[ -n "${AI_GATEWAY_TOKEN:-}" ]' 2>/dev/null; then + ok "AI_GATEWAY_TOKEN resolves in a login shell" +else + warn "AI_GATEWAY_TOKEN is EMPTY in a login shell." + warn "Any profile whose tokenEnv is AI_GATEWAY_TOKEN will get an empty token and 401 at the gateway." + warn "This only matters once a profile points at llm.ltms.dev — harmless before that." +fi + if [ "$CHECK_ONLY" = 1 ]; then say "--check: nothing changed" exit 0