# CB-591 — move the fleet onto the LLM and MCP gateway **Status: DONE — the fleet is on the gateway as of 2026-08-15.** `local` runs on `/anthropic` and `gx` on `/v1`, both at `weight: 100`; `local-direct` stays at `weight: 0` as the escape hatch. Getting here took a revert and two upstream fixes — see §7.1, which is the useful part of this document. One risk is **accepted rather than solved**: a stream cut by any mid-response timer arrives as HTTP 200 with no terminator, and our third-party members cannot detect it (§7.2). · **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway) · **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31) The gateway went live on 2026-08-15 and replaced Bifrost. This plan says what that means for a **member definition** in `bridged.yaml`, because that is the part of this repo the change actually touches. --- ## 1. What changed upstream One front door for every LLM and MCP client: `https://llm.ltms.dev`, one token per consumer. | Surface | URL | |---|---| | OpenAI chat | `https://llm.ltms.dev/v1/chat/completions` | | OpenAI models | `https://llm.ltms.dev/v1/models` | | **Anthropic messages** | `https://llm.ltms.dev/anthropic/v1/messages` | | MCP, all servers multiplexed | `https://llm.ltms.dev/mcp` | Anything outside that list returns **404 before any token is checked**, on purpose — the gateway must never become a blanket proxy. The model backend is unchanged: GX10 vLLM at `10.10.10.26:8000` (`gx00.gw`), model name exactly `deepseek-v4-flash`. The direct LAN path stays open on purpose as an escape hatch. --- ## 2. Where claude-bridge sits today We do **not** use the gateway. The `local` profile talks straight to the vLLM: ```yaml local: kind: claude-code baseUrl: http://gx00.gw:8000 # direct vLLM — no auth, LAN only model: deepseek-v4-flash configDir: /Users/dai.ha/.ccs/instances/gx10 ``` Three facts about our side that decide the shape of this work: 1. **`baseUrl` becomes `ANTHROPIC_BASE_URL`** in the member's environment, and `tokenEnv` becomes `ANTHROPIC_AUTH_TOKEN` (the value is read from a host env var and never stored in config). `local` sets no `tokenEnv` today, because a direct vLLM needs no token. 2. **`SubscriptionGuard` refuses any host not on an allowlist**, and that allowlist is `guard.offSubscriptionHosts: [gx00.gw]`. It is built once in `Bridged.java:93` and handed to the launcher, so **it is a restart-required key**, not a hot one. Changing `baseUrl` without changing this makes every `local` spawn throw. 3. **The wiki names us as a blocker.** Under *Not done yet*: retiring the shared `legacy` token is blocked because "kb, brain, **claude-bridge** and the workstation still share it. Each needs its own consumer first." Context7 is mounted twice today, both times straight at `https://ct7.ltms.dev/mcp` — once in `.mcp.json` (the primary) and once in `opencode.json` (the `sol` and `terra` members). ```mermaid flowchart LR subgraph now["Today"] M1["local member
claude-code"] -->|"ANTHROPIC_BASE_URL"| V1["vLLM gx00.gw:8000
no auth, LAN only"] M2["sol / terra
opencode"] --> CT1["ct7.ltms.dev/mcp"] P1["primary"] --> CT1 end subgraph after["Proposed"] M3["local member"] -->|"ANTHROPIC_BASE_URL
+ ANTHROPIC_AUTH_TOKEN"| G["llm.ltms.dev/anthropic
consumer: claude-bridge"] G --> V2["vLLM gx00.gw:8000"] M4["local-direct
weight 0, escape hatch"] --> V2 end ``` *The member definition is the only thing that moves. The model behind it does not.* --- ## 3. The member definition change The gateway serves an Anthropic surface *and* an OpenAI surface, so **both member kinds can point at it**. That is the main opportunity here, and it is bigger than the `local` profile alone. ### 3a. `local` — claude-code, on `/anthropic` | Key | Today | After | Note | |---|---|---|---| | `baseUrl` | `http://gx00.gw:8000` | `https://llm.ltms.dev/anthropic` | see the schema warning below | | `tokenEnv` | *(unset)* | `AI_GATEWAY_TOKEN` | new consumer token, `llmk-claude-bridge-<32 hex>` | | `model` | `deepseek-v4-flash` | unchanged | must stay **exact**; a regex match returns an empty `/v1/models` while completions keep working | | `guard.offSubscriptionHosts` | `[gx00.gw]` | `[gx00.gw, llm.ltms.dev]` | **restart required** | ### 3b. A new opencode profile on `/v1` — no code needed `OpenCodeLauncher` already supports a pinned OpenAI-compatible endpoint (CB-508). Given `baseUrl` it writes a custom provider block into the worker's opencode config: - `baseUrl` → `options.baseURL`. `openAiBaseUrl` appends `/v1` to a bare host, and takes a URL that already has a path **as-is** — so `https://llm.ltms.dev/v1` works unchanged. - `tokenEnv` → `options.apiKey` (falls back to a placeholder when unset, since a local vLLM ignores it). - `model:` **must** be `/` when `baseUrl` is set — a bare name is rejected loudly rather than silently falling back to opencode's default gateway. So the profile is pure config: ```yaml gx: kind: opencode baseUrl: https://llm.ltms.dev/v1 tokenEnv: AI_GATEWAY_TOKEN model: gx/deepseek-v4-flash # provider id is ours to choose; the half after / is the model argv: ["opencode"] mcpUrl: http://127.0.0.1:8765/mcp gitTokenEnv: WORKER_GITEA_TOKEN weight: 100 # same tier as `local` — free maxLoad: 2 # deliberately NO credentialId — this is our own box, not the shared OpenAI account ``` **Why this matters more than it looks.** Today every opencode member is `sol` or `terra`, and those are two models on **one** OpenAI account sharing `credentialId: openai-shared` — so an exhaustion on either locks out both, and half the fleet's opencode capacity dies at once. A gateway-backed opencode profile is free, is not on that credential, and therefore is not in that quarantine pair. It removes a single point of failure rather than just adding capacity. **Note the asymmetry, it is deliberate:** `SubscriptionGuard` does not apply to opencode at all — the guard exists to stop a *Claude* worker borrowing the operator's subscription, and opencode reads its own provider credentials. So 3b needs **no allowlist change**; only 3a does. **Both still need a restart, for a different reason.** `tokenEnv` is resolved by `HerdrPeerLauncher.resolveEnv` → `env.apply(name)`, which reads the **daemon's own process environment**. The running `bridged` inherited its environment when it started, so a variable added to `secrets.sh` afterwards is simply not there — the launcher would inject an empty token and the gateway would answer 401. This is the same failure as trap 1 in `scripts/redeploy-bridged.sh` (`WORKER_GITEA_TOKEN`), and it has the same fix: **restart from a login shell**, and use `scripts/redeploy-bridged.sh --check` to confirm the name resolves before restarting anything. ### 3c. What this does to `ccs` Once a profile carries `baseUrl`, `tokenEnv` and `model` itself, the ccs instance stops being what routes a member. Be precise about what is left, though: `configDir` still supplies **folder trust** and `settings.json`, and dropping it is what produced the trust dialog and the wrong-model error recorded in `bridged.yaml`. So ccs goes from *deciding where the tokens go* to *holding client-side state*. Less load-bearing, not removable. ### Why `/anthropic` and never `/v1/chat/completions` The gateway declares its Anthropic backend as `schema.name: Anthropic`, which means **no translation** — streaming, tool use and thinking blocks pass through exactly as they do against vLLM directly. Declared as `OpenAI`, Envoy's translator looks for a `thinking_blocks` field that our vLLM does not send (it sends `reasoning_content`), and **every thinking delta disappears silently**. Claude Code speaks the Anthropic protocol, so `/anthropic` is both correct and the only safe choice. This is the exact failure shape this repo keeps hitting: it compiles, it answers, it looks healthy, and a capability is quietly off. Treat it as a `silent-default` risk, not a config preference. **Open question for 3b — ANSWERED, 2026-08-15.** The worry was that the OpenAI surface might drop reasoning the way the wiki documents for a mis-declared Anthropic backend. It does not. Checked at the API before any profile was switched: | surface | request | result | |---|---|---| | `/anthropic/v1/messages` | `deepseek-v4-flash`, 64 tokens | 200, response carries a real `"type":"thinking"` block | | `/v1/chat/completions` | same | 200, message carries a populated `reasoning_content` (and a `reasoning` field) | | `/v1/models` | — | 200, exactly `["deepseek-v4-flash"]` — the exact-name trap is clear | | `/v1/models`, **no token** | — | **401** — Caddy is gating, as designed | So reasoning survives on **both** surfaces, and the `/anthropic` choice for `local` is about protocol correctness rather than a repair for a known loss. The last row matters on its own: the wiki warns the gateway's own `SecurityPolicy` fails open, so it is worth knowing the proxy in front really does refuse an unauthenticated request here. --- ## 4. Decisions ### D1 — switch, but keep the direct path as an explicit profile · **recommended** Switching buys four things we do not have: - **Free opencode capacity, off the shared credential.** The largest single win. See §3b — it retires a real single point of failure, not just a cost line. - **Per-consumer usage figures.** The cockpit counts requests per consumer. That is the first real measurement of what the fleet consumes, and it feeds [CB-589](https://git.ltms.dev/fleet/fleetd/issues/74) Gap 2 directly. - **Our own revocable token.** One consumer to revoke if a worker ever leaks it, instead of a shared `legacy` token used by four systems. - **It works off-LAN.** `gx00.gw` resolves on the LAN only. The cost is honest and worth stating: we add a TLS edge, an auth proxy and a gateway to the path of every member spawn. The wiki keeps the direct route open precisely because "if the gateway breaks, nothing that matters is blocked." So keep it. Add a second profile `local-direct` pointing at `http://gx00.gw:8000` with **`weight: 0`** — never auto-selected, still spawnable with an explicit `fleet_spawn{profile: "local-direct"}`. That is exactly what CB-554 made `weight: 0` mean, and it turns the escape hatch into something the lead can actually reach during an incident. ### D2 — do members also mount the gateway's `/mcp`? · **OPEN, operator's call** Not a detail. `CLAUDE.md` states in two places that a member mounts **only** the bridge MCP, and a worker's honesty rule leans on it ("never claim the result of a check you had no way to run"). - **Keep bridge-only.** The invariant stays true and simple. Workers stay cheap and narrow. - **Add the gateway MCP.** Implementers get context7 documentation lookups, which is genuinely useful for library work. But `mcpUrl` in `BridgedConfig.Profile` is a **single `String`**, so a claude-code member can mount exactly one MCP — this needs a code change, not a config edit. Note the invariant is **already inaccurate**: `opencode.json` gives `sol` and `terra` both context7 and gitea. So the choice is really "make the rule true" or "make the rule match reality". Either is defensible; picking one is not mine to do. ### D3 — token scope One consumer, `claude-bridge`, its token in `${SHARED_ENV}/tools/secrets.sh` as `AI_GATEWAY_TOKEN`, referenced by name only. Never the literal value in `bridged.yaml` — `tokenEnv` exists for this. --- ## 5. Units of work ```mermaid flowchart TB U1["U1 · consumer token
issue via cockpit, add to secrets.sh"] U2["U2 · profile + guard
bridged.yaml, restart"] U3["U3 · verify live
spawn, prove thinking survives"] U4["U4 · context7 via gateway
.mcp.json + opencode.json"] U5["U5 · docs
CLAUDE.md, wiki 11-Features"] U1 --> U2 --> U3 U4 --> U5 U3 --> U5 ``` | # | Scope | Who | Why | |---|---|---|---| | U1 | Issue the `claude-bridge` consumer at `auth.ltms.dev`; store as `AI_GATEWAY_TOKEN` | **operator** | touches secrets and a host we do not own | | U2a | New `gx` opencode profile on `/v1` — pure config, no guard change | **lead** | `bridged.yaml` is gitignored, so a worker cannot see or edit it | | U2b | `local` → `/anthropic`; add `local-direct` weight 0; add `llm.ltms.dev` to the guard allowlist | **lead** | same | | U2c | One restart from a **login shell**, after U2a and U2b | **lead** | picks up `AI_GATEWAY_TOKEN` into the daemon env *and* the guard allowlist, in one stop | | U3 | Live spawn on both new profiles; confirm reasoning survives on each surface | **lead** | needs real spawns and the running daemon | | U4 | Point `.mcp.json` and `opencode.json` context7 at the gateway `/mcp`; rename pinned tools | delegatable | tracked files, self-contained | | U5 | Fix the "members mount only the bridge" claim; add a `wiki/11-Features.md` entry | delegatable | writing, clear criteria | U1 blocks U2a, U2b and U3. U4 and U5 do not depend on it. **Write U2a and U2b, then restart once (U2c), then verify `gx` before `local`.** Since both profiles need the same restart there is no reason to do two, but there is still a reason to *verify* in order: `gx` exercises the token and the gateway with no guard involved, so if it fails the cause is upstream. `local` adds the guard allowlist on top, so a failure there points at our config instead. Testing them in that order separates the two causes instead of confusing them. > **U1 status, 2026-08-15:** the operator issued the consumer and exported it as `AI_GATEWAY_TOKEN` > (one key for every agent and MCP client behind `llm.ltms.dev`). Confirmed: it resolves in a login > shell, is 48 characters and carries the documented `llmk-` prefix. The value was never printed. --- ## 6. Traps carried over from the wiki Each of these cost someone real debugging time upstream. They apply to us. 1. **Rotating a token restarts the auth proxy, which drops in-flight streaming responses.** For us that means rotating `AI_GATEWAY_TOKEN` kills every live member mid-turn, and an async ticket's report goes with it. This is the same rule as a daemon redeploy: **drain the fleet first** (`fleet_list` → `fleet_poll` anything wanted → `fleet_stop`), then rotate. 2. **The gateway's own `SecurityPolicy` fails open.** Standalone `aigw run` accepts it and silently ignores it — an unauthenticated request returned **200**. Auth is the Caddy proxy in front, and nothing else. Never reason as if the gateway authenticates. 3. **Exact model name.** A regex match routes fine but returns an **empty** `/v1/models` list while completions keep working. A wrong name returns a bare 404 that reads exactly like a dead gateway. 4. **MCP tool names changed prefix separator.** Bifrost used one dash (`ct7-resolve-library-id`); the gateway uses **two underscores** (`ct7__resolve-library-id`). Relevant only if U4 is done. 5. **`/v1/models` 404 vs empty list are different faults.** 404 means no route loaded at all; empty means the model match is a regex. Do not conflate them when diagnosing. --- ## 7. Verification — what would prove this works Merging config is not proving it. The checks, in order: 1. `fleet_spawn{profile: "gx"}` succeeds and the member completes a real turn ending in `fleet_reply`. This is the first proof of the token, the URL and the model name, and it risks nothing the fleet depends on. 2. `fleet_spawn{profile: "local"}` succeeds. If the guard allowlist was missed, this **throws** — a loud, self-correcting failure, which is the good kind. If the restart was missed, it also throws, for the same reason. 3. A `local` member completes a turn. That exercises streaming through two TLS edges, the auth proxy and the gateway. 4. **Reasoning survives, checked separately on each surface.** For `local` on `/anthropic` this is the check that catches the `/v1` versus `/anthropic` mistake, and it is the only one that does — nothing else distinguishes a working passthrough from a translator quietly dropping thinking deltas. For `gx` on `/v1`, this answers the open question in §3 rather than assuming it. 5. The cockpit at `auth.ltms.dev` shows requests counted against the `claude-bridge` consumer, not `legacy`. That is the whole point of taking our own token. 6. `fleet_spawn{profile: "local-direct"}` still works, so the escape hatch is real rather than theoretical. 7. `fleet_list` shows `gx` carrying no `credentialId`, so a `sol`/`terra` exhaustion cannot quarantine it. This is the single-point-of-failure claim in §3b, checked rather than asserted. --- ## 7.1 What the live run actually found — 2026-08-15 U1–U2c were done, the daemon restarted onto them, and both new profiles were spawned for real. The migration was then **reverted**. This section is the result, so none of it has to be re-derived. ### The blocker `llm.ltms.dev` answers **HTTP 413 Request Entity Too Large** above **32 KiB (32768 bytes)**, on both surfaces: ``` /v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200 /v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413 ``` 32 KiB is far below one real agent turn. **Root cause — confirmed by the systems/vms side, 2026-08-15.** My guess that it was a Caddy `request_body max_size` was **wrong**. It is Envoy, inside `aigw` on `llm.vm`. Envoy Gateway defaults a listener's `per_connection_buffer_limit_bytes` to **32768**, and the AI Gateway buffers the *whole* request body before it can route on the model name — so that default is not a network tuning knob here, it is a hard ceiling on prompt size. Read out of the live Envoy `config_dump`: ``` listener default/llm/http per_connection_buffer_limit_bytes: 32768 ``` Nobody chose 32 KiB; it was inherited from the default. Both TLS edges are innocent: the same boundary reproduces on the LAN path and the internet path, and both 413s carry an `x-llm-consumer` header their auth proxy sets only *after* authenticating — so the body cleared both edges and the auth. Directly on `llm.vm`, `aigw` 413s at 39 KB while the vLLM backend accepts the same 39 KB and answers 200. **Do not plan around 32 KiB.** The intended ceiling is far higher. Their fix — a `ClientTrafficPolicy` setting `bufferLimit: 8Mi` — is written but **not deployed** as of this note, pending their operator's approval. I have not re-tested and will not until they confirm, so as not to measure a half-changed system. Fixed in **systems/vms**, not here. ### The part worth remembering Two members were spawned at the same moment with the same message: | | `local` (claude-code, `/anthropic`) | `gx` (opencode, `/v1`) | |---|---|---| | READY → BUSY | 19:07:26 | 19:07:45 | | BUSY → DONE | **19:08:51 (66s)** | **never — 10+ min, ticket FAILED** | **`local` passed.** It passed only because the probe was three trivial questions in a fresh session, so the request fit under 32 KiB. The profile looked healthy and was a landmine set to fire on the first turn that reads a file. So §7's checklist was not wrong, it was **too easy**. Any future run of it must use a task that reads a real file. A liveness probe proves the token and the URL; it does not prove the path. `gx` did not fail loudly either. Reproduced outside the bridge by running `opencode` by hand with the launcher's own generated config: ``` Error: Request Entity Too Large ...compacts context, retries... Error: Request Entity Too Large ``` opencode **catches the 413, compacts, and retries — indefinitely**. A member that fails loudly costs one turn; this one costs the whole task and is indistinguishable from a slow worker. > **Diagnosing a stuck opencode member.** Do not read its pane. The launcher writes its config to a > temp dir and passes it as `OPENCODE_CONFIG` — find it with > `ls -dt /var/folders/*/*/T/bridged-opencode-* | head -1`, check the provider block and the key's > length and prefix (never its value), then reproduce with `opencode run --auto -m /` > using the same `OPENCODE_CONFIG`. That is what turned "it hangs" into a one-line error. ### What checked out, and needs no re-testing - Token accepted on both surfaces. **Unauthenticated → 401**, so the Caddy proxy really does gate — the wiki's "SecurityPolicy fails open" warning is about the gateway itself, not the edge. - `/v1/models` returns exactly `["deepseek-v4-flash"]`, so trap 3 is clear. - **Reasoning survives both surfaces** — see §3b above. - The launcher's generated opencode provider block is correct, carrying a real 48-character `llmk-` key rather than the `bridged-local-noauth` placeholder. - `SubscriptionGuard` accepted `llm.ltms.dev` after the allowlist edit and the restart: `local` spawned without throwing, which is the check that catches a missed restart. ### Resolution — both ceilings fixed, migration completed systems/vms fixed both, and each was re-checked from this side rather than taken on trust: | ceiling | was | now | our own check | |---|---|---|---| | listener buffer | 32 KiB | 32 Mi | 1.2 MB body → **200** (was 413) | | LLM route timeout | 60s | 86400s | the request that truncated: **101s, `message_stop` present, 4000/4000** | The timeout moved in two steps on 2026-08-15: 60s → 1800s, then 1800s → **86400s (24 hours)** after the truncation risk below was discussed. They tried `request: 0s` first, which removes the total-duration timer completely. It works, but on an `AIGatewayRoute` the **idle timeout is derived from the request timeout**, so `0s` also removed any bound on a stalled connection. 86400s keeps a reaper for dead connections while putting the truncation timer out of practical reach. Neither was deliberate. The 32 KiB was Envoy Gateway's default `per_connection_buffer_limit_bytes`; the 60s was Envoy AI Gateway's own documented default. The 60s bounded **generation** as well as prompt size — a tiny prompt with a long answer returned 504 at 60.05s. Two configuration facts worth keeping, from their bisection: - **`ClientTrafficPolicy` is honoured in standalone `aigw run`; `BackendTrafficPolicy` is NOT.** A `BackendTrafficPolicy` setting `requestTimeout` is accepted, logs nothing, and leaves the routes unchanged (upstream `envoyproxy/gateway#9513`). What works is `timeouts: {request: …}` on each `AIGatewayRoute` rule. Nothing from the outside distinguishes the two — the same silent-default shape as their `SecurityPolicy` caveat. - In that stack, "the config was accepted" proves nothing. Read the live `config_dump`. ## 7.2 The risk we accepted, and why we could not remove it Raising the timeout made the failure **rare, not impossible**, and the residual failure is silent. On a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked encoding *cleanly* instead of resetting the connection, so the client receives what looks like a complete transfer (`envoyproxy/envoy#17186` — acknowledged as a bug in 2021, closed by a stale bot, never fixed). The December 2025 fix `envoyproxy/envoy#42269` changes locally-originated resets from `NO_ERROR` to `INTERNAL_ERROR`, but it is **HTTP/2 only** and SSE clients here speak HTTP/1.1. Measured on our side while the timeout was still 60s: ``` HTTP 200 61.07s 141992 bytes message_stop 0 message_delta 0 error events 0 emitted 2473 of 4000, ending on a WELL-FORMED SSE frame ``` A syntactically valid stream that simply stops. Any timer firing mid-stream — route timeout, idle timeout, `max_stream_duration` — fails this same way. **The recommended defence does not transfer to us.** The right fix is to treat a stream with no `message_stop` / `[DONE]` / `finish_reason` as failed. We cannot: our members are Claude Code and opencode, third-party clients whose SSE parsing we do not own, and there is no seam to insert the check. Whether either detects a missing terminator is unverified — and opencode's handling of the 413 (swallow, compact, retry forever, never surface an error) does not suggest it is strict. So the honest statement of our position: > Gateway traffic is acceptable at 86400s because a single request would have to run for 24 hours to > trip the bug — **not** because we could detect it if it did. At 86400s our **own** limit binds first, which is the ordering we want. `MessageService.ASYNC_TIMEOUT_MS` caps a turn at 30 minutes, so a runaway request ends as a clean `FAILED` ticket that we raised, rather than as a silently truncated `200` that we cannot see. While the gateway sat at 1800s the two numbers were equal and did not nest, so a gateway-side stall could have been misread as a bug in our own ticket handling. That ambiguity is now gone. **If a member ever returns a confident but truncated answer, suspect this before anything in our own code.** That is the whole reason this section exists. --- ## 8. Related - [CB-589 / #74](https://git.ltms.dev/fleet/fleetd/issues/74) — cost-first placement and a gateway that reports live capacity. The per-consumer figures this migration unlocks are the first input that ticket actually needs. - `docs/CB-500-Multi-Tier-Coordination.md` §11 — the distributed-sandbox topology this gateway is part of.