diff --git a/docs/CB-591-Gateway-Migration.md b/docs/CB-591-Gateway-Migration.md index e82909c..6962686 100644 --- a/docs/CB-591-Gateway-Migration.md +++ b/docs/CB-591-Gateway-Migration.md @@ -1,6 +1,10 @@ # CB-591 — move the fleet onto the LLM and MCP gateway -**Status:** plan, not started · **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway) +**Status: BLOCKED — deployed 2026-08-15, verified live, then REVERTED.** The gateway rejects request +bodies over **32 KiB** with HTTP 413, on both surfaces, which is far below one agent turn. `local` is +back on the direct vLLM and `gx` is held at `weight: 0`. Everything else about the migration checked +out — see §7.1. The fix is an edge limit in **systems/vms**, not in this repo. +· **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway) · **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31) The gateway went live on 2026-08-15 and replaced Bifrost. This plan says what that means for a @@ -154,12 +158,21 @@ speaks the Anthropic protocol, so `/anthropic` is both correct and the only safe This is the exact failure shape this repo keeps hitting: it compiles, it answers, it looks healthy, and a capability is quietly off. Treat it as a `silent-default` risk, not a config preference. -**Open question for 3b, stated as open.** The opencode profile must use the OpenAI surface, because -that is the only thing an OpenAI-compatible provider can speak. vLLM serves that API natively, so I -*expect* it to be a passthrough as well — but the wiki documents the translation trap only for the -Anthropic path, and I have not checked how `reasoning_content` behaves through `/v1`. Verify it on -first spawn (§7) rather than assuming. If reasoning is dropped there, that is a limitation of the -opencode profile, not a reason to abandon it — opencode members do dev work, not deep reasoning. +**Open question for 3b — ANSWERED, 2026-08-15.** The worry was that the OpenAI surface might drop +reasoning the way the wiki documents for a mis-declared Anthropic backend. It does not. Checked at +the API before any profile was switched: + +| surface | request | result | +|---|---|---| +| `/anthropic/v1/messages` | `deepseek-v4-flash`, 64 tokens | 200, response carries a real `"type":"thinking"` block | +| `/v1/chat/completions` | same | 200, message carries a populated `reasoning_content` (and a `reasoning` field) | +| `/v1/models` | — | 200, exactly `["deepseek-v4-flash"]` — the exact-name trap is clear | +| `/v1/models`, **no token** | — | **401** — Caddy is gating, as designed | + +So reasoning survives on **both** surfaces, and the `/anthropic` choice for `local` is about protocol +correctness rather than a repair for a known loss. The last row matters on its own: the wiki warns +the gateway's own `SecurityPolicy` fails open, so it is worth knowing the proxy in front really does +refuse an unauthenticated request here. --- @@ -290,6 +303,77 @@ Merging config is not proving it. The checks, in order: --- +## 7.1 What the live run actually found — 2026-08-15 + +U1–U2c were done, the daemon restarted onto them, and both new profiles were spawned for real. The +migration was then **reverted**. This section is the result, so none of it has to be re-derived. + +### The blocker + +`llm.ltms.dev` answers **HTTP 413 Request Entity Too Large** above **32 KiB (32768 bytes)**, on both +surfaces: + +``` +/v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200 +/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413 +``` + +32 KiB is far below one real agent turn. This is an edge limit — Caddy `request_body max_size` and/or +Envoy's own — so it is fixed in **systems/vms**, not here. + +### The part worth remembering + +Two members were spawned at the same moment with the same message: + +| | `local` (claude-code, `/anthropic`) | `gx` (opencode, `/v1`) | +|---|---|---| +| READY → BUSY | 19:07:26 | 19:07:45 | +| BUSY → DONE | **19:08:51 (66s)** | **never — 10+ min, ticket FAILED** | + +**`local` passed.** It passed only because the probe was three trivial questions in a fresh session, +so the request fit under 32 KiB. The profile looked healthy and was a landmine set to fire on the +first turn that reads a file. + +So §7's checklist was not wrong, it was **too easy**. Any future run of it must use a task that reads +a real file. A liveness probe proves the token and the URL; it does not prove the path. + +`gx` did not fail loudly either. Reproduced outside the bridge by running `opencode` by hand with the +launcher's own generated config: + +``` +Error: Request Entity Too Large +...compacts context, retries... +Error: Request Entity Too Large +``` + +opencode **catches the 413, compacts, and retries — indefinitely**. A member that fails loudly costs +one turn; this one costs the whole task and is indistinguishable from a slow worker. + +> **Diagnosing a stuck opencode member.** Do not read its pane. The launcher writes its config to a +> temp dir and passes it as `OPENCODE_CONFIG` — find it with +> `ls -dt /var/folders/*/*/T/bridged-opencode-* | head -1`, check the provider block and the key's +> length and prefix (never its value), then reproduce with `opencode run --auto -m /` +> using the same `OPENCODE_CONFIG`. That is what turned "it hangs" into a one-line error. + +### What checked out, and needs no re-testing + +- Token accepted on both surfaces. **Unauthenticated → 401**, so the Caddy proxy really does gate — + the wiki's "SecurityPolicy fails open" warning is about the gateway itself, not the edge. +- `/v1/models` returns exactly `["deepseek-v4-flash"]`, so trap 3 is clear. +- **Reasoning survives both surfaces** — see §3b above. +- The launcher's generated opencode provider block is correct, carrying a real 48-character `llmk-` + key rather than the `bridged-local-noauth` placeholder. +- `SubscriptionGuard` accepted `llm.ltms.dev` after the allowlist edit and the restart: `local` + spawned without throwing, which is the check that catches a missed restart. + +### To resume + +Once the edge limit is raised: set `baseUrl` + `tokenEnv` on `local`, raise `gx` to `weight: 100`, +restart, and re-run §7 **with a file-reading task**. `bridged.yaml` carries the exact two-key edit +and these numbers inline. + +--- + ## 8. Related - [CB-589 / #74](https://git.ltms.dev/lms/claude-bridge/issues/74) — cost-first placement and a