From 1cc34888fd373b994952c13216522cf424ebf5c6 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Sat, 15 Aug 2026 19:27:45 +0200 Subject: [PATCH] =?UTF-8?q?CB-591:=20record=20the=20live=20result=20?= =?UTF-8?q?=E2=80=94=20blocked=20by=20a=2032=20KiB=20body=20limit=20at=20t?= =?UTF-8?q?he=20gateway?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Deployed U1-U2c, restarted, spawned both new profiles for real, then reverted. llm.ltms.dev answers HTTP 413 above 32 KiB (32768 bytes), on BOTH surfaces: /v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200 /v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413 That is far below one agent turn. It is an edge limit (Caddy request_body max_size, and/or Envoy), so the fix is in systems/vms, not here. The part worth recording is how it nearly passed. Two members, same message, same moment: `local` finished in 66s, `gx` never finished at all. `local` passed only because the probe was three trivial questions in a fresh session, so the request fit under 32 KiB — the profile looked healthy and was a landmine set to fire on the first turn that reads a file. So §7's checklist was not wrong, it was too easy; it now says to use a file-reading task. opencode's failure mode is worse than a crash: it catches the 413, compacts its context, retries, and loops. Observed 10+ minutes BUSY with no reply. From the lead's side that is indistinguishable from a slow worker. Reproduced outside the bridge with the launcher's own generated config, which is how it became a one-line error instead of a hang; §7.1 records that procedure. Everything else about the migration checked out and is recorded so it is not re-tested: token accepted on both surfaces, unauthenticated 401 (the Caddy proxy does gate, whatever the gateway's own fail-open policy does), /v1/models exactly ["deepseek-v4-flash"], the guard allowlist accepted llm.ltms.dev, and the generated opencode provider block is correct with a real llmk- key. Also answers §3b's open question: reasoning survives BOTH surfaces — /anthropic returns a real "type":"thinking" block and /v1 returns a populated reasoning_content. The feared /v1 translation loss did not happen. Config state (bridged.yaml is gitignored, so it is described rather than committed): `local` back on http://gx00.gw:8000, `gx` kept at weight 0, `local-direct` kept, llm.ltms.dev left in the guard allowlist. The file carries these numbers and the exact two-key edit to switch back. Verified after the revert with a task that reads two large files: correct on all three questions. Daemon pid 66745, jar f1fd659423e6. Refs: gitea #76 --- docs/CB-591-Gateway-Migration.md | 98 +++++++++++++++++++++++++++++--- 1 file changed, 91 insertions(+), 7 deletions(-) diff --git a/docs/CB-591-Gateway-Migration.md b/docs/CB-591-Gateway-Migration.md index e82909c..6962686 100644 --- a/docs/CB-591-Gateway-Migration.md +++ b/docs/CB-591-Gateway-Migration.md @@ -1,6 +1,10 @@ # CB-591 — move the fleet onto the LLM and MCP gateway -**Status:** plan, not started · **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway) +**Status: BLOCKED — deployed 2026-08-15, verified live, then REVERTED.** The gateway rejects request +bodies over **32 KiB** with HTTP 413, on both surfaces, which is far below one agent turn. `local` is +back on the direct vLLM and `gx` is held at `weight: 0`. Everything else about the migration checked +out — see §7.1. The fix is an edge limit in **systems/vms**, not in this repo. +· **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway) · **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31) The gateway went live on 2026-08-15 and replaced Bifrost. This plan says what that means for a @@ -154,12 +158,21 @@ speaks the Anthropic protocol, so `/anthropic` is both correct and the only safe This is the exact failure shape this repo keeps hitting: it compiles, it answers, it looks healthy, and a capability is quietly off. Treat it as a `silent-default` risk, not a config preference. -**Open question for 3b, stated as open.** The opencode profile must use the OpenAI surface, because -that is the only thing an OpenAI-compatible provider can speak. vLLM serves that API natively, so I -*expect* it to be a passthrough as well — but the wiki documents the translation trap only for the -Anthropic path, and I have not checked how `reasoning_content` behaves through `/v1`. Verify it on -first spawn (§7) rather than assuming. If reasoning is dropped there, that is a limitation of the -opencode profile, not a reason to abandon it — opencode members do dev work, not deep reasoning. +**Open question for 3b — ANSWERED, 2026-08-15.** The worry was that the OpenAI surface might drop +reasoning the way the wiki documents for a mis-declared Anthropic backend. It does not. Checked at +the API before any profile was switched: + +| surface | request | result | +|---|---|---| +| `/anthropic/v1/messages` | `deepseek-v4-flash`, 64 tokens | 200, response carries a real `"type":"thinking"` block | +| `/v1/chat/completions` | same | 200, message carries a populated `reasoning_content` (and a `reasoning` field) | +| `/v1/models` | — | 200, exactly `["deepseek-v4-flash"]` — the exact-name trap is clear | +| `/v1/models`, **no token** | — | **401** — Caddy is gating, as designed | + +So reasoning survives on **both** surfaces, and the `/anthropic` choice for `local` is about protocol +correctness rather than a repair for a known loss. The last row matters on its own: the wiki warns +the gateway's own `SecurityPolicy` fails open, so it is worth knowing the proxy in front really does +refuse an unauthenticated request here. --- @@ -290,6 +303,77 @@ Merging config is not proving it. The checks, in order: --- +## 7.1 What the live run actually found — 2026-08-15 + +U1–U2c were done, the daemon restarted onto them, and both new profiles were spawned for real. The +migration was then **reverted**. This section is the result, so none of it has to be re-derived. + +### The blocker + +`llm.ltms.dev` answers **HTTP 413 Request Entity Too Large** above **32 KiB (32768 bytes)**, on both +surfaces: + +``` +/v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200 +/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413 +``` + +32 KiB is far below one real agent turn. This is an edge limit — Caddy `request_body max_size` and/or +Envoy's own — so it is fixed in **systems/vms**, not here. + +### The part worth remembering + +Two members were spawned at the same moment with the same message: + +| | `local` (claude-code, `/anthropic`) | `gx` (opencode, `/v1`) | +|---|---|---| +| READY → BUSY | 19:07:26 | 19:07:45 | +| BUSY → DONE | **19:08:51 (66s)** | **never — 10+ min, ticket FAILED** | + +**`local` passed.** It passed only because the probe was three trivial questions in a fresh session, +so the request fit under 32 KiB. The profile looked healthy and was a landmine set to fire on the +first turn that reads a file. + +So §7's checklist was not wrong, it was **too easy**. Any future run of it must use a task that reads +a real file. A liveness probe proves the token and the URL; it does not prove the path. + +`gx` did not fail loudly either. Reproduced outside the bridge by running `opencode` by hand with the +launcher's own generated config: + +``` +Error: Request Entity Too Large +...compacts context, retries... +Error: Request Entity Too Large +``` + +opencode **catches the 413, compacts, and retries — indefinitely**. A member that fails loudly costs +one turn; this one costs the whole task and is indistinguishable from a slow worker. + +> **Diagnosing a stuck opencode member.** Do not read its pane. The launcher writes its config to a +> temp dir and passes it as `OPENCODE_CONFIG` — find it with +> `ls -dt /var/folders/*/*/T/bridged-opencode-* | head -1`, check the provider block and the key's +> length and prefix (never its value), then reproduce with `opencode run --auto -m /` +> using the same `OPENCODE_CONFIG`. That is what turned "it hangs" into a one-line error. + +### What checked out, and needs no re-testing + +- Token accepted on both surfaces. **Unauthenticated → 401**, so the Caddy proxy really does gate — + the wiki's "SecurityPolicy fails open" warning is about the gateway itself, not the edge. +- `/v1/models` returns exactly `["deepseek-v4-flash"]`, so trap 3 is clear. +- **Reasoning survives both surfaces** — see §3b above. +- The launcher's generated opencode provider block is correct, carrying a real 48-character `llmk-` + key rather than the `bridged-local-noauth` placeholder. +- `SubscriptionGuard` accepted `llm.ltms.dev` after the allowlist edit and the restart: `local` + spawned without throwing, which is the check that catches a missed restart. + +### To resume + +Once the edge limit is raised: set `baseUrl` + `tokenEnv` on `local`, raise `gx` to `weight: 100`, +restart, and re-run §7 **with a file-reading task**. `bridged.yaml` carries the exact two-key edit +and these numbers inline. + +--- + ## 8. Related - [CB-589 / #74](https://git.ltms.dev/lms/claude-bridge/issues/74) — cost-first placement and a