diff --git a/docs/CB-591-Gateway-Migration.md b/docs/CB-591-Gateway-Migration.md index b1cd90f..5025b67 100644 --- a/docs/CB-591-Gateway-Migration.md +++ b/docs/CB-591-Gateway-Migration.md @@ -1,9 +1,10 @@ # CB-591 — move the fleet onto the LLM and MCP gateway -**Status: BLOCKED — deployed 2026-08-15, verified live, then REVERTED.** The gateway rejects request -bodies over **32 KiB** with HTTP 413, on both surfaces, which is far below one agent turn. `local` is -back on the direct vLLM and `gx` is held at `weight: 0`. Everything else about the migration checked -out — see §7.1. The fix is an edge limit in **systems/vms**, not in this repo. +**Status: DONE — the fleet is on the gateway as of 2026-08-15.** `local` runs on `/anthropic` and +`gx` on `/v1`, both at `weight: 100`; `local-direct` stays at `weight: 0` as the escape hatch. Getting +here took a revert and two upstream fixes — see §7.1, which is the useful part of this document. One +risk is **accepted rather than solved**: a stream cut by any mid-response timer arrives as HTTP 200 +with no terminator, and our third-party members cannot detect it (§7.2). · **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway) · **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31) @@ -386,16 +387,62 @@ one turn; this one costs the whole task and is indistinguishable from a slow wor - `SubscriptionGuard` accepted `llm.ltms.dev` after the allowlist edit and the restart: `local` spawned without throwing, which is the check that catches a missed restart. -### To resume +### Resolution — both ceilings fixed, migration completed -**Wait for systems/vms to confirm the fix is deployed and verified.** Do not re-test before that — -measuring a half-changed system produces a result nobody can trust. They have said they will verify by -re-reading the live `config_dump` and by sending a large request, not by the absence of an error, -because in standalone `aigw run` a `SecurityPolicy` is accepted and then silently ignored. If -`ClientTrafficPolicy` turns out to be ignored the same way, the fix will need a different shape. +systems/vms fixed both, and each was re-checked from this side rather than taken on trust: -Then: set `baseUrl` + `tokenEnv` on `local`, raise `gx` to `weight: 100`, restart, and re-run §7 -**with a file-reading task**. `bridged.yaml` carries the exact two-key edit and these numbers inline. +| ceiling | was | now | our own check | +|---|---|---|---| +| listener buffer | 32 KiB | 32 Mi | 1.2 MB body → **200** (was 413) | +| LLM route timeout | 60s | 1800s | the request that truncated: **101s, `message_stop` present, 4000/4000** | + +Neither was deliberate. The 32 KiB was Envoy Gateway's default `per_connection_buffer_limit_bytes`; +the 60s was Envoy AI Gateway's own documented default. The 60s bounded **generation** as well as +prompt size — a tiny prompt with a long answer returned 504 at 60.05s. + +Two configuration facts worth keeping, from their bisection: + +- **`ClientTrafficPolicy` is honoured in standalone `aigw run`; `BackendTrafficPolicy` is NOT.** A + `BackendTrafficPolicy` setting `requestTimeout` is accepted, logs nothing, and leaves the routes + unchanged (upstream `envoyproxy/gateway#9513`). What works is `timeouts: {request: …}` on each + `AIGatewayRoute` rule. Nothing from the outside distinguishes the two — the same silent-default + shape as their `SecurityPolicy` caveat. +- In that stack, "the config was accepted" proves nothing. Read the live `config_dump`. + +## 7.2 The risk we accepted, and why we could not remove it + +Raising the timeout made the failure **rare, not impossible**, and the residual failure is silent. + +On a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked encoding *cleanly* instead of +resetting the connection, so the client receives what looks like a complete transfer +(`envoyproxy/envoy#17186` — acknowledged as a bug in 2021, closed by a stale bot, never fixed). The +December 2025 fix `envoyproxy/envoy#42269` changes locally-originated resets from `NO_ERROR` to +`INTERNAL_ERROR`, but it is **HTTP/2 only** and SSE clients here speak HTTP/1.1. + +Measured on our side while the timeout was still 60s: + +``` +HTTP 200 61.07s 141992 bytes + message_stop 0 message_delta 0 error events 0 + emitted 2473 of 4000, ending on a WELL-FORMED SSE frame +``` + +A syntactically valid stream that simply stops. Any timer firing mid-stream — route timeout, idle +timeout, `max_stream_duration` — fails this same way. + +**The recommended defence does not transfer to us.** The right fix is to treat a stream with no +`message_stop` / `[DONE]` / `finish_reason` as failed. We cannot: our members are Claude Code and +opencode, third-party clients whose SSE parsing we do not own, and there is no seam to insert the +check. Whether either detects a missing terminator is unverified — and opencode's handling of the 413 +(swallow, compact, retry forever, never surface an error) does not suggest it is strict. + +So the honest statement of our position: + +> Gateway traffic is acceptable at 1800s because a single request would have to run for 30 minutes to +> trip the bug — **not** because we could detect it if it did. + +**If a member ever returns a confident but truncated answer, suspect this before anything in our own +code.** That is the whole reason this section exists. ---