diff --git a/docs/CB-591-Gateway-Migration.md b/docs/CB-591-Gateway-Migration.md index 6962686..b1cd90f 100644 --- a/docs/CB-591-Gateway-Migration.md +++ b/docs/CB-591-Gateway-Migration.md @@ -318,8 +318,28 @@ surfaces: /v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413 ``` -32 KiB is far below one real agent turn. This is an edge limit — Caddy `request_body max_size` and/or -Envoy's own — so it is fixed in **systems/vms**, not here. +32 KiB is far below one real agent turn. + +**Root cause — confirmed by the systems/vms side, 2026-08-15.** My guess that it was a Caddy +`request_body max_size` was **wrong**. It is Envoy, inside `aigw` on `llm.vm`. Envoy Gateway defaults +a listener's `per_connection_buffer_limit_bytes` to **32768**, and the AI Gateway buffers the *whole* +request body before it can route on the model name — so that default is not a network tuning knob +here, it is a hard ceiling on prompt size. Read out of the live Envoy `config_dump`: + +``` +listener default/llm/http per_connection_buffer_limit_bytes: 32768 +``` + +Nobody chose 32 KiB; it was inherited from the default. Both TLS edges are innocent: the same +boundary reproduces on the LAN path and the internet path, and both 413s carry an `x-llm-consumer` +header their auth proxy sets only *after* authenticating — so the body cleared both edges and the +auth. Directly on `llm.vm`, `aigw` 413s at 39 KB while the vLLM backend accepts the same 39 KB and +answers 200. + +**Do not plan around 32 KiB.** The intended ceiling is far higher. Their fix — a `ClientTrafficPolicy` +setting `bufferLimit: 8Mi` — is written but **not deployed** as of this note, pending their operator's +approval. I have not re-tested and will not until they confirm, so as not to measure a half-changed +system. Fixed in **systems/vms**, not here. ### The part worth remembering @@ -368,9 +388,14 @@ one turn; this one costs the whole task and is indistinguishable from a slow wor ### To resume -Once the edge limit is raised: set `baseUrl` + `tokenEnv` on `local`, raise `gx` to `weight: 100`, -restart, and re-run §7 **with a file-reading task**. `bridged.yaml` carries the exact two-key edit -and these numbers inline. +**Wait for systems/vms to confirm the fix is deployed and verified.** Do not re-test before that — +measuring a half-changed system produces a result nobody can trust. They have said they will verify by +re-reading the live `config_dump` and by sending a large request, not by the absence of an error, +because in standalone `aigw run` a `SecurityPolicy` is accepted and then silently ignored. If +`ClientTrafficPolicy` turns out to be ignored the same way, the fix will need a different shape. + +Then: set `baseUrl` + `tokenEnv` on `local`, raise `gx` to `weight: 100`, restart, and re-run §7 +**with a file-reading task**. `bridged.yaml` carries the exact two-key edit and these numbers inline. ---