CB-591: correct the root cause — Envoy's buffer limit, not Caddy
I wrote "Caddy request_body max_size and/or Envoy's own" and marked it
unverified. The Caddy half was wrong, and an unverified guess still points the
next reader at the wrong component.
Confirmed by the systems/vms side: Envoy Gateway defaults a listener's
per_connection_buffer_limit_bytes to 32768, and aigw buffers the WHOLE request
body before it can route on the model name. So that default is not a network
tuning knob — it is a hard ceiling on prompt size. From the live config_dump:
listener default/llm/http per_connection_buffer_limit_bytes: 32768
Nobody chose 32 KiB; it was inherited.
Both TLS edges are innocent, and the technique that showed it is better than
mine: both 413s carry x-llm-consumer, a header their auth proxy sets only AFTER
authenticating, so the body cleared both edges and the auth. On llm.vm, aigw
413s at 39 KB while the vLLM backend answers 200 at the same size. I found the
boundary; they found the component, by reading the failure's response headers.
Consequences recorded in the doc:
* DO NOT plan around 32 KiB. The intended ceiling is far higher, so sizing our
profiles to it would be designing around a bug.
* Their fix (ClientTrafficPolicy, bufferLimit: 8Mi) is written but NOT
deployed, pending their operator's approval. We do not re-test until they
confirm — a half-changed system gives a number neither side can trust.
* In standalone `aigw run` a SecurityPolicy is accepted and then silently
ignored, so "the config was accepted" proves nothing there. They will verify
by re-reading the live config_dump and sending a large request. Same
silent-default shape this repo keeps hitting, one layer down.
bridged.yaml carries the same correction (gitignored, so not in this commit).
Refs: gitea #76
This commit is contained in:
@@ -318,8 +318,28 @@ surfaces:
|
||||
/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413
|
||||
```
|
||||
|
||||
32 KiB is far below one real agent turn. This is an edge limit — Caddy `request_body max_size` and/or
|
||||
Envoy's own — so it is fixed in **systems/vms**, not here.
|
||||
32 KiB is far below one real agent turn.
|
||||
|
||||
**Root cause — confirmed by the systems/vms side, 2026-08-15.** My guess that it was a Caddy
|
||||
`request_body max_size` was **wrong**. It is Envoy, inside `aigw` on `llm.vm`. Envoy Gateway defaults
|
||||
a listener's `per_connection_buffer_limit_bytes` to **32768**, and the AI Gateway buffers the *whole*
|
||||
request body before it can route on the model name — so that default is not a network tuning knob
|
||||
here, it is a hard ceiling on prompt size. Read out of the live Envoy `config_dump`:
|
||||
|
||||
```
|
||||
listener default/llm/http per_connection_buffer_limit_bytes: 32768
|
||||
```
|
||||
|
||||
Nobody chose 32 KiB; it was inherited from the default. Both TLS edges are innocent: the same
|
||||
boundary reproduces on the LAN path and the internet path, and both 413s carry an `x-llm-consumer`
|
||||
header their auth proxy sets only *after* authenticating — so the body cleared both edges and the
|
||||
auth. Directly on `llm.vm`, `aigw` 413s at 39 KB while the vLLM backend accepts the same 39 KB and
|
||||
answers 200.
|
||||
|
||||
**Do not plan around 32 KiB.** The intended ceiling is far higher. Their fix — a `ClientTrafficPolicy`
|
||||
setting `bufferLimit: 8Mi` — is written but **not deployed** as of this note, pending their operator's
|
||||
approval. I have not re-tested and will not until they confirm, so as not to measure a half-changed
|
||||
system. Fixed in **systems/vms**, not here.
|
||||
|
||||
### The part worth remembering
|
||||
|
||||
@@ -368,9 +388,14 @@ one turn; this one costs the whole task and is indistinguishable from a slow wor
|
||||
|
||||
### To resume
|
||||
|
||||
Once the edge limit is raised: set `baseUrl` + `tokenEnv` on `local`, raise `gx` to `weight: 100`,
|
||||
restart, and re-run §7 **with a file-reading task**. `bridged.yaml` carries the exact two-key edit
|
||||
and these numbers inline.
|
||||
**Wait for systems/vms to confirm the fix is deployed and verified.** Do not re-test before that —
|
||||
measuring a half-changed system produces a result nobody can trust. They have said they will verify by
|
||||
re-reading the live `config_dump` and by sending a large request, not by the absence of an error,
|
||||
because in standalone `aigw run` a `SecurityPolicy` is accepted and then silently ignored. If
|
||||
`ClientTrafficPolicy` turns out to be ignored the same way, the fix will need a different shape.
|
||||
|
||||
Then: set `baseUrl` + `tokenEnv` on `local`, raise `gx` to `weight: 100`, restart, and re-run §7
|
||||
**with a file-reading task**. `bridged.yaml` carries the exact two-key edit and these numbers inline.
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user