CB-591: record the final 86400s route timeout and the request:0s trap
systems/vms moved the LLM route timeout again, 1800s -> 86400s (24h), after the silent-truncation risk was discussed. They tried request: 0s first: it removes the total-duration timer, but on an AIGatewayRoute the idle timeout is derived from the request timeout, so 0s also removed any bound on a stalled connection. At 86400s our own MessageService.ASYNC_TIMEOUT_MS (30 min) binds first, so a runaway request now ends as a clean FAILED ticket we raised instead of a silently truncated 200. While the gateway sat at 1800s the two numbers were equal and did not nest.
This commit is contained in:
@@ -394,7 +394,13 @@ systems/vms fixed both, and each was re-checked from this side rather than taken
|
||||
| ceiling | was | now | our own check |
|
||||
|---|---|---|---|
|
||||
| listener buffer | 32 KiB | 32 Mi | 1.2 MB body → **200** (was 413) |
|
||||
| LLM route timeout | 60s | 1800s | the request that truncated: **101s, `message_stop` present, 4000/4000** |
|
||||
| LLM route timeout | 60s | 86400s | the request that truncated: **101s, `message_stop` present, 4000/4000** |
|
||||
|
||||
The timeout moved in two steps on 2026-08-15: 60s → 1800s, then 1800s → **86400s (24 hours)** after
|
||||
the truncation risk below was discussed. They tried `request: 0s` first, which removes the
|
||||
total-duration timer completely. It works, but on an `AIGatewayRoute` the **idle timeout is derived
|
||||
from the request timeout**, so `0s` also removed any bound on a stalled connection. 86400s keeps a
|
||||
reaper for dead connections while putting the truncation timer out of practical reach.
|
||||
|
||||
Neither was deliberate. The 32 KiB was Envoy Gateway's default `per_connection_buffer_limit_bytes`;
|
||||
the 60s was Envoy AI Gateway's own documented default. The 60s bounded **generation** as well as
|
||||
@@ -438,9 +444,15 @@ check. Whether either detects a missing terminator is unverified — and opencod
|
||||
|
||||
So the honest statement of our position:
|
||||
|
||||
> Gateway traffic is acceptable at 1800s because a single request would have to run for 30 minutes to
|
||||
> Gateway traffic is acceptable at 86400s because a single request would have to run for 24 hours to
|
||||
> trip the bug — **not** because we could detect it if it did.
|
||||
|
||||
At 86400s our **own** limit binds first, which is the ordering we want. `MessageService.ASYNC_TIMEOUT_MS`
|
||||
caps a turn at 30 minutes, so a runaway request ends as a clean `FAILED` ticket that we raised, rather
|
||||
than as a silently truncated `200` that we cannot see. While the gateway sat at 1800s the two numbers
|
||||
were equal and did not nest, so a gateway-side stall could have been misread as a bug in our own ticket
|
||||
handling. That ambiguity is now gone.
|
||||
|
||||
**If a member ever returns a confident but truncated answer, suspect this before anything in our own
|
||||
code.** That is the whole reason this section exists.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user