CB-591: point member profiles at the new LLM/MCP gateway (claude-code and opencode) #76

Closed
opened 2026-08-15 17:00:38 +02:00 by ltms · 5 comments
Owner

Full plan: docs/CB-591-Gateway-Migration.md. Upstream: systems/vms wiki → LLM and MCP Gateway, systems/vms#31.

The gateway went live 2026-08-15 and replaced Bifrost. It serves an Anthropic surface and an
OpenAI surface, so both member kinds can point at it.

Why

The upstream wiki names us as a blocker for retiring the shared legacy token: "kb, brain,
claude-bridge and the workstation still share it. Each needs its own consumer first."

But the bigger win is on the opencode side. Every opencode member today is sol or terra, and both
sit on one OpenAI account via credentialId: openai-shared — an exhaustion on either locks out
both. A gateway-backed opencode profile is free, is not on that credential, and is therefore not in
that quarantine pair. That retires a single point of failure rather than only adding capacity.

Two profiles, and they are not symmetric

local (claude-code) gx (opencode, new)
surface https://llm.ltms.dev/anthropic https://llm.ltms.dev/v1
SubscriptionGuard applies — needs llm.ltms.dev on guard.offSubscriptionHosts does not apply (guard is Claude-private by design)
daemon restart required (the guard is built once at Bridged.java:93) not required
code change none none — OpenCodeLauncher already pins an OpenAI-compatible endpoint (CB-508)

Do the opencode profile first: it proves the token, URL and model name against a live gateway
while nothing the fleet depends on is touched.

Acceptance criteria

  1. A gx opencode profile spawns, completes a turn and ends in bridge_reply.
  2. local points at /anthropic, llm.ltms.dev is on the guard allowlist, and a local member
    completes a turn after a daemon restart.
  3. Reasoning survives, checked on each surface separately. On /anthropic this is the only check
    that catches an OpenAI-schema mistake — Envoy's translator looks for thinking_blocks while our
    vLLM sends reasoning_content, and every thinking delta then disappears silently. On /v1
    this is an open question the plan deliberately does not assume.
  4. A local-direct profile at weight: 0 keeps the direct gx00.gw:8000 path reachable by explicit
    spawn, so the documented escape hatch is real rather than theoretical.
  5. bridge_list shows gx with no credentialId, so a sol/terra exhaustion cannot quarantine it.
  6. The cockpit at auth.ltms.dev counts requests against a claude-bridge consumer, not legacy.
  7. wiki/11-Features.md gets an entry: what it does, the knob, why it exists, the gotcha.
  8. The claim in CLAUDE.md that a member mounts "only the bridge MCP" is corrected or made true — it
    is already false for opencode members, which get context7 and gitea via opencode.json.

Blocked on

Criterion 1–3, 5, 6 need a claude-bridge consumer token issued at auth.ltms.dev and stored in
${SHARED_ENV}/tools/secrets.sh as LLM_BRIDGE_TOKEN. That is the operator's step — it touches
secrets and a host this repo does not own.

Open decision (operator)

Whether members should also mount the gateway's multiplexed /mcp. mcpUrl is a single String in
BridgedConfig.Profile, so a claude-code member can mount exactly one MCP — adding the gateway MCP is
a code change, and it changes an invariant CLAUDE.md states twice. Not decided in this plan.

Carried-over trap worth repeating

Rotating a consumer token restarts the auth proxy, which drops in-flight streaming responses. For
us that kills every live member mid-turn and takes their async ticket reports with them. Same rule as
a daemon redeploy: drain the fleet first (bridge_list → bridge_poll → bridge_stop), then rotate.

Full plan: `docs/CB-591-Gateway-Migration.md`. Upstream: [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway), [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31). The gateway went live 2026-08-15 and replaced Bifrost. It serves an Anthropic surface **and** an OpenAI surface, so both member kinds can point at it. ## Why The upstream wiki names us as a blocker for retiring the shared `legacy` token: *"kb, brain, **claude-bridge** and the workstation still share it. Each needs its own consumer first."* But the bigger win is on the opencode side. Every opencode member today is `sol` or `terra`, and both sit on **one** OpenAI account via `credentialId: openai-shared` — an exhaustion on either locks out both. A gateway-backed opencode profile is free, is not on that credential, and is therefore not in that quarantine pair. That retires a single point of failure rather than only adding capacity. ## Two profiles, and they are not symmetric | | `local` (claude-code) | `gx` (opencode, new) | |---|---|---| | surface | `https://llm.ltms.dev/anthropic` | `https://llm.ltms.dev/v1` | | `SubscriptionGuard` | applies — needs `llm.ltms.dev` on `guard.offSubscriptionHosts` | **does not apply** (guard is Claude-private by design) | | daemon restart | **required** (the guard is built once at `Bridged.java:93`) | not required | | code change | none | none — `OpenCodeLauncher` already pins an OpenAI-compatible endpoint (CB-508) | Do the opencode profile **first**: it proves the token, URL and model name against a live gateway while nothing the fleet depends on is touched. ## Acceptance criteria 1. A `gx` opencode profile spawns, completes a turn and ends in `bridge_reply`. 2. `local` points at `/anthropic`, `llm.ltms.dev` is on the guard allowlist, and a `local` member completes a turn after a daemon restart. 3. **Reasoning survives, checked on each surface separately.** On `/anthropic` this is the only check that catches an OpenAI-schema mistake — Envoy's translator looks for `thinking_blocks` while our vLLM sends `reasoning_content`, and every thinking delta then disappears **silently**. On `/v1` this is an open question the plan deliberately does not assume. 4. A `local-direct` profile at `weight: 0` keeps the direct `gx00.gw:8000` path reachable by explicit spawn, so the documented escape hatch is real rather than theoretical. 5. `bridge_list` shows `gx` with no `credentialId`, so a `sol`/`terra` exhaustion cannot quarantine it. 6. The cockpit at `auth.ltms.dev` counts requests against a `claude-bridge` consumer, not `legacy`. 7. `wiki/11-Features.md` gets an entry: what it does, the knob, **why it exists**, the gotcha. 8. The claim in `CLAUDE.md` that a member mounts "only the bridge MCP" is corrected or made true — it is already false for opencode members, which get context7 and gitea via `opencode.json`. ## Blocked on Criterion 1–3, 5, 6 need a `claude-bridge` consumer token issued at `auth.ltms.dev` and stored in `${SHARED_ENV}/tools/secrets.sh` as `LLM_BRIDGE_TOKEN`. That is the operator's step — it touches secrets and a host this repo does not own. ## Open decision (operator) Whether members should also mount the gateway's multiplexed `/mcp`. `mcpUrl` is a single `String` in `BridgedConfig.Profile`, so a claude-code member can mount exactly one MCP — adding the gateway MCP is a code change, and it changes an invariant `CLAUDE.md` states twice. Not decided in this plan. ## Carried-over trap worth repeating Rotating a consumer token **restarts the auth proxy**, which drops in-flight streaming responses. For us that kills every live member mid-turn and takes their async ticket reports with them. Same rule as a daemon redeploy: drain the fleet first (`bridge_list` → `bridge_poll` → `bridge_stop`), then rotate.
Author
Owner

Token name settled, and one correction to the plan.

The operator issued the consumer and exported it as AI_GATEWAY_TOKEN — one key for every agent
and MCP client behind llm.ltms.dev, not a bridge-specific one. The plan doc used
LLM_BRIDGE_TOKEN; that name is now wrong everywhere and has been updated
(docs/CB-591-Gateway-Migration.md, commit f0095bf).

Confirmed without printing the value: it resolves in a login shell, is 48 characters, and carries the
documented llmk- prefix.

Correction — the opencode profile also needs a restart. The plan claimed U2a needed none because
SubscriptionGuard does not apply to opencode. The guard part is right, the conclusion was not.
tokenEnv is resolved by HerdrPeerLauncher.resolveEnv → env.apply(name), which reads the
daemon's own process environment. The running bridged inherited its environment at startup, so a
variable added to secrets.sh afterwards is absent — the launcher would inject an empty token and the
gateway would answer 401, long after the restart and with nothing connecting the two.

That is the same trap as WORKER_GITEA_TOKEN, so it now gets the same treatment:
scripts/redeploy-bridged.sh --check reports whether AI_GATEWAY_TOKEN resolves in a login shell,
and never prints it. Verified against the live daemon:

   ok    WORKER_GITEA_TOKEN resolves in a login shell
   ok    AI_GATEWAY_TOKEN resolves in a login shell

Revised order: write both profiles, restart once from a login shell, then verify gx before
local. Both profiles need the restart, so there is no reason to do two — but there is still a reason
to verify in that order. gx exercises the token and the gateway with no guard involved, so a failure
there is upstream; local adds the guard allowlist on top, so a failure there is ours. Testing in that
order separates the two causes instead of confusing them.

**Token name settled, and one correction to the plan.** The operator issued the consumer and exported it as **`AI_GATEWAY_TOKEN`** — one key for every agent and MCP client behind `llm.ltms.dev`, not a bridge-specific one. The plan doc used `LLM_BRIDGE_TOKEN`; that name is now wrong everywhere and has been updated (`docs/CB-591-Gateway-Migration.md`, commit `f0095bf`). Confirmed without printing the value: it resolves in a login shell, is 48 characters, and carries the documented `llmk-` prefix. **Correction — the opencode profile also needs a restart.** The plan claimed U2a needed none because `SubscriptionGuard` does not apply to opencode. The guard part is right, the conclusion was not. `tokenEnv` is resolved by `HerdrPeerLauncher.resolveEnv` → `env.apply(name)`, which reads the **daemon's own process environment**. The running `bridged` inherited its environment at startup, so a variable added to `secrets.sh` afterwards is absent — the launcher would inject an empty token and the gateway would answer 401, long after the restart and with nothing connecting the two. That is the same trap as `WORKER_GITEA_TOKEN`, so it now gets the same treatment: `scripts/redeploy-bridged.sh --check` reports whether `AI_GATEWAY_TOKEN` resolves in a login shell, and never prints it. Verified against the live daemon: ``` ok WORKER_GITEA_TOKEN resolves in a login shell ok AI_GATEWAY_TOKEN resolves in a login shell ``` Revised order: write both profiles, restart **once** from a login shell, then verify `gx` before `local`. Both profiles need the restart, so there is no reason to do two — but there is still a reason to verify in that order. `gx` exercises the token and the gateway with no guard involved, so a failure there is upstream; `local` adds the guard allowlist on top, so a failure there is ours. Testing in that order separates the two causes instead of confusing them.
Author
Owner

Deployed, tested live, and REVERTED — blocked on a 32 KiB request body limit at the edge

U1–U2c were done and the daemon restarted onto them. Then live verification found a blocker that no
amount of config on our side can fix, so local is back on the direct vLLM and gx is at
weight: 0.

What is actually wrong

llm.ltms.dev answers HTTP 413 Request Entity Too Large above 32 KiB (32768 bytes), on
both surfaces. Measured, not inferred:

/v1        32695 bytes -> 200        /anthropic  32095 bytes -> 200
/v1        32795 bytes -> 413        /anthropic  32855 bytes -> 413

32 KiB is far below one real agent turn. This is an edge limit — Caddy request_body max_size
and/or Envoy's own limit — so the fix is in systems/vms, not here.

How it presented, and why that is the dangerous part

Two members were spawned at the same time, same message:

local (claude-code, /anthropic) gx (opencode, /v1)
READY → BUSY 19:07:26 19:07:45
BUSY → DONE 19:08:51 (66s) never — 10+ min, ticket FAILED

local passed. It passed because the probe was three trivial questions in a fresh session, so
the request squeaked under 32 KiB. That is the trap: the profile looks healthy and then fails on the
first turn that reads a file.

gx did not fail loudly either. Reproduced outside the bridge, running opencode by hand with the
launcher's own generated config:

Error: Request Entity Too Large
...compacts context, retries...
Error: Request Entity Too Large

opencode catches the 413, compacts, and retries — forever. A member that fails loudly costs one
turn. This one costs the whole task and is indistinguishable from a hang.

What was verified good, and stays good

Everything except the size limit. Worth recording so none of it is re-tested later:

  • Token accepted on both surfaces; unauthenticated → 401, so the Caddy proxy really is gating
    (the wiki warns the gateway's own SecurityPolicy fails open — that warning is about the gateway,
    and the proxy in front does its job).
  • /v1/models returns exactly ["deepseek-v4-flash"] — the exact-name trap is clear.
  • Reasoning survives both surfaces. /anthropic returns a real "type":"thinking" block; /v1
    returns a populated reasoning_content. That closes the open question in §3b of the plan doc — the
    answer is that /v1 does not drop reasoning here.
  • The launcher's generated opencode provider block is correct: right baseURL, right model, and a
    real 48-char llmk- key rather than the bridged-local-noauth placeholder.
  • SubscriptionGuard accepted llm.ltms.dev after the allowlist edit — local spawned without
    throwing, which is the check that would have caught a missed restart.

Current state

  • local → http://gx00.gw:8000 (direct vLLM), weight: 100. Working.
  • gx → kept, weight: 0, so nothing auto-selects it. Still spawnable explicitly to test the fix.
  • local-direct → kept, temporarily duplicates local.
  • llm.ltms.dev stays in guard.offSubscriptionHosts — it only permits, and it is needed again on
    the way back.
  • Daemon pid 66745, jar f1fd659423e6.

bridged.yaml carries the numbers above and the exact two-key edit to switch back, so whoever picks
this up does not have to re-derive any of it.

Blocked on

Raising the request body limit at the edge in systems/vms. Once that lands: set baseUrl +
tokenEnv on local, raise gx to weight: 100, restart, and re-run §7 — but this time with a
task that reads a file, not a three-question probe. The trivial probe is what made this look fine
the first time.

## Deployed, tested live, and REVERTED — blocked on a 32 KiB request body limit at the edge U1–U2c were done and the daemon restarted onto them. Then live verification found a blocker that no amount of config on our side can fix, so `local` is back on the direct vLLM and `gx` is at `weight: 0`. ### What is actually wrong `llm.ltms.dev` answers **HTTP 413 Request Entity Too Large** above **32 KiB (32768 bytes)**, on **both** surfaces. Measured, not inferred: ``` /v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200 /v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413 ``` 32 KiB is far below one real agent turn. This is an edge limit — Caddy `request_body max_size` and/or Envoy's own limit — so the fix is in **systems/vms**, not here. ### How it presented, and why that is the dangerous part Two members were spawned at the same time, same message: | | `local` (claude-code, `/anthropic`) | `gx` (opencode, `/v1`) | |---|---|---| | READY → BUSY | 19:07:26 | 19:07:45 | | BUSY → DONE | **19:08:51 (66s)** | **never — 10+ min, ticket FAILED** | `local` **passed**. It passed because the probe was three trivial questions in a fresh session, so the request squeaked under 32 KiB. That is the trap: the profile looks healthy and then fails on the first turn that reads a file. `gx` did not fail loudly either. Reproduced outside the bridge, running `opencode` by hand with the launcher's own generated config: ``` Error: Request Entity Too Large ...compacts context, retries... Error: Request Entity Too Large ``` opencode **catches the 413, compacts, and retries — forever**. A member that fails loudly costs one turn. This one costs the whole task and is indistinguishable from a hang. ### What was verified good, and stays good Everything except the size limit. Worth recording so none of it is re-tested later: - Token accepted on both surfaces; **unauthenticated → 401**, so the Caddy proxy really is gating (the wiki warns the gateway's own `SecurityPolicy` fails open — that warning is about the gateway, and the proxy in front does its job). - `/v1/models` returns exactly `["deepseek-v4-flash"]` — the exact-name trap is clear. - **Reasoning survives both surfaces.** `/anthropic` returns a real `"type":"thinking"` block; `/v1` returns a populated `reasoning_content`. That closes the open question in §3b of the plan doc — the answer is that `/v1` does *not* drop reasoning here. - The launcher's generated opencode provider block is correct: right `baseURL`, right model, and a real 48-char `llmk-` key rather than the `bridged-local-noauth` placeholder. - `SubscriptionGuard` accepted `llm.ltms.dev` after the allowlist edit — `local` spawned without throwing, which is the check that would have caught a missed restart. ### Current state - `local` → `http://gx00.gw:8000` (direct vLLM), `weight: 100`. Working. - `gx` → kept, `weight: 0`, so nothing auto-selects it. Still spawnable explicitly to test the fix. - `local-direct` → kept, temporarily duplicates `local`. - `llm.ltms.dev` stays in `guard.offSubscriptionHosts` — it only permits, and it is needed again on the way back. - Daemon pid 66745, jar `f1fd659423e6`. `bridged.yaml` carries the numbers above and the exact two-key edit to switch back, so whoever picks this up does not have to re-derive any of it. ### Blocked on Raising the request body limit at the edge in **systems/vms**. Once that lands: set `baseUrl` + `tokenEnv` on `local`, raise `gx` to `weight: 100`, restart, and re-run §7 — but this time with a task that **reads a file**, not a three-question probe. The trivial probe is what made this look fine the first time.
Author
Owner

Root cause confirmed by systems/vms — it is Envoy, not Caddy, and it is not deliberate

Correcting my previous comment. I wrote that the limit was "Caddy request_body max_size and/or
Envoy's own" and marked it unverified. The Caddy half was wrong.

The actual cause

Envoy, inside aigw on llm.vm. Envoy Gateway defaults a listener's
per_connection_buffer_limit_bytes to 32768, and the AI Gateway buffers the whole request
body
before it can route on the model name. So that default is not a network tuning knob here — it
is a hard ceiling on prompt size. Read out of the running Envoy's own config_dump:

listener default/llm/http    per_connection_buffer_limit_bytes: 32768

Nobody chose 32 KiB. It was inherited from the default.

How they proved both TLS edges were innocent

Worth recording, because it is a better technique than the one I used:

  • The same 32 KiB boundary reproduces on the LAN path (gw.vm nginx) and the internet path (Caddy).
  • Both 413 responses carry x-llm-consumer — a header their auth proxy sets only after it has
    authenticated the request. So the body cleared both edges and the auth proxy before anything
    rejected it.
  • On llm.vm itself: aigw returns 200 at 32086 bytes and 413 at 39086 bytes, while the vLLM backend
    answers 200 at the same 39086 bytes. The backend was never the constraint.

I found the boundary; they found the component. Reading the failure's response headers is what
separated the two, and I did not think to do it.

What this changes for us

Do not plan around 32 KiB. The intended ceiling is far higher, so sizing our profiles to the
current limit would be designing around a bug.

Their fix is a ClientTrafficPolicy setting bufferLimit: 8Mi — written and committed, not
deployed
, pending their operator's approval. At 8Mi the model's context window becomes the binding
limit instead of the buffer, which takes this off the list of numbers anyone has to reason about.

We are holding off

No re-testing until they confirm it is live and verified. Measuring a half-changed system gives a
result neither side can trust. This is now written into docs/CB-591-Gateway-Migration.md so it is not
picked up later out of impatience.

One caveat they flagged that we should carry: in standalone aigw run, a SecurityPolicy is
accepted and then silently ignored. So "the config was accepted" proves nothing in that stack —
they will verify by re-reading the live config_dump and by sending a large request, not by the
absence of an error. If ClientTrafficPolicy is ignored the same way, the fix needs a different shape.
That is the same silent-default failure mode this repo keeps hitting, one layer down.

Unchanged

local stays on the direct vLLM, gx stays at weight: 0. When the fix lands, the re-test is a
member that reads real files — our first check passed at 32 KiB and hid this for an afternoon.

## Root cause confirmed by systems/vms — it is Envoy, not Caddy, and it is not deliberate Correcting my previous comment. I wrote that the limit was "Caddy `request_body max_size` and/or Envoy's own" and marked it unverified. The Caddy half was **wrong**. ### The actual cause Envoy, inside `aigw` on `llm.vm`. Envoy Gateway defaults a listener's `per_connection_buffer_limit_bytes` to **32768**, and the AI Gateway buffers the **whole request body** before it can route on the model name. So that default is not a network tuning knob here — it is a hard ceiling on prompt size. Read out of the running Envoy's own `config_dump`: ``` listener default/llm/http per_connection_buffer_limit_bytes: 32768 ``` Nobody chose 32 KiB. It was inherited from the default. ### How they proved both TLS edges were innocent Worth recording, because it is a better technique than the one I used: - The same 32 KiB boundary reproduces on the LAN path (gw.vm nginx) **and** the internet path (Caddy). - **Both 413 responses carry `x-llm-consumer`** — a header their auth proxy sets only *after* it has authenticated the request. So the body cleared both edges and the auth proxy before anything rejected it. - On `llm.vm` itself: `aigw` returns 200 at 32086 bytes and 413 at 39086 bytes, while the vLLM backend answers **200 at the same 39086 bytes**. The backend was never the constraint. I found the boundary; they found the component. Reading the failure's response headers is what separated the two, and I did not think to do it. ### What this changes for us **Do not plan around 32 KiB.** The intended ceiling is far higher, so sizing our profiles to the current limit would be designing around a bug. Their fix is a `ClientTrafficPolicy` setting `bufferLimit: 8Mi` — written and committed, **not deployed**, pending their operator's approval. At 8Mi the model's context window becomes the binding limit instead of the buffer, which takes this off the list of numbers anyone has to reason about. ### We are holding off **No re-testing until they confirm it is live and verified.** Measuring a half-changed system gives a result neither side can trust. This is now written into `docs/CB-591-Gateway-Migration.md` so it is not picked up later out of impatience. One caveat they flagged that we should carry: in standalone `aigw run`, a `SecurityPolicy` is **accepted and then silently ignored**. So "the config was accepted" proves nothing in that stack — they will verify by re-reading the live `config_dump` and by sending a large request, not by the absence of an error. If `ClientTrafficPolicy` is ignored the same way, the fix needs a different shape. That is the same silent-default failure mode this repo keeps hitting, one layer down. ### Unchanged `local` stays on the direct vLLM, `gx` stays at `weight: 0`. When the fix lands, the re-test is a member that **reads real files** — our first check passed at 32 KiB and hid this for an afternoon.
Author
Owner

DONE — the fleet is on the gateway, both ceilings fixed and re-verified

systems/vms fixed both blockers. I re-checked each from this side rather than taking them on trust,
then moved the fleet across and proved it with real workloads.

ceiling was now my own check
listener buffer 32 KiB 32 Mi 1.2 MB body → 200 (was 413)
LLM route timeout 60s 1800s the request that truncated: 101s, message_stop present, 4000/4000

Neither was deliberate. 32 KiB was Envoy Gateway's default per_connection_buffer_limit_bytes; 60s
was Envoy AI Gateway's own default. The 60s bounded generation as well as prompt size — a tiny
prompt with a long answer returned 504 at 60.05s.

Live verification — real workloads, not liveness probes

profile surface check result
local /anthropic read this plan + CLAUDE.md in full, answer 4 questions ✅ all correct
gx /v1 same two files, 3 questions ✅ all correct
local-direct direct vLLM escape hatch reachable ✅ hatch-ok

The decisive number: those two files are 48,344 bytes of content, comfortably past the old 32,768
ceiling. That exact request was a 413 this morning. The local member also reported the document's
length as "roughly 456 lines" against an actual 455 — it genuinely read the whole thing.

A trivial 3-question probe is what hid the 32 KiB ceiling for an afternoon, so it no longer counts as
proof for this profile. Any future re-test uses a file-reading task.

Final state

local         claude-code   llm.ltms.dev/anthropic   AI_GATEWAY_TOKEN   weight 100
gx            opencode      llm.ltms.dev/v1          AI_GATEWAY_TOKEN   weight 100
local-direct  claude-code   gx00.gw:8000             (none)             weight 0   <- escape hatch
guard: [gx00.gw, llm.ltms.dev]

Daemon pid 17702. local-direct is now genuinely load-bearing rather than decorative: it is the only
claude-code path with no TLS edge, no auth proxy and no gateway in it.

What this bought

  • Free opencode capacity off the shared OpenAI credential. gx carries no credentialId, so a
    sol/terra exhaustion can no longer quarantine it. This retires a single point of failure — the
    actual point of the migration, more than cost.
  • Per-consumer usage metering, the first real input #74 (CB-589) needs.
  • Our own revocable token instead of the shared legacy one.

One risk ACCEPTED, not solved — read §7.2 before debugging a weird member

On a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked encoding cleanly rather
than resetting, so a truncated answer arrives as HTTP 200 with no error and no terminator
(envoyproxy/envoy#17186 — acknowledged 2021, closed by a stale bot, never fixed; the Dec 2025 fix
#42269 is HTTP/2 only, and SSE here is HTTP/1.1). Measured at the old 60s:

HTTP 200   61.07s   2473 of 4000 emitted
message_stop 0   message_delta 0   error events 0   final frame well-formed

The recommended defence — reject a stream with no terminator — does not transfer to us. Claude
Code and opencode are third-party clients and we do not own their SSE parsing. So:

Gateway traffic is acceptable at 1800s because a request would have to run 30 minutes to trip the
bug — not because we could detect it if it did.

If a member ever returns a confident but truncated answer, suspect this before anything in our own
code.

Also worth keeping from the upstream bisection: ClientTrafficPolicy is honoured in standalone
aigw run, BackendTrafficPolicy is silently ignored, and nothing external distinguishes them
(envoyproxy/gateway#9513). The same silent-default shape this repo keeps hitting.

Still open upstream, not blocking us

systems/vms is asking their operator about dropping the route timeout to 0 bounded by the existing
3600s idle timeout — a better fit for this bug than any total-duration limit — and about adding a
long-running probe. Their new 40 KB probe is live and passing.

Closing. Remaining CB-591 units U4 (context7 via the gateway /mcp) and U5 (docs) are separate and
tracked in #79.

## DONE — the fleet is on the gateway, both ceilings fixed and re-verified systems/vms fixed both blockers. I re-checked each from this side rather than taking them on trust, then moved the fleet across and proved it with real workloads. | ceiling | was | now | my own check | |---|---|---|---| | listener buffer | 32 KiB | 32 Mi | 1.2 MB body → **200** (was 413) | | LLM route timeout | 60s | 1800s | the request that truncated: **101s, `message_stop` present, 4000/4000** | Neither was deliberate. 32 KiB was Envoy Gateway's default `per_connection_buffer_limit_bytes`; 60s was Envoy AI Gateway's own default. The 60s bounded **generation** as well as prompt size — a tiny prompt with a long answer returned 504 at 60.05s. ### Live verification — real workloads, not liveness probes | profile | surface | check | result | |---|---|---|---| | `local` | `/anthropic` | read this plan + `CLAUDE.md` in full, answer 4 questions | ✅ all correct | | `gx` | `/v1` | same two files, 3 questions | ✅ all correct | | `local-direct` | direct vLLM | escape hatch reachable | ✅ `hatch-ok` | The decisive number: those two files are **48,344 bytes** of content, comfortably past the old 32,768 ceiling. That exact request was a 413 this morning. The `local` member also reported the document's length as "roughly 456 lines" against an actual 455 — it genuinely read the whole thing. A trivial 3-question probe is what hid the 32 KiB ceiling for an afternoon, so it no longer counts as proof for this profile. Any future re-test uses a file-reading task. ### Final state ``` local claude-code llm.ltms.dev/anthropic AI_GATEWAY_TOKEN weight 100 gx opencode llm.ltms.dev/v1 AI_GATEWAY_TOKEN weight 100 local-direct claude-code gx00.gw:8000 (none) weight 0 <- escape hatch guard: [gx00.gw, llm.ltms.dev] ``` Daemon pid 17702. `local-direct` is now genuinely load-bearing rather than decorative: it is the only claude-code path with no TLS edge, no auth proxy and no gateway in it. ### What this bought - **Free opencode capacity off the shared OpenAI credential.** `gx` carries no `credentialId`, so a `sol`/`terra` exhaustion can no longer quarantine it. This retires a single point of failure — the actual point of the migration, more than cost. - **Per-consumer usage metering**, the first real input #74 (CB-589) needs. - Our own revocable token instead of the shared `legacy` one. ### One risk ACCEPTED, not solved — read §7.2 before debugging a weird member On a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked encoding *cleanly* rather than resetting, so a truncated answer arrives as **HTTP 200 with no error and no terminator** (`envoyproxy/envoy#17186` — acknowledged 2021, closed by a stale bot, never fixed; the Dec 2025 fix `#42269` is HTTP/2 only, and SSE here is HTTP/1.1). Measured at the old 60s: ``` HTTP 200 61.07s 2473 of 4000 emitted message_stop 0 message_delta 0 error events 0 final frame well-formed ``` The recommended defence — reject a stream with no terminator — **does not transfer to us**. Claude Code and opencode are third-party clients and we do not own their SSE parsing. So: > Gateway traffic is acceptable at 1800s because a request would have to run 30 minutes to trip the > bug — **not** because we could detect it if it did. **If a member ever returns a confident but truncated answer, suspect this before anything in our own code.** Also worth keeping from the upstream bisection: `ClientTrafficPolicy` is honoured in standalone `aigw run`, `BackendTrafficPolicy` is **silently ignored**, and nothing external distinguishes them (`envoyproxy/gateway#9513`). The same silent-default shape this repo keeps hitting. ### Still open upstream, not blocking us systems/vms is asking their operator about dropping the route timeout to `0` bounded by the existing 3600s idle timeout — a better fit for this bug than any total-duration limit — and about adding a long-running probe. Their new 40 KB probe is live and passing. Closing. Remaining CB-591 units U4 (context7 via the gateway `/mcp`) and U5 (docs) are separate and tracked in #79.
ltms closed this issue 2026-08-15 20:49:05 +02:00
Author
Owner

Correction: one acceptance criterion is NOT verified

Closing this, I listed per-consumer usage metering as a benefit. That was a capability claim, not a
measurement, and I should separate the two.

Criterion 6 — "the cockpit at auth.ltms.dev counts requests against a claude-bridge consumer,
not legacy" — is UNVERIFIED.
The cockpit needs credentials I do not have:

https://auth.ltms.dev/                -> 401
https://auth.ltms.dev/api/consumers   -> 401
https://auth.ltms.dev/stats           -> 401

What I can say from evidence: the fleet authenticates with AI_GATEWAY_TOKEN, an llmk- consumer
token issued for this purpose, and requests carrying it are accepted while unauthenticated ones get
401. So traffic is attributed to a consumer token. Whether the cockpit reports it under
claude-bridge specifically, and whether legacy can now be retired, I have not seen.

That matters beyond bookkeeping: those figures are the input #74 (CB-589) needs for cost-first
placement. Building on a number nobody has looked at would be the same mistake as trusting a config
that was merely accepted.

Whoever has cockpit access should confirm it. It is a one-minute check and it is the difference
between "we took our own token" and "our own token is actually being counted".

The rest of the criteria stand as reported — 1, 2, 3, 4 and 5 were each checked live, and 7 and 8 are
tracked in #79. This is the only one carried on trust, and it should not have been folded into a
benefits list.

## Correction: one acceptance criterion is NOT verified Closing this, I listed per-consumer usage metering as a benefit. That was a capability claim, not a measurement, and I should separate the two. **Criterion 6 — "the cockpit at `auth.ltms.dev` counts requests against a `claude-bridge` consumer, not `legacy`" — is UNVERIFIED.** The cockpit needs credentials I do not have: ``` https://auth.ltms.dev/ -> 401 https://auth.ltms.dev/api/consumers -> 401 https://auth.ltms.dev/stats -> 401 ``` What I can say from evidence: the fleet authenticates with `AI_GATEWAY_TOKEN`, an `llmk-` consumer token issued for this purpose, and requests carrying it are accepted while unauthenticated ones get 401. So traffic is *attributed to a consumer token*. Whether the cockpit reports it under `claude-bridge` specifically, and whether `legacy` can now be retired, I have not seen. That matters beyond bookkeeping: those figures are the input #74 (CB-589) needs for cost-first placement. Building on a number nobody has looked at would be the same mistake as trusting a config that was merely accepted. **Whoever has cockpit access should confirm it.** It is a one-minute check and it is the difference between "we took our own token" and "our own token is actually being counted". The rest of the criteria stand as reported — 1, 2, 3, 4 and 5 were each checked live, and 7 and 8 are tracked in #79. This is the only one carried on trust, and it should not have been folded into a benefits list.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#76