555715ced9
The repo moved lms/claude-bridge -> fleet/claude-bridge -> fleet/fleetd. Gitea redirects hold, so most of this is not urgent, but one line was a real break: the implementer skill posts a worker's PR to a hardcoded repo path, so every worker PR would have gone to the old address. .claude/skills/implementer/SKILL.md the worker PR endpoint (functional) .gitmodules wiki submodule URL CLAUDE.md + wiki/7-Use-Cases.md the canonical block, kept byte-identical README.md clone command and wiki link deploy/bridged.service Documentation= plugin/.claude-plugin/plugin.json homepage + repository docs/*.md issue and wiki links The wiki is not a separate repo. /repos/lms/claude-bridge.wiki returns 404 and lms owned no .wiki entity, so the wiki moved with the repo; both the old and the new wiki SSH URLs resolve to the same sha. Ticket step 4 assumed a second transfer that does not exist.
468 lines
25 KiB
Markdown
468 lines
25 KiB
Markdown
# CB-591 — move the fleet onto the LLM and MCP gateway
|
||
|
||
**Status: DONE — the fleet is on the gateway as of 2026-08-15.** `local` runs on `/anthropic` and
|
||
`gx` on `/v1`, both at `weight: 100`; `local-direct` stays at `weight: 0` as the escape hatch. Getting
|
||
here took a revert and two upstream fixes — see §7.1, which is the useful part of this document. One
|
||
risk is **accepted rather than solved**: a stream cut by any mid-response timer arrives as HTTP 200
|
||
with no terminator, and our third-party members cannot detect it (§7.2).
|
||
· **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway)
|
||
· **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31)
|
||
|
||
The gateway went live on 2026-08-15 and replaced Bifrost. This plan says what that means for a
|
||
**member definition** in `bridged.yaml`, because that is the part of this repo the change actually
|
||
touches.
|
||
|
||
---
|
||
|
||
## 1. What changed upstream
|
||
|
||
One front door for every LLM and MCP client: `https://llm.ltms.dev`, one token per consumer.
|
||
|
||
| Surface | URL |
|
||
|---|---|
|
||
| OpenAI chat | `https://llm.ltms.dev/v1/chat/completions` |
|
||
| OpenAI models | `https://llm.ltms.dev/v1/models` |
|
||
| **Anthropic messages** | `https://llm.ltms.dev/anthropic/v1/messages` |
|
||
| MCP, all servers multiplexed | `https://llm.ltms.dev/mcp` |
|
||
|
||
Anything outside that list returns **404 before any token is checked**, on purpose — the gateway must
|
||
never become a blanket proxy.
|
||
|
||
The model backend is unchanged: GX10 vLLM at `10.10.10.26:8000` (`gx00.gw`), model name exactly
|
||
`deepseek-v4-flash`. The direct LAN path stays open on purpose as an escape hatch.
|
||
|
||
---
|
||
|
||
## 2. Where claude-bridge sits today
|
||
|
||
We do **not** use the gateway. The `local` profile talks straight to the vLLM:
|
||
|
||
```yaml
|
||
local:
|
||
kind: claude-code
|
||
baseUrl: http://gx00.gw:8000 # direct vLLM — no auth, LAN only
|
||
model: deepseek-v4-flash
|
||
configDir: /Users/dai.ha/.ccs/instances/gx10
|
||
```
|
||
|
||
Three facts about our side that decide the shape of this work:
|
||
|
||
1. **`baseUrl` becomes `ANTHROPIC_BASE_URL`** in the member's environment, and `tokenEnv` becomes
|
||
`ANTHROPIC_AUTH_TOKEN` (the value is read from a host env var and never stored in config).
|
||
`local` sets no `tokenEnv` today, because a direct vLLM needs no token.
|
||
2. **`SubscriptionGuard` refuses any host not on an allowlist**, and that allowlist is
|
||
`guard.offSubscriptionHosts: [gx00.gw]`. It is built once in `Bridged.java:93` and handed to the
|
||
launcher, so **it is a restart-required key**, not a hot one. Changing `baseUrl` without changing
|
||
this makes every `local` spawn throw.
|
||
3. **The wiki names us as a blocker.** Under *Not done yet*: retiring the shared `legacy` token is
|
||
blocked because "kb, brain, **claude-bridge** and the workstation still share it. Each needs its
|
||
own consumer first."
|
||
|
||
Context7 is mounted twice today, both times straight at `https://ct7.ltms.dev/mcp` — once in
|
||
`.mcp.json` (the primary) and once in `opencode.json` (the `sol` and `terra` members).
|
||
|
||
```mermaid
|
||
flowchart LR
|
||
subgraph now["Today"]
|
||
M1["local member<br/>claude-code"] -->|"ANTHROPIC_BASE_URL"| V1["vLLM gx00.gw:8000<br/>no auth, LAN only"]
|
||
M2["sol / terra<br/>opencode"] --> CT1["ct7.ltms.dev/mcp"]
|
||
P1["primary"] --> CT1
|
||
end
|
||
subgraph after["Proposed"]
|
||
M3["local member"] -->|"ANTHROPIC_BASE_URL<br/>+ ANTHROPIC_AUTH_TOKEN"| G["llm.ltms.dev/anthropic<br/>consumer: claude-bridge"]
|
||
G --> V2["vLLM gx00.gw:8000"]
|
||
M4["local-direct<br/>weight 0, escape hatch"] --> V2
|
||
end
|
||
```
|
||
|
||
*The member definition is the only thing that moves. The model behind it does not.*
|
||
|
||
---
|
||
|
||
## 3. The member definition change
|
||
|
||
The gateway serves an Anthropic surface *and* an OpenAI surface, so **both member kinds can point at
|
||
it**. That is the main opportunity here, and it is bigger than the `local` profile alone.
|
||
|
||
### 3a. `local` — claude-code, on `/anthropic`
|
||
|
||
| Key | Today | After | Note |
|
||
|---|---|---|---|
|
||
| `baseUrl` | `http://gx00.gw:8000` | `https://llm.ltms.dev/anthropic` | see the schema warning below |
|
||
| `tokenEnv` | *(unset)* | `AI_GATEWAY_TOKEN` | new consumer token, `llmk-claude-bridge-<32 hex>` |
|
||
| `model` | `deepseek-v4-flash` | unchanged | must stay **exact**; a regex match returns an empty `/v1/models` while completions keep working |
|
||
| `guard.offSubscriptionHosts` | `[gx00.gw]` | `[gx00.gw, llm.ltms.dev]` | **restart required** |
|
||
|
||
### 3b. A new opencode profile on `/v1` — no code needed
|
||
|
||
`OpenCodeLauncher` already supports a pinned OpenAI-compatible endpoint (CB-508). Given `baseUrl` it
|
||
writes a custom provider block into the worker's opencode config:
|
||
|
||
- `baseUrl` → `options.baseURL`. `openAiBaseUrl` appends `/v1` to a bare host, and takes a URL that
|
||
already has a path **as-is** — so `https://llm.ltms.dev/v1` works unchanged.
|
||
- `tokenEnv` → `options.apiKey` (falls back to a placeholder when unset, since a local vLLM ignores it).
|
||
- `model:` **must** be `<provider>/<model>` when `baseUrl` is set — a bare name is rejected loudly
|
||
rather than silently falling back to opencode's default gateway.
|
||
|
||
So the profile is pure config:
|
||
|
||
```yaml
|
||
gx:
|
||
kind: opencode
|
||
baseUrl: https://llm.ltms.dev/v1
|
||
tokenEnv: AI_GATEWAY_TOKEN
|
||
model: gx/deepseek-v4-flash # provider id is ours to choose; the half after / is the model
|
||
argv: ["opencode"]
|
||
mcpUrl: http://127.0.0.1:8765/mcp
|
||
gitTokenEnv: WORKER_GITEA_TOKEN
|
||
weight: 100 # same tier as `local` — free
|
||
maxLoad: 2
|
||
# deliberately NO credentialId — this is our own box, not the shared OpenAI account
|
||
```
|
||
|
||
**Why this matters more than it looks.** Today every opencode member is `sol` or `terra`, and those
|
||
are two models on **one** OpenAI account sharing `credentialId: openai-shared` — so an exhaustion on
|
||
either locks out both, and half the fleet's opencode capacity dies at once. A gateway-backed opencode
|
||
profile is free, is not on that credential, and therefore is not in that quarantine pair. It removes
|
||
a single point of failure rather than just adding capacity.
|
||
|
||
**Note the asymmetry, it is deliberate:** `SubscriptionGuard` does not apply to opencode at all — the
|
||
guard exists to stop a *Claude* worker borrowing the operator's subscription, and opencode reads its
|
||
own provider credentials. So 3b needs **no allowlist change**; only 3a does.
|
||
|
||
**Both still need a restart, for a different reason.** `tokenEnv` is resolved by
|
||
`HerdrPeerLauncher.resolveEnv` → `env.apply(name)`, which reads the **daemon's own process
|
||
environment**. The running `bridged` inherited its environment when it started, so a variable added to
|
||
`secrets.sh` afterwards is simply not there — the launcher would inject an empty token and the
|
||
gateway would answer 401. This is the same failure as trap 1 in `scripts/redeploy-bridged.sh`
|
||
(`WORKER_GITEA_TOKEN`), and it has the same fix: **restart from a login shell**, and use
|
||
`scripts/redeploy-bridged.sh --check` to confirm the name resolves before restarting anything.
|
||
|
||
### 3c. What this does to `ccs`
|
||
|
||
Once a profile carries `baseUrl`, `tokenEnv` and `model` itself, the ccs instance stops being what
|
||
routes a member. Be precise about what is left, though: `configDir` still supplies **folder trust**
|
||
and `settings.json`, and dropping it is what produced the trust dialog and the wrong-model error
|
||
recorded in `bridged.yaml`. So ccs goes from *deciding where the tokens go* to *holding client-side
|
||
state*. Less load-bearing, not removable.
|
||
|
||
### Why `/anthropic` and never `/v1/chat/completions`
|
||
|
||
The gateway declares its Anthropic backend as `schema.name: Anthropic`, which means **no
|
||
translation** — streaming, tool use and thinking blocks pass through exactly as they do against vLLM
|
||
directly.
|
||
|
||
Declared as `OpenAI`, Envoy's translator looks for a `thinking_blocks` field that our vLLM does not
|
||
send (it sends `reasoning_content`), and **every thinking delta disappears silently**. Claude Code
|
||
speaks the Anthropic protocol, so `/anthropic` is both correct and the only safe choice.
|
||
|
||
This is the exact failure shape this repo keeps hitting: it compiles, it answers, it looks healthy,
|
||
and a capability is quietly off. Treat it as a `silent-default` risk, not a config preference.
|
||
|
||
**Open question for 3b — ANSWERED, 2026-08-15.** The worry was that the OpenAI surface might drop
|
||
reasoning the way the wiki documents for a mis-declared Anthropic backend. It does not. Checked at
|
||
the API before any profile was switched:
|
||
|
||
| surface | request | result |
|
||
|---|---|---|
|
||
| `/anthropic/v1/messages` | `deepseek-v4-flash`, 64 tokens | 200, response carries a real `"type":"thinking"` block |
|
||
| `/v1/chat/completions` | same | 200, message carries a populated `reasoning_content` (and a `reasoning` field) |
|
||
| `/v1/models` | — | 200, exactly `["deepseek-v4-flash"]` — the exact-name trap is clear |
|
||
| `/v1/models`, **no token** | — | **401** — Caddy is gating, as designed |
|
||
|
||
So reasoning survives on **both** surfaces, and the `/anthropic` choice for `local` is about protocol
|
||
correctness rather than a repair for a known loss. The last row matters on its own: the wiki warns
|
||
the gateway's own `SecurityPolicy` fails open, so it is worth knowing the proxy in front really does
|
||
refuse an unauthenticated request here.
|
||
|
||
---
|
||
|
||
## 4. Decisions
|
||
|
||
### D1 — switch, but keep the direct path as an explicit profile · **recommended**
|
||
|
||
Switching buys four things we do not have:
|
||
|
||
- **Free opencode capacity, off the shared credential.** The largest single win. See §3b — it retires
|
||
a real single point of failure, not just a cost line.
|
||
- **Per-consumer usage figures.** The cockpit counts requests per consumer. That is the first real
|
||
measurement of what the fleet consumes, and it feeds [CB-589](https://git.ltms.dev/fleet/fleetd/issues/74) Gap 2 directly.
|
||
- **Our own revocable token.** One consumer to revoke if a worker ever leaks it, instead of a shared
|
||
`legacy` token used by four systems.
|
||
- **It works off-LAN.** `gx00.gw` resolves on the LAN only.
|
||
|
||
The cost is honest and worth stating: we add a TLS edge, an auth proxy and a gateway to the path of
|
||
every member spawn. The wiki keeps the direct route open precisely because "if the gateway breaks,
|
||
nothing that matters is blocked."
|
||
|
||
So keep it. Add a second profile `local-direct` pointing at `http://gx00.gw:8000` with **`weight: 0`**
|
||
— never auto-selected, still spawnable with an explicit `fleet_spawn{profile: "local-direct"}`.
|
||
That is exactly what CB-554 made `weight: 0` mean, and it turns the escape hatch into something the
|
||
lead can actually reach during an incident.
|
||
|
||
### D2 — do members also mount the gateway's `/mcp`? · **OPEN, operator's call**
|
||
|
||
Not a detail. `CLAUDE.md` states in two places that a member mounts **only** the bridge MCP, and a
|
||
worker's honesty rule leans on it ("never claim the result of a check you had no way to run").
|
||
|
||
- **Keep bridge-only.** The invariant stays true and simple. Workers stay cheap and narrow.
|
||
- **Add the gateway MCP.** Implementers get context7 documentation lookups, which is genuinely useful
|
||
for library work. But `mcpUrl` in `BridgedConfig.Profile` is a **single `String`**, so a
|
||
claude-code member can mount exactly one MCP — this needs a code change, not a config edit.
|
||
|
||
Note the invariant is **already inaccurate**: `opencode.json` gives `sol` and `terra` both context7
|
||
and gitea. So the choice is really "make the rule true" or "make the rule match reality". Either is
|
||
defensible; picking one is not mine to do.
|
||
|
||
### D3 — token scope
|
||
|
||
One consumer, `claude-bridge`, its token in `${SHARED_ENV}/tools/secrets.sh` as `AI_GATEWAY_TOKEN`,
|
||
referenced by name only. Never the literal value in `bridged.yaml` — `tokenEnv` exists for this.
|
||
|
||
---
|
||
|
||
## 5. Units of work
|
||
|
||
```mermaid
|
||
flowchart TB
|
||
U1["U1 · consumer token<br/>issue via cockpit, add to secrets.sh"]
|
||
U2["U2 · profile + guard<br/>bridged.yaml, restart"]
|
||
U3["U3 · verify live<br/>spawn, prove thinking survives"]
|
||
U4["U4 · context7 via gateway<br/>.mcp.json + opencode.json"]
|
||
U5["U5 · docs<br/>CLAUDE.md, wiki 11-Features"]
|
||
U1 --> U2 --> U3
|
||
U4 --> U5
|
||
U3 --> U5
|
||
```
|
||
|
||
| # | Scope | Who | Why |
|
||
|---|---|---|---|
|
||
| U1 | Issue the `claude-bridge` consumer at `auth.ltms.dev`; store as `AI_GATEWAY_TOKEN` | **operator** | touches secrets and a host we do not own |
|
||
| U2a | New `gx` opencode profile on `/v1` — pure config, no guard change | **lead** | `bridged.yaml` is gitignored, so a worker cannot see or edit it |
|
||
| U2b | `local` → `/anthropic`; add `local-direct` weight 0; add `llm.ltms.dev` to the guard allowlist | **lead** | same |
|
||
| U2c | One restart from a **login shell**, after U2a and U2b | **lead** | picks up `AI_GATEWAY_TOKEN` into the daemon env *and* the guard allowlist, in one stop |
|
||
| U3 | Live spawn on both new profiles; confirm reasoning survives on each surface | **lead** | needs real spawns and the running daemon |
|
||
| U4 | Point `.mcp.json` and `opencode.json` context7 at the gateway `/mcp`; rename pinned tools | delegatable | tracked files, self-contained |
|
||
| U5 | Fix the "members mount only the bridge" claim; add a `wiki/11-Features.md` entry | delegatable | writing, clear criteria |
|
||
|
||
U1 blocks U2a, U2b and U3. U4 and U5 do not depend on it.
|
||
|
||
**Write U2a and U2b, then restart once (U2c), then verify `gx` before `local`.** Since both profiles
|
||
need the same restart there is no reason to do two, but there is still a reason to *verify* in order:
|
||
`gx` exercises the token and the gateway with no guard involved, so if it fails the cause is upstream.
|
||
`local` adds the guard allowlist on top, so a failure there points at our config instead. Testing them
|
||
in that order separates the two causes instead of confusing them.
|
||
|
||
> **U1 status, 2026-08-15:** the operator issued the consumer and exported it as `AI_GATEWAY_TOKEN`
|
||
> (one key for every agent and MCP client behind `llm.ltms.dev`). Confirmed: it resolves in a login
|
||
> shell, is 48 characters and carries the documented `llmk-` prefix. The value was never printed.
|
||
|
||
---
|
||
|
||
## 6. Traps carried over from the wiki
|
||
|
||
Each of these cost someone real debugging time upstream. They apply to us.
|
||
|
||
1. **Rotating a token restarts the auth proxy, which drops in-flight streaming responses.** For us
|
||
that means rotating `AI_GATEWAY_TOKEN` kills every live member mid-turn, and an async ticket's
|
||
report goes with it. This is the same rule as a daemon redeploy: **drain the fleet first**
|
||
(`fleet_list` → `fleet_poll` anything wanted → `fleet_stop`), then rotate.
|
||
2. **The gateway's own `SecurityPolicy` fails open.** Standalone `aigw run` accepts it and silently
|
||
ignores it — an unauthenticated request returned **200**. Auth is the Caddy proxy in front, and
|
||
nothing else. Never reason as if the gateway authenticates.
|
||
3. **Exact model name.** A regex match routes fine but returns an **empty** `/v1/models` list while
|
||
completions keep working. A wrong name returns a bare 404 that reads exactly like a dead gateway.
|
||
4. **MCP tool names changed prefix separator.** Bifrost used one dash (`ct7-resolve-library-id`); the
|
||
gateway uses **two underscores** (`ct7__resolve-library-id`). Relevant only if U4 is done.
|
||
5. **`/v1/models` 404 vs empty list are different faults.** 404 means no route loaded at all; empty
|
||
means the model match is a regex. Do not conflate them when diagnosing.
|
||
|
||
---
|
||
|
||
## 7. Verification — what would prove this works
|
||
|
||
Merging config is not proving it. The checks, in order:
|
||
|
||
1. `fleet_spawn{profile: "gx"}` succeeds and the member completes a real turn ending in
|
||
`fleet_reply`. This is the first proof of the token, the URL and the model name, and it risks
|
||
nothing the fleet depends on.
|
||
2. `fleet_spawn{profile: "local"}` succeeds. If the guard allowlist was missed, this **throws** — a
|
||
loud, self-correcting failure, which is the good kind. If the restart was missed, it also throws,
|
||
for the same reason.
|
||
3. A `local` member completes a turn. That exercises streaming through two TLS edges, the auth proxy
|
||
and the gateway.
|
||
4. **Reasoning survives, checked separately on each surface.** For `local` on `/anthropic` this is
|
||
the check that catches the `/v1` versus `/anthropic` mistake, and it is the only one that does —
|
||
nothing else distinguishes a working passthrough from a translator quietly dropping thinking
|
||
deltas. For `gx` on `/v1`, this answers the open question in §3 rather than assuming it.
|
||
5. The cockpit at `auth.ltms.dev` shows requests counted against the `claude-bridge` consumer, not
|
||
`legacy`. That is the whole point of taking our own token.
|
||
6. `fleet_spawn{profile: "local-direct"}` still works, so the escape hatch is real rather than
|
||
theoretical.
|
||
7. `fleet_list` shows `gx` carrying no `credentialId`, so a `sol`/`terra` exhaustion cannot
|
||
quarantine it. This is the single-point-of-failure claim in §3b, checked rather than asserted.
|
||
|
||
---
|
||
|
||
## 7.1 What the live run actually found — 2026-08-15
|
||
|
||
U1–U2c were done, the daemon restarted onto them, and both new profiles were spawned for real. The
|
||
migration was then **reverted**. This section is the result, so none of it has to be re-derived.
|
||
|
||
### The blocker
|
||
|
||
`llm.ltms.dev` answers **HTTP 413 Request Entity Too Large** above **32 KiB (32768 bytes)**, on both
|
||
surfaces:
|
||
|
||
```
|
||
/v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200
|
||
/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413
|
||
```
|
||
|
||
32 KiB is far below one real agent turn.
|
||
|
||
**Root cause — confirmed by the systems/vms side, 2026-08-15.** My guess that it was a Caddy
|
||
`request_body max_size` was **wrong**. It is Envoy, inside `aigw` on `llm.vm`. Envoy Gateway defaults
|
||
a listener's `per_connection_buffer_limit_bytes` to **32768**, and the AI Gateway buffers the *whole*
|
||
request body before it can route on the model name — so that default is not a network tuning knob
|
||
here, it is a hard ceiling on prompt size. Read out of the live Envoy `config_dump`:
|
||
|
||
```
|
||
listener default/llm/http per_connection_buffer_limit_bytes: 32768
|
||
```
|
||
|
||
Nobody chose 32 KiB; it was inherited from the default. Both TLS edges are innocent: the same
|
||
boundary reproduces on the LAN path and the internet path, and both 413s carry an `x-llm-consumer`
|
||
header their auth proxy sets only *after* authenticating — so the body cleared both edges and the
|
||
auth. Directly on `llm.vm`, `aigw` 413s at 39 KB while the vLLM backend accepts the same 39 KB and
|
||
answers 200.
|
||
|
||
**Do not plan around 32 KiB.** The intended ceiling is far higher. Their fix — a `ClientTrafficPolicy`
|
||
setting `bufferLimit: 8Mi` — is written but **not deployed** as of this note, pending their operator's
|
||
approval. I have not re-tested and will not until they confirm, so as not to measure a half-changed
|
||
system. Fixed in **systems/vms**, not here.
|
||
|
||
### The part worth remembering
|
||
|
||
Two members were spawned at the same moment with the same message:
|
||
|
||
| | `local` (claude-code, `/anthropic`) | `gx` (opencode, `/v1`) |
|
||
|---|---|---|
|
||
| READY → BUSY | 19:07:26 | 19:07:45 |
|
||
| BUSY → DONE | **19:08:51 (66s)** | **never — 10+ min, ticket FAILED** |
|
||
|
||
**`local` passed.** It passed only because the probe was three trivial questions in a fresh session,
|
||
so the request fit under 32 KiB. The profile looked healthy and was a landmine set to fire on the
|
||
first turn that reads a file.
|
||
|
||
So §7's checklist was not wrong, it was **too easy**. Any future run of it must use a task that reads
|
||
a real file. A liveness probe proves the token and the URL; it does not prove the path.
|
||
|
||
`gx` did not fail loudly either. Reproduced outside the bridge by running `opencode` by hand with the
|
||
launcher's own generated config:
|
||
|
||
```
|
||
Error: Request Entity Too Large
|
||
...compacts context, retries...
|
||
Error: Request Entity Too Large
|
||
```
|
||
|
||
opencode **catches the 413, compacts, and retries — indefinitely**. A member that fails loudly costs
|
||
one turn; this one costs the whole task and is indistinguishable from a slow worker.
|
||
|
||
> **Diagnosing a stuck opencode member.** Do not read its pane. The launcher writes its config to a
|
||
> temp dir and passes it as `OPENCODE_CONFIG` — find it with
|
||
> `ls -dt /var/folders/*/*/T/bridged-opencode-* | head -1`, check the provider block and the key's
|
||
> length and prefix (never its value), then reproduce with `opencode run --auto -m <provider>/<model>`
|
||
> using the same `OPENCODE_CONFIG`. That is what turned "it hangs" into a one-line error.
|
||
|
||
### What checked out, and needs no re-testing
|
||
|
||
- Token accepted on both surfaces. **Unauthenticated → 401**, so the Caddy proxy really does gate —
|
||
the wiki's "SecurityPolicy fails open" warning is about the gateway itself, not the edge.
|
||
- `/v1/models` returns exactly `["deepseek-v4-flash"]`, so trap 3 is clear.
|
||
- **Reasoning survives both surfaces** — see §3b above.
|
||
- The launcher's generated opencode provider block is correct, carrying a real 48-character `llmk-`
|
||
key rather than the `bridged-local-noauth` placeholder.
|
||
- `SubscriptionGuard` accepted `llm.ltms.dev` after the allowlist edit and the restart: `local`
|
||
spawned without throwing, which is the check that catches a missed restart.
|
||
|
||
### Resolution — both ceilings fixed, migration completed
|
||
|
||
systems/vms fixed both, and each was re-checked from this side rather than taken on trust:
|
||
|
||
| ceiling | was | now | our own check |
|
||
|---|---|---|---|
|
||
| listener buffer | 32 KiB | 32 Mi | 1.2 MB body → **200** (was 413) |
|
||
| LLM route timeout | 60s | 86400s | the request that truncated: **101s, `message_stop` present, 4000/4000** |
|
||
|
||
The timeout moved in two steps on 2026-08-15: 60s → 1800s, then 1800s → **86400s (24 hours)** after
|
||
the truncation risk below was discussed. They tried `request: 0s` first, which removes the
|
||
total-duration timer completely. It works, but on an `AIGatewayRoute` the **idle timeout is derived
|
||
from the request timeout**, so `0s` also removed any bound on a stalled connection. 86400s keeps a
|
||
reaper for dead connections while putting the truncation timer out of practical reach.
|
||
|
||
Neither was deliberate. The 32 KiB was Envoy Gateway's default `per_connection_buffer_limit_bytes`;
|
||
the 60s was Envoy AI Gateway's own documented default. The 60s bounded **generation** as well as
|
||
prompt size — a tiny prompt with a long answer returned 504 at 60.05s.
|
||
|
||
Two configuration facts worth keeping, from their bisection:
|
||
|
||
- **`ClientTrafficPolicy` is honoured in standalone `aigw run`; `BackendTrafficPolicy` is NOT.** A
|
||
`BackendTrafficPolicy` setting `requestTimeout` is accepted, logs nothing, and leaves the routes
|
||
unchanged (upstream `envoyproxy/gateway#9513`). What works is `timeouts: {request: …}` on each
|
||
`AIGatewayRoute` rule. Nothing from the outside distinguishes the two — the same silent-default
|
||
shape as their `SecurityPolicy` caveat.
|
||
- In that stack, "the config was accepted" proves nothing. Read the live `config_dump`.
|
||
|
||
## 7.2 The risk we accepted, and why we could not remove it
|
||
|
||
Raising the timeout made the failure **rare, not impossible**, and the residual failure is silent.
|
||
|
||
On a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked encoding *cleanly* instead of
|
||
resetting the connection, so the client receives what looks like a complete transfer
|
||
(`envoyproxy/envoy#17186` — acknowledged as a bug in 2021, closed by a stale bot, never fixed). The
|
||
December 2025 fix `envoyproxy/envoy#42269` changes locally-originated resets from `NO_ERROR` to
|
||
`INTERNAL_ERROR`, but it is **HTTP/2 only** and SSE clients here speak HTTP/1.1.
|
||
|
||
Measured on our side while the timeout was still 60s:
|
||
|
||
```
|
||
HTTP 200 61.07s 141992 bytes
|
||
message_stop 0 message_delta 0 error events 0
|
||
emitted 2473 of 4000, ending on a WELL-FORMED SSE frame
|
||
```
|
||
|
||
A syntactically valid stream that simply stops. Any timer firing mid-stream — route timeout, idle
|
||
timeout, `max_stream_duration` — fails this same way.
|
||
|
||
**The recommended defence does not transfer to us.** The right fix is to treat a stream with no
|
||
`message_stop` / `[DONE]` / `finish_reason` as failed. We cannot: our members are Claude Code and
|
||
opencode, third-party clients whose SSE parsing we do not own, and there is no seam to insert the
|
||
check. Whether either detects a missing terminator is unverified — and opencode's handling of the 413
|
||
(swallow, compact, retry forever, never surface an error) does not suggest it is strict.
|
||
|
||
So the honest statement of our position:
|
||
|
||
> Gateway traffic is acceptable at 86400s because a single request would have to run for 24 hours to
|
||
> trip the bug — **not** because we could detect it if it did.
|
||
|
||
At 86400s our **own** limit binds first, which is the ordering we want. `MessageService.ASYNC_TIMEOUT_MS`
|
||
caps a turn at 30 minutes, so a runaway request ends as a clean `FAILED` ticket that we raised, rather
|
||
than as a silently truncated `200` that we cannot see. While the gateway sat at 1800s the two numbers
|
||
were equal and did not nest, so a gateway-side stall could have been misread as a bug in our own ticket
|
||
handling. That ambiguity is now gone.
|
||
|
||
**If a member ever returns a confident but truncated answer, suspect this before anything in our own
|
||
code.** That is the whole reason this section exists.
|
||
|
||
---
|
||
|
||
## 8. Related
|
||
|
||
- [CB-589 / #74](https://git.ltms.dev/fleet/fleetd/issues/74) — cost-first placement and a
|
||
gateway that reports live capacity. The per-consumer figures this migration unlocks are the first
|
||
input that ticket actually needs.
|
||
- `docs/CB-500-Multi-Tier-Coordination.md` §11 — the distributed-sandbox topology this gateway is
|
||
part of.
|