2e138a199b
Two changes ship together here.
1. One shared herdr workspace. The lead and every worker now live in one
workspace called "fleet", so the operator sees one "session" with many
windows, not two. Before, the lead sat in a "leads" workspace and workers
in "bridged-workers", which read as two sessions. The lead is still told
apart from workers by its exact tab label ("lead: <name>"), so putting them
in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
for split layouts; Fleetd now passes an empty exclude set.
2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
launchd/systemd units, module dir, and MCP mount).
- Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
- Log line, comments, docs, and CLAUDE.md updated to say fleetd.
- Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
- Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
- Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
still read as a fallback, and still gitignored.
- MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
server name is "fleet". The mount name in the local .mcp.json becomes
"fleet" (gitignored, not in this commit).
- Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.
Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.
Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.
The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).
949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
468 lines
25 KiB
Markdown
468 lines
25 KiB
Markdown
# CB-591 — move the fleet onto the LLM and MCP gateway
|
||
|
||
**Status: DONE — the fleet is on the gateway as of 2026-08-15.** `local` runs on `/anthropic` and
|
||
`gx` on `/v1`, both at `weight: 100`; `local-direct` stays at `weight: 0` as the escape hatch. Getting
|
||
here took a revert and two upstream fixes — see §7.1, which is the useful part of this document. One
|
||
risk is **accepted rather than solved**: a stream cut by any mid-response timer arrives as HTTP 200
|
||
with no terminator, and our third-party members cannot detect it (§7.2).
|
||
· **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway)
|
||
· **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31)
|
||
|
||
The gateway went live on 2026-08-15 and replaced Bifrost. This plan says what that means for a
|
||
**member definition** in `fleetd.yaml`, because that is the part of this repo the change actually
|
||
touches.
|
||
|
||
---
|
||
|
||
## 1. What changed upstream
|
||
|
||
One front door for every LLM and MCP client: `https://llm.ltms.dev`, one token per consumer.
|
||
|
||
| Surface | URL |
|
||
|---|---|
|
||
| OpenAI chat | `https://llm.ltms.dev/v1/chat/completions` |
|
||
| OpenAI models | `https://llm.ltms.dev/v1/models` |
|
||
| **Anthropic messages** | `https://llm.ltms.dev/anthropic/v1/messages` |
|
||
| MCP, all servers multiplexed | `https://llm.ltms.dev/mcp` |
|
||
|
||
Anything outside that list returns **404 before any token is checked**, on purpose — the gateway must
|
||
never become a blanket proxy.
|
||
|
||
The model backend is unchanged: GX10 vLLM at `10.10.10.26:8000` (`gx00.gw`), model name exactly
|
||
`deepseek-v4-flash`. The direct LAN path stays open on purpose as an escape hatch.
|
||
|
||
---
|
||
|
||
## 2. Where claude-bridge sits today
|
||
|
||
We do **not** use the gateway. The `local` profile talks straight to the vLLM:
|
||
|
||
```yaml
|
||
local:
|
||
kind: claude-code
|
||
baseUrl: http://gx00.gw:8000 # direct vLLM — no auth, LAN only
|
||
model: deepseek-v4-flash
|
||
configDir: /Users/dai.ha/.ccs/instances/gx10
|
||
```
|
||
|
||
Three facts about our side that decide the shape of this work:
|
||
|
||
1. **`baseUrl` becomes `ANTHROPIC_BASE_URL`** in the member's environment, and `tokenEnv` becomes
|
||
`ANTHROPIC_AUTH_TOKEN` (the value is read from a host env var and never stored in config).
|
||
`local` sets no `tokenEnv` today, because a direct vLLM needs no token.
|
||
2. **`SubscriptionGuard` refuses any host not on an allowlist**, and that allowlist is
|
||
`guard.offSubscriptionHosts: [gx00.gw]`. It is built once in `Fleetd.java:93` and handed to the
|
||
launcher, so **it is a restart-required key**, not a hot one. Changing `baseUrl` without changing
|
||
this makes every `local` spawn throw.
|
||
3. **The wiki names us as a blocker.** Under *Not done yet*: retiring the shared `legacy` token is
|
||
blocked because "kb, brain, **claude-bridge** and the workstation still share it. Each needs its
|
||
own consumer first."
|
||
|
||
Context7 is mounted twice today, both times straight at `https://ct7.ltms.dev/mcp` — once in
|
||
`.mcp.json` (the primary) and once in `opencode.json` (the `sol` and `terra` members).
|
||
|
||
```mermaid
|
||
flowchart LR
|
||
subgraph now["Today"]
|
||
M1["local member<br/>claude-code"] -->|"ANTHROPIC_BASE_URL"| V1["vLLM gx00.gw:8000<br/>no auth, LAN only"]
|
||
M2["sol / terra<br/>opencode"] --> CT1["ct7.ltms.dev/mcp"]
|
||
P1["primary"] --> CT1
|
||
end
|
||
subgraph after["Proposed"]
|
||
M3["local member"] -->|"ANTHROPIC_BASE_URL<br/>+ ANTHROPIC_AUTH_TOKEN"| G["llm.ltms.dev/anthropic<br/>consumer: claude-bridge"]
|
||
G --> V2["vLLM gx00.gw:8000"]
|
||
M4["local-direct<br/>weight 0, escape hatch"] --> V2
|
||
end
|
||
```
|
||
|
||
*The member definition is the only thing that moves. The model behind it does not.*
|
||
|
||
---
|
||
|
||
## 3. The member definition change
|
||
|
||
The gateway serves an Anthropic surface *and* an OpenAI surface, so **both member kinds can point at
|
||
it**. That is the main opportunity here, and it is bigger than the `local` profile alone.
|
||
|
||
### 3a. `local` — claude-code, on `/anthropic`
|
||
|
||
| Key | Today | After | Note |
|
||
|---|---|---|---|
|
||
| `baseUrl` | `http://gx00.gw:8000` | `https://llm.ltms.dev/anthropic` | see the schema warning below |
|
||
| `tokenEnv` | *(unset)* | `AI_GATEWAY_TOKEN` | new consumer token, `llmk-claude-bridge-<32 hex>` |
|
||
| `model` | `deepseek-v4-flash` | unchanged | must stay **exact**; a regex match returns an empty `/v1/models` while completions keep working |
|
||
| `guard.offSubscriptionHosts` | `[gx00.gw]` | `[gx00.gw, llm.ltms.dev]` | **restart required** |
|
||
|
||
### 3b. A new opencode profile on `/v1` — no code needed
|
||
|
||
`OpenCodeLauncher` already supports a pinned OpenAI-compatible endpoint (CB-508). Given `baseUrl` it
|
||
writes a custom provider block into the worker's opencode config:
|
||
|
||
- `baseUrl` → `options.baseURL`. `openAiBaseUrl` appends `/v1` to a bare host, and takes a URL that
|
||
already has a path **as-is** — so `https://llm.ltms.dev/v1` works unchanged.
|
||
- `tokenEnv` → `options.apiKey` (falls back to a placeholder when unset, since a local vLLM ignores it).
|
||
- `model:` **must** be `<provider>/<model>` when `baseUrl` is set — a bare name is rejected loudly
|
||
rather than silently falling back to opencode's default gateway.
|
||
|
||
So the profile is pure config:
|
||
|
||
```yaml
|
||
gx:
|
||
kind: opencode
|
||
baseUrl: https://llm.ltms.dev/v1
|
||
tokenEnv: AI_GATEWAY_TOKEN
|
||
model: gx/deepseek-v4-flash # provider id is ours to choose; the half after / is the model
|
||
argv: ["opencode"]
|
||
mcpUrl: http://127.0.0.1:8765/mcp
|
||
gitTokenEnv: WORKER_GITEA_TOKEN
|
||
weight: 100 # same tier as `local` — free
|
||
maxLoad: 2
|
||
# deliberately NO credentialId — this is our own box, not the shared OpenAI account
|
||
```
|
||
|
||
**Why this matters more than it looks.** Today every opencode member is `sol` or `terra`, and those
|
||
are two models on **one** OpenAI account sharing `credentialId: openai-shared` — so an exhaustion on
|
||
either locks out both, and half the fleet's opencode capacity dies at once. A gateway-backed opencode
|
||
profile is free, is not on that credential, and therefore is not in that quarantine pair. It removes
|
||
a single point of failure rather than just adding capacity.
|
||
|
||
**Note the asymmetry, it is deliberate:** `SubscriptionGuard` does not apply to opencode at all — the
|
||
guard exists to stop a *Claude* worker borrowing the operator's subscription, and opencode reads its
|
||
own provider credentials. So 3b needs **no allowlist change**; only 3a does.
|
||
|
||
**Both still need a restart, for a different reason.** `tokenEnv` is resolved by
|
||
`HerdrPeerLauncher.resolveEnv` → `env.apply(name)`, which reads the **daemon's own process
|
||
environment**. The running `fleetd` inherited its environment when it started, so a variable added to
|
||
`secrets.sh` afterwards is simply not there — the launcher would inject an empty token and the
|
||
gateway would answer 401. This is the same failure as trap 1 in `scripts/redeploy-fleetd.sh`
|
||
(`WORKER_GITEA_TOKEN`), and it has the same fix: **restart from a login shell**, and use
|
||
`scripts/redeploy-fleetd.sh --check` to confirm the name resolves before restarting anything.
|
||
|
||
### 3c. What this does to `ccs`
|
||
|
||
Once a profile carries `baseUrl`, `tokenEnv` and `model` itself, the ccs instance stops being what
|
||
routes a member. Be precise about what is left, though: `configDir` still supplies **folder trust**
|
||
and `settings.json`, and dropping it is what produced the trust dialog and the wrong-model error
|
||
recorded in `fleetd.yaml`. So ccs goes from *deciding where the tokens go* to *holding client-side
|
||
state*. Less load-bearing, not removable.
|
||
|
||
### Why `/anthropic` and never `/v1/chat/completions`
|
||
|
||
The gateway declares its Anthropic backend as `schema.name: Anthropic`, which means **no
|
||
translation** — streaming, tool use and thinking blocks pass through exactly as they do against vLLM
|
||
directly.
|
||
|
||
Declared as `OpenAI`, Envoy's translator looks for a `thinking_blocks` field that our vLLM does not
|
||
send (it sends `reasoning_content`), and **every thinking delta disappears silently**. Claude Code
|
||
speaks the Anthropic protocol, so `/anthropic` is both correct and the only safe choice.
|
||
|
||
This is the exact failure shape this repo keeps hitting: it compiles, it answers, it looks healthy,
|
||
and a capability is quietly off. Treat it as a `silent-default` risk, not a config preference.
|
||
|
||
**Open question for 3b — ANSWERED, 2026-08-15.** The worry was that the OpenAI surface might drop
|
||
reasoning the way the wiki documents for a mis-declared Anthropic backend. It does not. Checked at
|
||
the API before any profile was switched:
|
||
|
||
| surface | request | result |
|
||
|---|---|---|
|
||
| `/anthropic/v1/messages` | `deepseek-v4-flash`, 64 tokens | 200, response carries a real `"type":"thinking"` block |
|
||
| `/v1/chat/completions` | same | 200, message carries a populated `reasoning_content` (and a `reasoning` field) |
|
||
| `/v1/models` | — | 200, exactly `["deepseek-v4-flash"]` — the exact-name trap is clear |
|
||
| `/v1/models`, **no token** | — | **401** — Caddy is gating, as designed |
|
||
|
||
So reasoning survives on **both** surfaces, and the `/anthropic` choice for `local` is about protocol
|
||
correctness rather than a repair for a known loss. The last row matters on its own: the wiki warns
|
||
the gateway's own `SecurityPolicy` fails open, so it is worth knowing the proxy in front really does
|
||
refuse an unauthenticated request here.
|
||
|
||
---
|
||
|
||
## 4. Decisions
|
||
|
||
### D1 — switch, but keep the direct path as an explicit profile · **recommended**
|
||
|
||
Switching buys four things we do not have:
|
||
|
||
- **Free opencode capacity, off the shared credential.** The largest single win. See §3b — it retires
|
||
a real single point of failure, not just a cost line.
|
||
- **Per-consumer usage figures.** The cockpit counts requests per consumer. That is the first real
|
||
measurement of what the fleet consumes, and it feeds [CB-589](https://git.ltms.dev/fleet/fleetd/issues/74) Gap 2 directly.
|
||
- **Our own revocable token.** One consumer to revoke if a worker ever leaks it, instead of a shared
|
||
`legacy` token used by four systems.
|
||
- **It works off-LAN.** `gx00.gw` resolves on the LAN only.
|
||
|
||
The cost is honest and worth stating: we add a TLS edge, an auth proxy and a gateway to the path of
|
||
every member spawn. The wiki keeps the direct route open precisely because "if the gateway breaks,
|
||
nothing that matters is blocked."
|
||
|
||
So keep it. Add a second profile `local-direct` pointing at `http://gx00.gw:8000` with **`weight: 0`**
|
||
— never auto-selected, still spawnable with an explicit `fleet_spawn{profile: "local-direct"}`.
|
||
That is exactly what CB-554 made `weight: 0` mean, and it turns the escape hatch into something the
|
||
lead can actually reach during an incident.
|
||
|
||
### D2 — do members also mount the gateway's `/mcp`? · **OPEN, operator's call**
|
||
|
||
Not a detail. `CLAUDE.md` states in two places that a member mounts **only** the bridge MCP, and a
|
||
worker's honesty rule leans on it ("never claim the result of a check you had no way to run").
|
||
|
||
- **Keep bridge-only.** The invariant stays true and simple. Workers stay cheap and narrow.
|
||
- **Add the gateway MCP.** Implementers get context7 documentation lookups, which is genuinely useful
|
||
for library work. But `mcpUrl` in `FleetConfig.Profile` is a **single `String`**, so a
|
||
claude-code member can mount exactly one MCP — this needs a code change, not a config edit.
|
||
|
||
Note the invariant is **already inaccurate**: `opencode.json` gives `sol` and `terra` both context7
|
||
and gitea. So the choice is really "make the rule true" or "make the rule match reality". Either is
|
||
defensible; picking one is not mine to do.
|
||
|
||
### D3 — token scope
|
||
|
||
One consumer, `claude-bridge`, its token in `${SHARED_ENV}/tools/secrets.sh` as `AI_GATEWAY_TOKEN`,
|
||
referenced by name only. Never the literal value in `fleetd.yaml` — `tokenEnv` exists for this.
|
||
|
||
---
|
||
|
||
## 5. Units of work
|
||
|
||
```mermaid
|
||
flowchart TB
|
||
U1["U1 · consumer token<br/>issue via cockpit, add to secrets.sh"]
|
||
U2["U2 · profile + guard<br/>fleetd.yaml, restart"]
|
||
U3["U3 · verify live<br/>spawn, prove thinking survives"]
|
||
U4["U4 · context7 via gateway<br/>.mcp.json + opencode.json"]
|
||
U5["U5 · docs<br/>CLAUDE.md, wiki 11-Features"]
|
||
U1 --> U2 --> U3
|
||
U4 --> U5
|
||
U3 --> U5
|
||
```
|
||
|
||
| # | Scope | Who | Why |
|
||
|---|---|---|---|
|
||
| U1 | Issue the `claude-bridge` consumer at `auth.ltms.dev`; store as `AI_GATEWAY_TOKEN` | **operator** | touches secrets and a host we do not own |
|
||
| U2a | New `gx` opencode profile on `/v1` — pure config, no guard change | **lead** | `fleetd.yaml` is gitignored, so a worker cannot see or edit it |
|
||
| U2b | `local` → `/anthropic`; add `local-direct` weight 0; add `llm.ltms.dev` to the guard allowlist | **lead** | same |
|
||
| U2c | One restart from a **login shell**, after U2a and U2b | **lead** | picks up `AI_GATEWAY_TOKEN` into the daemon env *and* the guard allowlist, in one stop |
|
||
| U3 | Live spawn on both new profiles; confirm reasoning survives on each surface | **lead** | needs real spawns and the running daemon |
|
||
| U4 | Point `.mcp.json` and `opencode.json` context7 at the gateway `/mcp`; rename pinned tools | delegatable | tracked files, self-contained |
|
||
| U5 | Fix the "members mount only the bridge" claim; add a `wiki/11-Features.md` entry | delegatable | writing, clear criteria |
|
||
|
||
U1 blocks U2a, U2b and U3. U4 and U5 do not depend on it.
|
||
|
||
**Write U2a and U2b, then restart once (U2c), then verify `gx` before `local`.** Since both profiles
|
||
need the same restart there is no reason to do two, but there is still a reason to *verify* in order:
|
||
`gx` exercises the token and the gateway with no guard involved, so if it fails the cause is upstream.
|
||
`local` adds the guard allowlist on top, so a failure there points at our config instead. Testing them
|
||
in that order separates the two causes instead of confusing them.
|
||
|
||
> **U1 status, 2026-08-15:** the operator issued the consumer and exported it as `AI_GATEWAY_TOKEN`
|
||
> (one key for every agent and MCP client behind `llm.ltms.dev`). Confirmed: it resolves in a login
|
||
> shell, is 48 characters and carries the documented `llmk-` prefix. The value was never printed.
|
||
|
||
---
|
||
|
||
## 6. Traps carried over from the wiki
|
||
|
||
Each of these cost someone real debugging time upstream. They apply to us.
|
||
|
||
1. **Rotating a token restarts the auth proxy, which drops in-flight streaming responses.** For us
|
||
that means rotating `AI_GATEWAY_TOKEN` kills every live member mid-turn, and an async ticket's
|
||
report goes with it. This is the same rule as a daemon redeploy: **drain the fleet first**
|
||
(`fleet_list` → `fleet_poll` anything wanted → `fleet_stop`), then rotate.
|
||
2. **The gateway's own `SecurityPolicy` fails open.** Standalone `aigw run` accepts it and silently
|
||
ignores it — an unauthenticated request returned **200**. Auth is the Caddy proxy in front, and
|
||
nothing else. Never reason as if the gateway authenticates.
|
||
3. **Exact model name.** A regex match routes fine but returns an **empty** `/v1/models` list while
|
||
completions keep working. A wrong name returns a bare 404 that reads exactly like a dead gateway.
|
||
4. **MCP tool names changed prefix separator.** Bifrost used one dash (`ct7-resolve-library-id`); the
|
||
gateway uses **two underscores** (`ct7__resolve-library-id`). Relevant only if U4 is done.
|
||
5. **`/v1/models` 404 vs empty list are different faults.** 404 means no route loaded at all; empty
|
||
means the model match is a regex. Do not conflate them when diagnosing.
|
||
|
||
---
|
||
|
||
## 7. Verification — what would prove this works
|
||
|
||
Merging config is not proving it. The checks, in order:
|
||
|
||
1. `fleet_spawn{profile: "gx"}` succeeds and the member completes a real turn ending in
|
||
`fleet_reply`. This is the first proof of the token, the URL and the model name, and it risks
|
||
nothing the fleet depends on.
|
||
2. `fleet_spawn{profile: "local"}` succeeds. If the guard allowlist was missed, this **throws** — a
|
||
loud, self-correcting failure, which is the good kind. If the restart was missed, it also throws,
|
||
for the same reason.
|
||
3. A `local` member completes a turn. That exercises streaming through two TLS edges, the auth proxy
|
||
and the gateway.
|
||
4. **Reasoning survives, checked separately on each surface.** For `local` on `/anthropic` this is
|
||
the check that catches the `/v1` versus `/anthropic` mistake, and it is the only one that does —
|
||
nothing else distinguishes a working passthrough from a translator quietly dropping thinking
|
||
deltas. For `gx` on `/v1`, this answers the open question in §3 rather than assuming it.
|
||
5. The cockpit at `auth.ltms.dev` shows requests counted against the `claude-bridge` consumer, not
|
||
`legacy`. That is the whole point of taking our own token.
|
||
6. `fleet_spawn{profile: "local-direct"}` still works, so the escape hatch is real rather than
|
||
theoretical.
|
||
7. `fleet_list` shows `gx` carrying no `credentialId`, so a `sol`/`terra` exhaustion cannot
|
||
quarantine it. This is the single-point-of-failure claim in §3b, checked rather than asserted.
|
||
|
||
---
|
||
|
||
## 7.1 What the live run actually found — 2026-08-15
|
||
|
||
U1–U2c were done, the daemon restarted onto them, and both new profiles were spawned for real. The
|
||
migration was then **reverted**. This section is the result, so none of it has to be re-derived.
|
||
|
||
### The blocker
|
||
|
||
`llm.ltms.dev` answers **HTTP 413 Request Entity Too Large** above **32 KiB (32768 bytes)**, on both
|
||
surfaces:
|
||
|
||
```
|
||
/v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200
|
||
/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413
|
||
```
|
||
|
||
32 KiB is far below one real agent turn.
|
||
|
||
**Root cause — confirmed by the systems/vms side, 2026-08-15.** My guess that it was a Caddy
|
||
`request_body max_size` was **wrong**. It is Envoy, inside `aigw` on `llm.vm`. Envoy Gateway defaults
|
||
a listener's `per_connection_buffer_limit_bytes` to **32768**, and the AI Gateway buffers the *whole*
|
||
request body before it can route on the model name — so that default is not a network tuning knob
|
||
here, it is a hard ceiling on prompt size. Read out of the live Envoy `config_dump`:
|
||
|
||
```
|
||
listener default/llm/http per_connection_buffer_limit_bytes: 32768
|
||
```
|
||
|
||
Nobody chose 32 KiB; it was inherited from the default. Both TLS edges are innocent: the same
|
||
boundary reproduces on the LAN path and the internet path, and both 413s carry an `x-llm-consumer`
|
||
header their auth proxy sets only *after* authenticating — so the body cleared both edges and the
|
||
auth. Directly on `llm.vm`, `aigw` 413s at 39 KB while the vLLM backend accepts the same 39 KB and
|
||
answers 200.
|
||
|
||
**Do not plan around 32 KiB.** The intended ceiling is far higher. Their fix — a `ClientTrafficPolicy`
|
||
setting `bufferLimit: 8Mi` — is written but **not deployed** as of this note, pending their operator's
|
||
approval. I have not re-tested and will not until they confirm, so as not to measure a half-changed
|
||
system. Fixed in **systems/vms**, not here.
|
||
|
||
### The part worth remembering
|
||
|
||
Two members were spawned at the same moment with the same message:
|
||
|
||
| | `local` (claude-code, `/anthropic`) | `gx` (opencode, `/v1`) |
|
||
|---|---|---|
|
||
| READY → BUSY | 19:07:26 | 19:07:45 |
|
||
| BUSY → DONE | **19:08:51 (66s)** | **never — 10+ min, ticket FAILED** |
|
||
|
||
**`local` passed.** It passed only because the probe was three trivial questions in a fresh session,
|
||
so the request fit under 32 KiB. The profile looked healthy and was a landmine set to fire on the
|
||
first turn that reads a file.
|
||
|
||
So §7's checklist was not wrong, it was **too easy**. Any future run of it must use a task that reads
|
||
a real file. A liveness probe proves the token and the URL; it does not prove the path.
|
||
|
||
`gx` did not fail loudly either. Reproduced outside the bridge by running `opencode` by hand with the
|
||
launcher's own generated config:
|
||
|
||
```
|
||
Error: Request Entity Too Large
|
||
...compacts context, retries...
|
||
Error: Request Entity Too Large
|
||
```
|
||
|
||
opencode **catches the 413, compacts, and retries — indefinitely**. A member that fails loudly costs
|
||
one turn; this one costs the whole task and is indistinguishable from a slow worker.
|
||
|
||
> **Diagnosing a stuck opencode member.** Do not read its pane. The launcher writes its config to a
|
||
> temp dir and passes it as `OPENCODE_CONFIG` — find it with
|
||
> `ls -dt /var/folders/*/*/T/fleetd-opencode-* | head -1`, check the provider block and the key's
|
||
> length and prefix (never its value), then reproduce with `opencode run --auto -m <provider>/<model>`
|
||
> using the same `OPENCODE_CONFIG`. That is what turned "it hangs" into a one-line error.
|
||
|
||
### What checked out, and needs no re-testing
|
||
|
||
- Token accepted on both surfaces. **Unauthenticated → 401**, so the Caddy proxy really does gate —
|
||
the wiki's "SecurityPolicy fails open" warning is about the gateway itself, not the edge.
|
||
- `/v1/models` returns exactly `["deepseek-v4-flash"]`, so trap 3 is clear.
|
||
- **Reasoning survives both surfaces** — see §3b above.
|
||
- The launcher's generated opencode provider block is correct, carrying a real 48-character `llmk-`
|
||
key rather than the `fleetd-local-noauth` placeholder.
|
||
- `SubscriptionGuard` accepted `llm.ltms.dev` after the allowlist edit and the restart: `local`
|
||
spawned without throwing, which is the check that catches a missed restart.
|
||
|
||
### Resolution — both ceilings fixed, migration completed
|
||
|
||
systems/vms fixed both, and each was re-checked from this side rather than taken on trust:
|
||
|
||
| ceiling | was | now | our own check |
|
||
|---|---|---|---|
|
||
| listener buffer | 32 KiB | 32 Mi | 1.2 MB body → **200** (was 413) |
|
||
| LLM route timeout | 60s | 86400s | the request that truncated: **101s, `message_stop` present, 4000/4000** |
|
||
|
||
The timeout moved in two steps on 2026-08-15: 60s → 1800s, then 1800s → **86400s (24 hours)** after
|
||
the truncation risk below was discussed. They tried `request: 0s` first, which removes the
|
||
total-duration timer completely. It works, but on an `AIGatewayRoute` the **idle timeout is derived
|
||
from the request timeout**, so `0s` also removed any bound on a stalled connection. 86400s keeps a
|
||
reaper for dead connections while putting the truncation timer out of practical reach.
|
||
|
||
Neither was deliberate. The 32 KiB was Envoy Gateway's default `per_connection_buffer_limit_bytes`;
|
||
the 60s was Envoy AI Gateway's own documented default. The 60s bounded **generation** as well as
|
||
prompt size — a tiny prompt with a long answer returned 504 at 60.05s.
|
||
|
||
Two configuration facts worth keeping, from their bisection:
|
||
|
||
- **`ClientTrafficPolicy` is honoured in standalone `aigw run`; `BackendTrafficPolicy` is NOT.** A
|
||
`BackendTrafficPolicy` setting `requestTimeout` is accepted, logs nothing, and leaves the routes
|
||
unchanged (upstream `envoyproxy/gateway#9513`). What works is `timeouts: {request: …}` on each
|
||
`AIGatewayRoute` rule. Nothing from the outside distinguishes the two — the same silent-default
|
||
shape as their `SecurityPolicy` caveat.
|
||
- In that stack, "the config was accepted" proves nothing. Read the live `config_dump`.
|
||
|
||
## 7.2 The risk we accepted, and why we could not remove it
|
||
|
||
Raising the timeout made the failure **rare, not impossible**, and the residual failure is silent.
|
||
|
||
On a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked encoding *cleanly* instead of
|
||
resetting the connection, so the client receives what looks like a complete transfer
|
||
(`envoyproxy/envoy#17186` — acknowledged as a bug in 2021, closed by a stale bot, never fixed). The
|
||
December 2025 fix `envoyproxy/envoy#42269` changes locally-originated resets from `NO_ERROR` to
|
||
`INTERNAL_ERROR`, but it is **HTTP/2 only** and SSE clients here speak HTTP/1.1.
|
||
|
||
Measured on our side while the timeout was still 60s:
|
||
|
||
```
|
||
HTTP 200 61.07s 141992 bytes
|
||
message_stop 0 message_delta 0 error events 0
|
||
emitted 2473 of 4000, ending on a WELL-FORMED SSE frame
|
||
```
|
||
|
||
A syntactically valid stream that simply stops. Any timer firing mid-stream — route timeout, idle
|
||
timeout, `max_stream_duration` — fails this same way.
|
||
|
||
**The recommended defence does not transfer to us.** The right fix is to treat a stream with no
|
||
`message_stop` / `[DONE]` / `finish_reason` as failed. We cannot: our members are Claude Code and
|
||
opencode, third-party clients whose SSE parsing we do not own, and there is no seam to insert the
|
||
check. Whether either detects a missing terminator is unverified — and opencode's handling of the 413
|
||
(swallow, compact, retry forever, never surface an error) does not suggest it is strict.
|
||
|
||
So the honest statement of our position:
|
||
|
||
> Gateway traffic is acceptable at 86400s because a single request would have to run for 24 hours to
|
||
> trip the bug — **not** because we could detect it if it did.
|
||
|
||
At 86400s our **own** limit binds first, which is the ordering we want. `MessageService.ASYNC_TIMEOUT_MS`
|
||
caps a turn at 30 minutes, so a runaway request ends as a clean `FAILED` ticket that we raised, rather
|
||
than as a silently truncated `200` that we cannot see. While the gateway sat at 1800s the two numbers
|
||
were equal and did not nest, so a gateway-side stall could have been misread as a bug in our own ticket
|
||
handling. That ambiguity is now gone.
|
||
|
||
**If a member ever returns a confident but truncated answer, suspect this before anything in our own
|
||
code.** That is the whole reason this section exists.
|
||
|
||
---
|
||
|
||
## 8. Related
|
||
|
||
- [CB-589 / #74](https://git.ltms.dev/fleet/fleetd/issues/74) — cost-first placement and a
|
||
gateway that reports live capacity. The per-consumer figures this migration unlocks are the first
|
||
input that ticket actually needs.
|
||
- `docs/CB-500-Multi-Tier-Coordination.md` §11 — the distributed-sandbox topology this gateway is
|
||
part of.
|