I wrote "Caddy request_body max_size and/or Envoy's own" and marked it
unverified. The Caddy half was wrong, and an unverified guess still points the
next reader at the wrong component.
Confirmed by the systems/vms side: Envoy Gateway defaults a listener's
per_connection_buffer_limit_bytes to 32768, and aigw buffers the WHOLE request
body before it can route on the model name. So that default is not a network
tuning knob — it is a hard ceiling on prompt size. From the live config_dump:
listener default/llm/http per_connection_buffer_limit_bytes: 32768
Nobody chose 32 KiB; it was inherited.
Both TLS edges are innocent, and the technique that showed it is better than
mine: both 413s carry x-llm-consumer, a header their auth proxy sets only AFTER
authenticating, so the body cleared both edges and the auth. On llm.vm, aigw
413s at 39 KB while the vLLM backend answers 200 at the same size. I found the
boundary; they found the component, by reading the failure's response headers.
Consequences recorded in the doc:
* DO NOT plan around 32 KiB. The intended ceiling is far higher, so sizing our
profiles to it would be designing around a bug.
* Their fix (ClientTrafficPolicy, bufferLimit: 8Mi) is written but NOT
deployed, pending their operator's approval. We do not re-test until they
confirm — a half-changed system gives a number neither side can trust.
* In standalone `aigw run` a SecurityPolicy is accepted and then silently
ignored, so "the config was accepted" proves nothing there. They will verify
by re-reading the live config_dump and sending a large request. Same
silent-default shape this repo keeps hitting, one layer down.
bridged.yaml carries the same correction (gitignored, so not in this commit).
Refs: gitea #76
22 KiB
CB-591 — move the fleet onto the LLM and MCP gateway
Status: BLOCKED — deployed 2026-08-15, verified live, then REVERTED. The gateway rejects request
bodies over 32 KiB with HTTP 413, on both surfaces, which is far below one agent turn. local is
back on the direct vLLM and gx is held at weight: 0. Everything else about the migration checked
out — see §7.1. The fix is an edge limit in systems/vms, not in this repo.
· Upstream: systems/vms wiki → LLM and MCP Gateway
· Upstream issue: systems/vms#31
The gateway went live on 2026-08-15 and replaced Bifrost. This plan says what that means for a
member definition in bridged.yaml, because that is the part of this repo the change actually
touches.
1. What changed upstream
One front door for every LLM and MCP client: https://llm.ltms.dev, one token per consumer.
| Surface | URL |
|---|---|
| OpenAI chat | https://llm.ltms.dev/v1/chat/completions |
| OpenAI models | https://llm.ltms.dev/v1/models |
| Anthropic messages | https://llm.ltms.dev/anthropic/v1/messages |
| MCP, all servers multiplexed | https://llm.ltms.dev/mcp |
Anything outside that list returns 404 before any token is checked, on purpose — the gateway must never become a blanket proxy.
The model backend is unchanged: GX10 vLLM at 10.10.10.26:8000 (gx00.gw), model name exactly
deepseek-v4-flash. The direct LAN path stays open on purpose as an escape hatch.
2. Where claude-bridge sits today
We do not use the gateway. The local profile talks straight to the vLLM:
local:
kind: claude-code
baseUrl: http://gx00.gw:8000 # direct vLLM — no auth, LAN only
model: deepseek-v4-flash
configDir: /Users/dai.ha/.ccs/instances/gx10
Three facts about our side that decide the shape of this work:
baseUrlbecomesANTHROPIC_BASE_URLin the member's environment, andtokenEnvbecomesANTHROPIC_AUTH_TOKEN(the value is read from a host env var and never stored in config).localsets notokenEnvtoday, because a direct vLLM needs no token.SubscriptionGuardrefuses any host not on an allowlist, and that allowlist isguard.offSubscriptionHosts: [gx00.gw]. It is built once inBridged.java:93and handed to the launcher, so it is a restart-required key, not a hot one. ChangingbaseUrlwithout changing this makes everylocalspawn throw.- The wiki names us as a blocker. Under Not done yet: retiring the shared
legacytoken is blocked because "kb, brain, claude-bridge and the workstation still share it. Each needs its own consumer first."
Context7 is mounted twice today, both times straight at https://ct7.ltms.dev/mcp — once in
.mcp.json (the primary) and once in opencode.json (the sol and terra members).
flowchart LR
subgraph now["Today"]
M1["local member<br/>claude-code"] -->|"ANTHROPIC_BASE_URL"| V1["vLLM gx00.gw:8000<br/>no auth, LAN only"]
M2["sol / terra<br/>opencode"] --> CT1["ct7.ltms.dev/mcp"]
P1["primary"] --> CT1
end
subgraph after["Proposed"]
M3["local member"] -->|"ANTHROPIC_BASE_URL<br/>+ ANTHROPIC_AUTH_TOKEN"| G["llm.ltms.dev/anthropic<br/>consumer: claude-bridge"]
G --> V2["vLLM gx00.gw:8000"]
M4["local-direct<br/>weight 0, escape hatch"] --> V2
end
The member definition is the only thing that moves. The model behind it does not.
3. The member definition change
The gateway serves an Anthropic surface and an OpenAI surface, so both member kinds can point at
it. That is the main opportunity here, and it is bigger than the local profile alone.
3a. local — claude-code, on /anthropic
| Key | Today | After | Note |
|---|---|---|---|
baseUrl |
http://gx00.gw:8000 |
https://llm.ltms.dev/anthropic |
see the schema warning below |
tokenEnv |
(unset) | AI_GATEWAY_TOKEN |
new consumer token, llmk-claude-bridge-<32 hex> |
model |
deepseek-v4-flash |
unchanged | must stay exact; a regex match returns an empty /v1/models while completions keep working |
guard.offSubscriptionHosts |
[gx00.gw] |
[gx00.gw, llm.ltms.dev] |
restart required |
3b. A new opencode profile on /v1 — no code needed
OpenCodeLauncher already supports a pinned OpenAI-compatible endpoint (CB-508). Given baseUrl it
writes a custom provider block into the worker's opencode config:
baseUrl→options.baseURL.openAiBaseUrlappends/v1to a bare host, and takes a URL that already has a path as-is — sohttps://llm.ltms.dev/v1works unchanged.tokenEnv→options.apiKey(falls back to a placeholder when unset, since a local vLLM ignores it).model:must be<provider>/<model>whenbaseUrlis set — a bare name is rejected loudly rather than silently falling back to opencode's default gateway.
So the profile is pure config:
gx:
kind: opencode
baseUrl: https://llm.ltms.dev/v1
tokenEnv: AI_GATEWAY_TOKEN
model: gx/deepseek-v4-flash # provider id is ours to choose; the half after / is the model
argv: ["opencode"]
mcpUrl: http://127.0.0.1:8765/mcp
gitTokenEnv: WORKER_GITEA_TOKEN
weight: 100 # same tier as `local` — free
maxLoad: 2
# deliberately NO credentialId — this is our own box, not the shared OpenAI account
Why this matters more than it looks. Today every opencode member is sol or terra, and those
are two models on one OpenAI account sharing credentialId: openai-shared — so an exhaustion on
either locks out both, and half the fleet's opencode capacity dies at once. A gateway-backed opencode
profile is free, is not on that credential, and therefore is not in that quarantine pair. It removes
a single point of failure rather than just adding capacity.
Note the asymmetry, it is deliberate: SubscriptionGuard does not apply to opencode at all — the
guard exists to stop a Claude worker borrowing the operator's subscription, and opencode reads its
own provider credentials. So 3b needs no allowlist change; only 3a does.
Both still need a restart, for a different reason. tokenEnv is resolved by
HerdrPeerLauncher.resolveEnv → env.apply(name), which reads the daemon's own process
environment. The running bridged inherited its environment when it started, so a variable added to
secrets.sh afterwards is simply not there — the launcher would inject an empty token and the
gateway would answer 401. This is the same failure as trap 1 in scripts/redeploy-bridged.sh
(WORKER_GITEA_TOKEN), and it has the same fix: restart from a login shell, and use
scripts/redeploy-bridged.sh --check to confirm the name resolves before restarting anything.
3c. What this does to ccs
Once a profile carries baseUrl, tokenEnv and model itself, the ccs instance stops being what
routes a member. Be precise about what is left, though: configDir still supplies folder trust
and settings.json, and dropping it is what produced the trust dialog and the wrong-model error
recorded in bridged.yaml. So ccs goes from deciding where the tokens go to holding client-side
state. Less load-bearing, not removable.
Why /anthropic and never /v1/chat/completions
The gateway declares its Anthropic backend as schema.name: Anthropic, which means no
translation — streaming, tool use and thinking blocks pass through exactly as they do against vLLM
directly.
Declared as OpenAI, Envoy's translator looks for a thinking_blocks field that our vLLM does not
send (it sends reasoning_content), and every thinking delta disappears silently. Claude Code
speaks the Anthropic protocol, so /anthropic is both correct and the only safe choice.
This is the exact failure shape this repo keeps hitting: it compiles, it answers, it looks healthy,
and a capability is quietly off. Treat it as a silent-default risk, not a config preference.
Open question for 3b — ANSWERED, 2026-08-15. The worry was that the OpenAI surface might drop reasoning the way the wiki documents for a mis-declared Anthropic backend. It does not. Checked at the API before any profile was switched:
| surface | request | result |
|---|---|---|
/anthropic/v1/messages |
deepseek-v4-flash, 64 tokens |
200, response carries a real "type":"thinking" block |
/v1/chat/completions |
same | 200, message carries a populated reasoning_content (and a reasoning field) |
/v1/models |
— | 200, exactly ["deepseek-v4-flash"] — the exact-name trap is clear |
/v1/models, no token |
— | 401 — Caddy is gating, as designed |
So reasoning survives on both surfaces, and the /anthropic choice for local is about protocol
correctness rather than a repair for a known loss. The last row matters on its own: the wiki warns
the gateway's own SecurityPolicy fails open, so it is worth knowing the proxy in front really does
refuse an unauthenticated request here.
4. Decisions
D1 — switch, but keep the direct path as an explicit profile · recommended
Switching buys four things we do not have:
- Free opencode capacity, off the shared credential. The largest single win. See §3b — it retires a real single point of failure, not just a cost line.
- Per-consumer usage figures. The cockpit counts requests per consumer. That is the first real measurement of what the fleet consumes, and it feeds CB-589 Gap 2 directly.
- Our own revocable token. One consumer to revoke if a worker ever leaks it, instead of a shared
legacytoken used by four systems. - It works off-LAN.
gx00.gwresolves on the LAN only.
The cost is honest and worth stating: we add a TLS edge, an auth proxy and a gateway to the path of every member spawn. The wiki keeps the direct route open precisely because "if the gateway breaks, nothing that matters is blocked."
So keep it. Add a second profile local-direct pointing at http://gx00.gw:8000 with weight: 0
— never auto-selected, still spawnable with an explicit bridge_spawn{profile: "local-direct"}.
That is exactly what CB-554 made weight: 0 mean, and it turns the escape hatch into something the
lead can actually reach during an incident.
D2 — do members also mount the gateway's /mcp? · OPEN, operator's call
Not a detail. CLAUDE.md states in two places that a member mounts only the bridge MCP, and a
worker's honesty rule leans on it ("never claim the result of a check you had no way to run").
- Keep bridge-only. The invariant stays true and simple. Workers stay cheap and narrow.
- Add the gateway MCP. Implementers get context7 documentation lookups, which is genuinely useful
for library work. But
mcpUrlinBridgedConfig.Profileis a singleString, so a claude-code member can mount exactly one MCP — this needs a code change, not a config edit.
Note the invariant is already inaccurate: opencode.json gives sol and terra both context7
and gitea. So the choice is really "make the rule true" or "make the rule match reality". Either is
defensible; picking one is not mine to do.
D3 — token scope
One consumer, claude-bridge, its token in ${SHARED_ENV}/tools/secrets.sh as AI_GATEWAY_TOKEN,
referenced by name only. Never the literal value in bridged.yaml — tokenEnv exists for this.
5. Units of work
flowchart TB
U1["U1 · consumer token<br/>issue via cockpit, add to secrets.sh"]
U2["U2 · profile + guard<br/>bridged.yaml, restart"]
U3["U3 · verify live<br/>spawn, prove thinking survives"]
U4["U4 · context7 via gateway<br/>.mcp.json + opencode.json"]
U5["U5 · docs<br/>CLAUDE.md, wiki 11-Features"]
U1 --> U2 --> U3
U4 --> U5
U3 --> U5
| # | Scope | Who | Why |
|---|---|---|---|
| U1 | Issue the claude-bridge consumer at auth.ltms.dev; store as AI_GATEWAY_TOKEN |
operator | touches secrets and a host we do not own |
| U2a | New gx opencode profile on /v1 — pure config, no guard change |
lead | bridged.yaml is gitignored, so a worker cannot see or edit it |
| U2b | local → /anthropic; add local-direct weight 0; add llm.ltms.dev to the guard allowlist |
lead | same |
| U2c | One restart from a login shell, after U2a and U2b | lead | picks up AI_GATEWAY_TOKEN into the daemon env and the guard allowlist, in one stop |
| U3 | Live spawn on both new profiles; confirm reasoning survives on each surface | lead | needs real spawns and the running daemon |
| U4 | Point .mcp.json and opencode.json context7 at the gateway /mcp; rename pinned tools |
delegatable | tracked files, self-contained |
| U5 | Fix the "members mount only the bridge" claim; add a wiki/11-Features.md entry |
delegatable | writing, clear criteria |
U1 blocks U2a, U2b and U3. U4 and U5 do not depend on it.
Write U2a and U2b, then restart once (U2c), then verify gx before local. Since both profiles
need the same restart there is no reason to do two, but there is still a reason to verify in order:
gx exercises the token and the gateway with no guard involved, so if it fails the cause is upstream.
local adds the guard allowlist on top, so a failure there points at our config instead. Testing them
in that order separates the two causes instead of confusing them.
U1 status, 2026-08-15: the operator issued the consumer and exported it as
AI_GATEWAY_TOKEN(one key for every agent and MCP client behindllm.ltms.dev). Confirmed: it resolves in a login shell, is 48 characters and carries the documentedllmk-prefix. The value was never printed.
6. Traps carried over from the wiki
Each of these cost someone real debugging time upstream. They apply to us.
- Rotating a token restarts the auth proxy, which drops in-flight streaming responses. For us
that means rotating
AI_GATEWAY_TOKENkills every live member mid-turn, and an async ticket's report goes with it. This is the same rule as a daemon redeploy: drain the fleet first (bridge_list→bridge_pollanything wanted →bridge_stop), then rotate. - The gateway's own
SecurityPolicyfails open. Standaloneaigw runaccepts it and silently ignores it — an unauthenticated request returned 200. Auth is the Caddy proxy in front, and nothing else. Never reason as if the gateway authenticates. - Exact model name. A regex match routes fine but returns an empty
/v1/modelslist while completions keep working. A wrong name returns a bare 404 that reads exactly like a dead gateway. - MCP tool names changed prefix separator. Bifrost used one dash (
ct7-resolve-library-id); the gateway uses two underscores (ct7__resolve-library-id). Relevant only if U4 is done. /v1/models404 vs empty list are different faults. 404 means no route loaded at all; empty means the model match is a regex. Do not conflate them when diagnosing.
7. Verification — what would prove this works
Merging config is not proving it. The checks, in order:
bridge_spawn{profile: "gx"}succeeds and the member completes a real turn ending inbridge_reply. This is the first proof of the token, the URL and the model name, and it risks nothing the fleet depends on.bridge_spawn{profile: "local"}succeeds. If the guard allowlist was missed, this throws — a loud, self-correcting failure, which is the good kind. If the restart was missed, it also throws, for the same reason.- A
localmember completes a turn. That exercises streaming through two TLS edges, the auth proxy and the gateway. - Reasoning survives, checked separately on each surface. For
localon/anthropicthis is the check that catches the/v1versus/anthropicmistake, and it is the only one that does — nothing else distinguishes a working passthrough from a translator quietly dropping thinking deltas. Forgxon/v1, this answers the open question in §3 rather than assuming it. - The cockpit at
auth.ltms.devshows requests counted against theclaude-bridgeconsumer, notlegacy. That is the whole point of taking our own token. bridge_spawn{profile: "local-direct"}still works, so the escape hatch is real rather than theoretical.bridge_listshowsgxcarrying nocredentialId, so asol/terraexhaustion cannot quarantine it. This is the single-point-of-failure claim in §3b, checked rather than asserted.
7.1 What the live run actually found — 2026-08-15
U1–U2c were done, the daemon restarted onto them, and both new profiles were spawned for real. The migration was then reverted. This section is the result, so none of it has to be re-derived.
The blocker
llm.ltms.dev answers HTTP 413 Request Entity Too Large above 32 KiB (32768 bytes), on both
surfaces:
/v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200
/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413
32 KiB is far below one real agent turn.
Root cause — confirmed by the systems/vms side, 2026-08-15. My guess that it was a Caddy
request_body max_size was wrong. It is Envoy, inside aigw on llm.vm. Envoy Gateway defaults
a listener's per_connection_buffer_limit_bytes to 32768, and the AI Gateway buffers the whole
request body before it can route on the model name — so that default is not a network tuning knob
here, it is a hard ceiling on prompt size. Read out of the live Envoy config_dump:
listener default/llm/http per_connection_buffer_limit_bytes: 32768
Nobody chose 32 KiB; it was inherited from the default. Both TLS edges are innocent: the same
boundary reproduces on the LAN path and the internet path, and both 413s carry an x-llm-consumer
header their auth proxy sets only after authenticating — so the body cleared both edges and the
auth. Directly on llm.vm, aigw 413s at 39 KB while the vLLM backend accepts the same 39 KB and
answers 200.
Do not plan around 32 KiB. The intended ceiling is far higher. Their fix — a ClientTrafficPolicy
setting bufferLimit: 8Mi — is written but not deployed as of this note, pending their operator's
approval. I have not re-tested and will not until they confirm, so as not to measure a half-changed
system. Fixed in systems/vms, not here.
The part worth remembering
Two members were spawned at the same moment with the same message:
local (claude-code, /anthropic) |
gx (opencode, /v1) |
|
|---|---|---|
| READY → BUSY | 19:07:26 | 19:07:45 |
| BUSY → DONE | 19:08:51 (66s) | never — 10+ min, ticket FAILED |
local passed. It passed only because the probe was three trivial questions in a fresh session,
so the request fit under 32 KiB. The profile looked healthy and was a landmine set to fire on the
first turn that reads a file.
So §7's checklist was not wrong, it was too easy. Any future run of it must use a task that reads a real file. A liveness probe proves the token and the URL; it does not prove the path.
gx did not fail loudly either. Reproduced outside the bridge by running opencode by hand with the
launcher's own generated config:
Error: Request Entity Too Large
...compacts context, retries...
Error: Request Entity Too Large
opencode catches the 413, compacts, and retries — indefinitely. A member that fails loudly costs one turn; this one costs the whole task and is indistinguishable from a slow worker.
Diagnosing a stuck opencode member. Do not read its pane. The launcher writes its config to a temp dir and passes it as
OPENCODE_CONFIG— find it withls -dt /var/folders/*/*/T/bridged-opencode-* | head -1, check the provider block and the key's length and prefix (never its value), then reproduce withopencode run --auto -m <provider>/<model>using the sameOPENCODE_CONFIG. That is what turned "it hangs" into a one-line error.
What checked out, and needs no re-testing
- Token accepted on both surfaces. Unauthenticated → 401, so the Caddy proxy really does gate — the wiki's "SecurityPolicy fails open" warning is about the gateway itself, not the edge.
/v1/modelsreturns exactly["deepseek-v4-flash"], so trap 3 is clear.- Reasoning survives both surfaces — see §3b above.
- The launcher's generated opencode provider block is correct, carrying a real 48-character
llmk-key rather than thebridged-local-noauthplaceholder. SubscriptionGuardacceptedllm.ltms.devafter the allowlist edit and the restart:localspawned without throwing, which is the check that catches a missed restart.
To resume
Wait for systems/vms to confirm the fix is deployed and verified. Do not re-test before that —
measuring a half-changed system produces a result nobody can trust. They have said they will verify by
re-reading the live config_dump and by sending a large request, not by the absence of an error,
because in standalone aigw run a SecurityPolicy is accepted and then silently ignored. If
ClientTrafficPolicy turns out to be ignored the same way, the fix will need a different shape.
Then: set baseUrl + tokenEnv on local, raise gx to weight: 100, restart, and re-run §7
with a file-reading task. bridged.yaml carries the exact two-key edit and these numbers inline.
8. Related
- CB-589 / #74 — cost-first placement and a gateway that reports live capacity. The per-consumer figures this migration unlocks are the first input that ticket actually needs.
docs/CB-500-Multi-Tier-Coordination.md§11 — the distributed-sandbox topology this gateway is part of.