Files
fleetd/docs/CB-591-Gateway-Migration.md
T
Dai Ha f0095bf8b2
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m38s
CB-591: plan the move onto the LLM/MCP gateway, and check AI_GATEWAY_TOKEN
The gateway (llm.ltms.dev) replaced Bifrost on 2026-08-15 and serves an Anthropic
surface and an OpenAI surface, so both member kinds can point at it. The plan is in
docs/CB-591-Gateway-Migration.md; gitea #76 tracks the work.

The opencode half needs no code: OpenCodeLauncher already pins an OpenAI-compatible
endpoint (CB-508), so baseUrl + tokenEnv + provider/model is a config change. That
matters more than it looks — every opencode member today is sol or terra, and both sit
on one OpenAI account via credentialId: openai-shared, so an exhaustion on either locks
out both. A gateway-backed opencode profile is free and off that credential, which
retires a single point of failure rather than only adding capacity.

Also extends the redeploy script's --check to AI_GATEWAY_TOKEN. A profile's tokenEnv is
resolved from the DAEMON's own environment by HerdrPeerLauncher.resolveEnv, so a token
added to secrets.sh after the daemon started is simply absent: the launcher injects an
empty token and the gateway answers 401, long after the restart and with nothing tying
the two together. That is the same trap as WORKER_GITEA_TOKEN, and it gets the same
login-shell check that never prints the value.
2026-08-15 17:06:31 +02:00

16 KiB

CB-591 — move the fleet onto the LLM and MCP gateway

Status: plan, not started · Upstream: systems/vms wiki → LLM and MCP Gateway · Upstream issue: systems/vms#31

The gateway went live on 2026-08-15 and replaced Bifrost. This plan says what that means for a member definition in bridged.yaml, because that is the part of this repo the change actually touches.


1. What changed upstream

One front door for every LLM and MCP client: https://llm.ltms.dev, one token per consumer.

Surface URL
OpenAI chat https://llm.ltms.dev/v1/chat/completions
OpenAI models https://llm.ltms.dev/v1/models
Anthropic messages https://llm.ltms.dev/anthropic/v1/messages
MCP, all servers multiplexed https://llm.ltms.dev/mcp

Anything outside that list returns 404 before any token is checked, on purpose — the gateway must never become a blanket proxy.

The model backend is unchanged: GX10 vLLM at 10.10.10.26:8000 (gx00.gw), model name exactly deepseek-v4-flash. The direct LAN path stays open on purpose as an escape hatch.


2. Where claude-bridge sits today

We do not use the gateway. The local profile talks straight to the vLLM:

local:
  kind: claude-code
  baseUrl: http://gx00.gw:8000     # direct vLLM — no auth, LAN only
  model: deepseek-v4-flash
  configDir: /Users/dai.ha/.ccs/instances/gx10

Three facts about our side that decide the shape of this work:

  1. baseUrl becomes ANTHROPIC_BASE_URL in the member's environment, and tokenEnv becomes ANTHROPIC_AUTH_TOKEN (the value is read from a host env var and never stored in config). local sets no tokenEnv today, because a direct vLLM needs no token.
  2. SubscriptionGuard refuses any host not on an allowlist, and that allowlist is guard.offSubscriptionHosts: [gx00.gw]. It is built once in Bridged.java:93 and handed to the launcher, so it is a restart-required key, not a hot one. Changing baseUrl without changing this makes every local spawn throw.
  3. The wiki names us as a blocker. Under Not done yet: retiring the shared legacy token is blocked because "kb, brain, claude-bridge and the workstation still share it. Each needs its own consumer first."

Context7 is mounted twice today, both times straight at https://ct7.ltms.dev/mcp — once in .mcp.json (the primary) and once in opencode.json (the sol and terra members).

flowchart LR
    subgraph now["Today"]
        M1["local member<br/>claude-code"] -->|"ANTHROPIC_BASE_URL"| V1["vLLM gx00.gw:8000<br/>no auth, LAN only"]
        M2["sol / terra<br/>opencode"] --> CT1["ct7.ltms.dev/mcp"]
        P1["primary"] --> CT1
    end
    subgraph after["Proposed"]
        M3["local member"] -->|"ANTHROPIC_BASE_URL<br/>+ ANTHROPIC_AUTH_TOKEN"| G["llm.ltms.dev/anthropic<br/>consumer: claude-bridge"]
        G --> V2["vLLM gx00.gw:8000"]
        M4["local-direct<br/>weight 0, escape hatch"] --> V2
    end

The member definition is the only thing that moves. The model behind it does not.


3. The member definition change

The gateway serves an Anthropic surface and an OpenAI surface, so both member kinds can point at it. That is the main opportunity here, and it is bigger than the local profile alone.

3a. local — claude-code, on /anthropic

Key Today After Note
baseUrl http://gx00.gw:8000 https://llm.ltms.dev/anthropic see the schema warning below
tokenEnv (unset) AI_GATEWAY_TOKEN new consumer token, llmk-claude-bridge-<32 hex>
model deepseek-v4-flash unchanged must stay exact; a regex match returns an empty /v1/models while completions keep working
guard.offSubscriptionHosts [gx00.gw] [gx00.gw, llm.ltms.dev] restart required

3b. A new opencode profile on /v1 — no code needed

OpenCodeLauncher already supports a pinned OpenAI-compatible endpoint (CB-508). Given baseUrl it writes a custom provider block into the worker's opencode config:

  • baseUrl → options.baseURL. openAiBaseUrl appends /v1 to a bare host, and takes a URL that already has a path as-is — so https://llm.ltms.dev/v1 works unchanged.
  • tokenEnv → options.apiKey (falls back to a placeholder when unset, since a local vLLM ignores it).
  • model: must be <provider>/<model> when baseUrl is set — a bare name is rejected loudly rather than silently falling back to opencode's default gateway.

So the profile is pure config:

gx:
  kind: opencode
  baseUrl: https://llm.ltms.dev/v1
  tokenEnv: AI_GATEWAY_TOKEN
  model: gx/deepseek-v4-flash        # provider id is ours to choose; the half after / is the model
  argv: ["opencode"]
  mcpUrl: http://127.0.0.1:8765/mcp
  gitTokenEnv: WORKER_GITEA_TOKEN
  weight: 100                        # same tier as `local` — free
  maxLoad: 2
  # deliberately NO credentialId — this is our own box, not the shared OpenAI account

Why this matters more than it looks. Today every opencode member is sol or terra, and those are two models on one OpenAI account sharing credentialId: openai-shared — so an exhaustion on either locks out both, and half the fleet's opencode capacity dies at once. A gateway-backed opencode profile is free, is not on that credential, and therefore is not in that quarantine pair. It removes a single point of failure rather than just adding capacity.

Note the asymmetry, it is deliberate: SubscriptionGuard does not apply to opencode at all — the guard exists to stop a Claude worker borrowing the operator's subscription, and opencode reads its own provider credentials. So 3b needs no allowlist change; only 3a does.

Both still need a restart, for a different reason. tokenEnv is resolved by HerdrPeerLauncher.resolveEnv → env.apply(name), which reads the daemon's own process environment. The running bridged inherited its environment when it started, so a variable added to secrets.sh afterwards is simply not there — the launcher would inject an empty token and the gateway would answer 401. This is the same failure as trap 1 in scripts/redeploy-bridged.sh (WORKER_GITEA_TOKEN), and it has the same fix: restart from a login shell, and use scripts/redeploy-bridged.sh --check to confirm the name resolves before restarting anything.

3c. What this does to ccs

Once a profile carries baseUrl, tokenEnv and model itself, the ccs instance stops being what routes a member. Be precise about what is left, though: configDir still supplies folder trust and settings.json, and dropping it is what produced the trust dialog and the wrong-model error recorded in bridged.yaml. So ccs goes from deciding where the tokens go to holding client-side state. Less load-bearing, not removable.

Why /anthropic and never /v1/chat/completions

The gateway declares its Anthropic backend as schema.name: Anthropic, which means no translation — streaming, tool use and thinking blocks pass through exactly as they do against vLLM directly.

Declared as OpenAI, Envoy's translator looks for a thinking_blocks field that our vLLM does not send (it sends reasoning_content), and every thinking delta disappears silently. Claude Code speaks the Anthropic protocol, so /anthropic is both correct and the only safe choice.

This is the exact failure shape this repo keeps hitting: it compiles, it answers, it looks healthy, and a capability is quietly off. Treat it as a silent-default risk, not a config preference.

Open question for 3b, stated as open. The opencode profile must use the OpenAI surface, because that is the only thing an OpenAI-compatible provider can speak. vLLM serves that API natively, so I expect it to be a passthrough as well — but the wiki documents the translation trap only for the Anthropic path, and I have not checked how reasoning_content behaves through /v1. Verify it on first spawn (§7) rather than assuming. If reasoning is dropped there, that is a limitation of the opencode profile, not a reason to abandon it — opencode members do dev work, not deep reasoning.


4. Decisions

Switching buys four things we do not have:

  • Free opencode capacity, off the shared credential. The largest single win. See §3b — it retires a real single point of failure, not just a cost line.
  • Per-consumer usage figures. The cockpit counts requests per consumer. That is the first real measurement of what the fleet consumes, and it feeds CB-589 Gap 2 directly.
  • Our own revocable token. One consumer to revoke if a worker ever leaks it, instead of a shared legacy token used by four systems.
  • It works off-LAN. gx00.gw resolves on the LAN only.

The cost is honest and worth stating: we add a TLS edge, an auth proxy and a gateway to the path of every member spawn. The wiki keeps the direct route open precisely because "if the gateway breaks, nothing that matters is blocked."

So keep it. Add a second profile local-direct pointing at http://gx00.gw:8000 with weight: 0 — never auto-selected, still spawnable with an explicit bridge_spawn{profile: "local-direct"}. That is exactly what CB-554 made weight: 0 mean, and it turns the escape hatch into something the lead can actually reach during an incident.

D2 — do members also mount the gateway's /mcp? · OPEN, operator's call

Not a detail. CLAUDE.md states in two places that a member mounts only the bridge MCP, and a worker's honesty rule leans on it ("never claim the result of a check you had no way to run").

  • Keep bridge-only. The invariant stays true and simple. Workers stay cheap and narrow.
  • Add the gateway MCP. Implementers get context7 documentation lookups, which is genuinely useful for library work. But mcpUrl in BridgedConfig.Profile is a single String, so a claude-code member can mount exactly one MCP — this needs a code change, not a config edit.

Note the invariant is already inaccurate: opencode.json gives sol and terra both context7 and gitea. So the choice is really "make the rule true" or "make the rule match reality". Either is defensible; picking one is not mine to do.

D3 — token scope

One consumer, claude-bridge, its token in ${SHARED_ENV}/tools/secrets.sh as AI_GATEWAY_TOKEN, referenced by name only. Never the literal value in bridged.yaml — tokenEnv exists for this.


5. Units of work

flowchart TB
    U1["U1 · consumer token<br/>issue via cockpit, add to secrets.sh"]
    U2["U2 · profile + guard<br/>bridged.yaml, restart"]
    U3["U3 · verify live<br/>spawn, prove thinking survives"]
    U4["U4 · context7 via gateway<br/>.mcp.json + opencode.json"]
    U5["U5 · docs<br/>CLAUDE.md, wiki 11-Features"]
    U1 --> U2 --> U3
    U4 --> U5
    U3 --> U5
# Scope Who Why
U1 Issue the claude-bridge consumer at auth.ltms.dev; store as AI_GATEWAY_TOKEN operator touches secrets and a host we do not own
U2a New gx opencode profile on /v1 — pure config, no guard change lead bridged.yaml is gitignored, so a worker cannot see or edit it
U2b local → /anthropic; add local-direct weight 0; add llm.ltms.dev to the guard allowlist lead same
U2c One restart from a login shell, after U2a and U2b lead picks up AI_GATEWAY_TOKEN into the daemon env and the guard allowlist, in one stop
U3 Live spawn on both new profiles; confirm reasoning survives on each surface lead needs real spawns and the running daemon
U4 Point .mcp.json and opencode.json context7 at the gateway /mcp; rename pinned tools delegatable tracked files, self-contained
U5 Fix the "members mount only the bridge" claim; add a wiki/11-Features.md entry delegatable writing, clear criteria

U1 blocks U2a, U2b and U3. U4 and U5 do not depend on it.

Write U2a and U2b, then restart once (U2c), then verify gx before local. Since both profiles need the same restart there is no reason to do two, but there is still a reason to verify in order: gx exercises the token and the gateway with no guard involved, so if it fails the cause is upstream. local adds the guard allowlist on top, so a failure there points at our config instead. Testing them in that order separates the two causes instead of confusing them.

U1 status, 2026-08-15: the operator issued the consumer and exported it as AI_GATEWAY_TOKEN (one key for every agent and MCP client behind llm.ltms.dev). Confirmed: it resolves in a login shell, is 48 characters and carries the documented llmk- prefix. The value was never printed.


6. Traps carried over from the wiki

Each of these cost someone real debugging time upstream. They apply to us.

  1. Rotating a token restarts the auth proxy, which drops in-flight streaming responses. For us that means rotating AI_GATEWAY_TOKEN kills every live member mid-turn, and an async ticket's report goes with it. This is the same rule as a daemon redeploy: drain the fleet first (bridge_list → bridge_poll anything wanted → bridge_stop), then rotate.
  2. The gateway's own SecurityPolicy fails open. Standalone aigw run accepts it and silently ignores it — an unauthenticated request returned 200. Auth is the Caddy proxy in front, and nothing else. Never reason as if the gateway authenticates.
  3. Exact model name. A regex match routes fine but returns an empty /v1/models list while completions keep working. A wrong name returns a bare 404 that reads exactly like a dead gateway.
  4. MCP tool names changed prefix separator. Bifrost used one dash (ct7-resolve-library-id); the gateway uses two underscores (ct7__resolve-library-id). Relevant only if U4 is done.
  5. /v1/models 404 vs empty list are different faults. 404 means no route loaded at all; empty means the model match is a regex. Do not conflate them when diagnosing.

7. Verification — what would prove this works

Merging config is not proving it. The checks, in order:

  1. bridge_spawn{profile: "gx"} succeeds and the member completes a real turn ending in bridge_reply. This is the first proof of the token, the URL and the model name, and it risks nothing the fleet depends on.
  2. bridge_spawn{profile: "local"} succeeds. If the guard allowlist was missed, this throws — a loud, self-correcting failure, which is the good kind. If the restart was missed, it also throws, for the same reason.
  3. A local member completes a turn. That exercises streaming through two TLS edges, the auth proxy and the gateway.
  4. Reasoning survives, checked separately on each surface. For local on /anthropic this is the check that catches the /v1 versus /anthropic mistake, and it is the only one that does — nothing else distinguishes a working passthrough from a translator quietly dropping thinking deltas. For gx on /v1, this answers the open question in §3 rather than assuming it.
  5. The cockpit at auth.ltms.dev shows requests counted against the claude-bridge consumer, not legacy. That is the whole point of taking our own token.
  6. bridge_spawn{profile: "local-direct"} still works, so the escape hatch is real rather than theoretical.
  7. bridge_list shows gx carrying no credentialId, so a sol/terra exhaustion cannot quarantine it. This is the single-point-of-failure claim in §3b, checked rather than asserted.

  • CB-589 / #74 — cost-first placement and a gateway that reports live capacity. The per-consumer figures this migration unlocks are the first input that ticket actually needs.
  • docs/CB-500-Multi-Tier-Coordination.md §11 — the distributed-sandbox topology this gateway is part of.