CB-591: point member profiles at the new LLM/MCP gateway (claude-code and opencode) #76
Closed
opened 2026-08-15 17:00:38 +02:00 by ltms
·
5 comments
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#76
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Full plan:
docs/CB-591-Gateway-Migration.md. Upstream: systems/vms wiki → LLM and MCP Gateway, systems/vms#31.The gateway went live 2026-08-15 and replaced Bifrost. It serves an Anthropic surface and an
OpenAI surface, so both member kinds can point at it.
Why
The upstream wiki names us as a blocker for retiring the shared
legacytoken: "kb, brain,claude-bridge and the workstation still share it. Each needs its own consumer first."
But the bigger win is on the opencode side. Every opencode member today is
solorterra, and bothsit on one OpenAI account via
credentialId: openai-shared— an exhaustion on either locks outboth. A gateway-backed opencode profile is free, is not on that credential, and is therefore not in
that quarantine pair. That retires a single point of failure rather than only adding capacity.
Two profiles, and they are not symmetric
local(claude-code)gx(opencode, new)https://llm.ltms.dev/anthropichttps://llm.ltms.dev/v1SubscriptionGuardllm.ltms.devonguard.offSubscriptionHostsBridged.java:93)OpenCodeLauncheralready pins an OpenAI-compatible endpoint (CB-508)Do the opencode profile first: it proves the token, URL and model name against a live gateway
while nothing the fleet depends on is touched.
Acceptance criteria
gxopencode profile spawns, completes a turn and ends inbridge_reply.localpoints at/anthropic,llm.ltms.devis on the guard allowlist, and alocalmembercompletes a turn after a daemon restart.
/anthropicthis is the only checkthat catches an OpenAI-schema mistake — Envoy's translator looks for
thinking_blockswhile ourvLLM sends
reasoning_content, and every thinking delta then disappears silently. On/v1this is an open question the plan deliberately does not assume.
local-directprofile atweight: 0keeps the directgx00.gw:8000path reachable by explicitspawn, so the documented escape hatch is real rather than theoretical.
bridge_listshowsgxwith nocredentialId, so asol/terraexhaustion cannot quarantine it.auth.ltms.devcounts requests against aclaude-bridgeconsumer, notlegacy.wiki/11-Features.mdgets an entry: what it does, the knob, why it exists, the gotcha.CLAUDE.mdthat a member mounts "only the bridge MCP" is corrected or made true — itis already false for opencode members, which get context7 and gitea via
opencode.json.Blocked on
Criterion 1–3, 5, 6 need a
claude-bridgeconsumer token issued atauth.ltms.devand stored in${SHARED_ENV}/tools/secrets.shasLLM_BRIDGE_TOKEN. That is the operator's step — it touchessecrets and a host this repo does not own.
Open decision (operator)
Whether members should also mount the gateway's multiplexed
/mcp.mcpUrlis a singleStringinBridgedConfig.Profile, so a claude-code member can mount exactly one MCP — adding the gateway MCP isa code change, and it changes an invariant
CLAUDE.mdstates twice. Not decided in this plan.Carried-over trap worth repeating
Rotating a consumer token restarts the auth proxy, which drops in-flight streaming responses. For
us that kills every live member mid-turn and takes their async ticket reports with them. Same rule as
a daemon redeploy: drain the fleet first (
bridge_list→bridge_poll→bridge_stop), then rotate.Token name settled, and one correction to the plan.
The operator issued the consumer and exported it as
AI_GATEWAY_TOKEN— one key for every agentand MCP client behind
llm.ltms.dev, not a bridge-specific one. The plan doc usedLLM_BRIDGE_TOKEN; that name is now wrong everywhere and has been updated(
docs/CB-591-Gateway-Migration.md, commitf0095bf).Confirmed without printing the value: it resolves in a login shell, is 48 characters, and carries the
documented
llmk-prefix.Correction — the opencode profile also needs a restart. The plan claimed U2a needed none because
SubscriptionGuarddoes not apply to opencode. The guard part is right, the conclusion was not.tokenEnvis resolved byHerdrPeerLauncher.resolveEnv→env.apply(name), which reads thedaemon's own process environment. The running
bridgedinherited its environment at startup, so avariable added to
secrets.shafterwards is absent — the launcher would inject an empty token and thegateway would answer 401, long after the restart and with nothing connecting the two.
That is the same trap as
WORKER_GITEA_TOKEN, so it now gets the same treatment:scripts/redeploy-bridged.sh --checkreports whetherAI_GATEWAY_TOKENresolves in a login shell,and never prints it. Verified against the live daemon:
Revised order: write both profiles, restart once from a login shell, then verify
gxbeforelocal. Both profiles need the restart, so there is no reason to do two — but there is still a reasonto verify in that order.
gxexercises the token and the gateway with no guard involved, so a failurethere is upstream;
localadds the guard allowlist on top, so a failure there is ours. Testing in thatorder separates the two causes instead of confusing them.
Deployed, tested live, and REVERTED — blocked on a 32 KiB request body limit at the edge
U1–U2c were done and the daemon restarted onto them. Then live verification found a blocker that no
amount of config on our side can fix, so
localis back on the direct vLLM andgxis atweight: 0.What is actually wrong
llm.ltms.devanswers HTTP 413 Request Entity Too Large above 32 KiB (32768 bytes), onboth surfaces. Measured, not inferred:
32 KiB is far below one real agent turn. This is an edge limit — Caddy
request_body max_sizeand/or Envoy's own limit — so the fix is in systems/vms, not here.
How it presented, and why that is the dangerous part
Two members were spawned at the same time, same message:
local(claude-code,/anthropic)gx(opencode,/v1)localpassed. It passed because the probe was three trivial questions in a fresh session, sothe request squeaked under 32 KiB. That is the trap: the profile looks healthy and then fails on the
first turn that reads a file.
gxdid not fail loudly either. Reproduced outside the bridge, runningopencodeby hand with thelauncher's own generated config:
opencode catches the 413, compacts, and retries — forever. A member that fails loudly costs one
turn. This one costs the whole task and is indistinguishable from a hang.
What was verified good, and stays good
Everything except the size limit. Worth recording so none of it is re-tested later:
(the wiki warns the gateway's own
SecurityPolicyfails open — that warning is about the gateway,and the proxy in front does its job).
/v1/modelsreturns exactly["deepseek-v4-flash"]— the exact-name trap is clear./anthropicreturns a real"type":"thinking"block;/v1returns a populated
reasoning_content. That closes the open question in §3b of the plan doc — theanswer is that
/v1does not drop reasoning here.baseURL, right model, and areal 48-char
llmk-key rather than thebridged-local-noauthplaceholder.SubscriptionGuardacceptedllm.ltms.devafter the allowlist edit —localspawned withoutthrowing, which is the check that would have caught a missed restart.
Current state
local→http://gx00.gw:8000(direct vLLM),weight: 100. Working.gx→ kept,weight: 0, so nothing auto-selects it. Still spawnable explicitly to test the fix.local-direct→ kept, temporarily duplicateslocal.llm.ltms.devstays inguard.offSubscriptionHosts— it only permits, and it is needed again onthe way back.
f1fd659423e6.bridged.yamlcarries the numbers above and the exact two-key edit to switch back, so whoever picksthis up does not have to re-derive any of it.
Blocked on
Raising the request body limit at the edge in systems/vms. Once that lands: set
baseUrl+tokenEnvonlocal, raisegxtoweight: 100, restart, and re-run §7 — but this time with atask that reads a file, not a three-question probe. The trivial probe is what made this look fine
the first time.
Root cause confirmed by systems/vms — it is Envoy, not Caddy, and it is not deliberate
Correcting my previous comment. I wrote that the limit was "Caddy
request_body max_sizeand/orEnvoy's own" and marked it unverified. The Caddy half was wrong.
The actual cause
Envoy, inside
aigwonllm.vm. Envoy Gateway defaults a listener'sper_connection_buffer_limit_bytesto 32768, and the AI Gateway buffers the whole requestbody before it can route on the model name. So that default is not a network tuning knob here — it
is a hard ceiling on prompt size. Read out of the running Envoy's own
config_dump:Nobody chose 32 KiB. It was inherited from the default.
How they proved both TLS edges were innocent
Worth recording, because it is a better technique than the one I used:
x-llm-consumer— a header their auth proxy sets only after it hasauthenticated the request. So the body cleared both edges and the auth proxy before anything
rejected it.
llm.vmitself:aigwreturns 200 at 32086 bytes and 413 at 39086 bytes, while the vLLM backendanswers 200 at the same 39086 bytes. The backend was never the constraint.
I found the boundary; they found the component. Reading the failure's response headers is what
separated the two, and I did not think to do it.
What this changes for us
Do not plan around 32 KiB. The intended ceiling is far higher, so sizing our profiles to the
current limit would be designing around a bug.
Their fix is a
ClientTrafficPolicysettingbufferLimit: 8Mi— written and committed, notdeployed, pending their operator's approval. At 8Mi the model's context window becomes the binding
limit instead of the buffer, which takes this off the list of numbers anyone has to reason about.
We are holding off
No re-testing until they confirm it is live and verified. Measuring a half-changed system gives a
result neither side can trust. This is now written into
docs/CB-591-Gateway-Migration.mdso it is notpicked up later out of impatience.
One caveat they flagged that we should carry: in standalone
aigw run, aSecurityPolicyisaccepted and then silently ignored. So "the config was accepted" proves nothing in that stack —
they will verify by re-reading the live
config_dumpand by sending a large request, not by theabsence of an error. If
ClientTrafficPolicyis ignored the same way, the fix needs a different shape.That is the same silent-default failure mode this repo keeps hitting, one layer down.
Unchanged
localstays on the direct vLLM,gxstays atweight: 0. When the fix lands, the re-test is amember that reads real files — our first check passed at 32 KiB and hid this for an afternoon.
DONE — the fleet is on the gateway, both ceilings fixed and re-verified
systems/vms fixed both blockers. I re-checked each from this side rather than taking them on trust,
then moved the fleet across and proved it with real workloads.
message_stoppresent, 4000/4000Neither was deliberate. 32 KiB was Envoy Gateway's default
per_connection_buffer_limit_bytes; 60swas Envoy AI Gateway's own default. The 60s bounded generation as well as prompt size — a tiny
prompt with a long answer returned 504 at 60.05s.
Live verification — real workloads, not liveness probes
local/anthropicCLAUDE.mdin full, answer 4 questionsgx/v1local-directhatch-okThe decisive number: those two files are 48,344 bytes of content, comfortably past the old 32,768
ceiling. That exact request was a 413 this morning. The
localmember also reported the document'slength as "roughly 456 lines" against an actual 455 — it genuinely read the whole thing.
A trivial 3-question probe is what hid the 32 KiB ceiling for an afternoon, so it no longer counts as
proof for this profile. Any future re-test uses a file-reading task.
Final state
Daemon pid 17702.
local-directis now genuinely load-bearing rather than decorative: it is the onlyclaude-code path with no TLS edge, no auth proxy and no gateway in it.
What this bought
gxcarries nocredentialId, so asol/terraexhaustion can no longer quarantine it. This retires a single point of failure — theactual point of the migration, more than cost.
legacyone.One risk ACCEPTED, not solved — read §7.2 before debugging a weird member
On a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked encoding cleanly rather
than resetting, so a truncated answer arrives as HTTP 200 with no error and no terminator
(
envoyproxy/envoy#17186— acknowledged 2021, closed by a stale bot, never fixed; the Dec 2025 fix#42269is HTTP/2 only, and SSE here is HTTP/1.1). Measured at the old 60s:The recommended defence — reject a stream with no terminator — does not transfer to us. Claude
Code and opencode are third-party clients and we do not own their SSE parsing. So:
If a member ever returns a confident but truncated answer, suspect this before anything in our own
code.
Also worth keeping from the upstream bisection:
ClientTrafficPolicyis honoured in standaloneaigw run,BackendTrafficPolicyis silently ignored, and nothing external distinguishes them(
envoyproxy/gateway#9513). The same silent-default shape this repo keeps hitting.Still open upstream, not blocking us
systems/vms is asking their operator about dropping the route timeout to
0bounded by the existing3600s idle timeout — a better fit for this bug than any total-duration limit — and about adding a
long-running probe. Their new 40 KB probe is live and passing.
Closing. Remaining CB-591 units U4 (context7 via the gateway
/mcp) and U5 (docs) are separate andtracked in #79.
Correction: one acceptance criterion is NOT verified
Closing this, I listed per-consumer usage metering as a benefit. That was a capability claim, not a
measurement, and I should separate the two.
Criterion 6 — "the cockpit at
auth.ltms.devcounts requests against aclaude-bridgeconsumer,not
legacy" — is UNVERIFIED. The cockpit needs credentials I do not have:What I can say from evidence: the fleet authenticates with
AI_GATEWAY_TOKEN, anllmk-consumertoken issued for this purpose, and requests carrying it are accepted while unauthenticated ones get
401. So traffic is attributed to a consumer token. Whether the cockpit reports it under
claude-bridgespecifically, and whetherlegacycan now be retired, I have not seen.That matters beyond bookkeeping: those figures are the input #74 (CB-589) needs for cost-first
placement. Building on a number nobody has looked at would be the same mistake as trusting a config
that was merely accepted.
Whoever has cockpit access should confirm it. It is a one-minute check and it is the difference
between "we took our own token" and "our own token is actually being counted".
The rest of the criteria stand as reported — 1, 2, 3, 4 and 5 were each checked live, and 7 and 8 are
tracked in #79. This is the only one carried on trust, and it should not have been folded into a
benefits list.