555715ced99b4802b8b47bbb7cd88657d7c2bfa5
7 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
555715ced9 |
CB-623: point every path at fleet/fleetd after the org transfer
The repo moved lms/claude-bridge -> fleet/claude-bridge -> fleet/fleetd. Gitea redirects hold, so most of this is not urgent, but one line was a real break: the implementer skill posts a worker's PR to a hardcoded repo path, so every worker PR would have gone to the old address. .claude/skills/implementer/SKILL.md the worker PR endpoint (functional) .gitmodules wiki submodule URL CLAUDE.md + wiki/7-Use-Cases.md the canonical block, kept byte-identical README.md clone command and wiki link deploy/bridged.service Documentation= plugin/.claude-plugin/plugin.json homepage + repository docs/*.md issue and wiki links The wiki is not a separate repo. /repos/lms/claude-bridge.wiki returns 404 and lms owned no .wiki entity, so the wiki moved with the repo; both the old and the new wiki SSH URLs resolve to the same sha. Ticket step 4 assumed a second transfer that does not exist. |
||
|
|
76e4b577a0 |
CB-622: rename bridge_* tool names to fleet_* in Markdown docs
Rename the eleven MCP tool names (bridge_ack/ask/list/poll/profiles/reply/ send/spawn/status/stop/whoami) to their fleet_* names across the Markdown documentation. fleet_* is written as the normal name; one deprecation note in README.md says bridge_* still works for one release. docs/MCP-Contract.md is renamed only inside section 6 (lines 210-293): sections 1-5 and 7-11 are stale pre-build design text (CB-609) and are deliberately left with old names so dead text does not look maintained. Also renames e2e/bridge_ask_transcript.md to e2e/fleet_ask_transcript.md to match its content. CLAUDE.md, wiki/, plugin/skills/setup/SKILL.md and .claude/skills/port-to-opencode/SKILL.md are owned by other units and are untouched. |
||
|
|
e4f3620acb |
CB-591: record the final 86400s route timeout and the request:0s trap
systems/vms moved the LLM route timeout again, 1800s -> 86400s (24h), after the silent-truncation risk was discussed. They tried request: 0s first: it removes the total-duration timer, but on an AIGatewayRoute the idle timeout is derived from the request timeout, so 0s also removed any bound on a stalled connection. At 86400s our own MessageService.ASYNC_TIMEOUT_MS (30 min) binds first, so a runaway request now ends as a clean FAILED ticket we raised instead of a silently truncated 200. While the gateway sat at 1800s the two numbers were equal and did not nest. |
||
|
|
032a59a34d |
CB-591: fleet moved onto the gateway — both ceilings fixed and re-verified
`local` runs on /anthropic and `gx` on /v1, both weight 100; `local-direct`
stays weight 0 as the escape hatch.
systems/vms fixed both blockers, and each was re-checked from this side rather
than taken on trust:
listener buffer 32 KiB -> 32 Mi ours: 1.2 MB body -> 200 (was 413)
LLM route timeout 60s -> 1800s ours: 101s stream -> 200,
message_stop present, 4000/4000
Neither was deliberate: 32 KiB was Envoy Gateway's default
per_connection_buffer_limit_bytes, and 60s was Envoy AI Gateway's own default.
The 60s bounded GENERATION as well as prompt size — a tiny prompt with a long
answer returned 504 at 60.05s.
Verified with real workloads, not liveness probes. A `local` member read this
document and CLAUDE.md in full — 48,344 bytes of file content, comfortably past
the old 32,768 ceiling — and answered four questions correctly, including the
document's length (said ~456, actual 455). A `gx` member did the same. The
trivial 3-question probe is what hid the 32 KiB ceiling for an afternoon, so it
no longer counts as proof here.
§7.2 is new and is the part that matters later. One risk is ACCEPTED, not
solved: on a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked
encoding cleanly instead of resetting, so a truncated answer arrives as HTTP 200
with no error and no terminator (envoyproxy/envoy#17186, acknowledged 2021,
never fixed; the Dec 2025 fix #42269 is HTTP/2 only and SSE here is HTTP/1.1).
Measured at the old 60s: 200, 61.07s, 2473 of 4000 emitted, message_stop 0,
error events 0, ending on a well-formed frame.
The recommended defence — reject a stream with no terminator — does NOT
transfer to us: Claude Code and opencode are third-party clients and we do not
own their SSE parsing. So this is acceptable because a request would have to run
1800s to trip it, not because we could detect it. If a member ever returns a
confident but truncated answer, suspect this before anything in our own code.
Also recorded, from the upstream bisection: ClientTrafficPolicy is honoured in
standalone `aigw run` but BackendTrafficPolicy is silently ignored, and nothing
external distinguishes them (envoyproxy/gateway#9513). Same silent-default shape
this repo keeps hitting.
bridged.yaml carries the same notes inline (gitignored, so not in this commit).
Refs: gitea #76
|
||
|
|
e689090024 |
CB-591: correct the root cause — Envoy's buffer limit, not Caddy
I wrote "Caddy request_body max_size and/or Envoy's own" and marked it
unverified. The Caddy half was wrong, and an unverified guess still points the
next reader at the wrong component.
Confirmed by the systems/vms side: Envoy Gateway defaults a listener's
per_connection_buffer_limit_bytes to 32768, and aigw buffers the WHOLE request
body before it can route on the model name. So that default is not a network
tuning knob — it is a hard ceiling on prompt size. From the live config_dump:
listener default/llm/http per_connection_buffer_limit_bytes: 32768
Nobody chose 32 KiB; it was inherited.
Both TLS edges are innocent, and the technique that showed it is better than
mine: both 413s carry x-llm-consumer, a header their auth proxy sets only AFTER
authenticating, so the body cleared both edges and the auth. On llm.vm, aigw
413s at 39 KB while the vLLM backend answers 200 at the same size. I found the
boundary; they found the component, by reading the failure's response headers.
Consequences recorded in the doc:
* DO NOT plan around 32 KiB. The intended ceiling is far higher, so sizing our
profiles to it would be designing around a bug.
* Their fix (ClientTrafficPolicy, bufferLimit: 8Mi) is written but NOT
deployed, pending their operator's approval. We do not re-test until they
confirm — a half-changed system gives a number neither side can trust.
* In standalone `aigw run` a SecurityPolicy is accepted and then silently
ignored, so "the config was accepted" proves nothing there. They will verify
by re-reading the live config_dump and sending a large request. Same
silent-default shape this repo keeps hitting, one layer down.
bridged.yaml carries the same correction (gitignored, so not in this commit).
Refs: gitea #76
|
||
|
|
1cc34888fd |
CB-591: record the live result — blocked by a 32 KiB body limit at the gateway
Deployed U1-U2c, restarted, spawned both new profiles for real, then reverted.
llm.ltms.dev answers HTTP 413 above 32 KiB (32768 bytes), on BOTH surfaces:
/v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200
/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413
That is far below one agent turn. It is an edge limit (Caddy request_body
max_size, and/or Envoy), so the fix is in systems/vms, not here.
The part worth recording is how it nearly passed. Two members, same message,
same moment: `local` finished in 66s, `gx` never finished at all. `local`
passed only because the probe was three trivial questions in a fresh session,
so the request fit under 32 KiB — the profile looked healthy and was a
landmine set to fire on the first turn that reads a file. So §7's checklist
was not wrong, it was too easy; it now says to use a file-reading task.
opencode's failure mode is worse than a crash: it catches the 413, compacts
its context, retries, and loops. Observed 10+ minutes BUSY with no reply. From
the lead's side that is indistinguishable from a slow worker. Reproduced
outside the bridge with the launcher's own generated config, which is how it
became a one-line error instead of a hang; §7.1 records that procedure.
Everything else about the migration checked out and is recorded so it is not
re-tested: token accepted on both surfaces, unauthenticated 401 (the Caddy
proxy does gate, whatever the gateway's own fail-open policy does),
/v1/models exactly ["deepseek-v4-flash"], the guard allowlist accepted
llm.ltms.dev, and the generated opencode provider block is correct with a real
llmk- key.
Also answers §3b's open question: reasoning survives BOTH surfaces —
/anthropic returns a real "type":"thinking" block and /v1 returns a populated
reasoning_content. The feared /v1 translation loss did not happen.
Config state (bridged.yaml is gitignored, so it is described rather than
committed): `local` back on http://gx00.gw:8000, `gx` kept at weight 0,
`local-direct` kept, llm.ltms.dev left in the guard allowlist. The file
carries these numbers and the exact two-key edit to switch back.
Verified after the revert with a task that reads two large files: correct on
all three questions. Daemon pid 66745, jar f1fd659423e6.
Refs: gitea #76
|
||
|
|
f0095bf8b2 |
CB-591: plan the move onto the LLM/MCP gateway, and check AI_GATEWAY_TOKEN
The gateway (llm.ltms.dev) replaced Bifrost on 2026-08-15 and serves an Anthropic surface and an OpenAI surface, so both member kinds can point at it. The plan is in docs/CB-591-Gateway-Migration.md; gitea #76 tracks the work. The opencode half needs no code: OpenCodeLauncher already pins an OpenAI-compatible endpoint (CB-508), so baseUrl + tokenEnv + provider/model is a config change. That matters more than it looks — every opencode member today is sol or terra, and both sit on one OpenAI account via credentialId: openai-shared, so an exhaustion on either locks out both. A gateway-backed opencode profile is free and off that credential, which retires a single point of failure rather than only adding capacity. Also extends the redeploy script's --check to AI_GATEWAY_TOKEN. A profile's tokenEnv is resolved from the DAEMON's own environment by HerdrPeerLauncher.resolveEnv, so a token added to secrets.sh after the daemon started is simply absent: the launcher injects an empty token and the gateway answers 401, long after the restart and with nothing tying the two together. That is the same trap as WORKER_GITEA_TOKEN, and it gets the same login-shell check that never prints the value. |