9daf1ec5ba
Closes out single-host before the cross-host work. Sequenced BEFORE CB-308 deliberately: federation's own gating concern is the trust model, and it inherits whatever identity shape lands here. The finding this stage is built around: bridged had exactly ONE security control, the loopback bind. ConnectionIdentity resolves a worker from its connection (unforgeable), but every caller that was not a recognised worker pane fell through to being treated as the PRIMARY -- the most privileged role on the bus. Latent today; load-bearing the moment a bind widens. CB-501 auth: - Role/Principal/CallerResolver: connection identity first, bearer token second, ANONYMOUS third. Inverts the old default so absence of identity means nothing, not everything. - Worker identity is never token-gated, so enabling auth cannot lock the fleet out of bridge_reply. - Constant-time token compare (MessageDigest.isEqual). - validateAuthExposure(): a non-loopback bind under loopback-trust now REFUSES TO START. Makes the dangerous config unrepresentable rather than merely documented. - TLS terminates at a reverse proxy by design (D3), not in the JVM. CB-505 authz + audit, enforced on BOTH entry paths: - The docs describe MCP as "a thin adapter over the REST core"; at code level it is not. BridgeMcp calls MessageService directly, and /mcp is a raw servlet on Jetty's context handler that never traverses Javalin's before filter. Enforcing only at REST would have left /mcp open. - Load-bearing rule is own-session-only: a worker may reply/ask only as itself. Structurally true over MCP already; over REST the session id in the URL path had simply been trusted. - Audit: JSON lines to a dedicated appender, additivity=false. Never records message content -- this bus carries source and prompts. CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer instead of the specced Micrometer, because this pom already hand-pins jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated dependency CVE gate could not be run (no JetBrains MCP server connected). Instrumented at MessageService, the single funnel both surfaces share. CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner. Needs no contract-exclusion flag -- the pom's default-excludes profile already sets excludedGroups=contract, so plain `mvn clean install` IS the mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older). CB-504 supervision: launchd agent (the real target -- this host is macOS, there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds. Ordering directives are advisory, so the actual fix is that startup now waits up to 30s for the herdr socket and then serves degraded, instead of crashing into a restart loop on a boot-order race. Also fixes drift found while surveying: - bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config binds via plain Jackson with ignoreUnknown, so uncommenting it would have been silently dropped and the default kept. Now camelCase, with a test that loads the shipped example and one that pins every documented knob's spelling -- no test had ever loaded that file. - Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay, gitTokenEnv, gitHostEnv, configDir, primary:). - README "Next" listed bridge_ask and session lifecycle as upcoming; both shipped long ago. - docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre- implementation" for work already merged. 307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS. Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check could not be run -- no JetBrains/intellij-index MCP server is connected this session. mvn clean install is the only gate that ran.
157 lines
9.4 KiB
YAML
157 lines
9.4 KiB
YAML
# bridged configuration (example). Copy to bridged.yaml and adjust.
|
|
#
|
|
# bridged is the sole gateway between primary/worker Claude sessions and herdr.
|
|
# It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL.
|
|
|
|
# REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token`
|
|
# below — bridged REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
|
|
bind:
|
|
host: 127.0.0.1
|
|
port: 8765
|
|
|
|
# API authentication (CB-501). Governs how a caller that is NOT an on-host worker pane proves it
|
|
# is the primary. Worker identity never depends on this: a loopback peer PID that maps to a herdr
|
|
# pane is unforgeable and is always honoured, so turning auth on cannot lock the fleet out.
|
|
#
|
|
# mode: loopback-trust → DEFAULT, and the historical behaviour: any loopback caller that is not
|
|
# a worker is the primary, no credential needed. Sound ONLY because the
|
|
# OS refuses remote connections to a loopback socket.
|
|
# mode: token → such a caller must send `Authorization: Bearer <token>`; without it it
|
|
# is anonymous and authorized for nothing. REQUIRED for a non-loopback
|
|
# bind — the daemon fails fast otherwise, because "unauthenticated ⇒
|
|
# primary" on a reachable port would hand spawn/stop/send to anyone.
|
|
# tokenEnv → host env var holding the token (never the literal value). Default
|
|
# BRIDGED_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
|
|
#
|
|
# TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and
|
|
# let it own certificate lifecycle, e.g.
|
|
# location / { proxy_pass http://127.0.0.1:8765; proxy_set_header Authorization $http_authorization; }
|
|
# The broker link gets TLS from its own URI (amqps://…) — see `broker` below.
|
|
# auth:
|
|
# mode: token
|
|
# tokenEnv: BRIDGED_API_TOKEN
|
|
|
|
# herdr Unix socket. Omit to use the client default
|
|
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
|
|
herdrSocket: ~/.config/herdr/herdr.sock
|
|
|
|
# How worker sessions are spawned. Define one or more named profiles (backends) under
|
|
# `workers`; each key is the profile name (also the ccs profile). `defaultWorker` picks
|
|
# which one a no-argument spawn uses (bridge_spawn with no profile / POST /workers).
|
|
#
|
|
# Shared knobs (placement/workspace/tabLabel) can be repeated per profile; they usually match.
|
|
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
|
|
# Use `pane` for the legacy behaviour (split the focused tab).
|
|
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
|
|
# (--append-system-prompt) as launch flags; nothing is written to the profile.
|
|
# tokenEnv → host env var holding the worker's auth token (value never stored in config);
|
|
# omit for a backend that needs no token (e.g. a local ollama).
|
|
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
|
|
# cwd on an MCP spawn, else the daemon's cwd — never $HOME. See
|
|
# docs/Worker-Startup-and-Trust.md.
|
|
# configDir → CLAUDE_CONFIG_DIR for the worker, so it inherits that profile's
|
|
# skills/MCP/hooks. Omit to leave the worker on the host default.
|
|
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
|
|
# worktree sees the same local config (CB-301-ext). Omit for the default set:
|
|
# [.mcp.json, .claude/settings.local.json, .env, .envrc].
|
|
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
|
|
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
|
|
# Opt-in by design — omit and the worker gets no PR-create grant (push over
|
|
# SSH is unaffected). The token value itself is never stored in this file.
|
|
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
|
|
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
|
|
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
|
|
workers:
|
|
gx10: # ccs profile name (NOT a hostname)
|
|
kind: claude-code # which adapter spawns this profile (default; may omit)
|
|
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
|
|
model: coder
|
|
placement: tab
|
|
workspace: bridged-workers
|
|
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
|
|
mcpUrl: http://127.0.0.1:8765/mcp
|
|
tokenEnv: BRIDGED_WORKER_TOKEN
|
|
argv: ["ccs", "gx10"]
|
|
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
|
|
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
|
|
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
|
|
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
|
|
# parityOverlay: [".mcp.json", ".claude/settings.local.json", ".env", ".envrc"]
|
|
ollama:
|
|
baseUrl: http://ollama.ltms.dev # local/self-hosted; usually no token
|
|
placement: tab
|
|
workspace: bridged-workers
|
|
tabLabel: "worker: {profile} #{n}"
|
|
mcpUrl: http://127.0.0.1:8765/mcp
|
|
argv: ["ccs", "ollama"]
|
|
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
|
|
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
|
|
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
|
|
# The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a
|
|
# `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude.
|
|
# opencode-gemini:
|
|
# kind: opencode
|
|
# model: google/gemini-2.5-pro # opencode `provider/model` selector, injected as `-m`
|
|
# placement: tab
|
|
# workspace: bridged-workers
|
|
# tabLabel: "opencode: {model} #{n}"
|
|
# mcpUrl: http://127.0.0.1:8765/mcp
|
|
# argv: ["opencode"]
|
|
defaultWorker: gx10
|
|
|
|
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
|
|
# must carry none. Every profile above must have its host listed here.
|
|
guard:
|
|
offSubscriptionHosts:
|
|
- gx00.gw
|
|
- gx01.gw
|
|
- ollama.ltms.dev
|
|
|
|
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
|
|
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
|
|
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
|
|
# unknown keys are ignored, so a snake_case key would be silently dropped (default kept).
|
|
# spawnReadyTimeoutMs: 20000
|
|
# spawnReadyPollMs: 300
|
|
|
|
# Worktree provisioning root (CB-301-ext). Where per-worker git worktrees are checked out so
|
|
# each worker owns an isolated branch instead of sharing the primary's tree. Omit to default
|
|
# to a sibling directory of the repo root.
|
|
# worktreeRoot: /Users/me/src/.bridged-worktrees
|
|
|
|
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
|
|
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
|
|
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
|
|
# contextCap → force-release a session after this many delegated turns
|
|
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
|
|
# lifecycle:
|
|
# idleTtlSeconds: 300
|
|
# contextCap: 10
|
|
# drainTimeoutSeconds: 5
|
|
|
|
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
|
|
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
|
|
# Set a broker uri to swap in the AMQP-backed inbox: worker replies with no open send are held
|
|
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
|
|
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
|
|
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
|
|
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
|
|
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
|
|
# broker:
|
|
# uri: amqp://guest:guest@127.0.0.1:5672
|
|
|
|
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
|
|
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
|
|
# pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop
|
|
# stops as soon as the primary's inbox is empty.
|
|
# terminal → pin the primary's herdr terminal id. Omit to learn it from the connection on
|
|
# the first orchestration-side MCP call (the normal case). An off-host or
|
|
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
|
|
# degrades to pull; the reply is still never lost.
|
|
# pushReminders → max nudges before giving up (default 5)
|
|
# pushBackoffMs → delay between nudges in ms (default 15000)
|
|
# primary:
|
|
# terminal: term_65619bd6174568
|
|
# pushReminders: 5
|
|
# pushBackoffMs: 15000
|