bf616e192a
profiles.<name>.subscription is read in five places (BridgedConfig record + isSubscription + validateSubscriptionProfiles, ClaudeCodeLauncher, and the startup secret check that deliberately skips it) and was in bridged.example.yaml nowhere. bridged.yaml is gitignored, so the example is the only place an operator can learn a key exists - which meant a fresh host had no way to discover the one switch that moves cost onto the operator's own subscription. Two profiles here have it set. Documents what changes when it is true, that it is mutually exclusive with baseUrl and refused at load, that the startup secret check skips such profiles so a clean secrets report says nothing about them, and that maxLoad is the only throttle against the operator's plan - there is no metering or budget refusal. 94 config tests pass; the example still loads.
626 lines
41 KiB
YAML
626 lines
41 KiB
YAML
# bridged configuration (example). Copy to bridged.yaml and adjust.
|
|
#
|
|
# bridged is the sole gateway between primary/worker Claude sessions and herdr.
|
|
# It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL.
|
|
|
|
# REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token`
|
|
# below — bridged REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
|
|
bind:
|
|
host: 127.0.0.1
|
|
port: 8765
|
|
|
|
# API authentication (CB-501). Governs how a caller that is NOT an on-host worker pane proves it
|
|
# is the primary. Worker identity never depends on this: a loopback peer PID that maps to a herdr
|
|
# pane is unforgeable and is always honoured, so turning auth on cannot lock the fleet out.
|
|
#
|
|
# mode: loopback-trust → DEFAULT, and the historical behaviour: any loopback caller that is not
|
|
# a worker is the primary, no credential needed. Sound ONLY because the
|
|
# OS refuses remote connections to a loopback socket.
|
|
# mode: token → such a caller must send `Authorization: Bearer <token>`; without it it
|
|
# is anonymous and authorized for nothing. REQUIRED for a non-loopback
|
|
# bind — the daemon fails fast otherwise, because "unauthenticated ⇒
|
|
# primary" on a reachable port would hand spawn/stop/send to anyone.
|
|
# tokenEnv → host env var holding the token (never the literal value). Default
|
|
# BRIDGED_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
|
|
#
|
|
# TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and
|
|
# let it own certificate lifecycle, e.g.
|
|
# location / { proxy_pass http://127.0.0.1:8765; proxy_set_header Authorization $http_authorization; }
|
|
# The broker link gets TLS from its own URI (amqps://…) — see `broker` below.
|
|
# auth:
|
|
# mode: token
|
|
# tokenEnv: BRIDGED_API_TOKEN
|
|
|
|
# Optional pinned primary terminal (CB-307). Names the herdr pane the PRIMARY itself runs in:
|
|
# a caller whose connection maps to this pane resolves as the primary (no credential needed —
|
|
# the pane mapping is as unforgeable as a worker's), and reply nudges are pushed to it.
|
|
# REQUIRED when the primary runs inside a herdr pane — without it the pane match reads the
|
|
# primary as a worker and refuses spawn/send/stop. Get the id from bridge_whoami; re-pin if
|
|
# the primary moves panes.
|
|
# primary:
|
|
# terminal: term_0123456789abcd
|
|
# pushReminders: 5 # max nudges before giving up (default 5)
|
|
# pushBackoffMs: 15000 # delay between nudges (default 15000)
|
|
|
|
# CB-530: MORE THAN ONE LEAD. `primary:` above is singular by construction — every other pane
|
|
# resolves as a worker — which is right for one lead driving a fleet and wrong the moment two leads
|
|
# (say a Claude lead and an opencode lead) work as peers: the second is silently demoted and refused
|
|
# every orchestration call. List each lead's pane here and all of them resolve as leads.
|
|
#
|
|
# tab → the ONLY field identity depends on (CB-579); the exact label of the tab hosting the lead.
|
|
# Label the tab yourself, or let bridged label one it launches — see `fleet.leaders:` below.
|
|
# kind/model → descriptive; they document what runs in the pane and are echoed by bridge_whoami
|
|
#
|
|
# A lead's tab must already carry its label (or be launched by bridged, which labels it) — there is
|
|
# no terminal id to paste in and nothing to re-pin when the session restarts: the tab survives, so
|
|
# the same label resolves the same lead again on the next scan.
|
|
# `bridge_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
|
|
# IS a primary for authorization, so nothing that keys on the role breaks.
|
|
#
|
|
# KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single
|
|
# destination for its nudges, and is a separate mechanism from lead identity — see `fleet.leaders:`.
|
|
#
|
|
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
|
|
#
|
|
# Two things stop the tab-name convention from becoming a way to claim leadership: the configured
|
|
# member spaces are excluded from the scan, so nothing bridged places can land in a matching tab;
|
|
# and startup REFUSES a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel`
|
|
# override, also matches — so the two namespaces cannot overlap by accident. The label is a NAME,
|
|
# never a capability: what a pane may do is decided by the role the daemon resolves for it.
|
|
|
|
# CB-551: IDLE-LEAD HEARTBEAT — nudge the single lead back to work when it has been continuously
|
|
# idle (no open bridge_send driving it) past the quiet period. The fleet is one lead + architects +
|
|
# workers, so a lead that stalls is a single point of failure; the ReplyPushLoop only nudges when a
|
|
# reply lands, and this timer catches the gap where nothing lands and the lead just sits idle.
|
|
#
|
|
# Opt-in on purpose — it SPENDS the operator's subscription on its own initiative (each nudge starts
|
|
# a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent
|
|
# block = feature off, exactly as before.
|
|
#
|
|
# Three knobs, each with a default that errs on the side of not burning context:
|
|
# idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 —
|
|
# # absorbs normal post-turn pauses; re-prompting every pause burns context)
|
|
# backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000)
|
|
# quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops
|
|
# # until real state appears (default 3 — never nag an empty fleet forever)
|
|
# leadHeartbeat:
|
|
# idleAfterSeconds: 300
|
|
# backoffMs: 60000
|
|
# quietNudgeCap: 3
|
|
|
|
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
|
|
# per tick.
|
|
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
|
|
# Math.max(15, intervalSeconds), so a lower value is silently raised, not
|
|
# rejected.
|
|
# workingSuspectAfterSeconds, paneProbeIntervalSeconds → accepted and parsed, but NOT YET READ by
|
|
# anything — the dormant monitor only consumes intervalSeconds today (CB-573
|
|
# shipped ahead of the evidence publishers these two knobs are for). Setting
|
|
# them changes nothing right now, and no minimum is enforced on either, because
|
|
# nothing reads them to enforce one. They exist so a later build can start
|
|
# honouring them without another config-shape change.
|
|
# notifications.mode → "webhook" flips what bridge_list REPORTS (healthCoverage: "full" instead
|
|
# of "detection-only") — it does NOT make bridged send any webhook call; no
|
|
# delivery mechanism is implemented yet. Any other value, or omitting the
|
|
# block, reports "detection-only".
|
|
# health:
|
|
# enabled: true
|
|
# intervalSeconds: 30
|
|
# workingSuspectAfterSeconds: 600
|
|
# paneProbeIntervalSeconds: 60
|
|
# notifications:
|
|
# mode: disabled
|
|
|
|
# herdr Unix socket. Omit to use the client default
|
|
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
|
|
herdrSocket: ~/.config/herdr/herdr.sock
|
|
|
|
# How member sessions are spawned. Define one or more named profiles (backends) under
|
|
# `profiles`; each key is the profile name (also the ccs profile). A profile says only WHICH
|
|
# BACKEND — model, CLI adapter, credentials, cost. It says nothing about what a member spawned on
|
|
# it is for; that is the member's role, and roles live under `fleet:` below. Which profile an
|
|
# unqualified spawn lands on comes from that role's pool, not from a global default.
|
|
#
|
|
# Shared knobs (placement/workspace) can be repeated per profile; they usually match.
|
|
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
|
|
# Use `pane` for the legacy behaviour (split the focused tab).
|
|
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
|
|
# (--append-system-prompt) as launch flags; nothing is written to the profile.
|
|
# tokenEnv → host env var holding the worker's auth token (value never stored in config);
|
|
# omit for a backend that needs no token (e.g. a local ollama).
|
|
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
|
|
# cwd on an MCP spawn, else the daemon's cwd — never $HOME. See
|
|
# docs/Worker-Startup-and-Trust.md.
|
|
# configDir → CLAUDE_CONFIG_DIR for the worker, so it inherits that profile's
|
|
# skills/MCP/hooks. Omit to leave the worker on the host default.
|
|
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
|
|
# worktree sees the same local config (CB-301-ext). Omit for the default set:
|
|
# [.claude/settings.local.json, .env, .envrc].
|
|
#
|
|
# Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher
|
|
# mounts — the bridge, and nothing else. Replicating the primary's MCP config
|
|
# handed a worker the primary's IDE servers, which are bound to the primary's
|
|
# checkout, so its navigation returned paths OUTSIDE its own worktree: one
|
|
# worker made all 59 of its edits in the primary tree while compiling its
|
|
# worktree, and every build it ran was of code that did not contain them.
|
|
# bridged neutralizes a provisioned worktree's .mcp.json for this reason;
|
|
# listing it here would copy the primary's back over that.
|
|
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
|
|
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
|
|
# Opt-in by design — omit and the worker gets no PR-create grant (push over
|
|
# SSH is unaffected). The token value itself is never stored in this file.
|
|
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
|
|
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
|
|
# exhaustedPattern → regex matched against a completion-fallback scrape (CB-578 stage A) to
|
|
# classify a turn that ended with no bridge_reply as the backend having
|
|
# refused on a subscription usage limit, rather than a real answer. Opt-in —
|
|
# omit and this profile's completion fallback behaves exactly as before.
|
|
# Every backend words its refusal differently, so this is config, never a
|
|
# vendor string baked into bridged itself.
|
|
# DEFERRED: compiled once into a startup pattern map — editing it needs a
|
|
# daemon restart, same as this profile's model/baseUrl/argv.
|
|
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
|
|
# BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME
|
|
# credentialId share one quarantine — the case this exists for is two models
|
|
# on one account (e.g. sol and terra both billing one OpenAI credential): an
|
|
# exhaustion on either one must lock out both, or the fleet just walks onto
|
|
# the same dead account under the sibling's name. Opt-in — omit and this
|
|
# profile quarantines alone, under its own name, exactly as if the field did
|
|
# not exist. Cooldown length is the top-level quarantineCooldownSeconds below.
|
|
# HOT: read live at every spawn/exhaustion check — no restart needed.
|
|
# env → extra environment for this profile's workers, as a literal key/value map
|
|
# (CB-511). Use it to give workers a toolchain.
|
|
#
|
|
# A worker's environment does NOT come from your shell. bridged hands herdr an
|
|
# explicit env map and herdr merges it into ITS OWN process env — so before
|
|
# CB-511 a worker inherited whatever PATH the herdr server happened to be
|
|
# started with, which on a long-lived herdr can predate your toolchain entirely
|
|
# and leave workers unable to run `mvn` or `java` at all.
|
|
# bridged now propagates ITS OWN PATH to every worker by default; set `env:`
|
|
# only to override that or add more (JAVA_HOME, …). Since the default is the
|
|
# daemon's PATH, make sure the daemon is started with a good one — see the PATH
|
|
# lines in deploy/dev.ltms.bridged.plist and deploy/bridged.service.
|
|
#
|
|
# Adapter-owned variables always win over `env:`: ANTHROPIC_BASE_URL and the
|
|
# rest of the ANTHROPIC_*/CLAUDE_* wiring are applied after it, so an `env:`
|
|
# entry cannot repoint a worker past the SubscriptionGuard — which is checked
|
|
# against `baseUrl` alone.
|
|
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
|
|
profiles:
|
|
gx10: # ccs profile name (NOT a hostname)
|
|
kind: claude-code # which adapter spawns this profile (default; may omit)
|
|
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
|
|
model: coder
|
|
placement: tab
|
|
workspace: bridged-workers
|
|
# tabLabel: an optional per-profile override; the fleet template usually covers it
|
|
mcpUrl: http://127.0.0.1:8765/mcp
|
|
tokenEnv: BRIDGED_WORKER_TOKEN
|
|
argv: ["ccs", "gx10"]
|
|
# weight: relative selection weight for automatic placement (weighted, round-robin, and
|
|
# fixed's fallback walk). Absent defaults to 1.0. An explicit 0 or negative value means
|
|
# "never auto-select this profile" (CB-554) — it stays reachable via an explicit
|
|
# `bridge_spawn{profile:"gx10"}`, which bypasses placement entirely; only automatic
|
|
# selection skips it.
|
|
weight: 0.5
|
|
# maxLoad: max live workers on this profile. Omit for unlimited. An explicit 0 (CB-585) caps
|
|
# the profile at zero live members — it is excluded from automatic placement and an explicit
|
|
# `bridge_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
|
|
# is named directly. Negative is refused at config load — there is no sane meaning for it.
|
|
maxLoad: 2
|
|
# subscription: true
|
|
# THE KNOB THAT DECIDES WHO PAYS (CB-539). Default false. When true, this profile's members
|
|
# run on the OPERATOR'S OWN Claude subscription instead of a metered endpoint — every spawn
|
|
# bills your plan and eats your usage limit. Off-subscription is the whole point of this
|
|
# daemon, so treat `true` as a deliberate exception, not a convenience.
|
|
#
|
|
# What changes when it is set (ClaudeCodeLauncher):
|
|
# - no ANTHROPIC_BASE_URL and no ANTHROPIC_AUTH_TOKEN are injected — the member inherits
|
|
# the operator's own Claude Code auth, which is exactly why it bills the plan;
|
|
# - SubscriptionGuard never vets it, because there is no baseUrl to vet;
|
|
# - no token is required, so `tokenEnv` is irrelevant here.
|
|
#
|
|
# MUTUALLY EXCLUSIVE with `baseUrl` — setting both is refused at config load (CB-542). On the
|
|
# subscription path no guard would vet the URL, so allowing both would be a way around the
|
|
# guard rather than a configuration.
|
|
#
|
|
# GOTCHA 1 — it is invisible to the startup secret check. `Bridged.reportRequiredSecrets`
|
|
# skips subscription profiles on purpose (they need no token), so a boot log that reports
|
|
# every secret as fine says nothing about these profiles.
|
|
#
|
|
# GOTCHA 2 — `maxLoad` is the ONLY throttle you have here. There is no metering, no budget
|
|
# and no refusal on cost; the cap on live members is the single thing standing between a
|
|
# fan-out and your monthly limit. Set it deliberately and keep it small.
|
|
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
|
|
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
|
|
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
|
|
# credentialId: shared-openai # opt-in: quarantine together with every other profile sharing this id (CB-578)
|
|
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
|
|
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
|
|
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
|
|
gx11: # a second backend, so `placement: weighted` has a choice
|
|
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
|
|
placement: tab
|
|
workspace: bridged-workers
|
|
# tabLabel: an optional per-profile override; the fleet template usually covers it
|
|
mcpUrl: http://127.0.0.1:8765/mcp
|
|
argv: ["ccs", "gx11"]
|
|
weight: 0.5
|
|
maxLoad: 2
|
|
# Pin an auto-compact window BELOW the served model's context ceiling. The global
|
|
# ~/.claude/settings.json value is shared by every ccs instance and the primary, so the
|
|
# per-profile override belongs here. Equal to the ceiling means auto-compact never fires
|
|
# before the server rejects the prompt, which kills a worker mid-turn (CB-523).
|
|
env:
|
|
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "280000"
|
|
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
|
|
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
|
|
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
|
|
# The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a
|
|
# `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude.
|
|
#
|
|
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → bridge_send →
|
|
# structured bridge_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
|
|
# and need NO credentials — check `opencode models` for the current free list, since the names
|
|
# change. That also makes the worker off-subscription by construction.
|
|
# opencode-free:
|
|
# kind: opencode
|
|
# model: opencode/north-mini-code-free # `provider/model` selector, injected as `-m`
|
|
# placement: tab
|
|
# workspace: bridged-workers
|
|
# tabLabel: "opencode: {profile} #{n}"
|
|
# mcpUrl: http://127.0.0.1:8765/mcp
|
|
# argv: ["opencode"]
|
|
#
|
|
# CB-508: point an opencode profile at your OWN OpenAI-compatible endpoint (local vLLM, llama.cpp,
|
|
# LM Studio, TGI…) instead of opencode's gateway. Setting `baseUrl` on a `kind: opencode` profile
|
|
# makes the bridge emit a custom `provider` block into the generated opencode.json — opencode has
|
|
# no ANTHROPIC_BASE_URL seam, so this is how the endpoint is pinned.
|
|
# baseUrl → a bare host:port gets `/v1` appended (where these servers mount the API); a URL that
|
|
# already has a path is used verbatim, so a custom mount point still works.
|
|
# model → MUST be "<provider>/<model>". The provider half names the generated block; the model
|
|
# half must match an id the server reports at /v1/models. One field drives both the
|
|
# declaration and the `-m` flag, so they cannot drift apart. A bare model name with a
|
|
# baseUrl set is rejected at spawn rather than silently using the default gateway.
|
|
# tokenEnv → optional; its value becomes the provider apiKey. Most local servers ignore the key,
|
|
# so a placeholder is used when unset (the AI SDK still requires a non-empty one).
|
|
# NOTE: no `guard` entry is needed even with a baseUrl set. The SubscriptionGuard exists to stop a
|
|
# worker borrowing the primary's Anthropic subscription, and an opencode process has no Anthropic
|
|
# credential path at all.
|
|
# opencode-local:
|
|
# kind: opencode
|
|
# baseUrl: http://127.0.0.1:8000
|
|
# model: local-vllm/deepseek-v4-flash
|
|
# placement: tab
|
|
# workspace: bridged-workers
|
|
# tabLabel: "opencode: {profile} #{n}"
|
|
# mcpUrl: http://127.0.0.1:8765/mcp
|
|
# argv: ["opencode"]
|
|
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
|
|
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
|
|
#
|
|
# `weighted` IS NOT "cheapest first" — read this before you set weights (CB-589).
|
|
# It is smooth weighted round-robin: it spreads spawns across EVERY profile that has a free slot,
|
|
# in weight ratio. It has no idea which profile costs money. So with local:10 / paid:2 you do not
|
|
# get "use local, overflow to paid" — you get roughly one spawn in six going to the paid profile
|
|
# while the local box still has a free slot.
|
|
#
|
|
# There is a sharper second effect. The policy's running score map lives for the daemon's whole
|
|
# life. While a profile is at maxLoad it is filtered out and its score FREEZES, so the paid
|
|
# profiles keep accumulating against it. When the local slot frees up it returns with a stale
|
|
# score and can LOSE the next pick — a paid spawn while the free box sits idle.
|
|
#
|
|
# Until a real cost-first policy exists, the workaround is to make the ratio decisive rather than
|
|
# proportional: give the free profile a weight so large that it wins every pick it is eligible
|
|
# for, and paid profiles only ever take genuine overflow. On this host that is local weight 100
|
|
# against paid weights of ~1.
|
|
#
|
|
# The gotcha with that workaround: it expresses a PREFERENCE ORDER through a RATIO knob. Add a
|
|
# future profile at weight 150 and it silently outranks the free box, with nothing to warn you.
|
|
# Re-check the weights whenever you add a profile.
|
|
placement: weighted
|
|
|
|
# How long a credential sits out after a BACKEND_EXHAUSTED classification (CB-578 stage B), in
|
|
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
|
|
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
|
|
# 1800 (30 minutes) when omitted or non-positive.
|
|
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
|
|
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
|
|
# restart. Editing this needs a daemon restart to take effect.
|
|
# quarantineCooldownSeconds: 1800
|
|
|
|
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
|
|
# upgraded bridged keeps the old behaviour: the file is read once at boot and never again.
|
|
# enabled → turn the watch on. bridged checks the file's modified time on a timer and
|
|
# reloads when it moves.
|
|
# intervalSeconds → how often to check (default 10). One `stat` per tick, so this is cheap.
|
|
#
|
|
# Not every key can move under a running daemon, and the difference is about what already exists
|
|
# when the reload happens — not about how important the key is:
|
|
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
|
|
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad
|
|
# / credentialId. Those are hot because the placement policy (and, for credentialId,
|
|
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
|
|
# not by itself enough to make a key hot.
|
|
# EXCEPT `fleet.leaders`: Bridged.main reads it once at startup to build the lead tab
|
|
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
|
|
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
|
|
# with nothing in the deferred list — but has NO effect until you restart. Treat it
|
|
# as deferred in practice, even though today's reload output does not say so.
|
|
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
|
|
# until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`,
|
|
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, `quarantineCooldownSeconds` (CB-578
|
|
# stage B — baked once into the quarantine tracker built at startup), ADDING or
|
|
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
|
|
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
|
|
# configDir, mcpUrl, tabLabel, exhaustedPattern. The launcher takes a copy of
|
|
# `profiles:` at startup and resolves every spawn out of that copy, so those never
|
|
# reach a launch until you restart. The reload logs them by name rather than
|
|
# pretending they applied.
|
|
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
|
|
# bound, the broker connection is open, and the auth mode decides who may reach the
|
|
# port that is already listening.
|
|
#
|
|
# A changed COLD key refuses the WHOLE reload — not the hot half applied and the cold half warned
|
|
# about. A half-applied reload would leave the daemon matching no file on disk, which is the worst
|
|
# thing a reload can do to an operator debugging one. A file that fails to parse or fails a startup
|
|
# validator is refused the same way, and the running config stays live.
|
|
# configReload:
|
|
# enabled: true
|
|
# intervalSeconds: 10
|
|
|
|
# THE FLEET (CB-557) — who the daemon may run, and under which role. This one block replaced four
|
|
# older keys: `leaders:`, `members:`, `leadScan:` and `defaultProfile:`.
|
|
#
|
|
# A member is anything a lead spawns, and every member has two INDEPENDENT attributes:
|
|
# role — which contract: architect, dev or reviewer. It picks the launch charter, the role
|
|
# file, the playbook skill and the authz row.
|
|
# profile — which backend: one of the `profiles:` keys above (model, CLI adapter, cost).
|
|
# They vary on their own. A reviewer may run on the same profile as the dev whose diff it reads,
|
|
# which is why the two cannot be one field.
|
|
#
|
|
# The ROLE IS THE CONTAINING KEY, not a `role:` field. That is not only tidier: a misspelled role
|
|
# used to parse into a member with no contract at all, while a misspelled pool name here simply
|
|
# declares nothing.
|
|
#
|
|
# Each pool lists the profiles that role MAY run on — these are pools, not identities. That is also
|
|
# what replaced `defaultProfile:`: an unqualified spawn names a role, and that role's pool supplies
|
|
# the candidates, in definition order. A dev and a reviewer staying anonymous is exactly compatible
|
|
# with being listed here; the entry key just names the entry.
|
|
fleet:
|
|
# Optional launch-charter text, keyed only by the singular role wire names: architect, dev,
|
|
# reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put secrets
|
|
# here: a later launch step writes this text to a world-readable temp file, and ${ENV} interpolation
|
|
# is deliberately not supported.
|
|
charters:
|
|
architect: |-
|
|
You are an architect in this fleet. You refine work before anyone builds it:
|
|
scope, acceptance criteria, risks, and a unit split. You read the repo and
|
|
write analysis. You never commit production code and never open a PR.
|
|
A design task is worked by two architects. Design alone first, then exchange
|
|
and say plainly where you disagree. Do not concede just to agree.
|
|
dev: |-
|
|
You implement the one unit you were given, and nothing else. You test it,
|
|
commit it, and open your own pull request. You never merge.
|
|
reviewer: |-
|
|
You review the diff you were given. You report bugs, risks and missing tests.
|
|
You do not change code.
|
|
|
|
# Optional. Template for a member tab's label; {role}, {profile}, {model} and {n} are substituted.
|
|
# {n} counts per role+profile, so `dev: sonnet #2` really is the second sonnet dev. Because {role}
|
|
# comes from a closed enum, a generated label can never begin with a lead's tabPrefix.
|
|
# tabLabel: "{role}: {profile} #{n}"
|
|
|
|
# Panes that orchestrate rather than are orchestrated. A lead may now be CREATED as well as
|
|
# recognised: give it a `profile:` and the daemon launches the shortfall when fewer than
|
|
# `instances` are live. Omit `profile:` and it is recognise-only, as before.
|
|
#
|
|
# `tab:` (CB-579) is REQUIRED and is the only field identity depends on — the exact label of the
|
|
# tab hosting the lead, matched case-insensitively. Label the tab yourself and put that same
|
|
# string here, and the pane is recognised on the next rescan. Reopen the tab later, or the session
|
|
# inside it restarts — the terminal id changes; the tab, and its label, do not, so no config edit
|
|
# follows a restart.
|
|
#
|
|
# A lead the daemon launches is labelled BY the daemon with this same `tab:` value, so it is found
|
|
# by the same scan. A lead counts as live only when herdr also reports a running agent in that
|
|
# tab — a label left behind by a session that died does not block the relaunch, and a tab that is
|
|
# gone entirely drops out of the next scan rather than being remembered forever.
|
|
#
|
|
# An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with
|
|
# the session lifecycle (the idle reaper would kill your orchestrator), and stays on the
|
|
# subscription — ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.
|
|
#
|
|
# GET THE `tab:` VALUE RIGHT. A pane that does not match any configured `tab:` (a typo, a renamed
|
|
# tab, a pane no entry names at all) is not recognised as a lead — it resolves as an ordinary
|
|
# WORKER instead, silently, and every orchestration call it makes (spawn/stop/send/drain) is
|
|
# refused. There is no error at startup for this: an unmatched pane is simply not a lead. If your
|
|
# primary suddenly can't spawn or send, check this section first.
|
|
# leaders:
|
|
# opus-5.0:
|
|
# profile: opus # omit to never create this lead, only recognise it
|
|
# instances: 1 # desired live count; only the shortfall is launched. 0 = off
|
|
# tab: "lead: opus-5.0" # REQUIRED — the exact tab label this lead lives in
|
|
# tabPrefix: "lead:" # only used to guard against a worker tabLabel colliding with
|
|
# # this convention at startup; plays no part in matching a lead
|
|
# scanIntervalSeconds: 10 # rescan cadence, and the worst case before a new tab is seen
|
|
# workspace: leads # where a launched lead's tab is created (default "leads").
|
|
# # MUST NOT be a member workspace — those are excluded from the
|
|
# # scan, so a lead placed in one is never found again.
|
|
# cwd: /path/to/repo # the launched lead's working directory (default: bridged's own)
|
|
# kind: claude # descriptive; reported by bridge_whoami
|
|
# gpt-sol-5.6:
|
|
# tab: "lead: gpt-sol-5.6"
|
|
# kind: opencode
|
|
# model: openai/gpt-5.6-terra
|
|
|
|
# architects:
|
|
# architect-1:
|
|
# profile: opus # a strong model, on the operator's subscription
|
|
# architect-2:
|
|
# profile: sol # a different vendor on purpose — two architects that share a
|
|
# # model share its blind spots
|
|
developers:
|
|
gx10:
|
|
profile: gx10
|
|
# reviewers:
|
|
# gx10:
|
|
# profile: gx10 # the same backend may serve two roles; that is the point
|
|
|
|
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
|
|
# must carry none. Every profile above must have its host listed here.
|
|
guard:
|
|
offSubscriptionHosts:
|
|
- gx00.gw
|
|
- gx01.gw
|
|
|
|
# Member credential policy (CB-596, gitea issue #82). A herdr pane runs a LOGIN shell, and that
|
|
# shell re-sources the operator's own secret store — so a spawned member inherits every credential
|
|
# the operator's shell holds, not just the ones bridged means to give it. Measured on this host:
|
|
# 31 credential names, all set, with only ONE (GITEA_ACCESS_TOKEN) blocked before this — and that
|
|
# block was a single name hardcoded in HerdrPeerLauncher.java, not driven by this file. This block
|
|
# replaces that hardcoded shadow with a config-driven list of names.
|
|
#
|
|
# ROUND-2 CORRECTION, measured live: the pane-creation env overlay below (applied at tab.create /
|
|
# pane.split, BEFORE the pane's login shell runs) does NOT survive that login shell for any name
|
|
# secrets.sh actually exports — the shell re-exports it afterwards and overwrites the sentinel.
|
|
# Proof: GITEA_ACCESS_TOKEN comes back blocked only because secrets.sh itself carries a guarded
|
|
# export (`[ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...`) — that guard, not this
|
|
# file, is what wins. No other name in `known` below has a matching guard in secrets.sh yet (1
|
|
# guard measured against 33 export lines there). So today this block's overlay is REAL protection
|
|
# only for a name secrets.sh does not export, or a peer kind whose pane never runs a login shell —
|
|
# for everything secrets.sh exports and guards, the guard in secrets.sh (out of scope for this
|
|
# ticket) is what actually blocks it, not this list. An exec-time fix (winning after the login
|
|
# shell finishes, before the agent process starts) was attempted and found to have no seam in the
|
|
# current herdr protocol — AgentControl.start takes a fixed `kind` (herdr resolves the executable)
|
|
# plus trailing CLI args for that binary, not an arbitrary argv or an env map; only tab.create /
|
|
# pane.split accept `env`, and that is this same pane-creation overlay. See gitea #82 for the open
|
|
# design question this leaves.
|
|
#
|
|
# DENY-BY-DEFAULT, NOT A DENY-LIST. A deny-list (name the bad ones, let everything else through) is
|
|
# silently wrong the moment the operator's store gains a new secret — nothing would ever report it.
|
|
# Deny-by-default inverts that: `known` bounds the blast radius to names actually enumerated below,
|
|
# and EVERY one of them is blocked UNLESS it is also in `allow`. Omitting this block entirely (the
|
|
# shipped default) blocks NOTHING — unlike most optional blocks in this file, absence here is a real
|
|
# gap, not a safe "feature off". A name that is neither `known` nor `allow`-ed is not silently let
|
|
# through either: the daemon logs a WARN naming any credential-shaped env var it finds on neither
|
|
# list (never its value), so a secret added to the store later does not go unnoticed forever.
|
|
#
|
|
# policy → only "deny-by-default" exists today (an operator-authored deny-list was deliberately
|
|
# rejected — see above). An unrecognized value refuses to start, naming it.
|
|
# allow → credential names a member legitimately needs. Left OUT of the pane's env overlay
|
|
# entirely, so the value the pane's own (login) shell exports passes through untouched.
|
|
# known → every credential name the operator's store is known to export. Every name here NOT
|
|
# also in `allow` is overlaid with a non-secret sentinel value before the pane's login
|
|
# shell runs — real protection only for names the login shell does not itself re-export
|
|
# (see the ROUND-2 CORRECTION note above for the ones it does).
|
|
#
|
|
# HOT-RELOADABLE the same way `fleet:` is (CB-559): read fresh on every spawn, so editing this list
|
|
# and reloading config (or restarting) changes what the NEXT spawn inherits; already-running members
|
|
# are unaffected either way.
|
|
# memberCredentials:
|
|
# policy: deny-by-default
|
|
# allow:
|
|
# - AI_GATEWAY_TOKEN # named in a profile's tokenEnv (local/gx) — a member reaching the
|
|
# # gateway is by design, not a leak
|
|
# - WORKER_GITEA_TOKEN # the repo-scoped forge token a member needs to open its own PR (CB-302)
|
|
# - CONTEXT7_TOKEN # already decided as allowed by CB-593
|
|
# - GITEA_HOST # not a credential — a hostname, paired with the forge token above
|
|
# known:
|
|
# - AI_GATEWAY_TOKEN
|
|
# - BESZEL_ADMIN_EMAIL
|
|
# - BESZEL_ADMIN_PASSWORD
|
|
# - BESZEL_HUB_URL
|
|
# - BESZEL_KEY
|
|
# - BESZEL_UNIVERSAL_TOKEN
|
|
# - BRAIN_MCP_TOKEN
|
|
# - CF_ACCOUNT_ID
|
|
# - CF_API_TOKEN
|
|
# - CF_USER_TOKEN
|
|
# - CONFLUENCE_API_TOKEN
|
|
# - CONFLUENCE_USERNAME
|
|
# - CONTEXT7_TOKEN
|
|
# - GITEA_ACCESS_TOKEN
|
|
# - GITEA_HOST
|
|
# - GITLAB_OAUTH_CLIENT_SECRET
|
|
# - GITLAB_PERSONAL_ACCESS_TOKEN
|
|
# - GRAFANA_ADMIN_PASSWORD
|
|
# - GRAFANA_ADMIN_USER
|
|
# - HASS_TOKEN
|
|
# - HW_PASSWORD
|
|
# - HW_USER
|
|
# - LTMS_API_KEY
|
|
# - MEMORY_MCP_TOKEN
|
|
# - METRICS_PUSH_TOKEN
|
|
# - OPENCODE_AUTOMODE_MODEL
|
|
# - TELEGRAM_BOT_TOKEN
|
|
# - TELEGRAM_CHAT_ID
|
|
# - TS_API_KEY
|
|
# - TS_AUTHKEY
|
|
# - WORKER_GITEA_TOKEN
|
|
|
|
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
|
|
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
|
|
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
|
|
# unknown keys are ignored, so a snake_case key would be silently dropped (default kept).
|
|
# spawnReadyTimeoutMs: 20000
|
|
# spawnReadyPollMs: 300
|
|
|
|
# Worktree provisioning root (CB-301-ext). Where per-worker git worktrees are checked out so
|
|
# each worker owns an isolated branch instead of sharing the primary's tree. Omit to default
|
|
# to a sibling directory of the repo root.
|
|
# worktreeRoot: /Users/me/src/.bridged-worktrees
|
|
|
|
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
|
|
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
|
|
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
|
|
# contextCap → force-release a session after this many delegated turns
|
|
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
|
|
# clearAfterTurn → whether a reusable worker discards its conversation context after every
|
|
# completed delegated turn (default false). Works for claude-code workers
|
|
# only — any other peer kind (e.g. opencode) logs "context reset is
|
|
# unsupported for peer kind …" once and the reset is a no-op.
|
|
# lifecycle:
|
|
# idleTtlSeconds: 300
|
|
# contextCap: 10
|
|
# drainTimeoutSeconds: 5
|
|
# clearAfterTurn: false
|
|
|
|
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
|
|
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
|
|
# Set a broker uri to swap in the AMQP-backed inbox: worker replies with no open send are held
|
|
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
|
|
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
|
|
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
|
|
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
|
|
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
|
|
# prefetch → CB-527: consumer basicQos, capping how many unacked messages the inbox holds
|
|
# in-heap per owned target (the rest sits on the broker's durable queue instead of
|
|
# growing the JVM heap). Default 32 when omitted.
|
|
# broker:
|
|
# uri: amqp://guest:guest@127.0.0.1:5672
|
|
# prefetch: 32
|
|
|
|
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
|
|
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
|
|
# pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop
|
|
# stops as soon as the primary's inbox is empty.
|
|
# terminal → pin the primary's herdr terminal id. Omit to learn it from the connection on
|
|
# the first orchestration-side MCP call (the normal case). An off-host or
|
|
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
|
|
# degrades to pull; the reply is still never lost.
|
|
#
|
|
# REQUIRED (CB-522) if the primary itself runs inside a herdr pane. Caller
|
|
# identity resolves a loopback PID to its herdr pane, and PaneLocator scans
|
|
# EVERY pane — not just bridged-spawned ones — so such a primary is otherwise
|
|
# classified as a WORKER and refused SPAWN/SEND/STOP. That failure is
|
|
# self-locking: the learned terminal is populated by the very orchestration
|
|
# calls being refused, so only this pinned value can break the cycle. Read the
|
|
# id off bridge_whoami (it reports the current terminal even while
|
|
# misclassified) and re-pin whenever the primary moves panes.
|
|
# pushReminders → max nudges before giving up (default 5)
|
|
# pushBackoffMs → delay between nudges in ms (default 15000)
|
|
# primary:
|
|
# terminal: term_65619bd6174568
|
|
# pushReminders: 5
|
|
# pushBackoffMs: 15000
|