d811b30df3
Adds `ideMcpUrl` to FleetConfig.Profile (default off). When set, the Claude Code launcher mounts the IDE Index MCP as a second inline --mcp-config server named `intellij`, and appends an IDE charter that pins every ide_* call to the member's own worktree (spec.cwd()). The charter order is role -> ide -> reply, one --append-system-prompt-file, reply last (CB-618). The mount gate now fires on ideMcpUrl alone, not only mcpUrl. ConfigRef treats an ideMcpUrl change as deferred, like the other launch flags. Never touches .mcp.json or CLAUDE.md — the mount and the rule arrive as launch flags, so a project's own config is untouched. Not yet done (see fleetd #162): the bridged-owned IDE lifecycle (open on provision, close before worktree removal), and the opencode adapter (separate ticket). fleetd.example.yaml documents ideMcpUrl and fixes the stale parityOverlay default. 911 tests green.
668 lines
45 KiB
YAML
668 lines
45 KiB
YAML
# bridged configuration (example). Copy to bridged.yaml and adjust.
|
|
#
|
|
# bridged is the sole gateway between primary/worker Claude sessions and herdr.
|
|
# It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL.
|
|
|
|
# REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token`
|
|
# below — bridged REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
|
|
bind:
|
|
host: 127.0.0.1
|
|
port: 8765
|
|
|
|
# API authentication (CB-501). Governs how a caller that is NOT an on-host worker pane proves it
|
|
# is the primary. Worker identity never depends on this: a loopback peer PID that maps to a herdr
|
|
# pane is unforgeable and is always honoured, so turning auth on cannot lock the fleet out.
|
|
#
|
|
# mode: loopback-trust → DEFAULT, and the historical behaviour: any loopback caller that is not
|
|
# a worker is the primary, no credential needed. Sound ONLY because the
|
|
# OS refuses remote connections to a loopback socket.
|
|
# mode: token → such a caller must send `Authorization: Bearer <token>`; without it it
|
|
# is anonymous and authorized for nothing. REQUIRED for a non-loopback
|
|
# bind — the daemon fails fast otherwise, because "unauthenticated ⇒
|
|
# primary" on a reachable port would hand spawn/stop/send to anyone.
|
|
# tokenEnv → host env var holding the token (never the literal value). Default
|
|
# BRIDGED_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
|
|
#
|
|
# TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and
|
|
# let it own certificate lifecycle, e.g.
|
|
# location / { proxy_pass http://127.0.0.1:8765; proxy_set_header Authorization $http_authorization; }
|
|
# The broker link gets TLS from its own URI (amqps://…) — see `broker` below.
|
|
# auth:
|
|
# mode: token
|
|
# tokenEnv: BRIDGED_API_TOKEN
|
|
|
|
# Optional pinned primary terminal (CB-307). Names the herdr pane the PRIMARY itself runs in:
|
|
# a caller whose connection maps to this pane resolves as the primary (no credential needed —
|
|
# the pane mapping is as unforgeable as a worker's), and reply nudges are pushed to it.
|
|
# REQUIRED when the primary runs inside a herdr pane — without it the pane match reads the
|
|
# primary as a worker and refuses spawn/send/stop. Get the id from fleet_whoami; re-pin if
|
|
# the primary moves panes.
|
|
# primary:
|
|
# terminal: term_0123456789abcd
|
|
# pushReminders: 5 # max nudges before giving up (default 5)
|
|
# pushBackoffMs: 15000 # delay between nudges (default 15000)
|
|
|
|
# CB-530: MORE THAN ONE LEAD. `primary:` above is singular by construction — every other pane
|
|
# resolves as a worker — which is right for one lead driving a fleet and wrong the moment two leads
|
|
# (say a Claude lead and an opencode lead) work as peers: the second is silently demoted and refused
|
|
# every orchestration call. List each lead's pane here and all of them resolve as leads.
|
|
#
|
|
# tab → the ONLY field identity depends on (CB-579); the exact label of the tab hosting the lead.
|
|
# Label the tab yourself, or let bridged label one it launches — see `fleet.leaders:` below.
|
|
# kind/model → descriptive; they document what runs in the pane and are echoed by fleet_whoami
|
|
#
|
|
# A lead's tab must already carry its label (or be launched by bridged, which labels it) — there is
|
|
# no terminal id to paste in and nothing to re-pin when the session restarts: the tab survives, so
|
|
# the same label resolves the same lead again on the next scan.
|
|
# `fleet_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
|
|
# IS a primary for authorization, so nothing that keys on the role breaks.
|
|
#
|
|
# KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single
|
|
# destination for its nudges, and is a separate mechanism from lead identity — see `fleet.leaders:`.
|
|
#
|
|
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
|
|
#
|
|
# Two things stop the tab-name convention from becoming a way to claim leadership: the configured
|
|
# member spaces are excluded from the scan, so nothing bridged places can land in a matching tab;
|
|
# and startup REFUSES a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel`
|
|
# override, also matches — so the two namespaces cannot overlap by accident. The label is a NAME,
|
|
# never a capability: what a pane may do is decided by the role the daemon resolves for it.
|
|
|
|
# CB-551: IDLE-LEAD HEARTBEAT — nudge the single lead back to work when it has been continuously
|
|
# idle (no open fleet_send driving it) past the quiet period. The fleet is one lead + architects +
|
|
# workers, so a lead that stalls is a single point of failure; the ReplyPushLoop only nudges when a
|
|
# reply lands, and this timer catches the gap where nothing lands and the lead just sits idle.
|
|
#
|
|
# Opt-in on purpose — it SPENDS the operator's subscription on its own initiative (each nudge starts
|
|
# a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent
|
|
# block = feature off, exactly as before.
|
|
#
|
|
# Three knobs, each with a default that errs on the side of not burning context:
|
|
# idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 —
|
|
# # absorbs normal post-turn pauses; re-prompting every pause burns context)
|
|
# backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000)
|
|
# quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops
|
|
# # until real state appears (default 3 — never nag an empty fleet forever)
|
|
# leadHeartbeat:
|
|
# idleAfterSeconds: 300
|
|
# backoffMs: 60000
|
|
# quietNudgeCap: 3
|
|
|
|
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
|
|
# per tick.
|
|
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
|
|
# Math.max(15, intervalSeconds), so a lower value is silently raised, not
|
|
# rejected.
|
|
# workingSuspectAfterSeconds, paneProbeIntervalSeconds → accepted and parsed, but NOT YET READ by
|
|
# anything — the dormant monitor only consumes intervalSeconds today (CB-573
|
|
# shipped ahead of the evidence publishers these two knobs are for). Setting
|
|
# them changes nothing right now, and no minimum is enforced on either, because
|
|
# nothing reads them to enforce one. They exist so a later build can start
|
|
# honouring them without another config-shape change.
|
|
# notifications.mode → "webhook" flips what fleet_list REPORTS (healthCoverage: "full" instead
|
|
# of "detection-only") — it does NOT make bridged send any webhook call; no
|
|
# delivery mechanism is implemented yet. Any other value, or omitting the
|
|
# block, reports "detection-only".
|
|
# health:
|
|
# enabled: true
|
|
# intervalSeconds: 30
|
|
# workingSuspectAfterSeconds: 600
|
|
# paneProbeIntervalSeconds: 60
|
|
# notifications:
|
|
# mode: disabled
|
|
|
|
# herdr Unix socket. Omit to use the client default
|
|
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
|
|
herdrSocket: ~/.config/herdr/herdr.sock
|
|
|
|
# How member sessions are spawned. Define one or more named profiles (backends) under
|
|
# `profiles`; each key is the profile name (also the ccs profile). A profile says only WHICH
|
|
# BACKEND — model, CLI adapter, credentials, cost. It says nothing about what a member spawned on
|
|
# it is for; that is the member's role, and roles live under `fleet:` below. Which profile an
|
|
# unqualified spawn lands on comes from that role's pool, not from a global default.
|
|
#
|
|
# Shared knobs (placement/workspace) can be repeated per profile; they usually match.
|
|
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
|
|
# Use `pane` for the legacy behaviour (split the focused tab).
|
|
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
|
|
# (--append-system-prompt) as launch flags; nothing is written to the profile.
|
|
# ideMcpUrl → opt-in (CB-634), default off. When set, bridged mounts the IDE Index MCP as a
|
|
# second inline server named `intellij`, and adds an IDE charter that pins every
|
|
# ide_* call to the member's own worktree. A URL, not a boolean — host and port
|
|
# are host-specific. Set it only on a host where the IDE actually runs.
|
|
# tokenEnv → host env var holding the worker's auth token (value never stored in config);
|
|
# omit for a backend that needs no token (e.g. a local ollama).
|
|
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
|
|
# cwd on an MCP spawn, else the daemon's cwd — never $HOME. See
|
|
# docs/Worker-Startup-and-Trust.md.
|
|
# configDir → CLAUDE_CONFIG_DIR for the worker, so it inherits that profile's
|
|
# skills/MCP/hooks. Omit to leave the worker on the host default.
|
|
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
|
|
# worktree sees the same local config (CB-301-ext). Omit for the default set:
|
|
# [.env, .envrc]. (.claude/settings.local.json is NOT in the default — it
|
|
# pre-approves IDE/tool grants a member must not hold ambiently; CB-525/CB-634.)
|
|
#
|
|
# Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher
|
|
# mounts — the bridge, and nothing else. Replicating the primary's MCP config
|
|
# handed a worker the primary's IDE servers, which are bound to the primary's
|
|
# checkout, so its navigation returned paths OUTSIDE its own worktree: one
|
|
# worker made all 59 of its edits in the primary tree while compiling its
|
|
# worktree, and every build it ran was of code that did not contain them.
|
|
# bridged neutralizes a provisioned worktree's .mcp.json for this reason;
|
|
# listing it here would copy the primary's back over that.
|
|
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
|
|
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
|
|
# Opt-in by design — omit and the worker gets no PR-create grant (push over
|
|
# SSH is unaffected). The token value itself is never stored in this file.
|
|
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
|
|
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
|
|
# exhaustedPattern → regex matched against a completion-fallback scrape (CB-578 stage A) to
|
|
# classify a turn that ended with no fleet_reply as the backend having
|
|
# refused on a subscription usage limit, rather than a real answer. Opt-in —
|
|
# omit and this profile's completion fallback behaves exactly as before.
|
|
# Every backend words its refusal differently, so this is config, never a
|
|
# vendor string baked into bridged itself.
|
|
# DEFERRED: compiled once into a startup pattern map — editing it needs a
|
|
# daemon restart, same as this profile's model/baseUrl/argv.
|
|
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
|
|
# BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME
|
|
# credentialId share one quarantine — the case this exists for is two models
|
|
# on one account (e.g. sol and terra both billing one OpenAI credential): an
|
|
# exhaustion on either one must lock out both, or the fleet just walks onto
|
|
# the same dead account under the sibling's name. Opt-in — omit and this
|
|
# profile quarantines alone, under its own name, exactly as if the field did
|
|
# not exist. Cooldown length is the top-level quarantineCooldownSeconds below.
|
|
# HOT: read live at every spawn/exhaustion check — no restart needed.
|
|
# env → extra environment for this profile's workers, as a literal key/value map
|
|
# (CB-511). Use it to give workers a toolchain.
|
|
#
|
|
# A worker's environment does NOT come from your shell. bridged hands herdr an
|
|
# explicit env map and herdr merges it into ITS OWN process env — so before
|
|
# CB-511 a worker inherited whatever PATH the herdr server happened to be
|
|
# started with, which on a long-lived herdr can predate your toolchain entirely
|
|
# and leave workers unable to run `mvn` or `java` at all.
|
|
# bridged now propagates ITS OWN PATH to every worker by default; set `env:`
|
|
# only to override that or add more (JAVA_HOME, …). Since the default is the
|
|
# daemon's PATH, make sure the daemon is started with a good one — see the PATH
|
|
# lines in deploy/dev.ltms.fleet.plist and deploy/bridged.service.
|
|
#
|
|
# Adapter-owned variables always win over `env:`: ANTHROPIC_BASE_URL and the
|
|
# rest of the ANTHROPIC_*/CLAUDE_* wiring are applied after it, so an `env:`
|
|
# entry cannot repoint a worker past the SubscriptionGuard — which is checked
|
|
# against `baseUrl` alone.
|
|
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
|
|
profiles:
|
|
gx10: # ccs profile name (NOT a hostname)
|
|
kind: claude-code # which adapter spawns this profile (default; may omit)
|
|
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
|
|
model: coder
|
|
placement: tab
|
|
workspace: bridged-workers
|
|
# tabLabel: an optional per-profile override; the fleet template usually covers it
|
|
mcpUrl: http://127.0.0.1:8765/mcp
|
|
tokenEnv: BRIDGED_WORKER_TOKEN
|
|
argv: ["ccs", "gx10"]
|
|
# weight: relative selection weight for automatic placement (weighted, round-robin, and
|
|
# fixed's fallback walk). Absent defaults to 1.0. An explicit 0 or negative value means
|
|
# "never auto-select this profile" (CB-554) — it stays reachable via an explicit
|
|
# `fleet_spawn{profile:"gx10"}`, which bypasses placement entirely; only automatic
|
|
# selection skips it.
|
|
weight: 0.5
|
|
# maxLoad: max live workers on this profile. Omit for unlimited. An explicit 0 (CB-585) caps
|
|
# the profile at zero live members — it is excluded from automatic placement and an explicit
|
|
# `fleet_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
|
|
# is named directly. Negative is refused at config load — there is no sane meaning for it.
|
|
maxLoad: 2
|
|
# subscription: true
|
|
# THE KNOB THAT DECIDES WHO PAYS (CB-539). Default false. When true, this profile's members
|
|
# run on the OPERATOR'S OWN Claude subscription instead of a metered endpoint — every spawn
|
|
# bills your plan and eats your usage limit. Off-subscription is the whole point of this
|
|
# daemon, so treat `true` as a deliberate exception, not a convenience.
|
|
#
|
|
# What changes when it is set (ClaudeCodeLauncher):
|
|
# - no ANTHROPIC_BASE_URL and no ANTHROPIC_AUTH_TOKEN are injected — the member inherits
|
|
# the operator's own Claude Code auth, which is exactly why it bills the plan;
|
|
# - SubscriptionGuard never vets it, because there is no baseUrl to vet;
|
|
# - no token is required, so `tokenEnv` is irrelevant here.
|
|
#
|
|
# MUTUALLY EXCLUSIVE with `baseUrl` — setting both is refused at config load (CB-542). On the
|
|
# subscription path no guard would vet the URL, so allowing both would be a way around the
|
|
# guard rather than a configuration.
|
|
#
|
|
# GOTCHA 1 — it is invisible to the startup secret check. `Fleetd.reportRequiredSecrets`
|
|
# skips subscription profiles on purpose (they need no token), so a boot log that reports
|
|
# every secret as fine says nothing about these profiles.
|
|
#
|
|
# GOTCHA 2 — `maxLoad` is the ONLY throttle you have here. There is no metering, no budget
|
|
# and no refusal on cost; the cap on live members is the single thing standing between a
|
|
# fan-out and your monthly limit. Set it deliberately and keep it small.
|
|
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
|
|
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
|
|
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
|
|
# credentialId: shared-openai # opt-in: quarantine together with every other profile sharing this id (CB-578)
|
|
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
|
|
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
|
|
# parityOverlay: [".env", ".envrc"] # the default; never add .mcp.json or .claude/settings.local.json — see above
|
|
# ideMcpUrl: http://127.0.0.1:29170/index-mcp/streamable-http # opt-in (CB-634): IDE code intelligence, pinned to the worktree
|
|
gx11: # a second backend, so `placement: weighted` has a choice
|
|
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
|
|
placement: tab
|
|
workspace: bridged-workers
|
|
# tabLabel: an optional per-profile override; the fleet template usually covers it
|
|
mcpUrl: http://127.0.0.1:8765/mcp
|
|
argv: ["ccs", "gx11"]
|
|
weight: 0.5
|
|
maxLoad: 2
|
|
# Pin an auto-compact window BELOW the served model's context ceiling. The global
|
|
# ~/.claude/settings.json value is shared by every ccs instance and the primary, so the
|
|
# per-profile override belongs here. Equal to the ceiling means auto-compact never fires
|
|
# before the server rejects the prompt, which kills a worker mid-turn (CB-523).
|
|
env:
|
|
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "280000"
|
|
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
|
|
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
|
|
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
|
|
# The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a
|
|
# `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude.
|
|
#
|
|
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → fleet_send →
|
|
# structured fleet_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
|
|
# and need NO credentials — check `opencode models` for the current free list, since the names
|
|
# change. That also makes the worker off-subscription by construction.
|
|
# opencode-free:
|
|
# kind: opencode
|
|
# model: opencode/north-mini-code-free # `provider/model` selector, injected as `-m`
|
|
# placement: tab
|
|
# workspace: bridged-workers
|
|
# tabLabel: "opencode: {profile} #{n}"
|
|
# mcpUrl: http://127.0.0.1:8765/mcp
|
|
# argv: ["opencode"]
|
|
#
|
|
# CB-508: point an opencode profile at your OWN OpenAI-compatible endpoint (local vLLM, llama.cpp,
|
|
# LM Studio, TGI…) instead of opencode's gateway. Setting `baseUrl` on a `kind: opencode` profile
|
|
# makes the bridge emit a custom `provider` block into the generated opencode.json — opencode has
|
|
# no ANTHROPIC_BASE_URL seam, so this is how the endpoint is pinned.
|
|
# baseUrl → a bare host:port gets `/v1` appended (where these servers mount the API); a URL that
|
|
# already has a path is used verbatim, so a custom mount point still works.
|
|
# model → MUST be "<provider>/<model>". The provider half names the generated block; the model
|
|
# half must match an id the server reports at /v1/models. One field drives both the
|
|
# declaration and the `-m` flag, so they cannot drift apart. A bare model name with a
|
|
# baseUrl set is rejected at spawn rather than silently using the default gateway.
|
|
# tokenEnv → optional; its value becomes the provider apiKey. Most local servers ignore the key,
|
|
# so a placeholder is used when unset (the AI SDK still requires a non-empty one).
|
|
# NOTE: no `guard` entry is needed even with a baseUrl set. The SubscriptionGuard exists to stop a
|
|
# worker borrowing the primary's Anthropic subscription, and an opencode process has no Anthropic
|
|
# credential path at all.
|
|
# opencode-local:
|
|
# kind: opencode
|
|
# baseUrl: http://127.0.0.1:8000
|
|
# model: local-vllm/deepseek-v4-flash
|
|
# placement: tab
|
|
# workspace: bridged-workers
|
|
# tabLabel: "opencode: {profile} #{n}"
|
|
# mcpUrl: http://127.0.0.1:8765/mcp
|
|
# argv: ["opencode"]
|
|
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
|
|
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
|
|
#
|
|
# `weighted` IS NOT "cheapest first" — read this before you set weights (CB-589).
|
|
# It is smooth weighted round-robin: it spreads spawns across EVERY profile that has a free slot,
|
|
# in weight ratio. It has no idea which profile costs money. So with local:10 / paid:2 you do not
|
|
# get "use local, overflow to paid" — you get roughly one spawn in six going to the paid profile
|
|
# while the local box still has a free slot.
|
|
#
|
|
# There is a sharper second effect. The policy's running score map lives for the daemon's whole
|
|
# life. While a profile is at maxLoad it is filtered out and its score FREEZES, so the paid
|
|
# profiles keep accumulating against it. When the local slot frees up it returns with a stale
|
|
# score and can LOSE the next pick — a paid spawn while the free box sits idle.
|
|
#
|
|
# Until a real cost-first policy exists, the workaround is to make the ratio decisive rather than
|
|
# proportional: give the free profile a weight so large that it wins every pick it is eligible
|
|
# for, and paid profiles only ever take genuine overflow. On this host that is local weight 100
|
|
# against paid weights of ~1.
|
|
#
|
|
# The gotcha with that workaround: it expresses a PREFERENCE ORDER through a RATIO knob. Add a
|
|
# future profile at weight 150 and it silently outranks the free box, with nothing to warn you.
|
|
# Re-check the weights whenever you add a profile.
|
|
placement: weighted
|
|
|
|
# How long a credential sits out after a BACKEND_EXHAUSTED classification (CB-578 stage B), in
|
|
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
|
|
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
|
|
# 1800 (30 minutes) when omitted or non-positive.
|
|
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
|
|
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
|
|
# restart. Editing this needs a daemon restart to take effect.
|
|
# quarantineCooldownSeconds: 1800
|
|
|
|
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
|
|
# upgraded bridged keeps the old behaviour: the file is read once at boot and never again.
|
|
# enabled → turn the watch on. bridged checks the file's modified time on a timer and
|
|
# reloads when it moves.
|
|
# intervalSeconds → how often to check (default 10). One `stat` per tick, so this is cheap.
|
|
#
|
|
# Not every key can move under a running daemon, and the difference is about what already exists
|
|
# when the reload happens — not about how important the key is:
|
|
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
|
|
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad
|
|
# / credentialId. Those are hot because the placement policy (and, for credentialId,
|
|
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
|
|
# not by itself enough to make a key hot.
|
|
# EXCEPT `fleet.leaders`: Fleetd.main reads it once at startup to build the lead tab
|
|
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
|
|
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
|
|
# with nothing in the deferred list — but has NO effect until you restart. Treat it
|
|
# as deferred in practice, even though today's reload output does not say so.
|
|
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
|
|
# until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`,
|
|
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, `quarantineCooldownSeconds` (CB-578
|
|
# stage B — baked once into the quarantine tracker built at startup), ADDING or
|
|
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
|
|
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
|
|
# configDir, mcpUrl, tabLabel, exhaustedPattern. The launcher takes a copy of
|
|
# `profiles:` at startup and resolves every spawn out of that copy, so those never
|
|
# reach a launch until you restart. The reload logs them by name rather than
|
|
# pretending they applied.
|
|
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
|
|
# bound, the broker connection is open, and the auth mode decides who may reach the
|
|
# port that is already listening.
|
|
#
|
|
# A changed COLD key refuses the WHOLE reload — not the hot half applied and the cold half warned
|
|
# about. A half-applied reload would leave the daemon matching no file on disk, which is the worst
|
|
# thing a reload can do to an operator debugging one. A file that fails to parse or fails a startup
|
|
# validator is refused the same way, and the running config stays live.
|
|
# configReload:
|
|
# enabled: true
|
|
# intervalSeconds: 10
|
|
|
|
# THE FLEET (CB-557) — who the daemon may run, and under which role. This one block replaced four
|
|
# older keys: `leaders:`, `members:`, `leadScan:` and `defaultProfile:`.
|
|
#
|
|
# A member is anything a lead spawns, and every member has two INDEPENDENT attributes:
|
|
# role — which contract: architect, dev or reviewer. It picks the launch charter, the role
|
|
# file, the playbook skill and the authz row.
|
|
# profile — which backend: one of the `profiles:` keys above (model, CLI adapter, cost).
|
|
# They vary on their own. A reviewer may run on the same profile as the dev whose diff it reads,
|
|
# which is why the two cannot be one field.
|
|
#
|
|
# The ROLE IS THE CONTAINING KEY, not a `role:` field. That is not only tidier: a misspelled role
|
|
# used to parse into a member with no contract at all, while a misspelled pool name here simply
|
|
# declares nothing.
|
|
#
|
|
# Each pool lists the profiles that role MAY run on — these are pools, not identities. That is also
|
|
# what replaced `defaultProfile:`: an unqualified spawn names a role, and that role's pool supplies
|
|
# the candidates, in definition order. A dev and a reviewer staying anonymous is exactly compatible
|
|
# with being listed here; the entry key just names the entry.
|
|
fleet:
|
|
# Optional launch-charter text, keyed only by the singular role wire names: architect, dev,
|
|
# reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put secrets
|
|
# here: a later launch step writes this text to a world-readable temp file, and ${ENV} interpolation
|
|
# is deliberately not supported.
|
|
charters:
|
|
architect: |-
|
|
You are an architect in this fleet. You refine work before anyone builds it:
|
|
scope, acceptance criteria, risks, and a unit split. You read the repo and
|
|
write analysis. You never commit production code and never open a PR.
|
|
A design task is worked by two architects. Design alone first, then exchange
|
|
and say plainly where you disagree. Do not concede just to agree.
|
|
dev: |-
|
|
You implement the one unit you were given, and nothing else. You test it,
|
|
commit it, and open your own pull request. You never merge.
|
|
reviewer: |-
|
|
You review the diff you were given. You report bugs, risks and missing tests.
|
|
You do not change code.
|
|
|
|
# Optional. Template for a member tab's label; {role}, {profile}, {model} and {n} are substituted.
|
|
# {n} counts per role+profile, so `dev: sonnet #2` really is the second sonnet dev. Because {role}
|
|
# comes from a closed enum, a generated label can never begin with a lead's tabPrefix.
|
|
# tabLabel: "{role}: {profile} #{n}"
|
|
|
|
# Panes that orchestrate rather than are orchestrated. A lead may now be CREATED as well as
|
|
# recognised: give it a `profile:` and the daemon launches the shortfall when fewer than
|
|
# `instances` are live. Omit `profile:` and it is recognise-only, as before.
|
|
#
|
|
# `tab:` (CB-579) is REQUIRED and is the only field identity depends on — the exact label of the
|
|
# tab hosting the lead, matched case-insensitively. Label the tab yourself and put that same
|
|
# string here, and the pane is recognised on the next rescan. Reopen the tab later, or the session
|
|
# inside it restarts — the terminal id changes; the tab, and its label, do not, so no config edit
|
|
# follows a restart.
|
|
#
|
|
# A lead the daemon launches is labelled BY the daemon with this same `tab:` value, so it is found
|
|
# by the same scan. A lead counts as live only when herdr also reports a running agent in that
|
|
# tab — a label left behind by a session that died does not block the relaunch, and a tab that is
|
|
# gone entirely drops out of the next scan rather than being remembered forever.
|
|
#
|
|
# An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with
|
|
# the session lifecycle (the idle reaper would kill your orchestrator), and stays on the
|
|
# subscription — ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.
|
|
#
|
|
# GET THE `tab:` VALUE RIGHT. A pane that does not match any configured `tab:` (a typo, a renamed
|
|
# tab, a pane no entry names at all) is not recognised as a lead — it resolves as an ordinary
|
|
# WORKER instead, silently, and every orchestration call it makes (spawn/stop/send/drain) is
|
|
# refused. There is no error at startup for this: an unmatched pane is simply not a lead. If your
|
|
# primary suddenly can't spawn or send, check this section first.
|
|
# leaders:
|
|
# opus-5.0:
|
|
# profile: opus # omit to never create this lead, only recognise it
|
|
# instances: 1 # desired live count; only the shortfall is launched. 0 = off
|
|
# tab: "lead: opus-5.0" # REQUIRED — the exact tab label this lead lives in
|
|
# tabPrefix: "lead:" # only used to guard against a worker tabLabel colliding with
|
|
# # this convention at startup; plays no part in matching a lead
|
|
# scanIntervalSeconds: 10 # rescan cadence, and the worst case before a new tab is seen
|
|
# workspace: leads # where a launched lead's tab is created (default "leads").
|
|
# # MUST NOT be a member workspace — those are excluded from the
|
|
# # scan, so a lead placed in one is never found again.
|
|
# cwd: /path/to/repo # the launched lead's working directory (default: bridged's own)
|
|
# kind: claude # descriptive; reported by fleet_whoami
|
|
# gpt-sol-5.6:
|
|
# tab: "lead: gpt-sol-5.6"
|
|
# kind: opencode
|
|
# model: openai/gpt-5.6-terra
|
|
|
|
# architects:
|
|
# architect-1:
|
|
# profile: opus # a strong model, on the operator's subscription
|
|
# architect-2:
|
|
# profile: sol # a different vendor on purpose — two architects that share a
|
|
# # model share its blind spots
|
|
developers:
|
|
gx10:
|
|
profile: gx10
|
|
# reviewers:
|
|
# gx10:
|
|
# profile: gx10 # the same backend may serve two roles; that is the point
|
|
|
|
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
|
|
# must carry none. Every profile above must have its host listed here.
|
|
guard:
|
|
offSubscriptionHosts:
|
|
- gx00.gw
|
|
- gx01.gw
|
|
|
|
# Member credential policy (CB-596, gitea issue #82). A herdr pane runs a LOGIN shell, and that
|
|
# shell re-sources the operator's own secret store — so a spawned member inherits every credential
|
|
# the operator's shell holds, not just the ones bridged means to give it. Measured on this host:
|
|
# 31 credential names, all set, with only ONE (GITEA_ACCESS_TOKEN) blocked before this — and that
|
|
# block was a single name hardcoded in HerdrPeerLauncher.java, not driven by this file. This block
|
|
# replaces that hardcoded shadow with a config-driven list of names.
|
|
#
|
|
# ROUND-2 CORRECTION, measured live: the pane-creation env overlay below (applied at tab.create /
|
|
# pane.split, BEFORE the pane's login shell runs) does NOT survive that login shell for any name
|
|
# secrets.sh actually exports — the shell re-exports it afterwards and overwrites the sentinel.
|
|
# Proof: GITEA_ACCESS_TOKEN comes back blocked only because secrets.sh itself carries a guarded
|
|
# export (`[ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...`) — that guard, not this
|
|
# file, is what wins. No other name in `known` below has a matching guard in secrets.sh yet (1
|
|
# guard measured against 33 export lines there). So today this block's overlay is REAL protection
|
|
# only for a name secrets.sh does not export, or a peer kind whose pane never runs a login shell —
|
|
# for everything secrets.sh exports and guards, the guard in secrets.sh (out of scope for this
|
|
# ticket) is what actually blocks it, not this list. An exec-time fix (winning after the login
|
|
# shell finishes, before the agent process starts) was attempted and found to have no seam in the
|
|
# current herdr protocol — AgentControl.start takes a fixed `kind` (herdr resolves the executable)
|
|
# plus trailing CLI args for that binary, not an arbitrary argv or an env map; only tab.create /
|
|
# pane.split accept `env`, and that is this same pane-creation overlay. See gitea #82 for the open
|
|
# design question this leaves.
|
|
#
|
|
# DENY-BY-DEFAULT, NOT A DENY-LIST. A deny-list (name the bad ones, let everything else through) is
|
|
# silently wrong the moment the operator's store gains a new secret — nothing would ever report it.
|
|
# Deny-by-default inverts that: `known` bounds the blast radius to names actually enumerated below,
|
|
# and EVERY one of them is blocked UNLESS it is also in `allow`. Omitting this block entirely (the
|
|
# shipped default) blocks NOTHING — unlike most optional blocks in this file, absence here is a real
|
|
# gap, not a safe "feature off". A name that is neither `known` nor `allow`-ed is not silently let
|
|
# through either: the daemon logs a WARN naming any credential-shaped env var it finds on neither
|
|
# list (never its value), so a secret added to the store later does not go unnoticed forever.
|
|
#
|
|
# policy → "deny-by-default" (the default; also accepted spelled "deny-list") overlays each
|
|
# known-but-not-allowed name BEFORE the pane's login shell runs — real protection only
|
|
# where that shell does not re-export the name (see ROUND-2 CORRECTION above). An
|
|
# unrecognized value refuses to start, naming it.
|
|
# policy → "allow-list" (CB-633) moves the control to a per-spawn ZDOTDIR directory the daemon
|
|
# generates and passes through tab.create's env map. Each generated startup file sources
|
|
# its ~/ counterpart FIRST and then runs the scrub, so the scrub happens after the
|
|
# operator's whole chain and no sourced file can undo it.
|
|
# The scrub is sourced from BOTH the generated .zshrc and the generated .zlogin, because
|
|
# herdr does not open the same kind of shell everywhere: macOS panes run a LOGIN zsh (so
|
|
# .zlogin runs), Linux panes run a plain interactive zsh (so .zlogin never runs at all).
|
|
# A scrub in .zlogin alone would be a control that silently does nothing on Linux.
|
|
# The allow-list is DERIVED, never typed:
|
|
# every profile's tokenEnv/gitTokenEnv/gitHostEnv values and env-map keys, plus an
|
|
# infrastructure set (PATH HOME SHELL TERM LANG LC_* TMPDIR USER LOGNAME PWD SHLVL EDITOR
|
|
# PAGER JAVA_HOME XDG_* ZDOTDIR), plus whatever keys this spawn's own env overlay carries.
|
|
# Adding a profile can therefore only widen the list, never break another spawn's scrub.
|
|
# Under this policy `known`/`allow` below become REPORTING ONLY — they feed the gap WARN,
|
|
# they are no longer a control. If the member's login shell is NOT zsh, the daemon logs a
|
|
# loud WARN saying protection is off and falls back to deny-by-default's overlay.
|
|
# Each pane writes a scrub-report.txt naming how many variables it kept of how many it
|
|
# saw; the daemon logs that "allowed N of M" line when the pane stops. If the report is
|
|
# MISSING the daemon logs a WARN instead — the scrub cannot then be confirmed to have
|
|
# run, and a silently-dead control is exactly what this policy exists to prevent.
|
|
# allow → credential names a member legitimately needs. Under deny-by-default, left OUT of the
|
|
# pane's env overlay entirely, so the value the pane's own (login) shell exports passes
|
|
# through untouched. Under allow-list: reporting only.
|
|
# known → every credential name the operator's store is known to export. Under deny-by-default,
|
|
# every name here NOT also in `allow` is overlaid with a non-secret sentinel value before
|
|
# the pane's login shell runs — real protection only for names that shell does not itself
|
|
# re-export (see the ROUND-2 CORRECTION note above). Under allow-list: reporting only.
|
|
# sshAuthSock → whether SSH_AUTH_SOCK may pass through under allow-list ("allow") or must be
|
|
# blanked like any other non-derived name ("block", the default). This is a decision you
|
|
# have to make explicitly: SSH_AUTH_SOCK is a handle to YOUR ssh-agent, and a member
|
|
# holding it can sign with your keys — it sits in no secret file and looks like no
|
|
# credential, which is why it slipped past three earlier tickets (gitea #110). Blocking
|
|
# it breaks git over SSH inside members (push/fetch authenticate as you); use HTTPS
|
|
# remotes or scoped deploy keys instead of allowing it lightly.
|
|
#
|
|
# HOT-RELOADABLE the same way `fleet:` is (CB-559): read fresh on every spawn, so editing this list
|
|
# and reloading config (or restarting) changes what the NEXT spawn inherits; already-running members
|
|
# are unaffected either way.
|
|
# memberCredentials:
|
|
# policy: deny-by-default # or "deny-list", or "allow-list" (CB-633) — see above
|
|
# sshAuthSock: block # allow-list only; see the sshAuthSock note above
|
|
# allow:
|
|
# - AI_GATEWAY_TOKEN # named in a profile's tokenEnv (local/gx) — a member reaching the
|
|
# # gateway is by design, not a leak
|
|
# - WORKER_GITEA_TOKEN # the repo-scoped forge token a member needs to open its own PR (CB-302)
|
|
# - CONTEXT7_TOKEN # already decided as allowed by CB-593
|
|
# - GITEA_HOST # not a credential — a hostname, paired with the forge token above
|
|
# known:
|
|
# - AI_GATEWAY_TOKEN
|
|
# - BESZEL_ADMIN_EMAIL
|
|
# - BESZEL_ADMIN_PASSWORD
|
|
# - BESZEL_HUB_URL
|
|
# - BESZEL_KEY
|
|
# - BESZEL_UNIVERSAL_TOKEN
|
|
# - BRAIN_MCP_TOKEN
|
|
# - CF_ACCOUNT_ID
|
|
# - CF_API_TOKEN
|
|
# - CF_USER_TOKEN
|
|
# - CONFLUENCE_API_TOKEN
|
|
# - CONFLUENCE_USERNAME
|
|
# - CONTEXT7_TOKEN
|
|
# - GITEA_ACCESS_TOKEN
|
|
# - GITEA_HOST
|
|
# - GITLAB_OAUTH_CLIENT_SECRET
|
|
# - GITLAB_PERSONAL_ACCESS_TOKEN
|
|
# - GRAFANA_ADMIN_PASSWORD
|
|
# - GRAFANA_ADMIN_USER
|
|
# - HASS_TOKEN
|
|
# - HW_PASSWORD
|
|
# - HW_USER
|
|
# - LTMS_API_KEY
|
|
# - MEMORY_MCP_TOKEN
|
|
# - METRICS_PUSH_TOKEN
|
|
# - OPENCODE_AUTOMODE_MODEL
|
|
# - TELEGRAM_BOT_TOKEN
|
|
# - TELEGRAM_CHAT_ID
|
|
# - TS_API_KEY
|
|
# - TS_AUTHKEY
|
|
# - WORKER_GITEA_TOKEN
|
|
|
|
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
|
|
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
|
|
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
|
|
# unknown keys are ignored, so a snake_case key would be silently dropped (default kept).
|
|
# spawnReadyTimeoutMs: 20000
|
|
# spawnReadyPollMs: 300
|
|
|
|
# Worktree provisioning root (CB-301-ext). Where per-worker git worktrees are checked out so
|
|
# each worker owns an isolated branch instead of sharing the primary's tree. Omit to default
|
|
# to a sibling directory of the repo root.
|
|
# worktreeRoot: /Users/me/src/.bridged-worktrees
|
|
|
|
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
|
|
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
|
|
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
|
|
# contextCap → force-release a session after this many delegated turns
|
|
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
|
|
# clearAfterTurn → whether a reusable worker discards its conversation context after every
|
|
# completed delegated turn (default false). Works for claude-code workers
|
|
# only — any other peer kind (e.g. opencode) logs "context reset is
|
|
# unsupported for peer kind …" once and the reset is a no-op.
|
|
# lifecycle:
|
|
# idleTtlSeconds: 300
|
|
# contextCap: 10
|
|
# drainTimeoutSeconds: 5
|
|
# clearAfterTurn: false
|
|
|
|
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
|
|
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
|
|
# Set a broker uri to swap in the AMQP-backed inbox: worker replies with no open send are held
|
|
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
|
|
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
|
|
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
|
|
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
|
|
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
|
|
# uriEnv → CB-151: name of a host env var holding the AMQP URI, preferred over `uri` (wins
|
|
# whenever set). The URI carries `user:pass@` inline, so naming a variable keeps the
|
|
# password out of fleetd.yaml — same pattern as auth.tokenEnv/Profile.tokenEnv. A
|
|
# uriEnv that resolves to an unset or blank variable is treated as NOT configured and
|
|
# the daemon falls back to the in-memory inbox, warning loudly.
|
|
# prefetch → CB-527: consumer basicQos, capping how many unacked messages the inbox holds
|
|
# in-heap per owned target (the rest sits on the broker's durable queue instead of
|
|
# growing the JVM heap). Default 32 when omitted.
|
|
# broker:
|
|
# uriEnv: LAVINMQ_URI
|
|
# prefetch: 32
|
|
|
|
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open fleet_send,
|
|
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
|
|
# pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop
|
|
# stops as soon as the primary's inbox is empty.
|
|
# terminal → pin the primary's herdr terminal id. Omit to learn it from the connection on
|
|
# the first orchestration-side MCP call (the normal case). An off-host or
|
|
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
|
|
# degrades to pull; the reply is still never lost.
|
|
#
|
|
# REQUIRED (CB-522) if the primary itself runs inside a herdr pane. Caller
|
|
# identity resolves a loopback PID to its herdr pane, and PaneLocator scans
|
|
# EVERY pane — not just bridged-spawned ones — so such a primary is otherwise
|
|
# classified as a WORKER and refused SPAWN/SEND/STOP. That failure is
|
|
# self-locking: the learned terminal is populated by the very orchestration
|
|
# calls being refused, so only this pinned value can break the cycle. Read the
|
|
# id off fleet_whoami (it reports the current terminal even while
|
|
# misclassified) and re-pin whenever the primary moves panes.
|
|
# pushReminders → max nudges before giving up (default 5)
|
|
# pushBackoffMs → delay between nudges in ms (default 15000)
|
|
# primary:
|
|
# terminal: term_65619bd6174568
|
|
# pushReminders: 5
|
|
# pushBackoffMs: 15000
|