# bridged configuration (example). Copy to bridged.yaml and adjust. # # bridged is the sole gateway between primary/worker Claude sessions and herdr. # It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL. # REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token` # below — bridged REFUSES TO START on a non-loopback bind under loopback-trust (see auth). bind: host: 127.0.0.1 port: 8765 # API authentication (CB-501). Governs how a caller that is NOT an on-host worker pane proves it # is the primary. Worker identity never depends on this: a loopback peer PID that maps to a herdr # pane is unforgeable and is always honoured, so turning auth on cannot lock the fleet out. # # mode: loopback-trust → DEFAULT, and the historical behaviour: any loopback caller that is not # a worker is the primary, no credential needed. Sound ONLY because the # OS refuses remote connections to a loopback socket. # mode: token → such a caller must send `Authorization: Bearer `; without it it # is anonymous and authorized for nothing. REQUIRED for a non-loopback # bind — the daemon fails fast otherwise, because "unauthenticated ⇒ # primary" on a reachable port would hand spawn/stop/send to anyone. # tokenEnv → host env var holding the token (never the literal value). Default # BRIDGED_API_TOKEN. Read only in token mode; empty ⇒ startup fails. # # TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and # let it own certificate lifecycle, e.g. # location / { proxy_pass http://127.0.0.1:8765; proxy_set_header Authorization $http_authorization; } # The broker link gets TLS from its own URI (amqps://…) — see `broker` below. # auth: # mode: token # tokenEnv: BRIDGED_API_TOKEN # Optional pinned primary terminal (CB-307). Names the herdr pane the PRIMARY itself runs in: # a caller whose connection maps to this pane resolves as the primary (no credential needed — # the pane mapping is as unforgeable as a worker's), and reply nudges are pushed to it. # REQUIRED when the primary runs inside a herdr pane — without it the pane match reads the # primary as a worker and refuses spawn/send/stop. Get the id from fleet_whoami; re-pin if # the primary moves panes. # primary: # terminal: term_0123456789abcd # pushReminders: 5 # max nudges before giving up (default 5) # pushBackoffMs: 15000 # delay between nudges (default 15000) # CB-530: MORE THAN ONE LEAD. `primary:` above is singular by construction — every other pane # resolves as a worker — which is right for one lead driving a fleet and wrong the moment two leads # (say a Claude lead and an opencode lead) work as peers: the second is silently demoted and refused # every orchestration call. List each lead's pane here and all of them resolve as leads. # # tab → the ONLY field identity depends on (CB-579); the exact label of the tab hosting the lead. # Label the tab yourself, or let bridged label one it launches — see `fleet.leaders:` below. # kind/model → descriptive; they document what runs in the pane and are echoed by fleet_whoami # # A lead's tab must already carry its label (or be launched by bridged, which labels it) — there is # no terminal id to paste in and nothing to re-pin when the session restarts: the tab survives, so # the same label resolves the same lead again on the next scan. # `fleet_whoami` reports `{"role":"primary","leader":""}`; role stays "primary" because a lead # IS a primary for authorization, so nothing that keys on the role breaks. # # KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single # destination for its nudges, and is a separate mechanism from lead identity — see `fleet.leaders:`. # # Leads are configured under `fleet.leaders:` — see THE FLEET further down. # # Two things stop the tab-name convention from becoming a way to claim leadership: the configured # member spaces are excluded from the scan, so nothing bridged places can land in a matching tab; # and startup REFUSES a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel` # override, also matches — so the two namespaces cannot overlap by accident. The label is a NAME, # never a capability: what a pane may do is decided by the role the daemon resolves for it. # CB-551: IDLE-LEAD HEARTBEAT — nudge the single lead back to work when it has been continuously # idle (no open fleet_send driving it) past the quiet period. The fleet is one lead + architects + # workers, so a lead that stalls is a single point of failure; the ReplyPushLoop only nudges when a # reply lands, and this timer catches the gap where nothing lands and the lead just sits idle. # # Opt-in on purpose — it SPENDS the operator's subscription on its own initiative (each nudge starts # a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent # block = feature off, exactly as before. # # Three knobs, each with a default that errs on the side of not burning context: # idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 — # # absorbs normal post-turn pauses; re-prompting every pause burns context) # backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000) # quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops # # until real state appears (default 3 — never nag an empty fleet forever) # leadHeartbeat: # idleAfterSeconds: 300 # backoffMs: 60000 # quietNudgeCap: 3 # Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list # per tick. # intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes # Math.max(15, intervalSeconds), so a lower value is silently raised, not # rejected. # workingSuspectAfterSeconds, paneProbeIntervalSeconds → accepted and parsed, but NOT YET READ by # anything — the dormant monitor only consumes intervalSeconds today (CB-573 # shipped ahead of the evidence publishers these two knobs are for). Setting # them changes nothing right now, and no minimum is enforced on either, because # nothing reads them to enforce one. They exist so a later build can start # honouring them without another config-shape change. # notifications.mode → "webhook" flips what fleet_list REPORTS (healthCoverage: "full" instead # of "detection-only") — it does NOT make bridged send any webhook call; no # delivery mechanism is implemented yet. Any other value, or omitting the # block, reports "detection-only". # health: # enabled: true # intervalSeconds: 30 # workingSuspectAfterSeconds: 600 # paneProbeIntervalSeconds: 60 # notifications: # mode: disabled # herdr Unix socket. Omit to use the client default # (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}). herdrSocket: ~/.config/herdr/herdr.sock # How member sessions are spawned. Define one or more named profiles (backends) under # `profiles`; each key is the profile name (also the ccs profile). A profile says only WHICH # BACKEND — model, CLI adapter, credentials, cost. It says nothing about what a member spawned on # it is for; that is the member's role, and roles live under `fleet:` below. Which profile an # unqualified spawn lands on comes from that role's pool, not from a global default. # # Shared knobs (placement/workspace) can be repeated per profile; they usually match. # placement: tab → each worker lands in its OWN tab in a dedicated worker space (default). # Use `pane` for the legacy behaviour (split the focused tab). # mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter # (--append-system-prompt) as launch flags; nothing is written to the profile. # ideMcpUrl → opt-in (CB-634), default off. When set, bridged mounts the IDE Index MCP as a # second inline server named `intellij`, and adds an IDE charter that pins every # ide_* call to the member's own worktree. A URL, not a boolean — host and port # are host-specific. Set it only on a host where the IDE actually runs. # ideProjectDir → repo-relative module dir the IDE opens and the overlay pins (CB-634). Only read # when ideMcpUrl is set. This repo's Maven pom lives in `bridged/`, not at the # worktree root, so opening the root imports no module and ide_* resolves nothing; # set this to `bridged`. Omit for a repo whose project is the worktree root. # ideOpenCommand → host command that opens ideProjectDir in the IDE at spawn (CB-634 auto-open). # Only read when ideMcpUrl is set. `{dir}` is replaced with the absolute module # dir and the command runs through `/bin/sh -c`, so set env inline if needed — # e.g. `env DISPLAY=:10.0 idea {dir}`. Best-effort: a failure is logged, never # fails the spawn. Omit to open the member's module by hand. There is no close # half yet — an opened module stays open until the operator closes it. # autoCompactWindow → opt-in, default off. A bounded token window that forces a spawned member to # compact its context instead of running on the backend's own default and dying # mid-turn (losing its fleet_reply — the whole point of the turn — with it). # Validated at config load to [100000, 1000000] — the band Claude Code's own # --autocompact flag accepts. # CROSS-BACKEND SEMANTICS DIFFER: on claude-code this is a launch-time # `--autocompact ` flag — the member compacts AT this window. opencode # has no equivalent flag (it only forces `compaction.auto: true`, unconditionally, # already), so this is instead applied as the model's `limit.context` in the # generated opencode.json — the member compacts WITHIN this window, not exactly # at it — and only when this profile's `model:` is in `provider/model` form; if it # isn't, bridged logs a WARN naming the profile rather than silently doing nothing. # tokenEnv → host env var holding the worker's auth token (value never stored in config); # omit for a backend that needs no token (e.g. a local ollama). # cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's # cwd on an MCP spawn, else the daemon's cwd — never $HOME. See # docs/Worker-Startup-and-Trust.md. # configDir → CLAUDE_CONFIG_DIR for the worker, so it inherits that profile's # skills/MCP/hooks. Omit to leave the worker on the host default. # parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned # worktree sees the same local config (CB-301-ext). Omit for the default set: # [.env, .envrc]. (.claude/settings.local.json is NOT in the default — it # pre-approves IDE/tool grants a member must not hold ambiently; CB-525/CB-634.) # # Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher # mounts — the bridge, and nothing else. Replicating the primary's MCP config # handed a worker the primary's IDE servers, which are bound to the primary's # checkout, so its navigation returned paths OUTSIDE its own worktree: one # worker made all 59 of its edits in the primary tree while compiling its # worktree, and every build it ran was of code that did not contain them. # bridged neutralizes a provisioned worktree's .mcp.json for this reason; # listing it here would copy the primary's back over that. # gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected # as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302). # Opt-in by design — omit and the worker gets no PR-create grant (push over # SSH is unaffected). The token value itself is never stored in this file. # gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as # GITEA_HOST *only* alongside a resolved gitTokenEnv. # exhaustedPattern → regex matched against a completion-fallback scrape (CB-578 stage A) to # classify a turn that ended with no fleet_reply as the backend having # refused on a subscription usage limit, rather than a real answer. Opt-in — # omit and this profile's completion fallback behaves exactly as before. # Every backend words its refusal differently, so this is config, never a # vendor string baked into bridged itself. # DEFERRED: compiled once into a startup pattern map — editing it needs a # daemon restart, same as this profile's model/baseUrl/argv. # credentialId → CB-578 stage B: the credential this profile quarantines WITH when a # BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME # credentialId share one quarantine — the case this exists for is two models # on one account (e.g. sol and terra both billing one OpenAI credential): an # exhaustion on either one must lock out both, or the fleet just walks onto # the same dead account under the sibling's name. Opt-in — omit and this # profile quarantines alone, under its own name, exactly as if the field did # not exist. Cooldown length is the top-level quarantineCooldownSeconds below. # HOT: read live at every spawn/exhaustion check — no restart needed. # env → extra environment for this profile's workers, as a literal key/value map # (CB-511). Use it to give workers a toolchain. # # A worker's environment does NOT come from your shell. bridged hands herdr an # explicit env map and herdr merges it into ITS OWN process env — so before # CB-511 a worker inherited whatever PATH the herdr server happened to be # started with, which on a long-lived herdr can predate your toolchain entirely # and leave workers unable to run `mvn` or `java` at all. # bridged now propagates ITS OWN PATH to every worker by default; set `env:` # only to override that or add more (JAVA_HOME, …). Since the default is the # daemon's PATH, make sure the daemon is started with a good one — see the PATH # lines in deploy/dev.ltms.fleet.plist and deploy/bridged.service. # # Adapter-owned variables always win over `env:`: ANTHROPIC_BASE_URL and the # rest of the ANTHROPIC_*/CLAUDE_* wiring are applied after it, so an `env:` # entry cannot repoint a worker past the SubscriptionGuard — which is checked # against `baseUrl` alone. # Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously. profiles: gx10: # ccs profile name (NOT a hostname) kind: claude-code # which adapter spawns this profile (default; may omit) baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw) model: coder placement: tab workspace: bridged-workers # tabLabel: an optional per-profile override; the fleet template usually covers it mcpUrl: http://127.0.0.1:8765/mcp tokenEnv: BRIDGED_WORKER_TOKEN argv: ["ccs", "gx10"] # weight: relative selection weight for automatic placement (weighted, round-robin, and # fixed's fallback walk). Absent defaults to 1.0. An explicit 0 or negative value means # "never auto-select this profile" (CB-554) — it stays reachable via an explicit # `fleet_spawn{profile:"gx10"}`, which bypasses placement entirely; only automatic # selection skips it. weight: 0.5 # maxLoad: max live workers on this profile. Omit for unlimited. An explicit 0 (CB-585) caps # the profile at zero live members — it is excluded from automatic placement and an explicit # `fleet_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile # is named directly. Negative is refused at config load — there is no sane meaning for it. maxLoad: 2 # subscription: true # THE KNOB THAT DECIDES WHO PAYS (CB-539). Default false. When true, this profile's members # run on the OPERATOR'S OWN Claude subscription instead of a metered endpoint — every spawn # bills your plan and eats your usage limit. Off-subscription is the whole point of this # daemon, so treat `true` as a deliberate exception, not a convenience. # # What changes when it is set (ClaudeCodeLauncher): # - no ANTHROPIC_BASE_URL and no ANTHROPIC_AUTH_TOKEN are injected — the member inherits # the operator's own Claude Code auth, which is exactly why it bills the plan; # - SubscriptionGuard never vets it, because there is no baseUrl to vet; # - no token is required, so `tokenEnv` is irrelevant here. # # MUTUALLY EXCLUSIVE with `baseUrl` — setting both is refused at config load (CB-542). On the # subscription path no guard would vet the URL, so allowing both would be a way around the # guard rather than a configuration. # # GOTCHA 1 — it is invisible to the startup secret check. `Fleetd.reportRequiredSecrets` # skips subscription profiles on purpose (they need no token), so a boot log that reports # every secret as fine says nothing about these profiles. # # GOTCHA 2 — `maxLoad` is the ONLY throttle you have here. There is no metering, no budget # and no refusal on cost; the cap on live members is the single thing standing between a # fan-out and your monthly limit. Set it deliberately and keep it small. # gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302) # gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv # exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578) # credentialId: shared-openai # opt-in: quarantine together with every other profile sharing this id (CB-578) # configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP # cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's # parityOverlay: [".env", ".envrc"] # the default; never add .mcp.json or .claude/settings.local.json — see above # ideMcpUrl: http://127.0.0.1:29170/index-mcp/streamable-http # opt-in (CB-634): IDE code intelligence, pinned to the worktree # ideProjectDir: bridged # CB-634: module dir the IDE opens + the overlay pins (this repo's pom is in bridged/) # ideOpenCommand: env DISPLAY=:10.0 idea {dir} # CB-634 auto-open: opens {dir} in the IDE at spawn; omit to open by hand # autoCompactWindow: 250000 # opt-in: bound member context; claude-code compacts AT this, opencode within it (model limit.context) gx11: # a second backend, so `placement: weighted` has a choice baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token placement: tab workspace: bridged-workers # tabLabel: an optional per-profile override; the fleet template usually covers it mcpUrl: http://127.0.0.1:8765/mcp argv: ["ccs", "gx11"] weight: 0.5 maxLoad: 2 # Pin an auto-compact window BELOW the served model's context ceiling. The global # ~/.claude/settings.json value is shared by every ccs instance and the primary, so the # per-profile override belongs here. Equal to the ceiling means auto-compact never fires # before the server rejects the prompt, which kills a worker mid-turn (CB-523). env: CLAUDE_CODE_AUTO_COMPACT_WINDOW: "280000" # CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral. # opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL / # SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt. # The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a # `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude. # # Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → fleet_send → # structured fleet_reply → teardown). The `opencode/*-free` models run on opencode's own gateway # and need NO credentials — check `opencode models` for the current free list, since the names # change. That also makes the worker off-subscription by construction. # opencode-free: # kind: opencode # model: opencode/north-mini-code-free # `provider/model` selector, injected as `-m` # placement: tab # workspace: bridged-workers # tabLabel: "opencode: {profile} #{n}" # mcpUrl: http://127.0.0.1:8765/mcp # argv: ["opencode"] # # CB-508: point an opencode profile at your OWN OpenAI-compatible endpoint (local vLLM, llama.cpp, # LM Studio, TGI…) instead of opencode's gateway. Setting `baseUrl` on a `kind: opencode` profile # makes the bridge emit a custom `provider` block into the generated opencode.json — opencode has # no ANTHROPIC_BASE_URL seam, so this is how the endpoint is pinned. # baseUrl → a bare host:port gets `/v1` appended (where these servers mount the API); a URL that # already has a path is used verbatim, so a custom mount point still works. # model → MUST be "/". The provider half names the generated block; the model # half must match an id the server reports at /v1/models. One field drives both the # declaration and the `-m` flag, so they cannot drift apart. A bare model name with a # baseUrl set is rejected at spawn rather than silently using the default gateway. # tokenEnv → optional; its value becomes the provider apiKey. Most local servers ignore the key, # so a placeholder is used when unset (the AI SDK still requires a non-empty one). # NOTE: no `guard` entry is needed even with a baseUrl set. The SubscriptionGuard exists to stop a # worker borrowing the primary's Anthropic subscription, and an opencode process has no Anthropic # credential path at all. # opencode-local: # kind: opencode # baseUrl: http://127.0.0.1:8000 # model: local-vllm/deepseek-v4-flash # placement: tab # workspace: bridged-workers # tabLabel: "opencode: {profile} #{n}" # mcpUrl: http://127.0.0.1:8765/mcp # argv: ["opencode"] # How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour), # round-robin, or weighted. Omitting this key is a strict no-op for existing configs. # # `weighted` IS NOT "cheapest first" — read this before you set weights (CB-589). # It is smooth weighted round-robin: it spreads spawns across EVERY profile that has a free slot, # in weight ratio. It has no idea which profile costs money. So with local:10 / paid:2 you do not # get "use local, overflow to paid" — you get roughly one spawn in six going to the paid profile # while the local box still has a free slot. # # There is a sharper second effect. The policy's running score map lives for the daemon's whole # life. While a profile is at maxLoad it is filtered out and its score FREEZES, so the paid # profiles keep accumulating against it. When the local slot frees up it returns with a stale # score and can LOSE the next pick — a paid spawn while the free box sits idle. # # Until a real cost-first policy exists, the workaround is to make the ratio decisive rather than # proportional: give the free profile a weight so large that it wins every pick it is eligible # for, and paid profiles only ever take genuine overflow. On this host that is local weight 100 # against paid weights of ~1. # # The gotcha with that workaround: it expresses a PREFERENCE ORDER through a RATIO knob. Add a # future profile at weight 150 and it silently outranks the free box, with nothing to warn you. # Re-check the weights whenever you add a profile. placement: weighted # How long a credential sits out after a BACKEND_EXHAUSTED classification (CB-578 stage B), in # seconds, before a spawn may land on it again. Applies to every profile's effective credential # (its own name, or its credentialId if set above) — there is no per-profile override. Default # 1800 (30 minutes) when omitted or non-positive. # DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps # its original cooldown regardless; a new value only applies to a quarantine that starts after a # restart. Editing this needs a daemon restart to take effect. # quarantineCooldownSeconds: 1800 # Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an # upgraded bridged keeps the old behaviour: the file is read once at boot and never again. # enabled → turn the watch on. bridged checks the file's modified time on a timer and # reloads when it moves. # intervalSeconds → how often to check (default 10). One `stat` per tick, so this is cheap. # # Not every key can move under a running daemon, and the difference is about what already exists # when the reload happens — not about how important the key is: # HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool, # `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad # / credentialId. Those are hot because the placement policy (and, for credentialId, # the CB-578 stage B quarantine check) reads them through a supplier — being config is # not by itself enough to make a key hot. # EXCEPT `fleet.leaders`: Fleetd.main reads it once at startup to build the lead tab # scanner and launcher, and neither is rebuilt on reload. A changed/added/removed # `fleet.leaders` entry is silently accepted — the reload reports "config reloaded" # with nothing in the deferred list — but has NO effect until you restart. Treat it # as deferred in practice, even though today's reload output does not say so. # DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value # until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`, # `spawnReadyTimeoutMs` / `spawnReadyPollMs`, `quarantineCooldownSeconds` (CB-578 # stage B — baked once into the quarantine tracker built at startup), ADDING or # REMOVING a profile (a new backend needs its own launcher, and launchers are built # once), AND an existing profile's launch settings — model, baseUrl, argv, env, # configDir, mcpUrl, tabLabel, exhaustedPattern. The launcher takes a copy of # `profiles:` at startup and resolves every spawn out of that copy, so those never # reach a launch until you restart. The reload logs them by name rather than # pretending they applied. # COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is # bound, the broker connection is open, and the auth mode decides who may reach the # port that is already listening. # # A changed COLD key refuses the WHOLE reload — not the hot half applied and the cold half warned # about. A half-applied reload would leave the daemon matching no file on disk, which is the worst # thing a reload can do to an operator debugging one. A file that fails to parse or fails a startup # validator is refused the same way, and the running config stays live. # configReload: # enabled: true # intervalSeconds: 10 # THE FLEET (CB-557) — who the daemon may run, and under which role. This one block replaced four # older keys: `leaders:`, `members:`, `leadScan:` and `defaultProfile:`. # # A member is anything a lead spawns, and every member has two INDEPENDENT attributes: # role — which contract: architect, dev or reviewer. It picks the launch charter, the role # file, the playbook skill and the authz row. # profile — which backend: one of the `profiles:` keys above (model, CLI adapter, cost). # They vary on their own. A reviewer may run on the same profile as the dev whose diff it reads, # which is why the two cannot be one field. # # The ROLE IS THE CONTAINING KEY, not a `role:` field. That is not only tidier: a misspelled role # used to parse into a member with no contract at all, while a misspelled pool name here simply # declares nothing. # # Each pool lists the profiles that role MAY run on — these are pools, not identities. That is also # what replaced `defaultProfile:`: an unqualified spawn names a role, and that role's pool supplies # the candidates, in definition order. A dev and a reviewer staying anonymous is exactly compatible # with being listed here; the entry key just names the entry. fleet: # Optional launch-charter text, keyed only by the singular role wire names: architect, dev, # reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put secrets # here: a later launch step writes this text to a world-readable temp file, and ${ENV} interpolation # is deliberately not supported. charters: architect: |- You are an architect in this fleet. You refine work before anyone builds it: scope, acceptance criteria, risks, and a unit split. You read the repo and write analysis. You never commit production code and never open a PR. A design task is worked by two architects. Design alone first, then exchange and say plainly where you disagree. Do not concede just to agree. dev: |- You implement the one unit you were given, and nothing else. You test it, commit it, and open your own pull request. You never merge. reviewer: |- You review the diff you were given. You report bugs, risks and missing tests. You do not change code. # Optional. Template for a member tab's label; {role}, {profile}, {model} and {n} are substituted. # {n} counts per role+profile, so `dev: sonnet #2` really is the second sonnet dev. Because {role} # comes from a closed enum, a generated label can never begin with a lead's tabPrefix. # tabLabel: "{role}: {profile} #{n}" # Panes that orchestrate rather than are orchestrated. A lead may now be CREATED as well as # recognised: give it a `profile:` and the daemon launches the shortfall when fewer than # `instances` are live. Omit `profile:` and it is recognise-only, as before. # # `tab:` (CB-579) is REQUIRED and is the only field identity depends on — the exact label of the # tab hosting the lead, matched case-insensitively. Label the tab yourself and put that same # string here, and the pane is recognised on the next rescan. Reopen the tab later, or the session # inside it restarts — the terminal id changes; the tab, and its label, do not, so no config edit # follows a restart. # # A lead the daemon launches is labelled BY the daemon with this same `tab:` value, so it is found # by the same scan. A lead counts as live only when herdr also reports a running agent in that # tab — a label left behind by a session that died does not block the relaunch, and a tab that is # gone entirely drops out of the next scan rather than being remembered forever. # # An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with # the session lifecycle (the idle reaper would kill your orchestrator), and stays on the # subscription — ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says. # # GET THE `tab:` VALUE RIGHT. A pane that does not match any configured `tab:` (a typo, a renamed # tab, a pane no entry names at all) is not recognised as a lead — it resolves as an ordinary # WORKER instead, silently, and every orchestration call it makes (spawn/stop/send/drain) is # refused. There is no error at startup for this: an unmatched pane is simply not a lead. If your # primary suddenly can't spawn or send, check this section first. # leaders: # opus-5.0: # profile: opus # omit to never create this lead, only recognise it # instances: 1 # desired live count; only the shortfall is launched. 0 = off # tab: "lead: opus-5.0" # REQUIRED — the exact tab label this lead lives in # tabPrefix: "lead:" # only used to guard against a worker tabLabel colliding with # # this convention at startup; plays no part in matching a lead # scanIntervalSeconds: 10 # rescan cadence, and the worst case before a new tab is seen # workspace: leads # where a launched lead's tab is created (default "leads"). # # MUST NOT be a member workspace — those are excluded from the # # scan, so a lead placed in one is never found again. # cwd: /path/to/repo # the launched lead's working directory (default: bridged's own) # kind: claude # descriptive; reported by fleet_whoami # gpt-sol-5.6: # tab: "lead: gpt-sol-5.6" # kind: opencode # model: openai/gpt-5.6-terra # architects: # architect-1: # profile: opus # a strong model, on the operator's subscription # architect-2: # profile: sol # a different vendor on purpose — two architects that share a # # model share its blind spots developers: gx10: profile: gx10 # reviewers: # gx10: # profile: gx10 # the same backend may serve two roles; that is the point # Subscription boundary. A worker's base_url host MUST be one of these; the primary # must carry none. Every profile above must have its host listed here. guard: offSubscriptionHosts: - gx00.gw - gx01.gw # Member credential policy (CB-596, gitea issue #82). A herdr pane runs a LOGIN shell, and that # shell re-sources the operator's own secret store — so a spawned member inherits every credential # the operator's shell holds, not just the ones bridged means to give it. Measured on this host: # 31 credential names, all set, with only ONE (GITEA_ACCESS_TOKEN) blocked before this — and that # block was a single name hardcoded in HerdrPeerLauncher.java, not driven by this file. This block # replaces that hardcoded shadow with a config-driven list of names. # # ROUND-2 CORRECTION, measured live: the pane-creation env overlay below (applied at tab.create / # pane.split, BEFORE the pane's login shell runs) does NOT survive that login shell for any name # secrets.sh actually exports — the shell re-exports it afterwards and overwrites the sentinel. # Proof: GITEA_ACCESS_TOKEN comes back blocked only because secrets.sh itself carries a guarded # export (`[ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...`) — that guard, not this # file, is what wins. No other name in `known` below has a matching guard in secrets.sh yet (1 # guard measured against 33 export lines there). So today this block's overlay is REAL protection # only for a name secrets.sh does not export, or a peer kind whose pane never runs a login shell — # for everything secrets.sh exports and guards, the guard in secrets.sh (out of scope for this # ticket) is what actually blocks it, not this list. An exec-time fix (winning after the login # shell finishes, before the agent process starts) was attempted and found to have no seam in the # current herdr protocol — AgentControl.start takes a fixed `kind` (herdr resolves the executable) # plus trailing CLI args for that binary, not an arbitrary argv or an env map; only tab.create / # pane.split accept `env`, and that is this same pane-creation overlay. See gitea #82 for the open # design question this leaves. # # DENY-BY-DEFAULT, NOT A DENY-LIST. A deny-list (name the bad ones, let everything else through) is # silently wrong the moment the operator's store gains a new secret — nothing would ever report it. # Deny-by-default inverts that: `known` bounds the blast radius to names actually enumerated below, # and EVERY one of them is blocked UNLESS it is also in `allow`. Omitting this block entirely (the # shipped default) blocks NOTHING — unlike most optional blocks in this file, absence here is a real # gap, not a safe "feature off". A name that is neither `known` nor `allow`-ed is not silently let # through either: the daemon logs a WARN naming any credential-shaped env var it finds on neither # list (never its value), so a secret added to the store later does not go unnoticed forever. # # policy → "deny-by-default" (the default; also accepted spelled "deny-list") overlays each # known-but-not-allowed name BEFORE the pane's login shell runs — real protection only # where that shell does not re-export the name (see ROUND-2 CORRECTION above). An # unrecognized value refuses to start, naming it. # policy → "allow-list" (CB-633) moves the control to a per-spawn ZDOTDIR directory the daemon # generates and passes through tab.create's env map. Each generated startup file sources # its ~/ counterpart FIRST and then runs the scrub, so the scrub happens after the # operator's whole chain and no sourced file can undo it. # The scrub is sourced from BOTH the generated .zshrc and the generated .zlogin, because # herdr does not open the same kind of shell everywhere: macOS panes run a LOGIN zsh (so # .zlogin runs), Linux panes run a plain interactive zsh (so .zlogin never runs at all). # A scrub in .zlogin alone would be a control that silently does nothing on Linux. # The allow-list is DERIVED, never typed: # every profile's tokenEnv/gitTokenEnv/gitHostEnv values and env-map keys, plus an # infrastructure set (PATH HOME SHELL TERM LANG LC_* TMPDIR USER LOGNAME PWD SHLVL EDITOR # PAGER JAVA_HOME XDG_* ZDOTDIR), plus whatever keys this spawn's own env overlay carries. # Adding a profile can therefore only widen the list, never break another spawn's scrub. # Under this policy `known`/`allow` below become REPORTING ONLY — they feed the gap WARN, # they are no longer a control. If the member's login shell is NOT zsh, the daemon logs a # loud WARN saying protection is off and falls back to deny-by-default's overlay. # Each pane writes a scrub-report.txt naming how many variables it kept of how many it # saw; the daemon logs that "allowed N of M" line when the pane stops. If the report is # MISSING the daemon logs a WARN instead — the scrub cannot then be confirmed to have # run, and a silently-dead control is exactly what this policy exists to prevent. # allow → credential names a member legitimately needs. Under deny-by-default, left OUT of the # pane's env overlay entirely, so the value the pane's own (login) shell exports passes # through untouched. Under allow-list: reporting only. # known → every credential name the operator's store is known to export. Under deny-by-default, # every name here NOT also in `allow` is overlaid with a non-secret sentinel value before # the pane's login shell runs — real protection only for names that shell does not itself # re-export (see the ROUND-2 CORRECTION note above). Under allow-list: reporting only. # sshAuthSock → whether SSH_AUTH_SOCK may pass through under allow-list ("allow") or must be # blanked like any other non-derived name ("block", the default). This is a decision you # have to make explicitly: SSH_AUTH_SOCK is a handle to YOUR ssh-agent, and a member # holding it can sign with your keys — it sits in no secret file and looks like no # credential, which is why it slipped past three earlier tickets (gitea #110). Blocking # it breaks git over SSH inside members (push/fetch authenticate as you); use HTTPS # remotes or scoped deploy keys instead of allowing it lightly. # # HOT-RELOADABLE the same way `fleet:` is (CB-559): read fresh on every spawn, so editing this list # and reloading config (or restarting) changes what the NEXT spawn inherits; already-running members # are unaffected either way. # memberCredentials: # policy: deny-by-default # or "deny-list", or "allow-list" (CB-633) — see above # sshAuthSock: block # allow-list only; see the sshAuthSock note above # allow: # - AI_GATEWAY_TOKEN # named in a profile's tokenEnv (local/gx) — a member reaching the # # gateway is by design, not a leak # - WORKER_GITEA_TOKEN # the repo-scoped forge token a member needs to open its own PR (CB-302) # - CONTEXT7_TOKEN # already decided as allowed by CB-593 # - GITEA_HOST # not a credential — a hostname, paired with the forge token above # known: # - AI_GATEWAY_TOKEN # - BESZEL_ADMIN_EMAIL # - BESZEL_ADMIN_PASSWORD # - BESZEL_HUB_URL # - BESZEL_KEY # - BESZEL_UNIVERSAL_TOKEN # - BRAIN_MCP_TOKEN # - CF_ACCOUNT_ID # - CF_API_TOKEN # - CF_USER_TOKEN # - CONFLUENCE_API_TOKEN # - CONFLUENCE_USERNAME # - CONTEXT7_TOKEN # - GITEA_ACCESS_TOKEN # - GITEA_HOST # - GITLAB_OAUTH_CLIENT_SECRET # - GITLAB_PERSONAL_ACCESS_TOKEN # - GRAFANA_ADMIN_PASSWORD # - GRAFANA_ADMIN_USER # - HASS_TOKEN # - HW_PASSWORD # - HW_USER # - LTMS_API_KEY # - MEMORY_MCP_TOKEN # - METRICS_PUSH_TOKEN # - OPENCODE_AUTOMODE_MODEL # - TELEGRAM_BOT_TOKEN # - TELEGRAM_CHAT_ID # - TS_API_KEY # - TS_AUTHKEY # - WORKER_GITEA_TOKEN # Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is # injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate. # NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and # unknown keys are ignored, so a snake_case key would be silently dropped (default kept). # spawnReadyTimeoutMs: 20000 # spawnReadyPollMs: 300 # Worktree provisioning root (CB-301-ext). Where per-worker git worktrees are checked out so # each worker owns an isolated branch instead of sharing the primary's tree. Omit to default # to a sibling directory of the repo root. # worktreeRoot: /Users/me/src/.bridged-worktrees # Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep # the feature disabled. By default the daemon never reaps, caps, or drains sessions. # idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING) # contextCap → force-release a session after this many delegated turns # drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown # clearAfterTurn → whether a reusable worker discards its conversation context after every # completed delegated turn (default false). Works for claude-code workers # only — any other peer kind (e.g. opencode) logs "context reset is # unsupported for peer kind …" once and the reset is a no-op. # lifecycle: # idleTtlSeconds: 300 # contextCap: 10 # drainTimeoutSeconds: 5 # clearAfterTurn: false # Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default # in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce). # Set a broker uri to swap in the AMQP-backed inbox: worker replies with no open send are held # on a durable per-target queue (agent..inbox) and survive a restart — the broker # redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock # RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap. # uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/") # is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost. # uriEnv → CB-151: name of a host env var holding the AMQP URI, preferred over `uri` (wins # whenever set). The URI carries `user:pass@` inline, so naming a variable keeps the # password out of fleetd.yaml — same pattern as auth.tokenEnv/Profile.tokenEnv. A # uriEnv that resolves to an unset or blank variable is treated as NOT configured and # the daemon falls back to the in-memory inbox, warning loudly. # prefetch → CB-527: consumer basicQos, capping how many unacked messages the inbox holds # in-heap per owned target (the rest sits on the broker's durable queue instead of # growing the JVM heap). Default 32 when omitted. # broker: # uriEnv: LAVINMQ_URI # prefetch: 32 # Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open fleet_send, # the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr # pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop # stops as soon as the primary's inbox is empty. # terminal → pin the primary's herdr terminal id. Omit to learn it from the connection on # the first orchestration-side MCP call (the normal case). An off-host or # non-herdr primary leaves this unresolved → the loop is a no-op and delivery # degrades to pull; the reply is still never lost. # # REQUIRED (CB-522) if the primary itself runs inside a herdr pane. Caller # identity resolves a loopback PID to its herdr pane, and PaneLocator scans # EVERY pane — not just bridged-spawned ones — so such a primary is otherwise # classified as a WORKER and refused SPAWN/SEND/STOP. That failure is # self-locking: the learned terminal is populated by the very orchestration # calls being refused, so only this pinned value can break the cycle. Read the # id off fleet_whoami (it reports the current terminal even while # misclassified) and re-pin whenever the primary moves panes. # pushReminders → max nudges before giving up (default 5) # pushBackoffMs → delay between nudges in ms (default 15000) # primary: # terminal: term_65619bd6174568 # pushReminders: 5 # pushBackoffMs: 15000