4bfab6b718
LeadRollover.open() now resolves handoverPath to an absolute path exactly once, against the calling lead's fleet.leaders.<name>.cwd (falling back to the daemon's own user.dir when that lead has none configured), matching the LeadLauncher#launch precedent. PendingRollover stores only the resolved absolute path, so checkHandover's exists/empty/fresh checks, the path handed back to the lead in the fleet_handover open response, and the default bootstrapText sentence all see the same absolute location instead of a value resolved against whatever directory the daemon process happened to start in. FleetConfig.LeadRollover.bootstrapText is no longer defaulted in the compact constructor (it would otherwise still bake in the raw, possibly-relative handoverPath); a new bootstrapTextFor (resolvedHandoverPath) method builds the default sentence from the resolved path instead. Fleetd.leadRollover(...) gains a required liveLeadTerminals parameter to build the terminal to lead-name to Leader.cwd lookup, read live through the existing `leads` supplier and ConfigRef on every call, never off a startup snapshot.
974 lines
70 KiB
YAML
974 lines
70 KiB
YAML
# fleetd configuration (example). Copy to fleetd.yaml and adjust.
|
|
#
|
|
# fleetd is the sole gateway between primary/worker Claude sessions and herdr.
|
|
# It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL.
|
|
|
|
# REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token`
|
|
# below — fleetd REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
|
|
bind:
|
|
host: 127.0.0.1
|
|
port: 8765
|
|
|
|
# API authentication (CB-501). Governs how a caller that is NOT an on-host worker pane proves it
|
|
# is the primary. Worker identity never depends on this: a loopback peer PID that maps to a herdr
|
|
# pane is unforgeable and is always honoured, so turning auth on cannot lock the fleet out.
|
|
#
|
|
# mode: loopback-trust → DEFAULT, and the historical behaviour: any loopback caller that is not
|
|
# a worker is the primary, no credential needed. Sound ONLY because the
|
|
# OS refuses remote connections to a loopback socket.
|
|
# mode: token → such a caller must send `Authorization: Bearer <token>`; without it it
|
|
# is anonymous and authorized for nothing. REQUIRED for a non-loopback
|
|
# bind — the daemon fails fast otherwise, because "unauthenticated ⇒
|
|
# primary" on a reachable port would hand spawn/stop/send to anyone.
|
|
# tokenEnv → host env var holding the token (never the literal value). Default
|
|
# FLEETD_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
|
|
#
|
|
# TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and
|
|
# let it own certificate lifecycle, e.g.
|
|
# location / { proxy_pass http://127.0.0.1:8765; proxy_set_header Authorization $http_authorization; }
|
|
# The broker link gets TLS from its own URI (amqps://…) — see `broker` below.
|
|
# auth:
|
|
# mode: token
|
|
# tokenEnv: FLEETD_API_TOKEN
|
|
|
|
# Optional pinned primary terminal (CB-307). Names the herdr pane the PRIMARY itself runs in:
|
|
# a caller whose connection maps to this pane resolves as the primary (no credential needed —
|
|
# the pane mapping is as unforgeable as a worker's), and reply nudges are pushed to it.
|
|
# REQUIRED when the primary runs inside a herdr pane — without it the pane match reads the
|
|
# primary as a worker and refuses spawn/send/stop. Get the id from fleet_whoami; re-pin if
|
|
# the primary moves panes.
|
|
# primary:
|
|
# terminal: term_0123456789abcd
|
|
# pushReminders: 5 # max nudges before giving up (default 5)
|
|
# pushBackoffMs: 15000 # delay between nudges (default 15000)
|
|
|
|
# CB-530: MORE THAN ONE LEAD. `primary:` above is singular by construction — every other pane
|
|
# resolves as a worker — which is right for one lead driving a fleet and wrong the moment two leads
|
|
# (say a Claude lead and an opencode lead) work as peers: the second is silently demoted and refused
|
|
# every orchestration call. List each lead's pane here and all of them resolve as leads.
|
|
#
|
|
# tab → the ONLY field identity depends on (CB-579); the exact label of the tab hosting the lead.
|
|
# Label the tab yourself, or let fleetd label one it launches — see `fleet.leaders:` below.
|
|
# kind/model → descriptive; they document what runs in the pane and are echoed by fleet_whoami
|
|
#
|
|
# A lead's tab must already carry its label (or be launched by fleetd, which labels it) — there is
|
|
# no terminal id to paste in and nothing to re-pin when the session restarts: the tab survives, so
|
|
# the same label resolves the same lead again on the next scan.
|
|
# `fleet_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
|
|
# IS a primary for authorization, so nothing that keys on the role breaks.
|
|
#
|
|
# KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single
|
|
# destination for its nudges, and is a separate mechanism from lead identity — see `fleet.leaders:`.
|
|
#
|
|
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
|
|
#
|
|
# Two things stop the tab-name convention from becoming a way to claim leadership: the configured
|
|
# member spaces are excluded from the scan, so nothing fleetd places can land in a matching tab;
|
|
# and startup REFUSES a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel`
|
|
# override, also matches — so the two namespaces cannot overlap by accident. The label is a NAME,
|
|
# never a capability: what a pane may do is decided by the role the daemon resolves for it.
|
|
|
|
# CB-551: IDLE-LEAD HEARTBEAT — nudge the single lead back to work when it has been continuously
|
|
# idle (no open fleet_send driving it) past the quiet period. The fleet is one lead + architects +
|
|
# workers, so a lead that stalls is a single point of failure; the ReplyPushLoop only nudges when a
|
|
# reply lands, and this timer catches the gap where nothing lands and the lead just sits idle.
|
|
#
|
|
# Opt-in on purpose — it SPENDS the operator's subscription on its own initiative (each nudge starts
|
|
# a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent
|
|
# block = feature off, exactly as before.
|
|
#
|
|
# Three knobs, each with a default that errs on the side of not burning context:
|
|
# idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 —
|
|
# # absorbs normal post-turn pauses; re-prompting every pause burns context)
|
|
# backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000)
|
|
# quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops
|
|
# # until real state appears (default 3 — never nag an empty fleet forever)
|
|
# leadHeartbeat:
|
|
# idleAfterSeconds: 300
|
|
# backoffMs: 60000
|
|
# quietNudgeCap: 3
|
|
|
|
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
|
|
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
|
|
# own pane and bootstrap a fresh session against that file.
|
|
#
|
|
# Opt-in on purpose — it clears the lead's own pane on request, so upgrading the daemon must never
|
|
# acquire that ability for you. Absent block = feature off, and nothing is constructed at all. Even
|
|
# once present, nothing but an explicit confirm() call — one that passes every check — can ever
|
|
# cause a /clear: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
|
|
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so
|
|
# it cannot clear the pane inline (that pane is still WORKING); instead it schedules a one-shot
|
|
# continuation that waits for the SAME confirm() call's turn to end, then does the actual work. See
|
|
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order (fleetd #480 correction).
|
|
#
|
|
# handoverPath: REQUIRED when this block is present — where the handover file a fresh lead session
|
|
# reads must live. No default (an operator-specific path); a present block with no
|
|
# handoverPath refuses to start. May be relative: it then resolves against the
|
|
# CALLING lead's own fleet.leaders.<name>.cwd (falling back to the daemon's own
|
|
# working directory when that lead has none configured) — never against whatever
|
|
# directory the daemon process happens to have been started in. An absolute path is
|
|
# used unchanged. Prefer an absolute path if the daemon and the lead's pane might not
|
|
# share a working directory (fleetd #480 follow-up).
|
|
# requireOperatorConfirm: true # default true — confirm() refuses unless the caller also passes
|
|
# # operatorConfirmed: true
|
|
# maxDocAgeSeconds: 3600 # default 3600 — refuse a handover file older than this
|
|
# turnSettleSeconds: 20 # default 20 — how long the deferred roll waits for the CALLING
|
|
# # lead's own turn to end (its pane to report injectable again)
|
|
# # before sending /clear at all. If this elapses, /clear is NEVER
|
|
# # sent — a lead that never goes idle is still doing real work.
|
|
# clearSettleSeconds: 20 # default 20 — how long to wait for the pane to become injectable
|
|
# # again AFTER /clear before giving up (never sends bootstrapText
|
|
# # if this elapses). A separate, second wait from turnSettleSeconds.
|
|
# bootstrapText: "..." # default names the RESOLVED (absolute) handoverPath — sent to
|
|
# # the lead once its pane settles after /clear
|
|
# leadRollover:
|
|
# handoverPath: /path/to/handover.md
|
|
# requireOperatorConfirm: true
|
|
# maxDocAgeSeconds: 3600
|
|
# turnSettleSeconds: 20
|
|
# clearSettleSeconds: 20
|
|
# bootstrapText: "Fresh lead session: read the handover file and carry on."
|
|
|
|
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
|
|
# per tick.
|
|
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
|
|
# Math.max(15, intervalSeconds), so a lower value is silently raised, not
|
|
# rejected.
|
|
# workingSuspectAfterSeconds → age before a BUSY member is suspected of a stall (default 600).
|
|
# ENFORCED floor of 300: a lower value is silently raised.
|
|
# paneProbeIntervalSeconds → accepted and parsed, but NOT YET READ by anything. Setting it changes
|
|
# nothing right now. It exists so a later build can start honouring it without
|
|
# another config-shape change.
|
|
# notifications.mode → "webhook" flips what fleet_list REPORTS (healthCoverage: "full" instead
|
|
# of "detection-only") — it does NOT make fleetd send any webhook call; no
|
|
# delivery mechanism is implemented yet. Any other value, or omitting the
|
|
# block, reports "detection-only".
|
|
# health:
|
|
# enabled: true
|
|
# intervalSeconds: 30 # floor 15
|
|
# workingSuspectAfterSeconds: 600 # floor 300 — how long BUSY with no activity means STALL_SUSPECTED
|
|
# paneProbeIntervalSeconds: 60 # parsed, but nothing reads it yet — changing it changes nothing
|
|
# notifications:
|
|
# mode: disabled
|
|
|
|
# Idle-sleep guard: while at least one member is live, hold an OS-level assertion against idle
|
|
# sleep (macOS only — a `caffeinate -i` child; a no-op elsewhere or if caffeinate is missing), so
|
|
# an unattended host does not idle-sleep out from under a member's long turn. Unlike health/
|
|
# configReload above, this is ON BY DEFAULT — omitting the block entirely leaves it enabled, the
|
|
# same as `enabled: true`. Uncomment only to turn it off:
|
|
# idleSleepGuard:
|
|
# enabled: false
|
|
|
|
# herdr Unix socket. Omit to use the client default
|
|
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
|
|
herdrSocket: ~/.config/herdr/herdr.sock
|
|
|
|
# Optional socket for member panes. Omit this to use herdrSocket for both leads and members.
|
|
# memberHerdrSocket: /Users/member/.config/herdr/herdr.sock
|
|
|
|
# fleetd #213: the login shell the member OS user (memberHerdrSocket above) actually runs. ONLY
|
|
# read when memberHerdrSocket is set — fleetd's own $SHELL says nothing about a pane running
|
|
# under a different OS user, and there is no channel to ask herdr for that user's shell, so this
|
|
# must be told rather than guessed. Absent, blank, or anything not ending in "zsh" is treated the
|
|
# same as "not zsh": the memberCredentials.policy: allow-list ZDOTDIR scrub (see worktreeGroup
|
|
# below) is skipped in favour of the weaker CB-596 sentinel overlay — a degraded control, never a
|
|
# refusal to spawn. When memberHerdrSocket is absent this key is never consulted at all.
|
|
# memberLoginShell: /bin/zsh
|
|
|
|
# How member sessions are spawned. Define one or more named profiles (backends) under
|
|
# `profiles`; each key is the profile name (also the ccs profile). A profile says only WHICH
|
|
# BACKEND — model, CLI adapter, credentials, cost. It says nothing about what a member spawned on
|
|
# it is for; that is the member's role, and roles live under `fleet:` below. Which profile an
|
|
# unqualified spawn lands on comes from that role's pool, not from a global default.
|
|
#
|
|
# Shared knobs (placement/workspace) can be repeated per profile; they usually match.
|
|
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
|
|
# Use `pane` for the legacy behaviour (split the focused tab).
|
|
# mcpUrl → fleetd mounts the bridge MCP (--mcp-config, inline) + reply charter
|
|
# (--append-system-prompt) as launch flags; nothing is written to the profile.
|
|
# ideMcpUrl → opt-in (CB-634), default off. When set, fleetd mounts the IDE Index MCP as a
|
|
# second inline server named `intellij`, and adds an IDE charter that pins every
|
|
# ide_* call to the member's own worktree. A URL, not a boolean — host and port
|
|
# are host-specific. Set it only on a host where the IDE actually runs.
|
|
# ideProjectDir → repo-relative module dir the IDE opens and the overlay pins (CB-634). Only read
|
|
# when ideMcpUrl is set. This repo's Maven pom lives in `fleetd/`, not at the
|
|
# worktree root, so opening the root imports no module and ide_* resolves nothing;
|
|
# set this to `fleetd`. Omit for a repo whose project is the worktree root.
|
|
# ideOpenCommand → host command that opens ideProjectDir in the IDE at spawn (CB-634 auto-open).
|
|
# Only read when ideMcpUrl is set. `{dir}` is replaced with the absolute module
|
|
# dir and the command runs through `/bin/sh -c`, so set env inline if needed —
|
|
# e.g. `env DISPLAY=:10.0 idea {dir}`. Best-effort: a failure is logged, never
|
|
# fails the spawn. Omit to open the member's module by hand. There is no close
|
|
# half yet — an opened module stays open until the operator closes it.
|
|
# autoCompactWindow → opt-in, default off. A bounded token window that forces a spawned member to
|
|
# compact its context instead of running on the backend's own default and dying
|
|
# mid-turn (losing its fleet_reply — the whole point of the turn — with it).
|
|
# Validated at config load to [100000, 1000000] — the band Claude Code's own
|
|
# --autocompact flag accepts.
|
|
# CROSS-BACKEND SEMANTICS DIFFER: on claude-code this is a launch-time
|
|
# `--autocompact <tokens>` flag — the member compacts AT this window. opencode
|
|
# has no equivalent flag (it only forces `compaction.auto: true`, unconditionally,
|
|
# already), so this is instead applied as the model's `limit.context` in the
|
|
# generated opencode.json — the member compacts WITHIN this window, not exactly
|
|
# at it — and only when this profile's `model:` is in `provider/model` form; if it
|
|
# isn't, fleetd logs a WARN naming the profile rather than silently doing nothing.
|
|
# tokenEnv → host env var holding the worker's auth token (value never stored in config);
|
|
# omit for a backend that needs no token (e.g. a local ollama).
|
|
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
|
|
# cwd on an MCP spawn, else the daemon's cwd — never $HOME. See
|
|
# docs/Worker-Startup-and-Trust.md.
|
|
# configDir → CLAUDE_CONFIG_DIR for the worker, so it inherits that profile's
|
|
# skills/MCP/hooks. Omit to leave the worker on the host default.
|
|
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
|
|
# worktree sees the same local config (CB-301-ext). Omit for the default set:
|
|
# [.env] only (CB-148). .envrc is left out of the default on purpose: it is
|
|
# executable shell that direnv runs on every cd, so copying it carries
|
|
# behaviour into the worker, not just values, unlike .env. An operator who
|
|
# wants it copied can still write parityOverlay: [.env, .envrc] explicitly.
|
|
# (.claude/settings.local.json is NOT in the default — it
|
|
# pre-approves IDE/tool grants a member must not hold ambiently; CB-525/CB-634.)
|
|
#
|
|
# Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher
|
|
# mounts — the bridge, and nothing else. Replicating the primary's MCP config
|
|
# handed a worker the primary's IDE servers, which are bound to the primary's
|
|
# checkout, so its navigation returned paths OUTSIDE its own worktree: one
|
|
# worker made all 59 of its edits in the primary tree while compiling its
|
|
# worktree, and every build it ran was of code that did not contain them.
|
|
# fleetd neutralizes a provisioned worktree's .mcp.json for this reason;
|
|
# listing it here would copy the primary's back over that.
|
|
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
|
|
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
|
|
# Opt-in by design — omit and the worker gets no PR-create grant (push over
|
|
# SSH is unaffected). The token value itself is never stored in this file.
|
|
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
|
|
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
|
|
# exhaustedPattern → regex matched against a completion-fallback scrape (CB-578 stage A) to
|
|
# classify a turn that ended with no fleet_reply as the backend having
|
|
# refused on a subscription usage limit, rather than a real answer. Opt-in —
|
|
# omit and this profile's completion fallback behaves exactly as before.
|
|
# Every backend words its refusal differently, so this is config, never a
|
|
# vendor string baked into fleetd itself.
|
|
# HOT (fleetd #446): read live, cached by profile name, at every
|
|
# completion-fallback check AND by fleet_profiles' exhaustionDetectionArmed —
|
|
# editing it and reloading arms or disarms usage-limit detection for this
|
|
# profile with no daemon restart. (Before fleetd #446 this was DEFERRED,
|
|
# compiled once into a startup pattern map like model/baseUrl/argv still are —
|
|
# see errorPattern below, which is still deferred that way on purpose.)
|
|
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
|
|
# BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME
|
|
# credentialId share one quarantine — the case this exists for is two models
|
|
# on one account (e.g. sol and terra both billing one OpenAI credential): an
|
|
# exhaustion on either one must lock out both, or the fleet just walks onto
|
|
# the same dead account under the sibling's name. Opt-in — omit and this
|
|
# profile quarantines alone, under its own name, exactly as if the field did
|
|
# not exist. Cooldown length is the top-level quarantineCooldownSeconds below.
|
|
# HOT: read live at every spawn/exhaustion check — no restart needed.
|
|
# errorPattern → fleetd #201 / #227: regex matched against a completion-fallback scrape to
|
|
# classify a turn that ended with no fleet_reply as a BACKEND ERROR — a
|
|
# credential outage or a provider 5xx — rather than a real answer or a
|
|
# usage-limit exhaustion (exhaustedPattern above always wins when a line
|
|
# matches both). Opt-in. Omit it and this profile falls back to fleetd's
|
|
# built-in legacy pattern `(?i)\bAPI Error\s*:` — classification still
|
|
# happens, just without a profile-specific match; every backend words its
|
|
# failure differently, so a hardcoded sentence would only ever match one
|
|
# of them.
|
|
# DEFERRED: compiled once into a startup pattern map — editing it needs a
|
|
# daemon restart. Unlike exhaustedPattern above (made hot by fleetd #446),
|
|
# errorPattern was scoped out of that ticket on purpose and stays deferred.
|
|
# # errorPattern: "503 Service Unavailable" # opt-in: classify a backend outage
|
|
#
|
|
# What happens once a match fires (BackendOutagePolicy, credentialId-keyed,
|
|
# SEPARATE from the CB-578 stage B quarantine above and never merged with it):
|
|
# - threshold 2 — TWO DISTINCT TARGETS (never raw events) on the same
|
|
# effective credential inside a 60-second window start an "incident" and a
|
|
# 60-second cool-off for that credential. One member repeating the same
|
|
# classified line twice never cools anything off — a real outage hits
|
|
# every target on that credential, so requiring a second, independent
|
|
# target loses nothing against the case this guards against, while
|
|
# protecting against a heuristic misfire on one flaky member.
|
|
# - a fresh error while a credential is already cooling off is ignored
|
|
# outright: it neither extends the 60s deadline nor starts a new incident.
|
|
# - `fleet_list`/`fleet_profiles` report a cooling credential with
|
|
# `coolingOffForSeconds` (never `quarantinedForSeconds`, unless CB-578
|
|
# exhaustion quarantine is ALSO independently active for the same
|
|
# credential — the two checks can both fire at once). A spawn onto a
|
|
# cooling profile is refused with a message naming the credential and
|
|
# remaining seconds — "cooling off", never "exhausted", so an operator can
|
|
# tell a short transient fault from a spent subscription at a glance.
|
|
# - the lead gets ONE nudge per incident (not one per affected target), via
|
|
# the same push loop that already delivers ticket/question reminders.
|
|
# env → extra environment for this profile's workers, as a literal key/value map
|
|
# (CB-511). Use it to give workers a toolchain.
|
|
#
|
|
# A worker's environment does NOT come from your shell. fleetd hands herdr an
|
|
# explicit env map and herdr merges it into ITS OWN process env — so before
|
|
# CB-511 a worker inherited whatever PATH the herdr server happened to be
|
|
# started with, which on a long-lived herdr can predate your toolchain entirely
|
|
# and leave workers unable to run `mvn` or `java` at all.
|
|
# fleetd now propagates ITS OWN PATH to every worker by default; set `env:`
|
|
# only to override that or add more (JAVA_HOME, …). Since the default is the
|
|
# daemon's PATH, make sure the daemon is started with a good one — see the PATH
|
|
# lines in deploy/dev.ltms.fleet.plist and deploy/fleetd.service.
|
|
#
|
|
# Adapter-owned variables always win over `env:`: ANTHROPIC_BASE_URL and the
|
|
# rest of the ANTHROPIC_*/CLAUDE_* wiring are applied after it, so an `env:`
|
|
# entry cannot repoint a worker past the SubscriptionGuard — which is checked
|
|
# against `baseUrl` alone.
|
|
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
|
|
profiles:
|
|
gx10: # ccs profile name (NOT a hostname)
|
|
kind: claude-code # which adapter spawns this profile (default; may omit)
|
|
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
|
|
model: coder
|
|
placement: tab
|
|
workspace: fleetd-workers
|
|
# tabLabel: an optional per-profile override; the fleet template usually covers it
|
|
mcpUrl: http://127.0.0.1:8765/mcp
|
|
tokenEnv: FLEETD_WORKER_TOKEN
|
|
argv: ["ccs", "gx10"]
|
|
# weight: relative selection weight for automatic placement (weighted, round-robin, and
|
|
# fixed's fallback walk). Absent defaults to 1.0. An explicit 0 or negative value means
|
|
# "never auto-select this profile" (CB-554) — it stays reachable via an explicit
|
|
# `fleet_spawn{profile:"gx10"}`, which bypasses placement entirely; only automatic
|
|
# selection skips it.
|
|
weight: 0.5
|
|
# maxLoad: max live workers on this profile. Omit for unlimited. An explicit 0 (CB-585) caps
|
|
# the profile at zero live members — it is excluded from automatic placement and an explicit
|
|
# `fleet_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
|
|
# is named directly. Negative is refused at config load — there is no sane meaning for it.
|
|
maxLoad: 2
|
|
# subscription: true
|
|
# THE KNOB THAT DECIDES WHO PAYS (CB-539). Default false. When true, this profile's members
|
|
# run on the OPERATOR'S OWN Claude subscription instead of a metered endpoint — every spawn
|
|
# bills your plan and eats your usage limit. Off-subscription is the whole point of this
|
|
# daemon, so treat `true` as a deliberate exception, not a convenience.
|
|
#
|
|
# What changes when it is set (ClaudeCodeLauncher):
|
|
# - no ANTHROPIC_BASE_URL and no ANTHROPIC_AUTH_TOKEN are injected — the member inherits
|
|
# the operator's own Claude Code auth, which is exactly why it bills the plan;
|
|
# - SubscriptionGuard never vets it, because there is no baseUrl to vet;
|
|
# - no token is required, so `tokenEnv` is irrelevant here.
|
|
#
|
|
# MUTUALLY EXCLUSIVE with `baseUrl` — setting both is refused at config load (CB-542). On the
|
|
# subscription path no guard would vet the URL, so allowing both would be a way around the
|
|
# guard rather than a configuration.
|
|
#
|
|
# GOTCHA 1 — it is invisible to the startup secret check. `Fleetd.reportRequiredSecrets`
|
|
# skips subscription profiles on purpose (they need no token), so a boot log that reports
|
|
# every secret as fine says nothing about these profiles.
|
|
#
|
|
# GOTCHA 2 — `maxLoad` is the ONLY throttle you have here. There is no metering, no budget
|
|
# and no refusal on cost; the cap on live members is the single thing standing between a
|
|
# fan-out and your monthly limit. Set it deliberately and keep it small.
|
|
#
|
|
# GOTCHA 3 (fleetd #176, corrected by fleetd #257) — `maxLoad` counts members, never the lead
|
|
# itself. The lead is a live `claude` session on this SAME account (a lead is never moved
|
|
# off-subscription, whatever its own profile says), so it already holds one seat before any
|
|
# member spawns. If a lead's `fleet.leaders.<name>.profile` names THIS profile — or ANY OTHER
|
|
# `subscription: true` profile that shares this one's account (see THE SENTINEL, just below,
|
|
# next to `credentialId:`) — `fleet_list` reports that seat count under `leadSeats`; see
|
|
# `profile:` under THE FLEET below. `free` itself is NEVER reduced by `leadSeats`: `free` means
|
|
# "what the real placement gate (`CompositePeerLauncher#enforceMaxLoad`) will actually grant a
|
|
# fresh `fleet_spawn` right now", and that gate only ever compares live members against
|
|
# `maxLoad` — it has no notion of the lead's own seat. An earlier cut of this feature
|
|
# subtracted `leadSeats` from `free` on the theory it made `free` describe the true ceiling on
|
|
# the account, but no backend seat ceiling shared with the lead has ever actually been
|
|
# measured, and the subtraction just made `free` disagree with the one thing it is supposed to
|
|
# describe — the fleetd #257 fix. `maxLoad: 3` means 3 member slots, full stop; a lead sharing
|
|
# the account is a fact you can see in `leadSeats`, not a reason `free` undercounts spawns that
|
|
# will, in practice, succeed.
|
|
#
|
|
# THE SENTINEL (fleetd #176 stage 2, correcting an inert stage 1 fix): every `subscription:
|
|
# true` profile that leaves `credentialId` unset shares ONE implicit account-wide credential
|
|
# id with every other such profile on this host — because a subscription profile doesn't
|
|
# authenticate with a credential of its own, it authenticates as the operator's own Claude
|
|
# login, and there is exactly one of those. So on a typical host, `opus` (the lead's profile)
|
|
# and `sonnet` (the members' profile) are linked automatically, with NOTHING to set here — that
|
|
# is what makes GOTCHA 3 above work without also writing matching `credentialId:` values on
|
|
# both. This linkage is not just cosmetic: it is the same key `BackendQuarantine`/cool-off use,
|
|
# so a usage-limit hit on `opus` now quarantines `sonnet` too (and vice versa) — correct, since
|
|
# they are one Claude account, but worth knowing before you wonder why an unrelated-looking
|
|
# profile went quarantined.
|
|
#
|
|
# WHEN TO OVERRIDE — set explicit, DIFFERENT `credentialId:` values on two `subscription: true`
|
|
# profiles only when they are genuinely two separate Claude logins on the same host (a real,
|
|
# if unusual, setup). An explicit `credentialId` always wins over the sentinel, so this is the
|
|
# one way to keep two subscription profiles from being treated as one account for lead-seat
|
|
# counting AND for quarantine/cool-off grouping alike.
|
|
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
|
|
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
|
|
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
|
|
# credentialId: shared-openai # opt-in: quarantine together with every other profile sharing this id (CB-578)
|
|
# errorPattern: "503 Service Unavailable" # opt-in: classify a backend outage (fleetd #201/#227) — see the key doc above
|
|
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
|
|
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
|
|
# parityOverlay: [".env"] # the default; add ".envrc" explicitly if you want it copied too (CB-148) — never add .mcp.json or .claude/settings.local.json — see above
|
|
# ideMcpUrl: http://127.0.0.1:29170/index-mcp/streamable-http # opt-in (CB-634): IDE code intelligence, pinned to the worktree
|
|
# ideProjectDir: fleetd # CB-634: module dir the IDE opens + the overlay pins (this repo's pom is in fleetd/)
|
|
# ideOpenCommand: env DISPLAY=:10.0 idea {dir} # CB-634 auto-open: opens {dir} in the IDE at spawn; omit to open by hand
|
|
# autoCompactWindow: 250000 # opt-in: bound member context; claude-code compacts AT this, opencode within it (model limit.context)
|
|
gx11: # a second backend, so `placement: weighted` has a choice
|
|
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
|
|
placement: tab
|
|
workspace: fleetd-workers
|
|
# tabLabel: an optional per-profile override; the fleet template usually covers it
|
|
mcpUrl: http://127.0.0.1:8765/mcp
|
|
argv: ["ccs", "gx11"]
|
|
weight: 0.5
|
|
maxLoad: 2
|
|
# Pin an auto-compact window BELOW the served model's context ceiling. The global
|
|
# ~/.claude/settings.json value is shared by every ccs instance and the primary, so the
|
|
# per-profile override belongs here. Equal to the ceiling means auto-compact never fires
|
|
# before the server rejects the prompt, which kills a worker mid-turn (CB-523).
|
|
env:
|
|
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "280000"
|
|
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
|
|
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
|
|
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
|
|
# The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a
|
|
# `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude.
|
|
#
|
|
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → fleet_send →
|
|
# structured fleet_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
|
|
# and need NO credentials — check `opencode models` for the current free list, since the names
|
|
# change. That also makes the worker off-subscription by construction.
|
|
# opencode-free:
|
|
# kind: opencode
|
|
# model: opencode/north-mini-code-free # `provider/model` selector, injected as `-m`
|
|
# placement: tab
|
|
# workspace: fleetd-workers
|
|
# tabLabel: "opencode: {profile} #{n}"
|
|
# mcpUrl: http://127.0.0.1:8765/mcp
|
|
# argv: ["opencode"]
|
|
#
|
|
# CB-508: point an opencode profile at your OWN OpenAI-compatible endpoint (local vLLM, llama.cpp,
|
|
# LM Studio, TGI…) instead of opencode's gateway. Setting `baseUrl` on a `kind: opencode` profile
|
|
# makes the bridge emit a custom `provider` block into the generated opencode.json — opencode has
|
|
# no ANTHROPIC_BASE_URL seam, so this is how the endpoint is pinned.
|
|
# baseUrl → a bare host:port gets `/v1` appended (where these servers mount the API); a URL that
|
|
# already has a path is used verbatim, so a custom mount point still works.
|
|
# model → MUST be "<provider>/<model>". The provider half names the generated block; the model
|
|
# half must match an id the server reports at /v1/models. One field drives both the
|
|
# declaration and the `-m` flag, so they cannot drift apart. A bare model name with a
|
|
# baseUrl set is rejected at spawn rather than silently using the default gateway.
|
|
# tokenEnv → optional; its value becomes the provider apiKey. Most local servers ignore the key,
|
|
# so a placeholder is used when unset (the AI SDK still requires a non-empty one).
|
|
# NOTE: no `guard` entry is needed even with a baseUrl set. The SubscriptionGuard exists to stop a
|
|
# worker borrowing the primary's Anthropic subscription, and an opencode process has no Anthropic
|
|
# credential path at all.
|
|
# opencode-local:
|
|
# kind: opencode
|
|
# baseUrl: http://127.0.0.1:8000
|
|
# model: local-vllm/deepseek-v4-flash
|
|
# placement: tab
|
|
# workspace: fleetd-workers
|
|
# tabLabel: "opencode: {profile} #{n}"
|
|
# mcpUrl: http://127.0.0.1:8765/mcp
|
|
# argv: ["opencode"]
|
|
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
|
|
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
|
|
#
|
|
# `weighted` IS NOT "cheapest first" — read this before you set weights (CB-589).
|
|
# It is smooth weighted round-robin: it spreads spawns across EVERY profile that has a free slot,
|
|
# in weight ratio. It has no idea which profile costs money. So with local:10 / paid:2 you do not
|
|
# get "use local, overflow to paid" — you get roughly one spawn in six going to the paid profile
|
|
# while the local box still has a free slot.
|
|
#
|
|
# There is a sharper second effect. The policy's running score map lives for the daemon's whole
|
|
# life. While a profile is at maxLoad it is filtered out and its score FREEZES, so the paid
|
|
# profiles keep accumulating against it. When the local slot frees up it returns with a stale
|
|
# score and can LOSE the next pick — a paid spawn while the free box sits idle.
|
|
#
|
|
# Until a real cost-first policy exists, the workaround is to make the ratio decisive rather than
|
|
# proportional: give the free profile a weight so large that it wins every pick it is eligible
|
|
# for, and paid profiles only ever take genuine overflow. On this host that is local weight 100
|
|
# against paid weights of ~1.
|
|
#
|
|
# The gotcha with that workaround: it expresses a PREFERENCE ORDER through a RATIO knob. Add a
|
|
# future profile at weight 150 and it silently outranks the free box, with nothing to warn you.
|
|
# Re-check the weights whenever you add a profile.
|
|
placement: weighted
|
|
|
|
# How long a credential sits out after a BACKEND_EXHAUSTED classification (CB-578 stage B), in
|
|
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
|
|
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
|
|
# 1800 (30 minutes) when omitted or non-positive.
|
|
#
|
|
# fleetd #466: this is now only the BASE of an escalating backoff, not a flat retry rate. A
|
|
# credential quarantined again within one base cooldown of the previous quarantine ending (still
|
|
# reporting exhausted — e.g. a weekly subscription limit that hasn't reset) backs off further:
|
|
# cooldown doubles each such time, capped at 12x this value (~6 hours at the 1800s default). A
|
|
# quarantine that starts after a base-cooldown's worth of quiet resets back to this value. Not
|
|
# configurable per se — the multiplier and ceiling are constants in BackendQuarantine, not new
|
|
# YAML keys; see its class doc for the exact formula and why there is no automatic probe to clear
|
|
# it early (the operator's own design constraint — a probe spends the quota it's measuring).
|
|
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
|
|
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
|
|
# restart. Editing this needs a daemon restart to take effect.
|
|
#
|
|
# This does NOT govern the fleetd #201 / #227 backend-error cool-off documented under errorPattern
|
|
# above — that mechanism is a separate, shorter-lived, NOT-configurable policy (threshold 2 distinct
|
|
# targets, 60-second window, 60-second cool-off), on purpose: it exists to survive a brief transient
|
|
# fault, not to replace this 30-minute exhaustion quarantine. Do not conflate the two when reading
|
|
# fleet_list/fleet_profiles — coolingOffForSeconds and quarantinedForSeconds are independent facts.
|
|
# quarantineCooldownSeconds: 1800
|
|
|
|
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
|
|
# upgraded fleetd keeps the old behaviour: the file is read once at boot and never again.
|
|
# enabled → turn the watch on. fleetd checks the file's modified time on a timer and
|
|
# reloads when it moves.
|
|
# intervalSeconds → how often to check (default 10). One `stat` per tick, so this is cheap.
|
|
#
|
|
# Not every key can move under a running daemon, and the difference is about what already exists
|
|
# when the reload happens — not about how important the key is:
|
|
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
|
|
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad
|
|
# / credentialId / exhaustedPattern. Those are hot because the placement policy (and,
|
|
# for credentialId, the CB-578 stage B quarantine check; for exhaustedPattern, fleetd
|
|
# #446's LiveExhaustedPatterns) reads them through a supplier — being config is not by
|
|
# itself enough to make a key hot.
|
|
# EXCEPT `fleet.leaders`: Fleetd.main reads it once at startup to build the lead tab
|
|
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
|
|
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
|
|
# with nothing in the deferred list — but has NO effect until you restart. Treat it
|
|
# as deferred in practice, even though today's reload output does not say so.
|
|
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
|
|
# until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`,
|
|
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, `quarantineCooldownSeconds` (CB-578
|
|
# stage B — baked once into the quarantine tracker built at startup), ADDING or
|
|
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
|
|
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
|
|
# configDir, mcpUrl, tabLabel, errorPattern (fleetd #201 / #227 — compiled once into
|
|
# a startup pattern map; exhaustedPattern used to be compiled the same way until
|
|
# fleetd #446 made it hot — see above). The launcher takes a copy of `profiles:` at
|
|
# startup and resolves every spawn out of that copy, so those never
|
|
# reach a launch until you restart. The reload logs them by name rather than
|
|
# pretending they applied.
|
|
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
|
|
# bound, the broker connection is open, and the auth mode decides who may reach the
|
|
# port that is already listening.
|
|
#
|
|
# A changed COLD key refuses the WHOLE reload — not the hot half applied and the cold half warned
|
|
# about. A half-applied reload would leave the daemon matching no file on disk, which is the worst
|
|
# thing a reload can do to an operator debugging one. A file that fails to parse or fails a startup
|
|
# validator is refused the same way, and the running config stays live.
|
|
# configReload:
|
|
# enabled: true
|
|
# intervalSeconds: 10
|
|
|
|
# THE FLEET (CB-557) — who the daemon may run, and under which role. This one block replaced four
|
|
# older keys: `leaders:`, `members:`, `leadScan:` and `defaultProfile:`.
|
|
#
|
|
# A member is anything a lead spawns, and every member has two INDEPENDENT attributes:
|
|
# role — which contract: architect, dev or reviewer. It picks the launch charter, the role
|
|
# file, the playbook skill and the authz row.
|
|
# profile — which backend: one of the `profiles:` keys above (model, CLI adapter, cost).
|
|
# They vary on their own. A reviewer may run on the same profile as the dev whose diff it reads,
|
|
# which is why the two cannot be one field.
|
|
#
|
|
# The ROLE IS THE CONTAINING KEY, not a `role:` field. That is not only tidier: a misspelled role
|
|
# used to parse into a member with no contract at all, while a misspelled pool name here simply
|
|
# declares nothing.
|
|
#
|
|
# Each pool lists the profiles that role MAY run on — these are pools, not identities. That is also
|
|
# what replaced `defaultProfile:`: an unqualified spawn names a role, and that role's pool supplies
|
|
# the candidates, in definition order. A dev and a reviewer staying anonymous is exactly compatible
|
|
# with being listed here; the entry key just names the entry.
|
|
fleet:
|
|
# Optional launch-charter text, keyed only by the singular role wire names: architect, dev,
|
|
# reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put secrets
|
|
# here: a later launch step writes this text to a world-readable temp file, and ${ENV} interpolation
|
|
# is deliberately not supported.
|
|
charters:
|
|
architect: |-
|
|
You are an architect in this fleet. You refine work before anyone builds it:
|
|
scope, acceptance criteria, risks, and a unit split. You read the repo and
|
|
write analysis. You never commit production code and never open a PR.
|
|
A design task is worked by two architects. Design alone first, then exchange
|
|
and say plainly where you disagree. Do not concede just to agree.
|
|
dev: |-
|
|
You implement the one unit you were given, and nothing else. You test it,
|
|
commit it, and open your own pull request. You never merge.
|
|
reviewer: |-
|
|
You review the diff you were given. You report bugs, risks and missing tests.
|
|
You do not change code.
|
|
|
|
# Optional. Template for a member tab's label; {role}, {profile}, {model} and {n} are substituted.
|
|
# {n} counts per role+profile, so `dev: sonnet #2` really is the second sonnet dev. Because {role}
|
|
# comes from a closed enum, a generated label can never begin with a lead's tabPrefix.
|
|
# tabLabel: "{role}: {profile} #{n}"
|
|
|
|
# Panes that orchestrate rather than are orchestrated. A lead may now be CREATED as well as
|
|
# recognised: give it a `profile:` and the daemon launches the shortfall when fewer than
|
|
# `instances` are live. Omit `profile:` and it is recognise-only, as before.
|
|
#
|
|
# `profile:` has a SECOND job as of fleetd #176, even for a recognise-only lead you never want
|
|
# auto-launched: it is also how fleetd learns which account this lead's own session shares. A
|
|
# `subscription: true` profile bills the operator's Claude account, and the lead itself is always
|
|
# a live `claude` session on that same account — `maxLoad` never counted that seat. If a lead
|
|
# entry here names a profile that shares a worker profile's account, `fleet_list` reports the
|
|
# lead's live seat(s) on that worker profile under `leadSeats` — informational only, as of fleetd
|
|
# #257 it is NEVER subtracted from `free` (see GOTCHA 3, next to `maxLoad:`, in THE WORKERS above,
|
|
# for why). "Shares the account" is decided by matching `effectiveCredentialId()`, which (fleetd
|
|
# #176 stage 2 — see THE SENTINEL, next to `credentialId:`, in THE WORKERS above) means: an
|
|
# explicit, matching `credentialId:` on both, OR — the common case, needing NO extra config — both
|
|
# being `subscription: true` with `credentialId` left unset, since those all share one implicit
|
|
# account-wide id. A lead on `opus` and workers on `sonnet` link automatically this way; they do
|
|
# NOT need the same profile name. Setting `profile:` on an already-running, recognise-only lead is
|
|
# safe — the daemon only launches the SHORTFALL below `instances`, so naming a profile here does
|
|
# not, by itself, start anything. Omit it and fleetd has no way to derive the sharing — there is
|
|
# no other reliable signal on the daemon's side — so that lead's seat never appears in `leadSeats`.
|
|
#
|
|
# `tab:` (CB-579) is REQUIRED and is the only field identity depends on — the exact label of the
|
|
# tab hosting the lead, matched case-insensitively. Label the tab yourself and put that same
|
|
# string here, and the pane is recognised on the next rescan. Reopen the tab later, or the session
|
|
# inside it restarts — the terminal id changes; the tab, and its label, do not, so no config edit
|
|
# follows a restart.
|
|
#
|
|
# A lead the daemon launches is labelled BY the daemon with this same `tab:` value, so it is found
|
|
# by the same scan. A lead counts as live only when herdr also reports a running agent in that
|
|
# tab — a label left behind by a session that died does not block the relaunch, and a tab that is
|
|
# gone entirely drops out of the next scan rather than being remembered forever.
|
|
#
|
|
# An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with
|
|
# the session lifecycle (the idle reaper would kill your orchestrator), and stays on the
|
|
# subscription — ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.
|
|
#
|
|
# GET THE `tab:` VALUE RIGHT. A pane that does not match any configured `tab:` (a typo, a renamed
|
|
# tab, a pane no entry names at all) is not recognised as a lead — it resolves as an ordinary
|
|
# WORKER instead, silently, and every orchestration call it makes (spawn/stop/send/drain) is
|
|
# refused. There is no error at startup for this: an unmatched pane is simply not a lead. If your
|
|
# primary suddenly can't spawn or send, check this section first.
|
|
# leaders:
|
|
# opus-5.0:
|
|
# profile: opus # omit to never create this lead, only recognise it
|
|
# instances: 1 # desired live count; only the shortfall is launched. 0 = off
|
|
# tab: "lead: opus-5.0" # REQUIRED — the exact tab label this lead lives in
|
|
# tabPrefix: "lead:" # only used to guard against a worker tabLabel colliding with
|
|
# # this convention at startup; plays no part in matching a lead
|
|
# scanIntervalSeconds: 10 # rescan cadence, and the worst case before a new tab is seen
|
|
# workspace: leads # where a launched lead's tab is created (default "leads").
|
|
# # MUST NOT be a member workspace — those are excluded from the
|
|
# # scan, so a lead placed in one is never found again.
|
|
# cwd: /path/to/repo # the launched lead's working directory (default: fleetd's own)
|
|
# kind: claude # descriptive; reported by fleet_whoami
|
|
# gpt-sol-5.6:
|
|
# tab: "lead: gpt-sol-5.6"
|
|
# kind: opencode
|
|
# model: openai/gpt-5.6-terra
|
|
|
|
# architects:
|
|
# architect-1:
|
|
# profile: opus # a strong model, on the operator's subscription
|
|
# architect-2:
|
|
# profile: sol # a different vendor on purpose — two architects that share a
|
|
# # model share its blind spots
|
|
developers:
|
|
gx10:
|
|
profile: gx10
|
|
# reviewers:
|
|
# gx10:
|
|
# profile: gx10 # the same backend may serve two roles; that is the point
|
|
|
|
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
|
|
# must carry none. Every profile above must have its host listed here.
|
|
guard:
|
|
offSubscriptionHosts:
|
|
- gx00.gw
|
|
- gx01.gw
|
|
|
|
# Member credential policy (CB-596, gitea issue #82). A herdr pane runs a LOGIN shell, and that
|
|
# shell re-sources the operator's own secret store — so a spawned member inherits every credential
|
|
# the operator's shell holds, not just the ones fleetd means to give it. Measured on this host:
|
|
# 31 credential names, all set, with only ONE (GITEA_ACCESS_TOKEN) blocked before this — and that
|
|
# block was a single name hardcoded in HerdrPeerLauncher.java, not driven by this file. This block
|
|
# replaces that hardcoded shadow with a config-driven list of names.
|
|
#
|
|
# ROUND-2 CORRECTION, measured live: the pane-creation env overlay below (applied at tab.create /
|
|
# pane.split, BEFORE the pane's login shell runs) does NOT survive that login shell for any name
|
|
# secrets.sh actually exports — the shell re-exports it afterwards and overwrites the sentinel.
|
|
# Proof: GITEA_ACCESS_TOKEN comes back blocked only because secrets.sh itself carries a guarded
|
|
# export (`[ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...`) — that guard, not this
|
|
# file, is what wins. No other name in `known` below has a matching guard in secrets.sh yet (1
|
|
# guard measured against 33 export lines there). So today this block's overlay is REAL protection
|
|
# only for a name secrets.sh does not export, or a peer kind whose pane never runs a login shell —
|
|
# for everything secrets.sh exports and guards, the guard in secrets.sh (out of scope for this
|
|
# ticket) is what actually blocks it, not this list. An exec-time fix (winning after the login
|
|
# shell finishes, before the agent process starts) was attempted and found to have no seam in the
|
|
# current herdr protocol — AgentControl.start takes a fixed `kind` (herdr resolves the executable)
|
|
# plus trailing CLI args for that binary, not an arbitrary argv or an env map; only tab.create /
|
|
# pane.split accept `env`, and that is this same pane-creation overlay. See gitea #82 for the open
|
|
# design question this leaves.
|
|
#
|
|
# DENY-BY-DEFAULT, NOT A DENY-LIST. A deny-list (name the bad ones, let everything else through) is
|
|
# silently wrong the moment the operator's store gains a new secret — nothing would ever report it.
|
|
# Deny-by-default inverts that: `known` bounds the blast radius to names actually enumerated below,
|
|
# and EVERY one of them is blocked UNLESS it is also in `allow`. Omitting this block entirely (the
|
|
# shipped default) blocks NOTHING — unlike most optional blocks in this file, absence here is a real
|
|
# gap, not a safe "feature off". A name that is neither `known` nor `allow`-ed is not silently let
|
|
# through either: the daemon logs a WARN naming any credential-shaped env var it finds on neither
|
|
# list (never its value), so a secret added to the store later does not go unnoticed forever.
|
|
#
|
|
# policy → "deny-by-default" (the default; also accepted spelled "deny-list") overlays each
|
|
# known-but-not-allowed name BEFORE the pane's login shell runs — real protection only
|
|
# where that shell does not re-export the name (see ROUND-2 CORRECTION above). An
|
|
# unrecognized value refuses to start, naming it.
|
|
# policy → "allow-list" (CB-633) moves the control to a per-spawn ZDOTDIR directory the daemon
|
|
# generates and passes through tab.create's env map. Each generated startup file sources
|
|
# its ~/ counterpart FIRST and then runs the scrub, so the scrub happens after the
|
|
# operator's whole chain and no sourced file can undo it.
|
|
# The scrub is sourced from BOTH the generated .zshrc and the generated .zlogin, because
|
|
# herdr does not open the same kind of shell everywhere: macOS panes run a LOGIN zsh (so
|
|
# .zlogin runs), Linux panes run a plain interactive zsh (so .zlogin never runs at all).
|
|
# A scrub in .zlogin alone would be a control that silently does nothing on Linux.
|
|
# The allow-list is DERIVED, never typed:
|
|
# every profile's tokenEnv/gitTokenEnv/gitHostEnv values and env-map keys, plus an
|
|
# infrastructure set (PATH HOME SHELL TERM LANG LC_* TMPDIR USER LOGNAME PWD SHLVL EDITOR
|
|
# PAGER JAVA_HOME XDG_* ZDOTDIR), plus whatever keys this spawn's own env overlay carries.
|
|
# Adding a profile can therefore only widen the list, never break another spawn's scrub.
|
|
# Under this policy `known`/`allow` below become REPORTING ONLY — they feed the gap WARN,
|
|
# they are no longer a control. If the member's login shell is NOT zsh, the daemon logs a
|
|
# loud WARN saying protection is off and falls back to deny-by-default's overlay.
|
|
# Each pane writes a scrub-report.txt naming how many variables it kept of how many it
|
|
# saw; the daemon logs that "allowed N of M" line when the pane stops. If the report is
|
|
# MISSING the daemon logs a WARN instead — the scrub cannot then be confirmed to have
|
|
# run, and a silently-dead control is exactly what this policy exists to prevent.
|
|
# allow → credential names a member legitimately needs. Under deny-by-default, left OUT of the
|
|
# pane's env overlay entirely, so the value the pane's own (login) shell exports passes
|
|
# through untouched. Under allow-list: reporting only.
|
|
# known → every credential name the operator's store is known to export. Under deny-by-default,
|
|
# every name here NOT also in `allow` is overlaid with a non-secret sentinel value before
|
|
# the pane's login shell runs — real protection only for names that shell does not itself
|
|
# re-export (see the ROUND-2 CORRECTION note above). Under allow-list: reporting only.
|
|
# sshAgentEnv → whether SSH_AUTH_SOCK may pass through under allow-list ("inherit") or is omitted
|
|
# from the member environment ("omit", the default). Omitting it only omits the
|
|
# inherited ssh-agent path. It discourages automatic use of the operator's agent.
|
|
# It does not deny same-user access to that socket. It also does not block SSH keys that
|
|
# are readable on disk. Git over SSH may still work from inside a member. Keep the block:
|
|
# it is correct and costs nothing, but it is not a control. A member runs as the same OS
|
|
# user as the lead. Inside one uid, ordinary Unix permissions provide no meaningful
|
|
# confidentiality boundary. A real boundary needs a different OS user or OS-level
|
|
# confinement, such as a container or VM. That is the open question in fleetd #184.
|
|
#
|
|
# Still do not set this to "inherit" casually. SSH_AUTH_SOCK is a live handle to YOUR
|
|
# ssh-agent, so a member holding it can sign with EVERY key the agent holds. It sits in
|
|
# no secret file and looks like no credential, which is why it slipped past three
|
|
# earlier tickets (gitea #110). Blocking it does not contain a member, but allowing it
|
|
# hands one a signing capability for no gain — the block costs nothing, so keep it.
|
|
#
|
|
# Both halves of this are measured, not argued. 2026-08-28: a member with
|
|
# SSH_AUTH_SOCK blanked pushed to the forge over SSH successfully, because `ssh -G`
|
|
# resolves an IdentityFile outside ~/.ssh that is readable and has no passphrase. An
|
|
# earlier version of this comment claimed blocking the socket BREAKS git over SSH. It
|
|
# does not. That claim came from looking only in ~/.ssh, which holds nothing but four
|
|
# `Include` lines — looking in one place and concluding about the whole host.
|
|
#
|
|
# HOT-RELOADABLE the same way `fleet:` is (CB-559): read fresh on every spawn, so editing this list
|
|
# and reloading config (or restarting) changes what the NEXT spawn inherits; already-running members
|
|
# are unaffected either way.
|
|
# memberCredentials:
|
|
# policy: deny-by-default # or "deny-list", or "allow-list" (CB-633) — see above
|
|
# sshAgentEnv: omit # allow-list only; see the sshAgentEnv note above
|
|
# allow:
|
|
# - AI_GATEWAY_TOKEN # named in a profile's tokenEnv (local/gx) — a member reaching the
|
|
# # gateway is by design, not a leak
|
|
# - WORKER_GITEA_TOKEN # the repo-scoped forge token a member needs to open its own PR (CB-302)
|
|
# - CONTEXT7_TOKEN # already decided as allowed by CB-593
|
|
# - GITEA_HOST # not a credential — a hostname, paired with the forge token above
|
|
# known:
|
|
# - AI_GATEWAY_TOKEN
|
|
# - BESZEL_ADMIN_EMAIL
|
|
# - BESZEL_ADMIN_PASSWORD
|
|
# - BESZEL_HUB_URL
|
|
# - BESZEL_KEY
|
|
# - BESZEL_UNIVERSAL_TOKEN
|
|
# - BRAIN_MCP_TOKEN
|
|
# - CF_ACCOUNT_ID
|
|
# - CF_API_TOKEN
|
|
# - CF_USER_TOKEN
|
|
# - CONFLUENCE_API_TOKEN
|
|
# - CONFLUENCE_USERNAME
|
|
# - CONTEXT7_TOKEN
|
|
# - GITEA_ACCESS_TOKEN
|
|
# - GITEA_HOST
|
|
# - GITLAB_OAUTH_CLIENT_SECRET
|
|
# - GITLAB_PERSONAL_ACCESS_TOKEN
|
|
# - GRAFANA_ADMIN_PASSWORD
|
|
# - GRAFANA_ADMIN_USER
|
|
# - HASS_TOKEN
|
|
# - HW_PASSWORD
|
|
# - HW_USER
|
|
# - LTMS_API_KEY
|
|
# - MEMORY_MCP_TOKEN
|
|
# - METRICS_PUSH_TOKEN
|
|
# - OPENCODE_AUTOMODE_MODEL
|
|
# - TELEGRAM_BOT_TOKEN
|
|
# - TELEGRAM_CHAT_ID
|
|
# - TS_API_KEY
|
|
# - TS_AUTHKEY
|
|
# - WORKER_GITEA_TOKEN
|
|
|
|
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
|
|
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
|
|
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
|
|
# unknown keys are ignored, so a snake_case key would be silently dropped (default kept).
|
|
# spawnReadyTimeoutMs: 20000
|
|
# spawnReadyPollMs: 300
|
|
|
|
# Worktree provisioning root (CB-301-ext). Where per-worker git worktrees are checked out so
|
|
# each worker owns an isolated branch instead of sharing the primary's tree. Omit to default
|
|
# to a sibling directory of the repo root.
|
|
# worktreeRoot: /Users/me/src/.bridged-worktrees
|
|
|
|
# Worktree group sharing (fleetd #185 stage 3). OPTIONAL, off by default. Names an OS group
|
|
# that a provisioned worktree's repo is made group-writable for (git config
|
|
# core.sharedRepository group, plus a one-time chgrp/chmod/setgid fix-up), so a member spawned
|
|
# under a DIFFERENT OS user (see memberHerdrSocket) can write its own worktree, its
|
|
# per-worktree git metadata, and its own commit objects — without it, every file GitWorktrees
|
|
# creates is owned by fleetd's own uid and unwritable by another user.
|
|
# CAUTION: this isolates credentials, not the repository — a member in the group can still
|
|
# write the operator's git objects and refs in the shared repo. The operator running fleetd
|
|
# must already be a member of the named group, or every provisioning spawn fails loudly.
|
|
#
|
|
# fleetd #213: this is also the ONE group the memberCredentials.policy: allow-list ZDOTDIR scrub
|
|
# reuses when memberHerdrSocket is set — deliberately not a second config key. Under
|
|
# memberHerdrSocket, the scrub directory is generated under worktreeRoot (never java.io.tmpdir,
|
|
# which the member OS user cannot reach) and shared read-only with this group. If worktreeGroup
|
|
# is unset while memberHerdrSocket is set, the scrub cannot be guaranteed reachable by the member,
|
|
# so fleetd falls back to the weaker CB-596 sentinel overlay instead (a WARN names the gap).
|
|
# worktreeGroup: fleet-workers
|
|
|
|
# fleetd #362: a directory of skill folders (each a subdirectory holding a SKILL.md, the same
|
|
# shape as this repo's own .claude/skills/) copied into every PROVISIONED worktree's
|
|
# .claude/skills/, so a member spawned against ANY repo — not only one that already ships its own
|
|
# copy — can load a bridge skill (e.g. implementer). Unset (the default): no worktree is touched
|
|
# beyond today's behaviour. A skill folder the target repo already carries under
|
|
# .claude/skills/<name> is never overwritten — the repo's own copy always wins. Best-effort like
|
|
# worktreeGroup above: a missing/unreadable directory here is logged and skipped, never a failed
|
|
# spawn. Every non-hidden subdirectory of this directory is copied wholesale, with no per-file
|
|
# allowlist — don't park scratch files or drafts alongside the real skill folders, they will be
|
|
# copied into every provisioned worktree too.
|
|
#
|
|
# fleetd #393: which member KINDS actually consume this once it is copied. kind: claude-code —
|
|
# the Claude Code CLI discovers .claude/skills/ on its own; nothing else is needed. kind: opencode
|
|
# — opencode has no such discovery, so OpenCodeLauncher reads whatever landed under
|
|
# .claude/skills/ and appends each seeded skill's SKILL.md to the generated instructions[] file
|
|
# (opencode's only channel for static guidance text; unlike Claude Code's Skill tool, the content
|
|
# is always part of the system prompt, not loaded on demand). Both kinds are covered as of #393 —
|
|
# earlier builds copied the files for every kind but only claude-code could read them, and the
|
|
# seeding log said "N of M" regardless. Check the per-spawn launcher log (not just the seeding
|
|
# log) to see what a given member actually got.
|
|
# memberSkills: /path/to/fleetd/checkout/.claude/skills
|
|
|
|
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
|
|
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
|
|
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
|
|
# contextCap → force-release a session after this many delegated turns
|
|
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
|
|
# clearAfterTurn → whether a reusable worker discards its conversation context after every
|
|
# completed delegated turn (default false). Works for claude-code workers
|
|
# only — any other peer kind (e.g. opencode) logs "context reset is
|
|
# unsupported for peer kind …" once and the reset is a no-op.
|
|
# lifecycle:
|
|
# idleTtlSeconds: 300
|
|
# contextCap: 10
|
|
# drainTimeoutSeconds: 5
|
|
# clearAfterTurn: false
|
|
|
|
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
|
|
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
|
|
# Set a broker uri to swap in the AMQP-backed inbox: worker replies with no open send are held
|
|
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
|
|
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
|
|
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
|
|
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
|
|
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
|
|
# uriEnv → CB-151: name of a host env var holding the AMQP URI, preferred over `uri` (wins
|
|
# whenever set). The URI carries `user:pass@` inline, so naming a variable keeps the
|
|
# password out of fleetd.yaml — same pattern as auth.tokenEnv/Profile.tokenEnv. A
|
|
# uriEnv that resolves to an unset or blank variable is treated as NOT configured and
|
|
# the daemon falls back to the in-memory inbox, warning loudly.
|
|
# prefetch → CB-527: consumer basicQos, capping how many unacked messages the inbox holds
|
|
# in-heap per owned target (the rest sits on the broker's durable queue instead of
|
|
# growing the JVM heap). Default 32 when omitted.
|
|
# broker:
|
|
# uriEnv: LAVINMQ_URI
|
|
# prefetch: 32
|
|
|
|
# Shared cross-host LEADER coordination broker. OMIT this block to leave lead-to-lead messaging
|
|
# off entirely (config-only in this ticket — nothing here wires it into a live LeadMailbox yet).
|
|
# This is a SEPARATE AMQP vhost from `broker:` above: member/worker inboxes always stay on the
|
|
# per-fleet `broker:` vhost, and this vhost carries only leader-to-leader traffic, so two fleets
|
|
# whose members must never see each other can still share one coordination vhost for their leads.
|
|
# uriEnv → name of a host env var holding the coordination AMQP URI, same convention as
|
|
# broker.uriEnv (keeps the credential out of fleetd.yaml). Wins over `uri` when set.
|
|
# selfId → this daemon's own lead coord-id — the name its mailbox is owned under
|
|
# (lead.<selfId>.inbox), e.g. "mac-opus" or "fleet01-lead". Must be globally unique
|
|
# across every daemon sharing this vhost.
|
|
# prefetch → consumer basicQos, capping how many unacked messages the mailbox holds in-heap.
|
|
# Default 32 when omitted.
|
|
# peers → fleetd #361: the coord-ids of the OTHER daemons on this vhost, declared by the
|
|
# operator (the daemon never guesses). fleet_list reports each one's live reachability
|
|
# (a passive queue check, never a presence protocol) alongside this daemon's own
|
|
# mailbox state. Omit, or leave empty, for a daemon with no known peers yet — an
|
|
# undeclared peer can still reach you and be reached by fleet_send, it just will not
|
|
# show up as a row in fleet_list.
|
|
# coordinator:
|
|
# uriEnv: LEAD_COORD_URI
|
|
# selfId: mac-opus
|
|
# prefetch: 32
|
|
# peers: [fleet01-lead]
|
|
|
|
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open fleet_send,
|
|
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
|
|
# pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop
|
|
# stops as soon as the primary's inbox is empty.
|
|
# terminal → pin the primary's herdr terminal id. Omit to learn it from the connection on
|
|
# the first orchestration-side MCP call (the normal case). An off-host or
|
|
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
|
|
# degrades to pull; the reply is still never lost.
|
|
#
|
|
# REQUIRED (CB-522) if the primary itself runs inside a herdr pane. Caller
|
|
# identity resolves a loopback PID to its herdr pane, and PaneLocator scans
|
|
# EVERY pane — not just fleetd-spawned ones — so such a primary is otherwise
|
|
# classified as a WORKER and refused SPAWN/SEND/STOP. That failure is
|
|
# self-locking: the learned terminal is populated by the very orchestration
|
|
# calls being refused, so only this pinned value can break the cycle. Read the
|
|
# id off fleet_whoami (it reports the current terminal even while
|
|
# misclassified) and re-pin whenever the primary moves panes.
|
|
# pushReminders → max nudges before giving up (default 5)
|
|
# pushBackoffMs → delay between nudges in ms (default 15000)
|
|
# primary:
|
|
# terminal: term_65619bd6174568
|
|
# pushReminders: 5
|
|
# pushBackoffMs: 15000
|
|
|
|
# Central allow-list of models any profiles: entry may name. Nothing checked a profile's model:
|
|
# value before this block existed — it was a free-form string handed straight to the backend
|
|
# adapter, and a withdrawn or misspelled name failed silently instead of at config load (opencode
|
|
# falls back to a default model rather than erroring on an unknown -m).
|
|
#
|
|
# Absent, or present with an empty allow:, is OFF: no profile's model: is checked, exactly like
|
|
# before this block existed. fleetd.yaml is gitignored on every host, so an upgrade must not force
|
|
# every operator to enumerate their models before the daemon will start.
|
|
#
|
|
# The list is the authority; profiles: is checked against it, never the reverse — adding or
|
|
# editing a profiles: entry cannot, by itself, widen what is permitted here.
|
|
#
|
|
# Enforcement is at CONFIG LOAD only (a bad model: fails the daemon at startup, naming both the
|
|
# model and the profile). There is no spawn-time enforcement, no runtime on/off switch, and no
|
|
# interaction with BackendQuarantine — those are separate, later units.
|
|
#
|
|
# allow → the permitted models. Each entry is its own block (not a bare string) so a later unit
|
|
# can add an on/off state or a load limit per model without changing this shape.
|
|
# model → the model id exactly as a profiles: entry's model: field would write it. One flat,
|
|
# opaque-string namespace: a bare Claude id (claude-sonnet-5) and an opencode
|
|
# provider-prefixed id (openai/gpt-5.6-terra) both fit here unchanged — the check is a
|
|
# plain string match, never a parse of the provider prefix or a branch on kind:.
|
|
# models:
|
|
# allow:
|
|
# - model: claude-sonnet-5
|
|
# - model: claude-opus-5
|
|
# - model: openai/gpt-5.6-terra
|
|
# - model: amazon.nova-pro-v1:0
|