Table of Contents
- 11 — Features
- Index
- Ask the bridge who you are
- Primary inside a herdr pane
- Weighted worker placement
- Run a worker on the subscription
- Advisory architect slots
- Keep a worker conversation
- Give workers a toolchain
- Isolated worktree per worker
- Worker tool-surface isolation
- Worktree-hostile config isolation
- No credentialed remote URL reaches a worktree
- Worker opens its own PR
- Session lifecycle caps
- Durable reply inbox
- Draining an inbox is lead-only, on both entry paths
- Broker password out of the config
- An unreachable broker does not stop the daemon
- Reject overlapping rendezvous
- Pin an opencode endpoint
- Onboard a project with the plugin
- Port a workspace to OpenCode
- More than one lead
- Find leads by tab name
- Leads talk to each other
- Leads are visible in fleet_list
- A lead can be delivered to
- Reply nudges follow the delegating lead
- Unknown config keys are named
- One fleet block, role as the key
- Role pools decide the backend
- Tab labels name the role
- Launch a lead when none is live
- Re-read the config without a restart
- Set a role's launch charter in config
- Know when a completion fallback is partial
- Reject a profile name as a send target
- See free fleet capacity
- Answer a question on an async delegation
- Learn why a delegation died
- Watch the fleet's health
- Keep a worktree that still holds work
- Fail a ticket when its member dies
- Tell a usage-limit refusal from a real reply
- Stop spawning onto an exhausted account
- See which charter a member got
- Redeploy the daemon safely
- Get told when an async delegation finishes
- Get told when a worker is waiting on your answer
- Members cannot use the operator's admin forge token
- Supervise the daemon without breaking the fleet
- See at startup which secrets the daemon actually got
- Bound the AMQP backlog, and know a reply was really published
- Three config keys that changed meaning — read this before upgrading
- Stop an explicit spawn from busting the cap
- Nudge an idle lead back to work
- Prove which charter a member got, without logging the prose
- Tell "busy" from "refusing" in the capacity view
- Resume a member onto its previous conversation
- opencode members run with --auto and forced auto-compaction
- Silent member failures now say why
- Sessions are one-shot — recycle() is gone
- A failed ticket names the conversation, not just the files
- weighted placement is not "cheapest first"
- Which inherited credentials a member may keep
- A member keeps only the credentials you name
- /metrics and /healthz
- Bearer-token auth and the non-loopback-bind fail-fast
- The per-session authz table and the audit log
- Multi-profile routing and kind: adapter selection
- Supervise the daemon on Linux (systemd)
- Every config value is checked against its valid set
- Old work-in-progress snapshots clean themselves up
- Set a member's auto-compact window
- Leads on different hosts talk over a shared broker
- A lead can read its own held peer mail
- Fleet health can finally reach its fault states
- /fleets-status — every fleet that shares one broker
- memberHerdrSocket — run members on a second herdr daemon
- worktreeGroup — let a member under another OS user write its worktree
- opencode session ids come from opencode.db
- One worker reply settles exactly one ticket
- Late-resolved member ids
- GET /members reports its rows under members
- The credential gap detector admits when it cannot see
- Backfill status
- Role agent definitions — how a member learns its role (CB-617)
- Found while cataloguing, not by looking for bugs
- A launch command that cannot fit is refused
- A failed spawn shows you the pane
- Every claude-code member is resumable
- An exhausted backend is quarantined even from a chrome-only pane
- The credential scrub follows the member's own user
- An opencode member's config follows the member's own user
- The model opencode actually ran is read back and checked
- Every file a member must read follows the member's own user
- A role fleetd cannot bind is refused, not quietly downgraded
- A backend that starts erroring cools off, instead of swallowing the next spawn
- A claude-code member no longer blocks forever on the workspace-trust dialog
- A worktree tells you which of its config files are stubs
- The credential probe asks the daemon what the policy is
- An allow-list policy refuses a spawn it cannot enforce
- What a member can actually reach — the honest boundary
- The ssh-agent setting is named after what it does — sshAgentEnv: omit
- fleetd states the member trust model at startup
- leadSeats reports the seat the lead itself holds
- The REST face — the operator's way in when MCP is not there
- A dead member's seat comes back
- A member's teardown no longer leaks a worktree or a branch
- A chained second question no longer kills its own ticket
- A ticket is not stranded when a dead member's question lapses
- The REST face reports what the MCP face reports
- A released reply goes back to the broker instead of being dropped
- A failed spawn no longer leaves a live pane behind
- A reply with no content is refused instead of resolving the waiter
- The member REST routes name a herdr failure instead of returning a bare 500
- One definition of loopback, so a worker cannot become the lead
- The idle reaper no longer stops a worker that just got a turn
- A wedged post-turn phase releases itself, like the turn phase already did
- A worker's real reply after a lapsed fleet_ask completes its ticket
- Shutdown refuses new spawns and sweeps up stragglers
- A failed git worktree add cleans up what it half-created
- fixed placement honours the failover loop's unreachable set
- A caller whose identity cannot be resolved is refused, not treated as the primary
- The check that authorises deleting a worktree is taken after the worker stops
- A reply that arrives while a target is being released is requeued, not stranded
- The reload classifier proves its own coverage instead of claiming it
- An answered fleet_ask no longer fails the lead's own call
- A reload now says a restart is needed for primary: and configReload:
- A reload can now say "half of this applied"
- Every top-level config key must be triaged before it ships
- An exception thrown after a ticket resolves now reaches the log
- A worker's real reply completes its async ticket more often
- A reload now reports a changed fleet.leaders as needing a restart
- Being listed as cold or split now proves a comparison exists
- A send that times out while still queued now cancels its message
- A member's prose about an error no longer records a credential outage
- Every distinct unprotected credential name gets its own warning
- The deferred key set now proves its own reporting coverage too
- Teardown now closes the tab a member really sits in
- A failed cleanup no longer strands the tickets behind it
- A member's prose about a usage limit no longer quarantines a credential
- The last window where a fleet_ask timeout stranded its ticket is closed
- fleet_list reports lead coordination state instead of guessing at it
- Bridge skills are seeded into every provisioned worktree
- Dead lead tabs are cleaned up, and a live one is never closed
- The shipped systemd units no longer disable the daemon they start
- The host no longer idle-sleeps while a member is working
- A reply says whether anything was waiting for it
- Startup says which profiles have usage-limit detection turned off
- A central allow-list of the models the fleet may use
- fleet_reply tells a lead which tool to use instead
- The credential-scrub receipt now measures the blank, not the attempt
- fleet_list no longer advertises a profile that fleet_spawn will refuse
- The startup log tells "off" apart from "running on the built-in default"
- Turning a model off at runtime, without editing profiles:
- Removing an architect slot now actually revokes it
- A worker's fleet_list no longer carries the lead's coordination state
- A usage limit now says which model to turn off, and detection can be armed without a restart
- A lead session can replace itself when its context fills up
- A timed-out send now says whether delivery was even attempted
- fleet_list and /healthz report whether the background loops are alive
- hunter is a member role, not just a skill
- The context-roll notice now obeys requireOperatorConfirm
11 — Features
What this page is for. The other chapters answer how is this built and why this way. This one answers "what can it do, and how do I turn it on" — one entry per operator-facing capability, so a feature that shipped six weeks ago is still findable without reading a design doc or a commit log.
What belongs here. A capability an operator can use, configure, or observe: an MCP tool, a
fleetd.yaml knob, an endpoint, or a behaviour visible from outside the daemon. Internal contract
changes go to Implementation; test and coverage work is a
Roadmap line. If a change adds none of those, it has no entry here — that is a normal
outcome, not an omission.
What every entry states, in this order: what it does · how you turn it on · why it exists · the gotcha. The why is the load-bearing line — it is what stops a decision being re-litigated in six weeks, and the table alone will not carry it.
Index
| Capability | Turn it on with | Since | Code |
|---|---|---|---|
| Ask the bridge who you are | fleet_whoami |
CB-517 | mcp/BridgeMcp |
| Primary inside a herdr pane | primary.terminal: |
CB-522 | auth/CallerResolver |
| More than one lead | fleet.leaders: |
CB-530 | auth/CallerResolver |
| Unknown config keys are named | (always on) | CB-530 | config/FleetdConfig |
| Find leads by tab name | fleet.leaders.<n>.tabPrefix |
CB-531 | herdr/LeadTabScanner |
| One fleet block, role as the key | fleet: |
CB-557 | config/FleetdConfig |
| Role pools decide the backend | fleet.developers: etc. |
CB-557 | member/CompositePeerLauncher |
| Tab labels name the role | fleet.tabLabel: |
CB-557 | member/HerdrPeerLauncher |
| Launch a lead when none is live | fleet.leaders.<n>.profile + instances |
CB-558 | lead/LeadLauncher |
| Re-read the config without a restart | configReload.enabled: true |
CB-559 | config/ConfigRef |
| Set a role's launch charter in config | fleet.charters: |
CB-566 | config/FleetdConfig |
| Leads talk to each other | (always on, two leads) | CB-532 | auth/Principal |
| A lead can be delivered to | automatic | CB-534 | Fleetd.deliverableTo |
| Know when a completion fallback is partial | automatic | CB-563 | inject/CompletionResolver |
| Leads are visible in fleet_list | automatic | CB-535 | mcp/BridgeMcp.listFleet |
| Reply nudges follow the delegating lead | automatic (retires primary:) |
CB-532 | mcp/PrimaryRegistry |
| Advisory architect slots | fleet.architects: |
CB-548 | auth/MemberRegistry |
| Weighted worker placement | placement: weighted + weight / maxLoad |
CB-518 | placement/ |
| Give workers a toolchain | per-profile env: |
CB-511 | worker/HerdrPeerLauncher |
| Run a worker on the subscription | profile subscription: true |
CB-539 | worker/ClaudeCodeLauncher |
| Keep a worker conversation | session name + resume id on spawn | CB-547 | peer/SpawnRequest |
| Isolated worktree per worker | fleet_spawn{worktree, ticket} |
CB-301-ext | session/GitWorktrees |
| Worker tool-surface isolation | automatic | CB-525 | session/GitWorktrees |
| Worktree-hostile config isolation | automatic | CB-543 | session/GitWorktrees |
| No credentialed remote URL reaches a worktree | automatic | CB-157 / CB-189 | session/GitWorktrees |
| Worker opens its own PR | gitTokenEnv: / gitHostEnv: |
CB-302 | worker/HerdrPeerLauncher |
| Session lifecycle caps | lifecycle: |
CB-303 | session/SessionManager |
| Keep a worktree that still holds work | automatic | CB-576 | session/GitWorktrees |
| Fail a ticket when its member dies | health.enabled: true |
CB-580 | health/FleetHealthMonitor |
| Durable reply inbox | broker: |
CB-307 | msg/AmqpReplyInbox |
| Set a member's auto-compact window | per-profile autoCompactWindow: |
CB-636 | member/ClaudeCodeLauncher, member/OpenCodeLauncher |
| Leads on different hosts talk over a shared broker | coordinator: |
CB-637 | msg/LeadMailbox, msg/LeadCoordLoop |
| A lead can read its own held peer mail | coordinator: |
#421 | mcp/FleetMcp, auth/Authz |
| Draining an inbox is lead-only, on both entry paths | (always on, with auth:) |
#272 | mcp/FleetMcp.pollAction |
| Broker password out of the config | broker.uriEnv: |
CB-635 | config/FleetConfig |
| An unreachable broker does not stop the daemon | (always on) | CB-635 | Fleetd.selectReplyInbox |
| Reject overlapping rendezvous | automatic | CB-548 | msg/Rendezvous |
| Pin an opencode endpoint | profile baseUrl: |
CB-508 | worker/OpenCodeLauncher |
| Onboard a project with the plugin | /plugin install claude-bridge → /claude-bridge:setup |
CB-527 | plugin/ |
| Port a workspace to OpenCode | port-to-opencode skill + opencode.json |
CB-529 | .claude/skills/port-to-opencode |
| Reject a profile name as a send target | automatic | CB-572 | mcp/BridgeMcp |
| See free fleet capacity | automatic | CB-573 | mcp/BridgeMcp |
| Answer a question on an async delegation | automatic | CB-574 | msg/MessageService |
| Learn why a delegation died | automatic | CB-568 | inject/CompletionResolver |
| Watch the fleet's health | health: |
CB-573 | health/FleetHealthMonitor |
| Tell a usage-limit refusal from a real reply | profile exhaustedPattern: |
CB-578 | inject/CompletionResolver |
| Stop spawning onto an exhausted account | quarantineCooldownSeconds: + profile credentialId: |
CB-578 | placement/BackendQuarantine |
| See which charter a member got | automatic | CB-571 | peer/CharterReceipt |
| Redeploy the daemon safely | redeploy-fleetd skill, then run the script |
— | scripts/redeploy-fleetd.sh |
| Run members on a second herdr daemon | memberHerdrSocket: |
CB-185 | herdr/HerdrRouter |
| Let a member under another OS user write its worktree | worktreeGroup: |
CB-185 | session/GitWorktrees |
| Resume an opencode member's prior session | automatic | CB-206 | member/OpenCodeSessionDiscovery |
| Fill in a member id the backend names late | automatic | CB-209 | session/SessionManager |
REST roster rows under members |
automatic | CB-199 | rest/FleetApp |
| Say "unknown" when the member env is unreadable | memberHerdrSocket: |
CB-185 | member/HerdrPeerLauncher |
| One reply settles one ticket | automatic | CB-137 | msg/MessageService |
| A launch command that cannot fit is refused | automatic | #220 | member/HerdrPeerLauncher |
| A failed spawn shows you the pane | automatic | #220 | member/HerdrPeerLauncher |
| Every claude-code member is resumable | automatic | #214 | member/ClaudeCodeLauncher |
| An exhausted backend is quarantined even from a chrome-only pane | automatic | #211 | inject/CompletionResolver |
| The credential scrub follows the member's own user | memberLoginShell: + worktreeGroup: |
#213 | member/HerdrPeerLauncher |
| An opencode member's config follows the member's own user | memberHerdrSocket: + worktreeGroup: |
#219 | member/OpenCodeLauncher |
| The model opencode actually ran is read back and checked | automatic (opencode profiles) | #175 | member/OpenCodeSessionDiscovery, member/OpenCodeLauncher |
| Every file a member must read follows the member's own user | memberHerdrSocket: + worktreeGroup: |
#222 #224 | member/ClaudeCodeLauncher, session/GitWorktrees |
| A role fleetd cannot bind is refused | automatic | #123 | auth/MemberRegistry |
| A reply says whether anything was waiting for it | automatic | #365 | msg/MessageService.ReplyOutcome |
Nearly every knob above lives in one file, on one profile:
flowchart LR
Y["fleetd.yaml"] --> G["bind / auth / guard"]
Y --> B["broker"]
Y --> L["lifecycle"]
Y --> W["profiles:"]
W --> P1["profile: gx10"]
W --> P2["profile: opus"]
P1 --> K["weight · maxLoad · model<br/>env · gitTokenEnv · parityOverlay<br/>configDir · cwd · argv"]
P2 --> K
Y --> F["fleet:"]
F --> FL["leaders"]
F --> FA["architects"]
F --> FD["developers"]
F --> FR["reviewers"]
FA --> K
FD --> K
FR --> K
The configuration surface has two halves. profiles: answers "which backend" — model, adapter,
cost. fleet: answers "who runs, and on which of those backends". A role pool holds profile names,
so the arrows meet: the same profile may serve several roles.
Ask the bridge who you are
What. fleet_whoami returns {"role":"primary"} or {"role":"worker", sessionId, profile, worktree, branch}.
On. Always available; no configuration.
Why. Every rule in the bridge charter is role-conditional, and both roles read the same
CLAUDE.md — a worker runs in a worktree of the same repo, so it inherits the file verbatim. Before
this, a session had to infer its role from side-channels (the mount name, ANTHROPIC_BASE_URL,
the system prompt), each of which is one-way and some of which are absent for Claude-model workers.
The daemon already resolves the role from the connection for its authorization gate; fleet_whoami
just exposes that same answer, so guessing is never necessary.
Gotcha. The answer comes from the connection and cannot be forged or overridden by an argument.
If it disagrees with what you expect, the daemon is right and your assumption is wrong — check
primary.terminal next.
Primary inside a herdr pane
What. Lets the orchestrating session run inside a herdr pane instead of an outside terminal.
On. primary: terminal: term_<id> — read the id from fleet_whoami, re-pin whenever the
primary moves panes.
Why. Caller identity resolves a loopback PID to its herdr pane, and the pane scan covers every
pane, not just fleetd-spawned ones. So a primary living in a pane classified itself as a worker and
was refused spawn/send/stop — every verb it exists to call. The failure is self-locking: the
daemon can also learn the primary's terminal, but only from fleet_send/fleet_spawn, the
exact calls being refused. Only an operator-set pin breaks the cycle, which is why the pinned value
is consulted and the learned one deliberately is not.
Gotcha. A stale pin is silent. You are simply demoted to worker and every orchestration call is refused. Re-pin after the primary changes panes, and note the daemon reads this at boot — a change needs a restart.
Weighted worker placement
What. Spreads unqualified spawns across profiles by weight, with a concurrency cap per profile and failover to the next candidate when one is unreachable.
On. placement: weighted plus per-profile weight: and maxLoad:. The default fixed policy
reproduces the historical always-the-default-profile behaviour.
Why. Profiles differ in model and cost, not tier. Without weights, every unqualified spawn piles onto one backend regardless of what it costs or how loaded it is.
Gotcha. An explicit profile: on fleet_spawn bypasses the policy entirely — placement only
governs unqualified spawns. Equal-weight candidates tie-break on YAML definition order, so
that order is load-bearing config, not cosmetics (CB-524).
Run a worker on the subscription
What. Lets a claude-code worker run on the operator's Claude subscription, on purpose, when no
off-subscription endpoint exists for its family (e.g. sonnet on ccs). The launcher injects
neither ANTHROPIC_BASE_URL nor ANTHROPIC_AUTH_TOKEN and skips SubscriptionGuard's base_url
requirement for that profile only — every other profile keeps the hard refusal.
On. Add subscription: true to a profile. The default (absent/false) keeps today's hard
boundary: a claude-code profile with no base_url may not spawn, because doing so would bill the
subscription.
workers:
sonnet:
kind: claude-code
subscription: true
model: claude-sonnet-5
argv: ["ccs", "sonnet"]
Why. Some model families (e.g. sonnet on ccs) have no off-subscription endpoint to point a
worker at. Rather than leave those profiles unspawnable, this is an explicit, visible opt-in — a
spawn under it logs a WARN naming the profile, so billing the subscription is never an accident.
Gotcha. subscription: true contradicts a baseUrl (the two state opposite intents) and is
refused at spawn if both are set. The same contradiction is refused at config load for the env:
map: on the subscription path the guard is skipped and the adapter writes neither Anthropic key, so
an ANTHROPIC_BASE_URL/ANTHROPIC_AUTH_TOKEN in env: would reach the worker having passed no
guard at all (CB-542). A subscription profile may still carry env: — just not those two keys.
weight: 0 cannot reserve this profile from unqualified weighted placement: worker-config
normalization changes non-positive weights to 1.0. Use an explicit profile: sonnet spawn until
the placement policy gains an exclusion setting.
Advisory architect slots
What. Declares named, strong-model advisory slots. A bound architect resolves as architect, can
send, reply, ask, and read, but cannot spawn, stop, or drain the fleet. The operating shape is one
human-driven lead, two short-lived architects, and N workers: the lead consults the architects sideways
and discards them, rather than creating a supervisor above the lead. Like every spawned member, an
architect becomes deliverable once it mounts the bridge MCP — until then a brief sent to it is held at
the readiness gate and never typed into its pane (CB-560).
On. Declare each slot against a configured worker profile:
architects:
sonnet-adviser:
profile: sonnet
gpt-adviser:
profile: gpt
Why. The rejected alternative put an orchestrator above the lead and made the lead a managed, resumable session. That breaks the actual control boundary: the human drives a lead that pre-exists, is recognised, and cannot be resumed by fleetd. Two advisory model families, Claude Sonnet 5 and GPT-5.6 through opencode, receive the same brief independently so agreement is evidence rather than correlated echo.
Gotcha. architects: declares slots, not sessions: no terminal is configured and nothing becomes
an architect until the lifecycle binds a live terminal to a slot. The profile name is validated at boot,
as are duplicate slot names. An architect has delegation authority but no lifecycle authority, by
design.
Shipping the binding is not the same as shipping the role. For one day an architect bound correctly,
resolved as architect, and could not receive a single message: the presence map that opens the
readiness gate was keyed on Role.WORKER, so an architect was never marked available. It built clean
and passed two reviewers. When a change adds a role, check every place that assumes a member is a
worker — fleet_status tells you whether a member ever became ready.
Keep a worker conversation
What. Carries a bridge logical session name and a peer-owned resume id through SpawnRequest, then
returns them from the peer handle where the adapter supports them. Claude Code mints a UUID for a fresh
named session, passes it as --session-id, exposes the logical name with -n, and resumes with -r.
OpenCode resumes with -s and discovers its id after launch from its on-disk session record by matching
the worker's unique worktree cwd.
On. Supply a session name and/or resume id when the spawn lifecycle has one to carry. Adapter
capabilities state the asymmetry: Claude Code offers SESSION_NAME and SESSION_RESUME; OpenCode
offers SESSION_RESUME only.
Why. A bridge session name is an operator-facing roster label, while a provider session id is the only handle that can resume the actual conversation. Treating them as peer-neutral values prevents the core from assuming Claude Code's flags are a universal protocol.
Gotcha. OpenCode has no name flag, and it writes its session record only after it persists a conversation. Its id is therefore discovered lazily and may be absent immediately after spawn; its slot name remains only in fleetd's roster.
Give workers a toolchain
What. Propagates the daemon's own PATH to every worker, plus a literal per-profile env: map.
On. Automatic for PATH; add env: {JAVA_HOME: ..., ...} on a profile to extend or override.
Why. A worker's environment does not come from your shell. fleetd hands herdr an explicit
env map and herdr merges it into its own process env — so before this, a worker inherited whatever
PATH the herdr server happened to be started with. On a long-lived herdr that can predate your
toolchain entirely, leaving workers unable to run mvn or java at all.
Gotcha. On an off-subscription profile, adapter-owned variables win over env: — the
ANTHROPIC_*/CLAUDE_* wiring is applied after it, so an env: entry cannot repoint a worker past
the SubscriptionGuard. A subscription: true profile is the exception: it skips the guard and
injects neither key, so an ANTHROPIC_BASE_URL/ANTHROPIC_AUTH_TOKEN in its env: would survive
unguarded — that configuration is refused at load (see Run a worker on the subscription).
Since the default PATH is the daemon's own, start the daemon with a good one (see the
PATH lines in deploy/dev.ltms.fleetd.plist and deploy/fleetd.service).
Isolated worktree per worker
What. Provisions a git worktree on its own branch per worker, and copies a configurable set of local config files in ("parity overlay") so the worker sees the same local setup.
On. fleet_spawn{worktree: true, ticket: "cb-123"}; overlay list via parityOverlay:
(default [.claude/settings.local.json, .env, .envrc]).
Why. Parallel workers editing one checkout collide. A worktree gives each its own branch and files at the cost of a checkout.
Gotcha. Release removes the checkout but keeps the branch — unmerged work survives a
teardown. Overlay-copied files that are tracked get --skip-worktree so they never read as pending
changes. Never add .mcp.json to the overlay; see the next entry for why.
A failed provision cleans up after itself (fleetd #274). Provisioning is more than
git worktree add: several steps run after the worktree and branch exist, including the
credential-free-origin check, which is an intended refusal rather than an IO accident. If any of
them fails, the worktree and its branch are now removed before the error is rethrown. Before #274
they leaked permanently — the caller never received the path, so its own cleanup could not fire,
and nothing else tracked the directory. Note the asymmetry with the paragraph above: a released
session keeps its branch, because a worker's branch is meant to outlive its worktree; a provision
that never completed has no session and no PR, so its branch goes too.
The isolation does not cover git stash. A worktree separates the working tree, the index and
HEAD — but refs/stash is a single stack shared by the primary's checkout and every worker
worktree of the repo. Two parallel workers stashing at the same time can pop each other's work,
silently. This happened on 2026-09-04 between the #273 and #274 workers; both noticed and recovered.
The implementer skill now forbids git stash and points at a wip: commit or a patch file.
Worker tool-surface isolation
What. A provisioned worktree's project .mcp.json is neutralized to an empty server map, so a
worker's tools are only what its launcher mounts (the bridge).
On. Automatic at provisioning. Nothing to configure.
Why. The repo commits a .mcp.json declaring the primary's IDE servers, so a fresh checkout
mounted them and the parity overlay copied the primary's own copy on top. Those servers are bound
to the primary's IDE project, so every path they return points into the primary's checkout. This
is not hypothetical: a worker made all 59 of its edits in the primary's tree while running mvn
against its worktree — so every build it ran was of code that did not contain its changes, and every
build passed.
Gotcha. This is why a worker cannot run IDE diagnostics and mvn is its only verification. That
is deliberate — the bridge is a message bus, and the primary is the gate. Never accept a worker's
claim about a check it had no way to run.
Worktree-hostile config isolation
What. Every provisioned worktree neutralizes tracked .mcp.json, opencode.json, and .autoenv
with a valid format-specific stub, then marks a tracked replacement --skip-worktree.
On. Automatic at worktree provisioning; no configuration.
Why. The tracked opencode.json mounts the primary's own gitea and context7 servers, with the
primary's credentials, so a worktree that keeps it hands a member access it must never hold. The same
isolation rule also keeps a primary-only MCP configuration or an autoenv authorization prompt out of a
worker's tool surface.
The reason used to be milder, and the change is worth recording. opencode.json referenced gitignored
.secrets/ files that a worktree never contains, and OpenCode refuses to start on that dangling
reference — a crash, but a loud one. Those credentials now live in one shell-level store and the file
reads them as {env:…}, so in a worktree the reference resolves instead of failing. A loud crash
became a quiet privilege leak, which makes this isolation load-bearing rather than a workaround.
Gotcha. The replacement is deliberately valid, not deleted: {} for opencode.json, an empty
server map for .mcp.json, and an empty .autoenv. A deletion could be undone by a later checkout; the
skip-worktree bit keeps the safe local replacement from looking like work for a worker to commit.
No credentialed remote URL reaches a worktree
What. Before provisioning a member's worktree, fleetd makes sure the repository's remote URLs carry no embedded credentials, and says so when they do.
Three things happen, in this order:
- Report. Every remote is enumerated, and both its fetch and its push URL are checked for user-info on any non-SSH scheme. Anything found is logged as a WARN naming the remote.
- Strip. User-info is removed from
origin's HTTPS URL. - Refuse. If the provisioned worktree still resolves a credentialed HTTPS
origin, the provision fails rather than handing the member a URL with a secret in it.
WARN member worktree shares a remote URL containing user-info: remote=upstream
repository=/Users/…/fleetd; remove credentials from the repository's git config
On. Always on. There is no key.
Why. A linked worktree shares its parent repository's git config. A token embedded in a
remote URL is therefore readable by the member the moment its worktree exists — and git remote -v
prints it in full, so it leaks again the first time anyone surveys the host. Members are meant to
push with the repo-scoped WORKER_GITEA_TOKEN injected at spawn, never with a credential baked into
config.
Gotcha — the report and the refusal cover different ground, on purpose. Only origin, and only
HTTPS, is stripped and refused. Every other remote, every pushurl, and every other non-SSH scheme
is reported and then left alone. Rewriting a remote nobody asked us to touch is not fleetd's
call; telling you it is there is. So a WARN here is work for you to do, not something the daemon has
already handled.
SSH URLs are deliberately exempt. There the user part selects an account and authentication happens
over the SSH transport, so git@host is not a credential the way user:token@host is.
The reason it warns before it strips. For the origin HTTPS case you will see the WARN
immediately followed by the strip's own INFO line. That ordering is intentional: it leaves an audit
trail that there was something to fix. Reporting after the strip would silence the one case the
daemon actually repairs.
Gotcha — a reporting check must never break a provision. An earlier attempt at this ran its git
calls unguarded at the top of add(), where any non-zero exit or the 30-second timeout would have
aborted the whole worktree provision. Every call here is wrapped, and a failure logs only the
exception's class, never its message — the message is read from the very config that may hold the
URL being looked for.
For the same reason, every git command whose stdout is a URL runs through a redacting variant
that keeps captured output out of exception messages. Without it a failing git remote get-url
copies the credentialed URL into the exception, and from there into the log — defeating the check by
way of its own error path.
Worker opens its own PR
What. Injects a repo-scoped forge token so a worker can commit, push over SSH, and open its own pull request at checkpoint.
On. gitTokenEnv: (host env var holding the token) and optionally gitHostEnv: on a profile.
Why. Opt-in by design: omit it and the worker gets no PR-create grant, while push over SSH still works.
Gotcha. Use a minimal write:repository token, never an admin one. The token can create a
PR but must not be able to merge — the primary is the gate, and a worker that can merge is not
gated.
Where the value comes from. gitTokenEnv: names a variable, and fleetd reads it from its own
process environment — so the token must be exported in the shell that launches the daemon, not
stored in the repo. Keep it beside the operator's other credentials in one sourced file; a per-repo
copy is a second copy of the same secret, and the copy you forget is the one that leaks or goes
stale. If the daemon was started before that export existed, it injects an empty token and every
worker push fails: restart it from a login shell.
Session lifecycle caps
What. Reaps idle sessions, caps turns per session, optionally clears a reused Claude Code worker's conversation after each delegation, and drains cleanly on shutdown.
On. lifecycle: { idleTtlSeconds, contextCap, drainTimeoutSeconds, clearAfterTurn }.
Why. Workers are disposable but not free; without caps an abandoned session holds a pane and a
context indefinitely. clearAfterTurn: true keeps the pane and process warm while preventing
delegation N+1 from inheriting delegation N's conversation.
Gotcha. clearAfterTurn defaults to false. It uses Claude Code's /clear command without
counting that housekeeping as a delegated turn, and waits for it to settle before delivering the
next task. Peer kinds without a known context-reset operation (currently opencode) treat the knob as
a no-op and log that once; fleetd never guesses a command. Lifecycle config is read at boot, so
changes need a daemon restart.
Durable reply inbox
What. A worker's reply survives with no waiter attached: it is queued and collected later by
fleet_poll / fleet_ack, over a real AMQP broker when one is configured.
On. broker: pointing at an AMQP URI; omit it for an in-memory inbox.
Why. A blocking fleet_send is capped by the caller's MCP client timeout (~60s), far below a
real task's runtime. Without a durable inbox, a reply arriving after that window lands nowhere.
Gotcha. The broker is LavinMQ, not RabbitMQ. Point it at a stray local RabbitMQ and you are writing into someone else's broker. Also see the known hole: an async ticket that times out while the session is still BUSY currently discards the later completion rather than parking it.
Draining an inbox is lead-only, on both entry paths
What. fleet_poll{target} empties that session's reply inbox — the replies are removed and a
second call returns nothing. It is therefore gated as a drain, exactly like fleet_ack, and only
the primary may call it. fleet_poll{ticket} is unaffected: it observes an async delegation and
changes nothing, so a worker or an architect may still poll a ticket it owns.
On. Always on, whenever authorization is enforced (an auth: block / a real CallerResolver).
Nothing to configure.
Why. Until fleetd #272 the MCP handler passed a constant READ for both branches, and READ is
open to every authenticated role. Any worker could read a peer's sessionId out of fleet_list and
destroy the replies that peer had queued for the primary — an unrecoverable loss, since a drained
reply is gone. The REST path had always checked DRAIN, and the wiki had always described the tool
as lead-only; the MCP gate was the one that disagreed.
Gotcha. An architect cannot drain either, even though it may fleet_send. That is
deliberate and matches fleet_ack: delegating is not a lifecycle right. If an architect delegates
with wait:false it still polls its own ticket, which is the branch that stayed open.
The general rule this came from is worth keeping: when one tool name covers two operations, the gate
belongs inside the branch that picks between them, not above it. fleet_poll was the only
handler in the server with that shape — the other ten were audited and are correct.
Broker password out of the config
What. broker.uriEnv: names an environment variable that holds the AMQP URI, instead of writing
the URI into the config file. An AMQP URI carries user:password@ inline, so the old broker.uri:
put a live password in clear text in fleetd.yaml.
On. broker: { uriEnv: LAVINMQ_URI }. The variable must be on the daemon's own environment,
so start the daemon from a login shell — scripts/redeploy-fleetd.sh --check reports whether the
named variable resolves, without ever printing its value.
Why. The same reason auth.tokenEnv and Profile.tokenEnv exist: a config file gets read,
copied and pasted into tickets far more often than a secret store does. This was the last credential
still living in clear text in the config.
Gotcha. uriEnv wins over uri whenever it is set, and it does not fall back. If the named
variable is unset or blank the daemon treats the broker as unconfigured and uses the in-memory inbox
— it does not quietly drop back onto a stale uri: left in the file. That is deliberate: an operator
who moved to the secret store must never be silently returned to clear text. Both set ⇒ the log says
broker.uri is ignored.
An unreachable broker does not stop the daemon
What. If the broker cannot be reached at boot, fleetd logs a loud warning and starts on the in-memory reply inbox for that process lifetime, instead of failing to start at all.
On. Always on. There is no background retry: fix the broker and restart to get durability back.
Why. AmqpReplyInbox.open throws IllegalStateException, and nothing caught it — so a broker
that was down took the whole daemon with it, and with the daemon went every lead, every member and
every pane. Losing durable replies is bad; losing the fleet because the reply store is down is far
worse. A broker that drops after startup already self-heals through the AMQP client's automatic
recovery, so only the boot path needed this.
Gotcha. The fallback is not silent, but it is easy to miss in a busy startup log. What you
lose is real: replies become soft-state and a held report does not survive the next restart. Grep the
startup log for reply inbox: — it says which adapter won, every time. The warning names the failing
URI with the credentials stripped.
Reject overlapping rendezvous
What. A second attempt to open a reply waiter for the same worker session fails atomically instead of replacing the first waiter.
On. Always on; no configuration.
Why. One session has one outstanding delegated turn. Replacing its waiter silently would strand the
first caller and let a reply resolve the wrong request. MessageService serializes normal sends, but
the atomic rejection is the tripwire that makes a future violation loud rather than corrupt.
Gotcha. This is not concurrent-turn support. A terminal send closes its own waiter before the next turn can open one; a double-open is an invariant failure that must be investigated.
Pin an opencode endpoint
What. An opencode profile can target its own OpenAI-compatible endpoint.
On. baseUrl: on a kind: opencode profile.
Why. opencode is provider-agnostic and shares none of Claude's private seams — no
ANTHROPIC_BASE_URL, no SubscriptionGuard, no --mcp-config. It is the adapter that proves the
PeerLauncher SPI is genuinely provider-neutral rather than Claude-shaped.
Gotcha. Because it bypasses SubscriptionGuard, the guard's allowlist does not protect this
path — the endpoint you name is the endpoint it uses.
Onboard a project with the plugin
What. A Claude Code plugin that makes any project fleet-ready: it mounts the fleetd MCP
gateway and ships a /fleet:setup skill that runs preflight, applies standard project
settings, names the environment variables the operator must export, and verifies the session
resolves as the primary.
On. /plugin marketplace add https://git.ltms.dev/fleet/fleetd then /plugin install fleet@fleetd; export FLEETD_MCP_URL (usually http://127.0.0.1:8765/mcp); run /fleet:setup
inside the project to onboard. Develop it locally with claude --plugin-dir ./plugin.
Why. The orchestration contract had no distributable form. Every consuming project had to
hand-copy a block of CLAUDE.md and hand-write an .mcp.json, and we maintained a script purely to
detect the copies drifting apart. A plugin is versioned, installed once, and updates in place — the
contract stops being something each project re-derives. It ships no credentials by design, so
the artifact is public-safe: every secret is referenced by environment-variable name and the value
never enters a file.
Gotcha. The plugin is client-side setup only — it mounts a daemon, it does not install one.
fleetd and herdr remain separate services, and the setup skill deliberately refuses to install
them (guessing at a system-service install is how you get two daemons on one socket). Note also the
plugin root is plugin/, not the repo root: an installed plugin's .mcp.json is a committed
file, while this repo's root .mcp.json is local-only and --skip-worktree, so rooting the plugin
at the repo would collide with the very isolation CB-525 exists to enforce.
Renamed in 0.2.0 (fleetd #362) — this is a breaking change. The plugin was claude-bridge and
mounted its server as fleetd; it is now fleet@fleetd and mounts fleet. The old name gave a
lead with both a project .mcp.json and the plugin two mounts of one daemon and a duplicated
fleet_* tool set, and it did not match PeerLauncher.MCP_MOUNT_NAME. A project that pre-allowed
mcp__fleetd__fleet_whoami in .claude/settings.json must be updated to mcp__fleet__*. The
.mcp.json URL is now ${FLEETD_MCP_URL} rather than a hardcoded address, so one plugin can serve
hosts running the daemon on different ports — the variable is required, and the setup skill
checks for it in preflight.
The plugin is lead-side only, and cannot be otherwise. It carries the mount, the setup skill
and the charter — never the worker playbook skills or the role agent definitions. Two structural
reasons, both measured: the launcher adds --agent only when <worktree>/.claude/agents/<role>.md
exists in the member's own tree (ClaudeCodeLauncher.java:371,391); and a member's
CLAUDE_CONFIG_DIR points at its profile's config directory (ClaudeCodeLauncher.java:285), so it
never reads the operator's plugin store. On this Mac every Claude profile sets configDir, and the
four ccs instances hold four separate copies of the plugin store — same md5, different inodes —
so a user-scope install lands in exactly one of them. Member-facing assets travel in the worktree.
This entry existed and still did not prevent a rebuild. In September 2026 a session planned the
whole plugin from scratch, because wiki/ is a submodule whose pointer is never advanced and no
session reads it by default. Documenting a feature here is necessary and not sufficient — a
capability an agent must not re-derive needs a line in CLAUDE.md, which is the only file every
session loads.
Port a workspace to OpenCode
What. A port-to-opencode skill that makes an OpenCode session a first-class participant in a
Claude Code workspace — same instructions, same MCP servers, same bridge mount — by writing a single
opencode.json and nothing else.
On. Load the port-to-opencode skill in the primary. It is a primary-side skill, not a
delegation playbook: a worker mounts only the bridge MCP and cannot run it.
Why. OpenCode reads CLAUDE.md natively — including ~/.claude/CLAUDE.md — so the rules cross
for free and only MCP servers need mapping. That is worth writing down because the obvious move is
the wrong one: the Codex attempt translated CLAUDE.md into a second rules file with a third-party
tool, and the translation silently corrupted a "never commit" rule into one naming a path that
cannot be committed at all. The skill exists to stop anyone reaching for a porting tool again.
Secrets cross by {env:VAR} reference, which is what makes opencode.json committable — and it
must be, because a peer in a worktree receives tracked files only.
Gotcha. An AGENTS.md left in the repo shadows CLAUDE.md — opencode takes the first match
walking up from the cwd, so a stale file from an earlier port silently wins over the live rules.
Delete it before anything else. And skills do not cross: .claude/skills/** is not read by opencode,
so a brief telling a peer to "load the implementer skill" is a no-op there — spell the procedure
out in the brief instead.
More than one lead
What. Recognises several panes as leads, so two orchestrators — say a Claude lead and an opencode lead — work as peers instead of one being demoted.
On.
leaders:
opus-5.0:
terminal: term_0123456789abcd
kind: claude
gpt-sol-5.6:
terminal: term_fedcba9876543
kind: opencode
model: openai/gpt-5.6-terra
terminal is the only field identity depends on; kind/model are descriptive and are echoed back
by fleet_whoami as leader: <name>. role still reads primary — a lead is a primary for
authorization, so nothing keying on the role breaks.
Why. primary.terminal is singular by construction: one pane is the lead and every other pane
resolving to a herdr terminal is a worker. That is right while one lead drives a fleet, and wrong the
moment two leads collaborate — the second is silently demoted and refused every orchestration call it
makes. Resolution is now a registry lookup rather than an equality test against one pin.
Gotcha. This entry originally said to keep primary: alongside leaders:, because the CB-307
push loop needed exactly one nudge destination while leaders: only widened who was recognised.
That is no longer true: reply nudges now follow the delegating
lead, and primary: is retired. Delete it. If both are
present and name the same terminal the leaders: entry wins. And a lead is never spawned — it
pre-exists, which is why it
must be named here rather than created; argv/placement are worker-profile keys and mean nothing
in this block. Read at boot, so a change needs a restart.
Find leads by tab name
What. Discovers leads by scanning herdr for tabs you labelled, instead of you pasting each
lead's terminal_id into leaders:. Label a tab lead: gpt-sol-5.6, start
an agent in it, and that pane resolves as a lead named gpt-sol-5.6 within one rescan — no config
edit, no daemon restart.
On.
leadScan:
tabPrefix: "lead:" # matched case-insensitively; the rest of the label is the lead's name
intervalSeconds: 10 # rescan cadence, and the worst case before a new tab is recognised
Opt-in: no block means leads come only from leaders:/primary:, exactly as before. Both sources
merge, and an explicit leaders: entry outranks a label for the same terminal.
Why. A lead is never spawned — a human opens a tab and starts an agent in it — so unlike a
worker, the daemon cannot learn its terminal_id at creation. leaders: therefore costs a
four-step ritual per lead: start the session, ask it fleet_whoami for its id, edit config,
restart. Naming the tab is one step, taken at the moment the operator is already there. The label
also survives what the id does not: close and reopen the tab and the terminal_id changes, while
the label is retyped as-is.
Gotcha. The direction of trust is what makes this safe, and it is one-way: fleetd reads
lead tab labels and never writes them, so what is in the tab bar is always what a human typed.
Two guards keep that from eroding — the configured worker spaces (where fleetd does write
labels, via tabLabel) are excluded from the scan wholesale, so nothing the bridge places can land
in a matching tab; and startup refuses a tabPrefix that any worker tabLabel also matches,
because overlapping those two namespaces would have the daemon label its own workers as leads and
promote the entire fleet. Every pane in a labelled tab is that lead, so split a lead tab only with
panes you mean to be leads. A failed scan keeps the leads already known rather than emptying the
registry — a herdr hiccup must not demote a live lead mid-session.
Leads talk to each other
What. A lead can message another lead and be answered. fleet_send{sessionId: <peer's terminal>} reaches a peer, and the peer closes the exchange with fleet_reply — the same
rendezvous a worker uses. fleet_whoami now reports a lead's own sessionId, which is how a lead
learns the address to give a peer.
On. Nothing to configure; it applies as soon as two panes resolve as leads (via
leaders: or leadScan:).
Why. CB-530 and CB-531 widened recognition — both leads are seen — but nothing had widened
addressing, so collaboration was one-way and silently so: the send was accepted, the peer's
fleet_reply was refused, and the sender waited out its timeout. The cause was that a lead's
Principal carried no terminal, so ownsSession() could never be true for it and the REPLY/ASK
rules excluded it by construction. A lead now carries the pane it was matched by, and the rule it
must satisfy is unchanged: you may act as the pane you occupy, and as no other. That was always
the real control — terminal.equals(sessionId), against a terminal that comes from the connection —
and the extra "…and you must be a worker" conjunct beside it protected nothing.
Gotcha. A lead replies only to answer a peer that messaged it — never to answer a worker, whose turn it is not, and never as a way to end its own turn. Widening who may reply did not widen what they may reply as: a lead still cannot act for another pane, and an unnamed primary (token mode, or off-host, with no pane at all) owns nothing and remains a sender only.
Leads are visible in fleet_list
What. fleet_list returns leads alongside workers. Each lead row carries its sessionId
(the address to fleet_send to), its name, its live status, and self: true on the caller's own
row.
On. Automatic, wherever a pane resolves as a lead.
Why. A lead had no way to discover a peer. fleet_list enumerated the worker roster alone, so a
lead asking "who else is here?" got an empty array — which reads as no peers but only ever meant
no workers spawned. The peer lead on this bridge drew exactly that wrong conclusion and reported
itself alone in a two-lead fleet. Addresses had to be carried between panes by a human, which is not
a protocol. Reporting both halves — even when a half is empty — also removes the ambiguity that
caused the misreading.
Gotcha. The lead rows come from the same registry that resolves identity
(CallerResolver.leads()), not from a second copy, so a listed address is
one that would actually resolve as a lead. A lead herdr is not tracking as an agent is listed with
status: unknown rather than hidden — it cannot be delivered to, and a would-be sender needs to see
that rather than infer it from silence.
A lead can be delivered to
What. The injector's readiness gate opens for a lead as well as for a booted worker, so a message addressed to a lead is actually typed into its pane.
On. Automatic, wherever a pane resolves as a lead.
Why. The gate (CB-113) holds a delivery out of a spawned worker's boot window: herdr reports
idle while the agent is still starting, and a paste into that window is lost. Membership in it
comes from MemberPresence, which BridgeMcp populates for every spawned member — worker or
architect, but never a lead, since that map doubles as the roster's availability signal and a lead
counted there would appear as an available member. The two rules composed into a dead end: a lead is never marked present, so the
gate never opened for one, so lead-to-lead messaging — shipped and
authorized in CB-532 — still could not deliver a single keystroke. The gate's premise simply does not
apply to a lead: a lead is never spawned, so it has no boot window to guard.
Gotcha. The failure this fixes was slow and mute, which is worth recognising if it recurs in
another form: the send was accepted, the pane stayed idle, nothing was ever typed, and ~60s later
(READINESS_GRACE_POLLS, 240 × 250ms) it failed through the same path as a stalled turn — so the log
said turn-stall fallback while the truth was that delivery had never been attempted. A repeating
62-second gap between send and failure is the signature of the gate, not of a peer that ignored you.
The gate is no longer mute (CB-562). When the grace expires it logs a WARN naming the target, the
poll count, the grace in seconds and how many queued messages it is failing because the target never
became deliverable — so a readiness failure now reads differently from a turn stall. The seconds are
derived from Injector.POLL_INTERVAL_MILLIS, the single source the poller is also built from, so a
cadence change cannot leave the log confidently stating a wrong duration.
It was also intermittently masked: presence is a sticky set, so a pane that was seen as a worker
before being recognised as a lead stayed deliverable until the next restart cleared the set.
Reply nudges follow the delegating lead
What. When a worker's reply lands with no fleet_send open, the CB-307 nudge goes to the lead
that delegated that worker — not to a globally-configured "the primary". This is what retires
primary.terminal:, which now logs a deprecation warning at startup.
On. Automatic. Delete primary: from fleetd.yaml; keep it only if you still want
pushReminders/pushBackoffMs, or a fallback nudge destination across restarts.
Why. PrimaryRegistry held one slot answering "who is the primary" — a question with no correct
answer once two leads drive one fleet. Whichever lead called fleet_send first captured every
nudge thereafter, so the other lead's results were announced to the wrong pane. The binding that
actually matters is per-delegation and is known exactly where it is created: at fleet_send, where
the target is the argument and the lead is resolved from the connection.
Gotcha. A restart loses the delegation map while the durable inbox keeps the reply. With one
lead the old pin (or the first lead to send) is an unambiguous fallback; with several and no
recorded delegation the daemon nudges nobody rather than guessing, and delivery degrades to
fleet_poll. That is the correct degradation — interrupting the wrong lead with someone else's
result is worse than a quiet inbox — but it does mean a post-restart reply may need an explicit
poll.
Second gotcha, fixed in fleetd #368. A binding was only ever dropped when the worker was
released. The map is keyed by the worker, so a lead that was closed, crashed or relaunched left its
bindings behind, pointing at a pane that no longer exists. A stale entry is not null, so it beat the
single-lead fallback above every time, and the nudge was sent into nothing. The push loop now checks
the recorded lead with agents.status before trusting it, and forgets a lead that is really gone so
resolution reaches the fallback. Only agent_not_found counts as gone: any other failure — a socket
blip, a decode error — is treated as live, because a wrong guess there unbinds a working lead
permanently, while a wrong guess the other way costs one retry on the next tick.
Unknown config keys are named
What. A top-level key in fleetd.yaml that this build does not understand is logged as a WARN
naming it, at load.
On. Always on; nothing to configure.
Why. Every config record is @JsonIgnoreProperties(ignoreUnknown = true) — deliberate, so config
may run ahead of the code and a rolled-back daemon still starts. The cost is that a whole block can be
written, parsed, dropped, and never mentioned again. That is exactly how a hand-written leaders:
registry came to look configured while being inert: the daemon started, nothing complained, and the
only way to find out was reading the config class. A dropped block and a working one were
indistinguishable.
Gotcha. A warning, not a failure — deliberately. Failing closed would turn "the config names something this build has not learned yet" into a daemon that will not boot, destroying the forward-compatibility the annotation exists for. Nested unknown keys are still silent; only the top level is checked.
One fleet block, role as the key
What. fleet: replaces four top-level keys — leaders:, members:, leadScan: and
defaultProfile:. A member's role is now the map key that contains it, not a role: field inside
it.
On. Write a fleet: block. The four old keys are hard errors that name what to use instead, so
an old config does not start silently changed.
fleet:
leaders:
opus: {profile: opus, instances: 1, tabPrefix: "lead:"}
architects:
opus: {profile: opus}
developers:
local: {profile: local}
sonnet: {profile: sonnet}
reviewers:
sonnet: {profile: sonnet}
Why. A misspelled role: architct used to parse into a member with a profile, a name, and no
contract at all — nothing rejected it, because role: was just a string. A misspelled pool name
declares nothing, which is a shape the loader can see. The four keys also had no relationship to
each other on the page, while all four describe one thing: who is in the fleet.
Gotcha. defaultProfile: has no single successor key, so its error message explains the new
model rather than pointing at a key that does not exist. Role pools took over its job — see below.
Role pools decide the backend
What. fleet.architects / developers / reviewers list the profiles that role may run
on. A spawn that names no profile is placed inside the pool of the role it asked for, instead of
across every configured profile.
On. List profile names under the role. Order matters under placement: fixed — the first entry
wins.
Why. Role and profile are separate axes, and collapsing them loses real cases: a reviewer may
run on the very same profile as the dev whose diff it reads. Before this, an unqualified spawn
ranged over all profiles, so a reviewer could land on the architect-only backend and quietly spend
the subscription. Pools also replaced the single global defaultProfile:, which could only ever
have one answer for a fleet that has three kinds of member.
Gotcha. A role with no pool is unconstrained, not blocked — it falls back to every profile,
so a config that pools some roles and not others keeps working. And an explicit profile is not
confined to the pool: fleet_spawn{profile:"opus"} carries no role, so it defaults to dev, and
judging it against the dev pool would refuse a spawn the operator asked for by name. maxLoad still
applies to it.
Tab labels name the role
What. A member's tab reads dev: sonnet #4 — role first, then backend, then a counter.
On. fleet.tabLabel: (default "{role}: {profile} #{n}"). {role}, {profile}, {model} and
{n} are substituted. A profile may override it with its own tabLabel:.
Why. The label lives on fleet: because a profile cannot know the role of the member launched
on it, and the role is the thing an operator scanning a tab bar actually wants. {n} counts per
role and profile, so a dev and a reviewer on one profile each start at #1 — a single fleet-wide
counter would make the number meaningless. Putting {role} first also turns the lead/member
namespace check into a structural guarantee: roles are a closed enum, so only a hand-written
template can still collide with a lead's tabPrefix.
Gotcha. The knob was inert on first release — HerdrPeerLauncher accepted the template and
nothing passed it, so the label only looked right because the fallback happened to match the default.
Fixed in CB-557; tests now pin the wiring rather than the coincidence.
Launch a lead when none is live
What. The daemon starts a lead declared under fleet.leaders: when fewer than instances are
running. It labels the tab by the same convention the scanner reads.
On. Give the lead a profile:. Omit it and the lead stays recognise-only, exactly as before.
instances: 0 is an off switch. workspace: (default leads) and cwd: control where it lands.
Why. A lead was the one pane a human had to open by hand before anything else worked, so a daemon restart after a reboot left a fleet with no orchestrator and no sign of why.
Gotcha — three, and they are the whole design. An auto-launched lead is not a member: it
gets no worker reply charter (that text tells its reader it is an off-subscription worker who must
end every turn with fleet_reply — the opposite of an orchestrator), it is never registered with
SessionManager (the idle reaper would kill it for being idle, which is a lead's normal state), and
ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.
A lead counts as live only when herdr reports a running agent — either in a tab labelled
lead: <name>, or on a pinned terminal:. Both are needed: label-only would relaunch a hand-opened
pinned lead on every boot, and pin-only would miss one the daemon started itself. A labelled tab with
nothing running in it is not a lead, so one crash does not disable auto-launch forever. If herdr
cannot be reached the daemon starts nothing — a second orchestrator is worse than none.
workspace: must not name a member workspace: those are excluded from the lead scan, so a lead
placed in one would never be found again and would be relaunched on every boot.
Re-read the config without a restart
What. fleetd watches fleetd.yaml's modified time and re-reads the file when it changes.
Consumers read the live config at the point of use, so a change reaches the next spawn without
rebuilding anything.
On. Add configReload: {enabled: true}. intervalSeconds: sets the poll period (default 10).
Absent the block nothing is constructed, so an upgraded daemon behaves exactly as before.
Why. Tuning a fleet meant restarting the daemon, and a restart tears down every lead and worker
it owns. Changing one pool's weight cost the whole fleet's state, so in practice nobody changed it.
Gotcha — three classes of key, and the difference is what already exists at reload time.
| Class | Keys | What a reload does |
|---|---|---|
| Hot | fleet: (pools + charters + tabLabel), placement:, an existing profile's weight / maxLoad |
takes effect on the next spawn |
| Deferred | lifecycle:, leadHeartbeat:, guard:, worktreeRoot:, spawnReadyTimeoutMs / spawnReadyPollMs, adding or removing a profile, and an existing profile's launch settings (model, baseUrl, argv, env, configDir, mcpUrl, tabLabel, exhaustedPattern) |
accepted into the new config, but the startup wiring keeps the old value; logged by name |
| Cold | bind:, herdrSocket:, broker:, auth: |
refuses the whole reload |
A changed cold key refuses everything, not just itself. Applying the hot half and warning about the cold half would leave the daemon in a state matching no file on disk — the worst outcome for an operator who is reading the file to work out what the daemon is doing. Refusing keeps one invariant: the live config is always some version of the file.
What decides the class is who reads the key and when, not how important it is. weight and
maxLoad are hot because the placement policy reads them live through a supplier. A profile's
model looks like it should behave the same way and does not: HerdrPeerLauncher takes
Map.copyOf(profiles) at construction and resolves every spawn out of that copy, so a reloaded
model never reaches a launch. Adding a profile is deferred for the same underlying reason — a new
backend needs its own launcher, and launchers are built once.
That distinction is worth stating plainly because getting it wrong is invisible. A reload that
reported a changed model as applied would log a clean "config reloaded" while every spawn kept
using the old one, and an operator would have no reason to doubt it. So the reload compares each
surviving profile's launch settings and names the profile in the deferred list instead.
A reload that fails to parse, or fails any of the startup validators, is refused the same way and the running config stays live. A config file being saved is sometimes read mid-write, and degrading a working daemon over a half-written file would be a bad trade. The watcher stamps the modified time before it reloads, so a refused file is not retried every tick — the next save earns a fresh attempt.
Set a role's launch charter in config
What. fleet.charters: holds the launch text for each member role, keyed by the singular role
wire name — architect, dev, reviewer. An operator edits the text and the next spawn uses
it. No rebuild, no daemon restart.
On. Add the block under fleet:. Every key is optional, so a config that never mentions
charters behaves exactly as before.
fleet:
charters:
architect: |-
You refine work before anyone builds it: scope, acceptance criteria,
risks, and a unit split. You never commit production code.
Why. The only per-role launch text used to be REPLY_CHARTER, a static final String in each
launcher. Changing what a role is told meant editing Java, rebuilding, and restarting the daemon —
which tears down every lead and member it owns. So in practice nobody changed it, and the text
drifted away from the truth: it still told every member "You are an off-subscription worker" long
after the daemon started resolving architects as well.
Gotcha — the type is a map on purpose, and blank is not the same as absent.
FleetdConfig.Fleet is @JsonIgnoreProperties(ignoreUnknown = true). A typed record field named
architetc would be dropped in silence, so the operator would see a clean "config reloaded" and no
charter. A Map<String,String> keeps every key the operator wrote, which lets validateCharters()
see the bad one and refuse the whole config, naming the key and listing the valid ones.
An absent key is fine — that is every deployment before CB-566, and it means "this role has no configured charter". A blank key is refused. The operator typed the key and expected something; treating it as absent would quietly strip that role of its contract.
Validation runs on both the startup path and the reload path. Wiring only one of the two is the whole bug: a config that a running daemon refuses but a restart accepts, or the reverse.
The charter's text is checked against the live tool surface too (fleetd #469). Validation used
to look only at the shape of the map: is the key a role name, is the text non-blank. It never read
what the text said. So a charter telling a member to call bridge_send — a tool name the CB-634
rename removed — started the daemon cleanly, and the member found out at run time by calling
something that was not there. Now CharterToolSurface pulls every fleet_* and bridge_* token
out of each charter and asks FleetTool, the one enum the MCP server derives its registrations
from, whether that tool exists. If one does not, the daemon refuses to start and the message names
both the charter key and the unknown tool.
A reload is checked too, not only startup (fleetd #474). For a short while this check had one
call site, in Fleetd.main, so editing fleet.charters: on a running daemon to name a tool that
does not exist was accepted, applied and delivered to the next member — while a restart would have
refused the very same file. Charters are hot and are read live at each spawn, so that gap was
reachable. ConfigRef.reload() now runs the same check, and a reload that fails it is refused whole
and keeps the running config, exactly like any other validation failure. The daemon logs
config reload from <path> refused, keeping the running config: <message>, and the message names
both the charter key and the unknown tool.
Why it was built this way: the check cannot live inside FleetConfig, because the canonical tool
set is in the mcp package and config is loaded before the MCP server exists. So ConfigRef takes
it as a Consumer<FleetConfig> and Fleetd.main supplies it — Fleetd is the one seam that
already holds both a loaded config and the mcp package. The gotcha: nothing yet stops someone
reverting Fleetd.main to the two-argument ConfigRef constructor. That one edit turns this gate
off with the whole test suite green, which is why the wiring is worth reading before you trust it.
Do not put secrets in charter text. There is deliberately no ${ENV} interpolation. The
OpenCode adapter writes the composed charter to a temp file so its CLI can read it, and that file is
world-readable.
Know when a completion fallback is partial
What. If a member ends a turn without fleet_reply, fleetd scrapes its pane and resolves the
waiting send with that tail, so the sender is not left hanging. The tail is capped at
MAX_SCRAPE_CHARS (4000). When the cap bites, the returned text now ends with
[Pane tail clipped: member did not call fleet_reply.] and fleetd logs a WARN with the original
length and the cap.
On. Automatic, whenever the completion fallback reads more than 4000 characters.
Why. A clipped transcript used to be indistinguishable from a complete report. A delegating lead
could act on an engineering report whose end had been cut off and never know — the only trace was a
DEBUG line reading (4000 chars scraped), which reads like a size, not a warning. The cap itself is
deliberate and unchanged; the defect was silence, not the number.
Gotcha. The marker does not recover the missing text, and it is not a licence to skip the reply.
The fallback is a liveness net, not a channel: it returns only the pane tail, stripped to the last
assistant block. A member must still end every delegated turn with exactly one fleet_reply. The
marker is appended after the CB-115 misattribution guard compares the scrape to its baseline, so
marking cannot make an unchanged pane look like new output.
Reject a profile name as a send target
What. fleet_send refuses a sessionId that exactly matches a configured profile name, on both
the blocking and the wait:false path, before any ticket is issued. The error names the value and
points the caller at fleet_list.
On. Always on; no configuration.
Why. A lead sent to sessionId: "sol" — the profile name, not the terminal id. The bridge accepted
it and returned accepted — task delegated, so the lead believed the work was dispatched. About 60
seconds later the injector logged sol is gone, dropping its queue, and twenty minutes after that the
ticket still reported pending — worker unknown. The whole delegation was lost and nothing told the
sender. Confusing a profile for a session id is the single easiest mistake to make with fleet_send,
because both are short names the operator sees side by side in fleet_profiles and fleet_list.
Gotcha. The check is deliberately narrow: it rejects only a value the bridge can prove is a
profile. A target missing from the member roster is still accepted, because it may be a peer lead's
terminal id or a herdr-owned pane the bridge did not spawn. So this catches one specific mistake well
and is not a general "is this target real" validation — an unreachable target still fails later, in the
injector. For the same reason profiles is a required parameter on both send methods rather than a
defaulted one: an overload that defaults it would silently turn the check off, which is exactly how
CB-561 disabled architect resolution while still compiling and passing its tests.
See free fleet capacity
What. fleet_list returns a top-level capacity block — for each profile its maxLoad, live,
free and reclaimable count — and adds idleForSeconds and reclaimable to every member row.
free is null when the profile is uncapped.
On. Always on; no configuration. The numbers come from profile maxLoad and the live roster.
Why. Free slots were invisible, so nothing told a lead when capacity was being wasted. A finished
member kept holding a terra slot until a later spawn was refused outright with worker profile 'terra' is at maxLoad: 2 live >= 2 cap; refusing spawn — no fallback to another profile. The bridge
already knew every fact needed to prevent that — the cap, the live count, and that the holder had
finished — and simply never reported them. This is a reporting gap, not a policy change.
Gotcha. reclaimable means only "this member holds capacity and has no open bridge work". It is
advisory. The bridge never spawns, stops or retasks a member to improve utilisation: it has
capacity facts but no work list, and choosing work needs authority it does not have. That decision
stays with the lead.
Two more things to know. live is read through the same liveCountRef function that placement
consumes, so an advertised free slot cannot drift from what fleet_spawn will actually accept —
a second count would eventually disagree, and a capacity view that lies is worse than none. And
lifecycle.idleTtlSeconds defaults to 1800, so a finished member holds its slot for 30 minutes
before the reaper takes it. That is far slower than slots turn over during active orchestration,
which is why the terra refusal happened even though a reaper exists — the view makes the wait
visible, it does not shorten it.
idleForSeconds derives from System.nanoTime(). It is monotonic within one daemon run and has no
wall-clock meaning across a restart.
Answer a question on an async delegation
What. When a member calls fleet_ask during a wait:false delegation, fleet_poll on that
ticket returns a non-terminal ASKING phase carrying the question text and its turnId. The lead
answers with fleet_send{turnId, content}; the member resumes the same turn and its real reply
still arrives on the original ticket.
On. Always on; no configuration.
Why. fleet_ask did not work at all on an async delegation. Outcome.QUESTION is deliberately
non-terminal, but the poll path tested "did this complete?", so a question fell into the failure
branch: the ticket was marked FAILED, and the question text and turnId were both discarded. The
member blocked for 55 seconds, gave up, and had to abandon its task and report the ambiguity in its
final reply instead. That is the mechanism working backwards — fleet_ask exists so a member can
resolve a decision without losing its turn. It also made the guidance self-contradictory: leads are
told to prefer wait:false for anything non-trivial and to answer an ask with
fleet_send{turnId, content}, and both could not be followed at once.
Gotcha. The ask still times out after 55 seconds by default (115s maximum). Those caps are not a
policy choice and raising them does not help: they exist because the member's own MCP client call
would time out, so a longer wait just moves the failure. A lead has to be reachable inside that
window. An unanswered ask returns the ticket to PENDING, not FAILED, because only the question
wait ended — the delegated turn keeps running.
Learn why a delegation died
What. When a member's queue is dropped — its pane is gone, or it was stopped — the send fails with
the real cause carried up from herdr, for example herdr error [agent_not_found]: agent target term_… not found. TurnListener.onTurnFailed takes that reason and CompletionResolver prefers it
over the pane scrape and over the old fixed text.
On. Always on; no configuration.
Why. The bridge already knew the exact cause and threw it away. A real log pair from a member being
stopped shows both halves two milliseconds apart: the injector logged dropping its queue … cause: herdr error [agent_not_found], while the sender was told only worker did not reply; its turn ended in an unrecoverable state (worker unreachable or stuck). Those two failures need opposite responses — a
dead pane means respawn, a stalled model means wait or kill — and the lead could not tell them apart.
Gotcha. drop now fires onTurnFailed unconditionally, where it used to fire only when a turn
was already in flight. That is the substantive half of the fix, not a tidy-up. A sender blocks on the
rendezvous waiter, never on the delivered future — delivered is only inspected to label a timeout
as queued or working. So completing delivered exceptionally never woke anybody, and a message that
was queued but not yet delivered sat until its timeout, which is 30 minutes on an async send.
Watch the fleet's health
What. An opt-in background observer. Each tick it takes one AgentControl.list() for the whole
fleet and one in-memory roster snapshot, joins them, and classifies every member. A member entering a
fault state logs one WARN; recovering logs one INFO. fleet_list reports healthCoverage, which is
off, detection-only, or full.
On. A health: block in fleetd.yaml with enabled: true. intervalSeconds defaults to 30 and
is floored at 15. With no block at all nothing is constructed and no herdr call is ever made.
Why. A fault is usually a disagreement between two views, not a value you can read from one of
them. A member that says BUSY in the session FSM while herdr says DONE has lost its turn boundary —
a real trace sat in that state for eighteen minutes with its ticket still PENDING and no fallback
firing. That is why both views must come from the same instant: one list call per tick, never one per
member. Reading panes is the exception, not the method.
Gotcha. Detection and notification are separate keys on purpose. An earlier design required a
webhook before health.enabled could be turned on, which would have removed real local detection to
avoid a narrower human-notification gap. So health runs with no sink configured and reports
detection-only — treat that value as "nobody will be paged", not as "health is off".
Two more. tick() catches Throwable and reschedules in a finally, because a
ScheduledExecutorService never re-runs a task that threw: the earlier version rescheduled as its last
statement, so the first agents.list() failure would have stopped health permanently and silently —
exactly when the control link is down, the highest-priority state in the model. And the snapshot fields
this unit cannot yet supply are the named constant NOT_YET_OBSERVED, not bare false, because to this
classifier false means "no fault" rather than "not known yet".
Keep a worktree that still holds work
What. Before a finished session's git worktree is deleted, fleetd checks whether it still holds uncommitted changes. If it does, the directory is kept and a WARN names its path, the pane and the release cause. A clean worktree is removed as before.
On. Always on, for every worktree-backed member. There is no knob.
Why. A member's uncommitted work exists in exactly one place — its worktree — so deleting it is
loss with no copy and no error. CB-544 already protected the shutdown drain for this reason, but left
the ordinary COMPLETED release deleting with --force. That gap fired: the idle reaper released two
members and deleted both worktrees, and only luck decided the work had already been pushed. A worker
that ends a turn without committing — because it stopped to ask a question, or refused the turn — is
the normal case, not the rare one.
Gotcha. The check is git status --porcelain with no --untracked-files=no, so an untracked
file counts as dirty. That is deliberate: the work at risk in the original incident was a new file that
was never git added, and ignoring untracked files would have missed exactly it. The cost is that a
profile whose parity overlay ever copies an untracked, non-gitignored file would make every release
preserve, and worktrees would pile up silently. Inert today — tracked overlay files carry
--skip-worktree so --porcelain cannot see them, and fleetd.yaml is gitignored — but it is a real
constraint on overlayParity, tracked in CB-581.
Second gotcha: hasUncommitted tolerates a worktree that is already gone and reports it clean. It has
to. It runs inside SessionManager.release() after the registry entry is dropped and before the
pane is stopped, so throwing there would orphan a live pane and strand a fleet_send caller on a
rendezvous nothing resolves. Anything added to that window needs the same tolerance.
Fail a ticket when its member dies
What. When fleet health sees a member reach a terminal state — GONE or NEVER_READY — every
ticket waiting on that member is failed straight away, naming the state as the reason, instead of
staying PENDING until something else notices.
On. The same health: block that turns on health watching. No separate
key.
Why. Detection without action just moves the silence. A lead that fires fleet_send{wait:false}
and polls its ticket gets pending forever when the member behind it is already gone — the failure is
known inside the daemon and invisible to the only caller who cares. Routing it through CB-568's
existing idempotent target-wide failure means the outcome is also counted, so a dead delegation stops
being invisible to /metrics.
Gotcha. It fires on the transition into the terminal state, not on every tick. An earlier
attempt put the call outside the transition guard, so a member that stayed GONE had the failure
operation invoked once per interval for as long as it remained in the roster; that commit was rejected.
The flip side is the honest limitation: the new state is recorded before the bounded retries run, so
if all three attempts throw, the tickets stay pending and no later tick retries. That path logs at WARN
and has its own test — it is a known edge, not an oversight. The retries also carry no backoff.
Stopping a member also fails its ticket now, even mid-fleet_ask (fleetd #275). There are two
ways a ticket loses its member, and they are not the same event. Health reporting GONE is a
guess read off the live agent list; fleet_stop or the idle reaper releasing a session is a
teardown the daemon performed, so it is certain. The sweep now takes a flag that says which one
it is. On a certain teardown it also fails a ticket parked in fleet_ask — the member is gone, so
nobody can ever answer that question — and closes the reverse rendezvous behind it. On a health
guess it still skips an asking ticket, because the member may be answerable by a live lead and a
guess must not kill it.
Gotcha for #275. The health path is deliberately left unable to self-heal one narrow case: a
member goes GONE while its session stays in the roster, its fleet_ask then lapses on its own,
and nothing re-fires the sweep — the transition already fired once, and FleetHealth.decide
returns GONE before it could ever return DELEGATION_ORPHANED. That is filed as fleetd #280, and
whether it is reachable at all depends on whether such a session is eventually released anyway,
which has not been checked.
Tell a usage-limit refusal from a real reply
What. A member can end its turn without calling fleet_reply. The bridge then scrapes the pane
and hands that text back as the answer. Sometimes that text is not an answer at all — it is the
backend refusing, because the account hit its usage limit. With this on, the bridge matches the scrape
against a pattern you configure. On a match it resolves the send as BACKEND_EXHAUSTED and carries the
matched line as the reason, instead of passing a refusal off as a completed reply.
On. Per profile, exhaustedPattern: — a regex. Opt-in: leave it out and that profile's completion
fallback behaves exactly as before. Startup logs one line naming which profiles have a pattern and
which do not, so you can see the coverage without reading the config by hand.
Why. The old behaviour lied in the worst direction. A lead asked for work, got back a block of text, and had no way to tell "here is your answer" from "my account is refusing to run". The lead would then treat a refusal as a result. Keeping the pattern in config, never in Java, is deliberate: every backend words its refusal differently, so a sentence baked into the code would only ever match one vendor.
Gotcha. The pattern map is built once at startup from the config snapshot, so exhaustedPattern
is a deferred key — adding one to a profile does nothing until the daemon restarts.
fleetd.example.yaml does not say this yet. Also, this stage only classifies. Nothing yet stops
the fleet spawning another member onto the same exhausted account, and nothing yet saves the work that
member was doing — those are stages B and C of CB-578.
Stop spawning onto an exhausted account
What. When a member's turn is classified as a usage-limit refusal (see
the entry above), the credential behind it is put
in quarantine for a cooldown. While it is quarantined, an explicit spawn onto it is refused with a
message naming the profile, the credential and roughly how many seconds are left; placement skips it
under every policy; and fleet_profiles shows it. The quarantine lifts itself — there is no manual
step.
Since fleetd #466 the cooldown escalates. Each consecutive exhaustion of the same credential doubles the wait, capped at 12x the base — about 6 hours at the 1800s default. A credential that keeps reporting exhausted is therefore retried roughly a dozen times a week instead of about 336 times.
On. quarantineCooldownSeconds: at the top level (default 1800, deferred — it is baked into
the tracker at startup). Since fleetd #466 it is the base of the backoff, not the whole of it. Per profile, credentialId: (hot) says which credential this profile
spends. Quarantine itself only ever fires for a profile that has an exhaustedPattern, so a fleet
with no patterns configured behaves exactly as before.
Why. Detecting the refusal was only half the problem. Without this, the fleet answers an exhausted account by spawning another member onto it, which fails the same way, and the operator sees a run of dead workers rather than one clear cause.
Gotcha. It quarantines the credential, not the profile name, and that distinction is the whole
point. In this fleet sol and terra are two different models billing one OpenAI account. Locking
only the profile that happened to report the refusal leaves its sibling live, and the next spawn walks
straight onto the same dead account under the other name. Profiles that share an account must share a
credentialId. A profile that sets none quarantines alone, under its own name — safe, but it will not
protect a sibling.
More gotchas, from the escalation (fleetd #466).
- The reset is a time proxy, not a success signal. Nothing in the daemon reports a successful spawn back to the quarantine tracker, so "the account started working again" cannot be observed there. What clears the streak is a base cooldown's worth of quiet — no further exhaustion report for that credential. That is the best available evidence, not proof. Read it as "we have not been told it is still broken", never as "it is fixed".
- The multiplier and the ceiling are constants, not config. Doubling (2.0) and the 12x cap live
in
BackendQuarantine, so there is no YAML knob for either and no new hot/cold question.quarantineCooldownSecondsstays deferred. - There is a ceiling on purpose. An unbounded backoff becomes a permanent outage that only a restart clears, which would be worse than the flat-rate retrying it replaced.
- Cooling-off is a different mechanism and is NOT escalated. A credential that throws repeated
non-exhaustion errors (an HTTP 5xx storm) gets
BackendOutagePolicy's flat 60s, with no repeat tracking. Escalating that would turn a transient storm into a multi-hour outage. The two states are reported separately and a profile can be in both at once.
See which charter a member got
What. Every member launch records a CharterReceipt: the role, where the charter came from, a
sha-256 digest of it, and its size in bytes. It is stored on the session and shown in the roster
(fleet_list and GET /members). The charter text itself is never recorded.
On. Automatic.
Why. Charters are per-role config and are re-read on every spawn, so two members of the same role
can get different text without anyone noticing. The digest answers "did this member actually get the
charter I think it got?" without printing prompt text into logs an operator may not be allowed to
keep. It also closed a real leak: the older pane-placement spawn log printed the whole argv, and the
charter travels inside argv. That argument is now replaced by its digest.
Gotcha. PeerHandle.charterReceipt() is deliberately not a default method. It used to be,
and OpenCodeLauncher's wrapping handle forgot to override it — so it answered null while the real
receipt sat on its delegate, and sol and terra silently showed no receipt at all while Claude Code
members showed one. Nothing failed; the roster field was just quietly missing. Removing the default
makes the compiler catch that, and any new adapter must now answer the question on purpose. If you add
a PeerHandle implementation, this is the line that will not let you skip it.
Redeploy the daemon safely
What. scripts/redeploy-fleetd.sh rebuilds the jar and restarts fleetd as one command. It
builds before it stops anything, waits for the old process to actually exit, restarts from a login
shell with cwd = fleetd/, then polls /healthz and reports the herdr protocol number, a fresh
fleetd listening line, the config keys accepted or deferred at boot, and any ERROR lines since
the restart. --check reports state and changes nothing; --yes skips the drain prompt;
--no-build restarts the jar already on disk.
On. Run it. Nothing is automatic — the daemon never restarts itself.
An agent gets there through the redeploy-fleetd skill (.claude/skills/redeploy-fleetd/). It is
a primary-side skill, and it holds the flags, the drain step, the operator's allow-list entry, and five
numbered checks — login shell, drain members, deferred config keys, re-check fleet_whoami, prove the
new jar runs — each of which has gone wrong here before. Workers must never load it: stopping the
daemon kills the worker's own channel mid-turn. CLAUDE.md used to carry all 58 lines, and paid for
them in every session's context. It now keeps only the two rules that must stay resident — a merge is
not a deployment, and workers never redeploy — plus the line that names the skill.
Why. A merge is not a deployment: the running daemon holds the jar it was started with, so merged
code does nothing until this runs. That gap has silently shipped inert features more than once — the
whole health: stack sat merged and doing nothing for two tickets. The script also exists so the
operator can allow-list one auditable command instead of approving a kill and a java -jar
separately every time, which is what a lead would otherwise have to ask for on every deploy.
Gotcha. The check that matters most has no log line anywhere in the daemon: fleetd inherits
WORKER_GITEA_TOKEN from the shell that starts it, via ${SHARED_ENV}/tools/secrets.sh. Start it
from a non-login shell and the variable is empty — the daemon boots normally, /healthz is green,
and the failure surfaces much later as workers that cannot open a PR. --check is the only thing
that reports this, and it tests whether the name resolves without ever printing the value.
Second gotcha: a green /healthz only proves herdr answers. If herdr's protocol number has moved,
every spawn can still fail. The script prints the protocol it saw so you can compare it; prove a real
spawn before trusting the fleet.
The script never escalates to kill -9. The shutdown hook releases sessions and worktrees in order,
and a hard kill can leave worktrees and panes behind; if the process will not exit it stops and tells
you rather than forcing it.
Get told when an async delegation finishes
What. A fleet_send{wait:false} ticket that reaches a terminal phase — replied, failed, wedged,
or abandoned — now injects a short nudge into the lead's own pane telling it to poll. Several
tickets finishing at once coalesce into one nudge naming the count. Polling a ticket marks it
collected, so a ticket you already read is never nudged about again.
On. Automatic when the lead is a herdr pane. Bounded by fleet.leaders.<name>.push_reminders
(default 5) and push_backoff_ms (default 15000). A lead that is not a herdr pane leaves the
registry empty, the loop becomes a no-op, and delivery degrades to pull — nothing is lost.
Since CB-590 there is one nudge schedule per lead, shared with the CB-307 reply nudge, so the two
can no longer inject into the same pane at once. push_reminders is a budget per source, not one
shared counter: reply work and ticket work each get their own, so a busy reply stream cannot spend the
budget a ticket needs. Worst case a lead sees up to twice push_reminders nudges, which is the
deliberate price of that isolation.
Why. The charter tells leads to prefer wait:false for anything non-trivial, because a blocking
fleet_send is capped by the caller's own MCP client timeout of about 60 seconds. But until this
landed, that preferred mode was the one mode with no notification at all: MessageService.reply
returns on the rendezvous fast path before the push loop hears anything, so an async ticket finished
in silence and the lead only found out by polling on a hunch.
Gotcha. The nudge goes to the lead's pane, not to its MCP session. A lead driving the REST
surface directly — which is exactly what you fall back to when the MCP mount drops — receives nothing.
Combine that with the 10-minute terminal-ticket TTL (MessageService.TICKET_TTL_NANOS) and a finished
worker's report can be pruned before it is ever read. That happened during this feature's own
close-out: a ticket returned 404 while its member still sat in done. The work survived only because
the implementer skill had opened a PR, and the PR body carried the report. The ticket is not the
durable artefact; the PR is. If you orchestrate over REST, poll on a timer.
Get told when a worker is waiting on your answer
What. A worker on an async (wait:false) delegation that pauses mid-turn in fleet_ask now
nudges the lead's pane by itself, naming the exact call that resumes it:
Worker term_a asked a question (ticket task-3) — answer it with
fleet_send(turnId="term_a#1", content=...) to resume its turn:
which config file?
fleet_status{sessionId} shows the same open question, and so does REST
GET /sessions/{id}/status (as question, turnId, ticket).
On. Automatic, on the same terms as the ticket nudge above — it is a third source in the same
per-lead schedule, not a new push path, so CB-590's one-schedule-per-lead guarantee still holds and
it spends from its own push_reminders budget.
Why. fleet_ask opens a reverse-rendezvous window of about 55 seconds. A lead polling on its
normal cadence of minutes never saw it, so the worker timed out and carried on without an answer —
the ask was, in practice, unusable on the delegation mode the charter tells leads to prefer. Two
smaller holes closed with it: REST GET /tasks/{ticket} dropped turnId on an ASKING phase, so a
REST caller could read the question and had no way to answer it, and fleet_status said nothing
about an open question at all.
Gotcha. This closes the window; it does not remove it. The worker still gets ~55 seconds, and a nudge only helps a lead that is injectable right now — a lead mid-turn for a minute still misses it. So the standing advice is unchanged: do not brief a worker to "ask me." Decide the question before you delegate, or give the worker an explicit default to use.
Why the window is not simply widened: DEFAULT_ASK_TIMEOUT_MS = 55_000 sits just under the worker's
own MCP client cap of about 60 seconds, so the daemon can return a clean typed timeout before the
client severs the call. Raising the server constant buys nothing — the worker's client kills the call
regardless.
Members cannot use the operator's admin forge token
What. Every member launch overwrites GITEA_ACCESS_TOKEN with a non-blank blocked sentinel, and
sets a BRIDGED_MEMBER=1 marker. Applied in HerdrPeerLauncher.baseEnv after the profile's own
env: map, so no profile — present or future — can name that key and restore the real value. CB-302's
separate GITEA_TOKEN grant is untouched, so a member can still push and open its own PR with a
repo-scoped token.
On. Automatic, both adapters, every profile.
Why. A member is an autonomous agent running arbitrary tool calls. It has no business holding the
operator's admin credential, and the rule the operator set is explicit: leads and architects may use
GITEA_ACCESS_TOKEN, members use WORKER_GITEA_TOKEN. A plain env overlay was not enough — the
pane's login shell re-sources the secret store afterwards and would put the real value back — which is
why the marker exists: the shell's export is guarded on BRIDGED_MEMBER being unset.
Gotcha, and it is the important one. This blocks one name. The pane's login shell sources a
secret store exporting about thirty, and a second forge token is among the ones nothing blocks.
Tracked as CB-596. The reason nobody noticed is worth internalising: a Claude Code member inherits
the operator's user-scope ~/.claude.json servers, so it mounts 45 mcp__gitea__* tools including
delete_branch and delete_file. They all fail — but only because the credential they read is
blocked. Remove the block and the same picture becomes 45 working destructive tools, with no visible
difference beforehand. Present-and-useless looks identical to absent.
Supervise the daemon without breaking the fleet
What. A launchd unit that restarts fleetd if it dies, and that still gets the fleet's secrets.
deploy/dev.ltms.fleetd.plist runs scripts/fleetd-launchd-wrapper.sh, which execs one login shell
in place (exec /bin/zsh -lc 'exec "$@"' -- "$@") and then execs the real java command. One exec
chain, so launchd keeps tracking the right PID. scripts/redeploy-fleetd.sh detects whether the
agent is loaded and switches stop/start to launchctl unload -w / load -w, falling back to its
original kill + nohup when it is not.
On. Not automatic, and deliberately so. Install it yourself:
cp deploy/dev.ltms.fleetd.plist ~/Library/LaunchAgents/
launchctl load -w ~/Library/LaunchAgents/dev.ltms.fleetd.plist
launchctl list | grep fleetd
scripts/redeploy-fleetd.sh --check reports whether the agent is installed and whether it is loaded.
It is read-only.
Why. The unit shipped with CB-504 was never installed, and could not have worked if it were.
launchd does not source a login shell, so a launchd-started daemon would have had no
WORKER_GITEA_TOKEN and no AI_GATEWAY_TOKEN. It would have started fine and looked healthy; the
failure would have appeared hours later as workers unable to open a PR. So the operator's real choice
was a supervised daemon with a broken fleet, or a working fleet with no supervision. The wrapper
removes that choice. The launchctl branch in the redeploy script matters just as much: a SIGTERMed
daemon exits 143 even when its shutdown hook completes normally, so under
KeepAlive{SuccessfulExit: false} a bare kill makes launchd restart the old jar, racing the
script's own restart.
Gotcha. The script computes its log path from where the script file sits; the plist hard-codes an
absolute StandardOutPath. Nothing checks that the two agree. If they ever diverge — a worktree, a
renamed clone — the script's post-restart ERROR check reads the wrong file, finds nothing, and reports
"ok" while the daemon crash-loops. And the loop really is unbounded: ThrottleInterval: 10 paces
restarts to one per ten seconds, it does not cap how many. Fix CB-600 before installing the agent.
See at startup which secrets the daemon actually got
What. fleetd logs, at startup, every secret environment variable it needs, and whether each one
resolved or is MISSING. The required set is derived from the loaded config — each non-subscription
profile's tokenEnv, plus every profile's gitTokenEnv — not hard-coded, so a new profile is covered
the day it is added.
On. Automatic. Read the startup secret … lines at the top of fleetd/fleetd.out.
Why. An empty token used to be completely invisible. The daemon started, /healthz went green,
and the first sign of trouble came much later and somewhere else — a worker that could not open a PR,
or a gateway profile that could not authenticate. Neither symptom points back at the shell the daemon
was started from, which is the actual cause. This turns a silent, delayed, misattributed failure into
one line at startup.
Gotcha. It reports names and set/MISSING only — never a value, a prefix, or a length. That is
deliberate and must stay that way; the log is not a secret store. Also, MISSING is a warning, not a
refusal: the daemon starts anyway, because refusing to boot over a credential that half the fleet may
not need would be worse. So the line has to actually be read. One known false alarm: a non-subscription
profile that never sets tokenEnv inherits the default name FLEETD_WORKER_TOKEN and is reported
missing — which is nearly always a real misconfiguration rather than a bug in the report.
Bound the AMQP backlog, and know a reply was really published
What. Two guarantees on the durable reply inbox. The consumer calls basicQos before
basicConsume, so unacked messages beyond the window stay on the queue instead of being pushed
into the daemon's heap. And publishing runs on its own confirm-mode channel with mandatory=true and
a return listener, so an unroutable or unconfirmed publish raises an error instead of vanishing.
On. broker.prefetch sets the window, default 32. The confirm behaviour is automatic whenever
broker: is configured at all.
Why. Both existed to make the word "durable" true. Without prefetch the queue sat near-empty while
the real backlog lived in an in-memory map with nothing capping it — so queue-depth metrics read
healthy, and any queue-level limit would have guarded an empty queue. Without confirms, a publish to a
queue that was never declared was a silent black hole, and deliveryMode(2) bought nothing, because
"persisted" is only true after the broker says so.
Gotcha. A broker Return always arrives before its matching Confirm, so an ack alone does
not mean routed — the confirm path has to check a per-message returned flag, and that ordering is easy
to get wrong when editing this code. Note also that these tests are @Tag("contract"): they are
excluded from the default build and run in a separate CI job against a real rabbitmq:3.13
container. A green default build says nothing about them. And the broker here is LavinMQ; the
RabbitMQ client library is used only because LavinMQ speaks the same protocol.
Three config keys that changed meaning — read this before upgrading
What. weight, maxLoad and fleet.leaders.* all changed what they mean, not just what they
do. A config file that worked before an upgrade can behave differently, or refuse to start, with no
edit to it. The three are grouped here because an upgrading operator meets them together.
| Key | Used to mean | Now means |
|---|---|---|
profiles.<n>.weight: 0 |
coerced to 1.0 — "pick me as often as anyone else" | excluded from automatic selection |
profiles.<n>.maxLoad: 0 |
coerced to unlimited | capped at zero live members |
fleet.leaders.<n>.terminal |
a terminal-id pin | rejected at startup — use tab |
On. Nothing to switch on. These are the meanings now.
Why. In the first two cases the old behaviour was the exact opposite of what the key reads like.
weight: 0 looked like "never pick this" and meant "pick it normally". maxLoad: 0 looked like
"never run anything here" and meant "unlimited" — and that was the only throttle a
subscription: true profile had against the operator's own paid plan. A key whose plain reading
inverts its behaviour is a trap, so both were fixed toward the reading. A negative maxLoad is now a
startup error rather than being silently normalised away.
fleet.leaders.*.terminal went further and is refused, naming the offending entries, instead of
being ignored. A silently-ignored lead pin means the lead is not recognised, gets demoted to worker,
and every orchestration call is refused — a failure that looks nothing like its cause. Failing at
startup is the kinder outcome.
Gotcha. weight: 0 excludes a profile from automatic selection only. An explicit
fleet_spawn{profile:"..."} bypasses placement entirely and still resolves it, so weight: 0 is
not a way to disable a profile. maxLoad: 0 is, because that cap now applies to explicit spawns
too. And a lead's tab must match a real herdr tab label exactly (case-insensitively) — the old
tabPrefix no longer finds a lead, it only guards against a worker's label colliding with the
convention.
Stop an explicit spawn from busting the cap
What. maxLoad is now an unconditional cap. An explicit fleet_spawn{profile:"X"} used to skip
the check entirely — the cap applied only to automatic placement — so naming a profile was a way
around it. Now an explicit spawn is refused when live >= cap, and it is not re-routed to
another profile.
On. profiles.<name>.maxLoad.
Why. A cap that any caller can opt out of by naming the profile is not a cap. The no-fallback part is deliberate: a caller who named a profile did so for a cost or model reason, so quietly moving the work to a different backend would defeat the reason they named it. Better to refuse and let the caller decide.
Gotcha. There is a known TOCTOU race: liveCount is read outside a lock, so two genuinely
concurrent spawns can both pass the check and briefly exceed the cap. It is documented in the code
rather than fixed. Also note the refusal message only became visible to callers in CB-599 — before
that it was a blank 500 with the reason in the log.
Nudge an idle lead back to work
What. An opt-in loop that nudges the one idle lead after it has been continuously injectable for a quiet period with nothing driving it. It fills the gap the reply push loop leaves: that loop only fires when a worker reply lands, so a lead that is simply sitting idle with nothing arriving is invisible to it.
On. The leadHeartbeat: block — idleAfterSeconds (default 300), backoffMs (default 60000),
quietNudgeCap (default 3). Absent means off, and the loop is not even constructed. Changing it
needs a restart.
Why. Off by default, deliberately. This loop spends the operator's model subscription on the daemon's own initiative, so it must never switch itself on during an upgrade. That constraint shaped the whole design.
Gotcha. It is status-gated (a WORKING lead is never touched), debounced, caps consecutive quiet
nudges, and stands aside while the reply push loop is already nudging. That last one depends on
ReplyPushLoop.isActive() being honest — see CB-598, where a schedule could die with work still
pending and report inactive.
Prove which charter a member got, without logging the prose
What. Per-role launch charters are read from the live config and composed once per spawn: the
role charter plus the reply charter. Claude Code gets it as an appended system prompt; opencode gets
an ephemeral member-charter.md referenced from its generated config. A CharterReceipt records a
sha256 and byte count of the exact composed bytes, and rides on both the spawn log and the roster.
On. fleet.charters.<role>, where role is dev, architect or reviewer. Read live per spawn,
so a change applies to the next spawn with no restart.
Why. Two problems at once. Charters had been hardcoded, which meant changing what a role is told required a rebuild. And an operator asking "what was this member actually told?" had no answer that did not involve printing the prose into a log, where it does not belong. The receipt answers the question with a digest instead.
Gotcha. The role charter is combined with the reply charter only when the profile mounts the
bridge MCP. A non-MCP profile gets the role charter alone — which is correct, since the reply
charter tells a member to call fleet_reply, and a member with no bridge cannot.
Tell "busy" from "refusing" in the capacity view
What. In fleet_list's capacity view, a quarantined profile's row is forced to free: 0
whatever its maxLoad and live counts say, and gains credentialId and quarantinedForSeconds.
On. No new knob. The quarantine facts come from the existing exhaustedPattern, credentialId
and quarantineCooldownSeconds keys.
Why. A quarantined profile used to show free slots it would refuse to fill. A lead reading that row would keep trying and keep failing. The two states need different reactions — "busy, will free up" means wait, "refusing for N seconds" means go elsewhere — and the row could not express the difference.
Gotcha. The two new fields appear only when a profile is actually quarantined, so an ordinary
fleet's rows are byte-identical to before. Do not write a client that expects them. The view shares
its QuarantineSource with fleet_profiles so the two surfaces cannot disagree.
The row also says which attempt this is, not only how long is left (fleetd #466, #473). Both
fleet_profiles and fleet_list's capacity rows carry quarantineAttempt beside
quarantinedForSeconds: 1 for a first refusal, 2 for the second in a row, and so on. The
cooldown escalates with that count, so quarantinedForSeconds: 1800, quarantineAttempt: 1 and
quarantinedForSeconds: 7200, quarantineAttempt: 3 mean very different things about the credential.
Why it exists: the seconds alone cannot tell a weekly subscription limit apart from a one-off
capacity blip. A lead seeing attempt 4 knows retrying is pointless and should move the work to
another credential, or tell the operator. Gotcha: the count keeps growing past the cooldown
ceiling, so a large quarantineAttempt with a flat quarantinedForSeconds is normal, not a bug —
the ceiling caps the wait, not the streak. A quiet gap long enough to clear the quarantine resets
the count to 1. The flat two-argument BackendQuarantine constructor still reports a real growing
count even though its own cooldown never escalates, because the streak is a fact either way. Both
values come from one status() call that does a single map read, so the reported count can
never disagree with the cooldown it describes.
Resume a member onto its previous conversation
What. A member's own agent-session id is captured at spawn, survives onto the roster as
agentSessionId in fleet_list and GET /members, and fleet_spawn accepts sessionName and
resumeSessionId to relaunch onto that same conversation.
On. fleet_spawn{sessionName, resumeSessionId}.
Why. The pieces existed but were connected at neither end: nothing persisted the id, and nothing exposed a way to pass it back. So a member that died took its context with it even though the backend could have resumed it.
Gotcha. A resumeSessionId requires an explicit profile, and that profile's adapter must
declare Capability.SESSION_RESUME. A placement-routed spawn carrying a resumeSessionId is
refused, because a resumed conversation is tied to the specific backend that started it. If the
adapter lacks the capability the spawn is refused naming it, rather than silently starting a cold
session that looks resumed.
opencode members run with --auto and forced auto-compaction
What. opencode workers launch with --auto, which auto-approves the permissions opencode does
not explicitly deny, and the bridge-generated config pins compaction.auto: true.
On. Neither is configurable. Both are unconditional for the opencode kind.
Why. A worker that stops to ask for permission on a tool call has no one to ask — the lead is not
watching its pane, and an unanswered prompt burns the whole turn. Auto-compaction is the same
argument: a worker that runs out of context dies mid-turn and loses its fleet_reply, so its report
is gone even though the work was done.
Gotcha, and it is a real trade. opencode's own help calls --auto "dangerous!". The blast radius
is bounded by everything else about a member — its own worktree, its own branch, off-subscription,
and it cannot merge — not by the flag. And compaction.auto: true in the generated config wins
over the operator's home config, so auto-compaction cannot be turned off for fleetd workers from
there. That was chosen knowingly: a lost report was judged worse than an unwanted compaction.
Silent member failures now say why
What. Logging only, no behaviour change. The paths that used to fail a member with no log, a bare DEBUG, or a message naming only the symptom now log at WARN with the real cause and the numbers — across the injector, the completion resolver, session acquire/reap/failure, and ticket abandonment.
On. Automatic.
Why. A member could fail and leave nothing to read. The lead saw a session that stopped responding and had no way to tell a crashed backend from a wedged pane from a reaped session.
Gotcha. You will see more WARN lines than before if you filter by level. That is the feature, not noise — but it will change what a log-volume alert sees.
Sessions are one-shot — recycle() is gone
What. SessionManager.recycle() — release plus immediate re-acquire under a new pane id — was
removed. A released or finished session is torn down, never pooled or reused. Callers must release
then acquire explicitly.
On. Nothing to configure; this is a removal.
Why. Recycling quietly reused a session across two unrelated pieces of work. The lifecycle is meant to be strictly one-shot, and a method that bypassed it was a standing invitation to leak one delegation's context into the next.
Gotcha. None left in-tree — no references remain. It had no external surface, so only internal callers were ever affected.
A failed ticket names the conversation, not just the files
What. When a member dies, the failed ticket's detail now carries the member's agentSessionId
alongside the worktree, branch and snapshot ref it already reported:
… worktree=/wt/x branch=worker/x snapshot=refs/wip/x agentSessionId=abc-123
A lead can pass that id back as fleet_spawn{resumeSessionId: "abc-123"} on the same profile.
On. Automatic, no knob. It appears whenever the released member's backend produced a session id.
Why. CB-578 stage C already lets a lead re-dispatch a fresh member onto the same worktree after a failure, and it states its own limit plainly: the files survive, the reasoning does not. A re-dispatch was a cold start re-reading a brief. Pairing the session id with the worktree is the other half — it turns "the work survives" into "the work and the thread both survive".
Gotcha. It is absent whenever the adapter never resolved an id: the member was spawned without
sessionName or resumeSessionId (the ordinary path), or its backend does not declare
Capability.SESSION_RESUME. Silence here is the expected case, not a fault — do not read a missing
agentSessionId= as a bug.
weighted placement is not "cheapest first"
What. placement: weighted is smooth weighted round-robin. It spreads unqualified spawns across
every profile that has a free slot, in weight ratio. It has no concept of cost. So local: 10
against terra: 2 and sonnet: 1 does not mean "use local, overflow to paid" — it means roughly a
quarter of spawns go to a paid profile while the free box still has a free slot.
On. placement: weighted in fleetd.yaml, with a weight: per profile. weight: 0 excludes a
profile from automatic placement entirely; an explicit fleet_spawn{profile:...} bypasses placement
either way.
Why. The operator's rule is cheapest-first: keep the free boxes busy and pay only for genuine
overflow. No shipped policy expresses that. FixedPlacementPolicy ignores maxLoad and throws
instead of overflowing; RoundRobinPlacementPolicy has no cost notion. CB-589 tracks a real
cost-first policy. Until then the workaround is to make the ratio decisive rather than
proportional — on this host local.weight is 100 against paid weights of about 1, so the free
box wins every pick it is eligible for.
Gotcha, two of them. First, the running score map lives for the daemon's whole life. While a
profile sits at maxLoad it is filtered out and its score freezes, so paid profiles keep
accumulating against it; when the free slot opens the profile returns with a stale score and can
lose the next pick — a paid spawn while the free box is idle. Second, and worse for the next
person: the workaround expresses a preference order through a ratio knob. Add a profile at
weight 150 later and it silently outranks the free box, with nothing to warn you. fleetd.yaml is
gitignored, so a fresh host starts without this workaround and quietly pays — the reasoning is
written into fleetd.example.yaml next to the key for exactly that reason.
Which inherited credentials a member may keep
What. A member's pane runs a login shell, which re-sources the operator's secret store and
exports about thirty names. Exactly one of them is blocked: GITEA_ACCESS_TOKEN, see Members cannot
use the operator's admin forge token above. This entry records the decision about the two that
matter most among the rest, so that "we chose to allow it" never again looks the same as "we never
noticed".
| Credential | A member may hold it | Why |
|---|---|---|
AI_GATEWAY_TOKEN |
yes | It is the key a member is meant to use. Paid-backend members reach their model through llm.ltms.dev, and that token is the single front-door key. Blocking it would stop those members working at all. |
CONTEXT7_TOKEN |
yes | It backs the context7 documentation MCP, a read-only docs lookup that members are meant to have. The worst case is documentation reads on the operator's quota. |
GITEA_ACCESS_TOKEN |
no | Admin scope on the forge. Blocked — see the entry above. Members push with the repo-scoped WORKER_GITEA_TOKEN instead. |
On. Nothing to turn on. The two allowed tokens flow through by default; the blocked one is overwritten at every launch.
Why this is written down at all. BRIDGED_MEMBER already exists, so splitting either of these
per-role is now a one-line guard in the secret store. The option is cheap and available — which is
exactly why the decision has to be explicit rather than implied by nobody having done it. The
original leak (CB-592) was found by accident. The same accident should not have to happen twice.
Gotcha. This decision covers two names out of about thirty. The rest were enumerated later by CB-596 — see A member keeps only the credentials you name below, which supersedes this entry's scope. Read this table as the two that were decided first, not as the whole policy.
A member keeps only the credentials you name
What. A memberCredentials: block in fleetd.yaml lists every credential-shaped variable on
the host, says which ones a member may keep, and blocks the rest. A blocked name is not unset — it is
overwritten with a fixed sentinel string, blocked-by-fleetd-cb596-see-gitea-issue-82, so a member
that reads it sees "deliberately blocked" rather than an empty variable it might quietly work
around. Before this, a member pane inherited the operator's whole secret store and exactly one
name was blocked.
memberCredentials:
policy: deny-by-default # the only policy today; named so a future allow-by-default is a change
allow: # names a member MAY keep
- AI_GATEWAY_TOKEN
- WORKER_GITEA_TOKEN
- CONTEXT7_TOKEN
known: # every credential-shaped name on this host
- GITEA_ACCESS_TOKEN
- ...
The blocked set is known minus allow, computed at load. A name in both is an error you cannot
make by accident — the intersection is empty by construction, because allow wins.
On. Add the block to fleetd.yaml (there is a commented template in fleetd.example.yaml) and
add the matching guarded export to the operator's secret store. Both halves are needed — see the
gotcha. The block is re-read on every spawn, so editing it takes effect without a restart; only the
startup summary line needs one.
Why it exists. CB-592 found that a member inherits the operator's credentials, and blocked one
name — GITEA_ACCESS_TOKEN — with a hardcoded string in the launcher. CB-593 then decided two more
by hand. That does not scale and, worse, it hides the shape of the problem: a hardcoded list of one
looks finished. Naming every variable in config turns "which secrets does a member hold?" from a
question nobody can answer into a list you can read, review and diff.
The daemon says so when the block is missing. With no memberCredentials: the daemon still
starts — refusing to boot would strand an operator who has not migrated — but logs a WARN saying
every member pane inherits the whole secret store unblocked. With the block present it logs the
counts instead:
memberCredentials: 34 known name(s), 5 allowed — blocking 29 on every spawn
It also reports what you forgot. On every spawn the launcher scans the host environment for
names shaped like credentials (TOKEN, SECRET, _KEY, APIKEY, PASSWORD, CREDENTIAL,
AUTH) and reports any that are on neither list. This is the part that pays for itself: the first
spawn after it deployed named two variables no hand-written list had ever contained, because neither
lives in the secret store —
WARN memberCredentials gap: 2 credential-shaped env var name(s) are on neither known: nor allow:
— every member pane inherits them UNBLOCKED — [CLAUDE_CODE_MESSAGING_TOKEN, SSH_AUTH_SOCK]
Since 2026-08-31 the report distinguishes three cases, because one wording was lying. A name
being on neither list does not by itself mean a member inherits it — under allow-list on a zsh
login shell the generated scrub blanks it anyway. The severity now follows what the scrub actually
does, decided with the same predicate the scrub itself evaluates:
| Situation | Line | Meaning |
|---|---|---|
deny-by-default, or allow-list on a non-zsh shell |
WARN … inherits them UNBLOCKED |
Nothing scrubs these. Real exposure. |
allow-list + zsh, name not on the derived allow-list |
INFO … the scrub blanks them anyway |
Contained. Listing it just makes that explicit. |
allow-list + zsh, name kept by the derived allow-list |
WARN … the derived allow-list keeps them anyway |
Real exposure, and easy to miss. |
That last row is the one worth knowing about. The derived allow-list is a superset of
known + allow: it also picks up every profile's gitTokenEnv, gitHostEnv, tokenEnv and
env: keys. So naming a variable in a profile silently grants it to members, whether or not it is on
either list — and until this change the daemon reported that case as contained. Each report kind
has its own once-per-daemon guard, so a benign INFO can no longer suppress a serious WARN.
SSH_AUTH_SOCK is now blocked, and that is a downgrade, not a win. It was once allowed on
purpose, because a worktree's remote was ssh://git@git.ltms.dev and a member could not push
without the agent socket. Members now push over HTTPS with WORKER_GITEA_TOKEN, so the live config
sets sshAgentEnv: omit (spelled sshAuthSock: block before #266; both still parse), and
MemberEnvAllowList.derive drops the name even if an operator lists it under allow: — the config
cannot re-grant it by accident.
Do not read that as the problem being solved. #184 measured a member with the socket blanked pushing
fine anyway: the forge key is a readable, passphrase-free file, and ssh -G finds it outside
~/.ssh. An environment control cannot remove a file. Blocking the socket closes one door in a room
with another door open; the actual fix is a separate OS user (#185).
Gotcha — under deny-by-default the config half alone does not hold. The launcher writes the
member's environment at spawn, and then the member's pane runs a login shell, which re-sources
the operator's secret store and overwrites it. So deny-by-default must be paired with a
BRIDGED_MEMBER-guarded block at the end of the secret store that re-applies the same sentinel. Two
lists that must agree — which is why the gap detector exists, and why the daemon logs its counts at
startup. If a member ever reports holding a name you blocked, the secret-store half is what is
missing.
That block must be the last thing in the secret store. It overwrites the blocked names, so anything that re-exports them afterwards silently undoes it.
policy: allow-list removes that whole problem (CB-633). It is the setting to prefer.
memberCredentials:
policy: allow-list # default is "deny-by-default"
sshAgentEnv: omit # "inherit" only if a member must use the operator's ssh-agent
Why it exists. A control that lives inside a sourced file can always be undone by a file sourced
later, and that is not a hypothetical: on this host .ltms sources mgnlSecrets.sh one line after
the guarded secrets.sh, so seven credentials reached every member in full — including an AWS key
with AdministratorAccess. Four of the seven were already on the block list. The list was correct
and it still failed. Making the list longer fixes nothing.
How it works. The daemon generates a throwaway ZDOTDIR directory per spawn and passes it in the
pane-creation env map, before the shell starts. Each generated startup file sources its $HOME
counterpart first and then runs the scrub, so the scrub happens after the operator's whole chain
and nothing sourced later can undo it. The kept-name set is derived, never typed: every profile's
tokenEnv/gitTokenEnv/gitHostEnv values and env: keys, plus an infrastructure set (PATH
HOME SHELL TERM LANG LC_* TMPDIR USER LOGNAME PWD SHLVL EDITOR PAGER
ZDOTDIR JAVA_HOME XDG_*), plus the exact keys this spawn's own env overlay carries. Adding a
profile can only widen the set, so it can never break another spawn's scrub. Under this policy
known:/allow: stop being a control and become reporting only — they still feed the gap warning.
Gotcha — it is zsh only. ZDOTDIR means nothing to bash. If the member's shell is not zsh the
daemon logs a loud WARN saying protection is off and falls back to the deny-by-default overlay.
That is weaker, so put members on a zsh account.
Gotcha — the platform nearly made this a dead control. The scrub first lived in the generated
.zlogin, and zsh reads .zlogin only for a login shell. herdr does not open the same kind of
shell everywhere: a macOS pane runs -zsh (login), a Linux pane runs a plain /usr/bin/zsh
(interactive, not login). So it protected the developer's Mac and would have protected nothing at all
on Linux, with no error anywhere. The scrub now runs from the generated .zshrc, .zlogin and
.zshenv (#388) — .zshenv is the only file zsh reads unconditionally, so the control no longer
depends on enumerating which kinds of shell exist. If you port this to another terminal backend,
check what kind of shell it opens before trusting it.
Do not infer the shell kind from argv[0]. A bare /usr/bin/zsh proves not login and says
nothing about interactive. Both this repo's javadoc and a host's fleetd.yaml once claimed
"interactive but NOT login" on that evidence alone, and neither had measured it. Worse, a probe run
from inside an agent measures the agent's own zsh -c child, not the pane — opencode forks a
fresh non-interactive shell per command, so it reports truthfully about the wrong process. Read the
pane shell's own /proc/<pid>/environ, or replicate the shape with script -qec zsh /dev/null.
How to tell it actually ran. Each pane writes a scrub-report.txt, and the daemon logs
allowed N of M environment variables when that pane stops — the denominator is the point. A
missing report is logged at WARN: the scrub then cannot be confirmed to have run at all, and a
silently dead control is the failure this policy exists to remove.
Gotcha — a missing report has more than one meaning, and reading it as one cost a day (#394).
The report block is the last statement in scrub.zsh, so its absence proves only that the script
did not reach the end. It does not tell you which shell ran. On one host the missing report was
read as evidence that the pane shell was neither login nor interactive; the real cause was that
export UID= in zsh is a fatal parameter error which terminates the whole sourced file. The
blanking loop is wrapped in { ... } 2>/dev/null, so the message was swallowed too.
The severity was the selection, not the count. env lists inherited names first and a startup
file's own exports last, so the loop blanked the harmless inherited half and died immediately
before the operator's own exports — the credentials the policy exists to remove. Measured there:
UID was name 42 of 57, and a ~/.zshrc decoy at 58 survived on 8 of 8 spawns. A partial scrub
got precisely the wrong half.
Fixed by routing each attempt through eval "export ${n}=" 2>/dev/null, which contains the error
to one iteration, rather than by skipping the known-fatal names (UID EUID GID EGID PPID LINENO).
A skip-list has to be complete forever; this is a security control, so it must not depend on an
enumeration being right. The report's first line now reads allowed N of M failed F, unblankable
names are listed with a ! prefix, and the daemon WARNs naming them — so "could not blank this
one" is now visible instead of being an abort you learn about from a missing file.
Two things to plant when you audit this yourself. First, a decoy: a non-secret variable the
operator's chain exports that is not on the allow-list is the only thing that separates "scrub
skipped" from "nothing was there to scrub" — allow-listed names come back blank either way, so
their blankness proves nothing. Second, check for .zcompdump in the pane's ZDOTDIR: its
presence proves an interactive zsh ran compinit there. A .zcompdump present and
scrub-report.txt absent is only explained by "sourced, then aborted partway".
Measured, not assumed. A real login zsh started from a clean parent kept 15 of 78 names with
zero profiles configured; an earlier prototype run kept 3 of 28. The test asserts equality between
the survivors and baseline ∩ derived allow-list, not a spot check of a few blocked names.
Verified, not assumed. Measured on 2026-08-17 inside a live member pane, once both halves were in
place: 29 of 29 blocked names hold the sentinel, and the allow-listed names present in that pane
keep their real values. The operator's own login shell is unchanged, because the whole block sits
inside if [ -n "${BRIDGED_MEMBER:-}" ]. To repeat the check, spawn a member and run
scripts/probe-member-credentials.sh; it refuses to run anywhere but in a member.
Third gotcha — the probe has the same defect it is checking for. It carries its own hardcoded name list and does not read the config, so it under-reported by three names and still exited 0. Read its count against the daemon's startup line rather than trusting the table. Tracked as CB-608.
Second gotcha — allow silences the warning. The detector treats known ∪ allow as covered, so
adding a name to allow makes its warning go away and lets the value through. That is correct
behaviour, but it means the quiet way to dismiss a gap warning is also the permissive one. Prefer
adding to known.
/metrics and /healthz
What. Two read-only diagnostic endpoints. GET /healthz calls herdr's ping and reports
only whether herdr answered: 200 {"status":"ok","herdr":{"version":…,"protocol":…}}, or
503 {"status":"degraded","herdr":"unreachable",…} on any HerdrException. GET /metrics
renders every registered series as Prometheus text (CB-502): counters
fleetd_sends_total, fleetd_replies_total, fleetd_push_nudges_total,
fleetd_lead_heartbeat_nudges_total, fleetd_spawns_total, fleetd_herdr_calls_total,
fleetd_auth_failures_total, plus gauges fleetd_sessions (one series per lifecycle state)
and fleetd_inbox_depth (one series per undrained target).
On. Both are always registered. /healthz carries no authorization check at all — it is
reachable by an unauthenticated caller by design. /metrics is gated on Authz.Action.METRICS:
open to PRIMARY/WORKER/ARCHITECT, refused to ANONYMOUS.
Why. FleetdMetrics's own class doc states the design intent directly: "each series maps
to a failure mode this project has actually hit, not to whatever was easy to count," and names
the two worth watching — a rising fleetd_sends_total{outcome="completion_fallback"} share
(turn detection degrading) and fleetd_push_nudges_total{outcome="exhausted"} (the primary
stopped draining its inbox). /healthz's narrow scope traces to CB-504: under supervision the
daemon must serve before herdr's socket even exists, so "degraded but alive" needed one cheap,
reliable signal.
Gotcha. Neither endpoint proves the fleet actually works. /healthz echoes back whatever
protocol number herdr reports, but nothing in the codebase compares that number against what
fleetd's own herdr calls need — and the two have already drifted apart in the source itself:
AgentControl's class doc says it was "ported to herdr protocol 19 (herdr 0.8.0, CB-521)," while
HerdrClient's class doc still says "protocol 14, herdr 0.7.0." This is the exact CB-521
incident: herdr answers ping correctly and /healthz goes green, while agent.start and the
rest of the protocol-19 surface fail because the adapter and the herdr binary disagree on
protocol version. /metrics has no counter or gauge for that mismatch either — a resulting spawn
failure only shows up as a fleetd_spawns_total{outcome=…} tick, and only once something
actually tries to spawn.
Bearer-token auth and the non-loopback-bind fail-fast
What. auth.mode decides how a non-worker caller proves it is the primary:
loopback-trust (default) — any loopback caller that is not a known worker/lead/architect pane
is the primary, no credential needed; or token — such a caller must present
Authorization: Bearer <token> or resolves to ANONYMOUS. A worker/lead/architect pane is
always resolved from the unforgeable loopback-PID→herdr-pane mapping regardless of auth.mode.
Separately, the daemon refuses to start at all when bind.host is non-loopback and auth.mode
is still loopback-trust.
On. auth.mode: loopback-trust | token (default loopback-trust); auth.tokenEnv names the
host env var holding the token (default FLEETD_API_TOKEN, read only under token mode). The
bind fail-fast has no separate switch — it always runs in main().
Why. Stated directly in the code: loopback-trust's safety depends entirely on the OS refusing
non-local connections to a loopback socket. Widen the bind without switching to token mode and
"not a known worker" silently becomes "any client that can reach this port is the primary" — the
most privileged role on the bus (spawn/stop/send/drain on any session). Rather than document the
hazard, the config makes it unrepresentable: it throws instead of starting.
Gotcha. There are two separate fail-fast throws, both inline in main(), both before the
daemon binds its port — so a bad config never opens the socket at all. Under launchd, that
repeats forever: the plist's own comment warns launchd retries a fast-failing job every
ThrottleInterval (10s) with no give-up count, until a human unloads the agent or fixes the
cause. And auth.tokenEnv naming an unset/empty var throws a different message than the
bind-mismatch check — don't assume one error class covers both.
The per-session authz table and the audit log
What. Authz.permits(Principal caller, Action action, String targetSession) is one static
table stating, for each of eight actions (SPAWN, STOP, SEND, REPLY, ASK, DRAIN,
READ, METRICS), which role may call it: SPAWN/STOP/DRAIN are the primary alone; SEND
is primary or architect; REPLY/ASK require the caller to own the target session (its own
pane, checked structurally, never by argument); READ/METRICS are open to any authenticated
role. It is the single gate behind both entry paths (REST and MCP) — FleetdApp.allow() and
BridgeMcp's own check both call into it, so the rule can't drift between the two surfaces.
Refusals are recorded by AuditLog, an append-only JSON-lines trail written by a dedicated
audit logger to logs/audit.log (daily rolling, 30-day retention, 100MB cap), independent of
the daemon's normal app log.
On. Always on; not configurable. Every request through FleetdApp or BridgeMcp passes
through Authz.permits().
Why. The class doc states this plainly: most of the rule was already true de facto — a worker's identity comes from its connection, never an argument, so it could never reply as another worker over MCP — but the REST surface used to trust the session id in the URL path outright, and neither surface checked role at all. This makes the invariant explicit and testable instead of emergent, and gives every privileged action one recorded outcome (allowed/denied/failed) instead of none.
Gotcha. Message content is never written to the audit log by design — only
who/what/target/outcome/reason and a correlation id, because the bus carries user source code,
diffs, and prompts, and an audit trail that quietly accumulated those would be a transcript
archive wearing a security control's clothing. READ actions are deliberately excluded from the
allowed audit trail ("reads would drown the trail") — only denials of READ/METRICS are
recorded, not successes. And the 401-vs-403 split matters if you're debugging a refusal: 401
means "you presented no usable identity" (fixable by the caller), 403 means "you are
authenticated, but this isn't yours" — a worker reaching for another worker's session, or for
orchestration it was never granted.
Multi-profile routing and kind: adapter selection
What. Each workers: profile carries a kind: field selecting which backend launcher spawns
it — claude-code (the default) or opencode (CB-402). At startup, main() partitions every
configured profile into two maps by Profile.isOpenCode(), builds one ClaudeCodeLauncher and/or
one OpenCodeLauncher accordingly, and wraps both in a CompositePeerLauncher that routes each
call to whichever adapter declares the profile the call names.
On. kind: claude-code | opencode on a profile; absent or blank defaults to claude-code.
The claude-code adapter is built even with zero claude-code profiles configured, unless opencode
is the only kind present — so a bridge with no workers: at all still has a well-defined base
adapter.
Why. Not stated as a single "why" comment beyond the CB-402 changelog note that opencode
"proves the PeerLauncher SPI is genuinely provider-neutral rather than Claude-shaped" (see the
existing Pin an opencode endpoint entry). The partition-by-kind design itself reads as the
natural consequence: profiles fully own their backend, so routing is a lookup, not a branch.
Validated at config load since CB-604 (2026-08-16). An unrecognized kind: now refuses to
start, naming the profile, the bad value, the accepted set, and what would otherwise happen:
refusing to start: profile(s) [gemini=opencod] set an unrecognized kind — accepted values are
claude-code, opencode (case-insensitive); an unrecognized kind would otherwise fall back to the
claude-code adapter and try to launch a program named after the typo.
Before that fix, kind: opencod was silently accepted, normalized, and — because it did not equal
"opencode" — routed into the claude-code adapter bucket. With argv: also unset, the launch
command defaulted to List.of(kind), literally the misspelled string, because the argv default only
special-cases the exact string "claude-code". Nothing caught it before the spawn failed.
Gotcha. The composite constructor separately refuses two adapters claiming the same profile name ("worker profile '…' is claimed by two peer adapters"), so that failure mode has always been loud.
Checking for the same shape elsewhere found three more fields, all fixed the same way by CB-606 (2026-08-16) — see Every config value is checked against its valid set below.
Supervise the daemon on Linux (systemd)
What. deploy/fleetd.service is a systemd user unit (not system-level — "fleetd drives
the user's herdr, not a system daemon") that runs java -jar target/fleetd.jar fleetd.yaml,
restarts on failure (Restart=on-failure, RestartSec=10s, capped at 5 restarts per 120s via
StartLimitBurst/StartLimitIntervalSec), waits on herdr.service only advisorially
(Wants=, not Requires=, so a herdr restart never takes fleetd down with it), and applies a
sandboxing profile (NoNewPrivileges, ProtectSystem=strict, ProtectHome=read-write, etc.).
On. Manual install: copy to ~/.config/systemd/user/, edit ExecStart/WorkingDirectory/
Environment, then systemctl --user daemon-reload && systemctl --user enable --now fleetd. Not
currently the live supervision target — the unit file's own header comment says the dogfooded
daemon runs on macOS under launchd; this unit is for the Linux gateways CB-308 introduces.
Why. Not stated beyond the practical need: a per-host gateway topology (CB-308, noted elsewhere in the wiki) needs Linux hosts, and those need systemd rather than launchd.
Gotcha. The unit hard-codes Environment=PATH=… with an explicit comment explaining why:
"systemd does not source a login shell, so without it the daemon — and every worker — gets a bare
default with no JDK/Maven" (same defect class CB-511/CB-594 already fixed for PATH on
launchd). For secrets, the unit's own comment says plainly "Secrets are NOT set here" and points
the operator at a systemctl --user edit fleetd drop-in or an EnvironmentFile=. It answers
the login-shell defect only for PATH, not for WORKER_GITEA_TOKEN/AI_GATEWAY_TOKEN.
Systemd answer — YES, it has the underlying defect, undocumented for those two variables
specifically. Comparing to the launchd side: launchd had the identical problem
(deploy/dev.ltms.fleetd.plist's own comment: "launchd does NOT source .zprofile/.zshrc") and it
was fixed by scripts/fleetd-launchd-wrapper.sh, which execs zsh -l so
${SHARED_ENV}/tools/secrets.sh gets sourced — its own header comment names exactly
WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN as the two secrets this closes the gap for. The
systemd unit has no equivalent wrapper and does not source that file at all. Its own comment
mentions only FLEETD_API_TOKEN as the secret to add via drop-in — it never names
WORKER_GITEA_TOKEN or AI_GATEWAY_TOKEN. An operator following the unit file's own guidance
verbatim would set FLEETD_API_TOKEN and stop there: the daemon boots fine, and the failure
surfaces only later as a worker that cannot open a PR or a gateway profile returning 401 — the
exact failure mode CB-594 fixed for launchd, left open here. Fleetd.reportRequiredSecrets
(called at the top of main()) does log which secret env-var names resolved on either
platform — but that log line can only tell the truth about the names it names; it doesn't fix
the sourcing gap, and the unit file gives the operator no prompt to look for it.
Every config value is checked against its valid set
What. A config value the daemon does not recognize now refuses to start, naming the field, the
value, the accepted set, and what would otherwise have happened. Four fields carried this defect and
all four are covered: a profile's kind: (CB-604), auth.mode, a profile's placement:, and the
top-level placement: policy (all CB-606).
refusing to start: auth.mode=toekn is not recognized — accepted values are loopback-trust,
token (case-insensitive); an unrecognized mode would otherwise silently fall back to
loopback-trust, which authenticates nobody.
On. Always on; nothing to configure. Every valid value still works case-insensitively, and every
absent value keeps its old default — kind: → claude-code, auth.mode → loopback-trust, a
profile's placement: → tab, the policy → fixed.
Why. All four fields had the same shape: the value was lower-cased in a compact constructor, then compared against exactly one string. Anything else fell through to the other branch silently. So a typo was accepted, did nothing the operator asked for, and reported no error.
auth.mode is why this is a release blocker rather than tidying. A typo of token behaved as
loopback-trust, and validateAuthExposure() only fires on a non-loopback bind — so on a
loopback bind, the common case, the mistake was invisible end to end. The daemon started cleanly and
authenticated nobody, while the operator's config said token and the operator believed it.
Gotcha — the general lesson, not the instances. CB-604 fixed kind: alone. The brief for it also
asked the worker to report any other field with the same shape and explicitly not to fix them. That
one question found the other three. When you fix a defect of shape, ask what else has that shape
— the instance you noticed is rarely the worst one. Here the field nobody had filed a ticket about
was the field that turned authentication off.
Two implementation notes worth keeping. The top-level policy was already validated, by
PlacementPolicies.fromName — but lazily, at first spawn, through the supplier in
CompositePeerLauncher that re-reads live config. So a bad name still started a daemon that looked
healthy and failed much later. It is now called eagerly at load as well, which also covers
ConfigRef.reload(). And the check calls fromName rather than copying its list, so there is still
one source of truth for what a valid policy name is.
Old work-in-progress snapshots clean themselves up
What. The daemon now deletes a refs/wip/<branch> snapshot once it is safe to — and GET /members reports how many are left and roughly what they cost:
{ "workers": [ ... ], "wipRefs": { "count": 12, "costBytes": 4183042 } }
The rule has two conditions and both must hold:
- The snapshot commit's tree is already reachable from
main. The content exists somewhere else, so deleting the ref loses nothing. - The snapshot is older than 24 hours.
Every deletion logs the ref name, the commit sha, the age, and the exact git update-ref command
that puts it back:
pruned snapshot ref refs/wip/worker/cb578c-92885c-1 commit=4a91f0c (age 51h): its tree is
already reachable from main, so the work is preserved; recover from reflog via
git update-ref refs/wip/worker/cb578c-92885c-1 4a91f0c
On. Always on; nothing to configure. The sweep rides the existing session reaper loop and runs
every 6 hours, plus once on the first pass after a restart. A fleet that has never snapshotted
anything sees no change, and wipRefs is simply absent from /members when no worktree session has
told the daemon which repo to look in.
Why. CB-578 stage C commits a dirty worktree to refs/wip/<branch> on release, so a worker's
uncommitted work is never lost. Nothing deleted those refs. A ref is a garbage-collection root, so
every snapshot pinned its whole tree and git gc could never reclaim any of it. Branch names are
unique per session, so the refs pile up rather than replace each other. It was never an error — it
would have shown up as a slow git gc and a large .git, months later.
Reachability is the safety floor, not the age. A snapshot exists because the work was not committed anywhere else, so a plain time-based sweep would throw away the only copy — exactly the failure the snapshots were built to stop.
Gotcha — the sweep shipped dead, and every test still passed. SessionReaper held the last sweep
time in a long that started at Long.MIN_VALUE to mean "never yet". That sentinel cannot be
compared by subtraction: System.nanoTime() is positive, so now - Long.MIN_VALUE overflows to a
large negative number, the "has 6 hours passed?" gate read it as swept moments ago, and the method
returned before the assignment that would have fixed the field. The sweep never ran once, for the
life of the process, with nothing in the log to say so.
The unit tests missed it because they all called the sweep method directly and walked around the gate. Two lessons worth keeping: a "never yet" sentinel belongs in its own boolean, not in a magic value of the same field, and a test that calls the seam does not prove the caller reaches it. The test that now guards this starts the real reaper loop and requires a prune call to arrive.
Second gotcha — costBytes is a rough figure, not disk usage. It sums the uncompressed sizes of
every blob in every snapshot's tree. Objects shared between snapshots, or shared with main, are
counted once per snapshot. So it reads much larger than the space actually reclaimable, and it is
useful for spotting growth over time, not for capacity planning. Nothing pushes refs/wip/* to the
forge; they are deliberately local and invisible to git branch.
Set a member's auto-compact window
What. A profile can set autoCompactWindow: <tokens>, the context size at which a spawned
member compacts on its own. For a claude-code member it becomes --autocompact <tokens> at
launch. For an opencode member it becomes a per-model provider.<p>.models.<m>.limit.context in
the generated config. Unset ⇒ the backend's own default (Claude Code's built-in auto-compact,
opencode's compaction.auto).
On. Per profile in fleetd.yaml, e.g. autoCompactWindow: 250000. The value must be in
[100000, 1000000] — the band Claude Code's flag accepts — and the daemon refuses to start with a
profile outside it, naming the profile. For opencode the key only takes effect when model: is in
provider/model form (a warning is logged otherwise), because the limit is written under that exact
provider and model.
Why. A member that runs out of context dies mid-turn and loses its fleet_reply, so the report
is gone even though the work was done. The backends already auto-compact, but the window was not in
the operator's hands — a long delegation on a big model could compact too late. This makes the
trigger a per-profile choice, so an expensive profile can be told to compact earlier than a cheap one.
Gotcha. The two backends take the window by different means. Claude Code's --autocompact is an
absolute token count and outranks env and settings. opencode has no absolute "compact at N" knob at
all; the only lever is the model's limit.context, and the generated config also writes
limit.output (16384) because opencode's schema requires both. On a gateway profile that is a
partial override of the model's real limits — confirm the live member still answers after a spawn.
Related: opencode members run with --auto.
Leads on different hosts talk over a shared broker
What. Two leads that share no herdr session — including leads on different hosts — can message
each other by a stable coordination id. A coordinator: block gives this daemon one id and a durable
AMQP mailbox lead.<id>.inbox on a shared vhost. fleet_send{coordId: <peer's id>, content}
publishes to the peer's mailbox; the peer's daemon drains its own mailbox and types the message into
its local lead's pane. fleet_list reports this daemon's own coord-id, so an operator knows the
address to hand a peer.
On. Add a coordinator: block with selfId: (this daemon's globally-unique id) and either
uri: or uriEnv: (the shared broker, ideally a separate vhost from the worker reply inbox). With
no block, nothing opens and lead-to-lead stays the existing same-host pane path. If the block is
present but selfId is blank, the daemon logs a warning and skips the mailbox — it does not crash.
If the broker is unreachable at boot, the daemon warns and carries on with the feature off.
Why. Leads talk to each other widened addressing, but only for leads that share one herdr session — the address was a bare herdr terminal id, which means nothing on another host. Cross-host coordination needs an address that means the same thing everywhere and a transport that is not a local pane. The shared broker is the one channel both hosts already reach, so a coord-id plus a per-lead mailbox is the smallest thing that carries a message across the gap.
Gotcha. A publish to a coord-id whose mailbox no daemon owns is reported as unroutable, not
silently dropped — the peer must be configured and running for the message to land. Delivery into the
local lead's pane is status-gated exactly like a worker reply: a message waits, unacked, until the
lead pane is idle, blocked or done, so it is never typed in mid-turn. A daemon with several leads
must name one lead after coordinator.selfId, or it cannot tell which pane a peer's message is for
and holds it. This is transport only — there is no peer discovery yet; a lead is addressed by a
coord-id it was already told.
A lead can read its own held peer mail
What. fleet_poll{coordId: <your own coord-id>} returns this daemon's held lead-to-lead
messages in full, and does not ack them. The messages stay held, so reading twice returns the same
bodies. It reads only your own mailbox: passing a peer's coord-id is refused with a message saying
so, because no route exists for reading another daemon's mail.
On. Nothing to configure beyond the coordinator: block that lead-to-lead messaging already
needs. Pass your own selfId — fleet_list reports it as coordinator.selfId — as fleet_poll's
coordId. fleet_list also now reports heldCount and heldDurable beside held[]. heldDurable is
derived from the queue's declared durability and the consume ack mode, not written as a
constant, so it can actually report false if either changes (fleetd #440). It is one boolean
over two inputs, so a false does not say which of the two moved.
Why. A lead could see that it had peer mail but not read it. fleet_list's held[] carries
only an 80-character preview, on purpose, so a constantly-called roster scan never dumps a
coordination body. fleet_poll{target} drained the wrong store — a worker's reply inbox, not the
coordinator mailbox — and returned an empty list with no error. The only other route, fleet_ack, looked
destructive but is not: it never touches the coordination channel at all, and reports
acknowledged <msgId> anyway (fleetd #437). So the full body was reachable in principle
and unreachable in practice, and the busiest leads were the ones most affected: delivery into a
lead's pane is status-gated, so a lead that never goes idle never drains its own mailbox.
Gotcha. This is primary-only, under a new authorization action COORD_READ. That includes
refusing an architect, which holds the ordinary READ action today. The reason is that READ's
grant rests on the roster carrying no secrets, and a lead-to-lead body is not the roster — it is
where leads discuss host shapes, credential states and unmerged work. Also do not read
mailbox.pending: 0 as an empty mailbox: pending counts only broker-ready messages, so a healthy
mailbox holding real mail reports zero. heldCount is the honest number.
Fleet health can finally reach its fault states
What. FleetHealthMonitor ticks on its own timer, well away from the 250ms delivery poller. Each
tick makes one agent.list call plus one roster snapshot, then classifies every member through the
pure FleetHealth.decide. There are 14 health states and 9 of them are faults. A fault is logged
when the state changes, never once per tick. Two terminal states, GONE and NEVER_READY, also
call failTarget (MessageService.abandon), which fails every ticket still waiting on that member.
On. The health: block in fleetd.yaml: enabled: true, intervalSeconds: 30 (floor 15),
workingSuspectAfterSeconds: 600 (floor 300 — how long a BUSY member with no activity must sit
before it is STALL_SUSPECTED), and notifications.mode. fleet_list reports healthCoverage:
off, detection-only, or full. It is full only for mode: webhook.
Why. The monitor shipped half-wired and nothing said so. tick passed a hardcoded
NOT_YET_OBSERVED = false for 7 of the 12 HealthSnapshot fields, so 8 of the 9 fault states
could never be reached — only TURN_BOUNDARY_LOST could fire. GONE and NEVER_READY were among
the dead ones, which means CB-580's repair never ran: a ticket waiting on a member that had vanished
still hung to the 30-minute async timeout. Everything compiled and every test passed, because each
test called the classifier directly and walked around the caller. Three units fixed it: CB-640 added
the message-layer facts (hasQueuedDelivery, hasStrandedReply, hasOrphanedDelegation), CB-641
added the herdr and clock facts (present, targetNotFound, controlLinkDown,
readinessGraceElapsed, stalled), and CB-643 joined them and deleted the placeholder. The rule
that came out of it: a snapshot field with no publisher is a dead state, and nothing else will tell
you.
Gotcha. Three things to know before you read the log.
- With
notifications.mode: disabledthe coverage isdetection-only. Every fault the monitor finds goes to the log and nowhere else. The daemon prints a warning about this at startup. - A member that is already faulty when the daemon starts logs its fault once, on the first tick, and then stays quiet. Grepping recent lines can therefore show nothing while the fault is live.
DELEGATION_ORPHANEDneeds the fact to hold for two consecutive ticks. An async ticket exists for a moment before its virtual thread opens the rendezvous waiter, so for that instant nothing is accepted or queued behind it and a single tick would report a fault that clears immediately. The gate costs one interval of delay on a real orphan.
/fleets-status — every fleet that shares one broker
What. A primary-side skill in .claude/skills/fleets-status/. It reports the state of every
fleet that shares one LavinMQ instance, in three tiers: (1) the local daemon — process, deployed jar
against HEAD, /healthz, fleet_list with capacity and healthCoverage, and the WARN/ERROR lines
since the last fleetd listening marker; (2) the broker itself over its management API — each
vhost's queues, depths and consumers; (3) cross-fleet lead coordination, read from the
lead.<coordId>.inbox queues. Every tier ends with what it cannot see, and a fleet it cannot
reach appears in the report as unknown with the reason, never as an omission.
On. Type /fleets-status. Tier 1 always runs. Tier 2 runs only when
LAVINMQ_MANAGEMENT_USER and LAVINMQ_MANAGEMENT_PASSWORD resolve in a login shell.
Why. There was no way to ask one question about the whole fleet estate. A lead could see its own daemon and nothing else, and the remote daemon's REST port is not reachable from here. The broker is the one thing both fleets touch, so its management API answers for a host you cannot log in to. The sharpest fact it gives you: a vhost with queues but zero consumers means that fleet's daemon is down while its durable state survives.
Gotcha. Tier 2 is blocked today. The AMQP user in LAVINMQ_URI connects fine on 5672 but gets
HTTP 401 from the management API on 15672 — an AMQP connection is not monitoring access. An operator
must create a separate read-only user with the monitoring tag, allowed on both vhosts, and store it
under the two names above in ${SHARED_ENV}/tools/secrets.sh. Also: the skill forbids printing
LAVINMQ_URI, because the password is inside it. Output that could carry it is redacted with
sed -E 's#://[^@]*@#://<redacted>@#g', and the g is not optional — without it a line holding two
URIs leaks the second one, which scripts/redeploy-fleetd.sh --check prints.
memberHerdrSocket — run members on a second herdr daemon
What. A new top-level config key. Set it to a second herdr API socket path and every member
is spawned on that daemon, while everything the lead does stays on the first one. Absent (the
default), both are the same client and nothing changes. Inside the daemon a HerdrRouter owns the
split: it hands each consumer the lead client, the member client, or — for consumers that see both
kinds of traffic — routes per target id.
memberHerdrSocket: /run/fleet/members.sock # omit for one daemon (the default)
On. Add the key and restart. It is read once at startup, so it is a deferred key, not a hot one.
Why. herdr forks every pane as its own OS user, and there is no user, uid or run-as parameter
anywhere in its socket API — measured against herdr 0.8.0's own docs, not just our client. So the
only way to make a member run as a different user than the operator is a second herdr server started
by that user. That matters because a member currently reads the operator's ssh key straight off the
filesystem: blocking SSH_AUTH_SOCK does not stop it, and no environment scrub can, because the
scrub removes variables and the key is a file (#184). A separate OS user is the control; this key is
the fleetd half of it. Design and measurements are in #185.
Gotcha — the feature is still not ready to switch on. #185 keeps the running list. Two of the blockers were cleared on 2026-08-31; the rest stand.
Cleared. Pane ownership lives only in memory, so a restart used to leave every surviving pane
unowned and fleet_stop refused it. CompositePeerLauncher now probes: it asks each daemon
which one knows the pane, routes to the single match and caches it, treats no match as already
stopped, and refuses only a genuine two-daemon collision — with a message that says that is what
happened. One unreachable daemon no longer breaks the probe for panes owned by a healthy one.
Still open. herdr creates its API socket mode 0600 and umask does not change it, so cross-user
access still needs a post-start chmod g+rw — a race on every start, and an upstream ask that has
not been made. Pane ids are still not daemon-qualified: fleet_stop{paneId} takes a bare
w1:p1, so the probe makes a collision fail loudly instead of silently closing the wrong pane, which
is a safety net, not the design fix. hostEnvNames still reads fleetd's own environment, which is
the wrong environment under a second user.
Worktree ownership now has an answer — worktreeGroup,
below — but a new blocker replaced it, and it is the harder one. ClaudeCodeLauncher writes the
role charter with Files.createTempFile, and on macOS the system temp directory is per-user:
$ ls -ld "$TMPDIR"
drwx------@ 219 dai.ha staff /var/folders/wf/…/T/
A member running as another user cannot even traverse it, so it never reads its charter — and the
charter carries the reply contract. EnvAllowListScrub has the same problem: it generates the
member's ZDOTDIR there, and the member's own login shell has to read it. No chmod fixes a
per-user 0700 directory safely; both writes have to move somewhere both uids can reach. Two
launcher writes (CLAUDE.local.md and <commonDir>/info/exclude) also land at spawn time, i.e.
after the group fix-up has run — harmless today because the member only reads them, but nothing
says so.
The shape to watch for. herdr's workspace, tab and pane ids are per-daemon sequential
counters. Two daemons really do both hold w1:p1, pointing at different panes owned by different
users — measured, not theorised. Four separate defects came from code that assumed one daemon, and
every one of them passed the whole test suite, because with the key unset both clients are the same
object. When reviewing anything on this path, ask each herdr call one question: which daemon does
this go to, and is that the daemon that owns the thing it is asking about?
Observable change even with one daemon. GET /healthz requires both daemons to answer
before it returns 200 (with one configured that is the single ping it always was). GET /sessions
merges workspace.list across both.
Since 2026-08-31 the 200 body also reports the member daemon's version and protocol, under its
own member key, and adds protocolMismatch: true when the two protocol numbers differ:
{ "status": "ok",
"herdr": { "version": "0.8.0", "protocol": 3 },
"member": { "version": "0.7.1", "protocol": 2 },
"protocolMismatch": true }
The member key and protocolMismatch appear only when a second daemon is configured, so a
single-daemon body is byte-identical to before. herdr still carries the lead daemon's values,
deliberately: scripts/redeploy-fleetd.sh and scripts/rename-checkout.sh both read this endpoint,
and folding the two daemons into one key would hide a mismatch from whichever reader looks at only
that key.
This closes a real trap. The member daemon was pinged and its answer thrown away, so a member herdr
running a mismatched protocol left /healthz green while every spawn failed — and it is the
member daemon's protocol, not the lead's, that decides whether a spawn works.
worktreeGroup — let a member under another OS user write its worktree
What. An optional top-level key naming an OS group. When set, every provisioned worktree and
the repository's git store are made group-writable, so a member spawned under a different OS user
(see memberHerdrSocket) can write its
own worktree, its per-worktree git metadata, and the objects its commits create. Absent — the
default — nothing runs and no process is spawned.
worktreeGroup: fleet-workers # omit for the single-user default
On. Add the key and restart. The operator running fleetd must already be a member of that group, or every provisioning spawn fails loudly, naming the group.
Why. GitWorktrees shells git worktree add, so it runs as fleetd's uid and everything it
creates is owned by the operator at 0644/0755. A member running as another user cannot write any
of it. That user is the whole point of #185: a member currently reads the operator's forge ssh key
straight off the filesystem, and no environment scrub can stop that, because the key is a file
(#184). This key is the git half of the fix.
It isolates credentials, not the repository. A member in the group can still write the
operator's git objects and refs. Say that plainly rather than implying more — the two uids share one
repository by design, which is what keeps the lead's refs/wip/* safety net working.
Gotchas, all three found the hard way.
The share pass must run after overlayParity, never inside add(). overlayParity copies more
files into the worktree after add() returns, so sharing any earlier leaves every overlay file
operator-owned and unwritable — with the whole suite still green. A test pins the order.
Missing paths are skipped, not passed to chgrp. .git/logs does not exist with
core.logAllRefUpdates=false or before the first ref update, and packed-refs does not exist until
refs are packed. chgrp on a missing path exits non-zero, which would fail every spawn with a
message blaming a group that is fine.
The git directory is asked for, never assumed. <repoRoot>/.git is a file, not a directory,
when the checkout is itself a linked worktree — which is what this code creates for every member. It
resolves git rev-parse --git-common-dir against repoRoot, because git answers relatively for an
ordinary checkout and absolutely for a linked one.
Cost. The fix-up re-runs on every spawn, deliberately. core.sharedRepository=group governs only
what git writes afterwards; the walk is what covers everything already on disk. Three walks of the
object store per spawn — about 3000 files in this repo, well under a second, growing with the repo.
It is not redundant work to optimise away.
opencode session ids come from opencode.db
What. fleet_list reports agentSessionId for an opencode member, and
fleet_spawn{resumeSessionId} can therefore relaunch one onto its prior conversation. Automatic —
no knob.
It took two fixes, not one, and the first one alone did nothing visible. Reading opencode.db
(below) was necessary but not sufficient: SessionManager asked the handle for the id once, at
spawn, and froze it into an immutable record. opencode has not written its row at that moment, so
the frozen answer was always null and the roster never showed one. The handle's own javadoc said
"the caller re-calls later" — no caller did, and SessionManager did not even keep the handle.
See Late-resolved member ids for the second half.
On. Nothing to turn on. It needs opencode's own storage root
(~/.local/share/opencode) to be readable, which it is.
Why it is an entry at all: it was silently dead. opencode moved its session store from a
one-JSON-file-per-session tree to a SQLite database in January 2026, and
OpenCodeSessionDiscovery went on scanning the frozen tree. It therefore returned null for every
member, forever. agentSessionId was never known, so resumeSessionId was unreachable for every
opencode profile — most of the fleet — and nothing said so, through 57 member spawns. It now reads
<storageRoot>/opencode.db, matching the session table's directory column against the worker's
cwd and taking the highest time_updated.
The gotcha is what the silence cost. A discovery that finds nothing and a store that is empty look identical. That is why a missing database is now a WARN naming the path searched — it means the layout moved again — while a present database with no matching row stays at DEBUG, because that is the normal answer right after a spawn.
Read-only, and pinned as such. opencode may be running and writing that database (WAL mode, 841MB
here), so the connection is opened with SQLITE_OPEN_READONLY. Beware the obvious test for this:
making the file unwritable proves nothing, because SQLite silently downgrades a read-write open
of an unwritable file to read-only. The test that works asks the connection to INSERT and requires
the refusal.
One dependency. org.xerial:sqlite-jdbc, which ships bundled native libraries and takes the
shaded jar to about 28.6MB. Chosen over shelling out to the sqlite3 binary so there is no
undeclared external-binary requirement on a headless host.
One worker reply settles exactly one ticket
What. When a worker's fleet_reply arrives with no send waiting for it, the bridge holds it and
later hands it to one delegation — never to several. If it cannot tell which one a held reply
answers, it queues the reply instead of guessing. Automatic — no knob.
On. Nothing to turn on.
Why it exists. Two separate paths used to assume a target had at most one delegation open, with no guard for the case where it had two:
fleet_send{turnId}(the answer to afleet_ask) waits only for its own bounded MCP window — 25s by default, 120s at most. A resumed turn can easily run longer than that. When the window expired, the worker's real reply had no waiter left, so it was held instead. The ticket stayedPENDINGuntilfleet_stop, which then reported "the worker session was released before it replied" — with a worktree, a branch and a snapshot commit, so it read like lost work. The reply had in fact arrived. Fixed first: a held reply now completes its own ticket (388aba7).- Teardown then drained that held reply once and reused it for every open ticket on the target.
Two open tickets both came back
DONEwith the same text, one of them an answer the worker never gave for that delegation. Fixed second (966c58a): the reply goes to the oldest open ticket and every other one keeps the ordinary failure path.
The gotcha: a confidently wrong status costs more than no status. Both defects produced a plausible answer, not an error. A lead that trusted the first one would go recover a snapshot of work that was already merged; a lead that trusted the second would act on a report its worker never wrote. That is why the ambiguous case queues the reply and logs a WARN naming every candidate ticket rather than picking one.
Known limit, on purpose. Only the second defect's broad case is reachable today. Two open
tickets on one target is reachable and was a real bug. The same ambiguity in the reply path, and the
combination of two open tickets with a held reply, are both blocked by existing guards — one send
holds a target's lock and its waiter together, so a held reply and a second open ticket cannot exist
at the same instant. The guards stay as defence in depth, and MessageService.abandon's javadoc
records why no test pins them: a test can force the state only by breaking that lock-and-waiter
invariant from outside the class, which no caller does.
Late-resolved member ids
What. A member whose backend cannot name its own session at spawn time gets its
agentSessionId filled in later, as soon as the backend writes it. fleet_list, fleet_status
and the REST roster all report the current answer rather than the one frozen at spawn. Automatic —
no knob.
On. Nothing to turn on. It only does work for a member whose id is still unknown.
Why it exists. agentSessionId is the handle fleet_spawn{resumeSessionId} needs. It was read
once at spawn and never again, which is fine for claude-code (it knows its session id up front) and
useless for opencode (it does not). The result was a documented, tested feature that could never
produce a value for most of the fleet. This is the second half of the
opencode.db fix.
The gotcha: the resolving read is not the cheap one. Asking an opencode handle for its id opens
a SQLite database — around 841MB here. roster() is the roster supplier for the heartbeat loop and
the health monitor, so resolving there would open that database on every tick, for every unresolved
member, forever for a member whose row never appears. So there are two reads:
| Method | Resolves? | Used by |
|---|---|---|
roster() |
no | LeadHeartbeatLoop, FleetHealthMonitor, placement and exhaustion checks, the /metrics scrape |
rosterResolved() |
yes | fleet_list, the REST roster — the surfaces that actually report the id |
get(paneId) (behind fleet_status) and the teardown path resolve too; both are caller-driven, not
timers. A test asserts the plain roster() never calls the handle again, so the split cannot be
quietly "tidied up" later. Once an id is found it is stored and never looked up again.
GET /members reports its rows under members
What. The REST roster's response body uses the key members, matching its path. The old key
workers is still emitted as a deprecated alias carrying the same rows.
On. Nothing to turn on.
Why it exists. The route was renamed /workers → /members, and the body key was not renamed
with it. A caller reading body["members"] — the obvious guess given the path — got an empty
list, which is a perfectly valid answer meaning "this fleet has no members". So a client reported
an idle fleet and nothing anywhere contradicted it.
The gotcha: do not just rename the key. The REST surface is the out-of-band path a lead falls
back to when its MCP mount drops, and at least one client reads workers today. Renaming outright
would break that client the same silent way, in the other direction. Both keys are emitted, and a
test asserts they carry identical rows so the alias cannot drift. Drop workers once nothing reads
it.
The credential gap detector admits when it cannot see
What. With memberHerdrSocket: set, the memberCredentials gap report stops drawing
conclusions about member panes and says the gap is unknown, not clean, naming the config key
that made it unknowable. One WARN per launcher, not per spawn.
On. Only in two-herdr mode. With memberHerdrSocket: absent — the default — every log line on
this path is byte-identical to before, pinned by a test.
Why it exists. The detector enumerates fleetd's own environment, on the assumption that the
member pane's login shell exports the same set. Under memberHerdrSocket: panes are routed to a
second herdr, and fleetd has no channel to confirm what OS user that herdr runs as — it may be
the same user or a different one, with a different $HOME and a different secret store. Either way
the assumption is no longer something fleetd can check, so the old output was reporting on a process
it could not see, and could call a gap clean while the member exported something dangerous. There is
no channel to read another user's environment, so "unknown" is the only honest answer.
Note the shape of that sentence: the fix is not "the member runs as a different user". Saying that would be the same mistake pointing the other way — fleetd cannot confirm a different uid any more than it can confirm the same one (fleetd #269). What changed is the claim, from a false certainty to an accurate unknown.
The count line beside it says whose environment it counted. The INFO line
member credentials: allowed N of M reads as a statement about the member's pane, and under
memberHerdrSocket: it is not one — the counts come from fleetd's own process. Since fleetd #276 the
line says so in place, because the caveat used to live only in a javadoc the operator reading
fleetd.out never sees. The counts themselves are unchanged and still real. With
memberHerdrSocket: absent the wording is byte-identical to before.
The gotcha: this fixes the report, not the control. Two related defects in the same area are
open under #213 — the ZDOTDIR scrub decides whether it can run by reading fleetd's $SHELL,
and writes its generated file into fleetd's own TMPDIR, which on macOS is a per-user 0700
directory another uid cannot even traverse. In two-user mode the scrub can therefore be silently
absent while the daemon believes it ran. Treat memberCredentials as unverified in that mode until
#213 lands.
Backfill status
This page was started after the fact, so it is not yet complete. Entries above are written from verified behaviour.
The CB-595 backfill (2026-08-16) cleared the largest gap: the ~14 operator-facing tickets shipped
between v1.0.0 and now that had landed nowhere. Those entries were written from a read of the code
on main, not from commit messages, and the three keys that changed meaning were pulled to the
front because an upgrading operator meets them first.
The second CB-595 pass (also 2026-08-16) cleared the five areas that pass had left open: /metrics
and /healthz, bearer auth and the bind fail-fast, the authz table and audit log, systemd
supervision, and multi-profile kind: routing. Every claim in those five was traced to a file:line
on main, and two of them turned up defects the catalogue work was not looking for — see below.
Still to catalogue:
- the reply push loop and its nudge budget (CB-307 step 2, CB-590, CB-598) — the budget is now tracked per pending item, not per lead and not per source. A newly-arrived item keeps its source eligible even when an older, still-undrained item has used up its own budget. The practical consequence to write up: the cap bounds nudges about one item, so a lead with a steady arrival of new work keeps being nudged — which is correct, but is not what the knob's name suggests
Role agent definitions — how a member learns its role (CB-617)
What it does. The architect, dev and reviewer contracts live as files in the repo, at both
.claude/agents/<role>.md and .opencode/agent/<role>.md. A member's worktree is a checkout of this
repo, so it arrives with its own contract. At spawn the launcher passes --agent <role> when the
matching file exists.
The knob. Nothing to turn on. The role you pass to fleet_spawn selects the file. Absent file ⇒
no --agent flag and the member still spawns, so this degrades rather than breaks.
Why it exists. The contract used to travel as an inline argv flag,
--append-system-prompt <text>. herdr refuses to shell-encode a multi-line argument, so every
claude-code member with a configured role charter failed to spawn (CB-616) — measured as an
architect that started on sol and died on opus. opencode escaped only because it already wrote
its charter to a file. Files fix that for both backends and make the contract reviewable in git.
The gotchas.
- Both directories, every time. Neither tool reads the other's: a
.claude/agents/file is invisible to opencode. Measured — a probe agent in.claude/agents/never appeared inopencode agent list. The two copies of a role must carry identical bodies or the backends work from different contracts. - Never put
model:in an agent file. Both CLIs honour it only when no launch flag is passed, and fleetd always passes one (--modelfor claude-code,-mfor opencode). Measured both ways: an agent pinned toopenai/gpt-5.6-terrarun with-m opencode/x-preview-f-freereported> pin · x-preview-f-free. Amodel:here is silently overridden on every spawn. The model stays infleetd.yaml, which also keeps role and backend as the separate axes the role pools need. - Both charters travel in one file, or Claude Code will not start (CB-618). The CLI refuses the
two prompt flags together:
Error: Cannot use both --append-system-prompt and --append-system-prompt-file. Please use only one.The first cut of this feature put the role charter on the file flag and left the reply charter inline, and every claude-code spawn with a role charter died at launch — reported by fleetd asspawn_timeout, which hides the cause completely. So when a role charter is present, both go in the one file with the reply charter last: last is where the reply rule must sit, because it is the rule that must survive. A member with no role charter keeps the inline flag, which is also the only form that reaches a member with no repo checkout. - A
.claude/agents/file needsname:in its frontmatter (CB-618). With onlydescription:the file is skipped in silence and the launch fails with--agent 'architect' not found. Available agents: claude, Explore, .... The OpenCode files take their name from the filename and need no such key, so the two formats are not interchangeable even though the bodies are identical. - Measured together, on the real binary.
claude --agent architect --append-system-prompt-file <both charters>answeredROLEOK, REPLYOK, yes— agent body, role charter and reply charter all in force at once. Both defects above shipped green because every test read the argv fleetd builds and none ran the binary that has to accept it (see #113).
Found while cataloguing, not by looking for bugs
Two of these are filed as their own tickets. They are recorded here because both are the same shape: a configuration that is accepted, does nothing useful, and reports no error.
kind:is never validated. A typo such askind: opencodis lower-cased, accepted, and — not matching"opencode"— routed to the claude-code adapter. Ifargv:is unset, the launch command defaults to the misspelled string itself rather thanclaude. No config-load check catches it. See Multi-profile routing above.- The systemd unit has the login-shell secret defect that launchd's had.
deploy/fleetd.servicefixesPATHexplicitly and names onlyFLEETD_API_TOKENas a secret to add. It never sources the secret store, and never mentionsWORKER_GITEA_TOKENorAI_GATEWAY_TOKEN. An operator following the unit's own guidance gets a daemon that boots cleanly and members that cannot open a PR. launchd gotscripts/fleetd-launchd-wrapper.shfor exactly this; the Linux unit has no equivalent.
One more piece of drift worth knowing while reading these entries: AgentControl's class doc says
herdr protocol 19 (herdr 0.8.0), while HerdrClient's still says protocol 14 (0.7.0). Nothing
compares the protocol number herdr reports against what fleetd actually needs — which is precisely
how /healthz once went green while every spawn failed.
A launch command that cannot fit is refused
What. Before starting a member, fleetd measures the command it is about to hand herdr. If it cannot fit the pane's line, the spawn is refused straight away with the byte count and the longest argument named. Automatic — no knob.
On. Nothing to turn on.
Why it exists. herdr does not exec a member's launch command. It types it into the pane,
and a pty line buffer holds 1024 bytes (BSD/macOS MAX_CANON). Everything past that byte is dropped,
and no layer says a word: herdr answers "agent started", the backend exits on the mangled argument it
was handed, the pane closes, and the only symptom is the readiness gate timing out twenty seconds
later with no reason at all.
That is not a theory. It took the whole claude-code half of the fleet down. The reply charter used to
ride inline on --append-system-prompt, so a sonnet launch command was already 978 bytes. Adding
--session-id <uuid> for #214 — fifty bytes — made it 1028, and the four bytes cut off the end
turned --autocompact 250000 into --autocompact 25. Claude Code refuses that value, so every
claude-code spawn died. The cut is at byte 1024 exactly; that was measured on a live pane, not
assumed.
Two things changed. The charter now always travels as --append-system-prompt-file, which takes
~800 bytes of prose off the command line for good. And this guard catches whatever grows next.
The gotcha: the estimate is deliberately too big. fleetd cannot see how herdr quotes each argument, so every argument is charged its own bytes plus a separator and a quote pair. An over-estimate costs you a clear error at a length that was already unsafe. An under-estimate would let the silent truncation back in, and that is the failure this exists to prevent.
It already caught a second one. A profile with ideMcpUrl: set assembles 1084 bytes — over the
limit before this ticket was written. The IDE mount was one config key away from the same silent
death. It cannot ship that way now.
A failed spawn shows you the pane
What. When the spawn-readiness gate gives up on a member, the WARN carries the last status herdr reported and the tail of the pane, read before the pane is closed. Automatic — no knob.
On. Nothing to turn on.
Why it exists. The gate used to log only that it had timed out, which is true of every possible cause: a backend that never launched, a binary that rejected an argument and exited, a trust prompt waiting for a keypress, a login shell that hung. The pane holds the one copy of that answer, and the next line of code destroyed it.
Finding #220 without this took a live process sampler, a hand-built pty and a byte count. With it,
the answer was one log line: ... --model claude-sonnet-5 --autocompact 25.
The gotcha: read it before stop(). The order matters and is easy to get backwards. stop(paneId)
closes the pane, and a pane read after that returns nothing — the log would be just as empty as
before, only slower.
Every claude-code member is resumable
What. Every claude-code spawn gets a session id, so fleet_list always reports an
agentSessionId you can hand back to fleet_spawn{resumeSessionId}. Automatic — no knob.
On. Nothing to turn on.
Why it exists. A session id used to be minted only when the caller passed sessionName or
resumeSessionId, and fleet_spawn treats both as optional. So an ordinary spawn minted nothing and
agentSessionId stayed empty for that member's whole life. Unlike opencode, nothing resolves it
afterwards: a claude-code session id is fixed at launch.
That made resume opt-in at spawn, while the moment you want it is after a member has done work worth keeping. By then it was too late, and an empty field in the roster read as "this backend does not support resume", which was not what was happening.
The gotcha: the binary takes a UUID and nothing else. claude --session-id rejects any other
value at argument parsing — "Invalid session ID. Must be a valid UUID." — so the random UUID is
required, not incidental. The one documented flag interaction is with -r, which claims the same id;
a resume spawn passes -r and never mints.
An exhausted backend is quarantined even from a chrome-only pane
What. When a member's pane yields no usable assistant block, fleetd still classifies
BACKEND_EXHAUSTED and BACKEND_ERROR from the raw screen — so an exhausted credential is
quarantined instead of being recorded as a member that did nothing. Automatic — no knob.
On. Nothing to turn on.
Why it exists. The completion resolver checked for an empty scrape before it looked for an
exhaustion line, and returned. lastAssistantBlock finds nothing whenever the pane carries no ⏺
marker and its first visible line is TUI chrome — a box border, a warning, the input box — because
the boundary scan breaks on that first line.
The exhaustion branch is the only caller of the quarantine sink. So a backend that refused the turn for a usage limit was filed as "produced nothing", the credential was never quarantined, and fleetd kept handing work to a backend that could not run it. Each attempt looked like another member silently doing nothing. That is the shape of a usage limit taking out several members with the roster giving no reason.
The gotcha: it is a fallback, not a reordering. The empty-scrape failure was not moved or weakened, and a pane that yields a usable block never reaches this path. Matching the raw screen widens what the patterns can hit, including scrollback that is not this turn's output — and a false positive here quarantines a working credential. That is why the raw match runs only where the alternative was a generic failure with no information in it at all.
The credential scrub follows the member's own user
What. When memberHerdrSocket: puts member panes under a different OS user, the allow-list
credential scrub uses that user's facts: memberLoginShell: decides whether the ZDOTDIR scrub can
run, and the generated directory goes under worktreeRoot, shared read-only with worktreeGroup.
On. memberLoginShell: and worktreeGroup: — both required once memberHerdrSocket: is set.
With memberHerdrSocket: absent, nothing changes: fleetd's own $SHELL still decides and the
directory still goes to java.io.tmpdir, byte-identical to before.
Why it exists. The scrub read fleetd's own $SHELL to decide whether a member's login shell
honours ZDOTDIR, and wrote the generated directory into fleetd's own java.io.tmpdir. Both
describe the wrong process once panes run as another user.
The dangerous combination was fleetd on zsh and the member user on anything else. fleetd took the
zsh branch, generated a ZDOTDIR, and set it on the pane; the member's shell ignored ZDOTDIR
entirely, so no scrub ran — and because the zsh branch deliberately skips the sentinel overlay, the
fallback never happened either. No protection at all, on the path fleetd believed was the protected
one. Even with the member on zsh it still failed: java.io.tmpdir is mode 0700 on macOS, so
another uid cannot even traverse it.
The gotcha: it degrades, it never refuses to spawn. If the member shell is not zsh, or
worktreeRoot/worktreeGroup is missing, fleetd warns loudly and falls back to the sentinel
overlay. A weakened credential control must not become an outage for an opt-in feature. The
permissions are deliberately tight in the other direction: the directory is rwxr-x--- and each file
rw-r----- — group read, no group write anywhere, no world bits. A member sources what it needs
and cannot alter fleetd's own scrub.
An opencode member's config follows the member's own user
What. When memberHerdrSocket: puts member panes under a different OS user, the ephemeral
opencode.json no longer goes into fleetd's own temp directory. It goes under worktreeRoot, shared
read-only with worktreeGroup. Session discovery, which cannot work across users, now says so once
instead of failing quietly.
On. memberHerdrSocket: plus worktreeRoot/worktreeGroup. With memberHerdrSocket: absent,
both paths are byte-identical to before.
Why it exists. The launcher decided two paths from its own process. The config directory came
from java.io.tmpdir, which is mode 0700 on macOS, so another uid cannot even traverse it. That file
is the only way a member learns where the bridge MCP is. A member that cannot read it still starts
and still holds a pane, but never becomes deliverable: fleet_send waits about 60 seconds on the
readiness gate and then fails, and nothing in that message points at a temp directory.
Session discovery had the same shape with a different result. It read fleetd's own user.home to
find opencode.db, but under this key that file lives under the member's home. It would find
nothing, forever, and agentSessionId would stay null — which reads as "this backend cannot
resume" rather than "fleetd looked in the wrong home".
The gotcha: this one refuses to spawn, where the credential scrub above degrades. The difference
is what each thing is. A weakened credential control still has value, so the scrub falls back to the
sentinel overlay. This config file is not a control — it is the member's only route to the bridge. So
a missing worktreeRoot or worktreeGroup refuses the spawn and names the missing key. Nothing is
lost by refusing: that member would not have worked either way, and the later failure carries no
clue. Discovery is switched off under this key on purpose, with one WARN — an honest "not available"
beats an answer read from the wrong directory.
The model opencode actually ran is read back and checked
What. After an opencode member's session row appears, fleetd reads the model opencode actually resolved and compares it to what the profile asked for. On a real mismatch it logs an ERROR naming both models and quarantines that profile's credential through the existing exhaustion path.
On. Automatic, for opencode profiles that configure a model:. claude-code profiles are not
affected — they have no opencode session row, and the check is structurally unreachable for them.
Why it exists. opencode does not fail on an unknown -m. It silently falls back to a default.
The xf profile named a model that had been withdrawn from the catalogue, so every xf member ran
on a paid model at the expensive high variant while the profile was configured to be free. That
was 97.6% of dev spawns for about a day.
The second harm is worse than the bill. xf declared no credentialId, because it was supposed to
be free — so five concurrent members could hammer the shared paid credential while fleet_list
reported the paid profiles as barely loaded. The accounting built to protect that credential was
looking the other way. Nothing logged any of it; it was found because the operator happened to notice
the member's own UI naming the wrong model.
The general shape is worth naming: fleetd set a backend option by flag and treated the process
starting as proof the option took effect. #232 tracks the same assumption for autoCompactWindow.
The gotcha: the check runs late, and "unknown" is never a mismatch. opencode writes its session
row only when the session is first persisted, so a check inside spawn() runs before the evidence
exists and can never fire — an earlier attempt shipped exactly that and was closed. This one runs
from the same late-resolve step that fills in agentSessionId, so a resolved agentSessionId is
also the sign the model check ran.
The comparison is deliberately narrow, because a false positive stops all work on a credential and is
worse than the bug it prevents. The profile string and the stored JSON are in different formats
(openai/gpt-5.6-terra against {"id":"gpt-5.6-terra","providerID":"openai"}), and a profile may
carry no provider prefix at all. So the id is always compared, the provider only when both sides
have one, and absent, incomplete or unparseable evidence is UNKNOWN — never a mismatch, never a
quarantine.
The bigger gotcha: for a long time this check did not run at all for most spawns, and said
nothing (#267). checkModelMatch has exactly one call site, and it sits after the #249 gate that
returns early when the spawn was given no fleetd-provisioned worktree — which the code itself calls
"the ordinary, expected shape of the large majority of spawns". So the detector built after the xf
incident was switched off for most of the spawns it existed to protect, silently.
Since #267 that case is no longer silent. Once per profile, when a model is configured:
opencode model-mismatch check (fleetd #175) cannot run for profile 'X': it was spawned
without a fleetd-provisioned worktree (fleetd #249), so its cwd may be shared with other
sessions and the actual model it is running cannot be safely told apart from a sibling's —
spawn with worktree:true to enable the check for this profile.
The check itself still does not run there, and that is deliberate. The model can only be read via
actualModelForSessionId, keyed on the resolved session id. The only other lookup is
sessionIdForDirectory, the "newest row for this directory" heuristic #249 exists to distrust: on a
shared cwd it can return a sibling's row, so a sibling running a different, correctly-configured
model would look like this profile's mismatch and quarantine an innocent credential. A false
quarantine takes real capacity away on bad evidence, which is worse than failing to detect. Spawn
with worktree:true if you want the check.
Worth naming as a shape, because it is not the same one as the entry above: a safety check rode on the same return value as an identity lookup. #249 tightened the identity question for good reasons and narrowed the safety check as a side effect. Neither change was wrong alone — the coupling was.
Every file a member must read follows the member's own user
What. Two more places where fleetd used to write a file into its own process's filesystem and
then hand a member the path. The claude-code role/reply charter now goes under worktreeRoot
instead of java.io.tmpdir, shared with worktreeGroup. And worktreeRoot itself is now made
group-traversable where it is created, not only the worktrees inside it.
On. memberHerdrSocket: plus worktreeRoot/worktreeGroup. With memberHerdrSocket: absent,
the charter path is byte-identical to before; with worktreeGroup: unset, the root is untouched.
Why it exists. This is the same shape as the two entries above, found twice more. The charter one
became load-bearing without anyone noticing: the charter used to ride inline on
--append-system-prompt, and the file was the exception. #220 made the file the only delivery
path, always, to keep the launch command inside the pane's 1024-byte line. So under
memberHerdrSocket: every claude-code member would have been handed a charter path it could not
read. Measured against the real binary (claude 2.1.258), that is the loud failure — it prints
Error reading append system prompt file: EACCES and exits 1 before it touches auth or the network,
so the pane dies and the spawn fails at the readiness gate. Loud, but with no clue in the message.
The worktreeRoot one is the floor the other three stand on. A different uid needs the execute bit
on every ancestor directory, so a root at 0700 makes every carefully-shared child unreachable —
the member cannot read the opencode config, cannot reach its own checkout, and cannot read the
credential scrub. Files.createDirectories respects the umask, so whether this bites depends on the
umask of whatever shell started the daemon. It would work on the machine it was developed on and fail
on the next one, and the failure looks like a member that never becomes deliverable.
The gotcha: both refuse rather than degrade, and both are invisible until you turn the key on.
A charter is not a control, it is the member's turn contract — the rule that ends every turn with
fleet_reply — so a member that cannot read it is broken, not weakened. Missing config refuses the
spawn and names the key; an untraversable root refuses and names the root, its current mode and the
group. None of this changes anything on a host where memberHerdrSocket: is unset, which is why it
can look like dead code until the #185 rollout starts.
A role fleetd cannot bind is refused, not quietly downgraded
What. A spawn asking for role: architect on a profile that no configured slot carries now fails
at once. The message names the role, the profile, and the profiles fleet.architects does carry. The
roster also reports the role a member really holds, never the role that was asked for.
On. Automatic.
Why it exists. Such a spawn used to succeed. The member started, read the architect charter and
the architect agent definition, and then fleet_whoami told it — correctly — that it is a worker,
because no slot could be bound. Meanwhile GET /members still said architect. Three sources of
truth disagreed about one live member, and the only log line was INFO.
That is the quiet direction of failure, which is the worse one. A member acting above its role is refused by the authorization gate, so the mistake announces itself. A member acting below its role simply does not do the job, and the lead reading the roster has no way to see why. An explicit operator request became the silent case.
Refusing was chosen over binding anyway for one reason: an architect's identity is the slot it was bound to. With no slot, there is nothing to bind an identity to, so binding would mean inventing one. The operator's fix is one line of config, and the refusal names it.
The race is closed too, by reserving first (#226). A pre-check cannot close it: "does the config carry a slot" is stable, but "will a slot still be free after the launch" depends on a spawn that may not have happened yet, so widening the check only moves the window. Instead fleetd now reserves a matching slot before it launches, binds that reservation after, and releases it if the launch fails. A spawn that would lose the race is refused before a process exists, so no member is ever started on a charter it will not hold.
Two things were deliberately kept. The old acquired fallback and its WARN stay, because a
reservation narrows the window and claiming it removes the window is the mistake to avoid. And a
failed launch must return its reservation, or the pool shrinks silently with every failure — that is
the way this shape usually goes wrong, so it is the case the tests pin.
The gotcha: a full architect pool now fails your spawn instead of demoting it. fleet_spawn{role: "architect"} used to hand back a plain dev when every slot was taken. It now throws. That is the
point — a refusal is recoverable and a mismatched charter is not — but it is a visible behaviour
change for anyone who relied on the old silence. The refusal stays scoped to architect alone: dev
and reviewer pools are placement candidates, not identity bindings, so an explicit profile outside
them stays a supported override.
A backend that starts erroring cools off, instead of swallowing the next spawn
What. When two different members fail inside 60 seconds with output matching a profile's
backend-error pattern, fleetd treats it as one incident rather than two failures. The credential
those profiles share cools off for 60 seconds. Automatic placement skips it, an explicit
fleet_spawn naming it is refused before the backend adapter is ever called, and the lead gets
exactly one nudge in its own pane naming the credential and the affected members. fleet_list and
fleet_profiles report coolingOffForSeconds, and a cooling profile's free drops to 0.
On. The correlation and the cool-off are automatic. The pattern is the knob:
profiles:
terra:
errorPattern: "503 Service Unavailable" # optional
A profile that sets nothing still gets the built-in compatibility pattern, so this is never silently
off. A malformed regex is rejected at config load, naming the profile, the key and the parser's
own message — and since fleetd #273 that covers exhaustedPattern too, not only errorPattern. Both
keys are checked in one pass, so a config with a bad value in each is told about both at once.
Before #273 a bad exhaustedPattern passed load and then crashed the daemon at startup with a raw
PatternSyntaxException naming neither the profile nor the key. The key is deferred: a running
daemon keeps the pattern it started with, so a change needs a restart.
Why it exists. On 2026-09-01 one backend outage killed both running members mid-turn. fleetd
handled each one correctly and separately, and never noticed it was one event. Both tickets were
reported honestly, both members showed as done and idle — the same rows a clean finish produces —
and capacity still advertised a free slot. A third spawn onto the same credential would have died the
same way. The evidence that it was an outage only exists across members, so a per-turn view
cannot see it however correct each turn's own handling is.
The pattern is per profile and configured for the same reason exhaustedPattern is: the previous
mechanism was one hard-coded API Error: string, which fails in the silent direction. When a
backend rewords its error the string stops matching, and a failed turn goes back to resolving as a
success. Wording belongs next to the backend definition that produces it.
The gotcha: cooling off is not quarantine, and the two can be true at once. Quarantine (CB-578)
means the backend said it is out of capacity — a long, 1800s-default cooldown. Cooling off means a
credential is erroring right now — short, fixed at 60s, and deliberately not configurable per
profile. They are reported as independent fields precisely so an operator can tell "spent" from
"flaky at the moment", and either one alone already forces free to 0. Read the refusal wording
too: "cooling off after repeated backend errors" is a different situation from an exhaustion refusal,
and they have different fixes.
One error changes nothing, on purpose. One member failing repeatedly is a member problem; two
different members failing together is a backend problem. Correlation keys on the credential, not
the profile, because the credential is the thing that actually runs out — so cooling terra also
cools sol when they share one.
A claude-code member no longer blocks forever on the workspace-trust dialog
What. Before starting a claude-code member in a worktree fleetd provisioned, fleetd marks that
directory as trusted in ~/.claude.json (or the profile's configDir). The member reaches idle
instead of sitting on a prompt nobody can answer.
On. Automatic, for kind: claude-code members, and only in a worktree fleetd itself
provisioned. Any failure is logged at debug and never blocks a spawn.
Why it exists. Claude Code asks for confirmation the first time it opens an unfamiliar directory. Every provisioned worktree is unfamiliar by construction — it is a fresh path with a nonce in it. The member came up, printed a dialog, and waited. Nothing could answer it: the lead cannot see a member's screen, and driving the pane directly is exactly what the bridge exists to prevent. The spawn then failed as a readiness timeout, which points an investigation at the wrong thing entirely.
The rejected fixes are worth naming, because each looks reasonable: --dangerously-skip-permissions
turns off a real control for a whole session; permissions.additionalDirectories widens what the
member may touch rather than answering the question asked; running non-interactively gives up the
pane the whole design depends on. Seeding the one flag the dialog sets is the narrow answer.
The gotcha: it writes a file fleetd does not own, and the guard is the worktree check. While this
was being built, a mutation test aimed at the seeding logic overwrote the operator's real
~/.claude.json, shrinking it from 72581 bytes to 919 and taking the account entry with it. That is
why the write is gated on the directory actually being a fleetd-provisioned worktree rather than on
the path merely being set — a default-resolved cwd is the operator's own checkout. The write is
additive and atomic, and a process-wide lock serialises concurrent spawns.
The second gotcha: the lock cannot reach the operator's own Claude Code (fleetd #247). That
in-process lock only serialises spawns inside one JVM. The operator's own running Claude Code writes
the same file, and on this host it is literally the same file: opus and sonnet both set
configDir to the lead's own config directory. So this was never a rare collision — it was a lost
update on every ordinary spawn of those profiles.
The write is now a compare-and-swap with a bounded retry. fleetd reads the file's bytes, builds its update, then re-reads the bytes immediately before the atomic move and compares. If they changed, another writer got in, so it throws its work away and rebuilds from the fresh bytes — up to five times.
On exhausting the retries it writes nothing and logs a WARN naming the cwd and the file. That is deliberate. An unseeded member shows the trust dialog and fails to reach an injectable state: visible, logged, and recoverable by retrying the spawn. Writing a stale copy over the operator's live config is silent and not recoverable. The code fails toward the recoverable outcome.
A second WARN now fires every time the seed is about to target the default ~/.claude.json, which
happens only when a profile sets no configDir. That is the one path reaching the operator's home
file, and it is the default, so a profile that simply forgot the setting used to get no signal at all.
What this still does not fix. The window is narrowed, not closed. A write landing between the final re-read and the move itself is still lost. There is no operating-system compare-and-swap on a plain file, only this cooperative narrowing. Do not read the CAS as making the race gone.
The third gotcha: under memberHerdrSocket the seed used to write the wrong home (fleetd #285).
When memberHerdrSocket: is set, the member pane runs as a different OS user with its own $HOME.
The seeding code could not see that at all — it was a static method, so it had no access to the
launcher's own memberHerdrSocketConfigured() check. With configDir unset it therefore wrote
fleetd's own ~/.claude.json — the operator's real file — while believing it was seeding the
member's. The member never got a seed, and the operator's file was edited for nothing.
The sibling method one line below, writeCharterFile, had already solved this: it refuses the spawn
when it cannot place a file where a different-uid member can read it. The trust seed now does the
same. Under memberHerdrSocket it requires both configDir (so there is a member-readable
target at all) and worktreeGroup (so the 0600 file it writes can be shared), and refuses the
spawn — naming which key is missing — before touching any file. When both are set, the written file
is chgrp'd and chmod'd to rw-r----- for that group, so the member's OS user can actually open it.
The refusal is deliberate rather than a degradation. A member with no readable trust seed is not
"slightly worse": it sits on the interactive dialog forever and never calls fleet_reply, which is
the exact failure this whole feature exists to prevent. With memberHerdrSocket absent — every
live fleet today — nothing about this changed.
A worktree tells you which of its config files are stubs
What. A provisioned worktree neutralizes .mcp.json, opencode.json and .autoenv: the copy in
the worktree is a stub, not the repo's committed file. The daemon now logs which ones it actually
neutralized and which were absent, and — the part a member can reach — records the list in
worktree-scoped git config:
git config --worktree --get-all fleet.neutralizedConfig
git config --worktree --get fleet.neutralizedConfigNote
The parity overlay reports itself the same way, naming what it copied and what it marked
--skip-worktree.
On. Automatic, for every provisioned worktree.
Why it exists. Both mechanisms were completely silent. A member opening .mcp.json saw a
plausible file and had no way to know it was a stub — so an edit to it was real work that could never
be committed, and nothing said so. The daemon side was no better: nobody could tell from the logs
whether a given member had been handed an .envrc at all, and a silent copy is what makes that kind
of problem hard to notice in the first place.
Both log lines carry a denominator — neutralized 1 of 3 configs: .mcp.json (opencode.json absent, .autoenv absent) — because copied 1 on its own reads as success whether the candidate list
had one entry or ten.
The gotcha: worktree-scoped config is invisible to git status, and that is why it was chosen.
It lives in .git/worktrees/<nonce>/config.worktree, never in the working tree, so it cannot show up
as a pending change or get swept into a commit. A marker file in the worktree would have needed a
gitignore entry, a name-collision check, and would still read as "this IS the file" to anything that
just opens it.
The parity overlay's default changed with this work: it is [.env] now, not [.env, .envrc]. .env
is data, so copying it can only move values; .envrc is executable shell that direnv runs on every
cd, so copying it moves behaviour. An operator who wants it can still list it explicitly, and then
owns that choice.
The credential probe asks the daemon what the policy is
What. scripts/probe-member-credentials.sh reads the policy's name list from the daemon over a
new read-only endpoint, GET /member-credentials, instead of carrying its own copy of the names. The
endpoint returns the policy mode, the known and allowed names, and the counts. Never a value —
the daemon does not hold the values, and MemberCredentialPolicyView reads no environment at all, so
there is nothing to redact by construction.
On. Automatic. The probe needs no flag; it refuses to run outside a member shell unless you pass
--allow-outside-member, which labels the reading as a comparison rather than a finding.
Why it exists. The probe used to hold a hardcoded array of 31 names. The live policy had 34. It reported 26 blocked against a policy that blocks 29, exited 0, and printed a table that looked complete. Both numbers were right; they were counting different sets, and nothing said so.
That is the worst direction for a verification tool to fail in. Its clean output is taken as evidence, so a silent gap stops anyone looking. The same "second hand-maintained copy" defect had already been fixed twice elsewhere, and the fix here is the same one: delete the copy rather than correct it.
The gotcha: it refuses rather than degrades. An unreachable daemon, an absent or empty policy, or
a knownCount that disagrees with the length of known[] all exit non-zero. There is deliberately
no local fallback — a checker that quietly drops to a weaker check is the thing this replaced.
It also prints its own denominator: "policy contains 34 name(s); this run checked 34 — they match." Two numbers a reader can compare beat one number they have to trust.
The startup log line and the endpoint now share one class, so the counting exists in one place. Watch
one subtlety if you ever recompute it: blocked is not known - allowed. Live, known is 34 and
allow is 7, but only 5 of those 7 appear in known, so blocked is 29 and the naive subtraction
gives 27.
An allow-list policy refuses a spawn it cannot enforce
What. memberCredentials.policy: allow-list is enforced by a generated .zlogin under a
per-member ZDOTDIR. A non-zsh login shell ignores ZDOTDIR entirely, so no scrub runs. Under that
policy the launcher now refuses the spawn, naming the shell it actually found and offering three
ways out. Under policy: deny-by-default nothing changed — that overlay is applied to the pane
before any shell runs, so it does not depend on the shell.
On. Automatic whenever policy: allow-list is set.
Why it exists. The launcher already detected a non-zsh shell and logged a warning — then degraded to the weaker overlay and spawned anyway. The operator had explicitly asked for the blocking control and silently received the weaker one. Detection existed; refusal did not, and a control that silently does nothing is worse than no control, because the config still says it is on.
The gotcha: memberLoginShell: and memberHerdrSocket: must be set together. Which shell is
checked depends on the routing. With memberHerdrSocket absent — today's normal mode — fleetd's own
$SHELL decides. With it set, member panes run as a different OS user, so fleetd's $SHELL
describes the wrong account and only the configured memberLoginShell: is read. If you set
memberHerdrSocket and leave memberLoginShell unset, the shell reads as <unset>, counts as
non-zsh, and every spawn is refused. The refusal message says so, but it is cheaper to know first.
What a member can actually reach — the honest boundary
What. This is not a feature. It is the frame every credential setting on this page sits inside, and it has been missing, so people have read those settings and concluded more than the settings say.
A member runs as the same OS user as the lead. That is a trust model, not a sandbox. Inside one uid, ordinary Unix file permissions give no confidentiality boundary. A member that can run commands can read any file this user can read.
What each control actually does:
| Channel | Control today | What it really buys |
|---|---|---|
| Environment | memberCredentials + the ZDOTDIR scrub |
Real, for what the member inherits. It does not protect the values — the member can source the same files again. |
| Files | worktree provisioning neutralizes some config files | Not a boundary. File modes do not separate processes with the same uid. A member can read the original by absolute path. |
| Sockets | sshAgentEnv: omit |
Not a boundary. It omits one variable. It does not revoke access to the socket, and it cannot remove a readable key file. |
| Git config | the worktree's HTTPS rewrite + cleared credential.helper |
Routing, not enforcement. Normal git commands go the intended way; a member can still call ssh directly. |
| Process table | secrets kept out of argv | No boundary. argv is visible to any local user, and a different uid would not fix that. |
Why it exists. The proven path is short: the scrub blanks SSH_AUTH_SOCK → git starts ssh → ssh
reads the user's config → it opens an IdentityFile that is readable and has no passphrase → the
forge accepts the operator's key. Measured on 2026-08-28: a member with the socket blanked pushed
successfully. The gate closes lead → child environment inheritance. It never closed member →
filesystem → credential.
That is worth naming as a shape, because it keeps recurring here: a gate written after an incident closes only the direction that incident came from. Ask which states open it, not just which it blocked.
The gotcha: no meaningful boundary exists inside one uid. Environment scrubbing, worktrees and
config rewrites reduce accidents, and that is worth having. They do not contain a member that
chooses to look. A real boundary needs a different OS user (memberHerdrSocket, above — but see
the next entry: fleetd cannot verify that one for you) or OS-level
confinement such as a container or VM — a different execution model, not a longer name list. Even a
different user does not protect argv, and does not protect shared git objects where worktreeGroup
grants access to them.
Open work, ranked, is on #184.
The ssh-agent setting is named after what it does — sshAgentEnv: omit
What. memberCredentials.sshAuthSock: block|allow is now
memberCredentials.sshAgentEnv: omit|inherit.
The knob. All eight spellings parse, silently, with no deprecation warning:
memberCredentials:
sshAgentEnv: omit # canonical; "inherit" passes SSH_AUTH_SOCK through
# sshAuthSock: block # the old key and old values still work, unchanged
Old configs need no edit. If both keys appear, sshAgentEnv wins. An unrecognised value
normalises to omit, so a typo fails closed rather than handing a member the agent.
Why it exists. block named an effect fleetd does not have. It omits the variable from the
member's environment; it does not deny access to the socket. A member runs as the same OS user, so
the socket stays reachable and the path is discoverable. An operator scanning a config file reads a
value name and stops — that is the point of a good one — so block was actively misleading, and the
honest explanation sat in a comment most readers never reach.
What the setting still buys is real and worth keeping: a member does not pick up the operator's agent by default, which removes a whole class of accident. It is an accident-reducer, not a deny.
The gotcha: the shim is load-bearing, not cosmetic. fleetd.yaml is gitignored, so no worker
can see it and no test in the repo covers it. The live config on this host still says
sshAuthSock: block. A rename without the read-both shim would have silently dropped the setting at
the next restart — trading a naming bug for a real credential regression, in a file the test suite
cannot reach. This was verified against the running daemon rather than argued: after redeploy,
GET /member-credentials still reported policy: allow-list with 39 known / 7 allowed / 34 blocked,
and a real member spawned on the new jar reported SSH_AUTH_SOCK: unset.
fleetd states the member trust model at startup
What. Every boot, fleetd logs one INFO line saying what kind of boundary members run inside.
The line changes with memberHerdrSocket. With the key unset:
member trust model: members run as the same OS user as fleetd, not in a sandbox. A member can
read any file this user can read, including SSH keys and credential stores, whatever
memberCredentials says. To add a real boundary, route members to a second herdr under a
different OS user with memberHerdrSocket.
With the key set, it says instead that members are routed to a separate herdr, and that fleetd cannot see that herdr's uid, so the operator must confirm it runs as a different OS user before treating it as a boundary.
The knob. None. It always logs, at INFO, from Fleetd.reportMemberTrustModel, right after the
startup secret report. Find it with:
grep 'member trust model' fleetd/fleetd.out | tail -1
Why it exists. The entry above documents the honest boundary, but a wiki page only reaches
someone who goes looking. memberCredentials reads like a security control, so an operator who
never opens this page reasonably assumes it contains a member. It does not. Putting the trust model
in the boot log means the claim arrives unprompted, in the same place the operator already checks
startup secret lines, on a machine where the config is whatever it is.
The gotcha: this line is honest about the uid, and one older WARN is not.
HerdrPeerLauncher.warnUnknownMemberEnvironment still says a configured memberHerdrSocket means
"member panes run under a different OS user". fleetd cannot see the uid at the other end of a unix
socket — it infers that from the key being set. The direction that costs something is a second
herdr running as the same user: members then inherit fleetd's environment, and the credential
gap the WARN reports as UNKNOWN is in fact exactly the count it just told you to disregard. A known
exposure reported as an unknown one. Tracked as item 5 on #184.
leadSeats reports the seat the lead itself holds
What. For a subscription: true profile, fleet_list's row carries a leadSeats field when the
lead's own session occupies a seat on that same account. Profiles are grouped by account, not by
profile name: a profile with no explicit credentialId that is subscription: true joins a shared
<subscription> group. An explicit credentialId still wins, so two genuinely separate Claude
logins on one host stay apart.
free is not reduced by leadSeats. free means one thing only: how many slots a fresh
fleet_spawn on that profile will actually be granted right now, which is max(0, maxLoad - live) —
the same check the spawn gate runs. leadSeats is a fact reported beside it, for a caller to use
however it likes.
On. Automatic for subscription: true profiles, once fleet.leaders.<name>.profile is set.
Why it exists. The lead is always a live claude session on the operator's subscription, and
maxLoad only ever counted members. So a fan-out that filled every member slot still left the lead's
seat unaccounted for, and the daemon advertised a slot that the account was already using.
The same grouping fixes a second, wider bug: quarantining one subscription profile now also refuses spawns on the other. One subscription hitting a usage limit really does take out every profile running on it.
The gotcha, and how it was resolved (fleetd #257). The first cut of this did subtract
leadSeats from free. That shipped, and was corrected the same day. The placement gate never
consulted the lead seat, so with maxLoad: 3 and two members up, fleet_list said free: 0 while a
third spawn still succeeded. A lead that believed the number gave up a slot the daemon would have
granted — the opposite of the overstatement this was filed to fix.
The subtraction was removed rather than pushed into the spawn gate, for two reasons. Making the gate
subtract it would change what maxLoad: 3 means in every existing config file, on every host. And no
backend seat ceiling shared with the lead has ever been measured: a test on 2026-08-29 ran three
concurrent interactive claude sessions with no trouble. Subtracting the seat therefore described an
accounting policy, not a constraint, and dressing a policy as a capacity fact is what made the two
numbers disagree.
Do not raise maxLoad to "win a slot back". Nothing takes one; you would be granting a real extra
member.
The first attempt at this shipped inert and every test passed: the matcher grouped by
effectiveCredentialId(), which fell back to the profile's own name, so a lead on opus and
members on sonnet never matched and zero seats were charged. Every test in that change put the lead
on the same profile name as the target — the one shape the live config does not have.
The REST face — the operator's way in when MCP is not there
What. The daemon serves 15 HTTP routes on the same port as the MCP mount. Every capability
behind an MCP tool is reachable there, so the fleet can be driven with plain curl. The full
reference — each route, the Authz role it needs, and its matching MCP tool — is
REST API Reference.
On. Always on. The listen address is the bind: block in fleetd.yaml (host and port);
this host uses 127.0.0.1:8765. GET /metrics is the one route that can be absent — it is
registered only when metrics are configured.
Why it exists. MCP is the agent channel, but an operator needs a way in that does not depend on
an agent session being healthy. A lead whose MCP mount has dropped cannot call a single fleet_*
tool, and that is exactly the moment you most need to see what the fleet is doing. REST is that
door, and it is also what a dashboard or an acceptance-test script would use.
The gotcha: GET /sessions/{id}/replies destroys what it returns. It drains the inbox, so the
first read is the only read. Send it to a file. Piping it through head, or any script that exits
early, loses the payload for good — and this is the route you reach for when a ticket has timed out
and a member's real answer is sitting in that inbox. There is no second copy.
It is not an agent channel, and members should not be pointed at it. Not because it skips the
authorization gate — it does not. Every route except /healthz resolves the caller through the same
CallerResolver and the same Authz table MCP uses, so a member calling REST gets a member's
rights. The reason is narrower: identity is resolved from the connection, and a member's own child
process is a connection the daemon must reason about. That has been wrong before — a member's curl
child was resolved as the primary until it was fixed. One channel for agents is one place to get that
right. CLAUDE.md therefore says nothing about REST, deliberately, and should keep saying nothing.
Why there is no route table on this page. There is already one, on chapter 15, and a second copy
is the defect — not the errors it accumulates. Chapter 15 was written on 2026-08-31 and was accurate
for all 14 routes that existed then. GET /member-credentials shipped three days later and the page
did not follow it, which is how the drift starts. A test in the repo now enumerates the routes
FleetApp registers and fails when they no longer match its inventory, naming chapter 15 as the page
to update.
A note for anyone briefing a worker to read this page
A worker cannot see the current version of this file. wiki/ is a submodule, and the parent
repo's recorded pointer is deliberately never updated (committing it is on the never-commit list). So
git submodule update --init in a worker's worktree checks out a months-old snapshot. During
this very backfill a worker reported that an entry "does not exist anywhere in the repo" when it had
been on the page for hours. Paste the relevant text into the brief instead of pointing at the file.
A dead member's seat comes back
What. A member whose backend errored (BACKEND_ERROR) or whose turn failed (FAILED) no longer
counts against its profile's maxLoad. Its row stays in fleet_list so you can still see what
happened, but a fresh fleet_spawn on that profile is granted.
On. Always on; no configuration.
Why it exists. Nothing ever removed these sessions, and the live counter had no state filter at
all — it counted every roster entry. That counter is the one the real spawn gate reads
(CompositePeerLauncher.enforceMaxLoad), so one backend error took a seat and never gave it back.
Only an explicit fleet_stop on that exact pane, or a daemon restart, freed it.
The harm ran in the direction that hurts. On a maxLoad: 1 profile — opus and sol on this host —
a single backend error put the profile out of service for good, and fleet_spawn answered "at
maxLoad: 1 live >= 1 cap; refusing spawn — no fallback to another profile". The lead saw a capacity
refusal with no reason to suspect a dead seat. It also outlived the cooldown that was supposed to be
the remedy: quarantine and cool-off both expire on their own, the dead seat did not, so the profile
was still refusing long after the credential recovered.
The gotcha: free and reclaimable count different things, and a freed seat belongs in only one
of them. reclaimable means "this member holds a seat and has no open bridge work — stop it and
you get the seat back". A terminal session holds no seat any more, so its seat is already in free.
Counting it as reclaimable too would report the same seat twice, and free + reclaimable would
read as more capacity than maxLoad allows. The ticket asked for exactly that, and that half of the
ticket was wrong. The dead session is still visible without it: its roster row carries
state: "backend_error" or "failed", which is what tells the lead to stop it.
Both places that report reclaimable — the per-member flag and the per-profile count, which travel
in the same fleet_list response — now call one shared predicate, so they cannot drift apart. A test
runs it over every MemberSession.State value, so a state added later cannot slip through
unconsidered.
What this does not do. Terminal sessions are still never reaped. They accumulate in the roster until stopped or until the daemon restarts. That is now cosmetic rather than a capacity loss, and it is deliberate: the roster entry is the only record of what went wrong.
A member's teardown no longer leaks a worktree or a branch
What. Two cleanup paths in SessionManager were one-sided, and both are closed. Stopping a
member whose worktree cannot be removed now completes the stop and logs a WARN instead of throwing.
A spawn that fails after its worktree was created now deletes the orphaned branch as well as the
worktree.
On. Always on; no configuration.
Why it exists. Every other cleanup step in release() was wrapped in a try/catch — the last one,
removing the worktree, was not. By the time it ran, the session was already out of the registry and
the pane already stopped, so a throw there escaped release() with the teardown in fact complete.
There was no retry path: a second stop on that pane is a no-op. The caller saw a failed stop for a
session that was gone, and the directory leaked with nothing left to point at it. git worktree remove --force can throw for ordinary reasons — a stale index lock, a slow filesystem, its own 30
second timeout — so this was not exotic.
The second leak is the same shape as the one fixed one layer down earlier: GitWorktrees already
deleted both the worktree and the branch when provisioning failed inside add(). The catch
that covers failures after add() returns — the parity overlay, the group share, the spawn itself
— removed only the worktree. A spawn failure there is routine (a quarantined credential, a backend
refusal), so every occurrence left an orphan worker/<slug>-<nonce> branch behind.
The gotcha: only the failed-provisioning path deletes a branch. A normal release deliberately
keeps the branch, so a worker's committed work can still be recovered after its pane is gone. Getting
these two backwards would destroy real work, so the tests pin the separation from both sides: the
normal-release test asserts the branch delete is never called, and the failed-spawn test asserts
the deleted branch is the exact one add() created.
A third door into the same leak, found later. Tearing a member down runs three steps: close the
pane, close its tab, then release the generated credential-scrub directory. The first and third were
wrapped; closing the tab was bare. By the time that step runs the pane is already gone, so a failure
there is cosmetic tidying — but it threw, and the throw travelled up into the release path before
the worktree removal that the fix above wraps. The session was already deregistered by then, so a
second stop is a no-op: same unrecoverable leak, plus the scrub directory and its allowed N of M
report. Closing the tab now logs a WARN naming the tab and carries on.
Closing the pane still propagates its failures, and that difference is the point. That step is the one whose failure means the teardown may genuinely not have happened, so reporting it as done would be a lie. The tab is not.
A chained second question no longer kills its own ticket
What. A member that calls fleet_ask twice in one turn — asks, gets an answer, then asks again —
keeps its ticket. Before, the second question silently ended the ticket, and the lead's fleet_poll
returned nothing for a member that was still working.
On. Always on; no configuration.
Why it exists. Each fleet_ask moves the async task to a new turnId. The thread answering the
first question then finished its work against the old turnId. Whether that killed the ticket
came down to which of two lines ran first — a race, not a decision. Answering now registers its
waiter before it resolves, and drops the stale key explicitly rather than relying on the order.
The gotcha: the guard that looks load-bearing is not. The fix also added a state guard on the
answering path, and the pull request claimed the ticket would still die without it. It would not: a
mutation removing only that guard left the test green, because ask() already moves the task to the
new turnId before the answering thread wakes. The guard is kept as defence in depth — it mirrors a
sibling guard whose own comment warns that this ordering is not something to rely on — and the
measurement is written into the code next to it, so nobody deletes it as dead code without
re-checking the ordering, and nobody trusts it as the only protection either.
A ticket is not stranded when a dead member's question lapses
What. A member whose pane dies while it is parked in fleet_ask no longer leaves its async
ticket at PENDING forever. About two minutes after health first classifies the member as GONE or
NEVER_READY, one delayed re-check sweeps the ticket to FAILED, so fleet_poll gives the lead an
answer instead of silence.
On. Always on when fleet health is enabled; no configuration.
Why it exists. An earlier fix closed the definite teardown road: an explicit fleet_stop or
the idle reaper now fails an ASKING ticket immediately. Health deliberately did not, and that is
right — a GONE reading is a guess from the live agent list, not a teardown the daemon performed,
and a guess must never kill a ticket whose worker a live lead could still answer.
But that left a second road to the same dead end. Health fires its sweep once, on the transition
into GONE, and skips the ticket because the member really is still asking. Between 55 and 115
seconds later the member's own fleet_ask lapses and clears its question — and now nothing fires
again. The tick loop returns early when the state has not changed, and the classifier answers GONE
before it could ever reach the orphan state. The ticket sat at PENDING for good.
Nothing else rescued it. A member parked in fleet_ask is BUSY, and the idle reaper only ever
touches READY or DONE sessions, so a GONE-but-never-stopped member is never released and the
teardown fix never runs for it.
The gotcha: this is one extra attempt, not a retry loop. An earlier attempt at this called the
sweep once per tick for as long as a member stayed terminal, and that was rejected. So the
transition schedules exactly one delayed follow-up. The delay is 120 seconds, chosen to clear
the 115-second worst case of the ask window; because the ask always starts before the GONE reading,
that ordering guarantees the question has lapsed by the time the re-check runs.
Two independent things keep the re-check from doing harm. It still passes sweepAsking: false, so a
member that is genuinely asking again is skipped exactly as on the first attempt. And it fires only
if the member is still classified in that same terminal state — a member that recovered, or was
released and dropped from the roster, is left alone rather than having some brand-new unrelated turn
failed underneath it.
One thing to know for maintenance. The 120-second delay is a constant in the code, not a config
knob. If either ask ceiling (FleetMcp's or the REST face's, both 115 seconds today) is ever raised
past it, this delay must be raised with it, or the re-check fires while the question is still open,
finds nothing to sweep, and the single attempt is spent.
The REST face reports what the MCP face reports
What. GET /profiles now carries the two backend outage states that fleet_profiles has always
reported: quarantined (the backend said it is out of capacity) and coolingOff (that credential
threw repeated non-exhaustion errors). Each names the credentialId and the seconds remaining. The
two checks are independent, so a profile can appear in both maps at once, and each map is present
only when at least one profile is in that state. Separately, GET /agents and GET /members now
map a herdr transport failure into the same {error, detail} envelope every other route uses,
instead of letting it escape as a bare 500.
On. Always on. Nothing to configure.
Why it exists. REST is the door a lead falls back to when its MCP mount drops — listMembers
says so in its own comment. So the weaker door was weakest exactly when it was load-bearing. A lead
on the fallback during an outage could not tell "busy, frees in 40 seconds" from "broken", and a
monitoring client that parses the JSON error envelope broke at the moment it was asking why the
fleet looked unreachable. Neither gap ever caused a wrong spawn: a REST spawn goes through the same
placement gate, so a refusal was still a refusal. These were visibility gaps, not state defects.
One thing to know for maintenance. Both doors render the profiles body from one method,
FleetMcp.profilesView, and are handed the same BackendQuarantine/BackendOutagePolicy
instances. Both halves are needed and neither is sufficient. Sharing the instances stops the doors
reading different facts; sharing the builder stops them reporting those facts differently. The first
version of this change shared the instances and copied the loop into FleetApp, which is the fleetd
#284 shape — one rule in two places, where the next edit lands on one and not the other. If you add
a field to that body, add it in profilesView and both doors get it.
A released reply goes back to the broker instead of being dropped
What. When a member's session is torn down, AmqpReplyInbox.release now cancels that target's
consumer and then nacks every delivery it still holds with requeue, so an undrained reply stays
recoverable. Before, it dropped its local record and left the message unacked on a still-open
channel — invisible to fleet_poll and peek, and freed only when the whole AMQP connection
happened to drop.
On. Always on, for the durable (AMQP) inbox. The in-memory inbox is unaffected and correctly still drops on release, because it is soft state with no broker behind it.
Why it exists. This was silent loss of the one artefact the bridge exists to carry. The path was ordinary: a member replies with no live waiter, so the reply is held; the lead sends that target a fresh task, which clears the stranded-reply flag but never drains the actual entry; the session is later stopped, so the teardown skips recovery because the flag says there is nothing to recover, and then drops the entry. The lead sees a member that finished and reported nothing, with no way to tell that from a member that genuinely never replied.
One thing to know for maintenance. The order is cancel first, then nack — not the reverse. A
nack-with-requeue while the consumer is still attached hands the message straight back to that same
consumer as soon as a prefetch slot frees, racing release's own cleanup and leaving a stale entry.
A delivery tag stays valid for basicNack on an open channel whether or not its consumer is
attached, so cancelling first costs nothing. This was found against a real broker, not by reading
the code, and it is why AmqpReplyInboxContractTest needs Docker.
A failed spawn no longer leaves a live pane behind
What. Two spawn-side exits created a pane and left it running when the spawn failed: the readiness gate propagated an unrelated herdr error without teardown, and pane placement never closed the pane it had just split when the peer failed to start. Both now close the pane, best-effort, and still hand the caller the original exception unchanged.
On. Always on.
Why it exists. Nothing downstream could clean these up. SessionManager.acquire never learns the
pane id — spawn throws before it returns one — so its failure path removed the worktree and the
branch and could do nothing about the pane. The result was a live backend process with a deleted
working directory, absent from fleet_list, holding a real backend seat that nothing decremented.
The pane-placement leak was the worse of the two: with no named agent started, the orphan reaper
cannot see it either, so no restart ever reclaimed it.
One thing to know for maintenance. The readiness gate still refuses to interpret an error it
does not recognise — it rethrows it unchanged, and a test pins that. Closing the pane and converting
the exception are different things, and only the first was missing. Do not "simplify" this by
mapping the error to PeerUnreachableException; that reintroduces the exact behaviour #176 wrote
failFastOnGoneBackend to keep distinct.
A reply with no content is refused instead of resolving the waiter
What. POST /sessions/{id}/reply read the content field with a silent default, so a body that
omitted the field became an empty reply. That empty reply resolved the lead's waiter and the turn
completed. The required-content check now lives in MessageService.reply, which both doors call, so
the REST door returns 400 bad_request and the MCP tool returns a tool error. Blank and
whitespace-only content are treated the same as missing.
On. Always on.
Why it exists. The lead could not tell an empty reply from a member that genuinely said nothing. This project already has several real conditions that look like that — a clipped completion scrape, a backend that died mid-turn — so the bad case hid among them. The guard was written in one handler instead of in the thing both handlers call, which is the same drift shape as #284 and #297.
One thing to know for maintenance. fleet_reply had a smaller version of the same hole:
its guard checked content == null, not blank, so whitespace went through. Moving the check into
MessageService.reply closed that too, but it also made the shared method throw. replyHandler in
FleetMcp is a bare BiFunction with no try/catch, so that throw would have escaped the tool
handler as an uncaught exception. The tool's own guard was widened to isBlank for that reason, and
FleetMcpTest.replyWithBlankContentIsACleanToolErrorNotAnUncaughtException pins it. If you ever move
the check again, check the handler between the door and the service, not only the two ends.
The member REST routes name a herdr failure instead of returning a bare 500
What. POST /members and DELETE /members/{paneId} were the last two routes in FleetApp
with no catch (HerdrException). FleetApp registers no Javalin exception mapper, so the exception
escaped as the default 500 with the body Server Error — no herdr code, no herdr message. Both now
go through the same herdrError mapper every other herdr-calling route uses: 404 when herdr says
the target is gone, 502 otherwise.
On. Always on.
Why it exists. fleet_spawn and fleet_stop already caught the same exception and reported a
named error, so the two doors disagreed about the same failure — the shape #297 exists to remove.
The stop route is the one that mattered: SessionManager.release deregisters the session, notifies
the release listener and preserves a dirty worktree before it calls launcher.stop, so a throw
from that stop arrives after the teardown the caller asked for has already happened. A 500 told the
caller to retry and gave it nothing to reason about.
One thing to know for maintenance. An already-gone pane is not one of these failures.
HerdrPeerLauncher.stop treats a *_not_found from pane.close as success, so a double stop still
returns 204, and stopToleratesAnAlreadyGonePane pins that. What propagates is any other herdr
failure — a transport error, or a code like herdr_busy. No new tests were added: two tests already
covered these paths and asserted 500, and neither was about the status code (one guards that a
failed spawn still closes its tab, the other that a failed teardown is not reported as 204). Their
expectations moved to 502 and gained a body check.
One definition of loopback, so a worker cannot become the lead
What. ConnectionIdentity and CallerResolver each kept their own isLoopback, and the two
disagreed: the identity resolver accepted only 127.0.0.1, the authorization check accepted all of
127.0.0.0/8. There is now one predicate, on ConnectionIdentity, that the other calls.
On. Always on. It matters most in auth.mode: loopback-trust, which is the default — auth: is
commented out in fleetd.example.yaml, and the daemon logs the mode at every boot.
Why it exists. The disagreement was a privilege escalation. A caller from 127.0.0.2 had its
identity resolution skipped, so it carried no terminal; CallerResolver reads a missing terminal as
"not a worker", and a same-host non-worker is the primary. A worker could take spawn, stop, send and
drain. The skip also happens before the PID ancestry walk, so the defence that stopped a member's
curl child being read as the primary was bypassed as well.
One thing to know for maintenance. Do not "tighten" ConnectionIdentity.isLoopback back to
127.0.0.1. It reads like the safe direction and it is the opposite. That predicate does not
decide whether a caller is trusted — it decides whether a caller's identity is resolved at all, and
resolution is what demotes a worker. Every address excluded there is an address on which a worker
becomes the lead. The range matters because on Linux the whole /8 is bound to lo, so a source of
127.0.0.2 is bindable; measured on the fleet host with curl --interface 127.0.0.2 returning exit
7 (connect refused) rather than 45 (bind failed). On macOS the same command fails at the bind, which
is why no test binds a real 127.0.0.2 source — it would fail for every developer on a Mac.
The idle reaper no longer stops a worker that just got a turn
What. reapIdle decided a session was idle from a roster snapshot and then released it without
re-reading. A delivery landing in between made the session BUSY, and it was torn down anyway —
contrary to the method's own javadoc, which promised it never reaps a BUSY worker. The reap is now a
compare-and-release: it tears the session down only while the registry still holds the exact record
it checked.
On. Always on.
Why it exists. onDelivered flips exactly the two states the reaper accepts (READY, DONE) to
BUSY, and bumpTurn refreshes the activity clock — so the very act that should save a session from
the reaper is the one that raced it. The worker was killed, the turn never ran, and the release
reason said nothing about a race. It was never a silent loss: the release listener still fails a
blocked send with a reason, so the lead is told something went wrong, just not what.
One thing to know for maintenance. The compare is ConcurrentHashMap.remove(key, value), and
that is the linearization point — do not "simplify" it to a re-read of the state followed by a plain
remove, which shrinks the window without closing it and leaves the code looking correct. The public
release path stays unconditional on purpose: an explicit fleet_stop must not be refused because
the worker happens to have just become busy. A skipped reap logs one debug line naming the pane;
without it the race is unobservable by construction, and a reaper that quietly stops reaping is very
hard to diagnose.
A wedged post-turn phase releases itself, like the turn phase already did
What. After a delegated turn finishes, the injector runs a housekeeping phase — it sends the
adapter's /clear and waits for the pane to pick it up. Four latches gate delivery to a member:
awaitingCompletion, postTurnPending, awaitingPostTurnPickup and postTurnObserved. Only the
first had a way out of a long run of unknown pane states. The other two now get the same escape: a
sustained unknown streak drops them and the member becomes deliverable again.
On. Always on, but only reachable when the post-turn /clear is configured
(lifecycle.clearAfterTurn). It is not set in the live fleetd.yaml today, so this is a latent
fix here rather than one that was costing us turns.
Why it exists. A pane whose state cannot be read stays unknown forever, and a latch with no
timer waits forever with it. The member never returns to idle, so every later fleet_send to it
waits and then fails without ever reaching the pane. The turn phase was given an escape when this
happened to it; the housekeeping phase was left with none, which is the same gate closed in one
direction only.
One thing to know for maintenance. The new escape uses its own counter, unknownSincePostTurn,
and deliberately does not set turnFailed. The delegated turn already completed and its waiter
already resolved — what is outstanding is adapter housekeeping, so failing the turn would report a
false failure for work that succeeded. When you fix a latch here, check its siblings in the same
method: this fix had to cover two latches, not the one that was reported, or the wedge simply moved
one notch further along.
A worker's real reply after a lapsed fleet_ask completes its ticket
What. A worker that calls fleet_ask and gets no answer within the ~55s window resumes on its
own, finishes, and ends the turn with fleet_reply. That reply used to land in the session inbox
with nothing tying it to the delegation: fleet_poll{ticket} stayed PENDING, and fleet_stop
later forced the ticket FAILED with the reason "session released before it replied". The reply now
completes its own ticket.
On. Always on.
Why it exists. The lead was told the opposite of what happened. The report was never destroyed —
it reached the inbox — but the ticket said the worker never replied, and a lead that believes that
re-does the work. An unanswered ask is the ordinary case on this fleet, not an edge, which is why
the charter says never to brief a worker to "ask me". The sibling timeout in answer() already had
this recovery; ask()'s did not.
One thing to know for maintenance. Do not "fix" this by passing forgetTurn=false on the ask
timeout. It looks like the one-line version of the same fix and it wedges the member: the stamped
turnId keeps hasAsyncQuestion reporting the target BUSY, so every later fleet_send to it is
refused. The fix instead sets a separate Task.askTimedOut flag before the forgetting, and never
re-adds the task to asyncTasksByTurn. One real consequence: the ambiguity branch in reply() used
to be unreachable and is now reachable, because a lapsed ask frees its target for a fresh
delegation that can lapse in turn. Two open tasks on one target fall back to the inbox rather than
guess — completing the wrong ticket would hand the lead a plausible answer to work nobody did.
Shutdown refuses new spawns and sweeps up stragglers
What. fleet_spawn is now refused once the daemon's shutdown drain has started — a named error
over MCP, HTTP 503 over REST — and the drain re-reads the registry after its main pass to tear down
anything that raced in anyway.
On. Always on.
Why it exists. The drain worked from a one-shot registry snapshot, and the MCP server stayed up
for a long time after it started: on the live config lifecycle.drainTimeoutSeconds is 120, so a
drain waiting on a busy worker could hold fleet_spawn open for two minutes. A session accepted in
that window was invisible to the drain — its pane kept running, its worktree was never preserved,
and the in-memory registry died with the process, so nothing else could ever reclaim either. The
caller got a normal sessionId and no way to know.
One thing to know for maintenance. The guard and the sweep are a pair and neither is redundant: a
guard alone still loses to a caller already inside launcher.spawn(), and a sweep alone hands the
caller a session that is then destroyed. The sweep shares the drain's one deadline rather than
taking a second budget — a per-session grace could push the daemon past launchd's exit window and
get it SIGKILLed mid-teardown. The draining flag is never reset, which is correct only because
drainAll is reachable from the shutdown hook alone; if a fleet_drain tool is ever added for a
live daemon, that flag becomes a permanent spawn outage.
A failed git worktree add cleans up what it half-created
What. The git worktree add command now runs inside the same cleanup scope as the provisioning
steps that follow it. If the command is killed — the 30-second timeout, or an interrupt — after Git
has begun writing worktree state, that state is removed instead of leaking.
On. Always on.
Why it exists. The cleanup added for #274 covered every step after the add and assumed the add
itself was atomic on failure. It is, for an error Git reports; it is not when fleetd calls
destroyForcibly on it. SessionManager.acquireWithWorktree never receives a path in that case, so
its own if (path != null) cleanup never fires either, and nothing at any layer reclaims the
directory or the branch. This is a resource leak, not data loss — no worker was ever started, so the
half-made worktree holds nobody's work.
One thing to know for maintenance. Widening the cleanup scope creates a data-loss risk that the
guard exists to stop. cleanupAfterAddFailure deletes the branch with -D, so running it after an
ordinary "a branch named X already exists" refusal would delete the operator's existing branch. The
Files.exists(worktreePath) check prevents that: measured in a throwaway repo, Git creates no
directory on that refusal, so no directory means nothing was created and the branch is not ours to
touch. Removing that check makes addFailureBeforeCreatingAWorktreeIsQuiet fail with the branch
actually deleted. Do not remove it.
fixed placement honours the failover loop's unreachable set
What. The fixed placement policy now skips a profile that the current spawn attempt has
already found unreachable. It checks ctx.unreachable() in both places it picks a profile: the
configured default, and the fallback walk over the remaining candidates.
On. Only when placement: fixed is set in fleetd.yaml. The live fleet runs weighted, which
already honoured the set, so this changed nothing here — it was reachable only for an operator who
had switched to fixed.
Why it exists. CompositePeerLauncher retries a failed spawn by adding the dead profile to the
unreachable set and asking the policy again, and its own comment says why: "Update the context for
the next selection so the policy excludes this profile." fixed ignored the set, so it handed back
the same dead profile on every attempt. The loop then spent all of its attempts on one profile and
reported failure, while healthy profiles sat idle and were never tried. The failover existed but
could not move.
One thing to know for maintenance. fixed ignoring the load caps is deliberate and stays —
that is what "fixed" means. Reachability is not a cap, and the javadoc now separates the two so the
next reader does not undo this as a bug fix. Also note what does not work: breaking out of the
retry loop early when the policy returns an already-unreachable profile only makes it fail faster on
the same dead profile, because only the policy chooses what comes next. The failure message now
counts distinct candidates, so "tried 1 distinct candidate(s)" is the visible symptom of this
whole class of problem.
A caller whose identity cannot be resolved is refused, not treated as the primary
What. In loopback-trust mode, fleetd grants the primary role only to a caller whose operating
system process id was actually found. A caller whose peer-PID lookup failed is now refused as
ANONYMOUS. ConnectionIdentity.Caller carries a resolved() predicate that says which case it is.
On. Always on in loopback-trust, which is the mode you get when fleetd.yaml has no auth:
block — this fleet's live configuration. token mode was never affected.
Why it exists. LsofPeerPidLookup returns -1 for every failure, including the silent one
where lsof runs fine and simply reports no matching process. No pane matches -1, so the caller
arrived at the resolver with a null terminal — the same shape the real primary has, because no
pane owns the primary either. The two were indistinguishable, and both were granted SPAWN, STOP,
SEND and DRAIN. A worker whose lookup failed became the lead. PaneLocator's javadoc had already
named this exact escalation for a different trigger, and CB-161's ancestry walk closed that one — but
the walk needs a candidate pid to walk, and a failed lookup has none.
One thing to know for maintenance. The signal that separates the two cases was already in the
data and simply not read: the primary has a real pid and no pane; an unresolved caller has neither.
That is why resolved() lives on the record next to the sentinel rather than as a pid > 0 test
copied into the resolver — one rule in two copies is what let #305 drift. This change is
deliberately fail-closed: if lsof ever fails for the primary's own connection, the primary is
refused until its next call resolves. That costs availability and it is the right trade, because the
old behaviour spent it on a silent escalation instead. A new DEBUG line in LsofPeerPidLookup now
names the previously-silent no-match case, so a refusal that does happen can be diagnosed.
The check that authorises deleting a worktree is taken after the worker stops
What. When a member is released, fleetd re-reads whether its worktree has uncommitted work
immediately before the git worktree remove --force, with the pane already closed. If the tree is
dirty by then, the removal is skipped and the work is also snapshotted into refs/wip/*.
On. Always on. It costs one extra git status, and only on a release that was about to delete
something — a release that already decided to preserve, and every SHUTDOWN, pay nothing.
Why it exists. The old code read hasUncommitted once, while the worker was still running, and
used that one boolean after launcher.stop to authorise the force-delete. A worker that committed,
was released, and then wrote one more file during teardown lost that file. Worse, the same stale
boolean gated the snapshot, so both defences — CB-576's preserve and CB-578 stage C's refs/wip
copy — failed together. There was no third layer: pane gone, registry entry gone, directory
force-deleted.
One thing to know for maintenance. The snapshot deliberately did not move after the stop.
notifyReleased fires before launcher.stop on purpose (CB-516/CB-581), so a caller blocked in a
fleet_ask rendezvous fails fast instead of waiting on the herdr RPC; moving the snapshot would
force that notification to move too or to carry a snapshotRef that was never computed. The late
snapshot's ref therefore reaches the log and not the listener — a known, accepted gap. Also note
what is not proven: whether herdr's pane.close returning means the worker process is really
dead is decided inside herdr, whose source is not in this repo. The re-check narrows the race; it is
not proof the race is gone. And the fail-toward-preserving rule in dirtyImmediatelyBeforeRemoval's
catch is the whole point — flipping it to return false turns this guard into a cause of the data
loss it prevents. releasePreservesAWorktreeWhoseLateRecheckCannotBeRead exists because that
mutation once passed the entire suite.
A reply that arrives while a target is being released is requeued, not stranded
What. AmqpReplyInbox.release now leaves a RELEASED tombstone in its held map instead of
removing the key. A delivery that lands during or after the release sees the tombstone and is
nacked with requeue, so a later owner or a connection drop can still recover it.
On. Always on, wherever the AMQP inbox is used.
Why it exists. basicCancel stops new dispatches but does not flush one already handed to the
client's consumer work pool. That delivery reached deliverCallback, found the key gone, and
created a brand-new map under it — one release had already walked past and would never read
again. The message then sat delivered-but-unacked until the whole inbox closed: never requeued,
never redelivered, and nothing peeks a released target again. A worker's real report disappeared
with no log line naming it. #298 fixed the case where the delivery was already held; this is the
case where it arrives during the release.
One thing to know for maintenance. The tombstone closes the window rather than narrowing it,
and the reason is specific: ConcurrentHashMap serializes compute and computeIfAbsent for the
same key against each other, so the swap and a racing insert cannot interleave. That is why the
tombstone lives in the same map rather than in a separate "released" set — a second structure would
have to be kept in sync, which is the shape that keeps producing defects here. RELEASED is one
shared mutable map instance, so every read site must compare it by reference before touching it:
peek, ack and deliverCallback all do, and own clears a stale tombstone with the two-argument
remove so it can never delete a real map. Both halves of the fix are separately load-bearing —
reverting either one alone fails AmqpReplyInboxReleaseRaceTest.
The reload classifier proves its own coverage instead of claiming it
What. A test enumerates every record component of a worker profile by reflection, changes each one in turn, and asserts the reload classifier actually notices. A component that is neither compared nor on a small, pinned exclusion list fails the build by name. The test prints its own denominator on every run:
ConfigRef.sameLaunchSettings coverage — 26 Profile components total, 23 compared, 3 excluded ([weight, maxLoad, credentialId])
On. Always on — it is a unit test, so it runs in every build.
Why it exists. sameLaunchSettings decides whether a changed profile key is reported to the
operator as deferred (needs a restart). Its javadoc said it "compares every component the launcher
reads at spawn". It did not: ideProjectDir, ideOpenCommand and autoCompactWindow were all
missing, and worktreeGroup was missing from the sibling top-level check. A reload of any of them
returned applied = true with an empty deferred list — a bare config reloaded — while the daemon
kept the old value. The surrounding code already calls that "the worst outcome a reload can produce,
because the operator has no reason to doubt it". The list was a hand-maintained second copy of "what
the launcher reads at spawn", and it drifted. On its first run the new test found a fifth gap nobody
had reported: the profile component itself.
One thing to know for maintenance. The exclusion list is this mechanism's own escape hatch, and
it is pinned for a reason. Measured while verifying the fix: moving autoCompactWindow and
ideOpenCommand out of the comparison and into the exclusion set left the whole suite green —
the loop simply skipped them and the denominator still balanced. That is the cheapest way to silence
a failing coverage test, and it silently restores the original bug. The test now asserts the
exclusion set equals exactly {weight, maxLoad, credentialId}, so growing it takes a visible,
deliberate edit. A component belongs there only if it is read live off the config supplier, never
because adding it makes the build pass. The general rule: when a ticket asks for a checker rather
than a fix, mutate the checker too, and ask what the cheapest way to pass it without doing the work
would be.
An answered fleet_ask no longer fails the lead's own call
What it does. When the lead answers a worker's fleet_ask at the same moment that ask's ~55
second window lapses, fleet_send{turnId, content} used to be able to throw
NullPointerException back at the lead — even though the answer had already been delivered. It
does not any more.
On. Always on. There is no knob; it is a correctness fix (fleetd #324, merged 02e6aef).
Why it exists. answer() does its work holding the target's sessionLocks entry. ask()'s own
timeout path holds no lock at all, and it nulls Task.turnId. finishAsyncTask read that field
twice — once to check it was not null, once as the key for asyncTasksByTurn.remove. volatile
makes each read fresh, but it does not make a pair of reads atomic. When the unlocked null-out
landed between them, the second read saw null and ConcurrentHashMap.remove(null, task) threw. The
lead was told its answer failed. It had not: task.future.complete(result) ran on the line above.
A lead that reacts by re-sending is acting on a false failure.
One thing to know for maintenance. The fix reads the field once into a local. That closes the
crash and not the asymmetry behind it. One side of this invariant is still locked and the other
is not, and two consequences of that are open in fleetd #329: a worker's real reply can leave its
async ticket PENDING for good, and reply() still contains the same double read. There is also a
production test seam here — a volatile Runnable hook plus two package-private setters, null and
unused outside tests. It exists because no public path reaches this interleaving without a real race,
so a deterministic test needs somewhere to stand.
A reload now says a restart is needed for primary: and configReload:
What it does. ConfigRef reports a changed primary: or configReload: block as deferred —
accepted into the new snapshot, but not in effect until the daemon restarts. Before this, changing
either one reported a bare config reloaded and the daemon quietly kept the old value.
On. Always on (fleetd #326, merged 823976c). Visible in the reload summary and in
ConfigRef.Outcome.deferred().
Why it exists. Both keys are read once, at startup. primary: feeds PrimaryRegistry's pinned
terminal and sizes ReplyPushLoop's reminder cap and backoff; neither is rebuilt. configReload:
decides whether a ConfigWatcher is built at all and with what interval — so the component that
would apply a later change is itself built once. Turning reload off through a reload reported
success and changed nothing. This is the same drift fleetd #323 fixed one level down, in the
per-profile list.
One thing to know for maintenance. Write down the denominator, because "not mentioned in
ConfigRef" looks identical for a key that is correctly hot and for a key nobody triaged.
FleetConfig has 22 top-level components; four are named nowhere in that file.
memberCredentials and memberLoginShell are hot and correctly absent — both are read live off
config.get() at spawn. health and coordinator are undecided: each is read both off the
startup snapshot and live, at different sites, so no single class fits either. A reload touching
them still under-claims. That count and both verdicts are now in ConfigRef's class doc so the next
person does not measure it again. Twice now — worktreeGroup, then primary/configReload — the
untriaged kind hid among the correct kind.
A reload can now say "half of this applied"
What it does. ConfigRef has a fourth reload class, split, for a key that is read both
off the startup snapshot and live off the config supplier, at different sites. health: and
coordinator: are both like that. A reload that changes either now names the key and says which
half is already live and which half waits for a restart. Outcome carries a new split list
beside deferred.
On. Always on (fleetd #330, merged 7b918c5). Visible in the reload summary:
config reloaded; partially live — coordinator: the LeadMailbox connection (uri, uriEnv, selfId,
prefetch) is opened once and needs a restart; the broker URI env-var name kept out of a member's
environment is read live on every spawn and already applied
Why it exists. Before this, changing either key reported a bare config reloaded. That
under-claims: it tells the operator a change applied when half of it did not. The three existing
classes could not express the truth, and forcing one of them would be wrong in one direction or the
other. The direction matters. Over-claiming a restart costs an unnecessary restart, which the
operator can see and recover from. Under-claiming is what this file's own doc calls "the worst
thing a reload can do to an operator debugging one". split is the only option that is simply
true. Making the frozen half live was rejected for coordinator: selfId names this daemon's own
AMQP inbox queue, so changing it live is a distributed-identity problem, not a reconnect.
One thing to know for maintenance. Membership in SPLIT_KEYS is not proof that any
reporting code exists for that key. ConfigRefTopLevelCoverageTest reads a name in the set as
"triaged" and stops there — it cannot see whether changedSplitKeys has a branch for it. Measured:
dropping the coordinator branch while leaving the name in the set left both the coverage test and
the assert silent; only three hand-written behavioural tests caught it. The same one-way shape is
older than this change — assert COLD_KEYS.containsAll(changed) catches "reported but not listed"
and never the reverse. fleet: is already an instance of the gap: it is split too and it sits in
the checker's hot-exclusion hatch, so a changed fleet.leaders still reports nothing. Both are open
in fleetd #333.
Every top-level config key must be triaged before it ships
What it does. ConfigRefTopLevelCoverageTest enumerates FleetConfig's record components by
reflection and requires each one to sit in exactly one of four buckets: COLD_KEYS, the deferred
set, SPLIT_KEYS, or a pinned hot-exclusion set. A new key in none of them fails the build by name.
It prints its own denominator every run:
FleetConfig top-level coverage — 22 components total: 5 cold [...], 11 deferred [...],
2 split [health, coordinator], 4 hot-excluded [...]
On. Always on — it is a unit test (fleetd #330).
Why it exists. Three times the same key-was-forgotten bug shipped: worktreeGroup (#323),
then primary and configReload (#326), then fleet (#333). Each time a reload reported success
for a change the daemon never picked up. "Not mentioned in ConfigRef" looks identical for a key
that is correctly hot and a key nobody triaged, so the forgotten kind kept hiding among the correct
kind. This is the top-level twin of the per-profile checker #323 added.
One thing to know for maintenance. Read the test's own javadoc before trusting a green run. It
says plainly what it cannot do: it proves the record's shape is triaged, and it cannot prove a
citation is true. "Compared in changedDeferredKeys" and "read live off config.get()" are facts
about other files that a reflection test over one record cannot inspect. That disclosure is not
modesty — it is what made the two gaps in #333 findable in minutes. Hold any checker in this repo to
the same standard: say what it does not cover, next to what it does.
An exception thrown after a ticket resolves now reaches the log
What it does. sendAsync's executor used to end with a bare
catch (Throwable t) { task.future.completeExceptionally(t); }. That future is already completed by
then, because finishAsyncTask completes it on its first line. completeExceptionally on a
completed future returns false and does nothing. The exception simply vanished. The catch now
checks that return value and logs at error with the ticket and the target when it is false.
On. Always on (fleetd #329). Nothing to configure — look for
async send task-N -> <target> threw after its ticket was already resolved in fleetd.out.
Why it exists. This was not one bug, it was a blind spot over the whole async region. Measured
during #324: a deliberately broken finishAsyncTask threw 19 real NullPointerExceptions on the
ordinary path while the suite reported 1340 tests green and not one log line. Any defect that
throws after the future completes produced a passing build and a silent daemon. The daemon could not
tell an operator, and no test could tell a developer.
One thing to know for maintenance. When a mutation you expect to fail passes, that can mean the
failure is invisible, not that the code is unpinned. Add one temporary log.error and re-run
before you conclude anything. That is the only reason this was found.
A worker's real reply completes its async ticket more often
What it does. answer() used to look the task up a second time, by turnId, when completing the
async ticket. That second lookup raced ask()'s own unlocked timeout cleanup, so a ticket could stay
PENDING forever after the worker had actually replied. answer() now completes from the Task it
already holds from its first lookup.
On. Always on (fleetd #329).
Why it exists. The failure was silent and it lied to the lead. fleet_poll{ticket} reported
PENDING for good, and teardown later resolved the ticket as WORKER_FAILED — "session released
before it replied" — long after the worker had replied. A lead reading that is told something false
about its own worker.
One thing to know for maintenance. This narrows the window; it does not close it. ask()
drops the asyncTasksByTurn entry in its catch (clearAsyncQuestion(turnId, true)) but closes the
ask later, in its finally. Between those two the ask is still answerable and the entry is already
gone, so answer()'s own first lookup returns null and the ticket is stranded one step earlier in
the same race. Measured on 2026-09-04: a probe firing only that first half printed
answer=REPLIED phase=PENDING reply=null. Open as fleetd #334, and the comment in answer() says so.
Do not read that task != null guard as complete.
A reload now reports a changed fleet.leaders as needing a restart
What it does. fleet: used to be filed as a fully hot key, so a reload that changed it said
nothing at all. It is really a split key. fleet.leaders is read once at startup, in two
places — Fleetd builds the tab-label→lead map from the startup snapshot, and LeadLauncher holds
a frozen FleetConfig rather than a supplier. Neither rebuilds on reload. The rest of fleet:
(role pools, charters, tabLabel) really is live. changedSplitKeys now compares fleet.leaders
specifically and names both halves.
On. Always on (fleetd #333).
Why it exists. A lead is found by its tab label, and a lead whose tab no longer matches is
demoted to worker and refuses every orchestration call. Before this, an operator could edit that
label, read config reloaded, and be left with a broken lead and no message saying why. The
comparison is deliberately on fleet.leaders and not on the whole fleet record: comparing the
whole record would claim "needs a restart" for a tabLabel-only change that is fully live. Over-
claiming a restart is cheap to recover from; it is still a wrong report, and this class exists to
stop wrong reports in both directions.
One thing to know for maintenance. This is the third time the same bug shipped — worktreeGroup
(#323), then primary and configReload (#326), now fleet. It was in the coverage checker's own
hot-exclusion hatch, blessed by the checker, on the checker's first commit. A name sitting in an
escape hatch is the place to look first, not last.
Being listed as cold or split now proves a comparison exists
What it does. ConfigRefTopLevelReportingCoverageTest builds FleetConfig pairs by reflection
that differ in exactly one top-level component, then calls the real changedColdKeys and
changedSplitKeys and requires that key to come back. A name added to COLD_KEYS or SPLIT_KEYS
with no if behind it now fails the build, by name.
On. Always on — a unit test (fleetd #333).
Why it exists. Every checker in this area had closed only one direction. The older
assert COLD_KEYS.containsAll(changed) catches "reported but not listed" and never the reverse, and
ConfigRefTopLevelCoverageTest reads a name in a set as "triaged" and stops there. Measured:
dropping the coordinator branch while leaving the name in SPLIT_KEYS left both of those silent.
So the cheapest way to pass the shape checker was to add one string and write no code — which is
exactly how fleet got through.
One thing to know for maintenance. It covers COLD_KEYS and SPLIT_KEYS only. The deferred set
— 11 of the 22 keys, the largest bucket — is not covered, and the worker wrote that gap into the
test's own javadoc instead of quietly leaving it. Measured on 2026-09-04: deleting guard's
comparison from changedDeferredKeys while leaving guard in the deferred set left all 1355 tests
green. Open as fleetd #337. Two tickets running, an honest caveat in a test's javadoc has been the
fastest route to the next bug — hold every checker here to that standard.
A send that times out while still queued now cancels its message
What it does. Injector.enqueue returns an identity Delivery handle. When a send's deadline
passes and its message was never delivered, MessageService cancels that exact Pending instead of
leaving it in the queue. queuedDeliveries still records the health fact — that half is unchanged.
On. Always on (fleetd #338).
Why it exists. Before this, TIMED_OUT_QUEUED told the sender the message did not go, and then
the message went anyway. The injector picked it up on the member's next injectable status and typed
it in — minutes or hours later, after the lead had moved on and usually re-sent the work elsewhere.
Nobody was told. This is not a theoretical path: it is a behaviour hit while operating the fleet,
and a test had been asserting it as correct.
One thing to know for maintenance. Cancelling is the strict direction and it costs something: a member that was about to go idle loses a message it could have taken, and the lead must send again. That is the right trade because it is loud and recoverable, but it is a trade. If a queued message starts disappearing more often than expected, this is why.
The race is resolved deliberately. Cancellation and Injector.onStatus share the target
monitor. If pickup wins, the text has already landed, cancel returns DELIVERED, and the result
is TIMED_OUT_WORKING — not TIMED_OUT_QUEUED. Claiming "queued" while the text landed would be
the original bug with a smaller window. Note that MessageService reading that return value is
not currently pinned by a test: InjectorTest covers cancel itself, and mutating the caller's
use of it left all 1357 tests green. Open as fleetd #345.
A member's prose about an error no longer records a credential outage
What it does. CompletionResolver used to notify backendErrorSink on any line of a member's
pane matching the backend-error pattern. The built-in fallback is (?i)\bAPI Error\s*: —
case-insensitive and unanchored — so a member that ended a turn without fleet_reply while merely
writing about an error matched it. The sink call now needs the pattern at the start of its
matched line, ignoring leading terminal chrome. Failing the send is unchanged: any match still fails
it and still carries the whole pane tail.
On. Always on (fleetd #339). The too-fast crash path still notifies on any match, because there the crash signature is corroboration.
Why it exists. The sink is not cosmetic. Two backend errors on one credential within 60 seconds
put it into cooling-off, and a fleet_spawn naming a cooling profile is refused before it reaches
the backend. Profiles share credentials here, so a member's own prose could block spawns on a
profile that never had a problem. The code's comment already admitted the false match and argued
that failing the send is still right — a sound argument that covers resolveFailure and says
nothing about the sink call sitting in the same block.
One thing to know for maintenance. The first version of this check used a bare lookingAt(),
and that rejected a genuine error line rendered as │ 503 Service Unavailable: ... — the send
failed and the outage went unrecorded. That is the worse direction: an unrecorded outage leaves the
fleet spawning into a dead credential. It was caught on merge by a probe, not by the suite, because
every existing test put the error line with no chrome in front of it. The check now skips a leading
run of non-letter, non-digit characters, and aRealErrorBehindTerminalChromeStillNotifiesTheSink
pins it. If you touch this check, test it against a chromed line, not only a clean one.
It is still a heuristic: prose that begins with API Error: will still notify the sink. The same
free-text shape drives ExhaustionSink, where a quarantine runs 1800s against this cooldown's fixed
60s — open as fleetd #348, and unproven.
Every distinct unprotected credential name gets its own warning
What it does. The memberCredentials gap WARN — the one naming environment variables that every
member pane inherits unblocked — was guarded by a single AtomicBoolean shared by two branches
that report different variable names. It is now a Set<String> of names already warned about, so
the guard is per name rather than per launcher.
On. Always on (fleetd #341).
Why it exists. memberCredentials is re-read on every spawn, so an operator can change the
policy with a reload and no restart. Spawn 1 under deny-by-default warned about one variable and
tripped the flag; after a reload, spawn 2's genuinely unprotected different variable was never
reported. The operator fixes the one name they were shown and reasonably believes the gap is closed.
The whole point of naming variables in these lines is so they can be acted on.
The same class already had the right reasoning written down: #192 split this flag from the allow-list INFO guard precisely so a harmless report could not suppress a real one. That reasoning was applied to INFO-versus-WARN and never to WARN-versus-WARN.
One thing to know for maintenance. The set has two duties and only one was pinned at first.
Measured on merge: replacing the .filter(...::add) with one that logs every name on every spawn
left all 1358 tests green — the noise control was correct and nothing held it there. Two tests now
cover both duties, including the reverse policy order (allow-list first, then deny-by-default),
because a guard fixed in one direction is not automatically fixed in the other.
The deferred key set now proves its own reporting coverage too
What it does. ConfigRefTopLevelReportingCoverageTest covered COLD_KEYS and SPLIT_KEYS only.
It now covers the deferred bucket as well. The deferred set moved out of the test and into
ConfigRef.DEFERRED_KEYS (package-private, 11 keys), and changedDeferredKeys became
package-private like changedColdKeys and changedSplitKeys, so the test calls the real method
instead of a copy of the list. A name in any of the three sets with no comparison behind it now
fails the build, by name.
On. Always on — a unit test (fleetd #337).
Why it exists. The deferred bucket is the largest of the four: 11 of the 22 top-level keys. It
was also the one bucket where a name could be added with no code behind it and every checker stayed
green. Measured before the fix: dropping guard's comparison out of changedDeferredKeys while
"guard" stayed in the set left all 1355 tests green. An operator who changes such a key gets no
"restart needed" line, so the daemon keeps running config that matches no file on disk and says
nothing.
One thing to know for maintenance. The real number of uncovered keys was 6 of 11 — guard,
worktreeRoot, spawnReadyTimeoutMs, spawnReadyPollMs, quarantineCooldownSeconds and
leadHeartbeat. My own ticket listed a different six. The worker re-derived the list by mutating
each key one at a time, as the brief asked, and contradicted me on two entries: lifecycle was
already covered, and worktreeRoot was uncovered and missing from my list. Both corrections were
checked again on merge with a third mutation. A list in a ticket is a starting point, not a
measurement — re-derive it, and say so when it disagrees.
Teardown now closes the tab a member really sits in
What it does. HerdrPeerLauncher.stop used to look for a tab to clean up only when one of its
own configured profiles used tab placement (usesTabPlacement()). It now resolves the pane's real
tab every time, and the single-occupant check decides whether that tab is closed. That check is
unchanged: a tab holding other panes is never closed.
On. Always on (fleetd #342).
Why it exists. CompositePeerLauncher routes a stop through the adapter recorded in
spawnedBy, and that map is in memory only — a daemon restart empties it. On a miss in a
one-daemon fleet, the stop goes to delegates.getFirst(). When two adapters share one herdr daemon
and disagree on tab versus pane placement, teardown could run through an adapter whose config
says nothing true about how that pane was placed, and the whole tab-cleanup block was skipped. The
member still stopped, but its now-empty tab stayed. Nothing reaps an orphaned tab, so the operator's
herdr session collected one more of them on every affected teardown, silently.
Why not fix the routing instead. Probing for the pane's real owner looks like the obvious fix
and does nothing here. probeOwner groups candidates in an IdentityHashMap keyed by
HerdrClient, so two delegates sharing one daemon collapse to whichever was inserted first — that
is delegates.getFirst() again, the same answer the shortcut already gave. The routing cannot tell
these two adapters apart at all.
One thing to know for maintenance. The single-occupant check is now the only thing protecting
a shared tab, so treat it as load-bearing. Measured on merge: weakening it from tabPaneCount() == 1
to >= 1 fails two tests, including the pane-placement one. One case is uncovered on purpose — a
pane-placement member that is the sole occupant of its tab will now have that tab closed.
spawnAsPane splits an existing tab, so the count is normally at least 2, and the tab is empty
after the member's pane goes anyway.
The same in-memory spawnedBy breaks clearContext a different way: it plainly no-ops on a cache
miss. Open as fleetd #352, and the consequence is not measured yet.
A failed cleanup no longer strands the tickets behind it
What it does. MessageService.abandon walks every task it must fail and completes each one. The
per-task cleanup after that completion is now inside a try/catch, and each task's own
future.complete(outcome) runs before the guard. A cleanup that throws costs that one task its
bookkeeping and nothing more; the loop still reaches every task behind it. The second half of the
same fix wraps pushLoop.onTicketTerminal inside sendAsync's whenComplete action.
On. Always on (fleetd #335).
Why it exists. abandon runs on teardown — a release, or the health monitor's GONE sweep — and
nothing comes along later to finish what it misses. One call in that loop reaches a broker:
inbox.publish puts a recovered reply back, and AmqpReplyInbox.publish throws
IllegalStateException on an unroutable publish, on a confirm timeout, and on an interrupt. An
uncaught throw there aborted the loop, so every task after it stayed PENDING forever and its lead
waited on a ticket that would never resolve. The whenComplete half is the same failure with a
different cause: the daemon's shutdown hook closes MessageService before ReplyPushLoop, and
messages.close() does not cancel a send already in flight, so a ticket completing in that window
made the push loop's scheduler throw RejectedExecutionException into a discarded stage — no log,
no metric, and the push loop never learned the ticket was terminal.
One thing to know for maintenance. A publish failure used to propagate out of abandon to its
caller. It is now logged at error with the ticket, target and turnId, and swallowed. That is
deliberate and matches what #293 already does for teardown in HerdrPeerLauncher: past the point
where the real work is done, a cleanup failure must not mask the steps behind it.
A third site was reported and is not a defect: the finally blocks in send() and answer()
call asyncTasksByWaiter.remove and Rendezvous.close, which is waiters.remove(session, waiter).
Neither can throw, so no guard was added. A guard that can never fire is worse than none — it reads
as evidence that somebody checked.
Measured on merge: keeping the catch but adding a break to it leaves the suite failing at
aPerTaskCleanupFailureDoesNotStrandTheRemainingMatchingTasks with expected: <FAILED> but was: <PENDING>. So the test pins the property that matters — the tasks behind the throwing one still
finish — and not merely that no exception escapes.
A member's prose about a usage limit no longer quarantines a credential
What it does. CompletionResolver classifies a turn that ended without fleet_reply and whose
pane text matches the profile's exhaustedPattern as BACKEND_EXHAUSTED, and tells
ExhaustionSink. The sink call now needs the match to sit before the first sentence ending on its
line. Failing the send is unchanged: any match still fails it and still carries the whole pane tail.
On. Always on, wherever a profile sets exhaustedPattern (fleetd #348).
Why it exists. This is #339's shape with a much heavier penalty. A backend-error cooldown is a fixed 60 seconds; an exhaustion quarantine defaults to 1800. Measured on this ticket rather than assumed: a normal member report — "I reviewed capacity handling. The usage limit has been reached means no more work can start." — notified the sink and would have taken a credential out for half an hour.
Why the check is looser than the backend-error one. An exhaustedPattern is written per profile
and may name only the decisive words, without the provider's leading "The". A start-of-line check
would then reject the genuine refusal, which is the worse direction — an unrecorded exhaustion
leaves the fleet spawning into a credential that has no capacity. Measured on merge: swapping in the
start-of-line check fails four tests, three of them pre-existing. The rule as written accepts a
superset of what a start-of-line check accepts, so it cannot add a false negative.
One thing to know for maintenance. It is a heuristic and the javadoc says exactly where it stops. Prose whose first sentence carries the pattern still notifies the sink; a genuine refusal behind an earlier full stop — a hostname, a version number — still does not.
The first version copied the backend-error check's leading-chrome loop. Measured: deleting that loop
left all 1369 tests green, and it must, because the scan only looks for ., ! and ? and no
chrome character is one of those. It is gone. A step that cannot change the result is worse than
no step — the next reader takes it as evidence that chrome was handled.
The last window where a fleet_ask timeout stranded its ticket is closed
What it does. When a worker's fleet_ask times out, ask() now closes the ask turn
(rendezvous.closeAsk) before it forgets that task's turnId mapping. The order used to be the
other way round, with the close happening later in the shared finally. The whole teardown is also
gated on ticket.fresh() now, matching the finally block and the NO_WAITER branch, which were
already gated that way.
On. Always on (fleetd #334, closing what #329 only narrowed).
Why it exists. Between the forget and the close, the ask was still answerable while its task
mapping was already gone. A lead calling fleet_send{turnId} in that window got a REPLIED answer,
while the ticket stayed PENDING with no reply — forever, because nothing revisits it. Measured
with a probe before the fix: answer=REPLIED phase=PENDING reply=null. With the new order, an
answer() either sees the ask open — and then the mapping is still there — or sees it closed and
returns STALE_TURN. There is no state in between, because both steps run on one thread with
nothing yielding.
The ticket.fresh() gate fixes a second door into the same failure. A coalesced duplicate ask
passes its own timeoutMillis, which says nothing about whether the shared ask is done, so a
duplicate timing out first could lapse an ask the fresh owner was still holding.
One thing to know for maintenance. The two halves are pinned by two different tests, and the
second one only exists because the first did not cover it. Measured on merge: removing the
ticket.fresh() gate while keeping the new order left all 1371 tests green. The gate shipped with
the reorder and nothing held it there. aCoalescedDuplicateAskTimingOutLeavesTheFreshOwnersAskOpen
now does, and aLateAnswerDuringAskTimeoutTeardownStillCompletesTheAsyncTicket covers the ordering.
MessageService now carries six test-only hooks, one per race of this kind. Each exists because its
window is unreachable through the public API — which is also why each bug was invisible — but six is
enough that the next one needs a harder look than "the file already does this".
fleet_list reports lead coordination state instead of guessing at it
What it does. A lead's fleet_list now answers three questions it could not answer before: is
my own coordination mailbox there and is anyone reading it, what is waiting in it, and what is the
state of each peer daemon I know about. Each mailbox row carries a status of exists, absent
or unknown. pending and consumers appear only when status is exists.
The reading is a passive AMQP queue declare, not a presence protocol. consumers: 1 means a daemon
is attached and consuming; pending: N is the backlog.
fleet_send{coordId} also stopped saying "delivered". It now says the message was published and
durably confirmed by the broker, which is what a publisher confirm actually proves. When the target
mailbox exists but has zero consumers, the (still successful) result adds a warning that nobody is
reading it right now.
On. Add peers: to the coordinator: block — the operator declares who exists, because the
daemon never guesses:
coordinator:
uriEnv: LEAD_COORD_URI
selfId: mac
peers: [fleet01]
Omit it and you still get your own mailbox row. An undeclared peer can still reach you and be
reached by fleet_send; it simply does not get a row. Like the rest of the coordinator: block,
peers is read once at boot, so a change needs a daemon restart.
Why it exists. The first real cross-host link (Mac ↔ fleet01, 2026-09-05) worked, and using it
showed the MCP surface only covered sending. A lead could not learn that a peer existed, could not
read its own inbox, and was told "delivered" for a message the peer never saw — that one sat
undelivered through three daemon restarts because of the duplicate-lead-tab bug (#359). Finding
that out needed an ssh to the other host and lavinmqctl list_queues. None of it was reachable
through MCP, and a lead on a host with no broker shell was blind to its own inbox. fleetd #361.
One thing to know for maintenance. The three-state status is the whole point, and it is easy
to collapse back into a boolean. The first cut of this feature did exactly that: one absent()
value stood for both "the broker said there is no such queue" and "I could not check", so a
self-probe timeout rendered as pending: 0, consumers: 0 — indistinguishable from a mailbox that
is genuinely empty and genuinely unread. That is the same overstatement the ticket exists to fix,
one level down.
LeadMailbox.isMissingQueue is the discriminator that keeps them apart, and it is narrow on
purpose: only an IOException whose cause is a ShutdownSignalException carrying an
AMQP.Channel.Close with reply code 404 counts as a confirmed absence. Measured on merge: making
it return true unconditionally restored the original defect and left 1389 tests green. It now has
five tests of its own, and the same mutation gives 4 failures.
Two other things this feature depends on, both easy to undo by accident. inspect runs its passive
declare on a throwaway channel, because in AMQP 0-9-1 a passive declare of a missing queue
closes the channel it ran on — reusing the publish channel would let one miss break every later
publish on that instance. And FleetMcp.probe cancels a timed-out probe rather than abandoning it;
without that, a hung (not down) broker would orphan one channel per fleet_list call until the
connection's channel-max ran out, breaking publish by a different route.
Bridge skills are seeded into every provisioned worktree
What it does. fleetd copies a directory of skill folders into each worker worktree it creates,
at <worktree>/.claude/skills/. A skill folder whose name the target repo already ships is never
touched — the repo's own copy wins, byte for byte.
On. Point memberSkills: at the directory:
memberSkills: /Users/dai.ha/LTMS/claude-bridge/.claude/skills
Leave it out and nothing is seeded, which is the old behaviour. This is a deferred key — the
GitWorktrees that reads it is built once at startup, so a change needs a daemon restart.
Why it exists. Every brief starts with Load the <name> skill., and outside this repo that
line was silently a no-op. A member spawned against any other repo — kb on fleet01, for example —
had no implementer, reviewer or hunter to load, and nothing said so. The skills could not
travel in the plugin either: ClaudeCodeLauncher exports CLAUDE_CONFIG_DIR, so a member never
reads the operator's plugin store. The worktree is the only channel that reaches a member.
fleetd #362 item 3.
Opencode members need a second step, and they now get it (fleetd #393). Copying the folders is
the whole feature for a Claude Code member, because Claude Code reads .claude/skills/ natively.
Opencode never reads that directory. So for the first weeks this key existed, an opencode member
was seeded correctly and read nothing: the copy succeeded, the files were right, and no test failed
because there was nothing to fail. The feature worked at the only layer it implemented.
An opencode member's only channel for static guidance text is the instructions[] array in the
config OpenCodeLauncher generates for it. Each seeded skill's SKILL.md is now added there, by
absolute path. A skill folder with no SKILL.md is never delivered, and the log names the folder
so a typo is visible instead of silent.
The gotcha that matters here is delivery versus activation. Opencode has no equivalent of
Claude Code's Skill tool. The text arrives as part of the system prompt from spawn and stays there;
a member cannot load one skill by name when it needs it. So Load the implementer skill. means
something different on the two backends: on Claude Code it is an instruction the member acts on, on
opencode the content is simply already present. Do not read "skills work on opencode now" as more
than that. This is opencode's design, not a fleetd limit, and it is why the delivery gap could be
closed here and the activation gap could not.
A second gotcha, for anyone adding a third kind of guidance file. Three writers append to
instructions[]: the role charter, the seeded skills, and the IDE rules. All three now use
Jackson's withArray (get-or-create). One of them used putArray (create-or-replace), which
was safe only because it happened to run first against an empty array — an ordering rule nothing
wrote down and nothing tested. Measured before the fix: switching the skills writer to putArray
left the whole suite green while silently deleting the charter entry, so an opencode member launched
with no role contract at all. If you add a writer, use withArray, and assert the array's
contents — a size assertion passes when putArray swaps two entries for two different ones.
One thing to know for maintenance. core.excludesFile is single-valued. The seeded paths
are hidden from git status by pointing that key at a fleetd-written file, scoped --worktree —
and a worktree-scoped value replaces the operator's global one rather than adding to it. The
first cut of this feature did exactly that, and the consequence was severe: this repo's own
.gitignore does not ignore target/, only an operator's global excludesFile does, so every
worker that ran mvn clean install made target/ untracked. GitWorktrees.hasUncommitted counts
untracked files on purpose (CB-576), so SessionManager would have preserved every worktree that
built, forever, with no error to notice.
So the file is composed, not replaced: whatever core.excludesFile resolved to beforehand is
copied in ahead of the seeded patterns, including git's own default ($XDG_CONFIG_HOME/git/ignore,
else $HOME/.config/git/ignore) when the key was unset. Measured on merge: removing the
composition fails 2 tests, and removing just the default-file fallback fails 1.
Two smaller things. The composed content is a snapshot taken at seed time, so an operator
editing their own excludesFile later does not change an already-seeded worktree. And the exclude
file itself lives under the worktree's private git dir (git rev-parse --absolute-git-dir), not in
the working tree, so it cannot be committed and git worktree remove --force deletes it along with
everything else.
.git/info/exclude was rejected as the mechanism, and this is worth knowing before someone tries
it again: from a linked worktree it resolves to the common git dir, so it would have hidden the
seeded paths in the primary checkout and every sibling worktree too.
Dead lead tabs are cleaned up, and a live one is never closed
What it does. On startup, LeadLauncher.ensureLeads() now looks for tabs that carry a lead's
configured label but have no running agent. It does not close them straight away. It renames the
tab, appending [fleetd:pending-close], and leaves it open. Only a later reconcile that still
finds the same tab dead actually closes it. LeadTabScanner does the matching cross-check on the
read side: it joins a labelled tab to a terminal only when agent.list says that terminal is
live, and it grants one grace scan to a terminal it already knew was live.
The knob that turns it on. None — it is always on for any lead slot declared under
fleet.leaders.*. The marker suffix is a constant in PendingCloseMarker, and every place that
matches a tab label strips it first, so a flagged tab is still recognised as that lead's tab.
Why it exists. Two separate defects met here. LeadTabScanner joined labelled tabs straight
to terminals with no liveness check at all, and its javadoc excused that ("a stale name costs
nothing here"). It cost plenty: LeadCoordLoop reads that map to choose which pane a peer lead's
message is delivered into, so a dead tab was a valid candidate. LeadLauncher had the opposite
problem — no cleanup path whatsoever, so every reconcile that found 0 live leads created another
tab and left the old one behind. Restart the daemon a few times and the tab bar fills up.
One thing to know for maintenance. The two-reading rule is not caution for its own sake. The
evidence that opened this ticket was fleet01's own daemon log: agent.list reported 0 live
while ps showed one real claude process. A first cut of this fix closed tabs on that single
reading, which would have closed the operator's live lead rather than tidying a spare tab. So the
accepted cost is stated plainly: ensureLeads() runs at startup only, so the second reading
arrives at the next restart, and a tab whose agent dies mid-session stays flagged and open until
then. That is deliberate. The bug is about repeated restarts, and one leftover tab is much cheaper
than closing a live session on evidence that has already been seen to lie.
Measured on merge: 1412 tests green; making PendingCloseMarker.strip() the identity function
fails 4 tests. fleetd #359.
The shipped systemd units no longer disable the daemon they start
What it does. deploy/fleetd.service, deploy/herdr.service and deploy/herdr-inner.sh are
the units and helper script that actually run fleet01. Copy them, edit the paths in the headers,
systemctl --user enable --now both. fleetd starts from a login shell, neither unit takes a mount
namespace, and neither gets a private /tmp.
The knob that turns it on. Nothing to switch on — these are the deployment artefacts. What matters is what must stay off, and the files say so in their own comments.
Why it exists. The old deploy/fleetd.service carried ProtectSystem=strict,
ProtectHome=read-write, ProtectKernelTunables=true, ProtectControlGroups=true,
PrivateTmp=true, and an ExecStart that ran java directly. Every one of those looks like good
hardening and each breaks the daemon silently:
- Each of the four
Protect*directives gives the unit its own mount namespace. fleetd resolves a caller's role by runninglsofto find the loopback peer PID (mcp/LsofPeerPidLookup). Inside such a namespacelsofreturns nothing, every caller falls back to ANONYMOUS, and the primary is refused every orchestration call with "unauthenticated: anonymous may not SPAWN". The daemon still starts./healthzstill returns ok. The only symptom is that the fleet cannot be driven at all. Bisected on fleet01 (lsof line count): no sandbox 3,ProtectSystem=strict0,ProtectHome=read-only0,ProtectKernelTunables0,ProtectControlGroups0,RestrictSUIDSGID3,NoNewPrivileges3 — so the last two are safe and are kept. PrivateTmp=truemust be false on both units. fleetd writes the member ZDOTDIR credential-scrub directory and the opencode config directory underjava.io.tmpdir, and the member pane — a child of the other unit — has to read them back. A private/tmpturns the credential scrub into a silent no-op.- systemd runs no login shell, and every secret the daemon needs lives in a file only the login
shell sources. Started with a bare
ExecStart=java, fleetd boots fine with empty credentials and the failure appears hours later as a member that cannot open a pull request.
deploy/herdr.service did not exist at all, even though fleetd.service's After=/Wants=
already named it. fleetd #360.
One thing to know for maintenance. A unit file has no compile step, so SystemdUnitSafetyTest
is the guard: it reads all three files and fails on an active forbidden directive, on
PrivateTmp=true, on an ExecStart that skips the login shell, on a herdr-inner.sh that does
not exec a login shell, and on one that does not set a non-zero pty size (stty rows N cols M
with both positive — a 0x0 pty makes every pane spawn fail with ghostty error -2). Each failure
message names the consequence rather than the rule, because the rule alone is what someone deletes.
A commented-out mention inside the file's own DO-NOT-add block must not trip the test — that
comment is the whole point of the ticket, and a vacuity guard pins that it is still there, that all
three files exist, and that none is trivially small. Measured on merge: 1420 green; adding
ProtectHome=read-only to herdr.service fails 1 test; removing the login shell from
herdr-inner.sh fails 1; removing its stty fails 1; deleting that file errors 3 and fails the
build rather than passing vacuously.
The host no longer idle-sleeps while a member is working
fleetd now keeps the machine awake for as long as at least one member is live, and lets it sleep
again once the last one goes. On macOS it does this by holding a caffeinate -i child process.
The knob. A new top-level block, on by default:
idleSleepGuard:
enabled: true # set false to turn the guard off
It is a deferred key: the daemon reads it once at startup, so a change needs a restart. A reload reports it as such rather than pretending it applied.
Why it exists. The Mac that runs this fleet was set to idle-sleep after one minute on battery
(pmset -g custom reported sleep 1). Over one night the daemon's AMQP link dropped 13 times, and
every drop had a sleep or wake event in pmset -g log in the same minute or the minute before. The
broken AMQP link is only the visible symptom. The real cost is a member that freezes mid-turn with
the host — and a long turn with nobody typing is exactly the case that goes idle. Before this, the
fleet needed a human sitting at the keyboard to keep running, which defeats the point of delegating
long work.
How it knows. The guard hangs off SessionManager's existing onAcquire/onRelease hooks and
its size(). It does not count members a second way, so it always agrees with the numbers
fleet_list reports. Only a real 0→1 or 1→0 crossing touches the OS.
Gotchas — three, and all of them are by design.
- It is macOS-only.
caffeinateships on no other platform, so on Linux — fleet01, for example — the guard is a clean no-op. It says so once at INFO and never again, so a daemon running for weeks does not fill its log. fleet01 has the same exposure and does not get the fix from this change; a Linux mechanism (systemd-inhibit) is a separate job. -iis idle sleep only. Closing the lid still sleeps the host, and so does an operator asking for sleep. That is deliberate: the guard stops an unattended host sleeping under a member's turn, it never overrides the operator.caffeinate -s/-dwould do that and are not used.- It fails safe, and silently. If the mechanism cannot start — binary missing, process table
full —
acquire()returns null and nothing is ever held. The guard never throws, and never blocks a spawn, a release or shutdown. So "the guard is enabled" is not proof the host is awake. The proof is a livecaffeinateprocess, or a member that survives an idle night.
A reply says whether anything was waiting for it
What. fleet_reply and POST /sessions/{id}/reply now report which of three things happened to
a worker's reply, instead of saying "delivered" for all of them:
| Outcome | What happened | REST delivered |
|---|---|---|
resolved_send |
A fleet_send or fleet_ask was actively waiting, and took the reply now |
true |
resolved_async_ticket |
No live waiter, but the reply completed a parked async ticket — a fleet_poll caller sees it at once |
true |
queued |
Nothing was waiting. The reply is held in the inbox for a later drain | false |
Over MCP the tool result carries the wording, for example queued — no send or ticket was waiting; held in the inbox for a later drain. Over REST the body gains an outcome field, and delivered
stops being a constant.
The knob. None. It is how both doors answer now.
Why it exists. All three outcomes are successes, but they are not the same fact, and the caller
could not tell them apart. "Delivered" for a reply nobody was waiting for is the report reading
better than the state — the same failure this project keeps finding in other places. A lead that
sees queued knows its send never opened, or had already timed out, which is a real and different
situation from a clean handoff. Before this, that difference was visible only in a metrics label
nobody reads during a task.
The nudge counters were renamed in the same change, for the same reason. fleet_push_nudges_total
and fleet_lead_heartbeat_nudges_total used to record an outcome called delivered; it is now
sent. Nothing about the count changed — only the word. That call is a one-way herdr
paste-and-submit into a pane, and there is no read-receipt concept at that layer, so "sent" is the
most the counter can ever honestly claim.
Gotchas.
delivered: falseover REST is not an error. The status is still200, and the reply is safely held. Any client that treatsdelivered: falseas a failure and retries will queue the reply twice. This is the one behaviour change that can break an existing caller, and it is the reason theoutcomefield exists: branch onoutcome, not ondelivered.queueddoes not mean the lead will never see it. It means nothing was waiting at that moment. The reply is drainable withfleet_poll{target}, and the push loop nudges the lead.- The wording is not a contract;
outcomeis. The human-readable text may be reworded. The threeoutcomevalues (resolved_send,resolved_async_ticket,queued) are the stable names.
fleetd #365.
Startup says which profiles have usage-limit detection turned off
What. exhaustedPattern is the per-profile regex that recognises "you are out of quota" in a
backend's own words. It is opt-in, and a profile without one has usage-limit detection off. Two
places now say so:
- At startup the daemon logs one aggregate
WARNnaming every profile with noexhaustedPattern, plus a second, louderWARNif any of them is a subscription profile. When every profile is armed it logs anINFOinstead, so the healthy case is also on the record. - In the API
fleet_profilesandGET /profilescarryexhaustionDetectionArmed, a boolean per profile. A lead can read the state without shell access to the daemon's log.
The knob. None to turn this on. The knob it reports on is profiles.<name>.exhaustedPattern;
set one to arm detection for that profile.
Why it exists. A missing exhaustedPattern failed in the worst way an opt-in can fail: nothing
was wrong, nothing was logged, and the profile kept accepting spawns. When the credential really
hit its limit, the backend said so in prose, fleetd did not recognise it, and no quarantine
started — so the fleet kept spawning members onto a dead credential. The operator's only clue was
members that spawn fine and produce nothing. A subscription profile is the sharp case, because a
subscription is the thing that actually runs out.
Gotchas.
- Armed is not correct.
exhaustionDetectionArmed: truemeans a pattern is configured, not that it matches what this backend prints. A wrong pattern reports as armed. - After a reload, the field lies — fleetd #404, open.
exhaustedPatternis a deferred key: the patterns are compiled once into a map at startup and a reload never re-reads them. The field reads the live config instead. So if you add a pattern tofleetd.yamland reload, the field flips totruewhile detection stays off until the daemon restarts. Until #404 lands, trust the field only on a freshly started daemon, and restart after editingexhaustedPattern— the reload report is the honest door here, and it namesprofilesas needing a restart. errorPatternhas the same opt-in shape and is not covered by this report. It is less severe: with none set, the code falls back to a narrow built-in pattern rather than going inert.- The startup call site is not pinned by a test. Deleting
reportExhaustedPatternGap(cfg)fromFleetd.javaleaves the suite green (measured at the merge: 1472 tests, 0 failures). The report's own behaviour is tested; that it is still called is not. This is true of all four startup reports, not just this one —reportGitHostShape,reportMemberTrustModel,reportMemberCredentialsGapandreportExhaustedPatternGapare each referenced by exactly one test file, and that test calls the method directly. fleetd #442 owns closing it. The sixFleetConfig.validateXxx()startup calls had the same defect and it is now FIXED: they were collapsed into onecfg.validateAll(), whichFleetdStartupValidationTestpins by calling the realFleetd.mainand asserting it refuses a bad config. (An earlier version of this line said "fleetd #398 owns closing it". That was wrong: #398 is the closed models allow-list PR.)
Two limits of the detection this reports on. Both bound what any recovery feature can do, so read them before designing one.
- Detection needs a worker turn that finished.
CompletionResolverclassifies a usage limit by matching the member's own terminal output when its turn completes. There is no HTTP status or header path into this. So fleetd cannot learn a limit has been hit until a member has run and ended, and it cannot cheaply ask "is the limit lifted yet?" — any auto-resume has to spend real backend work to find out. - No reset time is delivered, anywhere. The captured live refusal is
The usage limit has been reached. Try again later.— there is no time in it, and nothing in the refusal path reads aRetry-After.BackendQuarantinestoresnow + cooldownNanosand nothing else. Recovery can therefore only ever be a fixed timer or a backoff probe, never a resume scheduled for the real reset moment. - Every
subscription: trueprofile shares one quarantine key. With no explicitcredentialId,effectiveCredentialId()returns the<subscription>sentinel, which is right for a single Claude plan. The consequence is easy to miss: on this fleet the lead's own profile (opus) shares that key with the worker fan-out profile (sonnet). ArmexhaustedPatternonsonnetalone and worker exhaustion starts quarantining the lead's seat too. Arm them together, or not at all.
fleetd #395.
A central allow-list of the models the fleet may use
What. A top-level models: block names every model id the fleet is allowed to run. Once it is
non-empty, a profiles: entry naming a model that is not on the list refuses to start, and
refuses a reload too.
models:
allow:
- model: claude-sonnet-5
- model: openai/gpt-5.6-terra
- model: gx/deepseek-v4-flash
model: is one flat, opaque string namespace. A bare Claude id and a provider-prefixed opencode id
both fit unchanged, because the check is exact string equality — it never parses a provider prefix
and never branches on kind:.
The knob. models.allow. Absent or empty keeps the old behaviour, where no model was ever
checked, so this ships inert until an operator writes the block.
Why it exists. Before this there was no single place that said which models the fleet may use. Each profile named one, and a typo or a withdrawn model id reached the backend adapter as a free-form string. On an opencode profile that has a specific bad outcome, already recorded here: a withdrawn model name makes opencode fall back to a paid model silently. An allow-list turns that class of mistake into a refusal at load, which is the cheapest place to find it.
Why it is operator-owned and not checked against a vendor catalogue. For an opencode profile
fleetd synthesizes the provider from the provider/model selector plus baseUrl
(OpenCodeLauncher:562-583). So a perfectly valid fleetd model id can appear in no published
catalogue — gx/deepseek-v4-flash on this fleet is exactly that, absent from models.dev and
correct. Any attempt to validate the list against a vendor catalogue would reject working
configurations. The list is the source of truth; nothing else can be.
Gotchas.
- Adding a profile means adding its model in the same edit. Otherwise the daemon will not start. That is the intended trade: the failure is loud and immediate rather than silent and later.
models:used to be deferred. Since fleetd #422 it is hot. A reload validates the new block, so a bad edit is still refused and the running config kept — and a good edit now takes effect on the very next spawn, with no restart. If you are reading an older note that says amodels:edit needs a restart, that note is stale.- The list is a name gate, not a capability check. It says the operator permits this id. It does not say the backend serves it, that the credential may use it, or that the id is spelled the way the provider spells it. A model on the list can still fail at spawn.
- Cross-check the two lists rather than trusting either.
grep -E '^\s+model:' fleetd.yaml | awk '{print $2}' | sort -uagainst theallow:entries. If they differ, the daemon refuses to boot — which is the point, but it is better to know before a restart than during one.
fleetd #398.
fleet_reply tells a lead which tool to use instead
What. A lead that calls fleet_reply is refused, and the refusal names both working routes:
fleet_reply has no route to a peer lead. Use fleet_send{coordId: ...} for a peer on another
daemon or fleet_send{sessionId: ...} for a peer on this host. fleet_reply resolves a member's
blocked fleet_send, and a peer's coord-id message is durable and non-blocking, so there is
nothing for it to resolve.
The knob. None. The check runs on every fleet_reply call and needs no configuration.
Why it exists. A lead that gets a message from a peer reaches for "reply" — the word matches
what it is doing. But fleet_reply exists to resolve a member's blocked fleet_send, and a peer
lead never has one open. So the call used to be accepted and the message went into a worker inbox
nobody would ever drain. The peer waited on nothing, and no error said so.
The old refusal was worse than none, because it only covered the case where the caller could not be identified at all. A lead is identified, so it sailed past that check.
The message names the routes on purpose. A refusal that says only "not allowed" makes the lead
guess, and the two right answers depend on where the peer lives — coordId across daemons,
sessionId on this host.
Gotchas.
- The order of the two checks matters, and is now pinned. An unidentified caller gets the
"workers only" message even if its role is
PRIMARY. That is deliberate: "we do not know who you are" is the more useful thing to hear first. - A queued reply is still a success for a worker. This refusal is about leads only. A worker
whose
fleet_replyfinds no open send still gets its reply stored in the inbox, and that is normal. - Role comes from the connection, never from an argument, so a caller cannot present itself as a worker to get around this.
fleetd #391.
The credential-scrub receipt now measures the blank, not the attempt
What. The startup scrub that clears inherited credentials from a member's shell reports which names it actually blanked. It now checks each parameter's value after trying, instead of trusting the exit status of the attempt.
allowed 41 of 57 failed 3
<name> ← blanked, confirmed empty
!<name> ← could not be blanked
The knob. None. The receipt is part of the scrub and always runs.
Why it exists. The old loop counted a name as blanked when eval "export ${n}=" returned 0.
zsh has integer parameters — SECONDS RANDOM SHLVL HISTSIZE COLUMNS LINES USERNAME — and on those
an empty assignment is coerced to a number rather than failing. So eval returns 0 and the
value is unchanged. Measured across 10 names: 7 false receipts.
The general rule, which cost two tickets to learn: an attempt's exit status is not a measurement
of its effect. The fix verifies afterwards with [[ -z "${(P)n}" ]] and only then counts it.
Gotchas.
evalis required, not stylistic.export UID=is a fatal zsh parameter error that aborts the whole sourced file — it once killed the scrub at name 42 of 57 and left the operator's own exports untouched. Onlyeval "export ${n}=" 2>/dev/nullcontains that. The same applies toEUID GID EGID PPID LINENO.- The
!names are the ones to read. They are shell parameters that cannot be blanked, not credentials that leaked. A long!list is normal; a shrinking allowed count is the warning. - The scrub is zsh-only, like the secret store it defends against. It lives in
.zshenv, the one file zsh always reads. - The receipt reports the member's shell, so a claim about it measured from inside an agent's
zsh -cchild is measuring the wrong process.
fleet_list no longer advertises a profile that fleet_spawn will refuse
What. The capacity rows in fleet_list now list the profiles the daemon can really spawn on —
the set it read at startup. Before, they listed the set in the live config, so a profile added by
a hot reload showed up with free slots and every spawn onto it failed.
# operator adds profile "ghost" to fleetd.yaml and the config reloads
fleet_list -> capacity: [... {profile: ghost, maxLoad: 3, live: 0, free: 3}]
fleet_spawn{profile: "ghost"} -> refused: unknown worker profile
The knob. None. This is the capacity reporting in fleet_list, and it always runs.
Why it exists. profiles is a "both ways" key. Per-profile tunables — maxLoad, weight,
credentialId — are hot, so a reload changes them at once. The profile set is frozen,
because HerdrPeerLauncher takes Map.copyOf(profiles) once when it is built and never looks
again. So "profiles is deferred" and "this live read is fine" are both true, of different halves of
the same key. The old wiring read the live keySet() and the hot maxLoad() through one lambda,
which looked consistent and was half wrong.
The direction of the error is what made it worth a ticket: it overstated a capability. A lead
reading free: 3 had no way to tell that number from a real one, and only found out by spawning.
Gotchas.
maxLoadis still hot, and must stay hot. The fix must freeze only the set. A test pins that: it changesmaxLoadfrom 3 to 9 by reload and asserts the sameCapacitySourceinstance reports the new number. Freezing both would be the mirror-image regression.- The right rule was already written three lines below, for
coordinator.peers: read from the same snapshot the collaborator itself was opened from, rather than the live config, so the report follows one rule instead of half hot-reloading. - A fresh-daemon test can never catch this. On a daemon that has not reloaded, the live config and the startup snapshot are the same object, so both the wrong and the right wiring pass. The test has to perform a real reload and assert the reload applied, or it proves nothing.
- Both directions need a test. A single test that only adds a profile passes if the source returns a permanently empty set. The second test — a startup profile is listed — is what makes the first one load-bearing.
fleetd #416. Found by the fleet01 lead on a live daemon; the same shape as #404, where a status field read a different source than the behaviour it described.
The startup log tells "off" apart from "running on the built-in default"
What. Two startup lines report whether a backend-classification is running. The
backend-error line no longer says off when no profile sets an errorPattern, because that
classification is still running — it falls back to a built-in pattern. The backend-exhausted
line still says off, because for that key it is true.
backend-error classification: built-in default for all profiles
(no profile customises errorPattern; profiles: [gx, local, opus, xf])
backend-exhausted classification: off
(no profile has an exhaustedPattern configured; profiles: [gx, local, opus, xf])
The knob. None. Both lines are logged at INFO at every startup.
Why it exists. The old line was produced by one shared helper that knew only the key's name, so it worded both keys the same way. But the two keys disagree on what "unset" means:
- an unset
errorPatternfalls back toCompletionResolver's built-in(?i)\bAPI Error\s*:compatibility pattern, so classification keeps running; - an unset
exhaustedPatternhas no fallback, so it really is off.
One string, two meanings — and the false one was the reassuring direction. It told an operator a
live classifier was disabled while it was running, which is the worst thing to read when you are
debugging a false quarantine. The fact needed to word it correctly was written down about 1600
lines away, in FleetConfig's javadoc, and the method that got it wrong had no way to reach it.
Gotchas.
- The classification really is live even with no
errorPatternanywhere. The built-in pattern is narrow but it can still fire on a member's own prose that quotes anAPI Error:line. That is why the outage policy needs the line twice before it acts. - "no profile customises it" is the useful reading, not "nothing is configured". If you want a different pattern for a backend, set it; the absence of the key is a choice of the default, not an off switch.
offforexhaustedPatternis real. Do not assume symmetry between the two lines — that assumption is the bug this entry describes.- The wording is now the caller's responsibility.
coverage()takes a requiredUnsetMeaningargument, so a third pattern key cannot compile without stating what unset means for it. There is no permissive default to inherit.
Measured, and worth keeping. The first fix was correct and still not safe: its tests all passed
the meaning in themselves, so they proved the wording and not the pairing. Swapping the two
arguments at the two call sites recreates the original bug with the keys exchanged, and that left
all 1506 tests green with BUILD SUCCESS. The call sites are now extracted into two named
factories and pinned by FleetdPatternCoverageLineTest; the same swap gives 3 failures. The
general rule: when a fix adds a parameter so the caller can supply a missing fact, test the
caller's choice of value — a test that passes the value in itself tests the half that was never
broken.
fleetd #415. Found by the fleet01 lead on their own startup log.
Turning a model off at runtime, without editing profiles:
What. Each entry in models.allow now takes an enabled: flag. Set it to false and the
fleet stops spawning on every profile that names that model — at once, with no restart and no edit
to profiles:. Set it back to true and those profiles are usable again.
models:
allow:
- model: claude-sonnet-5
- model: openai/gpt-5.6-terra
enabled: false # every profile naming this model is now unspawnable
- model: gx/deepseek-v4-flash
enabled: is optional and defaults to on, so every existing models: block keeps working
untouched. An entry that is off stays in allow — do not delete it. allow answers "may the
fleet use this id at all", and enabled answers "may it use it right now". Removing the line
instead of turning it off makes every profile naming that model fail validation, and the whole
reload is refused.
Seeing the gate's own state
fleet_profiles and GET /profiles report the gate, so a lead can see it without reading the
config file. Two fields, and you need both:
| field | meaning |
|---|---|
modelGateArmed |
is there a models: block at all. Always present, true or false. |
modelsOff |
which model ids are off right now. Absent when nothing is off. |
The off set alone cannot answer the question you usually have. An empty off set means one of two
very different things — there is no models: block on this host, so nothing is gated and nothing
can be; or there is a block, it is working, and right now nothing is turned off. modelGateArmed
separates them, which is why it is reported even when it is false.
The daemon logs the same three states at startup, from the same read:
model gate (fleetd #422): not configured (no models: block — nothing is gated, and nothing can be)
model gate (fleetd #422): armed (models: block present; 0 models currently turned off)
model gate (fleetd #422): armed (2 model(s) turned off: [openai/gpt-5.6-terra, sol/x])
Both the log line and modelGateArmed come from one PeerLauncher.modelGateState() call, which is
the same accessor the spawn gate itself reads. So the status can never claim more or less than the
gate enforces, and a reload landing between two reads cannot make them disagree.
The knob. models.allow[].enabled. Absent means on.
Why it exists. A subscription runs out. When it does, the profiles on it must stop taking work
while the rest of the fleet keeps going. Before this the only ways to do that were to edit every
affected profiles: entry, or to let each spawn fail against the dead backend and wait for the
quarantine to catch up. One central switch keyed by model is the right shape, because a model
is what a subscription sells — several profiles usually share one.
Where the gate sits, and the one thing to know about it. fleet_spawn takes two different
paths, and they treat a bad profile in opposite ways:
flowchart TD
A["fleet_spawn"] --> B{"did the caller<br/>name a profile?"}
B -->|"yes — an operator override"| C["check quarantine, cooling off,<br/>maxLoad, model-off"]
C --> D["REFUSE the spawn"]
B -->|"no"| E["placement: build the<br/>quarantined / coolingOff / modelOff sets"]
E --> F["the policy SKIPS those profiles"]
F --> G["spawn on the next candidate"]
classDef stop fill:#b7791f,stroke:#7b341e,color:#ffffff;
classDef go fill:#2f855a,stroke:#22543d,color:#ffffff;
class D stop
class G go
Naming an off-model profile explicitly is refused, and that is deliberate — it is the operator
overriding, so a clear refusal beats a silent redirect. Leaving the profile blank routes around it
instead. Both behaviours are correct; they are just not the same, and code that resolves a profile
name before calling spawn silently moves itself from the second path to the first.
Gotchas.
- A host with no
models:block is not gated. The whole feature is inert there, and that is permitted, never fatal. On this fleet,fleet01has nomodels:block, so the gate ships green and does nothing on that host. Check withgrep -c '^models:' fleetd.yamlbefore assuming a second host is protected. - The gate is a name gate, like
allowitself. Turning a model off stops fleetd spawning on it. It does not reach a member that is already running. - Off is not quarantine, and the message says so. A model-off refusal names
models.allow; a quarantine names the credential and the seconds left. When a profile is both, quarantine is reported, because a backend fact outranks an operator preference. fixedis the default policy, and it applies this filter in two places — the default fast path and the fallback walk. The first round of this feature put the filter in the shared helperPlacementPolicyUtil.available(), whichfixednever calls, so it shipped green and inert under the default policy with 1519 tests passing. Both new tests had usedweighted().
Still missing. Detecting a subscription limit and flipping the switch automatically is not built. Today an operator turns the model off and back on by hand. That decision is open.
fleetd #422.
Removing an architect slot now actually revokes it
What. An architect's extra rights come from a slot in fleet.architects. Delete that slot from
the config and reload, and the session bound to it drops to worker rights on its very next request.
Before, it kept full architect rights until the daemon restarted.
The knob. None. This is how fleet.architects behaves on reload.
Why it exists. An architect can do things a worker cannot, so "revoke" has to mean revoke. The old code flattened the slot list once when the daemon started and never read it again, so removing a slot only refused the next spawn. The session already holding the privilege kept it — a revocation that did not revoke, with nothing in the logs to say so.
Gotchas.
- The privilege is revoked; the slot stays occupied. The demoted session keeps its slot key
until it unbinds. That is on purpose: freeing the key at once would let a second terminal claim
the slot the operator was actually trying to shut down, and it would break
unbind, which needs the original terminal-to-slot pair intact. - The demoted session can still finish its turn.
fleet_replyis authorised by terminal identity, not by role, so a demoted architect ends its turn normally instead of stalling. - Do not confuse the binding with the privilege. They are separate, and wording that treats them as one thing is what produced the first, wrong version of this fix — the ticket asked for a test that a bound architect "survives the rebuild", which is true of the binding and false of the privilege.
fleetd #424.
A worker's fleet_list no longer carries the lead's coordination state
What. fleet_list's reply used to include a coordinator block for every caller: this
daemon's own coord-id, its mailbox state, a preview of the peer mail held for it, and each
configured peer's live reachability. That is lead-to-lead state. Now a worker or an architect
calling fleet_list gets no coordinator key at all. The key is absent, not present and
empty.
The knob. None. It follows the caller's role, which the daemon resolves from the connection, not from anything the caller sends.
Why it exists. A worker has no use for peer names, held-mail previews, or this daemon's
coord-id, and it should not learn them from a roster call it makes for other reasons. Absent beats
empty on purpose: an empty object still tells the caller the feature is configured, and it makes a
client that tests if (coordinator) behave differently from one that tests
coordinator.peers.length. So the gate runs before the row is built.
Gotchas.
- A worker cannot tell "no coordination configured" from "not for you". Both look like a missing key. That is the intended trade: the alternative leaks the fact that peers exist.
- The lead sees no change. Same key, same fields.
- The compat overloads used to default to showing the row, and no longer do.
listFleethas seven declarations. Six of them do not take the caller's role, and they all inherited the default from one line. That default wastrue, so a call site that forgot the argument would have disclosed the row silently. It is nowfalse: a missing identity means a missing row, which is a visible bug rather than a quiet disclosure. Fixed in fleetd #463. A test calls a compat overload with no boolean and asserts thecoordinatorkey is absent.
fleetd #439.
A usage limit now says which model to turn off, and detection can be armed without a restart
What. Three related changes to how fleetd reacts when a backend reports that a subscription limit is reached.
-
exhaustedPatternis now a hot config key. Before, it was compiled once into a map at daemon startup, so arming detection for a profile needed a restart. Now the pattern is read live per profile, and compiled patterns are cached by pattern string so a check does not recompile a regex every time. -
The warning logged when a credential is quarantined names the fix, not just the fact:
usage-limit fix: profile 'X' runs model 'Y' — set `enabled: false` on that model's entry under models.allow in fleetd.yaml to stop new spawns landing on it (models: is hot, no restart needed); remove the line again once the subscription window resets -
fleet_profilesreports two new fields on a quarantined row:model(the model that profile runs) andreason(the backend text that triggered the most recent quarantine of that credential).
The knob. exhaustedPattern: on a profile, which is opt-in and off by default. It is now hot,
so a reload arms it. Turning a model off is enabled: false on its models.allow entry, which
was already hot.
Why it exists. The two halves of model gating were reloadable in opposite directions, and it
was the wrong way round. An operator could already turn a model off at runtime, but could not
arm the detector that tells them to — that needed a restart, on a feature whose whole point is
reacting to a limit while the fleet is running. The reason field exists because "this
credential is quarantined" does not say whether the cause was a usage limit or something else,
and the two need different responses. The warning names the fix because an operator reading a
quarantine line should not have to work out which of several configured model names to edit.
Gotchas.
reasonis only written where a quarantine actually happens, keyed by credential id — the same key the cooldown itself uses. So a reason can never be reported for a quarantine that did not occur.- A pattern has to match real text from that specific backend. Do not guess one. A guessed
regex gives you a profile that reports
exhaustionDetectionArmed: trueand silently never fires, which is worse than an honestfalse. - There is no
off:config key. The flip isenabled: false.offis the name of the reported set (modelsOff), which is an output. - Nothing re-enables a model automatically. "On again when the limit lifts" is a manual edit,
which is cheap because
models:is hot and needs no restart. An automatic backoff probe was considered and rejected: a probe spends quota to discover quota, so on a metered subscription it burns the first tokens of every new window on discovery instead of work. - Nothing proves the daemon wires this sink. The quarantine sink was extracted into a factory
and what it logs is pinned by a test, but replacing
main()'s call to that factory with an inert lambda leaves the whole suite green. Tracked in fleetd #460.
fleetd #446.
A lead session can replace itself when its context fills up
What. A lead (primary) session runs out of context and has to be replaced by a fresh one. Until now that was entirely manual: the lead wrote a handover file, told the operator where it was, and the operator started a new session by hand and pointed it at the file. Now the lead can ask fleetd to do the swap.
The cycle has three steps, and the order matters:
fleet_handover{action: "open"}— fleetd records a token and tells the lead thehandoverPathit must write to.- The lead writes the handover file. The
handoverskill is the procedure for what goes in it. fleet_handover{action: "confirm", token, operatorConfirmed}— fleetd checks every gate and, if all of them pass, schedules the roll: wait for the lead's own turn to end, send/clearto its pane, wait again, then send a bootstrap prompt naming the handover file. A fresh session reads the file and carries on.
{action: "cancel", token} drops a pending request without rolling.
The knob. A new top-level leadRollover: block in fleetd.yaml. It is opt-in and off by
default — with the block absent, fleet_handover is still registered but every action answers a
clean refusal naming NOT_CONFIGURED.
leadRollover:
handoverPath: .handover/HANDOVER.md # required when the block is present; relative is allowed
requireOperatorConfirm: true # default true
maxDocAgeSeconds: 3600 # default 3600
turnSettleSeconds: 20 # default 20
clearSettleSeconds: 20 # default 20
bootstrapText: "…" # default names handoverPath
Every field is read fresh on each call, so the values are hot. Adding the block where it was
absent at boot still needs a restart, because Fleetd.main only constructs the executor when the
block is present in the startup snapshot. That is an existence fact, not a stale-value one — the
same way a brand-new profiles: entry needs a restart while an existing profile's fields do not.
Why it exists. A lead that fills its context is the one agent nobody else can replace: it holds the orchestration state, and a worker cannot restart its own lead. The whole cycle previously stopped dead waiting for a human to notice. Three design choices were deliberate and are worth not re-litigating:
- The lead asks; nothing watches it. There is no timer, no heartbeat, and no background loop
that can decide on its own that a lead should be replaced. Only an explicit
confirm()call that passes every gate can ever cause a/clear. A context-pressure detector was considered and rejected — the cost of a false positive is a destroyed live session. /clearin the same pane, not stop-and-relaunch. Relaunching would lose the pane's identity, and a lead is pinned to its terminal by a tab label (fleet.leaders.<name>.tab), so a new pane is a new lead as far as the daemon is concerned.- The operator confirms too.
requireOperatorConfirmdefaults totrue, so the lead's own judgement is not enough to wipe a session.
Gotchas.
-
Write the file after
open, never before.confirmrefuses withHANDOVER_STALEunless the file's modified time is later than theopenrequest. That check exists so a leftover file from a previous session can never be accepted as this session's handover, and it means the obvious order — write the file, then ask — is the wrong one. -
A relative
handoverPathresolves against the lead's workspace, not the daemon's. The daemon and the lead run in different directories — on this host the daemon sits in<repo>/fleetdand the lead in<repo>— so a bare relative path would mean two different files.open()resolves it once, against the calling lead's configuredfleet.leaders.<name>.cwd(falling back to the daemon's own working directory when that lead has nocwd), and hands back an absolute path. Everything downstream — theopenresponse the lead writes to, the file the daemon stats, and the bootstrap prompt the fresh session reads — uses that one absolute path. An absolutehandoverPathis used unchanged. -
A relative
handoverPathlands in the LEAD's repo, which is usually not the repo that ignores it (#491). The handover file is a snapshot of live state — unpushed branches, open questions — so it must never be committed. The trap is which.gitignoreprotects it. On this host the lead'scwdIS thefleetdcheckout, so the rule infleetd/.gitignoreworks and the distinction is invisible. On fleet01 the lead'scwdis/home/ltms/LTMS/kbwhile fleetd sits in/home/ltms/LTMS/fleetd— two different repositories, andkb/.gitignorehas no.handoverrule. Measured 2026-09-12. So: put the ignore rule in the lead's workspace repo, or give an absolute path outside every repository. Re-check withgrep -n handover <lead workspace>/.gitignore— no output means the trap is live on that host. -
confirmreturningaccepteddoes NOT mean the pane has been cleared. It means every gate passed and the roll is scheduled. The roll itself runs after the calling turn ends, and it may still refuse at that point; those outcomes are logged only, because there is no caller left to answer. Look forlead-rollover:lines in the daemon log. -
A lead can only ever roll itself. The tool has no terminal, session or lead parameter of any kind. The pane comes from the caller's own connection. This daemon can hold more than one labelled lead tab, and an earlier draft that looked the pane up in a single-slot registry let one lead clear another lead's pane.
-
/clearmust never go throughInjector. It produces no turn boundary, so the Injector's turn never completes and every later message to that pane queues behind it for ever — the pane is wedged.LeadRollovercallsagents.send(...)directly, exactly likeClaudeCodeLauncher#clearContext. Measured, not assumed. -
The settle waits require
IDLEorDONE, notinjectable().injectable()also acceptsBLOCKED, which is a live turn paused on an approval prompt. Treating that as settled sent/clearinto an open prompt mid-turn. -
The first real roll failed, and the fix is merged but not yet proven live (#489). On 2026-09-12 the roll put one line into the pane —
/clearFresh lead session. …— and Claude Code answeredUnknown command: /clearFresh. That is/clearand the bootstrap text joined with no space, because a paste lands at the cursor. Two faults. The second settle wait was a no-op:/clearstarts no turn, so the pane never leavesIDLEand the wait returned on its first poll — the whole roll ran in 438 ms of a 20-second budget. Under that, the submit keystroke raced the paste, whichAgentControl.submit's own javadoc already records as CB-113;LeadRolloverbypassesInjectoron purpose, so it got none of the Enter-nudging that makes a/clearland everywhere else. The failure was safe: nothing was cleared and no context was lost. PR #490 replaced the second wait withwaitForClearPickupAndSettle, which nudges the submit keystroke while no pickup has been seen — the same patternInjectoralready ships for its own post-turn/clear(#306). Merged and deployed on 2026-09-12. Acceptance criterion 6 is still not met: no roll has yet bootstrapped a fresh session end to end, and only using the feature for real on a live lead can settle it. Until then, treat the manual path as the reliable one. -
A settle wait that never settles hangs the test suite instead of failing it (#486). The poll loop is bounded only by an injected clock, and the test seams pass a clock that never advances together with a no-op sleeper. Both
waitUntilAtTurnBoundaryandwaitForClearPickupAndSettlehave this shape. In production the bound is real; in the suite it is inert.
fleetd #480 (PRs #483, #484, #485). Related: #486, #489 (PR #490).
A timed-out send now says whether delivery was even attempted
Before this, a fleet_send that timed out reported one of two things: TIMED_OUT_WORKING if the
message was delivered, or TIMED_OUT_QUEUED if it was not. There was no third answer, so a case
that is neither got filed under "not delivered".
That case is real. When the injector cancels a delivery it reports DELIVERED, NOT_DELIVERED, or
ATTEMPTED — meaning the keystrokes were already going out and nobody can say whether they
landed. The send path collapsed ATTEMPTED into TIMED_OUT_QUEUED, which promises the message
never arrived. A lead reading that promise retries, and the worker gets the same brief twice.
What it does. MessageService.Outcome gains TIMED_OUT_UNCONFIRMED. The ATTEMPTED
cancellation now routes to it instead of to TIMED_OUT_QUEUED. Every reader handles it:
- MCP —
fleet_sendreturns[no reply within Nms — delivery unconfirmed; the message may already have reached the worker, so a retry risks sending it twice — poll status before resending]. It deliberately does not carry the "retry or poll status" wording the queued/working arm uses, because on this route a resend can double-deliver. - REST —
POST /sessions/{id}/messagesanswers 202 with"status":"unconfirmed". - Metrics — the
fleet_sendscounter still labels ittimeout, grouped with the other two timeout outcomes. That grouping is deliberate and was left alone.
The knob. None. It is a behaviour change on an existing path, live as soon as the daemon restarts.
Why it exists. A sentinel that means "no" and a sentinel that means "cannot tell" need opposite handling from the caller, and this code had only the confident one. Fixing it needs a third state, not a better guess — the same shape as fleetd #512.
The gotcha, and it is the interesting one. FleetApp.writeReply's inner switch carried
default -> "done", so any outcome it did not name told a REST caller the delegation completed.
Adding a constant would have been absorbed silently by that default. The fix deletes it and lists
all ten outcomes by name, so the compiler now catches the next missed one. Order matters if you
repeat this exercise: a default is exactly what suppresses the compile error you are trying to
provoke, so delete the defaults first, then add the new constant, or the proof comes back clean
and proves nothing.
Still open. The outcome's wire token is its constant name lowercased, not a pinned string — the
sibling enum ReplyOutcome does pin its own. That is fleetd #578, which has since grown: the same
enum emits four different tokens across four live surfaces, so the fix is not one accessor.
fleetd #571 (PR #580). Related: #512, #578, #586.
fleet_list and /healthz report whether the background loops are alive
Two loops keep the fleet honest: StatusPoller, which refreshes member status, and
SessionReaper, which retires idle members. Both have had a LoopWatchdog for a while. Nothing
outside the daemon could see it.
What it does. fleet_list gains a loopHealth object with keys statusPoller and
sessionReaper, each RUNNING, STALLED, or STOPPED. /healthz carries the same object in its
body. The 200/503 status codes are unchanged — a stalled loop does not turn the endpoint red.
The knob. None. Always on.
Why it exists. A stalled poller does not announce itself. The fleet keeps answering, member status quietly goes stale, and the first symptom is a lead acting on state that stopped updating hours ago. Exposing the watchdog turns a silent failure into a visible one.
The gotcha — the failure direction. This is monitoring, so ask what happens when the monitor
lies. The dangerous direction here is a false negative: report RUNNING while the loop is dead,
and the watchdog can never fire. That is worse than a false positive, because a false positive is
noisy and somebody mutes it, whereas a false negative produces no signal for anyone to notice is
missing. The wiring is what protects against it, and the wiring is now pinned:
FleetdLoopHealthSourceWiringTest fails if Fleetd.loopHealthSource stops asking the real poller
or the real reaper. That test exists because the first version of this feature shipped with five
green tests and the production wiring could still be replaced by a constant with nothing failing —
every one of the five built its own LoopHealthSource, which tests the consumer and can never be
evidence about the producer.
Reading it. STOPPED for sessionReaper is normal when no reaper is configured; that is the
null case, not a fault.
hunter is a member role, not just a skill
The repo already shipped a hunter playbook skill: sweep a package, report several ranked
findings, change nothing. There was no matching role. A hunt was run by spawning a dev and
telling it not to behave like one.
What it does. MemberRole.HUNTER joins ARCHITECT, DEV and REVIEWER. It has the wire
token hunter, the config pool fleet.hunters, and the agent file .claude/agents/hunter.md.
fleet_spawn{role: "hunter"} now picks a contract instead of borrowing one.
The knob. fleet.hunters in fleetd.yaml, listing the profiles a hunter may run on — the
same shape as fleet.developers and fleet.reviewers:
fleet:
hunters:
sonnet: {profile: sonnet}
terra: {profile: terra}
Optional. Leave it out and the role exists but no hunter can be placed.
Why it exists. A role carries prohibitions the brief should not have to repeat. dev is
allowed to edit, commit, push and open a pull request; a hunt must do none of those. Spawning a
dev for a hunt meant the only thing standing between the sweep and a surprise commit was a
sentence in the brief. Miss that sentence once and the member is inside its contract while doing
the wrong job. hunter.md forbids editing, committing, pushing and opening a PR, and — unlike
reviewer.md — explicitly permits running the build, because a hunter checks its findings.
The gotcha — the role ships inert. Merging the code does not create the pool. On a host whose
fleetd.yaml has no fleet.hunters, unknown nested keys are ignored rather than rejected, so
there is no warning at load and no error at merge: fleet_spawn{role: "hunter"} simply finds no
candidate profile. The sequence is merge, redeploy, add the pool, then spawn one hunter and read
the spawn log. Parsing the config proves nothing about placement.
The second gotcha — hunter and reviewer are still not interchangeable. reviewer caps its
answer at one finding; hunter reports several. Naming the wrong skill for the role hands the
member two contradictory output contracts, and the measured result is a member that writes a good
report to its terminal and ends the turn with no fleet_reply. The role does not fix that; the
brief still has to name the right skill.
The context-roll notice now obeys requireOperatorConfirm
What it does. When an idle lead's own context reads HIGH, the heartbeat loop appends a notice to
its nudge telling the lead how to hand over. The wording of that notice now follows the daemon's own
leadRollover.requireOperatorConfirm setting. With the default (true) it tells the lead to ask the
operator before confirming. With false it drops that instruction and tells the lead to decide for
itself, naming the gate that actually applies: the handover file must exist, must have been changed
after the open() request, and must not be older than maxDocAgeSeconds.
The knob. leadRollover.requireOperatorConfirm (default true). It is deferred, not hot —
read once at boot, so an edit does nothing until the daemon is redeployed.
Why it exists. The daemon's own gate, LeadRollover.confirm(...), already honoured this flag at
LeadRollover.java:480. So setting it to false did stop the daemon refusing a roll. But the text
the lead reads is built by LeadHeartbeatLoop.contextNotice(...), which took no config at all and
hardcoded "ask the operator … Only the operator can approve the roll". A lead follows the
instructions it is given, so it asked the operator anyway. The operator was interrupted for a routine
context roll exactly as before, and the config looked broken. Our operator asked for this directly:
"is it intended or? if yes, fix this behavior".
The gotcha — the knob used to be half a fix, and the failing half was silent. Nothing warned that the message and the policy disagreed. Changing config that the instruction text does not read is invisible to the agent reading that text, and no test or log line caught it. If you set this knob and the asking continues, check whether the daemon has actually been redeployed since — the config half and the code half each need their own restart to take effect.
The second gotcha — a config value that reaches a message needs its own test. The enforcement path and the wording path are two consumers of one setting, and pinning the first proves nothing about the second. The two tests added here assert on the returned string for both values of the flag. Inverting the branch condition kills three tests, including one written before this change.
📖 fleet
Home — overview & the decision
Chapters
- Architecture — system · 2 invariants · 2 modes
- Message Server — the
fleetddesign - Approaches — transports compared, why herdr
- Setup — ⚫ superseded by 13
- Operations — ⚫ superseded by 13
- Team — orchestrating a mixed fleet
- Use Cases — the review scenario + mechanisms
- Roadmap — delivery record: what is live, what is off, what was dropped
- Implementation — as-built code map · classes · flows · state machines
- Cross-Host Messaging — broker topology · exchanges · queues per entity
- Features — what it can do · the knob that turns it on · why · the gotcha
- Claude → OpenCode — porting a workspace to a second host
- User Guide — 🟢 install · configure · run · delegate · the traps
- Fleet Manager — many fleets on one host, over REST
- REST API Reference — all 14 routes, roles, and bodies
- Security & Trust Boundary — the guard · authz · what a member inherits
Design proposals (not built)
- CB-548 Lead Quorum — a deterministic decision procedure around a lead's judgment
🟢 herdr-centric fleetd · AgentAPI = research, never built