120
11 Features
lead edited this page 2026-09-22 11:31:36 +07:00
This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

11 — Features

What this page is for. The other chapters answer how is this built and why this way. This one answers "what can it do, and how do I turn it on" — one entry per operator-facing capability, so a feature that shipped six weeks ago is still findable without reading a design doc or a commit log.

What belongs here. A capability an operator can use, configure, or observe: an MCP tool, a fleetd.yaml knob, an endpoint, or a behaviour visible from outside the daemon. Internal contract changes go to Implementation; test and coverage work is a Roadmap line. If a change adds none of those, it has no entry here — that is a normal outcome, not an omission.

What every entry states, in this order: what it does · how you turn it on · why it exists · the gotcha. The why is the load-bearing line — it is what stops a decision being re-litigated in six weeks, and the table alone will not carry it.

Index

Capability Turn it on with Since Code
Ask the bridge who you are fleet_whoami CB-517 mcp/BridgeMcp
Primary inside a herdr pane primary.terminal: CB-522 auth/CallerResolver
More than one lead fleet.leaders: CB-530 auth/CallerResolver
Unknown config keys are named (always on) CB-530 config/FleetdConfig
Find leads by tab name fleet.leaders.<n>.tabPrefix CB-531 herdr/LeadTabScanner
One fleet block, role as the key fleet: CB-557 config/FleetdConfig
Role pools decide the backend fleet.developers: etc. CB-557 member/CompositePeerLauncher
Tab labels name the role fleet.tabLabel: CB-557 member/HerdrPeerLauncher
Launch a lead when none is live fleet.leaders.<n>.profile + instances CB-558 lead/LeadLauncher
Re-read the config without a restart configReload.enabled: true CB-559 config/ConfigRef
Set a role's launch charter in config fleet.charters: CB-566 config/FleetdConfig
Leads talk to each other (always on, two leads) CB-532 auth/Principal
A lead can be delivered to automatic CB-534 Fleetd.deliverableTo
Know when a completion fallback is partial automatic CB-563 inject/CompletionResolver
Leads are visible in fleet_list automatic CB-535 mcp/BridgeMcp.listFleet
Reply nudges follow the delegating lead automatic (retires primary:) CB-532 mcp/PrimaryRegistry
Advisory architect slots fleet.architects: CB-548 auth/MemberRegistry
Weighted worker placement placement: weighted + weight / maxLoad CB-518 placement/
Give workers a toolchain per-profile env: CB-511 worker/HerdrPeerLauncher
Run a worker on the subscription profile subscription: true CB-539 worker/ClaudeCodeLauncher
Keep a worker conversation session name + resume id on spawn CB-547 peer/SpawnRequest
Isolated worktree per worker fleet_spawn{worktree, ticket} CB-301-ext session/GitWorktrees
Worker tool-surface isolation automatic CB-525 session/GitWorktrees
Worktree-hostile config isolation automatic CB-543 session/GitWorktrees
No credentialed remote URL reaches a worktree automatic CB-157 / CB-189 session/GitWorktrees
Worker opens its own PR gitTokenEnv: / gitHostEnv: CB-302 worker/HerdrPeerLauncher
Session lifecycle caps lifecycle: CB-303 session/SessionManager
Keep a worktree that still holds work automatic CB-576 session/GitWorktrees
Fail a ticket when its member dies health.enabled: true CB-580 health/FleetHealthMonitor
Durable reply inbox broker: CB-307 msg/AmqpReplyInbox
Set a member's auto-compact window per-profile autoCompactWindow: CB-636 member/ClaudeCodeLauncher, member/OpenCodeLauncher
Leads on different hosts talk over a shared broker coordinator: CB-637 msg/LeadMailbox, msg/LeadCoordLoop
A lead can read its own held peer mail coordinator: #421 mcp/FleetMcp, auth/Authz
Draining an inbox is lead-only, on both entry paths (always on, with auth:) #272 mcp/FleetMcp.pollAction
Broker password out of the config broker.uriEnv: CB-635 config/FleetConfig
An unreachable broker does not stop the daemon (always on) CB-635 Fleetd.selectReplyInbox
Reject overlapping rendezvous automatic CB-548 msg/Rendezvous
Pin an opencode endpoint profile baseUrl: CB-508 worker/OpenCodeLauncher
Onboard a project with the plugin /plugin install claude-bridge → /claude-bridge:setup CB-527 plugin/
Port a workspace to OpenCode port-to-opencode skill + opencode.json CB-529 .claude/skills/port-to-opencode
Reject a profile name as a send target automatic CB-572 mcp/BridgeMcp
See free fleet capacity automatic CB-573 mcp/BridgeMcp
Answer a question on an async delegation automatic CB-574 msg/MessageService
Learn why a delegation died automatic CB-568 inject/CompletionResolver
Watch the fleet's health health: CB-573 health/FleetHealthMonitor
Tell a usage-limit refusal from a real reply profile exhaustedPattern: CB-578 inject/CompletionResolver
Stop spawning onto an exhausted account quarantineCooldownSeconds: + profile credentialId: CB-578 placement/BackendQuarantine
See which charter a member got automatic CB-571 peer/CharterReceipt
Redeploy the daemon safely redeploy-fleetd skill, then run the script — scripts/redeploy-fleetd.sh
Run members on a second herdr daemon memberHerdrSocket: CB-185 herdr/HerdrRouter
Let a member under another OS user write its worktree worktreeGroup: CB-185 session/GitWorktrees
Resume an opencode member's prior session automatic CB-206 member/OpenCodeSessionDiscovery
Fill in a member id the backend names late automatic CB-209 session/SessionManager
REST roster rows under members automatic CB-199 rest/FleetApp
Say "unknown" when the member env is unreadable memberHerdrSocket: CB-185 member/HerdrPeerLauncher
One reply settles one ticket automatic CB-137 msg/MessageService
A launch command that cannot fit is refused automatic #220 member/HerdrPeerLauncher
A failed spawn shows you the pane automatic #220 member/HerdrPeerLauncher
Every claude-code member is resumable automatic #214 member/ClaudeCodeLauncher
An exhausted backend is quarantined even from a chrome-only pane automatic #211 inject/CompletionResolver
The credential scrub follows the member's own user memberLoginShell: + worktreeGroup: #213 member/HerdrPeerLauncher
An opencode member's config follows the member's own user memberHerdrSocket: + worktreeGroup: #219 member/OpenCodeLauncher
The model opencode actually ran is read back and checked automatic (opencode profiles) #175 member/OpenCodeSessionDiscovery, member/OpenCodeLauncher
Every file a member must read follows the member's own user memberHerdrSocket: + worktreeGroup: #222 #224 member/ClaudeCodeLauncher, session/GitWorktrees
A role fleetd cannot bind is refused automatic #123 auth/MemberRegistry
A reply says whether anything was waiting for it automatic #365 msg/MessageService.ReplyOutcome

Nearly every knob above lives in one file, on one profile:

flowchart LR
    Y["fleetd.yaml"] --> G["bind / auth / guard"]
    Y --> B["broker"]
    Y --> L["lifecycle"]
    Y --> W["profiles:"]
    W --> P1["profile: gx10"]
    W --> P2["profile: opus"]
    P1 --> K["weight · maxLoad · model<br/>env · gitTokenEnv · parityOverlay<br/>configDir · cwd · argv"]
    P2 --> K
    Y --> F["fleet:"]
    F --> FL["leaders"]
    F --> FA["architects"]
    F --> FD["developers"]
    F --> FR["reviewers"]
    FA --> K
    FD --> K
    FR --> K

The configuration surface has two halves. profiles: answers "which backend" — model, adapter, cost. fleet: answers "who runs, and on which of those backends". A role pool holds profile names, so the arrows meet: the same profile may serve several roles.


Ask the bridge who you are

What. fleet_whoami returns {"role":"primary"} or {"role":"worker", sessionId, profile, worktree, branch}.

On. Always available; no configuration.

Why. Every rule in the bridge charter is role-conditional, and both roles read the same CLAUDE.md — a worker runs in a worktree of the same repo, so it inherits the file verbatim. Before this, a session had to infer its role from side-channels (the mount name, ANTHROPIC_BASE_URL, the system prompt), each of which is one-way and some of which are absent for Claude-model workers. The daemon already resolves the role from the connection for its authorization gate; fleet_whoami just exposes that same answer, so guessing is never necessary.

Gotcha. The answer comes from the connection and cannot be forged or overridden by an argument. If it disagrees with what you expect, the daemon is right and your assumption is wrong — check primary.terminal next.

Primary inside a herdr pane

What. Lets the orchestrating session run inside a herdr pane instead of an outside terminal.

On. primary: terminal: term_<id> — read the id from fleet_whoami, re-pin whenever the primary moves panes.

Why. Caller identity resolves a loopback PID to its herdr pane, and the pane scan covers every pane, not just fleetd-spawned ones. So a primary living in a pane classified itself as a worker and was refused spawn/send/stop — every verb it exists to call. The failure is self-locking: the daemon can also learn the primary's terminal, but only from fleet_send/fleet_spawn, the exact calls being refused. Only an operator-set pin breaks the cycle, which is why the pinned value is consulted and the learned one deliberately is not.

Gotcha. A stale pin is silent. You are simply demoted to worker and every orchestration call is refused. Re-pin after the primary changes panes, and note the daemon reads this at boot — a change needs a restart.

Weighted worker placement

What. Spreads unqualified spawns across profiles by weight, with a concurrency cap per profile and failover to the next candidate when one is unreachable.

On. placement: weighted plus per-profile weight: and maxLoad:. The default fixed policy reproduces the historical always-the-default-profile behaviour.

Why. Profiles differ in model and cost, not tier. Without weights, every unqualified spawn piles onto one backend regardless of what it costs or how loaded it is.

Gotcha. An explicit profile: on fleet_spawn bypasses the policy entirely — placement only governs unqualified spawns. Equal-weight candidates tie-break on YAML definition order, so that order is load-bearing config, not cosmetics (CB-524).

Run a worker on the subscription

What. Lets a claude-code worker run on the operator's Claude subscription, on purpose, when no off-subscription endpoint exists for its family (e.g. sonnet on ccs). The launcher injects neither ANTHROPIC_BASE_URL nor ANTHROPIC_AUTH_TOKEN and skips SubscriptionGuard's base_url requirement for that profile only — every other profile keeps the hard refusal.

On. Add subscription: true to a profile. The default (absent/false) keeps today's hard boundary: a claude-code profile with no base_url may not spawn, because doing so would bill the subscription.

workers:
  sonnet:
    kind: claude-code
    subscription: true
    model: claude-sonnet-5
    argv: ["ccs", "sonnet"]

Why. Some model families (e.g. sonnet on ccs) have no off-subscription endpoint to point a worker at. Rather than leave those profiles unspawnable, this is an explicit, visible opt-in — a spawn under it logs a WARN naming the profile, so billing the subscription is never an accident.

Gotcha. subscription: true contradicts a baseUrl (the two state opposite intents) and is refused at spawn if both are set. The same contradiction is refused at config load for the env: map: on the subscription path the guard is skipped and the adapter writes neither Anthropic key, so an ANTHROPIC_BASE_URL/ANTHROPIC_AUTH_TOKEN in env: would reach the worker having passed no guard at all (CB-542). A subscription profile may still carry env: — just not those two keys. weight: 0 cannot reserve this profile from unqualified weighted placement: worker-config normalization changes non-positive weights to 1.0. Use an explicit profile: sonnet spawn until the placement policy gains an exclusion setting.

Advisory architect slots

What. Declares named, strong-model advisory slots. A bound architect resolves as architect, can send, reply, ask, and read, but cannot spawn, stop, or drain the fleet. The operating shape is one human-driven lead, two short-lived architects, and N workers: the lead consults the architects sideways and discards them, rather than creating a supervisor above the lead. Like every spawned member, an architect becomes deliverable once it mounts the bridge MCP — until then a brief sent to it is held at the readiness gate and never typed into its pane (CB-560).

On. Declare each slot against a configured worker profile:

architects:
  sonnet-adviser:
    profile: sonnet
  gpt-adviser:
    profile: gpt

Why. The rejected alternative put an orchestrator above the lead and made the lead a managed, resumable session. That breaks the actual control boundary: the human drives a lead that pre-exists, is recognised, and cannot be resumed by fleetd. Two advisory model families, Claude Sonnet 5 and GPT-5.6 through opencode, receive the same brief independently so agreement is evidence rather than correlated echo.

Gotcha. architects: declares slots, not sessions: no terminal is configured and nothing becomes an architect until the lifecycle binds a live terminal to a slot. The profile name is validated at boot, as are duplicate slot names. An architect has delegation authority but no lifecycle authority, by design.

Shipping the binding is not the same as shipping the role. For one day an architect bound correctly, resolved as architect, and could not receive a single message: the presence map that opens the readiness gate was keyed on Role.WORKER, so an architect was never marked available. It built clean and passed two reviewers. When a change adds a role, check every place that assumes a member is a worker — fleet_status tells you whether a member ever became ready.

Keep a worker conversation

What. Carries a bridge logical session name and a peer-owned resume id through SpawnRequest, then returns them from the peer handle where the adapter supports them. Claude Code mints a UUID for a fresh named session, passes it as --session-id, exposes the logical name with -n, and resumes with -r. OpenCode resumes with -s and discovers its id after launch from its on-disk session record by matching the worker's unique worktree cwd.

On. Supply a session name and/or resume id when the spawn lifecycle has one to carry. Adapter capabilities state the asymmetry: Claude Code offers SESSION_NAME and SESSION_RESUME; OpenCode offers SESSION_RESUME only.

Why. A bridge session name is an operator-facing roster label, while a provider session id is the only handle that can resume the actual conversation. Treating them as peer-neutral values prevents the core from assuming Claude Code's flags are a universal protocol.

Gotcha. OpenCode has no name flag, and it writes its session record only after it persists a conversation. Its id is therefore discovered lazily and may be absent immediately after spawn; its slot name remains only in fleetd's roster.

Give workers a toolchain

What. Propagates the daemon's own PATH to every worker, plus a literal per-profile env: map.

On. Automatic for PATH; add env: {JAVA_HOME: ..., ...} on a profile to extend or override.

Why. A worker's environment does not come from your shell. fleetd hands herdr an explicit env map and herdr merges it into its own process env — so before this, a worker inherited whatever PATH the herdr server happened to be started with. On a long-lived herdr that can predate your toolchain entirely, leaving workers unable to run mvn or java at all.

Gotcha. On an off-subscription profile, adapter-owned variables win over env: — the ANTHROPIC_*/CLAUDE_* wiring is applied after it, so an env: entry cannot repoint a worker past the SubscriptionGuard. A subscription: true profile is the exception: it skips the guard and injects neither key, so an ANTHROPIC_BASE_URL/ANTHROPIC_AUTH_TOKEN in its env: would survive unguarded — that configuration is refused at load (see Run a worker on the subscription). Since the default PATH is the daemon's own, start the daemon with a good one (see the PATH lines in deploy/dev.ltms.fleetd.plist and deploy/fleetd.service).

Isolated worktree per worker

What. Provisions a git worktree on its own branch per worker, and copies a configurable set of local config files in ("parity overlay") so the worker sees the same local setup.

On. fleet_spawn{worktree: true, ticket: "cb-123"}; overlay list via parityOverlay: (default [.claude/settings.local.json, .env, .envrc]).

Why. Parallel workers editing one checkout collide. A worktree gives each its own branch and files at the cost of a checkout.

Gotcha. Release removes the checkout but keeps the branch — unmerged work survives a teardown. Overlay-copied files that are tracked get --skip-worktree so they never read as pending changes. Never add .mcp.json to the overlay; see the next entry for why.

A failed provision cleans up after itself (fleetd #274). Provisioning is more than git worktree add: several steps run after the worktree and branch exist, including the credential-free-origin check, which is an intended refusal rather than an IO accident. If any of them fails, the worktree and its branch are now removed before the error is rethrown. Before #274 they leaked permanently — the caller never received the path, so its own cleanup could not fire, and nothing else tracked the directory. Note the asymmetry with the paragraph above: a released session keeps its branch, because a worker's branch is meant to outlive its worktree; a provision that never completed has no session and no PR, so its branch goes too.

The isolation does not cover git stash. A worktree separates the working tree, the index and HEAD — but refs/stash is a single stack shared by the primary's checkout and every worker worktree of the repo. Two parallel workers stashing at the same time can pop each other's work, silently. This happened on 2026-09-04 between the #273 and #274 workers; both noticed and recovered. The implementer skill now forbids git stash and points at a wip: commit or a patch file.

Worker tool-surface isolation

What. A provisioned worktree's project .mcp.json is neutralized to an empty server map, so a worker's tools are only what its launcher mounts (the bridge).

On. Automatic at provisioning. Nothing to configure.

Why. The repo commits a .mcp.json declaring the primary's IDE servers, so a fresh checkout mounted them and the parity overlay copied the primary's own copy on top. Those servers are bound to the primary's IDE project, so every path they return points into the primary's checkout. This is not hypothetical: a worker made all 59 of its edits in the primary's tree while running mvn against its worktree — so every build it ran was of code that did not contain its changes, and every build passed.

Gotcha. This is why a worker cannot run IDE diagnostics and mvn is its only verification. That is deliberate — the bridge is a message bus, and the primary is the gate. Never accept a worker's claim about a check it had no way to run.

Worktree-hostile config isolation

What. Every provisioned worktree neutralizes tracked .mcp.json, opencode.json, and .autoenv with a valid format-specific stub, then marks a tracked replacement --skip-worktree.

On. Automatic at worktree provisioning; no configuration.

Why. The tracked opencode.json mounts the primary's own gitea and context7 servers, with the primary's credentials, so a worktree that keeps it hands a member access it must never hold. The same isolation rule also keeps a primary-only MCP configuration or an autoenv authorization prompt out of a worker's tool surface.

The reason used to be milder, and the change is worth recording. opencode.json referenced gitignored .secrets/ files that a worktree never contains, and OpenCode refuses to start on that dangling reference — a crash, but a loud one. Those credentials now live in one shell-level store and the file reads them as {env:…}, so in a worktree the reference resolves instead of failing. A loud crash became a quiet privilege leak, which makes this isolation load-bearing rather than a workaround.

Gotcha. The replacement is deliberately valid, not deleted: {} for opencode.json, an empty server map for .mcp.json, and an empty .autoenv. A deletion could be undone by a later checkout; the skip-worktree bit keeps the safe local replacement from looking like work for a worker to commit.

No credentialed remote URL reaches a worktree

What. Before provisioning a member's worktree, fleetd makes sure the repository's remote URLs carry no embedded credentials, and says so when they do.

Three things happen, in this order:

  1. Report. Every remote is enumerated, and both its fetch and its push URL are checked for user-info on any non-SSH scheme. Anything found is logged as a WARN naming the remote.
  2. Strip. User-info is removed from origin's HTTPS URL.
  3. Refuse. If the provisioned worktree still resolves a credentialed HTTPS origin, the provision fails rather than handing the member a URL with a secret in it.
WARN member worktree shares a remote URL containing user-info: remote=upstream
     repository=/Users/…/fleetd; remove credentials from the repository's git config

On. Always on. There is no key.

Why. A linked worktree shares its parent repository's git config. A token embedded in a remote URL is therefore readable by the member the moment its worktree exists — and git remote -v prints it in full, so it leaks again the first time anyone surveys the host. Members are meant to push with the repo-scoped WORKER_GITEA_TOKEN injected at spawn, never with a credential baked into config.

Gotcha — the report and the refusal cover different ground, on purpose. Only origin, and only HTTPS, is stripped and refused. Every other remote, every pushurl, and every other non-SSH scheme is reported and then left alone. Rewriting a remote nobody asked us to touch is not fleetd's call; telling you it is there is. So a WARN here is work for you to do, not something the daemon has already handled.

SSH URLs are deliberately exempt. There the user part selects an account and authentication happens over the SSH transport, so git@host is not a credential the way user:token@host is.

The reason it warns before it strips. For the origin HTTPS case you will see the WARN immediately followed by the strip's own INFO line. That ordering is intentional: it leaves an audit trail that there was something to fix. Reporting after the strip would silence the one case the daemon actually repairs.

Gotcha — a reporting check must never break a provision. An earlier attempt at this ran its git calls unguarded at the top of add(), where any non-zero exit or the 30-second timeout would have aborted the whole worktree provision. Every call here is wrapped, and a failure logs only the exception's class, never its message — the message is read from the very config that may hold the URL being looked for.

For the same reason, every git command whose stdout is a URL runs through a redacting variant that keeps captured output out of exception messages. Without it a failing git remote get-url copies the credentialed URL into the exception, and from there into the log — defeating the check by way of its own error path.


Worker opens its own PR

What. Injects a repo-scoped forge token so a worker can commit, push over SSH, and open its own pull request at checkpoint.

On. gitTokenEnv: (host env var holding the token) and optionally gitHostEnv: on a profile.

Why. Opt-in by design: omit it and the worker gets no PR-create grant, while push over SSH still works.

Gotcha. Use a minimal write:repository token, never an admin one. The token can create a PR but must not be able to merge — the primary is the gate, and a worker that can merge is not gated.

Where the value comes from. gitTokenEnv: names a variable, and fleetd reads it from its own process environment — so the token must be exported in the shell that launches the daemon, not stored in the repo. Keep it beside the operator's other credentials in one sourced file; a per-repo copy is a second copy of the same secret, and the copy you forget is the one that leaks or goes stale. If the daemon was started before that export existed, it injects an empty token and every worker push fails: restart it from a login shell.

Session lifecycle caps

What. Reaps idle sessions, caps turns per session, optionally clears a reused Claude Code worker's conversation after each delegation, and drains cleanly on shutdown.

On. lifecycle: { idleTtlSeconds, contextCap, drainTimeoutSeconds, clearAfterTurn }.

Why. Workers are disposable but not free; without caps an abandoned session holds a pane and a context indefinitely. clearAfterTurn: true keeps the pane and process warm while preventing delegation N+1 from inheriting delegation N's conversation.

Gotcha. clearAfterTurn defaults to false. It uses Claude Code's /clear command without counting that housekeeping as a delegated turn, and waits for it to settle before delivering the next task. Peer kinds without a known context-reset operation (currently opencode) treat the knob as a no-op and log that once; fleetd never guesses a command. Lifecycle config is read at boot, so changes need a daemon restart.

Durable reply inbox

What. A worker's reply survives with no waiter attached: it is queued and collected later by fleet_poll / fleet_ack, over a real AMQP broker when one is configured.

On. broker: pointing at an AMQP URI; omit it for an in-memory inbox.

Why. A blocking fleet_send is capped by the caller's MCP client timeout (~60s), far below a real task's runtime. Without a durable inbox, a reply arriving after that window lands nowhere.

Gotcha. The broker is LavinMQ, not RabbitMQ. Point it at a stray local RabbitMQ and you are writing into someone else's broker. Also see the known hole: an async ticket that times out while the session is still BUSY currently discards the later completion rather than parking it.

Draining an inbox is lead-only, on both entry paths

What. fleet_poll{target} empties that session's reply inbox — the replies are removed and a second call returns nothing. It is therefore gated as a drain, exactly like fleet_ack, and only the primary may call it. fleet_poll{ticket} is unaffected: it observes an async delegation and changes nothing, so a worker or an architect may still poll a ticket it owns.

On. Always on, whenever authorization is enforced (an auth: block / a real CallerResolver). Nothing to configure.

Why. Until fleetd #272 the MCP handler passed a constant READ for both branches, and READ is open to every authenticated role. Any worker could read a peer's sessionId out of fleet_list and destroy the replies that peer had queued for the primary — an unrecoverable loss, since a drained reply is gone. The REST path had always checked DRAIN, and the wiki had always described the tool as lead-only; the MCP gate was the one that disagreed.

Gotcha. An architect cannot drain either, even though it may fleet_send. That is deliberate and matches fleet_ack: delegating is not a lifecycle right. If an architect delegates with wait:false it still polls its own ticket, which is the branch that stayed open.

The general rule this came from is worth keeping: when one tool name covers two operations, the gate belongs inside the branch that picks between them, not above it. fleet_poll was the only handler in the server with that shape — the other ten were audited and are correct.

Broker password out of the config

What. broker.uriEnv: names an environment variable that holds the AMQP URI, instead of writing the URI into the config file. An AMQP URI carries user:password@ inline, so the old broker.uri: put a live password in clear text in fleetd.yaml.

On. broker: { uriEnv: LAVINMQ_URI }. The variable must be on the daemon's own environment, so start the daemon from a login shell — scripts/redeploy-fleetd.sh --check reports whether the named variable resolves, without ever printing its value.

Why. The same reason auth.tokenEnv and Profile.tokenEnv exist: a config file gets read, copied and pasted into tickets far more often than a secret store does. This was the last credential still living in clear text in the config.

Gotcha. uriEnv wins over uri whenever it is set, and it does not fall back. If the named variable is unset or blank the daemon treats the broker as unconfigured and uses the in-memory inbox — it does not quietly drop back onto a stale uri: left in the file. That is deliberate: an operator who moved to the secret store must never be silently returned to clear text. Both set ⇒ the log says broker.uri is ignored.

An unreachable broker does not stop the daemon

What. If the broker cannot be reached at boot, fleetd logs a loud warning and starts on the in-memory reply inbox for that process lifetime, instead of failing to start at all.

On. Always on. There is no background retry: fix the broker and restart to get durability back.

Why. AmqpReplyInbox.open throws IllegalStateException, and nothing caught it — so a broker that was down took the whole daemon with it, and with the daemon went every lead, every member and every pane. Losing durable replies is bad; losing the fleet because the reply store is down is far worse. A broker that drops after startup already self-heals through the AMQP client's automatic recovery, so only the boot path needed this.

Gotcha. The fallback is not silent, but it is easy to miss in a busy startup log. What you lose is real: replies become soft-state and a held report does not survive the next restart. Grep the startup log for reply inbox: — it says which adapter won, every time. The warning names the failing URI with the credentials stripped.

Reject overlapping rendezvous

What. A second attempt to open a reply waiter for the same worker session fails atomically instead of replacing the first waiter.

On. Always on; no configuration.

Why. One session has one outstanding delegated turn. Replacing its waiter silently would strand the first caller and let a reply resolve the wrong request. MessageService serializes normal sends, but the atomic rejection is the tripwire that makes a future violation loud rather than corrupt.

Gotcha. This is not concurrent-turn support. A terminal send closes its own waiter before the next turn can open one; a double-open is an invariant failure that must be investigated.

Pin an opencode endpoint

What. An opencode profile can target its own OpenAI-compatible endpoint.

On. baseUrl: on a kind: opencode profile.

Why. opencode is provider-agnostic and shares none of Claude's private seams — no ANTHROPIC_BASE_URL, no SubscriptionGuard, no --mcp-config. It is the adapter that proves the PeerLauncher SPI is genuinely provider-neutral rather than Claude-shaped.

Gotcha. Because it bypasses SubscriptionGuard, the guard's allowlist does not protect this path — the endpoint you name is the endpoint it uses.

Onboard a project with the plugin

What. A Claude Code plugin that makes any project fleet-ready: it mounts the fleetd MCP gateway and ships a /fleet:setup skill that runs preflight, applies standard project settings, names the environment variables the operator must export, and verifies the session resolves as the primary.

On. /plugin marketplace add https://git.ltms.dev/fleet/fleetd then /plugin install fleet@fleetd; export FLEETD_MCP_URL (usually http://127.0.0.1:8765/mcp); run /fleet:setup inside the project to onboard. Develop it locally with claude --plugin-dir ./plugin.

Why. The orchestration contract had no distributable form. Every consuming project had to hand-copy a block of CLAUDE.md and hand-write an .mcp.json, and we maintained a script purely to detect the copies drifting apart. A plugin is versioned, installed once, and updates in place — the contract stops being something each project re-derives. It ships no credentials by design, so the artifact is public-safe: every secret is referenced by environment-variable name and the value never enters a file.

Gotcha. The plugin is client-side setup only — it mounts a daemon, it does not install one. fleetd and herdr remain separate services, and the setup skill deliberately refuses to install them (guessing at a system-service install is how you get two daemons on one socket). Note also the plugin root is plugin/, not the repo root: an installed plugin's .mcp.json is a committed file, while this repo's root .mcp.json is local-only and --skip-worktree, so rooting the plugin at the repo would collide with the very isolation CB-525 exists to enforce.


Renamed in 0.2.0 (fleetd #362) — this is a breaking change. The plugin was claude-bridge and mounted its server as fleetd; it is now fleet@fleetd and mounts fleet. The old name gave a lead with both a project .mcp.json and the plugin two mounts of one daemon and a duplicated fleet_* tool set, and it did not match PeerLauncher.MCP_MOUNT_NAME. A project that pre-allowed mcp__fleetd__fleet_whoami in .claude/settings.json must be updated to mcp__fleet__*. The .mcp.json URL is now ${FLEETD_MCP_URL} rather than a hardcoded address, so one plugin can serve hosts running the daemon on different ports — the variable is required, and the setup skill checks for it in preflight.

The plugin is lead-side only, and cannot be otherwise. It carries the mount, the setup skill and the charter — never the worker playbook skills or the role agent definitions. Two structural reasons, both measured: the launcher adds --agent only when <worktree>/.claude/agents/<role>.md exists in the member's own tree (ClaudeCodeLauncher.java:371,391); and a member's CLAUDE_CONFIG_DIR points at its profile's config directory (ClaudeCodeLauncher.java:285), so it never reads the operator's plugin store. On this Mac every Claude profile sets configDir, and the four ccs instances hold four separate copies of the plugin store — same md5, different inodes — so a user-scope install lands in exactly one of them. Member-facing assets travel in the worktree.

This entry existed and still did not prevent a rebuild. In September 2026 a session planned the whole plugin from scratch, because wiki/ is a submodule whose pointer is never advanced and no session reads it by default. Documenting a feature here is necessary and not sufficient — a capability an agent must not re-derive needs a line in CLAUDE.md, which is the only file every session loads.

Port a workspace to OpenCode

What. A port-to-opencode skill that makes an OpenCode session a first-class participant in a Claude Code workspace — same instructions, same MCP servers, same bridge mount — by writing a single opencode.json and nothing else.

On. Load the port-to-opencode skill in the primary. It is a primary-side skill, not a delegation playbook: a worker mounts only the bridge MCP and cannot run it.

Why. OpenCode reads CLAUDE.md natively — including ~/.claude/CLAUDE.md — so the rules cross for free and only MCP servers need mapping. That is worth writing down because the obvious move is the wrong one: the Codex attempt translated CLAUDE.md into a second rules file with a third-party tool, and the translation silently corrupted a "never commit" rule into one naming a path that cannot be committed at all. The skill exists to stop anyone reaching for a porting tool again. Secrets cross by {env:VAR} reference, which is what makes opencode.json committable — and it must be, because a peer in a worktree receives tracked files only.

Gotcha. An AGENTS.md left in the repo shadows CLAUDE.md — opencode takes the first match walking up from the cwd, so a stale file from an earlier port silently wins over the live rules. Delete it before anything else. And skills do not cross: .claude/skills/** is not read by opencode, so a brief telling a peer to "load the implementer skill" is a no-op there — spell the procedure out in the brief instead.

More than one lead

What. Recognises several panes as leads, so two orchestrators — say a Claude lead and an opencode lead — work as peers instead of one being demoted.

On.

leaders:
  opus-5.0:
    terminal: term_0123456789abcd
    kind: claude
  gpt-sol-5.6:
    terminal: term_fedcba9876543
    kind: opencode
    model: openai/gpt-5.6-terra

terminal is the only field identity depends on; kind/model are descriptive and are echoed back by fleet_whoami as leader: <name>. role still reads primary — a lead is a primary for authorization, so nothing keying on the role breaks.

Why. primary.terminal is singular by construction: one pane is the lead and every other pane resolving to a herdr terminal is a worker. That is right while one lead drives a fleet, and wrong the moment two leads collaborate — the second is silently demoted and refused every orchestration call it makes. Resolution is now a registry lookup rather than an equality test against one pin.

Gotcha. This entry originally said to keep primary: alongside leaders:, because the CB-307 push loop needed exactly one nudge destination while leaders: only widened who was recognised. That is no longer true: reply nudges now follow the delegating lead, and primary: is retired. Delete it. If both are present and name the same terminal the leaders: entry wins. And a lead is never spawned — it pre-exists, which is why it must be named here rather than created; argv/placement are worker-profile keys and mean nothing in this block. Read at boot, so a change needs a restart.

Find leads by tab name

What. Discovers leads by scanning herdr for tabs you labelled, instead of you pasting each lead's terminal_id into leaders:. Label a tab lead: gpt-sol-5.6, start an agent in it, and that pane resolves as a lead named gpt-sol-5.6 within one rescan — no config edit, no daemon restart.

On.

leadScan:
  tabPrefix: "lead:"     # matched case-insensitively; the rest of the label is the lead's name
  intervalSeconds: 10    # rescan cadence, and the worst case before a new tab is recognised

Opt-in: no block means leads come only from leaders:/primary:, exactly as before. Both sources merge, and an explicit leaders: entry outranks a label for the same terminal.

Why. A lead is never spawned — a human opens a tab and starts an agent in it — so unlike a worker, the daemon cannot learn its terminal_id at creation. leaders: therefore costs a four-step ritual per lead: start the session, ask it fleet_whoami for its id, edit config, restart. Naming the tab is one step, taken at the moment the operator is already there. The label also survives what the id does not: close and reopen the tab and the terminal_id changes, while the label is retyped as-is.

Gotcha. The direction of trust is what makes this safe, and it is one-way: fleetd reads lead tab labels and never writes them, so what is in the tab bar is always what a human typed. Two guards keep that from eroding — the configured worker spaces (where fleetd does write labels, via tabLabel) are excluded from the scan wholesale, so nothing the bridge places can land in a matching tab; and startup refuses a tabPrefix that any worker tabLabel also matches, because overlapping those two namespaces would have the daemon label its own workers as leads and promote the entire fleet. Every pane in a labelled tab is that lead, so split a lead tab only with panes you mean to be leads. A failed scan keeps the leads already known rather than emptying the registry — a herdr hiccup must not demote a live lead mid-session.

Leads talk to each other

What. A lead can message another lead and be answered. fleet_send{sessionId: <peer's terminal>} reaches a peer, and the peer closes the exchange with fleet_reply — the same rendezvous a worker uses. fleet_whoami now reports a lead's own sessionId, which is how a lead learns the address to give a peer.

On. Nothing to configure; it applies as soon as two panes resolve as leads (via leaders: or leadScan:).

Why. CB-530 and CB-531 widened recognition — both leads are seen — but nothing had widened addressing, so collaboration was one-way and silently so: the send was accepted, the peer's fleet_reply was refused, and the sender waited out its timeout. The cause was that a lead's Principal carried no terminal, so ownsSession() could never be true for it and the REPLY/ASK rules excluded it by construction. A lead now carries the pane it was matched by, and the rule it must satisfy is unchanged: you may act as the pane you occupy, and as no other. That was always the real control — terminal.equals(sessionId), against a terminal that comes from the connection — and the extra "…and you must be a worker" conjunct beside it protected nothing.

Gotcha. A lead replies only to answer a peer that messaged it — never to answer a worker, whose turn it is not, and never as a way to end its own turn. Widening who may reply did not widen what they may reply as: a lead still cannot act for another pane, and an unnamed primary (token mode, or off-host, with no pane at all) owns nothing and remains a sender only.

Leads are visible in fleet_list

What. fleet_list returns leads alongside workers. Each lead row carries its sessionId (the address to fleet_send to), its name, its live status, and self: true on the caller's own row.

On. Automatic, wherever a pane resolves as a lead.

Why. A lead had no way to discover a peer. fleet_list enumerated the worker roster alone, so a lead asking "who else is here?" got an empty array — which reads as no peers but only ever meant no workers spawned. The peer lead on this bridge drew exactly that wrong conclusion and reported itself alone in a two-lead fleet. Addresses had to be carried between panes by a human, which is not a protocol. Reporting both halves — even when a half is empty — also removes the ambiguity that caused the misreading.

Gotcha. The lead rows come from the same registry that resolves identity (CallerResolver.leads()), not from a second copy, so a listed address is one that would actually resolve as a lead. A lead herdr is not tracking as an agent is listed with status: unknown rather than hidden — it cannot be delivered to, and a would-be sender needs to see that rather than infer it from silence.

A lead can be delivered to

What. The injector's readiness gate opens for a lead as well as for a booted worker, so a message addressed to a lead is actually typed into its pane.

On. Automatic, wherever a pane resolves as a lead.

Why. The gate (CB-113) holds a delivery out of a spawned worker's boot window: herdr reports idle while the agent is still starting, and a paste into that window is lost. Membership in it comes from MemberPresence, which BridgeMcp populates for every spawned member — worker or architect, but never a lead, since that map doubles as the roster's availability signal and a lead counted there would appear as an available member. The two rules composed into a dead end: a lead is never marked present, so the gate never opened for one, so lead-to-lead messaging — shipped and authorized in CB-532 — still could not deliver a single keystroke. The gate's premise simply does not apply to a lead: a lead is never spawned, so it has no boot window to guard.

Gotcha. The failure this fixes was slow and mute, which is worth recognising if it recurs in another form: the send was accepted, the pane stayed idle, nothing was ever typed, and ~60s later (READINESS_GRACE_POLLS, 240 × 250ms) it failed through the same path as a stalled turn — so the log said turn-stall fallback while the truth was that delivery had never been attempted. A repeating 62-second gap between send and failure is the signature of the gate, not of a peer that ignored you.

The gate is no longer mute (CB-562). When the grace expires it logs a WARN naming the target, the poll count, the grace in seconds and how many queued messages it is failing because the target never became deliverable — so a readiness failure now reads differently from a turn stall. The seconds are derived from Injector.POLL_INTERVAL_MILLIS, the single source the poller is also built from, so a cadence change cannot leave the log confidently stating a wrong duration. It was also intermittently masked: presence is a sticky set, so a pane that was seen as a worker before being recognised as a lead stayed deliverable until the next restart cleared the set.

Reply nudges follow the delegating lead

What. When a worker's reply lands with no fleet_send open, the CB-307 nudge goes to the lead that delegated that worker — not to a globally-configured "the primary". This is what retires primary.terminal:, which now logs a deprecation warning at startup.

On. Automatic. Delete primary: from fleetd.yaml; keep it only if you still want pushReminders/pushBackoffMs, or a fallback nudge destination across restarts.

Why. PrimaryRegistry held one slot answering "who is the primary" — a question with no correct answer once two leads drive one fleet. Whichever lead called fleet_send first captured every nudge thereafter, so the other lead's results were announced to the wrong pane. The binding that actually matters is per-delegation and is known exactly where it is created: at fleet_send, where the target is the argument and the lead is resolved from the connection.

Gotcha. A restart loses the delegation map while the durable inbox keeps the reply. With one lead the old pin (or the first lead to send) is an unambiguous fallback; with several and no recorded delegation the daemon nudges nobody rather than guessing, and delivery degrades to fleet_poll. That is the correct degradation — interrupting the wrong lead with someone else's result is worse than a quiet inbox — but it does mean a post-restart reply may need an explicit poll.

Second gotcha, fixed in fleetd #368. A binding was only ever dropped when the worker was released. The map is keyed by the worker, so a lead that was closed, crashed or relaunched left its bindings behind, pointing at a pane that no longer exists. A stale entry is not null, so it beat the single-lead fallback above every time, and the nudge was sent into nothing. The push loop now checks the recorded lead with agents.status before trusting it, and forgets a lead that is really gone so resolution reaches the fallback. Only agent_not_found counts as gone: any other failure — a socket blip, a decode error — is treated as live, because a wrong guess there unbinds a working lead permanently, while a wrong guess the other way costs one retry on the next tick.

Unknown config keys are named

What. A top-level key in fleetd.yaml that this build does not understand is logged as a WARN naming it, at load.

On. Always on; nothing to configure.

Why. Every config record is @JsonIgnoreProperties(ignoreUnknown = true) — deliberate, so config may run ahead of the code and a rolled-back daemon still starts. The cost is that a whole block can be written, parsed, dropped, and never mentioned again. That is exactly how a hand-written leaders: registry came to look configured while being inert: the daemon started, nothing complained, and the only way to find out was reading the config class. A dropped block and a working one were indistinguishable.

Gotcha. A warning, not a failure — deliberately. Failing closed would turn "the config names something this build has not learned yet" into a daemon that will not boot, destroying the forward-compatibility the annotation exists for. Nested unknown keys are still silent; only the top level is checked.

One fleet block, role as the key

What. fleet: replaces four top-level keys — leaders:, members:, leadScan: and defaultProfile:. A member's role is now the map key that contains it, not a role: field inside it.

On. Write a fleet: block. The four old keys are hard errors that name what to use instead, so an old config does not start silently changed.

fleet:
  leaders:
    opus: {profile: opus, instances: 1, tabPrefix: "lead:"}
  architects:
    opus:   {profile: opus}
  developers:
    local:  {profile: local}
    sonnet: {profile: sonnet}
  reviewers:
    sonnet: {profile: sonnet}

Why. A misspelled role: architct used to parse into a member with a profile, a name, and no contract at all — nothing rejected it, because role: was just a string. A misspelled pool name declares nothing, which is a shape the loader can see. The four keys also had no relationship to each other on the page, while all four describe one thing: who is in the fleet.

Gotcha. defaultProfile: has no single successor key, so its error message explains the new model rather than pointing at a key that does not exist. Role pools took over its job — see below.


Role pools decide the backend

What. fleet.architects / developers / reviewers list the profiles that role may run on. A spawn that names no profile is placed inside the pool of the role it asked for, instead of across every configured profile.

On. List profile names under the role. Order matters under placement: fixed — the first entry wins.

Why. Role and profile are separate axes, and collapsing them loses real cases: a reviewer may run on the very same profile as the dev whose diff it reads. Before this, an unqualified spawn ranged over all profiles, so a reviewer could land on the architect-only backend and quietly spend the subscription. Pools also replaced the single global defaultProfile:, which could only ever have one answer for a fleet that has three kinds of member.

Gotcha. A role with no pool is unconstrained, not blocked — it falls back to every profile, so a config that pools some roles and not others keeps working. And an explicit profile is not confined to the pool: fleet_spawn{profile:"opus"} carries no role, so it defaults to dev, and judging it against the dev pool would refuse a spawn the operator asked for by name. maxLoad still applies to it.


Tab labels name the role

What. A member's tab reads dev: sonnet #4 — role first, then backend, then a counter.

On. fleet.tabLabel: (default "{role}: {profile} #{n}"). {role}, {profile}, {model} and {n} are substituted. A profile may override it with its own tabLabel:.

Why. The label lives on fleet: because a profile cannot know the role of the member launched on it, and the role is the thing an operator scanning a tab bar actually wants. {n} counts per role and profile, so a dev and a reviewer on one profile each start at #1 — a single fleet-wide counter would make the number meaningless. Putting {role} first also turns the lead/member namespace check into a structural guarantee: roles are a closed enum, so only a hand-written template can still collide with a lead's tabPrefix.

Gotcha. The knob was inert on first release — HerdrPeerLauncher accepted the template and nothing passed it, so the label only looked right because the fallback happened to match the default. Fixed in CB-557; tests now pin the wiring rather than the coincidence.


Launch a lead when none is live

What. The daemon starts a lead declared under fleet.leaders: when fewer than instances are running. It labels the tab by the same convention the scanner reads.

On. Give the lead a profile:. Omit it and the lead stays recognise-only, exactly as before. instances: 0 is an off switch. workspace: (default leads) and cwd: control where it lands.

Why. A lead was the one pane a human had to open by hand before anything else worked, so a daemon restart after a reboot left a fleet with no orchestrator and no sign of why.

Gotcha — three, and they are the whole design. An auto-launched lead is not a member: it gets no worker reply charter (that text tells its reader it is an off-subscription worker who must end every turn with fleet_reply — the opposite of an orchestrator), it is never registered with SessionManager (the idle reaper would kill it for being idle, which is a lead's normal state), and ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.

A lead counts as live only when herdr reports a running agent — either in a tab labelled lead: <name>, or on a pinned terminal:. Both are needed: label-only would relaunch a hand-opened pinned lead on every boot, and pin-only would miss one the daemon started itself. A labelled tab with nothing running in it is not a lead, so one crash does not disable auto-launch forever. If herdr cannot be reached the daemon starts nothing — a second orchestrator is worse than none.

workspace: must not name a member workspace: those are excluded from the lead scan, so a lead placed in one would never be found again and would be relaunched on every boot.


Re-read the config without a restart

What. fleetd watches fleetd.yaml's modified time and re-reads the file when it changes. Consumers read the live config at the point of use, so a change reaches the next spawn without rebuilding anything.

On. Add configReload: {enabled: true}. intervalSeconds: sets the poll period (default 10). Absent the block nothing is constructed, so an upgraded daemon behaves exactly as before.

Why. Tuning a fleet meant restarting the daemon, and a restart tears down every lead and worker it owns. Changing one pool's weight cost the whole fleet's state, so in practice nobody changed it.

Gotcha — three classes of key, and the difference is what already exists at reload time.

Class Keys What a reload does
Hot fleet: (pools + charters + tabLabel), placement:, an existing profile's weight / maxLoad takes effect on the next spawn
Deferred lifecycle:, leadHeartbeat:, guard:, worktreeRoot:, spawnReadyTimeoutMs / spawnReadyPollMs, adding or removing a profile, and an existing profile's launch settings (model, baseUrl, argv, env, configDir, mcpUrl, tabLabel, exhaustedPattern) accepted into the new config, but the startup wiring keeps the old value; logged by name
Cold bind:, herdrSocket:, broker:, auth: refuses the whole reload

A changed cold key refuses everything, not just itself. Applying the hot half and warning about the cold half would leave the daemon in a state matching no file on disk — the worst outcome for an operator who is reading the file to work out what the daemon is doing. Refusing keeps one invariant: the live config is always some version of the file.

What decides the class is who reads the key and when, not how important it is. weight and maxLoad are hot because the placement policy reads them live through a supplier. A profile's model looks like it should behave the same way and does not: HerdrPeerLauncher takes Map.copyOf(profiles) at construction and resolves every spawn out of that copy, so a reloaded model never reaches a launch. Adding a profile is deferred for the same underlying reason — a new backend needs its own launcher, and launchers are built once.

That distinction is worth stating plainly because getting it wrong is invisible. A reload that reported a changed model as applied would log a clean "config reloaded" while every spawn kept using the old one, and an operator would have no reason to doubt it. So the reload compares each surviving profile's launch settings and names the profile in the deferred list instead.

A reload that fails to parse, or fails any of the startup validators, is refused the same way and the running config stays live. A config file being saved is sometimes read mid-write, and degrading a working daemon over a half-written file would be a bad trade. The watcher stamps the modified time before it reloads, so a refused file is not retried every tick — the next save earns a fresh attempt.


Set a role's launch charter in config

What. fleet.charters: holds the launch text for each member role, keyed by the singular role wire name — architect, dev, reviewer. An operator edits the text and the next spawn uses it. No rebuild, no daemon restart.

On. Add the block under fleet:. Every key is optional, so a config that never mentions charters behaves exactly as before.

fleet:
  charters:
    architect: |-
      You refine work before anyone builds it: scope, acceptance criteria,
      risks, and a unit split. You never commit production code.

Why. The only per-role launch text used to be REPLY_CHARTER, a static final String in each launcher. Changing what a role is told meant editing Java, rebuilding, and restarting the daemon — which tears down every lead and member it owns. So in practice nobody changed it, and the text drifted away from the truth: it still told every member "You are an off-subscription worker" long after the daemon started resolving architects as well.

Gotcha — the type is a map on purpose, and blank is not the same as absent.

FleetdConfig.Fleet is @JsonIgnoreProperties(ignoreUnknown = true). A typed record field named architetc would be dropped in silence, so the operator would see a clean "config reloaded" and no charter. A Map<String,String> keeps every key the operator wrote, which lets validateCharters() see the bad one and refuse the whole config, naming the key and listing the valid ones.

An absent key is fine — that is every deployment before CB-566, and it means "this role has no configured charter". A blank key is refused. The operator typed the key and expected something; treating it as absent would quietly strip that role of its contract.

Validation runs on both the startup path and the reload path. Wiring only one of the two is the whole bug: a config that a running daemon refuses but a restart accepts, or the reverse.

The charter's text is checked against the live tool surface too (fleetd #469). Validation used to look only at the shape of the map: is the key a role name, is the text non-blank. It never read what the text said. So a charter telling a member to call bridge_send — a tool name the CB-634 rename removed — started the daemon cleanly, and the member found out at run time by calling something that was not there. Now CharterToolSurface pulls every fleet_* and bridge_* token out of each charter and asks FleetTool, the one enum the MCP server derives its registrations from, whether that tool exists. If one does not, the daemon refuses to start and the message names both the charter key and the unknown tool.

A reload is checked too, not only startup (fleetd #474). For a short while this check had one call site, in Fleetd.main, so editing fleet.charters: on a running daemon to name a tool that does not exist was accepted, applied and delivered to the next member — while a restart would have refused the very same file. Charters are hot and are read live at each spawn, so that gap was reachable. ConfigRef.reload() now runs the same check, and a reload that fails it is refused whole and keeps the running config, exactly like any other validation failure. The daemon logs config reload from <path> refused, keeping the running config: <message>, and the message names both the charter key and the unknown tool.

Why it was built this way: the check cannot live inside FleetConfig, because the canonical tool set is in the mcp package and config is loaded before the MCP server exists. So ConfigRef takes it as a Consumer<FleetConfig> and Fleetd.main supplies it — Fleetd is the one seam that already holds both a loaded config and the mcp package. The gotcha: nothing yet stops someone reverting Fleetd.main to the two-argument ConfigRef constructor. That one edit turns this gate off with the whole test suite green, which is why the wiring is worth reading before you trust it.

Do not put secrets in charter text. There is deliberately no ${ENV} interpolation. The OpenCode adapter writes the composed charter to a temp file so its CLI can read it, and that file is world-readable.


Know when a completion fallback is partial

What. If a member ends a turn without fleet_reply, fleetd scrapes its pane and resolves the waiting send with that tail, so the sender is not left hanging. The tail is capped at MAX_SCRAPE_CHARS (4000). When the cap bites, the returned text now ends with [Pane tail clipped: member did not call fleet_reply.] and fleetd logs a WARN with the original length and the cap.

On. Automatic, whenever the completion fallback reads more than 4000 characters.

Why. A clipped transcript used to be indistinguishable from a complete report. A delegating lead could act on an engineering report whose end had been cut off and never know — the only trace was a DEBUG line reading (4000 chars scraped), which reads like a size, not a warning. The cap itself is deliberate and unchanged; the defect was silence, not the number.

Gotcha. The marker does not recover the missing text, and it is not a licence to skip the reply. The fallback is a liveness net, not a channel: it returns only the pane tail, stripped to the last assistant block. A member must still end every delegated turn with exactly one fleet_reply. The marker is appended after the CB-115 misattribution guard compares the scrape to its baseline, so marking cannot make an unchanged pane look like new output.


Reject a profile name as a send target

What. fleet_send refuses a sessionId that exactly matches a configured profile name, on both the blocking and the wait:false path, before any ticket is issued. The error names the value and points the caller at fleet_list.

On. Always on; no configuration.

Why. A lead sent to sessionId: "sol" — the profile name, not the terminal id. The bridge accepted it and returned accepted — task delegated, so the lead believed the work was dispatched. About 60 seconds later the injector logged sol is gone, dropping its queue, and twenty minutes after that the ticket still reported pending — worker unknown. The whole delegation was lost and nothing told the sender. Confusing a profile for a session id is the single easiest mistake to make with fleet_send, because both are short names the operator sees side by side in fleet_profiles and fleet_list.

Gotcha. The check is deliberately narrow: it rejects only a value the bridge can prove is a profile. A target missing from the member roster is still accepted, because it may be a peer lead's terminal id or a herdr-owned pane the bridge did not spawn. So this catches one specific mistake well and is not a general "is this target real" validation — an unreachable target still fails later, in the injector. For the same reason profiles is a required parameter on both send methods rather than a defaulted one: an overload that defaults it would silently turn the check off, which is exactly how CB-561 disabled architect resolution while still compiling and passing its tests.


See free fleet capacity

What. fleet_list returns a top-level capacity block — for each profile its maxLoad, live, free and reclaimable count — and adds idleForSeconds and reclaimable to every member row. free is null when the profile is uncapped.

On. Always on; no configuration. The numbers come from profile maxLoad and the live roster.

Why. Free slots were invisible, so nothing told a lead when capacity was being wasted. A finished member kept holding a terra slot until a later spawn was refused outright with worker profile 'terra' is at maxLoad: 2 live >= 2 cap; refusing spawn — no fallback to another profile. The bridge already knew every fact needed to prevent that — the cap, the live count, and that the holder had finished — and simply never reported them. This is a reporting gap, not a policy change.

Gotcha. reclaimable means only "this member holds capacity and has no open bridge work". It is advisory. The bridge never spawns, stops or retasks a member to improve utilisation: it has capacity facts but no work list, and choosing work needs authority it does not have. That decision stays with the lead.

Two more things to know. live is read through the same liveCountRef function that placement consumes, so an advertised free slot cannot drift from what fleet_spawn will actually accept — a second count would eventually disagree, and a capacity view that lies is worse than none. And lifecycle.idleTtlSeconds defaults to 1800, so a finished member holds its slot for 30 minutes before the reaper takes it. That is far slower than slots turn over during active orchestration, which is why the terra refusal happened even though a reaper exists — the view makes the wait visible, it does not shorten it.

idleForSeconds derives from System.nanoTime(). It is monotonic within one daemon run and has no wall-clock meaning across a restart.


Answer a question on an async delegation

What. When a member calls fleet_ask during a wait:false delegation, fleet_poll on that ticket returns a non-terminal ASKING phase carrying the question text and its turnId. The lead answers with fleet_send{turnId, content}; the member resumes the same turn and its real reply still arrives on the original ticket.

On. Always on; no configuration.

Why. fleet_ask did not work at all on an async delegation. Outcome.QUESTION is deliberately non-terminal, but the poll path tested "did this complete?", so a question fell into the failure branch: the ticket was marked FAILED, and the question text and turnId were both discarded. The member blocked for 55 seconds, gave up, and had to abandon its task and report the ambiguity in its final reply instead. That is the mechanism working backwards — fleet_ask exists so a member can resolve a decision without losing its turn. It also made the guidance self-contradictory: leads are told to prefer wait:false for anything non-trivial and to answer an ask with fleet_send{turnId, content}, and both could not be followed at once.

Gotcha. The ask still times out after 55 seconds by default (115s maximum). Those caps are not a policy choice and raising them does not help: they exist because the member's own MCP client call would time out, so a longer wait just moves the failure. A lead has to be reachable inside that window. An unanswered ask returns the ticket to PENDING, not FAILED, because only the question wait ended — the delegated turn keeps running.


Learn why a delegation died

What. When a member's queue is dropped — its pane is gone, or it was stopped — the send fails with the real cause carried up from herdr, for example herdr error [agent_not_found]: agent target term_… not found. TurnListener.onTurnFailed takes that reason and CompletionResolver prefers it over the pane scrape and over the old fixed text.

On. Always on; no configuration.

Why. The bridge already knew the exact cause and threw it away. A real log pair from a member being stopped shows both halves two milliseconds apart: the injector logged dropping its queue … cause: herdr error [agent_not_found], while the sender was told only worker did not reply; its turn ended in an unrecoverable state (worker unreachable or stuck). Those two failures need opposite responses — a dead pane means respawn, a stalled model means wait or kill — and the lead could not tell them apart.

Gotcha. drop now fires onTurnFailed unconditionally, where it used to fire only when a turn was already in flight. That is the substantive half of the fix, not a tidy-up. A sender blocks on the rendezvous waiter, never on the delivered future — delivered is only inspected to label a timeout as queued or working. So completing delivered exceptionally never woke anybody, and a message that was queued but not yet delivered sat until its timeout, which is 30 minutes on an async send.


Watch the fleet's health

What. An opt-in background observer. Each tick it takes one AgentControl.list() for the whole fleet and one in-memory roster snapshot, joins them, and classifies every member. A member entering a fault state logs one WARN; recovering logs one INFO. fleet_list reports healthCoverage, which is off, detection-only, or full.

On. A health: block in fleetd.yaml with enabled: true. intervalSeconds defaults to 30 and is floored at 15. With no block at all nothing is constructed and no herdr call is ever made.

Why. A fault is usually a disagreement between two views, not a value you can read from one of them. A member that says BUSY in the session FSM while herdr says DONE has lost its turn boundary — a real trace sat in that state for eighteen minutes with its ticket still PENDING and no fallback firing. That is why both views must come from the same instant: one list call per tick, never one per member. Reading panes is the exception, not the method.

Gotcha. Detection and notification are separate keys on purpose. An earlier design required a webhook before health.enabled could be turned on, which would have removed real local detection to avoid a narrower human-notification gap. So health runs with no sink configured and reports detection-only — treat that value as "nobody will be paged", not as "health is off".

Two more. tick() catches Throwable and reschedules in a finally, because a ScheduledExecutorService never re-runs a task that threw: the earlier version rescheduled as its last statement, so the first agents.list() failure would have stopped health permanently and silently — exactly when the control link is down, the highest-priority state in the model. And the snapshot fields this unit cannot yet supply are the named constant NOT_YET_OBSERVED, not bare false, because to this classifier false means "no fault" rather than "not known yet".


Keep a worktree that still holds work

What. Before a finished session's git worktree is deleted, fleetd checks whether it still holds uncommitted changes. If it does, the directory is kept and a WARN names its path, the pane and the release cause. A clean worktree is removed as before.

On. Always on, for every worktree-backed member. There is no knob.

Why. A member's uncommitted work exists in exactly one place — its worktree — so deleting it is loss with no copy and no error. CB-544 already protected the shutdown drain for this reason, but left the ordinary COMPLETED release deleting with --force. That gap fired: the idle reaper released two members and deleted both worktrees, and only luck decided the work had already been pushed. A worker that ends a turn without committing — because it stopped to ask a question, or refused the turn — is the normal case, not the rare one.

Gotcha. The check is git status --porcelain with no --untracked-files=no, so an untracked file counts as dirty. That is deliberate: the work at risk in the original incident was a new file that was never git added, and ignoring untracked files would have missed exactly it. The cost is that a profile whose parity overlay ever copies an untracked, non-gitignored file would make every release preserve, and worktrees would pile up silently. Inert today — tracked overlay files carry --skip-worktree so --porcelain cannot see them, and fleetd.yaml is gitignored — but it is a real constraint on overlayParity, tracked in CB-581.

Second gotcha: hasUncommitted tolerates a worktree that is already gone and reports it clean. It has to. It runs inside SessionManager.release() after the registry entry is dropped and before the pane is stopped, so throwing there would orphan a live pane and strand a fleet_send caller on a rendezvous nothing resolves. Anything added to that window needs the same tolerance.


Fail a ticket when its member dies

What. When fleet health sees a member reach a terminal state — GONE or NEVER_READY — every ticket waiting on that member is failed straight away, naming the state as the reason, instead of staying PENDING until something else notices.

On. The same health: block that turns on health watching. No separate key.

Why. Detection without action just moves the silence. A lead that fires fleet_send{wait:false} and polls its ticket gets pending forever when the member behind it is already gone — the failure is known inside the daemon and invisible to the only caller who cares. Routing it through CB-568's existing idempotent target-wide failure means the outcome is also counted, so a dead delegation stops being invisible to /metrics.

Gotcha. It fires on the transition into the terminal state, not on every tick. An earlier attempt put the call outside the transition guard, so a member that stayed GONE had the failure operation invoked once per interval for as long as it remained in the roster; that commit was rejected. The flip side is the honest limitation: the new state is recorded before the bounded retries run, so if all three attempts throw, the tickets stay pending and no later tick retries. That path logs at WARN and has its own test — it is a known edge, not an oversight. The retries also carry no backoff.

Stopping a member also fails its ticket now, even mid-fleet_ask (fleetd #275). There are two ways a ticket loses its member, and they are not the same event. Health reporting GONE is a guess read off the live agent list; fleet_stop or the idle reaper releasing a session is a teardown the daemon performed, so it is certain. The sweep now takes a flag that says which one it is. On a certain teardown it also fails a ticket parked in fleet_ask — the member is gone, so nobody can ever answer that question — and closes the reverse rendezvous behind it. On a health guess it still skips an asking ticket, because the member may be answerable by a live lead and a guess must not kill it.

Gotcha for #275. The health path is deliberately left unable to self-heal one narrow case: a member goes GONE while its session stays in the roster, its fleet_ask then lapses on its own, and nothing re-fires the sweep — the transition already fired once, and FleetHealth.decide returns GONE before it could ever return DELEGATION_ORPHANED. That is filed as fleetd #280, and whether it is reachable at all depends on whether such a session is eventually released anyway, which has not been checked.


Tell a usage-limit refusal from a real reply

What. A member can end its turn without calling fleet_reply. The bridge then scrapes the pane and hands that text back as the answer. Sometimes that text is not an answer at all — it is the backend refusing, because the account hit its usage limit. With this on, the bridge matches the scrape against a pattern you configure. On a match it resolves the send as BACKEND_EXHAUSTED and carries the matched line as the reason, instead of passing a refusal off as a completed reply.

On. Per profile, exhaustedPattern: — a regex. Opt-in: leave it out and that profile's completion fallback behaves exactly as before. Startup logs one line naming which profiles have a pattern and which do not, so you can see the coverage without reading the config by hand.

Why. The old behaviour lied in the worst direction. A lead asked for work, got back a block of text, and had no way to tell "here is your answer" from "my account is refusing to run". The lead would then treat a refusal as a result. Keeping the pattern in config, never in Java, is deliberate: every backend words its refusal differently, so a sentence baked into the code would only ever match one vendor.

Gotcha. The pattern map is built once at startup from the config snapshot, so exhaustedPattern is a deferred key — adding one to a profile does nothing until the daemon restarts. fleetd.example.yaml does not say this yet. Also, this stage only classifies. Nothing yet stops the fleet spawning another member onto the same exhausted account, and nothing yet saves the work that member was doing — those are stages B and C of CB-578.


Stop spawning onto an exhausted account

What. When a member's turn is classified as a usage-limit refusal (see the entry above), the credential behind it is put in quarantine for a cooldown. While it is quarantined, an explicit spawn onto it is refused with a message naming the profile, the credential and roughly how many seconds are left; placement skips it under every policy; and fleet_profiles shows it. The quarantine lifts itself — there is no manual step.

Since fleetd #466 the cooldown escalates. Each consecutive exhaustion of the same credential doubles the wait, capped at 12x the base — about 6 hours at the 1800s default. A credential that keeps reporting exhausted is therefore retried roughly a dozen times a week instead of about 336 times.

On. quarantineCooldownSeconds: at the top level (default 1800, deferred — it is baked into the tracker at startup). Since fleetd #466 it is the base of the backoff, not the whole of it. Per profile, credentialId: (hot) says which credential this profile spends. Quarantine itself only ever fires for a profile that has an exhaustedPattern, so a fleet with no patterns configured behaves exactly as before.

Why. Detecting the refusal was only half the problem. Without this, the fleet answers an exhausted account by spawning another member onto it, which fails the same way, and the operator sees a run of dead workers rather than one clear cause.

Gotcha. It quarantines the credential, not the profile name, and that distinction is the whole point. In this fleet sol and terra are two different models billing one OpenAI account. Locking only the profile that happened to report the refusal leaves its sibling live, and the next spawn walks straight onto the same dead account under the other name. Profiles that share an account must share a credentialId. A profile that sets none quarantines alone, under its own name — safe, but it will not protect a sibling.

More gotchas, from the escalation (fleetd #466).

  • The reset is a time proxy, not a success signal. Nothing in the daemon reports a successful spawn back to the quarantine tracker, so "the account started working again" cannot be observed there. What clears the streak is a base cooldown's worth of quiet — no further exhaustion report for that credential. That is the best available evidence, not proof. Read it as "we have not been told it is still broken", never as "it is fixed".
  • The multiplier and the ceiling are constants, not config. Doubling (2.0) and the 12x cap live in BackendQuarantine, so there is no YAML knob for either and no new hot/cold question. quarantineCooldownSeconds stays deferred.
  • There is a ceiling on purpose. An unbounded backoff becomes a permanent outage that only a restart clears, which would be worse than the flat-rate retrying it replaced.
  • Cooling-off is a different mechanism and is NOT escalated. A credential that throws repeated non-exhaustion errors (an HTTP 5xx storm) gets BackendOutagePolicy's flat 60s, with no repeat tracking. Escalating that would turn a transient storm into a multi-hour outage. The two states are reported separately and a profile can be in both at once.

See which charter a member got

What. Every member launch records a CharterReceipt: the role, where the charter came from, a sha-256 digest of it, and its size in bytes. It is stored on the session and shown in the roster (fleet_list and GET /members). The charter text itself is never recorded.

On. Automatic.

Why. Charters are per-role config and are re-read on every spawn, so two members of the same role can get different text without anyone noticing. The digest answers "did this member actually get the charter I think it got?" without printing prompt text into logs an operator may not be allowed to keep. It also closed a real leak: the older pane-placement spawn log printed the whole argv, and the charter travels inside argv. That argument is now replaced by its digest.

Gotcha. PeerHandle.charterReceipt() is deliberately not a default method. It used to be, and OpenCodeLauncher's wrapping handle forgot to override it — so it answered null while the real receipt sat on its delegate, and sol and terra silently showed no receipt at all while Claude Code members showed one. Nothing failed; the roster field was just quietly missing. Removing the default makes the compiler catch that, and any new adapter must now answer the question on purpose. If you add a PeerHandle implementation, this is the line that will not let you skip it.


Redeploy the daemon safely

What. scripts/redeploy-fleetd.sh rebuilds the jar and restarts fleetd as one command. It builds before it stops anything, waits for the old process to actually exit, restarts from a login shell with cwd = fleetd/, then polls /healthz and reports the herdr protocol number, a fresh fleetd listening line, the config keys accepted or deferred at boot, and any ERROR lines since the restart. --check reports state and changes nothing; --yes skips the drain prompt; --no-build restarts the jar already on disk.

On. Run it. Nothing is automatic — the daemon never restarts itself.

An agent gets there through the redeploy-fleetd skill (.claude/skills/redeploy-fleetd/). It is a primary-side skill, and it holds the flags, the drain step, the operator's allow-list entry, and five numbered checks — login shell, drain members, deferred config keys, re-check fleet_whoami, prove the new jar runs — each of which has gone wrong here before. Workers must never load it: stopping the daemon kills the worker's own channel mid-turn. CLAUDE.md used to carry all 58 lines, and paid for them in every session's context. It now keeps only the two rules that must stay resident — a merge is not a deployment, and workers never redeploy — plus the line that names the skill.

Why. A merge is not a deployment: the running daemon holds the jar it was started with, so merged code does nothing until this runs. That gap has silently shipped inert features more than once — the whole health: stack sat merged and doing nothing for two tickets. The script also exists so the operator can allow-list one auditable command instead of approving a kill and a java -jar separately every time, which is what a lead would otherwise have to ask for on every deploy.

Gotcha. The check that matters most has no log line anywhere in the daemon: fleetd inherits WORKER_GITEA_TOKEN from the shell that starts it, via ${SHARED_ENV}/tools/secrets.sh. Start it from a non-login shell and the variable is empty — the daemon boots normally, /healthz is green, and the failure surfaces much later as workers that cannot open a PR. --check is the only thing that reports this, and it tests whether the name resolves without ever printing the value.

Second gotcha: a green /healthz only proves herdr answers. If herdr's protocol number has moved, every spawn can still fail. The script prints the protocol it saw so you can compare it; prove a real spawn before trusting the fleet.

The script never escalates to kill -9. The shutdown hook releases sessions and worktrees in order, and a hard kill can leave worktrees and panes behind; if the process will not exit it stops and tells you rather than forcing it.


Get told when an async delegation finishes

What. A fleet_send{wait:false} ticket that reaches a terminal phase — replied, failed, wedged, or abandoned — now injects a short nudge into the lead's own pane telling it to poll. Several tickets finishing at once coalesce into one nudge naming the count. Polling a ticket marks it collected, so a ticket you already read is never nudged about again.

On. Automatic when the lead is a herdr pane. Bounded by fleet.leaders.<name>.push_reminders (default 5) and push_backoff_ms (default 15000). A lead that is not a herdr pane leaves the registry empty, the loop becomes a no-op, and delivery degrades to pull — nothing is lost.

Since CB-590 there is one nudge schedule per lead, shared with the CB-307 reply nudge, so the two can no longer inject into the same pane at once. push_reminders is a budget per source, not one shared counter: reply work and ticket work each get their own, so a busy reply stream cannot spend the budget a ticket needs. Worst case a lead sees up to twice push_reminders nudges, which is the deliberate price of that isolation.

Why. The charter tells leads to prefer wait:false for anything non-trivial, because a blocking fleet_send is capped by the caller's own MCP client timeout of about 60 seconds. But until this landed, that preferred mode was the one mode with no notification at all: MessageService.reply returns on the rendezvous fast path before the push loop hears anything, so an async ticket finished in silence and the lead only found out by polling on a hunch.

Gotcha. The nudge goes to the lead's pane, not to its MCP session. A lead driving the REST surface directly — which is exactly what you fall back to when the MCP mount drops — receives nothing. Combine that with the 10-minute terminal-ticket TTL (MessageService.TICKET_TTL_NANOS) and a finished worker's report can be pruned before it is ever read. That happened during this feature's own close-out: a ticket returned 404 while its member still sat in done. The work survived only because the implementer skill had opened a PR, and the PR body carried the report. The ticket is not the durable artefact; the PR is. If you orchestrate over REST, poll on a timer.


Get told when a worker is waiting on your answer

What. A worker on an async (wait:false) delegation that pauses mid-turn in fleet_ask now nudges the lead's pane by itself, naming the exact call that resumes it:

Worker term_a asked a question (ticket task-3) — answer it with
fleet_send(turnId="term_a#1", content=...) to resume its turn:
which config file?

fleet_status{sessionId} shows the same open question, and so does REST GET /sessions/{id}/status (as question, turnId, ticket).

On. Automatic, on the same terms as the ticket nudge above — it is a third source in the same per-lead schedule, not a new push path, so CB-590's one-schedule-per-lead guarantee still holds and it spends from its own push_reminders budget.

Why. fleet_ask opens a reverse-rendezvous window of about 55 seconds. A lead polling on its normal cadence of minutes never saw it, so the worker timed out and carried on without an answer — the ask was, in practice, unusable on the delegation mode the charter tells leads to prefer. Two smaller holes closed with it: REST GET /tasks/{ticket} dropped turnId on an ASKING phase, so a REST caller could read the question and had no way to answer it, and fleet_status said nothing about an open question at all.

Gotcha. This closes the window; it does not remove it. The worker still gets ~55 seconds, and a nudge only helps a lead that is injectable right now — a lead mid-turn for a minute still misses it. So the standing advice is unchanged: do not brief a worker to "ask me." Decide the question before you delegate, or give the worker an explicit default to use.

Why the window is not simply widened: DEFAULT_ASK_TIMEOUT_MS = 55_000 sits just under the worker's own MCP client cap of about 60 seconds, so the daemon can return a clean typed timeout before the client severs the call. Raising the server constant buys nothing — the worker's client kills the call regardless.


Members cannot use the operator's admin forge token

What. Every member launch overwrites GITEA_ACCESS_TOKEN with a non-blank blocked sentinel, and sets a BRIDGED_MEMBER=1 marker. Applied in HerdrPeerLauncher.baseEnv after the profile's own env: map, so no profile — present or future — can name that key and restore the real value. CB-302's separate GITEA_TOKEN grant is untouched, so a member can still push and open its own PR with a repo-scoped token.

On. Automatic, both adapters, every profile.

Why. A member is an autonomous agent running arbitrary tool calls. It has no business holding the operator's admin credential, and the rule the operator set is explicit: leads and architects may use GITEA_ACCESS_TOKEN, members use WORKER_GITEA_TOKEN. A plain env overlay was not enough — the pane's login shell re-sources the secret store afterwards and would put the real value back — which is why the marker exists: the shell's export is guarded on BRIDGED_MEMBER being unset.

Gotcha, and it is the important one. This blocks one name. The pane's login shell sources a secret store exporting about thirty, and a second forge token is among the ones nothing blocks. Tracked as CB-596. The reason nobody noticed is worth internalising: a Claude Code member inherits the operator's user-scope ~/.claude.json servers, so it mounts 45 mcp__gitea__* tools including delete_branch and delete_file. They all fail — but only because the credential they read is blocked. Remove the block and the same picture becomes 45 working destructive tools, with no visible difference beforehand. Present-and-useless looks identical to absent.


Supervise the daemon without breaking the fleet

What. A launchd unit that restarts fleetd if it dies, and that still gets the fleet's secrets. deploy/dev.ltms.fleetd.plist runs scripts/fleetd-launchd-wrapper.sh, which execs one login shell in place (exec /bin/zsh -lc 'exec "$@"' -- "$@") and then execs the real java command. One exec chain, so launchd keeps tracking the right PID. scripts/redeploy-fleetd.sh detects whether the agent is loaded and switches stop/start to launchctl unload -w / load -w, falling back to its original kill + nohup when it is not.

On. Not automatic, and deliberately so. Install it yourself:

cp deploy/dev.ltms.fleetd.plist ~/Library/LaunchAgents/
launchctl load -w ~/Library/LaunchAgents/dev.ltms.fleetd.plist
launchctl list | grep fleetd

scripts/redeploy-fleetd.sh --check reports whether the agent is installed and whether it is loaded. It is read-only.

Why. The unit shipped with CB-504 was never installed, and could not have worked if it were. launchd does not source a login shell, so a launchd-started daemon would have had no WORKER_GITEA_TOKEN and no AI_GATEWAY_TOKEN. It would have started fine and looked healthy; the failure would have appeared hours later as workers unable to open a PR. So the operator's real choice was a supervised daemon with a broken fleet, or a working fleet with no supervision. The wrapper removes that choice. The launchctl branch in the redeploy script matters just as much: a SIGTERMed daemon exits 143 even when its shutdown hook completes normally, so under KeepAlive{SuccessfulExit: false} a bare kill makes launchd restart the old jar, racing the script's own restart.

Gotcha. The script computes its log path from where the script file sits; the plist hard-codes an absolute StandardOutPath. Nothing checks that the two agree. If they ever diverge — a worktree, a renamed clone — the script's post-restart ERROR check reads the wrong file, finds nothing, and reports "ok" while the daemon crash-loops. And the loop really is unbounded: ThrottleInterval: 10 paces restarts to one per ten seconds, it does not cap how many. Fix CB-600 before installing the agent.


See at startup which secrets the daemon actually got

What. fleetd logs, at startup, every secret environment variable it needs, and whether each one resolved or is MISSING. The required set is derived from the loaded config — each non-subscription profile's tokenEnv, plus every profile's gitTokenEnv — not hard-coded, so a new profile is covered the day it is added.

On. Automatic. Read the startup secret … lines at the top of fleetd/fleetd.out.

Why. An empty token used to be completely invisible. The daemon started, /healthz went green, and the first sign of trouble came much later and somewhere else — a worker that could not open a PR, or a gateway profile that could not authenticate. Neither symptom points back at the shell the daemon was started from, which is the actual cause. This turns a silent, delayed, misattributed failure into one line at startup.

Gotcha. It reports names and set/MISSING only — never a value, a prefix, or a length. That is deliberate and must stay that way; the log is not a secret store. Also, MISSING is a warning, not a refusal: the daemon starts anyway, because refusing to boot over a credential that half the fleet may not need would be worse. So the line has to actually be read. One known false alarm: a non-subscription profile that never sets tokenEnv inherits the default name FLEETD_WORKER_TOKEN and is reported missing — which is nearly always a real misconfiguration rather than a bug in the report.


Bound the AMQP backlog, and know a reply was really published

What. Two guarantees on the durable reply inbox. The consumer calls basicQos before basicConsume, so unacked messages beyond the window stay on the queue instead of being pushed into the daemon's heap. And publishing runs on its own confirm-mode channel with mandatory=true and a return listener, so an unroutable or unconfirmed publish raises an error instead of vanishing.

On. broker.prefetch sets the window, default 32. The confirm behaviour is automatic whenever broker: is configured at all.

Why. Both existed to make the word "durable" true. Without prefetch the queue sat near-empty while the real backlog lived in an in-memory map with nothing capping it — so queue-depth metrics read healthy, and any queue-level limit would have guarded an empty queue. Without confirms, a publish to a queue that was never declared was a silent black hole, and deliveryMode(2) bought nothing, because "persisted" is only true after the broker says so.

Gotcha. A broker Return always arrives before its matching Confirm, so an ack alone does not mean routed — the confirm path has to check a per-message returned flag, and that ordering is easy to get wrong when editing this code. Note also that these tests are @Tag("contract"): they are excluded from the default build and run in a separate CI job against a real rabbitmq:3.13 container. A green default build says nothing about them. And the broker here is LavinMQ; the RabbitMQ client library is used only because LavinMQ speaks the same protocol.


Three config keys that changed meaning — read this before upgrading

What. weight, maxLoad and fleet.leaders.* all changed what they mean, not just what they do. A config file that worked before an upgrade can behave differently, or refuse to start, with no edit to it. The three are grouped here because an upgrading operator meets them together.

Key Used to mean Now means
profiles.<n>.weight: 0 coerced to 1.0 — "pick me as often as anyone else" excluded from automatic selection
profiles.<n>.maxLoad: 0 coerced to unlimited capped at zero live members
fleet.leaders.<n>.terminal a terminal-id pin rejected at startup — use tab

On. Nothing to switch on. These are the meanings now.

Why. In the first two cases the old behaviour was the exact opposite of what the key reads like. weight: 0 looked like "never pick this" and meant "pick it normally". maxLoad: 0 looked like "never run anything here" and meant "unlimited" — and that was the only throttle a subscription: true profile had against the operator's own paid plan. A key whose plain reading inverts its behaviour is a trap, so both were fixed toward the reading. A negative maxLoad is now a startup error rather than being silently normalised away.

fleet.leaders.*.terminal went further and is refused, naming the offending entries, instead of being ignored. A silently-ignored lead pin means the lead is not recognised, gets demoted to worker, and every orchestration call is refused — a failure that looks nothing like its cause. Failing at startup is the kinder outcome.

Gotcha. weight: 0 excludes a profile from automatic selection only. An explicit fleet_spawn{profile:"..."} bypasses placement entirely and still resolves it, so weight: 0 is not a way to disable a profile. maxLoad: 0 is, because that cap now applies to explicit spawns too. And a lead's tab must match a real herdr tab label exactly (case-insensitively) — the old tabPrefix no longer finds a lead, it only guards against a worker's label colliding with the convention.


Stop an explicit spawn from busting the cap

What. maxLoad is now an unconditional cap. An explicit fleet_spawn{profile:"X"} used to skip the check entirely — the cap applied only to automatic placement — so naming a profile was a way around it. Now an explicit spawn is refused when live >= cap, and it is not re-routed to another profile.

On. profiles.<name>.maxLoad.

Why. A cap that any caller can opt out of by naming the profile is not a cap. The no-fallback part is deliberate: a caller who named a profile did so for a cost or model reason, so quietly moving the work to a different backend would defeat the reason they named it. Better to refuse and let the caller decide.

Gotcha. There is a known TOCTOU race: liveCount is read outside a lock, so two genuinely concurrent spawns can both pass the check and briefly exceed the cap. It is documented in the code rather than fixed. Also note the refusal message only became visible to callers in CB-599 — before that it was a blank 500 with the reason in the log.


Nudge an idle lead back to work

What. An opt-in loop that nudges the one idle lead after it has been continuously injectable for a quiet period with nothing driving it. It fills the gap the reply push loop leaves: that loop only fires when a worker reply lands, so a lead that is simply sitting idle with nothing arriving is invisible to it.

On. The leadHeartbeat: block — idleAfterSeconds (default 300), backoffMs (default 60000), quietNudgeCap (default 3). Absent means off, and the loop is not even constructed. Changing it needs a restart.

Why. Off by default, deliberately. This loop spends the operator's model subscription on the daemon's own initiative, so it must never switch itself on during an upgrade. That constraint shaped the whole design.

Gotcha. It is status-gated (a WORKING lead is never touched), debounced, caps consecutive quiet nudges, and stands aside while the reply push loop is already nudging. That last one depends on ReplyPushLoop.isActive() being honest — see CB-598, where a schedule could die with work still pending and report inactive.


Prove which charter a member got, without logging the prose

What. Per-role launch charters are read from the live config and composed once per spawn: the role charter plus the reply charter. Claude Code gets it as an appended system prompt; opencode gets an ephemeral member-charter.md referenced from its generated config. A CharterReceipt records a sha256 and byte count of the exact composed bytes, and rides on both the spawn log and the roster.

On. fleet.charters.<role>, where role is dev, architect or reviewer. Read live per spawn, so a change applies to the next spawn with no restart.

Why. Two problems at once. Charters had been hardcoded, which meant changing what a role is told required a rebuild. And an operator asking "what was this member actually told?" had no answer that did not involve printing the prose into a log, where it does not belong. The receipt answers the question with a digest instead.

Gotcha. The role charter is combined with the reply charter only when the profile mounts the bridge MCP. A non-MCP profile gets the role charter alone — which is correct, since the reply charter tells a member to call fleet_reply, and a member with no bridge cannot.


Tell "busy" from "refusing" in the capacity view

What. In fleet_list's capacity view, a quarantined profile's row is forced to free: 0 whatever its maxLoad and live counts say, and gains credentialId and quarantinedForSeconds.

On. No new knob. The quarantine facts come from the existing exhaustedPattern, credentialId and quarantineCooldownSeconds keys.

Why. A quarantined profile used to show free slots it would refuse to fill. A lead reading that row would keep trying and keep failing. The two states need different reactions — "busy, will free up" means wait, "refusing for N seconds" means go elsewhere — and the row could not express the difference.

Gotcha. The two new fields appear only when a profile is actually quarantined, so an ordinary fleet's rows are byte-identical to before. Do not write a client that expects them. The view shares its QuarantineSource with fleet_profiles so the two surfaces cannot disagree.

The row also says which attempt this is, not only how long is left (fleetd #466, #473). Both fleet_profiles and fleet_list's capacity rows carry quarantineAttempt beside quarantinedForSeconds: 1 for a first refusal, 2 for the second in a row, and so on. The cooldown escalates with that count, so quarantinedForSeconds: 1800, quarantineAttempt: 1 and quarantinedForSeconds: 7200, quarantineAttempt: 3 mean very different things about the credential.

Why it exists: the seconds alone cannot tell a weekly subscription limit apart from a one-off capacity blip. A lead seeing attempt 4 knows retrying is pointless and should move the work to another credential, or tell the operator. Gotcha: the count keeps growing past the cooldown ceiling, so a large quarantineAttempt with a flat quarantinedForSeconds is normal, not a bug — the ceiling caps the wait, not the streak. A quiet gap long enough to clear the quarantine resets the count to 1. The flat two-argument BackendQuarantine constructor still reports a real growing count even though its own cooldown never escalates, because the streak is a fact either way. Both values come from one status() call that does a single map read, so the reported count can never disagree with the cooldown it describes.


Resume a member onto its previous conversation

What. A member's own agent-session id is captured at spawn, survives onto the roster as agentSessionId in fleet_list and GET /members, and fleet_spawn accepts sessionName and resumeSessionId to relaunch onto that same conversation.

On. fleet_spawn{sessionName, resumeSessionId}.

Why. The pieces existed but were connected at neither end: nothing persisted the id, and nothing exposed a way to pass it back. So a member that died took its context with it even though the backend could have resumed it.

Gotcha. A resumeSessionId requires an explicit profile, and that profile's adapter must declare Capability.SESSION_RESUME. A placement-routed spawn carrying a resumeSessionId is refused, because a resumed conversation is tied to the specific backend that started it. If the adapter lacks the capability the spawn is refused naming it, rather than silently starting a cold session that looks resumed.


opencode members run with --auto and forced auto-compaction

What. opencode workers launch with --auto, which auto-approves the permissions opencode does not explicitly deny, and the bridge-generated config pins compaction.auto: true.

On. Neither is configurable. Both are unconditional for the opencode kind.

Why. A worker that stops to ask for permission on a tool call has no one to ask — the lead is not watching its pane, and an unanswered prompt burns the whole turn. Auto-compaction is the same argument: a worker that runs out of context dies mid-turn and loses its fleet_reply, so its report is gone even though the work was done.

Gotcha, and it is a real trade. opencode's own help calls --auto "dangerous!". The blast radius is bounded by everything else about a member — its own worktree, its own branch, off-subscription, and it cannot merge — not by the flag. And compaction.auto: true in the generated config wins over the operator's home config, so auto-compaction cannot be turned off for fleetd workers from there. That was chosen knowingly: a lost report was judged worse than an unwanted compaction.


Silent member failures now say why

What. Logging only, no behaviour change. The paths that used to fail a member with no log, a bare DEBUG, or a message naming only the symptom now log at WARN with the real cause and the numbers — across the injector, the completion resolver, session acquire/reap/failure, and ticket abandonment.

On. Automatic.

Why. A member could fail and leave nothing to read. The lead saw a session that stopped responding and had no way to tell a crashed backend from a wedged pane from a reaped session.

Gotcha. You will see more WARN lines than before if you filter by level. That is the feature, not noise — but it will change what a log-volume alert sees.


Sessions are one-shot — recycle() is gone

What. SessionManager.recycle() — release plus immediate re-acquire under a new pane id — was removed. A released or finished session is torn down, never pooled or reused. Callers must release then acquire explicitly.

On. Nothing to configure; this is a removal.

Why. Recycling quietly reused a session across two unrelated pieces of work. The lifecycle is meant to be strictly one-shot, and a method that bypassed it was a standing invitation to leak one delegation's context into the next.

Gotcha. None left in-tree — no references remain. It had no external surface, so only internal callers were ever affected.


A failed ticket names the conversation, not just the files

What. When a member dies, the failed ticket's detail now carries the member's agentSessionId alongside the worktree, branch and snapshot ref it already reported:

… worktree=/wt/x branch=worker/x snapshot=refs/wip/x agentSessionId=abc-123

A lead can pass that id back as fleet_spawn{resumeSessionId: "abc-123"} on the same profile.

On. Automatic, no knob. It appears whenever the released member's backend produced a session id.

Why. CB-578 stage C already lets a lead re-dispatch a fresh member onto the same worktree after a failure, and it states its own limit plainly: the files survive, the reasoning does not. A re-dispatch was a cold start re-reading a brief. Pairing the session id with the worktree is the other half — it turns "the work survives" into "the work and the thread both survive".

Gotcha. It is absent whenever the adapter never resolved an id: the member was spawned without sessionName or resumeSessionId (the ordinary path), or its backend does not declare Capability.SESSION_RESUME. Silence here is the expected case, not a fault — do not read a missing agentSessionId= as a bug.


weighted placement is not "cheapest first"

What. placement: weighted is smooth weighted round-robin. It spreads unqualified spawns across every profile that has a free slot, in weight ratio. It has no concept of cost. So local: 10 against terra: 2 and sonnet: 1 does not mean "use local, overflow to paid" — it means roughly a quarter of spawns go to a paid profile while the free box still has a free slot.

On. placement: weighted in fleetd.yaml, with a weight: per profile. weight: 0 excludes a profile from automatic placement entirely; an explicit fleet_spawn{profile:...} bypasses placement either way.

Why. The operator's rule is cheapest-first: keep the free boxes busy and pay only for genuine overflow. No shipped policy expresses that. FixedPlacementPolicy ignores maxLoad and throws instead of overflowing; RoundRobinPlacementPolicy has no cost notion. CB-589 tracks a real cost-first policy. Until then the workaround is to make the ratio decisive rather than proportional — on this host local.weight is 100 against paid weights of about 1, so the free box wins every pick it is eligible for.

Gotcha, two of them. First, the running score map lives for the daemon's whole life. While a profile sits at maxLoad it is filtered out and its score freezes, so paid profiles keep accumulating against it; when the free slot opens the profile returns with a stale score and can lose the next pick — a paid spawn while the free box is idle. Second, and worse for the next person: the workaround expresses a preference order through a ratio knob. Add a profile at weight 150 later and it silently outranks the free box, with nothing to warn you. fleetd.yaml is gitignored, so a fresh host starts without this workaround and quietly pays — the reasoning is written into fleetd.example.yaml next to the key for exactly that reason.


Which inherited credentials a member may keep

What. A member's pane runs a login shell, which re-sources the operator's secret store and exports about thirty names. Exactly one of them is blocked: GITEA_ACCESS_TOKEN, see Members cannot use the operator's admin forge token above. This entry records the decision about the two that matter most among the rest, so that "we chose to allow it" never again looks the same as "we never noticed".

Credential A member may hold it Why
AI_GATEWAY_TOKEN yes It is the key a member is meant to use. Paid-backend members reach their model through llm.ltms.dev, and that token is the single front-door key. Blocking it would stop those members working at all.
CONTEXT7_TOKEN yes It backs the context7 documentation MCP, a read-only docs lookup that members are meant to have. The worst case is documentation reads on the operator's quota.
GITEA_ACCESS_TOKEN no Admin scope on the forge. Blocked — see the entry above. Members push with the repo-scoped WORKER_GITEA_TOKEN instead.

On. Nothing to turn on. The two allowed tokens flow through by default; the blocked one is overwritten at every launch.

Why this is written down at all. BRIDGED_MEMBER already exists, so splitting either of these per-role is now a one-line guard in the secret store. The option is cheap and available — which is exactly why the decision has to be explicit rather than implied by nobody having done it. The original leak (CB-592) was found by accident. The same accident should not have to happen twice.

Gotcha. This decision covers two names out of about thirty. The rest were enumerated later by CB-596 — see A member keeps only the credentials you name below, which supersedes this entry's scope. Read this table as the two that were decided first, not as the whole policy.


A member keeps only the credentials you name

What. A memberCredentials: block in fleetd.yaml lists every credential-shaped variable on the host, says which ones a member may keep, and blocks the rest. A blocked name is not unset — it is overwritten with a fixed sentinel string, blocked-by-fleetd-cb596-see-gitea-issue-82, so a member that reads it sees "deliberately blocked" rather than an empty variable it might quietly work around. Before this, a member pane inherited the operator's whole secret store and exactly one name was blocked.

memberCredentials:
  policy: deny-by-default     # the only policy today; named so a future allow-by-default is a change
  allow:                      # names a member MAY keep
    - AI_GATEWAY_TOKEN
    - WORKER_GITEA_TOKEN
    - CONTEXT7_TOKEN
  known:                      # every credential-shaped name on this host
    - GITEA_ACCESS_TOKEN
    - ...

The blocked set is known minus allow, computed at load. A name in both is an error you cannot make by accident — the intersection is empty by construction, because allow wins.

On. Add the block to fleetd.yaml (there is a commented template in fleetd.example.yaml) and add the matching guarded export to the operator's secret store. Both halves are needed — see the gotcha. The block is re-read on every spawn, so editing it takes effect without a restart; only the startup summary line needs one.

Why it exists. CB-592 found that a member inherits the operator's credentials, and blocked one name — GITEA_ACCESS_TOKEN — with a hardcoded string in the launcher. CB-593 then decided two more by hand. That does not scale and, worse, it hides the shape of the problem: a hardcoded list of one looks finished. Naming every variable in config turns "which secrets does a member hold?" from a question nobody can answer into a list you can read, review and diff.

The daemon says so when the block is missing. With no memberCredentials: the daemon still starts — refusing to boot would strand an operator who has not migrated — but logs a WARN saying every member pane inherits the whole secret store unblocked. With the block present it logs the counts instead:

memberCredentials: 34 known name(s), 5 allowed — blocking 29 on every spawn

It also reports what you forgot. On every spawn the launcher scans the host environment for names shaped like credentials (TOKEN, SECRET, _KEY, APIKEY, PASSWORD, CREDENTIAL, AUTH) and reports any that are on neither list. This is the part that pays for itself: the first spawn after it deployed named two variables no hand-written list had ever contained, because neither lives in the secret store —

WARN memberCredentials gap: 2 credential-shaped env var name(s) are on neither known: nor allow:
     — every member pane inherits them UNBLOCKED — [CLAUDE_CODE_MESSAGING_TOKEN, SSH_AUTH_SOCK]

Since 2026-08-31 the report distinguishes three cases, because one wording was lying. A name being on neither list does not by itself mean a member inherits it — under allow-list on a zsh login shell the generated scrub blanks it anyway. The severity now follows what the scrub actually does, decided with the same predicate the scrub itself evaluates:

Situation Line Meaning
deny-by-default, or allow-list on a non-zsh shell WARN … inherits them UNBLOCKED Nothing scrubs these. Real exposure.
allow-list + zsh, name not on the derived allow-list INFO … the scrub blanks them anyway Contained. Listing it just makes that explicit.
allow-list + zsh, name kept by the derived allow-list WARN … the derived allow-list keeps them anyway Real exposure, and easy to miss.

That last row is the one worth knowing about. The derived allow-list is a superset of known + allow: it also picks up every profile's gitTokenEnv, gitHostEnv, tokenEnv and env: keys. So naming a variable in a profile silently grants it to members, whether or not it is on either list — and until this change the daemon reported that case as contained. Each report kind has its own once-per-daemon guard, so a benign INFO can no longer suppress a serious WARN.

SSH_AUTH_SOCK is now blocked, and that is a downgrade, not a win. It was once allowed on purpose, because a worktree's remote was ssh://git@git.ltms.dev and a member could not push without the agent socket. Members now push over HTTPS with WORKER_GITEA_TOKEN, so the live config sets sshAgentEnv: omit (spelled sshAuthSock: block before #266; both still parse), and MemberEnvAllowList.derive drops the name even if an operator lists it under allow: — the config cannot re-grant it by accident.

Do not read that as the problem being solved. #184 measured a member with the socket blanked pushing fine anyway: the forge key is a readable, passphrase-free file, and ssh -G finds it outside ~/.ssh. An environment control cannot remove a file. Blocking the socket closes one door in a room with another door open; the actual fix is a separate OS user (#185).

Gotcha — under deny-by-default the config half alone does not hold. The launcher writes the member's environment at spawn, and then the member's pane runs a login shell, which re-sources the operator's secret store and overwrites it. So deny-by-default must be paired with a BRIDGED_MEMBER-guarded block at the end of the secret store that re-applies the same sentinel. Two lists that must agree — which is why the gap detector exists, and why the daemon logs its counts at startup. If a member ever reports holding a name you blocked, the secret-store half is what is missing.

That block must be the last thing in the secret store. It overwrites the blocked names, so anything that re-exports them afterwards silently undoes it.

policy: allow-list removes that whole problem (CB-633). It is the setting to prefer.

memberCredentials:
  policy: allow-list      # default is "deny-by-default"
  sshAgentEnv: omit       # "inherit" only if a member must use the operator's ssh-agent

Why it exists. A control that lives inside a sourced file can always be undone by a file sourced later, and that is not a hypothetical: on this host .ltms sources mgnlSecrets.sh one line after the guarded secrets.sh, so seven credentials reached every member in full — including an AWS key with AdministratorAccess. Four of the seven were already on the block list. The list was correct and it still failed. Making the list longer fixes nothing.

How it works. The daemon generates a throwaway ZDOTDIR directory per spawn and passes it in the pane-creation env map, before the shell starts. Each generated startup file sources its $HOME counterpart first and then runs the scrub, so the scrub happens after the operator's whole chain and nothing sourced later can undo it. The kept-name set is derived, never typed: every profile's tokenEnv/gitTokenEnv/gitHostEnv values and env: keys, plus an infrastructure set (PATH HOME SHELL TERM LANG LC_* TMPDIR USER LOGNAME PWD SHLVL EDITOR PAGER ZDOTDIR JAVA_HOME XDG_*), plus the exact keys this spawn's own env overlay carries. Adding a profile can only widen the set, so it can never break another spawn's scrub. Under this policy known:/allow: stop being a control and become reporting only — they still feed the gap warning.

Gotcha — it is zsh only. ZDOTDIR means nothing to bash. If the member's shell is not zsh the daemon logs a loud WARN saying protection is off and falls back to the deny-by-default overlay. That is weaker, so put members on a zsh account.

Gotcha — the platform nearly made this a dead control. The scrub first lived in the generated .zlogin, and zsh reads .zlogin only for a login shell. herdr does not open the same kind of shell everywhere: a macOS pane runs -zsh (login), a Linux pane runs a plain /usr/bin/zsh (interactive, not login). So it protected the developer's Mac and would have protected nothing at all on Linux, with no error anywhere. The scrub now runs from the generated .zshrc, .zlogin and .zshenv (#388) — .zshenv is the only file zsh reads unconditionally, so the control no longer depends on enumerating which kinds of shell exist. If you port this to another terminal backend, check what kind of shell it opens before trusting it.

Do not infer the shell kind from argv[0]. A bare /usr/bin/zsh proves not login and says nothing about interactive. Both this repo's javadoc and a host's fleetd.yaml once claimed "interactive but NOT login" on that evidence alone, and neither had measured it. Worse, a probe run from inside an agent measures the agent's own zsh -c child, not the pane — opencode forks a fresh non-interactive shell per command, so it reports truthfully about the wrong process. Read the pane shell's own /proc/<pid>/environ, or replicate the shape with script -qec zsh /dev/null.

How to tell it actually ran. Each pane writes a scrub-report.txt, and the daemon logs allowed N of M environment variables when that pane stops — the denominator is the point. A missing report is logged at WARN: the scrub then cannot be confirmed to have run at all, and a silently dead control is the failure this policy exists to remove.

Gotcha — a missing report has more than one meaning, and reading it as one cost a day (#394). The report block is the last statement in scrub.zsh, so its absence proves only that the script did not reach the end. It does not tell you which shell ran. On one host the missing report was read as evidence that the pane shell was neither login nor interactive; the real cause was that export UID= in zsh is a fatal parameter error which terminates the whole sourced file. The blanking loop is wrapped in { ... } 2>/dev/null, so the message was swallowed too.

The severity was the selection, not the count. env lists inherited names first and a startup file's own exports last, so the loop blanked the harmless inherited half and died immediately before the operator's own exports — the credentials the policy exists to remove. Measured there: UID was name 42 of 57, and a ~/.zshrc decoy at 58 survived on 8 of 8 spawns. A partial scrub got precisely the wrong half.

Fixed by routing each attempt through eval "export ${n}=" 2>/dev/null, which contains the error to one iteration, rather than by skipping the known-fatal names (UID EUID GID EGID PPID LINENO). A skip-list has to be complete forever; this is a security control, so it must not depend on an enumeration being right. The report's first line now reads allowed N of M failed F, unblankable names are listed with a ! prefix, and the daemon WARNs naming them — so "could not blank this one" is now visible instead of being an abort you learn about from a missing file.

Two things to plant when you audit this yourself. First, a decoy: a non-secret variable the operator's chain exports that is not on the allow-list is the only thing that separates "scrub skipped" from "nothing was there to scrub" — allow-listed names come back blank either way, so their blankness proves nothing. Second, check for .zcompdump in the pane's ZDOTDIR: its presence proves an interactive zsh ran compinit there. A .zcompdump present and scrub-report.txt absent is only explained by "sourced, then aborted partway".

Measured, not assumed. A real login zsh started from a clean parent kept 15 of 78 names with zero profiles configured; an earlier prototype run kept 3 of 28. The test asserts equality between the survivors and baseline ∩ derived allow-list, not a spot check of a few blocked names.

Verified, not assumed. Measured on 2026-08-17 inside a live member pane, once both halves were in place: 29 of 29 blocked names hold the sentinel, and the allow-listed names present in that pane keep their real values. The operator's own login shell is unchanged, because the whole block sits inside if [ -n "${BRIDGED_MEMBER:-}" ]. To repeat the check, spawn a member and run scripts/probe-member-credentials.sh; it refuses to run anywhere but in a member.

Third gotcha — the probe has the same defect it is checking for. It carries its own hardcoded name list and does not read the config, so it under-reported by three names and still exited 0. Read its count against the daemon's startup line rather than trusting the table. Tracked as CB-608.

Second gotcha — allow silences the warning. The detector treats known ∪ allow as covered, so adding a name to allow makes its warning go away and lets the value through. That is correct behaviour, but it means the quiet way to dismiss a gap warning is also the permissive one. Prefer adding to known.


/metrics and /healthz

What. Two read-only diagnostic endpoints. GET /healthz calls herdr's ping and reports only whether herdr answered: 200 {"status":"ok","herdr":{"version":…,"protocol":…}}, or 503 {"status":"degraded","herdr":"unreachable",…} on any HerdrException. GET /metrics renders every registered series as Prometheus text (CB-502): counters fleetd_sends_total, fleetd_replies_total, fleetd_push_nudges_total, fleetd_lead_heartbeat_nudges_total, fleetd_spawns_total, fleetd_herdr_calls_total, fleetd_auth_failures_total, plus gauges fleetd_sessions (one series per lifecycle state) and fleetd_inbox_depth (one series per undrained target).

On. Both are always registered. /healthz carries no authorization check at all — it is reachable by an unauthenticated caller by design. /metrics is gated on Authz.Action.METRICS: open to PRIMARY/WORKER/ARCHITECT, refused to ANONYMOUS.

Why. FleetdMetrics's own class doc states the design intent directly: "each series maps to a failure mode this project has actually hit, not to whatever was easy to count," and names the two worth watching — a rising fleetd_sends_total{outcome="completion_fallback"} share (turn detection degrading) and fleetd_push_nudges_total{outcome="exhausted"} (the primary stopped draining its inbox). /healthz's narrow scope traces to CB-504: under supervision the daemon must serve before herdr's socket even exists, so "degraded but alive" needed one cheap, reliable signal.

Gotcha. Neither endpoint proves the fleet actually works. /healthz echoes back whatever protocol number herdr reports, but nothing in the codebase compares that number against what fleetd's own herdr calls need — and the two have already drifted apart in the source itself: AgentControl's class doc says it was "ported to herdr protocol 19 (herdr 0.8.0, CB-521)," while HerdrClient's class doc still says "protocol 14, herdr 0.7.0." This is the exact CB-521 incident: herdr answers ping correctly and /healthz goes green, while agent.start and the rest of the protocol-19 surface fail because the adapter and the herdr binary disagree on protocol version. /metrics has no counter or gauge for that mismatch either — a resulting spawn failure only shows up as a fleetd_spawns_total{outcome=…} tick, and only once something actually tries to spawn.


Bearer-token auth and the non-loopback-bind fail-fast

What. auth.mode decides how a non-worker caller proves it is the primary: loopback-trust (default) — any loopback caller that is not a known worker/lead/architect pane is the primary, no credential needed; or token — such a caller must present Authorization: Bearer <token> or resolves to ANONYMOUS. A worker/lead/architect pane is always resolved from the unforgeable loopback-PID→herdr-pane mapping regardless of auth.mode. Separately, the daemon refuses to start at all when bind.host is non-loopback and auth.mode is still loopback-trust.

On. auth.mode: loopback-trust | token (default loopback-trust); auth.tokenEnv names the host env var holding the token (default FLEETD_API_TOKEN, read only under token mode). The bind fail-fast has no separate switch — it always runs in main().

Why. Stated directly in the code: loopback-trust's safety depends entirely on the OS refusing non-local connections to a loopback socket. Widen the bind without switching to token mode and "not a known worker" silently becomes "any client that can reach this port is the primary" — the most privileged role on the bus (spawn/stop/send/drain on any session). Rather than document the hazard, the config makes it unrepresentable: it throws instead of starting.

Gotcha. There are two separate fail-fast throws, both inline in main(), both before the daemon binds its port — so a bad config never opens the socket at all. Under launchd, that repeats forever: the plist's own comment warns launchd retries a fast-failing job every ThrottleInterval (10s) with no give-up count, until a human unloads the agent or fixes the cause. And auth.tokenEnv naming an unset/empty var throws a different message than the bind-mismatch check — don't assume one error class covers both.


The per-session authz table and the audit log

What. Authz.permits(Principal caller, Action action, String targetSession) is one static table stating, for each of eight actions (SPAWN, STOP, SEND, REPLY, ASK, DRAIN, READ, METRICS), which role may call it: SPAWN/STOP/DRAIN are the primary alone; SEND is primary or architect; REPLY/ASK require the caller to own the target session (its own pane, checked structurally, never by argument); READ/METRICS are open to any authenticated role. It is the single gate behind both entry paths (REST and MCP) — FleetdApp.allow() and BridgeMcp's own check both call into it, so the rule can't drift between the two surfaces. Refusals are recorded by AuditLog, an append-only JSON-lines trail written by a dedicated audit logger to logs/audit.log (daily rolling, 30-day retention, 100MB cap), independent of the daemon's normal app log.

On. Always on; not configurable. Every request through FleetdApp or BridgeMcp passes through Authz.permits().

Why. The class doc states this plainly: most of the rule was already true de facto — a worker's identity comes from its connection, never an argument, so it could never reply as another worker over MCP — but the REST surface used to trust the session id in the URL path outright, and neither surface checked role at all. This makes the invariant explicit and testable instead of emergent, and gives every privileged action one recorded outcome (allowed/denied/failed) instead of none.

Gotcha. Message content is never written to the audit log by design — only who/what/target/outcome/reason and a correlation id, because the bus carries user source code, diffs, and prompts, and an audit trail that quietly accumulated those would be a transcript archive wearing a security control's clothing. READ actions are deliberately excluded from the allowed audit trail ("reads would drown the trail") — only denials of READ/METRICS are recorded, not successes. And the 401-vs-403 split matters if you're debugging a refusal: 401 means "you presented no usable identity" (fixable by the caller), 403 means "you are authenticated, but this isn't yours" — a worker reaching for another worker's session, or for orchestration it was never granted.


Multi-profile routing and kind: adapter selection

What. Each workers: profile carries a kind: field selecting which backend launcher spawns it — claude-code (the default) or opencode (CB-402). At startup, main() partitions every configured profile into two maps by Profile.isOpenCode(), builds one ClaudeCodeLauncher and/or one OpenCodeLauncher accordingly, and wraps both in a CompositePeerLauncher that routes each call to whichever adapter declares the profile the call names.

On. kind: claude-code | opencode on a profile; absent or blank defaults to claude-code. The claude-code adapter is built even with zero claude-code profiles configured, unless opencode is the only kind present — so a bridge with no workers: at all still has a well-defined base adapter.

Why. Not stated as a single "why" comment beyond the CB-402 changelog note that opencode "proves the PeerLauncher SPI is genuinely provider-neutral rather than Claude-shaped" (see the existing Pin an opencode endpoint entry). The partition-by-kind design itself reads as the natural consequence: profiles fully own their backend, so routing is a lookup, not a branch.

Validated at config load since CB-604 (2026-08-16). An unrecognized kind: now refuses to start, naming the profile, the bad value, the accepted set, and what would otherwise happen:

refusing to start: profile(s) [gemini=opencod] set an unrecognized kind — accepted values are
claude-code, opencode (case-insensitive); an unrecognized kind would otherwise fall back to the
claude-code adapter and try to launch a program named after the typo.

Before that fix, kind: opencod was silently accepted, normalized, and — because it did not equal "opencode" — routed into the claude-code adapter bucket. With argv: also unset, the launch command defaulted to List.of(kind), literally the misspelled string, because the argv default only special-cases the exact string "claude-code". Nothing caught it before the spawn failed.

Gotcha. The composite constructor separately refuses two adapters claiming the same profile name ("worker profile '…' is claimed by two peer adapters"), so that failure mode has always been loud.

Checking for the same shape elsewhere found three more fields, all fixed the same way by CB-606 (2026-08-16) — see Every config value is checked against its valid set below.


Supervise the daemon on Linux (systemd)

What. deploy/fleetd.service is a systemd user unit (not system-level — "fleetd drives the user's herdr, not a system daemon") that runs java -jar target/fleetd.jar fleetd.yaml, restarts on failure (Restart=on-failure, RestartSec=10s, capped at 5 restarts per 120s via StartLimitBurst/StartLimitIntervalSec), waits on herdr.service only advisorially (Wants=, not Requires=, so a herdr restart never takes fleetd down with it), and applies a sandboxing profile (NoNewPrivileges, ProtectSystem=strict, ProtectHome=read-write, etc.).

On. Manual install: copy to ~/.config/systemd/user/, edit ExecStart/WorkingDirectory/ Environment, then systemctl --user daemon-reload && systemctl --user enable --now fleetd. Not currently the live supervision target — the unit file's own header comment says the dogfooded daemon runs on macOS under launchd; this unit is for the Linux gateways CB-308 introduces.

Why. Not stated beyond the practical need: a per-host gateway topology (CB-308, noted elsewhere in the wiki) needs Linux hosts, and those need systemd rather than launchd.

Gotcha. The unit hard-codes Environment=PATH=… with an explicit comment explaining why: "systemd does not source a login shell, so without it the daemon — and every worker — gets a bare default with no JDK/Maven" (same defect class CB-511/CB-594 already fixed for PATH on launchd). For secrets, the unit's own comment says plainly "Secrets are NOT set here" and points the operator at a systemctl --user edit fleetd drop-in or an EnvironmentFile=. It answers the login-shell defect only for PATH, not for WORKER_GITEA_TOKEN/AI_GATEWAY_TOKEN.

Systemd answer — YES, it has the underlying defect, undocumented for those two variables specifically. Comparing to the launchd side: launchd had the identical problem (deploy/dev.ltms.fleetd.plist's own comment: "launchd does NOT source .zprofile/.zshrc") and it was fixed by scripts/fleetd-launchd-wrapper.sh, which execs zsh -l so ${SHARED_ENV}/tools/secrets.sh gets sourced — its own header comment names exactly WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN as the two secrets this closes the gap for. The systemd unit has no equivalent wrapper and does not source that file at all. Its own comment mentions only FLEETD_API_TOKEN as the secret to add via drop-in — it never names WORKER_GITEA_TOKEN or AI_GATEWAY_TOKEN. An operator following the unit file's own guidance verbatim would set FLEETD_API_TOKEN and stop there: the daemon boots fine, and the failure surfaces only later as a worker that cannot open a PR or a gateway profile returning 401 — the exact failure mode CB-594 fixed for launchd, left open here. Fleetd.reportRequiredSecrets (called at the top of main()) does log which secret env-var names resolved on either platform — but that log line can only tell the truth about the names it names; it doesn't fix the sourcing gap, and the unit file gives the operator no prompt to look for it.


Every config value is checked against its valid set

What. A config value the daemon does not recognize now refuses to start, naming the field, the value, the accepted set, and what would otherwise have happened. Four fields carried this defect and all four are covered: a profile's kind: (CB-604), auth.mode, a profile's placement:, and the top-level placement: policy (all CB-606).

refusing to start: auth.mode=toekn is not recognized — accepted values are loopback-trust,
token (case-insensitive); an unrecognized mode would otherwise silently fall back to
loopback-trust, which authenticates nobody.

On. Always on; nothing to configure. Every valid value still works case-insensitively, and every absent value keeps its old default — kind: → claude-code, auth.mode → loopback-trust, a profile's placement: → tab, the policy → fixed.

Why. All four fields had the same shape: the value was lower-cased in a compact constructor, then compared against exactly one string. Anything else fell through to the other branch silently. So a typo was accepted, did nothing the operator asked for, and reported no error.

auth.mode is why this is a release blocker rather than tidying. A typo of token behaved as loopback-trust, and validateAuthExposure() only fires on a non-loopback bind — so on a loopback bind, the common case, the mistake was invisible end to end. The daemon started cleanly and authenticated nobody, while the operator's config said token and the operator believed it.

Gotcha — the general lesson, not the instances. CB-604 fixed kind: alone. The brief for it also asked the worker to report any other field with the same shape and explicitly not to fix them. That one question found the other three. When you fix a defect of shape, ask what else has that shape — the instance you noticed is rarely the worst one. Here the field nobody had filed a ticket about was the field that turned authentication off.

Two implementation notes worth keeping. The top-level policy was already validated, by PlacementPolicies.fromName — but lazily, at first spawn, through the supplier in CompositePeerLauncher that re-reads live config. So a bad name still started a daemon that looked healthy and failed much later. It is now called eagerly at load as well, which also covers ConfigRef.reload(). And the check calls fromName rather than copying its list, so there is still one source of truth for what a valid policy name is.


Old work-in-progress snapshots clean themselves up

What. The daemon now deletes a refs/wip/<branch> snapshot once it is safe to — and GET /members reports how many are left and roughly what they cost:

{ "workers": [ ... ], "wipRefs": { "count": 12, "costBytes": 4183042 } }

The rule has two conditions and both must hold:

  1. The snapshot commit's tree is already reachable from main. The content exists somewhere else, so deleting the ref loses nothing.
  2. The snapshot is older than 24 hours.

Every deletion logs the ref name, the commit sha, the age, and the exact git update-ref command that puts it back:

pruned snapshot ref refs/wip/worker/cb578c-92885c-1 commit=4a91f0c (age 51h): its tree is
already reachable from main, so the work is preserved; recover from reflog via
git update-ref refs/wip/worker/cb578c-92885c-1 4a91f0c

On. Always on; nothing to configure. The sweep rides the existing session reaper loop and runs every 6 hours, plus once on the first pass after a restart. A fleet that has never snapshotted anything sees no change, and wipRefs is simply absent from /members when no worktree session has told the daemon which repo to look in.

Why. CB-578 stage C commits a dirty worktree to refs/wip/<branch> on release, so a worker's uncommitted work is never lost. Nothing deleted those refs. A ref is a garbage-collection root, so every snapshot pinned its whole tree and git gc could never reclaim any of it. Branch names are unique per session, so the refs pile up rather than replace each other. It was never an error — it would have shown up as a slow git gc and a large .git, months later.

Reachability is the safety floor, not the age. A snapshot exists because the work was not committed anywhere else, so a plain time-based sweep would throw away the only copy — exactly the failure the snapshots were built to stop.

Gotcha — the sweep shipped dead, and every test still passed. SessionReaper held the last sweep time in a long that started at Long.MIN_VALUE to mean "never yet". That sentinel cannot be compared by subtraction: System.nanoTime() is positive, so now - Long.MIN_VALUE overflows to a large negative number, the "has 6 hours passed?" gate read it as swept moments ago, and the method returned before the assignment that would have fixed the field. The sweep never ran once, for the life of the process, with nothing in the log to say so.

The unit tests missed it because they all called the sweep method directly and walked around the gate. Two lessons worth keeping: a "never yet" sentinel belongs in its own boolean, not in a magic value of the same field, and a test that calls the seam does not prove the caller reaches it. The test that now guards this starts the real reaper loop and requires a prune call to arrive.

Second gotcha — costBytes is a rough figure, not disk usage. It sums the uncompressed sizes of every blob in every snapshot's tree. Objects shared between snapshots, or shared with main, are counted once per snapshot. So it reads much larger than the space actually reclaimable, and it is useful for spotting growth over time, not for capacity planning. Nothing pushes refs/wip/* to the forge; they are deliberately local and invisible to git branch.


Set a member's auto-compact window

What. A profile can set autoCompactWindow: <tokens>, the context size at which a spawned member compacts on its own. For a claude-code member it becomes --autocompact <tokens> at launch. For an opencode member it becomes a per-model provider.<p>.models.<m>.limit.context in the generated config. Unset ⇒ the backend's own default (Claude Code's built-in auto-compact, opencode's compaction.auto).

On. Per profile in fleetd.yaml, e.g. autoCompactWindow: 250000. The value must be in [100000, 1000000] — the band Claude Code's flag accepts — and the daemon refuses to start with a profile outside it, naming the profile. For opencode the key only takes effect when model: is in provider/model form (a warning is logged otherwise), because the limit is written under that exact provider and model.

Why. A member that runs out of context dies mid-turn and loses its fleet_reply, so the report is gone even though the work was done. The backends already auto-compact, but the window was not in the operator's hands — a long delegation on a big model could compact too late. This makes the trigger a per-profile choice, so an expensive profile can be told to compact earlier than a cheap one.

Gotcha. The two backends take the window by different means. Claude Code's --autocompact is an absolute token count and outranks env and settings. opencode has no absolute "compact at N" knob at all; the only lever is the model's limit.context, and the generated config also writes limit.output (16384) because opencode's schema requires both. On a gateway profile that is a partial override of the model's real limits — confirm the live member still answers after a spawn. Related: opencode members run with --auto.


Leads on different hosts talk over a shared broker

What. Two leads that share no herdr session — including leads on different hosts — can message each other by a stable coordination id. A coordinator: block gives this daemon one id and a durable AMQP mailbox lead.<id>.inbox on a shared vhost. fleet_send{coordId: <peer's id>, content} publishes to the peer's mailbox; the peer's daemon drains its own mailbox and types the message into its local lead's pane. fleet_list reports this daemon's own coord-id, so an operator knows the address to hand a peer.

On. Add a coordinator: block with selfId: (this daemon's globally-unique id) and either uri: or uriEnv: (the shared broker, ideally a separate vhost from the worker reply inbox). With no block, nothing opens and lead-to-lead stays the existing same-host pane path. If the block is present but selfId is blank, the daemon logs a warning and skips the mailbox — it does not crash. If the broker is unreachable at boot, the daemon warns and carries on with the feature off.

Why. Leads talk to each other widened addressing, but only for leads that share one herdr session — the address was a bare herdr terminal id, which means nothing on another host. Cross-host coordination needs an address that means the same thing everywhere and a transport that is not a local pane. The shared broker is the one channel both hosts already reach, so a coord-id plus a per-lead mailbox is the smallest thing that carries a message across the gap.

Gotcha. A publish to a coord-id whose mailbox no daemon owns is reported as unroutable, not silently dropped — the peer must be configured and running for the message to land. Delivery into the local lead's pane is status-gated exactly like a worker reply: a message waits, unacked, until the lead pane is idle, blocked or done, so it is never typed in mid-turn. A daemon with several leads must name one lead after coordinator.selfId, or it cannot tell which pane a peer's message is for and holds it. This is transport only — there is no peer discovery yet; a lead is addressed by a coord-id it was already told.


A lead can read its own held peer mail

What. fleet_poll{coordId: <your own coord-id>} returns this daemon's held lead-to-lead messages in full, and does not ack them. The messages stay held, so reading twice returns the same bodies. It reads only your own mailbox: passing a peer's coord-id is refused with a message saying so, because no route exists for reading another daemon's mail.

On. Nothing to configure beyond the coordinator: block that lead-to-lead messaging already needs. Pass your own selfId — fleet_list reports it as coordinator.selfId — as fleet_poll's coordId. fleet_list also now reports heldCount and heldDurable beside held[]. heldDurable is derived from the queue's declared durability and the consume ack mode, not written as a constant, so it can actually report false if either changes (fleetd #440). It is one boolean over two inputs, so a false does not say which of the two moved.

Why. A lead could see that it had peer mail but not read it. fleet_list's held[] carries only an 80-character preview, on purpose, so a constantly-called roster scan never dumps a coordination body. fleet_poll{target} drained the wrong store — a worker's reply inbox, not the coordinator mailbox — and returned an empty list with no error. The only other route, fleet_ack, looked destructive but is not: it never touches the coordination channel at all, and reports acknowledged <msgId> anyway (fleetd #437). So the full body was reachable in principle and unreachable in practice, and the busiest leads were the ones most affected: delivery into a lead's pane is status-gated, so a lead that never goes idle never drains its own mailbox.

Gotcha. This is primary-only, under a new authorization action COORD_READ. That includes refusing an architect, which holds the ordinary READ action today. The reason is that READ's grant rests on the roster carrying no secrets, and a lead-to-lead body is not the roster — it is where leads discuss host shapes, credential states and unmerged work. Also do not read mailbox.pending: 0 as an empty mailbox: pending counts only broker-ready messages, so a healthy mailbox holding real mail reports zero. heldCount is the honest number.


Fleet health can finally reach its fault states

What. FleetHealthMonitor ticks on its own timer, well away from the 250ms delivery poller. Each tick makes one agent.list call plus one roster snapshot, then classifies every member through the pure FleetHealth.decide. There are 14 health states and 9 of them are faults. A fault is logged when the state changes, never once per tick. Two terminal states, GONE and NEVER_READY, also call failTarget (MessageService.abandon), which fails every ticket still waiting on that member.

On. The health: block in fleetd.yaml: enabled: true, intervalSeconds: 30 (floor 15), workingSuspectAfterSeconds: 600 (floor 300 — how long a BUSY member with no activity must sit before it is STALL_SUSPECTED), and notifications.mode. fleet_list reports healthCoverage: off, detection-only, or full. It is full only for mode: webhook.

Why. The monitor shipped half-wired and nothing said so. tick passed a hardcoded NOT_YET_OBSERVED = false for 7 of the 12 HealthSnapshot fields, so 8 of the 9 fault states could never be reached — only TURN_BOUNDARY_LOST could fire. GONE and NEVER_READY were among the dead ones, which means CB-580's repair never ran: a ticket waiting on a member that had vanished still hung to the 30-minute async timeout. Everything compiled and every test passed, because each test called the classifier directly and walked around the caller. Three units fixed it: CB-640 added the message-layer facts (hasQueuedDelivery, hasStrandedReply, hasOrphanedDelegation), CB-641 added the herdr and clock facts (present, targetNotFound, controlLinkDown, readinessGraceElapsed, stalled), and CB-643 joined them and deleted the placeholder. The rule that came out of it: a snapshot field with no publisher is a dead state, and nothing else will tell you.

Gotcha. Three things to know before you read the log.

  • With notifications.mode: disabled the coverage is detection-only. Every fault the monitor finds goes to the log and nowhere else. The daemon prints a warning about this at startup.
  • A member that is already faulty when the daemon starts logs its fault once, on the first tick, and then stays quiet. Grepping recent lines can therefore show nothing while the fault is live.
  • DELEGATION_ORPHANED needs the fact to hold for two consecutive ticks. An async ticket exists for a moment before its virtual thread opens the rendezvous waiter, so for that instant nothing is accepted or queued behind it and a single tick would report a fault that clears immediately. The gate costs one interval of delay on a real orphan.

/fleets-status — every fleet that shares one broker

What. A primary-side skill in .claude/skills/fleets-status/. It reports the state of every fleet that shares one LavinMQ instance, in three tiers: (1) the local daemon — process, deployed jar against HEAD, /healthz, fleet_list with capacity and healthCoverage, and the WARN/ERROR lines since the last fleetd listening marker; (2) the broker itself over its management API — each vhost's queues, depths and consumers; (3) cross-fleet lead coordination, read from the lead.<coordId>.inbox queues. Every tier ends with what it cannot see, and a fleet it cannot reach appears in the report as unknown with the reason, never as an omission.

On. Type /fleets-status. Tier 1 always runs. Tier 2 runs only when LAVINMQ_MANAGEMENT_USER and LAVINMQ_MANAGEMENT_PASSWORD resolve in a login shell.

Why. There was no way to ask one question about the whole fleet estate. A lead could see its own daemon and nothing else, and the remote daemon's REST port is not reachable from here. The broker is the one thing both fleets touch, so its management API answers for a host you cannot log in to. The sharpest fact it gives you: a vhost with queues but zero consumers means that fleet's daemon is down while its durable state survives.

Gotcha. Tier 2 is blocked today. The AMQP user in LAVINMQ_URI connects fine on 5672 but gets HTTP 401 from the management API on 15672 — an AMQP connection is not monitoring access. An operator must create a separate read-only user with the monitoring tag, allowed on both vhosts, and store it under the two names above in ${SHARED_ENV}/tools/secrets.sh. Also: the skill forbids printing LAVINMQ_URI, because the password is inside it. Output that could carry it is redacted with sed -E 's#://[^@]*@#://<redacted>@#g', and the g is not optional — without it a line holding two URIs leaks the second one, which scripts/redeploy-fleetd.sh --check prints.

memberHerdrSocket — run members on a second herdr daemon

What. A new top-level config key. Set it to a second herdr API socket path and every member is spawned on that daemon, while everything the lead does stays on the first one. Absent (the default), both are the same client and nothing changes. Inside the daemon a HerdrRouter owns the split: it hands each consumer the lead client, the member client, or — for consumers that see both kinds of traffic — routes per target id.

memberHerdrSocket: /run/fleet/members.sock   # omit for one daemon (the default)

On. Add the key and restart. It is read once at startup, so it is a deferred key, not a hot one.

Why. herdr forks every pane as its own OS user, and there is no user, uid or run-as parameter anywhere in its socket API — measured against herdr 0.8.0's own docs, not just our client. So the only way to make a member run as a different user than the operator is a second herdr server started by that user. That matters because a member currently reads the operator's ssh key straight off the filesystem: blocking SSH_AUTH_SOCK does not stop it, and no environment scrub can, because the scrub removes variables and the key is a file (#184). A separate OS user is the control; this key is the fleetd half of it. Design and measurements are in #185.

Gotcha — the feature is still not ready to switch on. #185 keeps the running list. Two of the blockers were cleared on 2026-08-31; the rest stand.

Cleared. Pane ownership lives only in memory, so a restart used to leave every surviving pane unowned and fleet_stop refused it. CompositePeerLauncher now probes: it asks each daemon which one knows the pane, routes to the single match and caches it, treats no match as already stopped, and refuses only a genuine two-daemon collision — with a message that says that is what happened. One unreachable daemon no longer breaks the probe for panes owned by a healthy one.

Still open. herdr creates its API socket mode 0600 and umask does not change it, so cross-user access still needs a post-start chmod g+rw — a race on every start, and an upstream ask that has not been made. Pane ids are still not daemon-qualified: fleet_stop{paneId} takes a bare w1:p1, so the probe makes a collision fail loudly instead of silently closing the wrong pane, which is a safety net, not the design fix. hostEnvNames still reads fleetd's own environment, which is the wrong environment under a second user.

Worktree ownership now has an answer — worktreeGroup, below — but a new blocker replaced it, and it is the harder one. ClaudeCodeLauncher writes the role charter with Files.createTempFile, and on macOS the system temp directory is per-user:

$ ls -ld "$TMPDIR"
drwx------@ 219 dai.ha  staff  /var/folders/wf/…/T/

A member running as another user cannot even traverse it, so it never reads its charter — and the charter carries the reply contract. EnvAllowListScrub has the same problem: it generates the member's ZDOTDIR there, and the member's own login shell has to read it. No chmod fixes a per-user 0700 directory safely; both writes have to move somewhere both uids can reach. Two launcher writes (CLAUDE.local.md and <commonDir>/info/exclude) also land at spawn time, i.e. after the group fix-up has run — harmless today because the member only reads them, but nothing says so.

The shape to watch for. herdr's workspace, tab and pane ids are per-daemon sequential counters. Two daemons really do both hold w1:p1, pointing at different panes owned by different users — measured, not theorised. Four separate defects came from code that assumed one daemon, and every one of them passed the whole test suite, because with the key unset both clients are the same object. When reviewing anything on this path, ask each herdr call one question: which daemon does this go to, and is that the daemon that owns the thing it is asking about?

Observable change even with one daemon. GET /healthz requires both daemons to answer before it returns 200 (with one configured that is the single ping it always was). GET /sessions merges workspace.list across both.

Since 2026-08-31 the 200 body also reports the member daemon's version and protocol, under its own member key, and adds protocolMismatch: true when the two protocol numbers differ:

{ "status": "ok",
  "herdr":  { "version": "0.8.0", "protocol": 3 },
  "member": { "version": "0.7.1", "protocol": 2 },
  "protocolMismatch": true }

The member key and protocolMismatch appear only when a second daemon is configured, so a single-daemon body is byte-identical to before. herdr still carries the lead daemon's values, deliberately: scripts/redeploy-fleetd.sh and scripts/rename-checkout.sh both read this endpoint, and folding the two daemons into one key would hide a mismatch from whichever reader looks at only that key.

This closes a real trap. The member daemon was pinged and its answer thrown away, so a member herdr running a mismatched protocol left /healthz green while every spawn failed — and it is the member daemon's protocol, not the lead's, that decides whether a spawn works.



worktreeGroup — let a member under another OS user write its worktree

What. An optional top-level key naming an OS group. When set, every provisioned worktree and the repository's git store are made group-writable, so a member spawned under a different OS user (see memberHerdrSocket) can write its own worktree, its per-worktree git metadata, and the objects its commits create. Absent — the default — nothing runs and no process is spawned.

worktreeGroup: fleet-workers   # omit for the single-user default

On. Add the key and restart. The operator running fleetd must already be a member of that group, or every provisioning spawn fails loudly, naming the group.

Why. GitWorktrees shells git worktree add, so it runs as fleetd's uid and everything it creates is owned by the operator at 0644/0755. A member running as another user cannot write any of it. That user is the whole point of #185: a member currently reads the operator's forge ssh key straight off the filesystem, and no environment scrub can stop that, because the key is a file (#184). This key is the git half of the fix.

It isolates credentials, not the repository. A member in the group can still write the operator's git objects and refs. Say that plainly rather than implying more — the two uids share one repository by design, which is what keeps the lead's refs/wip/* safety net working.

Gotchas, all three found the hard way.

The share pass must run after overlayParity, never inside add(). overlayParity copies more files into the worktree after add() returns, so sharing any earlier leaves every overlay file operator-owned and unwritable — with the whole suite still green. A test pins the order.

Missing paths are skipped, not passed to chgrp. .git/logs does not exist with core.logAllRefUpdates=false or before the first ref update, and packed-refs does not exist until refs are packed. chgrp on a missing path exits non-zero, which would fail every spawn with a message blaming a group that is fine.

The git directory is asked for, never assumed. <repoRoot>/.git is a file, not a directory, when the checkout is itself a linked worktree — which is what this code creates for every member. It resolves git rev-parse --git-common-dir against repoRoot, because git answers relatively for an ordinary checkout and absolutely for a linked one.

Cost. The fix-up re-runs on every spawn, deliberately. core.sharedRepository=group governs only what git writes afterwards; the walk is what covers everything already on disk. Three walks of the object store per spawn — about 3000 files in this repo, well under a second, growing with the repo. It is not redundant work to optimise away.


opencode session ids come from opencode.db

What. fleet_list reports agentSessionId for an opencode member, and fleet_spawn{resumeSessionId} can therefore relaunch one onto its prior conversation. Automatic — no knob.

It took two fixes, not one, and the first one alone did nothing visible. Reading opencode.db (below) was necessary but not sufficient: SessionManager asked the handle for the id once, at spawn, and froze it into an immutable record. opencode has not written its row at that moment, so the frozen answer was always null and the roster never showed one. The handle's own javadoc said "the caller re-calls later" — no caller did, and SessionManager did not even keep the handle. See Late-resolved member ids for the second half.

On. Nothing to turn on. It needs opencode's own storage root (~/.local/share/opencode) to be readable, which it is.

Why it is an entry at all: it was silently dead. opencode moved its session store from a one-JSON-file-per-session tree to a SQLite database in January 2026, and OpenCodeSessionDiscovery went on scanning the frozen tree. It therefore returned null for every member, forever. agentSessionId was never known, so resumeSessionId was unreachable for every opencode profile — most of the fleet — and nothing said so, through 57 member spawns. It now reads <storageRoot>/opencode.db, matching the session table's directory column against the worker's cwd and taking the highest time_updated.

The gotcha is what the silence cost. A discovery that finds nothing and a store that is empty look identical. That is why a missing database is now a WARN naming the path searched — it means the layout moved again — while a present database with no matching row stays at DEBUG, because that is the normal answer right after a spawn.

Read-only, and pinned as such. opencode may be running and writing that database (WAL mode, 841MB here), so the connection is opened with SQLITE_OPEN_READONLY. Beware the obvious test for this: making the file unwritable proves nothing, because SQLite silently downgrades a read-write open of an unwritable file to read-only. The test that works asks the connection to INSERT and requires the refusal.

One dependency. org.xerial:sqlite-jdbc, which ships bundled native libraries and takes the shaded jar to about 28.6MB. Chosen over shelling out to the sqlite3 binary so there is no undeclared external-binary requirement on a headless host.


One worker reply settles exactly one ticket

What. When a worker's fleet_reply arrives with no send waiting for it, the bridge holds it and later hands it to one delegation — never to several. If it cannot tell which one a held reply answers, it queues the reply instead of guessing. Automatic — no knob.

On. Nothing to turn on.

Why it exists. Two separate paths used to assume a target had at most one delegation open, with no guard for the case where it had two:

  • fleet_send{turnId} (the answer to a fleet_ask) waits only for its own bounded MCP window — 25s by default, 120s at most. A resumed turn can easily run longer than that. When the window expired, the worker's real reply had no waiter left, so it was held instead. The ticket stayed PENDING until fleet_stop, which then reported "the worker session was released before it replied" — with a worktree, a branch and a snapshot commit, so it read like lost work. The reply had in fact arrived. Fixed first: a held reply now completes its own ticket (388aba7).
  • Teardown then drained that held reply once and reused it for every open ticket on the target. Two open tickets both came back DONE with the same text, one of them an answer the worker never gave for that delegation. Fixed second (966c58a): the reply goes to the oldest open ticket and every other one keeps the ordinary failure path.

The gotcha: a confidently wrong status costs more than no status. Both defects produced a plausible answer, not an error. A lead that trusted the first one would go recover a snapshot of work that was already merged; a lead that trusted the second would act on a report its worker never wrote. That is why the ambiguous case queues the reply and logs a WARN naming every candidate ticket rather than picking one.

Known limit, on purpose. Only the second defect's broad case is reachable today. Two open tickets on one target is reachable and was a real bug. The same ambiguity in the reply path, and the combination of two open tickets with a held reply, are both blocked by existing guards — one send holds a target's lock and its waiter together, so a held reply and a second open ticket cannot exist at the same instant. The guards stay as defence in depth, and MessageService.abandon's javadoc records why no test pins them: a test can force the state only by breaking that lock-and-waiter invariant from outside the class, which no caller does.

Late-resolved member ids

What. A member whose backend cannot name its own session at spawn time gets its agentSessionId filled in later, as soon as the backend writes it. fleet_list, fleet_status and the REST roster all report the current answer rather than the one frozen at spawn. Automatic — no knob.

On. Nothing to turn on. It only does work for a member whose id is still unknown.

Why it exists. agentSessionId is the handle fleet_spawn{resumeSessionId} needs. It was read once at spawn and never again, which is fine for claude-code (it knows its session id up front) and useless for opencode (it does not). The result was a documented, tested feature that could never produce a value for most of the fleet. This is the second half of the opencode.db fix.

The gotcha: the resolving read is not the cheap one. Asking an opencode handle for its id opens a SQLite database — around 841MB here. roster() is the roster supplier for the heartbeat loop and the health monitor, so resolving there would open that database on every tick, for every unresolved member, forever for a member whose row never appears. So there are two reads:

Method Resolves? Used by
roster() no LeadHeartbeatLoop, FleetHealthMonitor, placement and exhaustion checks, the /metrics scrape
rosterResolved() yes fleet_list, the REST roster — the surfaces that actually report the id

get(paneId) (behind fleet_status) and the teardown path resolve too; both are caller-driven, not timers. A test asserts the plain roster() never calls the handle again, so the split cannot be quietly "tidied up" later. Once an id is found it is stored and never looked up again.


GET /members reports its rows under members

What. The REST roster's response body uses the key members, matching its path. The old key workers is still emitted as a deprecated alias carrying the same rows.

On. Nothing to turn on.

Why it exists. The route was renamed /workers → /members, and the body key was not renamed with it. A caller reading body["members"] — the obvious guess given the path — got an empty list, which is a perfectly valid answer meaning "this fleet has no members". So a client reported an idle fleet and nothing anywhere contradicted it.

The gotcha: do not just rename the key. The REST surface is the out-of-band path a lead falls back to when its MCP mount drops, and at least one client reads workers today. Renaming outright would break that client the same silent way, in the other direction. Both keys are emitted, and a test asserts they carry identical rows so the alias cannot drift. Drop workers once nothing reads it.


The credential gap detector admits when it cannot see

What. With memberHerdrSocket: set, the memberCredentials gap report stops drawing conclusions about member panes and says the gap is unknown, not clean, naming the config key that made it unknowable. One WARN per launcher, not per spawn.

On. Only in two-herdr mode. With memberHerdrSocket: absent — the default — every log line on this path is byte-identical to before, pinned by a test.

Why it exists. The detector enumerates fleetd's own environment, on the assumption that the member pane's login shell exports the same set. Under memberHerdrSocket: panes are routed to a second herdr, and fleetd has no channel to confirm what OS user that herdr runs as — it may be the same user or a different one, with a different $HOME and a different secret store. Either way the assumption is no longer something fleetd can check, so the old output was reporting on a process it could not see, and could call a gap clean while the member exported something dangerous. There is no channel to read another user's environment, so "unknown" is the only honest answer.

Note the shape of that sentence: the fix is not "the member runs as a different user". Saying that would be the same mistake pointing the other way — fleetd cannot confirm a different uid any more than it can confirm the same one (fleetd #269). What changed is the claim, from a false certainty to an accurate unknown.

The count line beside it says whose environment it counted. The INFO line member credentials: allowed N of M reads as a statement about the member's pane, and under memberHerdrSocket: it is not one — the counts come from fleetd's own process. Since fleetd #276 the line says so in place, because the caveat used to live only in a javadoc the operator reading fleetd.out never sees. The counts themselves are unchanged and still real. With memberHerdrSocket: absent the wording is byte-identical to before.

The gotcha: this fixes the report, not the control. Two related defects in the same area are open under #213 — the ZDOTDIR scrub decides whether it can run by reading fleetd's $SHELL, and writes its generated file into fleetd's own TMPDIR, which on macOS is a per-user 0700 directory another uid cannot even traverse. In two-user mode the scrub can therefore be silently absent while the daemon believes it ran. Treat memberCredentials as unverified in that mode until #213 lands.


Backfill status

This page was started after the fact, so it is not yet complete. Entries above are written from verified behaviour.

The CB-595 backfill (2026-08-16) cleared the largest gap: the ~14 operator-facing tickets shipped between v1.0.0 and now that had landed nowhere. Those entries were written from a read of the code on main, not from commit messages, and the three keys that changed meaning were pulled to the front because an upgrading operator meets them first.

The second CB-595 pass (also 2026-08-16) cleared the five areas that pass had left open: /metrics and /healthz, bearer auth and the bind fail-fast, the authz table and audit log, systemd supervision, and multi-profile kind: routing. Every claim in those five was traced to a file:line on main, and two of them turned up defects the catalogue work was not looking for — see below.

Still to catalogue:

  • the reply push loop and its nudge budget (CB-307 step 2, CB-590, CB-598) — the budget is now tracked per pending item, not per lead and not per source. A newly-arrived item keeps its source eligible even when an older, still-undrained item has used up its own budget. The practical consequence to write up: the cap bounds nudges about one item, so a lead with a steady arrival of new work keeps being nudged — which is correct, but is not what the knob's name suggests

Role agent definitions — how a member learns its role (CB-617)

What it does. The architect, dev and reviewer contracts live as files in the repo, at both .claude/agents/<role>.md and .opencode/agent/<role>.md. A member's worktree is a checkout of this repo, so it arrives with its own contract. At spawn the launcher passes --agent <role> when the matching file exists.

The knob. Nothing to turn on. The role you pass to fleet_spawn selects the file. Absent file ⇒ no --agent flag and the member still spawns, so this degrades rather than breaks.

Why it exists. The contract used to travel as an inline argv flag, --append-system-prompt <text>. herdr refuses to shell-encode a multi-line argument, so every claude-code member with a configured role charter failed to spawn (CB-616) — measured as an architect that started on sol and died on opus. opencode escaped only because it already wrote its charter to a file. Files fix that for both backends and make the contract reviewable in git.

The gotchas.

  • Both directories, every time. Neither tool reads the other's: a .claude/agents/ file is invisible to opencode. Measured — a probe agent in .claude/agents/ never appeared in opencode agent list. The two copies of a role must carry identical bodies or the backends work from different contracts.
  • Never put model: in an agent file. Both CLIs honour it only when no launch flag is passed, and fleetd always passes one (--model for claude-code, -m for opencode). Measured both ways: an agent pinned to openai/gpt-5.6-terra run with -m opencode/x-preview-f-free reported > pin · x-preview-f-free. A model: here is silently overridden on every spawn. The model stays in fleetd.yaml, which also keeps role and backend as the separate axes the role pools need.
  • Both charters travel in one file, or Claude Code will not start (CB-618). The CLI refuses the two prompt flags together: Error: Cannot use both --append-system-prompt and --append-system-prompt-file. Please use only one. The first cut of this feature put the role charter on the file flag and left the reply charter inline, and every claude-code spawn with a role charter died at launch — reported by fleetd as spawn_timeout, which hides the cause completely. So when a role charter is present, both go in the one file with the reply charter last: last is where the reply rule must sit, because it is the rule that must survive. A member with no role charter keeps the inline flag, which is also the only form that reaches a member with no repo checkout.
  • A .claude/agents/ file needs name: in its frontmatter (CB-618). With only description: the file is skipped in silence and the launch fails with --agent 'architect' not found. Available agents: claude, Explore, .... The OpenCode files take their name from the filename and need no such key, so the two formats are not interchangeable even though the bodies are identical.
  • Measured together, on the real binary. claude --agent architect --append-system-prompt-file <both charters> answered ROLEOK, REPLYOK, yes — agent body, role charter and reply charter all in force at once. Both defects above shipped green because every test read the argv fleetd builds and none ran the binary that has to accept it (see #113).

Found while cataloguing, not by looking for bugs

Two of these are filed as their own tickets. They are recorded here because both are the same shape: a configuration that is accepted, does nothing useful, and reports no error.

  • kind: is never validated. A typo such as kind: opencod is lower-cased, accepted, and — not matching "opencode" — routed to the claude-code adapter. If argv: is unset, the launch command defaults to the misspelled string itself rather than claude. No config-load check catches it. See Multi-profile routing above.
  • The systemd unit has the login-shell secret defect that launchd's had. deploy/fleetd.service fixes PATH explicitly and names only FLEETD_API_TOKEN as a secret to add. It never sources the secret store, and never mentions WORKER_GITEA_TOKEN or AI_GATEWAY_TOKEN. An operator following the unit's own guidance gets a daemon that boots cleanly and members that cannot open a PR. launchd got scripts/fleetd-launchd-wrapper.sh for exactly this; the Linux unit has no equivalent.

One more piece of drift worth knowing while reading these entries: AgentControl's class doc says herdr protocol 19 (herdr 0.8.0), while HerdrClient's still says protocol 14 (0.7.0). Nothing compares the protocol number herdr reports against what fleetd actually needs — which is precisely how /healthz once went green while every spawn failed.


A launch command that cannot fit is refused

What. Before starting a member, fleetd measures the command it is about to hand herdr. If it cannot fit the pane's line, the spawn is refused straight away with the byte count and the longest argument named. Automatic — no knob.

On. Nothing to turn on.

Why it exists. herdr does not exec a member's launch command. It types it into the pane, and a pty line buffer holds 1024 bytes (BSD/macOS MAX_CANON). Everything past that byte is dropped, and no layer says a word: herdr answers "agent started", the backend exits on the mangled argument it was handed, the pane closes, and the only symptom is the readiness gate timing out twenty seconds later with no reason at all.

That is not a theory. It took the whole claude-code half of the fleet down. The reply charter used to ride inline on --append-system-prompt, so a sonnet launch command was already 978 bytes. Adding --session-id <uuid> for #214 — fifty bytes — made it 1028, and the four bytes cut off the end turned --autocompact 250000 into --autocompact 25. Claude Code refuses that value, so every claude-code spawn died. The cut is at byte 1024 exactly; that was measured on a live pane, not assumed.

Two things changed. The charter now always travels as --append-system-prompt-file, which takes ~800 bytes of prose off the command line for good. And this guard catches whatever grows next.

The gotcha: the estimate is deliberately too big. fleetd cannot see how herdr quotes each argument, so every argument is charged its own bytes plus a separator and a quote pair. An over-estimate costs you a clear error at a length that was already unsafe. An under-estimate would let the silent truncation back in, and that is the failure this exists to prevent.

It already caught a second one. A profile with ideMcpUrl: set assembles 1084 bytes — over the limit before this ticket was written. The IDE mount was one config key away from the same silent death. It cannot ship that way now.


A failed spawn shows you the pane

What. When the spawn-readiness gate gives up on a member, the WARN carries the last status herdr reported and the tail of the pane, read before the pane is closed. Automatic — no knob.

On. Nothing to turn on.

Why it exists. The gate used to log only that it had timed out, which is true of every possible cause: a backend that never launched, a binary that rejected an argument and exited, a trust prompt waiting for a keypress, a login shell that hung. The pane holds the one copy of that answer, and the next line of code destroyed it.

Finding #220 without this took a live process sampler, a hand-built pty and a byte count. With it, the answer was one log line: ... --model claude-sonnet-5 --autocompact 25.

The gotcha: read it before stop(). The order matters and is easy to get backwards. stop(paneId) closes the pane, and a pane read after that returns nothing — the log would be just as empty as before, only slower.


Every claude-code member is resumable

What. Every claude-code spawn gets a session id, so fleet_list always reports an agentSessionId you can hand back to fleet_spawn{resumeSessionId}. Automatic — no knob.

On. Nothing to turn on.

Why it exists. A session id used to be minted only when the caller passed sessionName or resumeSessionId, and fleet_spawn treats both as optional. So an ordinary spawn minted nothing and agentSessionId stayed empty for that member's whole life. Unlike opencode, nothing resolves it afterwards: a claude-code session id is fixed at launch.

That made resume opt-in at spawn, while the moment you want it is after a member has done work worth keeping. By then it was too late, and an empty field in the roster read as "this backend does not support resume", which was not what was happening.

The gotcha: the binary takes a UUID and nothing else. claude --session-id rejects any other value at argument parsing — "Invalid session ID. Must be a valid UUID." — so the random UUID is required, not incidental. The one documented flag interaction is with -r, which claims the same id; a resume spawn passes -r and never mints.


An exhausted backend is quarantined even from a chrome-only pane

What. When a member's pane yields no usable assistant block, fleetd still classifies BACKEND_EXHAUSTED and BACKEND_ERROR from the raw screen — so an exhausted credential is quarantined instead of being recorded as a member that did nothing. Automatic — no knob.

On. Nothing to turn on.

Why it exists. The completion resolver checked for an empty scrape before it looked for an exhaustion line, and returned. lastAssistantBlock finds nothing whenever the pane carries no ⏺ marker and its first visible line is TUI chrome — a box border, a warning, the input box — because the boundary scan breaks on that first line.

The exhaustion branch is the only caller of the quarantine sink. So a backend that refused the turn for a usage limit was filed as "produced nothing", the credential was never quarantined, and fleetd kept handing work to a backend that could not run it. Each attempt looked like another member silently doing nothing. That is the shape of a usage limit taking out several members with the roster giving no reason.

The gotcha: it is a fallback, not a reordering. The empty-scrape failure was not moved or weakened, and a pane that yields a usable block never reaches this path. Matching the raw screen widens what the patterns can hit, including scrollback that is not this turn's output — and a false positive here quarantines a working credential. That is why the raw match runs only where the alternative was a generic failure with no information in it at all.


The credential scrub follows the member's own user

What. When memberHerdrSocket: puts member panes under a different OS user, the allow-list credential scrub uses that user's facts: memberLoginShell: decides whether the ZDOTDIR scrub can run, and the generated directory goes under worktreeRoot, shared read-only with worktreeGroup.

On. memberLoginShell: and worktreeGroup: — both required once memberHerdrSocket: is set. With memberHerdrSocket: absent, nothing changes: fleetd's own $SHELL still decides and the directory still goes to java.io.tmpdir, byte-identical to before.

Why it exists. The scrub read fleetd's own $SHELL to decide whether a member's login shell honours ZDOTDIR, and wrote the generated directory into fleetd's own java.io.tmpdir. Both describe the wrong process once panes run as another user.

The dangerous combination was fleetd on zsh and the member user on anything else. fleetd took the zsh branch, generated a ZDOTDIR, and set it on the pane; the member's shell ignored ZDOTDIR entirely, so no scrub ran — and because the zsh branch deliberately skips the sentinel overlay, the fallback never happened either. No protection at all, on the path fleetd believed was the protected one. Even with the member on zsh it still failed: java.io.tmpdir is mode 0700 on macOS, so another uid cannot even traverse it.

The gotcha: it degrades, it never refuses to spawn. If the member shell is not zsh, or worktreeRoot/worktreeGroup is missing, fleetd warns loudly and falls back to the sentinel overlay. A weakened credential control must not become an outage for an opt-in feature. The permissions are deliberately tight in the other direction: the directory is rwxr-x--- and each file rw-r----- — group read, no group write anywhere, no world bits. A member sources what it needs and cannot alter fleetd's own scrub.


An opencode member's config follows the member's own user

What. When memberHerdrSocket: puts member panes under a different OS user, the ephemeral opencode.json no longer goes into fleetd's own temp directory. It goes under worktreeRoot, shared read-only with worktreeGroup. Session discovery, which cannot work across users, now says so once instead of failing quietly.

On. memberHerdrSocket: plus worktreeRoot/worktreeGroup. With memberHerdrSocket: absent, both paths are byte-identical to before.

Why it exists. The launcher decided two paths from its own process. The config directory came from java.io.tmpdir, which is mode 0700 on macOS, so another uid cannot even traverse it. That file is the only way a member learns where the bridge MCP is. A member that cannot read it still starts and still holds a pane, but never becomes deliverable: fleet_send waits about 60 seconds on the readiness gate and then fails, and nothing in that message points at a temp directory.

Session discovery had the same shape with a different result. It read fleetd's own user.home to find opencode.db, but under this key that file lives under the member's home. It would find nothing, forever, and agentSessionId would stay null — which reads as "this backend cannot resume" rather than "fleetd looked in the wrong home".

The gotcha: this one refuses to spawn, where the credential scrub above degrades. The difference is what each thing is. A weakened credential control still has value, so the scrub falls back to the sentinel overlay. This config file is not a control — it is the member's only route to the bridge. So a missing worktreeRoot or worktreeGroup refuses the spawn and names the missing key. Nothing is lost by refusing: that member would not have worked either way, and the later failure carries no clue. Discovery is switched off under this key on purpose, with one WARN — an honest "not available" beats an answer read from the wrong directory.


The model opencode actually ran is read back and checked

What. After an opencode member's session row appears, fleetd reads the model opencode actually resolved and compares it to what the profile asked for. On a real mismatch it logs an ERROR naming both models and quarantines that profile's credential through the existing exhaustion path.

On. Automatic, for opencode profiles that configure a model:. claude-code profiles are not affected — they have no opencode session row, and the check is structurally unreachable for them.

Why it exists. opencode does not fail on an unknown -m. It silently falls back to a default. The xf profile named a model that had been withdrawn from the catalogue, so every xf member ran on a paid model at the expensive high variant while the profile was configured to be free. That was 97.6% of dev spawns for about a day.

The second harm is worse than the bill. xf declared no credentialId, because it was supposed to be free — so five concurrent members could hammer the shared paid credential while fleet_list reported the paid profiles as barely loaded. The accounting built to protect that credential was looking the other way. Nothing logged any of it; it was found because the operator happened to notice the member's own UI naming the wrong model.

The general shape is worth naming: fleetd set a backend option by flag and treated the process starting as proof the option took effect. #232 tracks the same assumption for autoCompactWindow.

The gotcha: the check runs late, and "unknown" is never a mismatch. opencode writes its session row only when the session is first persisted, so a check inside spawn() runs before the evidence exists and can never fire — an earlier attempt shipped exactly that and was closed. This one runs from the same late-resolve step that fills in agentSessionId, so a resolved agentSessionId is also the sign the model check ran.

The comparison is deliberately narrow, because a false positive stops all work on a credential and is worse than the bug it prevents. The profile string and the stored JSON are in different formats (openai/gpt-5.6-terra against {"id":"gpt-5.6-terra","providerID":"openai"}), and a profile may carry no provider prefix at all. So the id is always compared, the provider only when both sides have one, and absent, incomplete or unparseable evidence is UNKNOWN — never a mismatch, never a quarantine.

The bigger gotcha: for a long time this check did not run at all for most spawns, and said nothing (#267). checkModelMatch has exactly one call site, and it sits after the #249 gate that returns early when the spawn was given no fleetd-provisioned worktree — which the code itself calls "the ordinary, expected shape of the large majority of spawns". So the detector built after the xf incident was switched off for most of the spawns it existed to protect, silently.

Since #267 that case is no longer silent. Once per profile, when a model is configured:

opencode model-mismatch check (fleetd #175) cannot run for profile 'X': it was spawned
without a fleetd-provisioned worktree (fleetd #249), so its cwd may be shared with other
sessions and the actual model it is running cannot be safely told apart from a sibling's —
spawn with worktree:true to enable the check for this profile.

The check itself still does not run there, and that is deliberate. The model can only be read via actualModelForSessionId, keyed on the resolved session id. The only other lookup is sessionIdForDirectory, the "newest row for this directory" heuristic #249 exists to distrust: on a shared cwd it can return a sibling's row, so a sibling running a different, correctly-configured model would look like this profile's mismatch and quarantine an innocent credential. A false quarantine takes real capacity away on bad evidence, which is worse than failing to detect. Spawn with worktree:true if you want the check.

Worth naming as a shape, because it is not the same one as the entry above: a safety check rode on the same return value as an identity lookup. #249 tightened the identity question for good reasons and narrowed the safety check as a side effect. Neither change was wrong alone — the coupling was.

Every file a member must read follows the member's own user

What. Two more places where fleetd used to write a file into its own process's filesystem and then hand a member the path. The claude-code role/reply charter now goes under worktreeRoot instead of java.io.tmpdir, shared with worktreeGroup. And worktreeRoot itself is now made group-traversable where it is created, not only the worktrees inside it.

On. memberHerdrSocket: plus worktreeRoot/worktreeGroup. With memberHerdrSocket: absent, the charter path is byte-identical to before; with worktreeGroup: unset, the root is untouched.

Why it exists. This is the same shape as the two entries above, found twice more. The charter one became load-bearing without anyone noticing: the charter used to ride inline on --append-system-prompt, and the file was the exception. #220 made the file the only delivery path, always, to keep the launch command inside the pane's 1024-byte line. So under memberHerdrSocket: every claude-code member would have been handed a charter path it could not read. Measured against the real binary (claude 2.1.258), that is the loud failure — it prints Error reading append system prompt file: EACCES and exits 1 before it touches auth or the network, so the pane dies and the spawn fails at the readiness gate. Loud, but with no clue in the message.

The worktreeRoot one is the floor the other three stand on. A different uid needs the execute bit on every ancestor directory, so a root at 0700 makes every carefully-shared child unreachable — the member cannot read the opencode config, cannot reach its own checkout, and cannot read the credential scrub. Files.createDirectories respects the umask, so whether this bites depends on the umask of whatever shell started the daemon. It would work on the machine it was developed on and fail on the next one, and the failure looks like a member that never becomes deliverable.

The gotcha: both refuse rather than degrade, and both are invisible until you turn the key on. A charter is not a control, it is the member's turn contract — the rule that ends every turn with fleet_reply — so a member that cannot read it is broken, not weakened. Missing config refuses the spawn and names the key; an untraversable root refuses and names the root, its current mode and the group. None of this changes anything on a host where memberHerdrSocket: is unset, which is why it can look like dead code until the #185 rollout starts.

A role fleetd cannot bind is refused, not quietly downgraded

What. A spawn asking for role: architect on a profile that no configured slot carries now fails at once. The message names the role, the profile, and the profiles fleet.architects does carry. The roster also reports the role a member really holds, never the role that was asked for.

On. Automatic.

Why it exists. Such a spawn used to succeed. The member started, read the architect charter and the architect agent definition, and then fleet_whoami told it — correctly — that it is a worker, because no slot could be bound. Meanwhile GET /members still said architect. Three sources of truth disagreed about one live member, and the only log line was INFO.

That is the quiet direction of failure, which is the worse one. A member acting above its role is refused by the authorization gate, so the mistake announces itself. A member acting below its role simply does not do the job, and the lead reading the roster has no way to see why. An explicit operator request became the silent case.

Refusing was chosen over binding anyway for one reason: an architect's identity is the slot it was bound to. With no slot, there is nothing to bind an identity to, so binding would mean inventing one. The operator's fix is one line of config, and the refusal names it.

The race is closed too, by reserving first (#226). A pre-check cannot close it: "does the config carry a slot" is stable, but "will a slot still be free after the launch" depends on a spawn that may not have happened yet, so widening the check only moves the window. Instead fleetd now reserves a matching slot before it launches, binds that reservation after, and releases it if the launch fails. A spawn that would lose the race is refused before a process exists, so no member is ever started on a charter it will not hold.

Two things were deliberately kept. The old acquired fallback and its WARN stay, because a reservation narrows the window and claiming it removes the window is the mistake to avoid. And a failed launch must return its reservation, or the pool shrinks silently with every failure — that is the way this shape usually goes wrong, so it is the case the tests pin.

The gotcha: a full architect pool now fails your spawn instead of demoting it. fleet_spawn{role: "architect"} used to hand back a plain dev when every slot was taken. It now throws. That is the point — a refusal is recoverable and a mismatched charter is not — but it is a visible behaviour change for anyone who relied on the old silence. The refusal stays scoped to architect alone: dev and reviewer pools are placement candidates, not identity bindings, so an explicit profile outside them stays a supported override.

A backend that starts erroring cools off, instead of swallowing the next spawn

What. When two different members fail inside 60 seconds with output matching a profile's backend-error pattern, fleetd treats it as one incident rather than two failures. The credential those profiles share cools off for 60 seconds. Automatic placement skips it, an explicit fleet_spawn naming it is refused before the backend adapter is ever called, and the lead gets exactly one nudge in its own pane naming the credential and the affected members. fleet_list and fleet_profiles report coolingOffForSeconds, and a cooling profile's free drops to 0.

On. The correlation and the cool-off are automatic. The pattern is the knob:

profiles:
  terra:
    errorPattern: "503 Service Unavailable"   # optional

A profile that sets nothing still gets the built-in compatibility pattern, so this is never silently off. A malformed regex is rejected at config load, naming the profile, the key and the parser's own message — and since fleetd #273 that covers exhaustedPattern too, not only errorPattern. Both keys are checked in one pass, so a config with a bad value in each is told about both at once. Before #273 a bad exhaustedPattern passed load and then crashed the daemon at startup with a raw PatternSyntaxException naming neither the profile nor the key. The key is deferred: a running daemon keeps the pattern it started with, so a change needs a restart.

Why it exists. On 2026-09-01 one backend outage killed both running members mid-turn. fleetd handled each one correctly and separately, and never noticed it was one event. Both tickets were reported honestly, both members showed as done and idle — the same rows a clean finish produces — and capacity still advertised a free slot. A third spawn onto the same credential would have died the same way. The evidence that it was an outage only exists across members, so a per-turn view cannot see it however correct each turn's own handling is.

The pattern is per profile and configured for the same reason exhaustedPattern is: the previous mechanism was one hard-coded API Error: string, which fails in the silent direction. When a backend rewords its error the string stops matching, and a failed turn goes back to resolving as a success. Wording belongs next to the backend definition that produces it.

The gotcha: cooling off is not quarantine, and the two can be true at once. Quarantine (CB-578) means the backend said it is out of capacity — a long, 1800s-default cooldown. Cooling off means a credential is erroring right now — short, fixed at 60s, and deliberately not configurable per profile. They are reported as independent fields precisely so an operator can tell "spent" from "flaky at the moment", and either one alone already forces free to 0. Read the refusal wording too: "cooling off after repeated backend errors" is a different situation from an exhaustion refusal, and they have different fixes.

One error changes nothing, on purpose. One member failing repeatedly is a member problem; two different members failing together is a backend problem. Correlation keys on the credential, not the profile, because the credential is the thing that actually runs out — so cooling terra also cools sol when they share one.

A claude-code member no longer blocks forever on the workspace-trust dialog

What. Before starting a claude-code member in a worktree fleetd provisioned, fleetd marks that directory as trusted in ~/.claude.json (or the profile's configDir). The member reaches idle instead of sitting on a prompt nobody can answer.

On. Automatic, for kind: claude-code members, and only in a worktree fleetd itself provisioned. Any failure is logged at debug and never blocks a spawn.

Why it exists. Claude Code asks for confirmation the first time it opens an unfamiliar directory. Every provisioned worktree is unfamiliar by construction — it is a fresh path with a nonce in it. The member came up, printed a dialog, and waited. Nothing could answer it: the lead cannot see a member's screen, and driving the pane directly is exactly what the bridge exists to prevent. The spawn then failed as a readiness timeout, which points an investigation at the wrong thing entirely.

The rejected fixes are worth naming, because each looks reasonable: --dangerously-skip-permissions turns off a real control for a whole session; permissions.additionalDirectories widens what the member may touch rather than answering the question asked; running non-interactively gives up the pane the whole design depends on. Seeding the one flag the dialog sets is the narrow answer.

The gotcha: it writes a file fleetd does not own, and the guard is the worktree check. While this was being built, a mutation test aimed at the seeding logic overwrote the operator's real ~/.claude.json, shrinking it from 72581 bytes to 919 and taking the account entry with it. That is why the write is gated on the directory actually being a fleetd-provisioned worktree rather than on the path merely being set — a default-resolved cwd is the operator's own checkout. The write is additive and atomic, and a process-wide lock serialises concurrent spawns.

The second gotcha: the lock cannot reach the operator's own Claude Code (fleetd #247). That in-process lock only serialises spawns inside one JVM. The operator's own running Claude Code writes the same file, and on this host it is literally the same file: opus and sonnet both set configDir to the lead's own config directory. So this was never a rare collision — it was a lost update on every ordinary spawn of those profiles.

The write is now a compare-and-swap with a bounded retry. fleetd reads the file's bytes, builds its update, then re-reads the bytes immediately before the atomic move and compares. If they changed, another writer got in, so it throws its work away and rebuilds from the fresh bytes — up to five times.

On exhausting the retries it writes nothing and logs a WARN naming the cwd and the file. That is deliberate. An unseeded member shows the trust dialog and fails to reach an injectable state: visible, logged, and recoverable by retrying the spawn. Writing a stale copy over the operator's live config is silent and not recoverable. The code fails toward the recoverable outcome.

A second WARN now fires every time the seed is about to target the default ~/.claude.json, which happens only when a profile sets no configDir. That is the one path reaching the operator's home file, and it is the default, so a profile that simply forgot the setting used to get no signal at all.

What this still does not fix. The window is narrowed, not closed. A write landing between the final re-read and the move itself is still lost. There is no operating-system compare-and-swap on a plain file, only this cooperative narrowing. Do not read the CAS as making the race gone.

The third gotcha: under memberHerdrSocket the seed used to write the wrong home (fleetd #285). When memberHerdrSocket: is set, the member pane runs as a different OS user with its own $HOME. The seeding code could not see that at all — it was a static method, so it had no access to the launcher's own memberHerdrSocketConfigured() check. With configDir unset it therefore wrote fleetd's own ~/.claude.json — the operator's real file — while believing it was seeding the member's. The member never got a seed, and the operator's file was edited for nothing.

The sibling method one line below, writeCharterFile, had already solved this: it refuses the spawn when it cannot place a file where a different-uid member can read it. The trust seed now does the same. Under memberHerdrSocket it requires both configDir (so there is a member-readable target at all) and worktreeGroup (so the 0600 file it writes can be shared), and refuses the spawn — naming which key is missing — before touching any file. When both are set, the written file is chgrp'd and chmod'd to rw-r----- for that group, so the member's OS user can actually open it.

The refusal is deliberate rather than a degradation. A member with no readable trust seed is not "slightly worse": it sits on the interactive dialog forever and never calls fleet_reply, which is the exact failure this whole feature exists to prevent. With memberHerdrSocket absent — every live fleet today — nothing about this changed.

A worktree tells you which of its config files are stubs

What. A provisioned worktree neutralizes .mcp.json, opencode.json and .autoenv: the copy in the worktree is a stub, not the repo's committed file. The daemon now logs which ones it actually neutralized and which were absent, and — the part a member can reach — records the list in worktree-scoped git config:

git config --worktree --get-all fleet.neutralizedConfig
git config --worktree --get  fleet.neutralizedConfigNote

The parity overlay reports itself the same way, naming what it copied and what it marked --skip-worktree.

On. Automatic, for every provisioned worktree.

Why it exists. Both mechanisms were completely silent. A member opening .mcp.json saw a plausible file and had no way to know it was a stub — so an edit to it was real work that could never be committed, and nothing said so. The daemon side was no better: nobody could tell from the logs whether a given member had been handed an .envrc at all, and a silent copy is what makes that kind of problem hard to notice in the first place.

Both log lines carry a denominator — neutralized 1 of 3 configs: .mcp.json (opencode.json absent, .autoenv absent) — because copied 1 on its own reads as success whether the candidate list had one entry or ten.

The gotcha: worktree-scoped config is invisible to git status, and that is why it was chosen. It lives in .git/worktrees/<nonce>/config.worktree, never in the working tree, so it cannot show up as a pending change or get swept into a commit. A marker file in the worktree would have needed a gitignore entry, a name-collision check, and would still read as "this IS the file" to anything that just opens it.

The parity overlay's default changed with this work: it is [.env] now, not [.env, .envrc]. .env is data, so copying it can only move values; .envrc is executable shell that direnv runs on every cd, so copying it moves behaviour. An operator who wants it can still list it explicitly, and then owns that choice.

The credential probe asks the daemon what the policy is

What. scripts/probe-member-credentials.sh reads the policy's name list from the daemon over a new read-only endpoint, GET /member-credentials, instead of carrying its own copy of the names. The endpoint returns the policy mode, the known and allowed names, and the counts. Never a value — the daemon does not hold the values, and MemberCredentialPolicyView reads no environment at all, so there is nothing to redact by construction.

On. Automatic. The probe needs no flag; it refuses to run outside a member shell unless you pass --allow-outside-member, which labels the reading as a comparison rather than a finding.

Why it exists. The probe used to hold a hardcoded array of 31 names. The live policy had 34. It reported 26 blocked against a policy that blocks 29, exited 0, and printed a table that looked complete. Both numbers were right; they were counting different sets, and nothing said so.

That is the worst direction for a verification tool to fail in. Its clean output is taken as evidence, so a silent gap stops anyone looking. The same "second hand-maintained copy" defect had already been fixed twice elsewhere, and the fix here is the same one: delete the copy rather than correct it.

The gotcha: it refuses rather than degrades. An unreachable daemon, an absent or empty policy, or a knownCount that disagrees with the length of known[] all exit non-zero. There is deliberately no local fallback — a checker that quietly drops to a weaker check is the thing this replaced.

It also prints its own denominator: "policy contains 34 name(s); this run checked 34 — they match." Two numbers a reader can compare beat one number they have to trust.

The startup log line and the endpoint now share one class, so the counting exists in one place. Watch one subtlety if you ever recompute it: blocked is not known - allowed. Live, known is 34 and allow is 7, but only 5 of those 7 appear in known, so blocked is 29 and the naive subtraction gives 27.

An allow-list policy refuses a spawn it cannot enforce

What. memberCredentials.policy: allow-list is enforced by a generated .zlogin under a per-member ZDOTDIR. A non-zsh login shell ignores ZDOTDIR entirely, so no scrub runs. Under that policy the launcher now refuses the spawn, naming the shell it actually found and offering three ways out. Under policy: deny-by-default nothing changed — that overlay is applied to the pane before any shell runs, so it does not depend on the shell.

On. Automatic whenever policy: allow-list is set.

Why it exists. The launcher already detected a non-zsh shell and logged a warning — then degraded to the weaker overlay and spawned anyway. The operator had explicitly asked for the blocking control and silently received the weaker one. Detection existed; refusal did not, and a control that silently does nothing is worse than no control, because the config still says it is on.

The gotcha: memberLoginShell: and memberHerdrSocket: must be set together. Which shell is checked depends on the routing. With memberHerdrSocket absent — today's normal mode — fleetd's own $SHELL decides. With it set, member panes run as a different OS user, so fleetd's $SHELL describes the wrong account and only the configured memberLoginShell: is read. If you set memberHerdrSocket and leave memberLoginShell unset, the shell reads as <unset>, counts as non-zsh, and every spawn is refused. The refusal message says so, but it is cheaper to know first.

What a member can actually reach — the honest boundary

What. This is not a feature. It is the frame every credential setting on this page sits inside, and it has been missing, so people have read those settings and concluded more than the settings say.

A member runs as the same OS user as the lead. That is a trust model, not a sandbox. Inside one uid, ordinary Unix file permissions give no confidentiality boundary. A member that can run commands can read any file this user can read.

What each control actually does:

Channel Control today What it really buys
Environment memberCredentials + the ZDOTDIR scrub Real, for what the member inherits. It does not protect the values — the member can source the same files again.
Files worktree provisioning neutralizes some config files Not a boundary. File modes do not separate processes with the same uid. A member can read the original by absolute path.
Sockets sshAgentEnv: omit Not a boundary. It omits one variable. It does not revoke access to the socket, and it cannot remove a readable key file.
Git config the worktree's HTTPS rewrite + cleared credential.helper Routing, not enforcement. Normal git commands go the intended way; a member can still call ssh directly.
Process table secrets kept out of argv No boundary. argv is visible to any local user, and a different uid would not fix that.

Why it exists. The proven path is short: the scrub blanks SSH_AUTH_SOCK → git starts ssh → ssh reads the user's config → it opens an IdentityFile that is readable and has no passphrase → the forge accepts the operator's key. Measured on 2026-08-28: a member with the socket blanked pushed successfully. The gate closes lead → child environment inheritance. It never closed member → filesystem → credential.

That is worth naming as a shape, because it keeps recurring here: a gate written after an incident closes only the direction that incident came from. Ask which states open it, not just which it blocked.

The gotcha: no meaningful boundary exists inside one uid. Environment scrubbing, worktrees and config rewrites reduce accidents, and that is worth having. They do not contain a member that chooses to look. A real boundary needs a different OS user (memberHerdrSocket, above — but see the next entry: fleetd cannot verify that one for you) or OS-level confinement such as a container or VM — a different execution model, not a longer name list. Even a different user does not protect argv, and does not protect shared git objects where worktreeGroup grants access to them.

Open work, ranked, is on #184.

The ssh-agent setting is named after what it does — sshAgentEnv: omit

What. memberCredentials.sshAuthSock: block|allow is now memberCredentials.sshAgentEnv: omit|inherit.

The knob. All eight spellings parse, silently, with no deprecation warning:

memberCredentials:
  sshAgentEnv: omit        # canonical; "inherit" passes SSH_AUTH_SOCK through
  # sshAuthSock: block     # the old key and old values still work, unchanged

Old configs need no edit. If both keys appear, sshAgentEnv wins. An unrecognised value normalises to omit, so a typo fails closed rather than handing a member the agent.

Why it exists. block named an effect fleetd does not have. It omits the variable from the member's environment; it does not deny access to the socket. A member runs as the same OS user, so the socket stays reachable and the path is discoverable. An operator scanning a config file reads a value name and stops — that is the point of a good one — so block was actively misleading, and the honest explanation sat in a comment most readers never reach.

What the setting still buys is real and worth keeping: a member does not pick up the operator's agent by default, which removes a whole class of accident. It is an accident-reducer, not a deny.

The gotcha: the shim is load-bearing, not cosmetic. fleetd.yaml is gitignored, so no worker can see it and no test in the repo covers it. The live config on this host still says sshAuthSock: block. A rename without the read-both shim would have silently dropped the setting at the next restart — trading a naming bug for a real credential regression, in a file the test suite cannot reach. This was verified against the running daemon rather than argued: after redeploy, GET /member-credentials still reported policy: allow-list with 39 known / 7 allowed / 34 blocked, and a real member spawned on the new jar reported SSH_AUTH_SOCK: unset.

fleetd states the member trust model at startup

What. Every boot, fleetd logs one INFO line saying what kind of boundary members run inside. The line changes with memberHerdrSocket. With the key unset:

member trust model: members run as the same OS user as fleetd, not in a sandbox. A member can
read any file this user can read, including SSH keys and credential stores, whatever
memberCredentials says. To add a real boundary, route members to a second herdr under a
different OS user with memberHerdrSocket.

With the key set, it says instead that members are routed to a separate herdr, and that fleetd cannot see that herdr's uid, so the operator must confirm it runs as a different OS user before treating it as a boundary.

The knob. None. It always logs, at INFO, from Fleetd.reportMemberTrustModel, right after the startup secret report. Find it with:

grep 'member trust model' fleetd/fleetd.out | tail -1

Why it exists. The entry above documents the honest boundary, but a wiki page only reaches someone who goes looking. memberCredentials reads like a security control, so an operator who never opens this page reasonably assumes it contains a member. It does not. Putting the trust model in the boot log means the claim arrives unprompted, in the same place the operator already checks startup secret lines, on a machine where the config is whatever it is.

The gotcha: this line is honest about the uid, and one older WARN is not. HerdrPeerLauncher.warnUnknownMemberEnvironment still says a configured memberHerdrSocket means "member panes run under a different OS user". fleetd cannot see the uid at the other end of a unix socket — it infers that from the key being set. The direction that costs something is a second herdr running as the same user: members then inherit fleetd's environment, and the credential gap the WARN reports as UNKNOWN is in fact exactly the count it just told you to disregard. A known exposure reported as an unknown one. Tracked as item 5 on #184.

leadSeats reports the seat the lead itself holds

What. For a subscription: true profile, fleet_list's row carries a leadSeats field when the lead's own session occupies a seat on that same account. Profiles are grouped by account, not by profile name: a profile with no explicit credentialId that is subscription: true joins a shared <subscription> group. An explicit credentialId still wins, so two genuinely separate Claude logins on one host stay apart.

free is not reduced by leadSeats. free means one thing only: how many slots a fresh fleet_spawn on that profile will actually be granted right now, which is max(0, maxLoad - live) — the same check the spawn gate runs. leadSeats is a fact reported beside it, for a caller to use however it likes.

On. Automatic for subscription: true profiles, once fleet.leaders.<name>.profile is set.

Why it exists. The lead is always a live claude session on the operator's subscription, and maxLoad only ever counted members. So a fan-out that filled every member slot still left the lead's seat unaccounted for, and the daemon advertised a slot that the account was already using.

The same grouping fixes a second, wider bug: quarantining one subscription profile now also refuses spawns on the other. One subscription hitting a usage limit really does take out every profile running on it.

The gotcha, and how it was resolved (fleetd #257). The first cut of this did subtract leadSeats from free. That shipped, and was corrected the same day. The placement gate never consulted the lead seat, so with maxLoad: 3 and two members up, fleet_list said free: 0 while a third spawn still succeeded. A lead that believed the number gave up a slot the daemon would have granted — the opposite of the overstatement this was filed to fix.

The subtraction was removed rather than pushed into the spawn gate, for two reasons. Making the gate subtract it would change what maxLoad: 3 means in every existing config file, on every host. And no backend seat ceiling shared with the lead has ever been measured: a test on 2026-08-29 ran three concurrent interactive claude sessions with no trouble. Subtracting the seat therefore described an accounting policy, not a constraint, and dressing a policy as a capacity fact is what made the two numbers disagree.

Do not raise maxLoad to "win a slot back". Nothing takes one; you would be granting a real extra member.

The first attempt at this shipped inert and every test passed: the matcher grouped by effectiveCredentialId(), which fell back to the profile's own name, so a lead on opus and members on sonnet never matched and zero seats were charged. Every test in that change put the lead on the same profile name as the target — the one shape the live config does not have.

The REST face — the operator's way in when MCP is not there

What. The daemon serves 15 HTTP routes on the same port as the MCP mount. Every capability behind an MCP tool is reachable there, so the fleet can be driven with plain curl. The full reference — each route, the Authz role it needs, and its matching MCP tool — is REST API Reference.

On. Always on. The listen address is the bind: block in fleetd.yaml (host and port); this host uses 127.0.0.1:8765. GET /metrics is the one route that can be absent — it is registered only when metrics are configured.

Why it exists. MCP is the agent channel, but an operator needs a way in that does not depend on an agent session being healthy. A lead whose MCP mount has dropped cannot call a single fleet_* tool, and that is exactly the moment you most need to see what the fleet is doing. REST is that door, and it is also what a dashboard or an acceptance-test script would use.

The gotcha: GET /sessions/{id}/replies destroys what it returns. It drains the inbox, so the first read is the only read. Send it to a file. Piping it through head, or any script that exits early, loses the payload for good — and this is the route you reach for when a ticket has timed out and a member's real answer is sitting in that inbox. There is no second copy.

It is not an agent channel, and members should not be pointed at it. Not because it skips the authorization gate — it does not. Every route except /healthz resolves the caller through the same CallerResolver and the same Authz table MCP uses, so a member calling REST gets a member's rights. The reason is narrower: identity is resolved from the connection, and a member's own child process is a connection the daemon must reason about. That has been wrong before — a member's curl child was resolved as the primary until it was fixed. One channel for agents is one place to get that right. CLAUDE.md therefore says nothing about REST, deliberately, and should keep saying nothing.

Why there is no route table on this page. There is already one, on chapter 15, and a second copy is the defect — not the errors it accumulates. Chapter 15 was written on 2026-08-31 and was accurate for all 14 routes that existed then. GET /member-credentials shipped three days later and the page did not follow it, which is how the drift starts. A test in the repo now enumerates the routes FleetApp registers and fails when they no longer match its inventory, naming chapter 15 as the page to update.

A note for anyone briefing a worker to read this page

A worker cannot see the current version of this file. wiki/ is a submodule, and the parent repo's recorded pointer is deliberately never updated (committing it is on the never-commit list). So git submodule update --init in a worker's worktree checks out a months-old snapshot. During this very backfill a worker reported that an entry "does not exist anywhere in the repo" when it had been on the page for hours. Paste the relevant text into the brief instead of pointing at the file.

A dead member's seat comes back

What. A member whose backend errored (BACKEND_ERROR) or whose turn failed (FAILED) no longer counts against its profile's maxLoad. Its row stays in fleet_list so you can still see what happened, but a fresh fleet_spawn on that profile is granted.

On. Always on; no configuration.

Why it exists. Nothing ever removed these sessions, and the live counter had no state filter at all — it counted every roster entry. That counter is the one the real spawn gate reads (CompositePeerLauncher.enforceMaxLoad), so one backend error took a seat and never gave it back. Only an explicit fleet_stop on that exact pane, or a daemon restart, freed it.

The harm ran in the direction that hurts. On a maxLoad: 1 profile — opus and sol on this host — a single backend error put the profile out of service for good, and fleet_spawn answered "at maxLoad: 1 live >= 1 cap; refusing spawn — no fallback to another profile". The lead saw a capacity refusal with no reason to suspect a dead seat. It also outlived the cooldown that was supposed to be the remedy: quarantine and cool-off both expire on their own, the dead seat did not, so the profile was still refusing long after the credential recovered.

The gotcha: free and reclaimable count different things, and a freed seat belongs in only one of them. reclaimable means "this member holds a seat and has no open bridge work — stop it and you get the seat back". A terminal session holds no seat any more, so its seat is already in free. Counting it as reclaimable too would report the same seat twice, and free + reclaimable would read as more capacity than maxLoad allows. The ticket asked for exactly that, and that half of the ticket was wrong. The dead session is still visible without it: its roster row carries state: "backend_error" or "failed", which is what tells the lead to stop it.

Both places that report reclaimable — the per-member flag and the per-profile count, which travel in the same fleet_list response — now call one shared predicate, so they cannot drift apart. A test runs it over every MemberSession.State value, so a state added later cannot slip through unconsidered.

What this does not do. Terminal sessions are still never reaped. They accumulate in the roster until stopped or until the daemon restarts. That is now cosmetic rather than a capacity loss, and it is deliberate: the roster entry is the only record of what went wrong.

A member's teardown no longer leaks a worktree or a branch

What. Two cleanup paths in SessionManager were one-sided, and both are closed. Stopping a member whose worktree cannot be removed now completes the stop and logs a WARN instead of throwing. A spawn that fails after its worktree was created now deletes the orphaned branch as well as the worktree.

On. Always on; no configuration.

Why it exists. Every other cleanup step in release() was wrapped in a try/catch — the last one, removing the worktree, was not. By the time it ran, the session was already out of the registry and the pane already stopped, so a throw there escaped release() with the teardown in fact complete. There was no retry path: a second stop on that pane is a no-op. The caller saw a failed stop for a session that was gone, and the directory leaked with nothing left to point at it. git worktree remove --force can throw for ordinary reasons — a stale index lock, a slow filesystem, its own 30 second timeout — so this was not exotic.

The second leak is the same shape as the one fixed one layer down earlier: GitWorktrees already deleted both the worktree and the branch when provisioning failed inside add(). The catch that covers failures after add() returns — the parity overlay, the group share, the spawn itself — removed only the worktree. A spawn failure there is routine (a quarantined credential, a backend refusal), so every occurrence left an orphan worker/<slug>-<nonce> branch behind.

The gotcha: only the failed-provisioning path deletes a branch. A normal release deliberately keeps the branch, so a worker's committed work can still be recovered after its pane is gone. Getting these two backwards would destroy real work, so the tests pin the separation from both sides: the normal-release test asserts the branch delete is never called, and the failed-spawn test asserts the deleted branch is the exact one add() created.

A third door into the same leak, found later. Tearing a member down runs three steps: close the pane, close its tab, then release the generated credential-scrub directory. The first and third were wrapped; closing the tab was bare. By the time that step runs the pane is already gone, so a failure there is cosmetic tidying — but it threw, and the throw travelled up into the release path before the worktree removal that the fix above wraps. The session was already deregistered by then, so a second stop is a no-op: same unrecoverable leak, plus the scrub directory and its allowed N of M report. Closing the tab now logs a WARN naming the tab and carries on.

Closing the pane still propagates its failures, and that difference is the point. That step is the one whose failure means the teardown may genuinely not have happened, so reporting it as done would be a lie. The tab is not.

A chained second question no longer kills its own ticket

What. A member that calls fleet_ask twice in one turn — asks, gets an answer, then asks again — keeps its ticket. Before, the second question silently ended the ticket, and the lead's fleet_poll returned nothing for a member that was still working.

On. Always on; no configuration.

Why it exists. Each fleet_ask moves the async task to a new turnId. The thread answering the first question then finished its work against the old turnId. Whether that killed the ticket came down to which of two lines ran first — a race, not a decision. Answering now registers its waiter before it resolves, and drops the stale key explicitly rather than relying on the order.

The gotcha: the guard that looks load-bearing is not. The fix also added a state guard on the answering path, and the pull request claimed the ticket would still die without it. It would not: a mutation removing only that guard left the test green, because ask() already moves the task to the new turnId before the answering thread wakes. The guard is kept as defence in depth — it mirrors a sibling guard whose own comment warns that this ordering is not something to rely on — and the measurement is written into the code next to it, so nobody deletes it as dead code without re-checking the ordering, and nobody trusts it as the only protection either.

A ticket is not stranded when a dead member's question lapses

What. A member whose pane dies while it is parked in fleet_ask no longer leaves its async ticket at PENDING forever. About two minutes after health first classifies the member as GONE or NEVER_READY, one delayed re-check sweeps the ticket to FAILED, so fleet_poll gives the lead an answer instead of silence.

On. Always on when fleet health is enabled; no configuration.

Why it exists. An earlier fix closed the definite teardown road: an explicit fleet_stop or the idle reaper now fails an ASKING ticket immediately. Health deliberately did not, and that is right — a GONE reading is a guess from the live agent list, not a teardown the daemon performed, and a guess must never kill a ticket whose worker a live lead could still answer.

But that left a second road to the same dead end. Health fires its sweep once, on the transition into GONE, and skips the ticket because the member really is still asking. Between 55 and 115 seconds later the member's own fleet_ask lapses and clears its question — and now nothing fires again. The tick loop returns early when the state has not changed, and the classifier answers GONE before it could ever reach the orphan state. The ticket sat at PENDING for good.

Nothing else rescued it. A member parked in fleet_ask is BUSY, and the idle reaper only ever touches READY or DONE sessions, so a GONE-but-never-stopped member is never released and the teardown fix never runs for it.

The gotcha: this is one extra attempt, not a retry loop. An earlier attempt at this called the sweep once per tick for as long as a member stayed terminal, and that was rejected. So the transition schedules exactly one delayed follow-up. The delay is 120 seconds, chosen to clear the 115-second worst case of the ask window; because the ask always starts before the GONE reading, that ordering guarantees the question has lapsed by the time the re-check runs.

Two independent things keep the re-check from doing harm. It still passes sweepAsking: false, so a member that is genuinely asking again is skipped exactly as on the first attempt. And it fires only if the member is still classified in that same terminal state — a member that recovered, or was released and dropped from the roster, is left alone rather than having some brand-new unrelated turn failed underneath it.

One thing to know for maintenance. The 120-second delay is a constant in the code, not a config knob. If either ask ceiling (FleetMcp's or the REST face's, both 115 seconds today) is ever raised past it, this delay must be raised with it, or the re-check fires while the question is still open, finds nothing to sweep, and the single attempt is spent.

The REST face reports what the MCP face reports

What. GET /profiles now carries the two backend outage states that fleet_profiles has always reported: quarantined (the backend said it is out of capacity) and coolingOff (that credential threw repeated non-exhaustion errors). Each names the credentialId and the seconds remaining. The two checks are independent, so a profile can appear in both maps at once, and each map is present only when at least one profile is in that state. Separately, GET /agents and GET /members now map a herdr transport failure into the same {error, detail} envelope every other route uses, instead of letting it escape as a bare 500.

On. Always on. Nothing to configure.

Why it exists. REST is the door a lead falls back to when its MCP mount drops — listMembers says so in its own comment. So the weaker door was weakest exactly when it was load-bearing. A lead on the fallback during an outage could not tell "busy, frees in 40 seconds" from "broken", and a monitoring client that parses the JSON error envelope broke at the moment it was asking why the fleet looked unreachable. Neither gap ever caused a wrong spawn: a REST spawn goes through the same placement gate, so a refusal was still a refusal. These were visibility gaps, not state defects.

One thing to know for maintenance. Both doors render the profiles body from one method, FleetMcp.profilesView, and are handed the same BackendQuarantine/BackendOutagePolicy instances. Both halves are needed and neither is sufficient. Sharing the instances stops the doors reading different facts; sharing the builder stops them reporting those facts differently. The first version of this change shared the instances and copied the loop into FleetApp, which is the fleetd #284 shape — one rule in two places, where the next edit lands on one and not the other. If you add a field to that body, add it in profilesView and both doors get it.

A released reply goes back to the broker instead of being dropped

What. When a member's session is torn down, AmqpReplyInbox.release now cancels that target's consumer and then nacks every delivery it still holds with requeue, so an undrained reply stays recoverable. Before, it dropped its local record and left the message unacked on a still-open channel — invisible to fleet_poll and peek, and freed only when the whole AMQP connection happened to drop.

On. Always on, for the durable (AMQP) inbox. The in-memory inbox is unaffected and correctly still drops on release, because it is soft state with no broker behind it.

Why it exists. This was silent loss of the one artefact the bridge exists to carry. The path was ordinary: a member replies with no live waiter, so the reply is held; the lead sends that target a fresh task, which clears the stranded-reply flag but never drains the actual entry; the session is later stopped, so the teardown skips recovery because the flag says there is nothing to recover, and then drops the entry. The lead sees a member that finished and reported nothing, with no way to tell that from a member that genuinely never replied.

One thing to know for maintenance. The order is cancel first, then nack — not the reverse. A nack-with-requeue while the consumer is still attached hands the message straight back to that same consumer as soon as a prefetch slot frees, racing release's own cleanup and leaving a stale entry. A delivery tag stays valid for basicNack on an open channel whether or not its consumer is attached, so cancelling first costs nothing. This was found against a real broker, not by reading the code, and it is why AmqpReplyInboxContractTest needs Docker.

A failed spawn no longer leaves a live pane behind

What. Two spawn-side exits created a pane and left it running when the spawn failed: the readiness gate propagated an unrelated herdr error without teardown, and pane placement never closed the pane it had just split when the peer failed to start. Both now close the pane, best-effort, and still hand the caller the original exception unchanged.

On. Always on.

Why it exists. Nothing downstream could clean these up. SessionManager.acquire never learns the pane id — spawn throws before it returns one — so its failure path removed the worktree and the branch and could do nothing about the pane. The result was a live backend process with a deleted working directory, absent from fleet_list, holding a real backend seat that nothing decremented. The pane-placement leak was the worse of the two: with no named agent started, the orphan reaper cannot see it either, so no restart ever reclaimed it.

One thing to know for maintenance. The readiness gate still refuses to interpret an error it does not recognise — it rethrows it unchanged, and a test pins that. Closing the pane and converting the exception are different things, and only the first was missing. Do not "simplify" this by mapping the error to PeerUnreachableException; that reintroduces the exact behaviour #176 wrote failFastOnGoneBackend to keep distinct.

A reply with no content is refused instead of resolving the waiter

What. POST /sessions/{id}/reply read the content field with a silent default, so a body that omitted the field became an empty reply. That empty reply resolved the lead's waiter and the turn completed. The required-content check now lives in MessageService.reply, which both doors call, so the REST door returns 400 bad_request and the MCP tool returns a tool error. Blank and whitespace-only content are treated the same as missing.

On. Always on.

Why it exists. The lead could not tell an empty reply from a member that genuinely said nothing. This project already has several real conditions that look like that — a clipped completion scrape, a backend that died mid-turn — so the bad case hid among them. The guard was written in one handler instead of in the thing both handlers call, which is the same drift shape as #284 and #297.

One thing to know for maintenance. fleet_reply had a smaller version of the same hole: its guard checked content == null, not blank, so whitespace went through. Moving the check into MessageService.reply closed that too, but it also made the shared method throw. replyHandler in FleetMcp is a bare BiFunction with no try/catch, so that throw would have escaped the tool handler as an uncaught exception. The tool's own guard was widened to isBlank for that reason, and FleetMcpTest.replyWithBlankContentIsACleanToolErrorNotAnUncaughtException pins it. If you ever move the check again, check the handler between the door and the service, not only the two ends.

The member REST routes name a herdr failure instead of returning a bare 500

What. POST /members and DELETE /members/{paneId} were the last two routes in FleetApp with no catch (HerdrException). FleetApp registers no Javalin exception mapper, so the exception escaped as the default 500 with the body Server Error — no herdr code, no herdr message. Both now go through the same herdrError mapper every other herdr-calling route uses: 404 when herdr says the target is gone, 502 otherwise.

On. Always on.

Why it exists. fleet_spawn and fleet_stop already caught the same exception and reported a named error, so the two doors disagreed about the same failure — the shape #297 exists to remove. The stop route is the one that mattered: SessionManager.release deregisters the session, notifies the release listener and preserves a dirty worktree before it calls launcher.stop, so a throw from that stop arrives after the teardown the caller asked for has already happened. A 500 told the caller to retry and gave it nothing to reason about.

One thing to know for maintenance. An already-gone pane is not one of these failures. HerdrPeerLauncher.stop treats a *_not_found from pane.close as success, so a double stop still returns 204, and stopToleratesAnAlreadyGonePane pins that. What propagates is any other herdr failure — a transport error, or a code like herdr_busy. No new tests were added: two tests already covered these paths and asserted 500, and neither was about the status code (one guards that a failed spawn still closes its tab, the other that a failed teardown is not reported as 204). Their expectations moved to 502 and gained a body check.

One definition of loopback, so a worker cannot become the lead

What. ConnectionIdentity and CallerResolver each kept their own isLoopback, and the two disagreed: the identity resolver accepted only 127.0.0.1, the authorization check accepted all of 127.0.0.0/8. There is now one predicate, on ConnectionIdentity, that the other calls.

On. Always on. It matters most in auth.mode: loopback-trust, which is the default — auth: is commented out in fleetd.example.yaml, and the daemon logs the mode at every boot.

Why it exists. The disagreement was a privilege escalation. A caller from 127.0.0.2 had its identity resolution skipped, so it carried no terminal; CallerResolver reads a missing terminal as "not a worker", and a same-host non-worker is the primary. A worker could take spawn, stop, send and drain. The skip also happens before the PID ancestry walk, so the defence that stopped a member's curl child being read as the primary was bypassed as well.

One thing to know for maintenance. Do not "tighten" ConnectionIdentity.isLoopback back to 127.0.0.1. It reads like the safe direction and it is the opposite. That predicate does not decide whether a caller is trusted — it decides whether a caller's identity is resolved at all, and resolution is what demotes a worker. Every address excluded there is an address on which a worker becomes the lead. The range matters because on Linux the whole /8 is bound to lo, so a source of 127.0.0.2 is bindable; measured on the fleet host with curl --interface 127.0.0.2 returning exit 7 (connect refused) rather than 45 (bind failed). On macOS the same command fails at the bind, which is why no test binds a real 127.0.0.2 source — it would fail for every developer on a Mac.

The idle reaper no longer stops a worker that just got a turn

What. reapIdle decided a session was idle from a roster snapshot and then released it without re-reading. A delivery landing in between made the session BUSY, and it was torn down anyway — contrary to the method's own javadoc, which promised it never reaps a BUSY worker. The reap is now a compare-and-release: it tears the session down only while the registry still holds the exact record it checked.

On. Always on.

Why it exists. onDelivered flips exactly the two states the reaper accepts (READY, DONE) to BUSY, and bumpTurn refreshes the activity clock — so the very act that should save a session from the reaper is the one that raced it. The worker was killed, the turn never ran, and the release reason said nothing about a race. It was never a silent loss: the release listener still fails a blocked send with a reason, so the lead is told something went wrong, just not what.

One thing to know for maintenance. The compare is ConcurrentHashMap.remove(key, value), and that is the linearization point — do not "simplify" it to a re-read of the state followed by a plain remove, which shrinks the window without closing it and leaves the code looking correct. The public release path stays unconditional on purpose: an explicit fleet_stop must not be refused because the worker happens to have just become busy. A skipped reap logs one debug line naming the pane; without it the race is unobservable by construction, and a reaper that quietly stops reaping is very hard to diagnose.

A wedged post-turn phase releases itself, like the turn phase already did

What. After a delegated turn finishes, the injector runs a housekeeping phase — it sends the adapter's /clear and waits for the pane to pick it up. Four latches gate delivery to a member: awaitingCompletion, postTurnPending, awaitingPostTurnPickup and postTurnObserved. Only the first had a way out of a long run of unknown pane states. The other two now get the same escape: a sustained unknown streak drops them and the member becomes deliverable again.

On. Always on, but only reachable when the post-turn /clear is configured (lifecycle.clearAfterTurn). It is not set in the live fleetd.yaml today, so this is a latent fix here rather than one that was costing us turns.

Why it exists. A pane whose state cannot be read stays unknown forever, and a latch with no timer waits forever with it. The member never returns to idle, so every later fleet_send to it waits and then fails without ever reaching the pane. The turn phase was given an escape when this happened to it; the housekeeping phase was left with none, which is the same gate closed in one direction only.

One thing to know for maintenance. The new escape uses its own counter, unknownSincePostTurn, and deliberately does not set turnFailed. The delegated turn already completed and its waiter already resolved — what is outstanding is adapter housekeeping, so failing the turn would report a false failure for work that succeeded. When you fix a latch here, check its siblings in the same method: this fix had to cover two latches, not the one that was reported, or the wedge simply moved one notch further along.

A worker's real reply after a lapsed fleet_ask completes its ticket

What. A worker that calls fleet_ask and gets no answer within the ~55s window resumes on its own, finishes, and ends the turn with fleet_reply. That reply used to land in the session inbox with nothing tying it to the delegation: fleet_poll{ticket} stayed PENDING, and fleet_stop later forced the ticket FAILED with the reason "session released before it replied". The reply now completes its own ticket.

On. Always on.

Why it exists. The lead was told the opposite of what happened. The report was never destroyed — it reached the inbox — but the ticket said the worker never replied, and a lead that believes that re-does the work. An unanswered ask is the ordinary case on this fleet, not an edge, which is why the charter says never to brief a worker to "ask me". The sibling timeout in answer() already had this recovery; ask()'s did not.

One thing to know for maintenance. Do not "fix" this by passing forgetTurn=false on the ask timeout. It looks like the one-line version of the same fix and it wedges the member: the stamped turnId keeps hasAsyncQuestion reporting the target BUSY, so every later fleet_send to it is refused. The fix instead sets a separate Task.askTimedOut flag before the forgetting, and never re-adds the task to asyncTasksByTurn. One real consequence: the ambiguity branch in reply() used to be unreachable and is now reachable, because a lapsed ask frees its target for a fresh delegation that can lapse in turn. Two open tasks on one target fall back to the inbox rather than guess — completing the wrong ticket would hand the lead a plausible answer to work nobody did.

Shutdown refuses new spawns and sweeps up stragglers

What. fleet_spawn is now refused once the daemon's shutdown drain has started — a named error over MCP, HTTP 503 over REST — and the drain re-reads the registry after its main pass to tear down anything that raced in anyway.

On. Always on.

Why it exists. The drain worked from a one-shot registry snapshot, and the MCP server stayed up for a long time after it started: on the live config lifecycle.drainTimeoutSeconds is 120, so a drain waiting on a busy worker could hold fleet_spawn open for two minutes. A session accepted in that window was invisible to the drain — its pane kept running, its worktree was never preserved, and the in-memory registry died with the process, so nothing else could ever reclaim either. The caller got a normal sessionId and no way to know.

One thing to know for maintenance. The guard and the sweep are a pair and neither is redundant: a guard alone still loses to a caller already inside launcher.spawn(), and a sweep alone hands the caller a session that is then destroyed. The sweep shares the drain's one deadline rather than taking a second budget — a per-session grace could push the daemon past launchd's exit window and get it SIGKILLed mid-teardown. The draining flag is never reset, which is correct only because drainAll is reachable from the shutdown hook alone; if a fleet_drain tool is ever added for a live daemon, that flag becomes a permanent spawn outage.

A failed git worktree add cleans up what it half-created

What. The git worktree add command now runs inside the same cleanup scope as the provisioning steps that follow it. If the command is killed — the 30-second timeout, or an interrupt — after Git has begun writing worktree state, that state is removed instead of leaking.

On. Always on.

Why it exists. The cleanup added for #274 covered every step after the add and assumed the add itself was atomic on failure. It is, for an error Git reports; it is not when fleetd calls destroyForcibly on it. SessionManager.acquireWithWorktree never receives a path in that case, so its own if (path != null) cleanup never fires either, and nothing at any layer reclaims the directory or the branch. This is a resource leak, not data loss — no worker was ever started, so the half-made worktree holds nobody's work.

One thing to know for maintenance. Widening the cleanup scope creates a data-loss risk that the guard exists to stop. cleanupAfterAddFailure deletes the branch with -D, so running it after an ordinary "a branch named X already exists" refusal would delete the operator's existing branch. The Files.exists(worktreePath) check prevents that: measured in a throwaway repo, Git creates no directory on that refusal, so no directory means nothing was created and the branch is not ours to touch. Removing that check makes addFailureBeforeCreatingAWorktreeIsQuiet fail with the branch actually deleted. Do not remove it.

fixed placement honours the failover loop's unreachable set

What. The fixed placement policy now skips a profile that the current spawn attempt has already found unreachable. It checks ctx.unreachable() in both places it picks a profile: the configured default, and the fallback walk over the remaining candidates.

On. Only when placement: fixed is set in fleetd.yaml. The live fleet runs weighted, which already honoured the set, so this changed nothing here — it was reachable only for an operator who had switched to fixed.

Why it exists. CompositePeerLauncher retries a failed spawn by adding the dead profile to the unreachable set and asking the policy again, and its own comment says why: "Update the context for the next selection so the policy excludes this profile." fixed ignored the set, so it handed back the same dead profile on every attempt. The loop then spent all of its attempts on one profile and reported failure, while healthy profiles sat idle and were never tried. The failover existed but could not move.

One thing to know for maintenance. fixed ignoring the load caps is deliberate and stays — that is what "fixed" means. Reachability is not a cap, and the javadoc now separates the two so the next reader does not undo this as a bug fix. Also note what does not work: breaking out of the retry loop early when the policy returns an already-unreachable profile only makes it fail faster on the same dead profile, because only the policy chooses what comes next. The failure message now counts distinct candidates, so "tried 1 distinct candidate(s)" is the visible symptom of this whole class of problem.

A caller whose identity cannot be resolved is refused, not treated as the primary

What. In loopback-trust mode, fleetd grants the primary role only to a caller whose operating system process id was actually found. A caller whose peer-PID lookup failed is now refused as ANONYMOUS. ConnectionIdentity.Caller carries a resolved() predicate that says which case it is.

On. Always on in loopback-trust, which is the mode you get when fleetd.yaml has no auth: block — this fleet's live configuration. token mode was never affected.

Why it exists. LsofPeerPidLookup returns -1 for every failure, including the silent one where lsof runs fine and simply reports no matching process. No pane matches -1, so the caller arrived at the resolver with a null terminal — the same shape the real primary has, because no pane owns the primary either. The two were indistinguishable, and both were granted SPAWN, STOP, SEND and DRAIN. A worker whose lookup failed became the lead. PaneLocator's javadoc had already named this exact escalation for a different trigger, and CB-161's ancestry walk closed that one — but the walk needs a candidate pid to walk, and a failed lookup has none.

One thing to know for maintenance. The signal that separates the two cases was already in the data and simply not read: the primary has a real pid and no pane; an unresolved caller has neither. That is why resolved() lives on the record next to the sentinel rather than as a pid > 0 test copied into the resolver — one rule in two copies is what let #305 drift. This change is deliberately fail-closed: if lsof ever fails for the primary's own connection, the primary is refused until its next call resolves. That costs availability and it is the right trade, because the old behaviour spent it on a silent escalation instead. A new DEBUG line in LsofPeerPidLookup now names the previously-silent no-match case, so a refusal that does happen can be diagnosed.

The check that authorises deleting a worktree is taken after the worker stops

What. When a member is released, fleetd re-reads whether its worktree has uncommitted work immediately before the git worktree remove --force, with the pane already closed. If the tree is dirty by then, the removal is skipped and the work is also snapshotted into refs/wip/*.

On. Always on. It costs one extra git status, and only on a release that was about to delete something — a release that already decided to preserve, and every SHUTDOWN, pay nothing.

Why it exists. The old code read hasUncommitted once, while the worker was still running, and used that one boolean after launcher.stop to authorise the force-delete. A worker that committed, was released, and then wrote one more file during teardown lost that file. Worse, the same stale boolean gated the snapshot, so both defences — CB-576's preserve and CB-578 stage C's refs/wip copy — failed together. There was no third layer: pane gone, registry entry gone, directory force-deleted.

One thing to know for maintenance. The snapshot deliberately did not move after the stop. notifyReleased fires before launcher.stop on purpose (CB-516/CB-581), so a caller blocked in a fleet_ask rendezvous fails fast instead of waiting on the herdr RPC; moving the snapshot would force that notification to move too or to carry a snapshotRef that was never computed. The late snapshot's ref therefore reaches the log and not the listener — a known, accepted gap. Also note what is not proven: whether herdr's pane.close returning means the worker process is really dead is decided inside herdr, whose source is not in this repo. The re-check narrows the race; it is not proof the race is gone. And the fail-toward-preserving rule in dirtyImmediatelyBeforeRemoval's catch is the whole point — flipping it to return false turns this guard into a cause of the data loss it prevents. releasePreservesAWorktreeWhoseLateRecheckCannotBeRead exists because that mutation once passed the entire suite.

A reply that arrives while a target is being released is requeued, not stranded

What. AmqpReplyInbox.release now leaves a RELEASED tombstone in its held map instead of removing the key. A delivery that lands during or after the release sees the tombstone and is nacked with requeue, so a later owner or a connection drop can still recover it.

On. Always on, wherever the AMQP inbox is used.

Why it exists. basicCancel stops new dispatches but does not flush one already handed to the client's consumer work pool. That delivery reached deliverCallback, found the key gone, and created a brand-new map under it — one release had already walked past and would never read again. The message then sat delivered-but-unacked until the whole inbox closed: never requeued, never redelivered, and nothing peeks a released target again. A worker's real report disappeared with no log line naming it. #298 fixed the case where the delivery was already held; this is the case where it arrives during the release.

One thing to know for maintenance. The tombstone closes the window rather than narrowing it, and the reason is specific: ConcurrentHashMap serializes compute and computeIfAbsent for the same key against each other, so the swap and a racing insert cannot interleave. That is why the tombstone lives in the same map rather than in a separate "released" set — a second structure would have to be kept in sync, which is the shape that keeps producing defects here. RELEASED is one shared mutable map instance, so every read site must compare it by reference before touching it: peek, ack and deliverCallback all do, and own clears a stale tombstone with the two-argument remove so it can never delete a real map. Both halves of the fix are separately load-bearing — reverting either one alone fails AmqpReplyInboxReleaseRaceTest.

The reload classifier proves its own coverage instead of claiming it

What. A test enumerates every record component of a worker profile by reflection, changes each one in turn, and asserts the reload classifier actually notices. A component that is neither compared nor on a small, pinned exclusion list fails the build by name. The test prints its own denominator on every run:

ConfigRef.sameLaunchSettings coverage — 26 Profile components total, 23 compared, 3 excluded ([weight, maxLoad, credentialId])

On. Always on — it is a unit test, so it runs in every build.

Why it exists. sameLaunchSettings decides whether a changed profile key is reported to the operator as deferred (needs a restart). Its javadoc said it "compares every component the launcher reads at spawn". It did not: ideProjectDir, ideOpenCommand and autoCompactWindow were all missing, and worktreeGroup was missing from the sibling top-level check. A reload of any of them returned applied = true with an empty deferred list — a bare config reloaded — while the daemon kept the old value. The surrounding code already calls that "the worst outcome a reload can produce, because the operator has no reason to doubt it". The list was a hand-maintained second copy of "what the launcher reads at spawn", and it drifted. On its first run the new test found a fifth gap nobody had reported: the profile component itself.

One thing to know for maintenance. The exclusion list is this mechanism's own escape hatch, and it is pinned for a reason. Measured while verifying the fix: moving autoCompactWindow and ideOpenCommand out of the comparison and into the exclusion set left the whole suite green — the loop simply skipped them and the denominator still balanced. That is the cheapest way to silence a failing coverage test, and it silently restores the original bug. The test now asserts the exclusion set equals exactly {weight, maxLoad, credentialId}, so growing it takes a visible, deliberate edit. A component belongs there only if it is read live off the config supplier, never because adding it makes the build pass. The general rule: when a ticket asks for a checker rather than a fix, mutate the checker too, and ask what the cheapest way to pass it without doing the work would be.


An answered fleet_ask no longer fails the lead's own call

What it does. When the lead answers a worker's fleet_ask at the same moment that ask's ~55 second window lapses, fleet_send{turnId, content} used to be able to throw NullPointerException back at the lead — even though the answer had already been delivered. It does not any more.

On. Always on. There is no knob; it is a correctness fix (fleetd #324, merged 02e6aef).

Why it exists. answer() does its work holding the target's sessionLocks entry. ask()'s own timeout path holds no lock at all, and it nulls Task.turnId. finishAsyncTask read that field twice — once to check it was not null, once as the key for asyncTasksByTurn.remove. volatile makes each read fresh, but it does not make a pair of reads atomic. When the unlocked null-out landed between them, the second read saw null and ConcurrentHashMap.remove(null, task) threw. The lead was told its answer failed. It had not: task.future.complete(result) ran on the line above. A lead that reacts by re-sending is acting on a false failure.

One thing to know for maintenance. The fix reads the field once into a local. That closes the crash and not the asymmetry behind it. One side of this invariant is still locked and the other is not, and two consequences of that are open in fleetd #329: a worker's real reply can leave its async ticket PENDING for good, and reply() still contains the same double read. There is also a production test seam here — a volatile Runnable hook plus two package-private setters, null and unused outside tests. It exists because no public path reaches this interleaving without a real race, so a deterministic test needs somewhere to stand.


A reload now says a restart is needed for primary: and configReload:

What it does. ConfigRef reports a changed primary: or configReload: block as deferred — accepted into the new snapshot, but not in effect until the daemon restarts. Before this, changing either one reported a bare config reloaded and the daemon quietly kept the old value.

On. Always on (fleetd #326, merged 823976c). Visible in the reload summary and in ConfigRef.Outcome.deferred().

Why it exists. Both keys are read once, at startup. primary: feeds PrimaryRegistry's pinned terminal and sizes ReplyPushLoop's reminder cap and backoff; neither is rebuilt. configReload: decides whether a ConfigWatcher is built at all and with what interval — so the component that would apply a later change is itself built once. Turning reload off through a reload reported success and changed nothing. This is the same drift fleetd #323 fixed one level down, in the per-profile list.

One thing to know for maintenance. Write down the denominator, because "not mentioned in ConfigRef" looks identical for a key that is correctly hot and for a key nobody triaged. FleetConfig has 22 top-level components; four are named nowhere in that file. memberCredentials and memberLoginShell are hot and correctly absent — both are read live off config.get() at spawn. health and coordinator are undecided: each is read both off the startup snapshot and live, at different sites, so no single class fits either. A reload touching them still under-claims. That count and both verdicts are now in ConfigRef's class doc so the next person does not measure it again. Twice now — worktreeGroup, then primary/configReload — the untriaged kind hid among the correct kind.


A reload can now say "half of this applied"

What it does. ConfigRef has a fourth reload class, split, for a key that is read both off the startup snapshot and live off the config supplier, at different sites. health: and coordinator: are both like that. A reload that changes either now names the key and says which half is already live and which half waits for a restart. Outcome carries a new split list beside deferred.

On. Always on (fleetd #330, merged 7b918c5). Visible in the reload summary:

config reloaded; partially live — coordinator: the LeadMailbox connection (uri, uriEnv, selfId,
prefetch) is opened once and needs a restart; the broker URI env-var name kept out of a member's
environment is read live on every spawn and already applied

Why it exists. Before this, changing either key reported a bare config reloaded. That under-claims: it tells the operator a change applied when half of it did not. The three existing classes could not express the truth, and forcing one of them would be wrong in one direction or the other. The direction matters. Over-claiming a restart costs an unnecessary restart, which the operator can see and recover from. Under-claiming is what this file's own doc calls "the worst thing a reload can do to an operator debugging one". split is the only option that is simply true. Making the frozen half live was rejected for coordinator: selfId names this daemon's own AMQP inbox queue, so changing it live is a distributed-identity problem, not a reconnect.

One thing to know for maintenance. Membership in SPLIT_KEYS is not proof that any reporting code exists for that key. ConfigRefTopLevelCoverageTest reads a name in the set as "triaged" and stops there — it cannot see whether changedSplitKeys has a branch for it. Measured: dropping the coordinator branch while leaving the name in the set left both the coverage test and the assert silent; only three hand-written behavioural tests caught it. The same one-way shape is older than this change — assert COLD_KEYS.containsAll(changed) catches "reported but not listed" and never the reverse. fleet: is already an instance of the gap: it is split too and it sits in the checker's hot-exclusion hatch, so a changed fleet.leaders still reports nothing. Both are open in fleetd #333.


Every top-level config key must be triaged before it ships

What it does. ConfigRefTopLevelCoverageTest enumerates FleetConfig's record components by reflection and requires each one to sit in exactly one of four buckets: COLD_KEYS, the deferred set, SPLIT_KEYS, or a pinned hot-exclusion set. A new key in none of them fails the build by name. It prints its own denominator every run:

FleetConfig top-level coverage — 22 components total: 5 cold [...], 11 deferred [...],
2 split [health, coordinator], 4 hot-excluded [...]

On. Always on — it is a unit test (fleetd #330).

Why it exists. Three times the same key-was-forgotten bug shipped: worktreeGroup (#323), then primary and configReload (#326), then fleet (#333). Each time a reload reported success for a change the daemon never picked up. "Not mentioned in ConfigRef" looks identical for a key that is correctly hot and a key nobody triaged, so the forgotten kind kept hiding among the correct kind. This is the top-level twin of the per-profile checker #323 added.

One thing to know for maintenance. Read the test's own javadoc before trusting a green run. It says plainly what it cannot do: it proves the record's shape is triaged, and it cannot prove a citation is true. "Compared in changedDeferredKeys" and "read live off config.get()" are facts about other files that a reflection test over one record cannot inspect. That disclosure is not modesty — it is what made the two gaps in #333 findable in minutes. Hold any checker in this repo to the same standard: say what it does not cover, next to what it does.


An exception thrown after a ticket resolves now reaches the log

What it does. sendAsync's executor used to end with a bare catch (Throwable t) { task.future.completeExceptionally(t); }. That future is already completed by then, because finishAsyncTask completes it on its first line. completeExceptionally on a completed future returns false and does nothing. The exception simply vanished. The catch now checks that return value and logs at error with the ticket and the target when it is false.

On. Always on (fleetd #329). Nothing to configure — look for async send task-N -> <target> threw after its ticket was already resolved in fleetd.out.

Why it exists. This was not one bug, it was a blind spot over the whole async region. Measured during #324: a deliberately broken finishAsyncTask threw 19 real NullPointerExceptions on the ordinary path while the suite reported 1340 tests green and not one log line. Any defect that throws after the future completes produced a passing build and a silent daemon. The daemon could not tell an operator, and no test could tell a developer.

One thing to know for maintenance. When a mutation you expect to fail passes, that can mean the failure is invisible, not that the code is unpinned. Add one temporary log.error and re-run before you conclude anything. That is the only reason this was found.


A worker's real reply completes its async ticket more often

What it does. answer() used to look the task up a second time, by turnId, when completing the async ticket. That second lookup raced ask()'s own unlocked timeout cleanup, so a ticket could stay PENDING forever after the worker had actually replied. answer() now completes from the Task it already holds from its first lookup.

On. Always on (fleetd #329).

Why it exists. The failure was silent and it lied to the lead. fleet_poll{ticket} reported PENDING for good, and teardown later resolved the ticket as WORKER_FAILED — "session released before it replied" — long after the worker had replied. A lead reading that is told something false about its own worker.

One thing to know for maintenance. This narrows the window; it does not close it. ask() drops the asyncTasksByTurn entry in its catch (clearAsyncQuestion(turnId, true)) but closes the ask later, in its finally. Between those two the ask is still answerable and the entry is already gone, so answer()'s own first lookup returns null and the ticket is stranded one step earlier in the same race. Measured on 2026-09-04: a probe firing only that first half printed answer=REPLIED phase=PENDING reply=null. Open as fleetd #334, and the comment in answer() says so. Do not read that task != null guard as complete.


A reload now reports a changed fleet.leaders as needing a restart

What it does. fleet: used to be filed as a fully hot key, so a reload that changed it said nothing at all. It is really a split key. fleet.leaders is read once at startup, in two places — Fleetd builds the tab-label→lead map from the startup snapshot, and LeadLauncher holds a frozen FleetConfig rather than a supplier. Neither rebuilds on reload. The rest of fleet: (role pools, charters, tabLabel) really is live. changedSplitKeys now compares fleet.leaders specifically and names both halves.

On. Always on (fleetd #333).

Why it exists. A lead is found by its tab label, and a lead whose tab no longer matches is demoted to worker and refuses every orchestration call. Before this, an operator could edit that label, read config reloaded, and be left with a broken lead and no message saying why. The comparison is deliberately on fleet.leaders and not on the whole fleet record: comparing the whole record would claim "needs a restart" for a tabLabel-only change that is fully live. Over- claiming a restart is cheap to recover from; it is still a wrong report, and this class exists to stop wrong reports in both directions.

One thing to know for maintenance. This is the third time the same bug shipped — worktreeGroup (#323), then primary and configReload (#326), now fleet. It was in the coverage checker's own hot-exclusion hatch, blessed by the checker, on the checker's first commit. A name sitting in an escape hatch is the place to look first, not last.


Being listed as cold or split now proves a comparison exists

What it does. ConfigRefTopLevelReportingCoverageTest builds FleetConfig pairs by reflection that differ in exactly one top-level component, then calls the real changedColdKeys and changedSplitKeys and requires that key to come back. A name added to COLD_KEYS or SPLIT_KEYS with no if behind it now fails the build, by name.

On. Always on — a unit test (fleetd #333).

Why it exists. Every checker in this area had closed only one direction. The older assert COLD_KEYS.containsAll(changed) catches "reported but not listed" and never the reverse, and ConfigRefTopLevelCoverageTest reads a name in a set as "triaged" and stops there. Measured: dropping the coordinator branch while leaving the name in SPLIT_KEYS left both of those silent. So the cheapest way to pass the shape checker was to add one string and write no code — which is exactly how fleet got through.

One thing to know for maintenance. It covers COLD_KEYS and SPLIT_KEYS only. The deferred set — 11 of the 22 keys, the largest bucket — is not covered, and the worker wrote that gap into the test's own javadoc instead of quietly leaving it. Measured on 2026-09-04: deleting guard's comparison from changedDeferredKeys while leaving guard in the deferred set left all 1355 tests green. Open as fleetd #337. Two tickets running, an honest caveat in a test's javadoc has been the fastest route to the next bug — hold every checker here to that standard.


A send that times out while still queued now cancels its message

What it does. Injector.enqueue returns an identity Delivery handle. When a send's deadline passes and its message was never delivered, MessageService cancels that exact Pending instead of leaving it in the queue. queuedDeliveries still records the health fact — that half is unchanged.

On. Always on (fleetd #338).

Why it exists. Before this, TIMED_OUT_QUEUED told the sender the message did not go, and then the message went anyway. The injector picked it up on the member's next injectable status and typed it in — minutes or hours later, after the lead had moved on and usually re-sent the work elsewhere. Nobody was told. This is not a theoretical path: it is a behaviour hit while operating the fleet, and a test had been asserting it as correct.

One thing to know for maintenance. Cancelling is the strict direction and it costs something: a member that was about to go idle loses a message it could have taken, and the lead must send again. That is the right trade because it is loud and recoverable, but it is a trade. If a queued message starts disappearing more often than expected, this is why.

The race is resolved deliberately. Cancellation and Injector.onStatus share the target monitor. If pickup wins, the text has already landed, cancel returns DELIVERED, and the result is TIMED_OUT_WORKING — not TIMED_OUT_QUEUED. Claiming "queued" while the text landed would be the original bug with a smaller window. Note that MessageService reading that return value is not currently pinned by a test: InjectorTest covers cancel itself, and mutating the caller's use of it left all 1357 tests green. Open as fleetd #345.


A member's prose about an error no longer records a credential outage

What it does. CompletionResolver used to notify backendErrorSink on any line of a member's pane matching the backend-error pattern. The built-in fallback is (?i)\bAPI Error\s*: — case-insensitive and unanchored — so a member that ended a turn without fleet_reply while merely writing about an error matched it. The sink call now needs the pattern at the start of its matched line, ignoring leading terminal chrome. Failing the send is unchanged: any match still fails it and still carries the whole pane tail.

On. Always on (fleetd #339). The too-fast crash path still notifies on any match, because there the crash signature is corroboration.

Why it exists. The sink is not cosmetic. Two backend errors on one credential within 60 seconds put it into cooling-off, and a fleet_spawn naming a cooling profile is refused before it reaches the backend. Profiles share credentials here, so a member's own prose could block spawns on a profile that never had a problem. The code's comment already admitted the false match and argued that failing the send is still right — a sound argument that covers resolveFailure and says nothing about the sink call sitting in the same block.

One thing to know for maintenance. The first version of this check used a bare lookingAt(), and that rejected a genuine error line rendered as │ 503 Service Unavailable: ... — the send failed and the outage went unrecorded. That is the worse direction: an unrecorded outage leaves the fleet spawning into a dead credential. It was caught on merge by a probe, not by the suite, because every existing test put the error line with no chrome in front of it. The check now skips a leading run of non-letter, non-digit characters, and aRealErrorBehindTerminalChromeStillNotifiesTheSink pins it. If you touch this check, test it against a chromed line, not only a clean one.

It is still a heuristic: prose that begins with API Error: will still notify the sink. The same free-text shape drives ExhaustionSink, where a quarantine runs 1800s against this cooldown's fixed 60s — open as fleetd #348, and unproven.


Every distinct unprotected credential name gets its own warning

What it does. The memberCredentials gap WARN — the one naming environment variables that every member pane inherits unblocked — was guarded by a single AtomicBoolean shared by two branches that report different variable names. It is now a Set<String> of names already warned about, so the guard is per name rather than per launcher.

On. Always on (fleetd #341).

Why it exists. memberCredentials is re-read on every spawn, so an operator can change the policy with a reload and no restart. Spawn 1 under deny-by-default warned about one variable and tripped the flag; after a reload, spawn 2's genuinely unprotected different variable was never reported. The operator fixes the one name they were shown and reasonably believes the gap is closed. The whole point of naming variables in these lines is so they can be acted on.

The same class already had the right reasoning written down: #192 split this flag from the allow-list INFO guard precisely so a harmless report could not suppress a real one. That reasoning was applied to INFO-versus-WARN and never to WARN-versus-WARN.

One thing to know for maintenance. The set has two duties and only one was pinned at first. Measured on merge: replacing the .filter(...::add) with one that logs every name on every spawn left all 1358 tests green — the noise control was correct and nothing held it there. Two tests now cover both duties, including the reverse policy order (allow-list first, then deny-by-default), because a guard fixed in one direction is not automatically fixed in the other.


The deferred key set now proves its own reporting coverage too

What it does. ConfigRefTopLevelReportingCoverageTest covered COLD_KEYS and SPLIT_KEYS only. It now covers the deferred bucket as well. The deferred set moved out of the test and into ConfigRef.DEFERRED_KEYS (package-private, 11 keys), and changedDeferredKeys became package-private like changedColdKeys and changedSplitKeys, so the test calls the real method instead of a copy of the list. A name in any of the three sets with no comparison behind it now fails the build, by name.

On. Always on — a unit test (fleetd #337).

Why it exists. The deferred bucket is the largest of the four: 11 of the 22 top-level keys. It was also the one bucket where a name could be added with no code behind it and every checker stayed green. Measured before the fix: dropping guard's comparison out of changedDeferredKeys while "guard" stayed in the set left all 1355 tests green. An operator who changes such a key gets no "restart needed" line, so the daemon keeps running config that matches no file on disk and says nothing.

One thing to know for maintenance. The real number of uncovered keys was 6 of 11 — guard, worktreeRoot, spawnReadyTimeoutMs, spawnReadyPollMs, quarantineCooldownSeconds and leadHeartbeat. My own ticket listed a different six. The worker re-derived the list by mutating each key one at a time, as the brief asked, and contradicted me on two entries: lifecycle was already covered, and worktreeRoot was uncovered and missing from my list. Both corrections were checked again on merge with a third mutation. A list in a ticket is a starting point, not a measurement — re-derive it, and say so when it disagrees.


Teardown now closes the tab a member really sits in

What it does. HerdrPeerLauncher.stop used to look for a tab to clean up only when one of its own configured profiles used tab placement (usesTabPlacement()). It now resolves the pane's real tab every time, and the single-occupant check decides whether that tab is closed. That check is unchanged: a tab holding other panes is never closed.

On. Always on (fleetd #342).

Why it exists. CompositePeerLauncher routes a stop through the adapter recorded in spawnedBy, and that map is in memory only — a daemon restart empties it. On a miss in a one-daemon fleet, the stop goes to delegates.getFirst(). When two adapters share one herdr daemon and disagree on tab versus pane placement, teardown could run through an adapter whose config says nothing true about how that pane was placed, and the whole tab-cleanup block was skipped. The member still stopped, but its now-empty tab stayed. Nothing reaps an orphaned tab, so the operator's herdr session collected one more of them on every affected teardown, silently.

Why not fix the routing instead. Probing for the pane's real owner looks like the obvious fix and does nothing here. probeOwner groups candidates in an IdentityHashMap keyed by HerdrClient, so two delegates sharing one daemon collapse to whichever was inserted first — that is delegates.getFirst() again, the same answer the shortcut already gave. The routing cannot tell these two adapters apart at all.

One thing to know for maintenance. The single-occupant check is now the only thing protecting a shared tab, so treat it as load-bearing. Measured on merge: weakening it from tabPaneCount() == 1 to >= 1 fails two tests, including the pane-placement one. One case is uncovered on purpose — a pane-placement member that is the sole occupant of its tab will now have that tab closed. spawnAsPane splits an existing tab, so the count is normally at least 2, and the tab is empty after the member's pane goes anyway.

The same in-memory spawnedBy breaks clearContext a different way: it plainly no-ops on a cache miss. Open as fleetd #352, and the consequence is not measured yet.


A failed cleanup no longer strands the tickets behind it

What it does. MessageService.abandon walks every task it must fail and completes each one. The per-task cleanup after that completion is now inside a try/catch, and each task's own future.complete(outcome) runs before the guard. A cleanup that throws costs that one task its bookkeeping and nothing more; the loop still reaches every task behind it. The second half of the same fix wraps pushLoop.onTicketTerminal inside sendAsync's whenComplete action.

On. Always on (fleetd #335).

Why it exists. abandon runs on teardown — a release, or the health monitor's GONE sweep — and nothing comes along later to finish what it misses. One call in that loop reaches a broker: inbox.publish puts a recovered reply back, and AmqpReplyInbox.publish throws IllegalStateException on an unroutable publish, on a confirm timeout, and on an interrupt. An uncaught throw there aborted the loop, so every task after it stayed PENDING forever and its lead waited on a ticket that would never resolve. The whenComplete half is the same failure with a different cause: the daemon's shutdown hook closes MessageService before ReplyPushLoop, and messages.close() does not cancel a send already in flight, so a ticket completing in that window made the push loop's scheduler throw RejectedExecutionException into a discarded stage — no log, no metric, and the push loop never learned the ticket was terminal.

One thing to know for maintenance. A publish failure used to propagate out of abandon to its caller. It is now logged at error with the ticket, target and turnId, and swallowed. That is deliberate and matches what #293 already does for teardown in HerdrPeerLauncher: past the point where the real work is done, a cleanup failure must not mask the steps behind it.

A third site was reported and is not a defect: the finally blocks in send() and answer() call asyncTasksByWaiter.remove and Rendezvous.close, which is waiters.remove(session, waiter). Neither can throw, so no guard was added. A guard that can never fire is worse than none — it reads as evidence that somebody checked.

Measured on merge: keeping the catch but adding a break to it leaves the suite failing at aPerTaskCleanupFailureDoesNotStrandTheRemainingMatchingTasks with expected: <FAILED> but was: <PENDING>. So the test pins the property that matters — the tasks behind the throwing one still finish — and not merely that no exception escapes.


A member's prose about a usage limit no longer quarantines a credential

What it does. CompletionResolver classifies a turn that ended without fleet_reply and whose pane text matches the profile's exhaustedPattern as BACKEND_EXHAUSTED, and tells ExhaustionSink. The sink call now needs the match to sit before the first sentence ending on its line. Failing the send is unchanged: any match still fails it and still carries the whole pane tail.

On. Always on, wherever a profile sets exhaustedPattern (fleetd #348).

Why it exists. This is #339's shape with a much heavier penalty. A backend-error cooldown is a fixed 60 seconds; an exhaustion quarantine defaults to 1800. Measured on this ticket rather than assumed: a normal member report — "I reviewed capacity handling. The usage limit has been reached means no more work can start." — notified the sink and would have taken a credential out for half an hour.

Why the check is looser than the backend-error one. An exhaustedPattern is written per profile and may name only the decisive words, without the provider's leading "The". A start-of-line check would then reject the genuine refusal, which is the worse direction — an unrecorded exhaustion leaves the fleet spawning into a credential that has no capacity. Measured on merge: swapping in the start-of-line check fails four tests, three of them pre-existing. The rule as written accepts a superset of what a start-of-line check accepts, so it cannot add a false negative.

One thing to know for maintenance. It is a heuristic and the javadoc says exactly where it stops. Prose whose first sentence carries the pattern still notifies the sink; a genuine refusal behind an earlier full stop — a hostname, a version number — still does not.

The first version copied the backend-error check's leading-chrome loop. Measured: deleting that loop left all 1369 tests green, and it must, because the scan only looks for ., ! and ? and no chrome character is one of those. It is gone. A step that cannot change the result is worse than no step — the next reader takes it as evidence that chrome was handled.


The last window where a fleet_ask timeout stranded its ticket is closed

What it does. When a worker's fleet_ask times out, ask() now closes the ask turn (rendezvous.closeAsk) before it forgets that task's turnId mapping. The order used to be the other way round, with the close happening later in the shared finally. The whole teardown is also gated on ticket.fresh() now, matching the finally block and the NO_WAITER branch, which were already gated that way.

On. Always on (fleetd #334, closing what #329 only narrowed).

Why it exists. Between the forget and the close, the ask was still answerable while its task mapping was already gone. A lead calling fleet_send{turnId} in that window got a REPLIED answer, while the ticket stayed PENDING with no reply — forever, because nothing revisits it. Measured with a probe before the fix: answer=REPLIED phase=PENDING reply=null. With the new order, an answer() either sees the ask open — and then the mapping is still there — or sees it closed and returns STALE_TURN. There is no state in between, because both steps run on one thread with nothing yielding.

The ticket.fresh() gate fixes a second door into the same failure. A coalesced duplicate ask passes its own timeoutMillis, which says nothing about whether the shared ask is done, so a duplicate timing out first could lapse an ask the fresh owner was still holding.

One thing to know for maintenance. The two halves are pinned by two different tests, and the second one only exists because the first did not cover it. Measured on merge: removing the ticket.fresh() gate while keeping the new order left all 1371 tests green. The gate shipped with the reorder and nothing held it there. aCoalescedDuplicateAskTimingOutLeavesTheFreshOwnersAskOpen now does, and aLateAnswerDuringAskTimeoutTeardownStillCompletesTheAsyncTicket covers the ordering.

MessageService now carries six test-only hooks, one per race of this kind. Each exists because its window is unreachable through the public API — which is also why each bug was invisible — but six is enough that the next one needs a harder look than "the file already does this".

fleet_list reports lead coordination state instead of guessing at it

What it does. A lead's fleet_list now answers three questions it could not answer before: is my own coordination mailbox there and is anyone reading it, what is waiting in it, and what is the state of each peer daemon I know about. Each mailbox row carries a status of exists, absent or unknown. pending and consumers appear only when status is exists.

The reading is a passive AMQP queue declare, not a presence protocol. consumers: 1 means a daemon is attached and consuming; pending: N is the backlog.

fleet_send{coordId} also stopped saying "delivered". It now says the message was published and durably confirmed by the broker, which is what a publisher confirm actually proves. When the target mailbox exists but has zero consumers, the (still successful) result adds a warning that nobody is reading it right now.

On. Add peers: to the coordinator: block — the operator declares who exists, because the daemon never guesses:

coordinator:
  uriEnv: LEAD_COORD_URI
  selfId: mac
  peers: [fleet01]

Omit it and you still get your own mailbox row. An undeclared peer can still reach you and be reached by fleet_send; it simply does not get a row. Like the rest of the coordinator: block, peers is read once at boot, so a change needs a daemon restart.

Why it exists. The first real cross-host link (Mac ↔ fleet01, 2026-09-05) worked, and using it showed the MCP surface only covered sending. A lead could not learn that a peer existed, could not read its own inbox, and was told "delivered" for a message the peer never saw — that one sat undelivered through three daemon restarts because of the duplicate-lead-tab bug (#359). Finding that out needed an ssh to the other host and lavinmqctl list_queues. None of it was reachable through MCP, and a lead on a host with no broker shell was blind to its own inbox. fleetd #361.

One thing to know for maintenance. The three-state status is the whole point, and it is easy to collapse back into a boolean. The first cut of this feature did exactly that: one absent() value stood for both "the broker said there is no such queue" and "I could not check", so a self-probe timeout rendered as pending: 0, consumers: 0 — indistinguishable from a mailbox that is genuinely empty and genuinely unread. That is the same overstatement the ticket exists to fix, one level down.

LeadMailbox.isMissingQueue is the discriminator that keeps them apart, and it is narrow on purpose: only an IOException whose cause is a ShutdownSignalException carrying an AMQP.Channel.Close with reply code 404 counts as a confirmed absence. Measured on merge: making it return true unconditionally restored the original defect and left 1389 tests green. It now has five tests of its own, and the same mutation gives 4 failures.

Two other things this feature depends on, both easy to undo by accident. inspect runs its passive declare on a throwaway channel, because in AMQP 0-9-1 a passive declare of a missing queue closes the channel it ran on — reusing the publish channel would let one miss break every later publish on that instance. And FleetMcp.probe cancels a timed-out probe rather than abandoning it; without that, a hung (not down) broker would orphan one channel per fleet_list call until the connection's channel-max ran out, breaking publish by a different route.

Bridge skills are seeded into every provisioned worktree

What it does. fleetd copies a directory of skill folders into each worker worktree it creates, at <worktree>/.claude/skills/. A skill folder whose name the target repo already ships is never touched — the repo's own copy wins, byte for byte.

On. Point memberSkills: at the directory:

memberSkills: /Users/dai.ha/LTMS/claude-bridge/.claude/skills

Leave it out and nothing is seeded, which is the old behaviour. This is a deferred key — the GitWorktrees that reads it is built once at startup, so a change needs a daemon restart.

Why it exists. Every brief starts with Load the <name> skill., and outside this repo that line was silently a no-op. A member spawned against any other repo — kb on fleet01, for example — had no implementer, reviewer or hunter to load, and nothing said so. The skills could not travel in the plugin either: ClaudeCodeLauncher exports CLAUDE_CONFIG_DIR, so a member never reads the operator's plugin store. The worktree is the only channel that reaches a member. fleetd #362 item 3.

Opencode members need a second step, and they now get it (fleetd #393). Copying the folders is the whole feature for a Claude Code member, because Claude Code reads .claude/skills/ natively. Opencode never reads that directory. So for the first weeks this key existed, an opencode member was seeded correctly and read nothing: the copy succeeded, the files were right, and no test failed because there was nothing to fail. The feature worked at the only layer it implemented.

An opencode member's only channel for static guidance text is the instructions[] array in the config OpenCodeLauncher generates for it. Each seeded skill's SKILL.md is now added there, by absolute path. A skill folder with no SKILL.md is never delivered, and the log names the folder so a typo is visible instead of silent.

The gotcha that matters here is delivery versus activation. Opencode has no equivalent of Claude Code's Skill tool. The text arrives as part of the system prompt from spawn and stays there; a member cannot load one skill by name when it needs it. So Load the implementer skill. means something different on the two backends: on Claude Code it is an instruction the member acts on, on opencode the content is simply already present. Do not read "skills work on opencode now" as more than that. This is opencode's design, not a fleetd limit, and it is why the delivery gap could be closed here and the activation gap could not.

A second gotcha, for anyone adding a third kind of guidance file. Three writers append to instructions[]: the role charter, the seeded skills, and the IDE rules. All three now use Jackson's withArray (get-or-create). One of them used putArray (create-or-replace), which was safe only because it happened to run first against an empty array — an ordering rule nothing wrote down and nothing tested. Measured before the fix: switching the skills writer to putArray left the whole suite green while silently deleting the charter entry, so an opencode member launched with no role contract at all. If you add a writer, use withArray, and assert the array's contents — a size assertion passes when putArray swaps two entries for two different ones.

One thing to know for maintenance. core.excludesFile is single-valued. The seeded paths are hidden from git status by pointing that key at a fleetd-written file, scoped --worktree — and a worktree-scoped value replaces the operator's global one rather than adding to it. The first cut of this feature did exactly that, and the consequence was severe: this repo's own .gitignore does not ignore target/, only an operator's global excludesFile does, so every worker that ran mvn clean install made target/ untracked. GitWorktrees.hasUncommitted counts untracked files on purpose (CB-576), so SessionManager would have preserved every worktree that built, forever, with no error to notice.

So the file is composed, not replaced: whatever core.excludesFile resolved to beforehand is copied in ahead of the seeded patterns, including git's own default ($XDG_CONFIG_HOME/git/ignore, else $HOME/.config/git/ignore) when the key was unset. Measured on merge: removing the composition fails 2 tests, and removing just the default-file fallback fails 1.

Two smaller things. The composed content is a snapshot taken at seed time, so an operator editing their own excludesFile later does not change an already-seeded worktree. And the exclude file itself lives under the worktree's private git dir (git rev-parse --absolute-git-dir), not in the working tree, so it cannot be committed and git worktree remove --force deletes it along with everything else.

.git/info/exclude was rejected as the mechanism, and this is worth knowing before someone tries it again: from a linked worktree it resolves to the common git dir, so it would have hidden the seeded paths in the primary checkout and every sibling worktree too.


Dead lead tabs are cleaned up, and a live one is never closed

What it does. On startup, LeadLauncher.ensureLeads() now looks for tabs that carry a lead's configured label but have no running agent. It does not close them straight away. It renames the tab, appending [fleetd:pending-close], and leaves it open. Only a later reconcile that still finds the same tab dead actually closes it. LeadTabScanner does the matching cross-check on the read side: it joins a labelled tab to a terminal only when agent.list says that terminal is live, and it grants one grace scan to a terminal it already knew was live.

The knob that turns it on. None — it is always on for any lead slot declared under fleet.leaders.*. The marker suffix is a constant in PendingCloseMarker, and every place that matches a tab label strips it first, so a flagged tab is still recognised as that lead's tab.

Why it exists. Two separate defects met here. LeadTabScanner joined labelled tabs straight to terminals with no liveness check at all, and its javadoc excused that ("a stale name costs nothing here"). It cost plenty: LeadCoordLoop reads that map to choose which pane a peer lead's message is delivered into, so a dead tab was a valid candidate. LeadLauncher had the opposite problem — no cleanup path whatsoever, so every reconcile that found 0 live leads created another tab and left the old one behind. Restart the daemon a few times and the tab bar fills up.

One thing to know for maintenance. The two-reading rule is not caution for its own sake. The evidence that opened this ticket was fleet01's own daemon log: agent.list reported 0 live while ps showed one real claude process. A first cut of this fix closed tabs on that single reading, which would have closed the operator's live lead rather than tidying a spare tab. So the accepted cost is stated plainly: ensureLeads() runs at startup only, so the second reading arrives at the next restart, and a tab whose agent dies mid-session stays flagged and open until then. That is deliberate. The bug is about repeated restarts, and one leftover tab is much cheaper than closing a live session on evidence that has already been seen to lie.

Measured on merge: 1412 tests green; making PendingCloseMarker.strip() the identity function fails 4 tests. fleetd #359.


The shipped systemd units no longer disable the daemon they start

What it does. deploy/fleetd.service, deploy/herdr.service and deploy/herdr-inner.sh are the units and helper script that actually run fleet01. Copy them, edit the paths in the headers, systemctl --user enable --now both. fleetd starts from a login shell, neither unit takes a mount namespace, and neither gets a private /tmp.

The knob that turns it on. Nothing to switch on — these are the deployment artefacts. What matters is what must stay off, and the files say so in their own comments.

Why it exists. The old deploy/fleetd.service carried ProtectSystem=strict, ProtectHome=read-write, ProtectKernelTunables=true, ProtectControlGroups=true, PrivateTmp=true, and an ExecStart that ran java directly. Every one of those looks like good hardening and each breaks the daemon silently:

  • Each of the four Protect* directives gives the unit its own mount namespace. fleetd resolves a caller's role by running lsof to find the loopback peer PID (mcp/LsofPeerPidLookup). Inside such a namespace lsof returns nothing, every caller falls back to ANONYMOUS, and the primary is refused every orchestration call with "unauthenticated: anonymous may not SPAWN". The daemon still starts. /healthz still returns ok. The only symptom is that the fleet cannot be driven at all. Bisected on fleet01 (lsof line count): no sandbox 3, ProtectSystem=strict 0, ProtectHome=read-only 0, ProtectKernelTunables 0, ProtectControlGroups 0, RestrictSUIDSGID 3, NoNewPrivileges 3 — so the last two are safe and are kept.
  • PrivateTmp=true must be false on both units. fleetd writes the member ZDOTDIR credential-scrub directory and the opencode config directory under java.io.tmpdir, and the member pane — a child of the other unit — has to read them back. A private /tmp turns the credential scrub into a silent no-op.
  • systemd runs no login shell, and every secret the daemon needs lives in a file only the login shell sources. Started with a bare ExecStart=java, fleetd boots fine with empty credentials and the failure appears hours later as a member that cannot open a pull request.

deploy/herdr.service did not exist at all, even though fleetd.service's After=/Wants= already named it. fleetd #360.

One thing to know for maintenance. A unit file has no compile step, so SystemdUnitSafetyTest is the guard: it reads all three files and fails on an active forbidden directive, on PrivateTmp=true, on an ExecStart that skips the login shell, on a herdr-inner.sh that does not exec a login shell, and on one that does not set a non-zero pty size (stty rows N cols M with both positive — a 0x0 pty makes every pane spawn fail with ghostty error -2). Each failure message names the consequence rather than the rule, because the rule alone is what someone deletes.

A commented-out mention inside the file's own DO-NOT-add block must not trip the test — that comment is the whole point of the ticket, and a vacuity guard pins that it is still there, that all three files exist, and that none is trivially small. Measured on merge: 1420 green; adding ProtectHome=read-only to herdr.service fails 1 test; removing the login shell from herdr-inner.sh fails 1; removing its stty fails 1; deleting that file errors 3 and fails the build rather than passing vacuously.

The host no longer idle-sleeps while a member is working

fleetd now keeps the machine awake for as long as at least one member is live, and lets it sleep again once the last one goes. On macOS it does this by holding a caffeinate -i child process.

The knob. A new top-level block, on by default:

idleSleepGuard:
  enabled: true    # set false to turn the guard off

It is a deferred key: the daemon reads it once at startup, so a change needs a restart. A reload reports it as such rather than pretending it applied.

Why it exists. The Mac that runs this fleet was set to idle-sleep after one minute on battery (pmset -g custom reported sleep 1). Over one night the daemon's AMQP link dropped 13 times, and every drop had a sleep or wake event in pmset -g log in the same minute or the minute before. The broken AMQP link is only the visible symptom. The real cost is a member that freezes mid-turn with the host — and a long turn with nobody typing is exactly the case that goes idle. Before this, the fleet needed a human sitting at the keyboard to keep running, which defeats the point of delegating long work.

How it knows. The guard hangs off SessionManager's existing onAcquire/onRelease hooks and its size(). It does not count members a second way, so it always agrees with the numbers fleet_list reports. Only a real 0→1 or 1→0 crossing touches the OS.

Gotchas — three, and all of them are by design.

  • It is macOS-only. caffeinate ships on no other platform, so on Linux — fleet01, for example — the guard is a clean no-op. It says so once at INFO and never again, so a daemon running for weeks does not fill its log. fleet01 has the same exposure and does not get the fix from this change; a Linux mechanism (systemd-inhibit) is a separate job.
  • -i is idle sleep only. Closing the lid still sleeps the host, and so does an operator asking for sleep. That is deliberate: the guard stops an unattended host sleeping under a member's turn, it never overrides the operator. caffeinate -s/-d would do that and are not used.
  • It fails safe, and silently. If the mechanism cannot start — binary missing, process table full — acquire() returns null and nothing is ever held. The guard never throws, and never blocks a spawn, a release or shutdown. So "the guard is enabled" is not proof the host is awake. The proof is a live caffeinate process, or a member that survives an idle night.

fleetd #355 / #374.

A reply says whether anything was waiting for it

What. fleet_reply and POST /sessions/{id}/reply now report which of three things happened to a worker's reply, instead of saying "delivered" for all of them:

Outcome What happened REST delivered
resolved_send A fleet_send or fleet_ask was actively waiting, and took the reply now true
resolved_async_ticket No live waiter, but the reply completed a parked async ticket — a fleet_poll caller sees it at once true
queued Nothing was waiting. The reply is held in the inbox for a later drain false

Over MCP the tool result carries the wording, for example queued — no send or ticket was waiting; held in the inbox for a later drain. Over REST the body gains an outcome field, and delivered stops being a constant.

The knob. None. It is how both doors answer now.

Why it exists. All three outcomes are successes, but they are not the same fact, and the caller could not tell them apart. "Delivered" for a reply nobody was waiting for is the report reading better than the state — the same failure this project keeps finding in other places. A lead that sees queued knows its send never opened, or had already timed out, which is a real and different situation from a clean handoff. Before this, that difference was visible only in a metrics label nobody reads during a task.

The nudge counters were renamed in the same change, for the same reason. fleet_push_nudges_total and fleet_lead_heartbeat_nudges_total used to record an outcome called delivered; it is now sent. Nothing about the count changed — only the word. That call is a one-way herdr paste-and-submit into a pane, and there is no read-receipt concept at that layer, so "sent" is the most the counter can ever honestly claim.

Gotchas.

  • delivered: false over REST is not an error. The status is still 200, and the reply is safely held. Any client that treats delivered: false as a failure and retries will queue the reply twice. This is the one behaviour change that can break an existing caller, and it is the reason the outcome field exists: branch on outcome, not on delivered.
  • queued does not mean the lead will never see it. It means nothing was waiting at that moment. The reply is drainable with fleet_poll{target}, and the push loop nudges the lead.
  • The wording is not a contract; outcome is. The human-readable text may be reworded. The three outcome values (resolved_send, resolved_async_ticket, queued) are the stable names.

fleetd #365.

Startup says which profiles have usage-limit detection turned off

What. exhaustedPattern is the per-profile regex that recognises "you are out of quota" in a backend's own words. It is opt-in, and a profile without one has usage-limit detection off. Two places now say so:

  • At startup the daemon logs one aggregate WARN naming every profile with no exhaustedPattern, plus a second, louder WARN if any of them is a subscription profile. When every profile is armed it logs an INFO instead, so the healthy case is also on the record.
  • In the API fleet_profiles and GET /profiles carry exhaustionDetectionArmed, a boolean per profile. A lead can read the state without shell access to the daemon's log.

The knob. None to turn this on. The knob it reports on is profiles.<name>.exhaustedPattern; set one to arm detection for that profile.

Why it exists. A missing exhaustedPattern failed in the worst way an opt-in can fail: nothing was wrong, nothing was logged, and the profile kept accepting spawns. When the credential really hit its limit, the backend said so in prose, fleetd did not recognise it, and no quarantine started — so the fleet kept spawning members onto a dead credential. The operator's only clue was members that spawn fine and produce nothing. A subscription profile is the sharp case, because a subscription is the thing that actually runs out.

Gotchas.

  • Armed is not correct. exhaustionDetectionArmed: true means a pattern is configured, not that it matches what this backend prints. A wrong pattern reports as armed.
  • After a reload, the field lies — fleetd #404, open. exhaustedPattern is a deferred key: the patterns are compiled once into a map at startup and a reload never re-reads them. The field reads the live config instead. So if you add a pattern to fleetd.yaml and reload, the field flips to true while detection stays off until the daemon restarts. Until #404 lands, trust the field only on a freshly started daemon, and restart after editing exhaustedPattern — the reload report is the honest door here, and it names profiles as needing a restart.
  • errorPattern has the same opt-in shape and is not covered by this report. It is less severe: with none set, the code falls back to a narrow built-in pattern rather than going inert.
  • The startup call site is not pinned by a test. Deleting reportExhaustedPatternGap(cfg) from Fleetd.java leaves the suite green (measured at the merge: 1472 tests, 0 failures). The report's own behaviour is tested; that it is still called is not. This is true of all four startup reports, not just this one — reportGitHostShape, reportMemberTrustModel, reportMemberCredentialsGap and reportExhaustedPatternGap are each referenced by exactly one test file, and that test calls the method directly. fleetd #442 owns closing it. The six FleetConfig.validateXxx() startup calls had the same defect and it is now FIXED: they were collapsed into one cfg.validateAll(), which FleetdStartupValidationTest pins by calling the real Fleetd.main and asserting it refuses a bad config. (An earlier version of this line said "fleetd #398 owns closing it". That was wrong: #398 is the closed models allow-list PR.)

Two limits of the detection this reports on. Both bound what any recovery feature can do, so read them before designing one.

  • Detection needs a worker turn that finished. CompletionResolver classifies a usage limit by matching the member's own terminal output when its turn completes. There is no HTTP status or header path into this. So fleetd cannot learn a limit has been hit until a member has run and ended, and it cannot cheaply ask "is the limit lifted yet?" — any auto-resume has to spend real backend work to find out.
  • No reset time is delivered, anywhere. The captured live refusal is The usage limit has been reached. Try again later. — there is no time in it, and nothing in the refusal path reads a Retry-After. BackendQuarantine stores now + cooldownNanos and nothing else. Recovery can therefore only ever be a fixed timer or a backoff probe, never a resume scheduled for the real reset moment.
  • Every subscription: true profile shares one quarantine key. With no explicit credentialId, effectiveCredentialId() returns the <subscription> sentinel, which is right for a single Claude plan. The consequence is easy to miss: on this fleet the lead's own profile (opus) shares that key with the worker fan-out profile (sonnet). Arm exhaustedPattern on sonnet alone and worker exhaustion starts quarantining the lead's seat too. Arm them together, or not at all.

fleetd #395.

A central allow-list of the models the fleet may use

What. A top-level models: block names every model id the fleet is allowed to run. Once it is non-empty, a profiles: entry naming a model that is not on the list refuses to start, and refuses a reload too.

models:
  allow:
    - model: claude-sonnet-5
    - model: openai/gpt-5.6-terra
    - model: gx/deepseek-v4-flash

model: is one flat, opaque string namespace. A bare Claude id and a provider-prefixed opencode id both fit unchanged, because the check is exact string equality — it never parses a provider prefix and never branches on kind:.

The knob. models.allow. Absent or empty keeps the old behaviour, where no model was ever checked, so this ships inert until an operator writes the block.

Why it exists. Before this there was no single place that said which models the fleet may use. Each profile named one, and a typo or a withdrawn model id reached the backend adapter as a free-form string. On an opencode profile that has a specific bad outcome, already recorded here: a withdrawn model name makes opencode fall back to a paid model silently. An allow-list turns that class of mistake into a refusal at load, which is the cheapest place to find it.

Why it is operator-owned and not checked against a vendor catalogue. For an opencode profile fleetd synthesizes the provider from the provider/model selector plus baseUrl (OpenCodeLauncher:562-583). So a perfectly valid fleetd model id can appear in no published catalogue — gx/deepseek-v4-flash on this fleet is exactly that, absent from models.dev and correct. Any attempt to validate the list against a vendor catalogue would reject working configurations. The list is the source of truth; nothing else can be.

Gotchas.

  • Adding a profile means adding its model in the same edit. Otherwise the daemon will not start. That is the intended trade: the failure is loud and immediate rather than silent and later.
  • models: used to be deferred. Since fleetd #422 it is hot. A reload validates the new block, so a bad edit is still refused and the running config kept — and a good edit now takes effect on the very next spawn, with no restart. If you are reading an older note that says a models: edit needs a restart, that note is stale.
  • The list is a name gate, not a capability check. It says the operator permits this id. It does not say the backend serves it, that the credential may use it, or that the id is spelled the way the provider spells it. A model on the list can still fail at spawn.
  • Cross-check the two lists rather than trusting either. grep -E '^\s+model:' fleetd.yaml | awk '{print $2}' | sort -u against the allow: entries. If they differ, the daemon refuses to boot — which is the point, but it is better to know before a restart than during one.

fleetd #398.

fleet_reply tells a lead which tool to use instead

What. A lead that calls fleet_reply is refused, and the refusal names both working routes:

fleet_reply has no route to a peer lead. Use fleet_send{coordId: ...} for a peer on another
daemon or fleet_send{sessionId: ...} for a peer on this host. fleet_reply resolves a member's
blocked fleet_send, and a peer's coord-id message is durable and non-blocking, so there is
nothing for it to resolve.

The knob. None. The check runs on every fleet_reply call and needs no configuration.

Why it exists. A lead that gets a message from a peer reaches for "reply" — the word matches what it is doing. But fleet_reply exists to resolve a member's blocked fleet_send, and a peer lead never has one open. So the call used to be accepted and the message went into a worker inbox nobody would ever drain. The peer waited on nothing, and no error said so.

The old refusal was worse than none, because it only covered the case where the caller could not be identified at all. A lead is identified, so it sailed past that check.

The message names the routes on purpose. A refusal that says only "not allowed" makes the lead guess, and the two right answers depend on where the peer lives — coordId across daemons, sessionId on this host.

Gotchas.

  • The order of the two checks matters, and is now pinned. An unidentified caller gets the "workers only" message even if its role is PRIMARY. That is deliberate: "we do not know who you are" is the more useful thing to hear first.
  • A queued reply is still a success for a worker. This refusal is about leads only. A worker whose fleet_reply finds no open send still gets its reply stored in the inbox, and that is normal.
  • Role comes from the connection, never from an argument, so a caller cannot present itself as a worker to get around this.

fleetd #391.

The credential-scrub receipt now measures the blank, not the attempt

What. The startup scrub that clears inherited credentials from a member's shell reports which names it actually blanked. It now checks each parameter's value after trying, instead of trusting the exit status of the attempt.

allowed 41 of 57 failed 3
<name>      ← blanked, confirmed empty
!<name>     ← could not be blanked

The knob. None. The receipt is part of the scrub and always runs.

Why it exists. The old loop counted a name as blanked when eval "export ${n}=" returned 0. zsh has integer parameters — SECONDS RANDOM SHLVL HISTSIZE COLUMNS LINES USERNAME — and on those an empty assignment is coerced to a number rather than failing. So eval returns 0 and the value is unchanged. Measured across 10 names: 7 false receipts.

The general rule, which cost two tickets to learn: an attempt's exit status is not a measurement of its effect. The fix verifies afterwards with [[ -z "${(P)n}" ]] and only then counts it.

Gotchas.

  • eval is required, not stylistic. export UID= is a fatal zsh parameter error that aborts the whole sourced file — it once killed the scrub at name 42 of 57 and left the operator's own exports untouched. Only eval "export ${n}=" 2>/dev/null contains that. The same applies to EUID GID EGID PPID LINENO.
  • The ! names are the ones to read. They are shell parameters that cannot be blanked, not credentials that leaked. A long ! list is normal; a shrinking allowed count is the warning.
  • The scrub is zsh-only, like the secret store it defends against. It lives in .zshenv, the one file zsh always reads.
  • The receipt reports the member's shell, so a claim about it measured from inside an agent's zsh -c child is measuring the wrong process.

fleetd #394, #400.

fleet_list no longer advertises a profile that fleet_spawn will refuse

What. The capacity rows in fleet_list now list the profiles the daemon can really spawn on — the set it read at startup. Before, they listed the set in the live config, so a profile added by a hot reload showed up with free slots and every spawn onto it failed.

# operator adds profile "ghost" to fleetd.yaml and the config reloads
fleet_list   -> capacity: [... {profile: ghost, maxLoad: 3, live: 0, free: 3}]
fleet_spawn{profile: "ghost"}  -> refused: unknown worker profile

The knob. None. This is the capacity reporting in fleet_list, and it always runs.

Why it exists. profiles is a "both ways" key. Per-profile tunables — maxLoad, weight, credentialId — are hot, so a reload changes them at once. The profile set is frozen, because HerdrPeerLauncher takes Map.copyOf(profiles) once when it is built and never looks again. So "profiles is deferred" and "this live read is fine" are both true, of different halves of the same key. The old wiring read the live keySet() and the hot maxLoad() through one lambda, which looked consistent and was half wrong.

The direction of the error is what made it worth a ticket: it overstated a capability. A lead reading free: 3 had no way to tell that number from a real one, and only found out by spawning.

Gotchas.

  • maxLoad is still hot, and must stay hot. The fix must freeze only the set. A test pins that: it changes maxLoad from 3 to 9 by reload and asserts the same CapacitySource instance reports the new number. Freezing both would be the mirror-image regression.
  • The right rule was already written three lines below, for coordinator.peers: read from the same snapshot the collaborator itself was opened from, rather than the live config, so the report follows one rule instead of half hot-reloading.
  • A fresh-daemon test can never catch this. On a daemon that has not reloaded, the live config and the startup snapshot are the same object, so both the wrong and the right wiring pass. The test has to perform a real reload and assert the reload applied, or it proves nothing.
  • Both directions need a test. A single test that only adds a profile passes if the source returns a permanently empty set. The second test — a startup profile is listed — is what makes the first one load-bearing.

fleetd #416. Found by the fleet01 lead on a live daemon; the same shape as #404, where a status field read a different source than the behaviour it described.

The startup log tells "off" apart from "running on the built-in default"

What. Two startup lines report whether a backend-classification is running. The backend-error line no longer says off when no profile sets an errorPattern, because that classification is still running — it falls back to a built-in pattern. The backend-exhausted line still says off, because for that key it is true.

backend-error classification: built-in default for all profiles
    (no profile customises errorPattern; profiles: [gx, local, opus, xf])
backend-exhausted classification: off
    (no profile has an exhaustedPattern configured; profiles: [gx, local, opus, xf])

The knob. None. Both lines are logged at INFO at every startup.

Why it exists. The old line was produced by one shared helper that knew only the key's name, so it worded both keys the same way. But the two keys disagree on what "unset" means:

  • an unset errorPattern falls back to CompletionResolver's built-in (?i)\bAPI Error\s*: compatibility pattern, so classification keeps running;
  • an unset exhaustedPattern has no fallback, so it really is off.

One string, two meanings — and the false one was the reassuring direction. It told an operator a live classifier was disabled while it was running, which is the worst thing to read when you are debugging a false quarantine. The fact needed to word it correctly was written down about 1600 lines away, in FleetConfig's javadoc, and the method that got it wrong had no way to reach it.

Gotchas.

  • The classification really is live even with no errorPattern anywhere. The built-in pattern is narrow but it can still fire on a member's own prose that quotes an API Error: line. That is why the outage policy needs the line twice before it acts.
  • "no profile customises it" is the useful reading, not "nothing is configured". If you want a different pattern for a backend, set it; the absence of the key is a choice of the default, not an off switch.
  • off for exhaustedPattern is real. Do not assume symmetry between the two lines — that assumption is the bug this entry describes.
  • The wording is now the caller's responsibility. coverage() takes a required UnsetMeaning argument, so a third pattern key cannot compile without stating what unset means for it. There is no permissive default to inherit.

Measured, and worth keeping. The first fix was correct and still not safe: its tests all passed the meaning in themselves, so they proved the wording and not the pairing. Swapping the two arguments at the two call sites recreates the original bug with the keys exchanged, and that left all 1506 tests green with BUILD SUCCESS. The call sites are now extracted into two named factories and pinned by FleetdPatternCoverageLineTest; the same swap gives 3 failures. The general rule: when a fix adds a parameter so the caller can supply a missing fact, test the caller's choice of value — a test that passes the value in itself tests the half that was never broken.

fleetd #415. Found by the fleet01 lead on their own startup log.

Turning a model off at runtime, without editing profiles:

What. Each entry in models.allow now takes an enabled: flag. Set it to false and the fleet stops spawning on every profile that names that model — at once, with no restart and no edit to profiles:. Set it back to true and those profiles are usable again.

models:
  allow:
    - model: claude-sonnet-5
    - model: openai/gpt-5.6-terra
      enabled: false          # every profile naming this model is now unspawnable
    - model: gx/deepseek-v4-flash

enabled: is optional and defaults to on, so every existing models: block keeps working untouched. An entry that is off stays in allow — do not delete it. allow answers "may the fleet use this id at all", and enabled answers "may it use it right now". Removing the line instead of turning it off makes every profile naming that model fail validation, and the whole reload is refused.

Seeing the gate's own state

fleet_profiles and GET /profiles report the gate, so a lead can see it without reading the config file. Two fields, and you need both:

field meaning
modelGateArmed is there a models: block at all. Always present, true or false.
modelsOff which model ids are off right now. Absent when nothing is off.

The off set alone cannot answer the question you usually have. An empty off set means one of two very different things — there is no models: block on this host, so nothing is gated and nothing can be; or there is a block, it is working, and right now nothing is turned off. modelGateArmed separates them, which is why it is reported even when it is false.

The daemon logs the same three states at startup, from the same read:

model gate (fleetd #422): not configured (no models: block — nothing is gated, and nothing can be)
model gate (fleetd #422): armed (models: block present; 0 models currently turned off)
model gate (fleetd #422): armed (2 model(s) turned off: [openai/gpt-5.6-terra, sol/x])

Both the log line and modelGateArmed come from one PeerLauncher.modelGateState() call, which is the same accessor the spawn gate itself reads. So the status can never claim more or less than the gate enforces, and a reload landing between two reads cannot make them disagree.

The knob. models.allow[].enabled. Absent means on.

Why it exists. A subscription runs out. When it does, the profiles on it must stop taking work while the rest of the fleet keeps going. Before this the only ways to do that were to edit every affected profiles: entry, or to let each spawn fail against the dead backend and wait for the quarantine to catch up. One central switch keyed by model is the right shape, because a model is what a subscription sells — several profiles usually share one.

Where the gate sits, and the one thing to know about it. fleet_spawn takes two different paths, and they treat a bad profile in opposite ways:

flowchart TD
    A["fleet_spawn"] --> B{"did the caller<br/>name a profile?"}
    B -->|"yes — an operator override"| C["check quarantine, cooling off,<br/>maxLoad, model-off"]
    C --> D["REFUSE the spawn"]
    B -->|"no"| E["placement: build the<br/>quarantined / coolingOff / modelOff sets"]
    E --> F["the policy SKIPS those profiles"]
    F --> G["spawn on the next candidate"]
    classDef stop fill:#b7791f,stroke:#7b341e,color:#ffffff;
    classDef go fill:#2f855a,stroke:#22543d,color:#ffffff;
    class D stop
    class G go

Naming an off-model profile explicitly is refused, and that is deliberate — it is the operator overriding, so a clear refusal beats a silent redirect. Leaving the profile blank routes around it instead. Both behaviours are correct; they are just not the same, and code that resolves a profile name before calling spawn silently moves itself from the second path to the first.

Gotchas.

  • A host with no models: block is not gated. The whole feature is inert there, and that is permitted, never fatal. On this fleet, fleet01 has no models: block, so the gate ships green and does nothing on that host. Check with grep -c '^models:' fleetd.yaml before assuming a second host is protected.
  • The gate is a name gate, like allow itself. Turning a model off stops fleetd spawning on it. It does not reach a member that is already running.
  • Off is not quarantine, and the message says so. A model-off refusal names models.allow; a quarantine names the credential and the seconds left. When a profile is both, quarantine is reported, because a backend fact outranks an operator preference.
  • fixed is the default policy, and it applies this filter in two places — the default fast path and the fallback walk. The first round of this feature put the filter in the shared helper PlacementPolicyUtil.available(), which fixed never calls, so it shipped green and inert under the default policy with 1519 tests passing. Both new tests had used weighted().

Still missing. Detecting a subscription limit and flipping the switch automatically is not built. Today an operator turns the model off and back on by hand. That decision is open.

fleetd #422.

Removing an architect slot now actually revokes it

What. An architect's extra rights come from a slot in fleet.architects. Delete that slot from the config and reload, and the session bound to it drops to worker rights on its very next request. Before, it kept full architect rights until the daemon restarted.

The knob. None. This is how fleet.architects behaves on reload.

Why it exists. An architect can do things a worker cannot, so "revoke" has to mean revoke. The old code flattened the slot list once when the daemon started and never read it again, so removing a slot only refused the next spawn. The session already holding the privilege kept it — a revocation that did not revoke, with nothing in the logs to say so.

Gotchas.

  • The privilege is revoked; the slot stays occupied. The demoted session keeps its slot key until it unbinds. That is on purpose: freeing the key at once would let a second terminal claim the slot the operator was actually trying to shut down, and it would break unbind, which needs the original terminal-to-slot pair intact.
  • The demoted session can still finish its turn. fleet_reply is authorised by terminal identity, not by role, so a demoted architect ends its turn normally instead of stalling.
  • Do not confuse the binding with the privilege. They are separate, and wording that treats them as one thing is what produced the first, wrong version of this fix — the ticket asked for a test that a bound architect "survives the rebuild", which is true of the binding and false of the privilege.

fleetd #424.

A worker's fleet_list no longer carries the lead's coordination state

What. fleet_list's reply used to include a coordinator block for every caller: this daemon's own coord-id, its mailbox state, a preview of the peer mail held for it, and each configured peer's live reachability. That is lead-to-lead state. Now a worker or an architect calling fleet_list gets no coordinator key at all. The key is absent, not present and empty.

The knob. None. It follows the caller's role, which the daemon resolves from the connection, not from anything the caller sends.

Why it exists. A worker has no use for peer names, held-mail previews, or this daemon's coord-id, and it should not learn them from a roster call it makes for other reasons. Absent beats empty on purpose: an empty object still tells the caller the feature is configured, and it makes a client that tests if (coordinator) behave differently from one that tests coordinator.peers.length. So the gate runs before the row is built.

Gotchas.

  • A worker cannot tell "no coordination configured" from "not for you". Both look like a missing key. That is the intended trade: the alternative leaks the fact that peers exist.
  • The lead sees no change. Same key, same fields.
  • The compat overloads used to default to showing the row, and no longer do. listFleet has seven declarations. Six of them do not take the caller's role, and they all inherited the default from one line. That default was true, so a call site that forgot the argument would have disclosed the row silently. It is now false: a missing identity means a missing row, which is a visible bug rather than a quiet disclosure. Fixed in fleetd #463. A test calls a compat overload with no boolean and asserts the coordinator key is absent.

fleetd #439.

A usage limit now says which model to turn off, and detection can be armed without a restart

What. Three related changes to how fleetd reacts when a backend reports that a subscription limit is reached.

  • exhaustedPattern is now a hot config key. Before, it was compiled once into a map at daemon startup, so arming detection for a profile needed a restart. Now the pattern is read live per profile, and compiled patterns are cached by pattern string so a check does not recompile a regex every time.

  • The warning logged when a credential is quarantined names the fix, not just the fact:

    usage-limit fix: profile 'X' runs model 'Y' — set `enabled: false` on that model's entry
    under models.allow in fleetd.yaml to stop new spawns landing on it (models: is hot, no
    restart needed); remove the line again once the subscription window resets
    
  • fleet_profiles reports two new fields on a quarantined row: model (the model that profile runs) and reason (the backend text that triggered the most recent quarantine of that credential).

The knob. exhaustedPattern: on a profile, which is opt-in and off by default. It is now hot, so a reload arms it. Turning a model off is enabled: false on its models.allow entry, which was already hot.

Why it exists. The two halves of model gating were reloadable in opposite directions, and it was the wrong way round. An operator could already turn a model off at runtime, but could not arm the detector that tells them to — that needed a restart, on a feature whose whole point is reacting to a limit while the fleet is running. The reason field exists because "this credential is quarantined" does not say whether the cause was a usage limit or something else, and the two need different responses. The warning names the fix because an operator reading a quarantine line should not have to work out which of several configured model names to edit.

Gotchas.

  • reason is only written where a quarantine actually happens, keyed by credential id — the same key the cooldown itself uses. So a reason can never be reported for a quarantine that did not occur.
  • A pattern has to match real text from that specific backend. Do not guess one. A guessed regex gives you a profile that reports exhaustionDetectionArmed: true and silently never fires, which is worse than an honest false.
  • There is no off: config key. The flip is enabled: false. off is the name of the reported set (modelsOff), which is an output.
  • Nothing re-enables a model automatically. "On again when the limit lifts" is a manual edit, which is cheap because models: is hot and needs no restart. An automatic backoff probe was considered and rejected: a probe spends quota to discover quota, so on a metered subscription it burns the first tokens of every new window on discovery instead of work.
  • Nothing proves the daemon wires this sink. The quarantine sink was extracted into a factory and what it logs is pinned by a test, but replacing main()'s call to that factory with an inert lambda leaves the whole suite green. Tracked in fleetd #460.

fleetd #446.

A lead session can replace itself when its context fills up

What. A lead (primary) session runs out of context and has to be replaced by a fresh one. Until now that was entirely manual: the lead wrote a handover file, told the operator where it was, and the operator started a new session by hand and pointed it at the file. Now the lead can ask fleetd to do the swap.

The cycle has three steps, and the order matters:

  1. fleet_handover{action: "open"} — fleetd records a token and tells the lead the handoverPath it must write to.
  2. The lead writes the handover file. The handover skill is the procedure for what goes in it.
  3. fleet_handover{action: "confirm", token, operatorConfirmed} — fleetd checks every gate and, if all of them pass, schedules the roll: wait for the lead's own turn to end, send /clear to its pane, wait again, then send a bootstrap prompt naming the handover file. A fresh session reads the file and carries on.

{action: "cancel", token} drops a pending request without rolling.

The knob. A new top-level leadRollover: block in fleetd.yaml. It is opt-in and off by default — with the block absent, fleet_handover is still registered but every action answers a clean refusal naming NOT_CONFIGURED.

leadRollover:
  handoverPath: .handover/HANDOVER.md  # required when the block is present; relative is allowed
  requireOperatorConfirm: true         # default true
  maxDocAgeSeconds: 3600               # default 3600
  turnSettleSeconds: 20                # default 20
  clearSettleSeconds: 20               # default 20
  bootstrapText: "…"                   # default names handoverPath

Every field is read fresh on each call, so the values are hot. Adding the block where it was absent at boot still needs a restart, because Fleetd.main only constructs the executor when the block is present in the startup snapshot. That is an existence fact, not a stale-value one — the same way a brand-new profiles: entry needs a restart while an existing profile's fields do not.

Why it exists. A lead that fills its context is the one agent nobody else can replace: it holds the orchestration state, and a worker cannot restart its own lead. The whole cycle previously stopped dead waiting for a human to notice. Three design choices were deliberate and are worth not re-litigating:

  • The lead asks; nothing watches it. There is no timer, no heartbeat, and no background loop that can decide on its own that a lead should be replaced. Only an explicit confirm() call that passes every gate can ever cause a /clear. A context-pressure detector was considered and rejected — the cost of a false positive is a destroyed live session.
  • /clear in the same pane, not stop-and-relaunch. Relaunching would lose the pane's identity, and a lead is pinned to its terminal by a tab label (fleet.leaders.<name>.tab), so a new pane is a new lead as far as the daemon is concerned.
  • The operator confirms too. requireOperatorConfirm defaults to true, so the lead's own judgement is not enough to wipe a session.

Gotchas.

  • Write the file after open, never before. confirm refuses with HANDOVER_STALE unless the file's modified time is later than the open request. That check exists so a leftover file from a previous session can never be accepted as this session's handover, and it means the obvious order — write the file, then ask — is the wrong one.

  • A relative handoverPath resolves against the lead's workspace, not the daemon's. The daemon and the lead run in different directories — on this host the daemon sits in <repo>/fleetd and the lead in <repo> — so a bare relative path would mean two different files. open() resolves it once, against the calling lead's configured fleet.leaders.<name>.cwd (falling back to the daemon's own working directory when that lead has no cwd), and hands back an absolute path. Everything downstream — the open response the lead writes to, the file the daemon stats, and the bootstrap prompt the fresh session reads — uses that one absolute path. An absolute handoverPath is used unchanged.

  • A relative handoverPath lands in the LEAD's repo, which is usually not the repo that ignores it (#491). The handover file is a snapshot of live state — unpushed branches, open questions — so it must never be committed. The trap is which .gitignore protects it. On this host the lead's cwd IS the fleetd checkout, so the rule in fleetd/.gitignore works and the distinction is invisible. On fleet01 the lead's cwd is /home/ltms/LTMS/kb while fleetd sits in /home/ltms/LTMS/fleetd — two different repositories, and kb/.gitignore has no .handover rule. Measured 2026-09-12. So: put the ignore rule in the lead's workspace repo, or give an absolute path outside every repository. Re-check with grep -n handover <lead workspace>/.gitignore — no output means the trap is live on that host.

  • confirm returning accepted does NOT mean the pane has been cleared. It means every gate passed and the roll is scheduled. The roll itself runs after the calling turn ends, and it may still refuse at that point; those outcomes are logged only, because there is no caller left to answer. Look for lead-rollover: lines in the daemon log.

  • A lead can only ever roll itself. The tool has no terminal, session or lead parameter of any kind. The pane comes from the caller's own connection. This daemon can hold more than one labelled lead tab, and an earlier draft that looked the pane up in a single-slot registry let one lead clear another lead's pane.

  • /clear must never go through Injector. It produces no turn boundary, so the Injector's turn never completes and every later message to that pane queues behind it for ever — the pane is wedged. LeadRollover calls agents.send(...) directly, exactly like ClaudeCodeLauncher#clearContext. Measured, not assumed.

  • The settle waits require IDLE or DONE, not injectable(). injectable() also accepts BLOCKED, which is a live turn paused on an approval prompt. Treating that as settled sent /clear into an open prompt mid-turn.

  • The first real roll failed, and the fix is merged but not yet proven live (#489). On 2026-09-12 the roll put one line into the pane — /clearFresh lead session. … — and Claude Code answered Unknown command: /clearFresh. That is /clear and the bootstrap text joined with no space, because a paste lands at the cursor. Two faults. The second settle wait was a no-op: /clear starts no turn, so the pane never leaves IDLE and the wait returned on its first poll — the whole roll ran in 438 ms of a 20-second budget. Under that, the submit keystroke raced the paste, which AgentControl.submit's own javadoc already records as CB-113; LeadRollover bypasses Injector on purpose, so it got none of the Enter-nudging that makes a /clear land everywhere else. The failure was safe: nothing was cleared and no context was lost. PR #490 replaced the second wait with waitForClearPickupAndSettle, which nudges the submit keystroke while no pickup has been seen — the same pattern Injector already ships for its own post-turn /clear (#306). Merged and deployed on 2026-09-12. Acceptance criterion 6 is still not met: no roll has yet bootstrapped a fresh session end to end, and only using the feature for real on a live lead can settle it. Until then, treat the manual path as the reliable one.

  • A settle wait that never settles hangs the test suite instead of failing it (#486). The poll loop is bounded only by an injected clock, and the test seams pass a clock that never advances together with a no-op sleeper. Both waitUntilAtTurnBoundary and waitForClearPickupAndSettle have this shape. In production the bound is real; in the suite it is inert.

fleetd #480 (PRs #483, #484, #485). Related: #486, #489 (PR #490).

A timed-out send now says whether delivery was even attempted

Before this, a fleet_send that timed out reported one of two things: TIMED_OUT_WORKING if the message was delivered, or TIMED_OUT_QUEUED if it was not. There was no third answer, so a case that is neither got filed under "not delivered".

That case is real. When the injector cancels a delivery it reports DELIVERED, NOT_DELIVERED, or ATTEMPTED — meaning the keystrokes were already going out and nobody can say whether they landed. The send path collapsed ATTEMPTED into TIMED_OUT_QUEUED, which promises the message never arrived. A lead reading that promise retries, and the worker gets the same brief twice.

What it does. MessageService.Outcome gains TIMED_OUT_UNCONFIRMED. The ATTEMPTED cancellation now routes to it instead of to TIMED_OUT_QUEUED. Every reader handles it:

  • MCP — fleet_send returns [no reply within Nms — delivery unconfirmed; the message may already have reached the worker, so a retry risks sending it twice — poll status before resending]. It deliberately does not carry the "retry or poll status" wording the queued/working arm uses, because on this route a resend can double-deliver.
  • REST — POST /sessions/{id}/messages answers 202 with "status":"unconfirmed".
  • Metrics — the fleet_sends counter still labels it timeout, grouped with the other two timeout outcomes. That grouping is deliberate and was left alone.

The knob. None. It is a behaviour change on an existing path, live as soon as the daemon restarts.

Why it exists. A sentinel that means "no" and a sentinel that means "cannot tell" need opposite handling from the caller, and this code had only the confident one. Fixing it needs a third state, not a better guess — the same shape as fleetd #512.

The gotcha, and it is the interesting one. FleetApp.writeReply's inner switch carried default -> "done", so any outcome it did not name told a REST caller the delegation completed. Adding a constant would have been absorbed silently by that default. The fix deletes it and lists all ten outcomes by name, so the compiler now catches the next missed one. Order matters if you repeat this exercise: a default is exactly what suppresses the compile error you are trying to provoke, so delete the defaults first, then add the new constant, or the proof comes back clean and proves nothing.

Still open. The outcome's wire token is its constant name lowercased, not a pinned string — the sibling enum ReplyOutcome does pin its own. That is fleetd #578, which has since grown: the same enum emits four different tokens across four live surfaces, so the fix is not one accessor.

fleetd #571 (PR #580). Related: #512, #578, #586.

fleet_list and /healthz report whether the background loops are alive

Two loops keep the fleet honest: StatusPoller, which refreshes member status, and SessionReaper, which retires idle members. Both have had a LoopWatchdog for a while. Nothing outside the daemon could see it.

What it does. fleet_list gains a loopHealth object with keys statusPoller and sessionReaper, each RUNNING, STALLED, or STOPPED. /healthz carries the same object in its body. The 200/503 status codes are unchanged — a stalled loop does not turn the endpoint red.

The knob. None. Always on.

Why it exists. A stalled poller does not announce itself. The fleet keeps answering, member status quietly goes stale, and the first symptom is a lead acting on state that stopped updating hours ago. Exposing the watchdog turns a silent failure into a visible one.

The gotcha — the failure direction. This is monitoring, so ask what happens when the monitor lies. The dangerous direction here is a false negative: report RUNNING while the loop is dead, and the watchdog can never fire. That is worse than a false positive, because a false positive is noisy and somebody mutes it, whereas a false negative produces no signal for anyone to notice is missing. The wiring is what protects against it, and the wiring is now pinned: FleetdLoopHealthSourceWiringTest fails if Fleetd.loopHealthSource stops asking the real poller or the real reaper. That test exists because the first version of this feature shipped with five green tests and the production wiring could still be replaced by a constant with nothing failing — every one of the five built its own LoopHealthSource, which tests the consumer and can never be evidence about the producer.

Reading it. STOPPED for sessionReaper is normal when no reaper is configured; that is the null case, not a fault.

fleetd #562 (PRs #579, #584).

hunter is a member role, not just a skill

The repo already shipped a hunter playbook skill: sweep a package, report several ranked findings, change nothing. There was no matching role. A hunt was run by spawning a dev and telling it not to behave like one.

What it does. MemberRole.HUNTER joins ARCHITECT, DEV and REVIEWER. It has the wire token hunter, the config pool fleet.hunters, and the agent file .claude/agents/hunter.md. fleet_spawn{role: "hunter"} now picks a contract instead of borrowing one.

The knob. fleet.hunters in fleetd.yaml, listing the profiles a hunter may run on — the same shape as fleet.developers and fleet.reviewers:

fleet:
  hunters:
    sonnet: {profile: sonnet}
    terra:  {profile: terra}

Optional. Leave it out and the role exists but no hunter can be placed.

Why it exists. A role carries prohibitions the brief should not have to repeat. dev is allowed to edit, commit, push and open a pull request; a hunt must do none of those. Spawning a dev for a hunt meant the only thing standing between the sweep and a surprise commit was a sentence in the brief. Miss that sentence once and the member is inside its contract while doing the wrong job. hunter.md forbids editing, committing, pushing and opening a PR, and — unlike reviewer.md — explicitly permits running the build, because a hunter checks its findings.

The gotcha — the role ships inert. Merging the code does not create the pool. On a host whose fleetd.yaml has no fleet.hunters, unknown nested keys are ignored rather than rejected, so there is no warning at load and no error at merge: fleet_spawn{role: "hunter"} simply finds no candidate profile. The sequence is merge, redeploy, add the pool, then spawn one hunter and read the spawn log. Parsing the config proves nothing about placement.

The second gotcha — hunter and reviewer are still not interchangeable. reviewer caps its answer at one finding; hunter reports several. Naming the wrong skill for the role hands the member two contradictory output contracts, and the measured result is a member that writes a good report to its terminal and ends the turn with no fleet_reply. The role does not fix that; the brief still has to name the right skill.

fleetd #568 (PR #596).

The context-roll notice now obeys requireOperatorConfirm

What it does. When an idle lead's own context reads HIGH, the heartbeat loop appends a notice to its nudge telling the lead how to hand over. The wording of that notice now follows the daemon's own leadRollover.requireOperatorConfirm setting. With the default (true) it tells the lead to ask the operator before confirming. With false it drops that instruction and tells the lead to decide for itself, naming the gate that actually applies: the handover file must exist, must have been changed after the open() request, and must not be older than maxDocAgeSeconds.

The knob. leadRollover.requireOperatorConfirm (default true). It is deferred, not hot — read once at boot, so an edit does nothing until the daemon is redeployed.

Why it exists. The daemon's own gate, LeadRollover.confirm(...), already honoured this flag at LeadRollover.java:480. So setting it to false did stop the daemon refusing a roll. But the text the lead reads is built by LeadHeartbeatLoop.contextNotice(...), which took no config at all and hardcoded "ask the operator … Only the operator can approve the roll". A lead follows the instructions it is given, so it asked the operator anyway. The operator was interrupted for a routine context roll exactly as before, and the config looked broken. Our operator asked for this directly: "is it intended or? if yes, fix this behavior".

The gotcha — the knob used to be half a fix, and the failing half was silent. Nothing warned that the message and the policy disagreed. Changing config that the instruction text does not read is invisible to the agent reading that text, and no test or log line caught it. If you set this knob and the asking continues, check whether the daemon has actually been redeployed since — the config half and the code half each need their own restart to take effect.

The second gotcha — a config value that reaches a message needs its own test. The enforcement path and the wording path are two consumers of one setting, and pinning the first proves nothing about the second. The two tests added here assert on the returned string for both values of the flag. Inverting the branch condition kills three tests, including one written before this change.

fleetd #621 (PR #622).