Features: loop health in fleet_list//healthz (#562), and TIMED_OUT_UNCONFIRMED (#571)
7-Use-Cases: re-sync the portable CLAUDE.md block after #562 added the
loopHealth row to the intent->tool table.
11-Features: two entries. The loop-health one records the failure
DIRECTION that matters -- a false negative where nothing ever fires is
worse than a muted false positive -- and why the wiring test exists.
The TIMED_OUT_UNCONFIRMED one records the default -> "done" that told
REST callers a delegation had completed, and the order dependence in
using a compile error to find enum readers.
fleetd #489: the fix is merged and deployed; criterion 6 is still unmet
Also records #486: a settle wait that never settles hangs the suite instead
of failing it, because the test seams pass a clock that never advances with
a no-op sleeper.
fleetd #489: record that the lead rollover's bootstrap prompt does not land
Criterion 6 of #480 was written as 'still unproven'. It is now measured
failing: the first real roll joined /clear and bootstrapText into one line
and Claude Code refused it as 'Unknown command: /clearFresh'. Nothing was
cleared, so the failure is safe, but the roll does nothing until #489 is
merged and redeployed.
Features: a lead session can replace itself when its context fills up
fleetd #480 (#483, #484, #485). Covers what the cycle is, the leadRollover:
knob and its defaults, why the three design decisions were made the way they
were (the lead asks and nothing watches it; /clear in the same pane rather than
stop-and-relaunch; the operator confirms too), and six gotchas.
Two of those gotchas are measured facts that cost real debugging: /clear routed
through Injector wedges the pane for ever because it produces no turn boundary,
and the settle waits must require IDLE or DONE rather than injectable(), which
also accepts BLOCKED — a live turn paused on an approval prompt.
The last gotcha records what is NOT proven: nothing has yet shown end-to-end
that a freshly cleared pane picks up the bootstrap prompt. Stated with the
bound on the damage if it does not, rather than left out.
7-Use-Cases.md carries the matching fleet_handover row in the portable
CLAUDE.md block template, byte-identical with the project's own copy.
Features: the redeploy entry names the redeploy-fleetd skill
The daemon-redeploy procedure moved out of CLAUDE.md, which loads in
every session, into .claude/skills/redeploy-fleetd/, which loads only
when invoked. The feature entry said "run the script" and did not say
how an agent finds the procedure, so a reader had no route from the
wiki to the skill.
Adds the route and what the skill carries. CLAUDE.md keeps the two
rules that must stay resident: a merge is not a deployment, and
workers never redeploy.
Features: the charter tool check now covers reload, and the quarantine row reports the attempt
fleetd #474 removed the limit this page warned about. The charter
tool-surface check ran at startup only, so a reload could install a
charter naming a tool the server does not register. ConfigRef.reload()
now runs the same check and refuses the whole reload. Replaced the
"until #474 lands" warning, which was about to tell a reader they could
not rely on something they now can.
fleetd #473 added quarantineAttempt beside quarantinedForSeconds in both
fleet_profiles and fleet_list capacity rows. Recorded what a large count
with a flat cooldown means (normal -- the ceiling caps the wait, not the
streak) and that both values come from one map read.
Features: opencode members actually receive seeded skills, and a charter's text is checked
Two entries updated, both for behaviour an operator can configure or observe.
memberSkills (fleetd #393): the entry said fleetd copies skill folders into
each worktree's .claude/skills/, which was the whole story only for Claude
Code members. Opencode never reads that directory, so an opencode member
was seeded correctly and read nothing. Added what now happens (each seeded
skill's SKILL.md goes into the member's instructions[] array), and two
gotchas: delivery is not activation (opencode has no Skill tool, so the
text is static system-prompt content from spawn and cannot be invoked by
name), and anyone adding a fourth writer to instructions[] must use
withArray and assert contents rather than size.
fleet.charters (fleetd #469 and #474): the entry described validateCharters()
checking the shape of the map only. Added the new check on the charter's
TEXT against the canonical tool set, and - more important - the limit of
it. That check runs at startup only. The entry's own existing sentence says
validation runs on both the startup and reload paths, and that is still
true of the shape checks but not of the new one. Charters are hot and read
live per spawn, so a charter edit without a restart is currently unchecked.
Said so plainly and pointed at #474.
Entry count unchanged at 157 - both are edits to existing entries, not new
ones.
Features: the exhaustion quarantine now escalates (fleetd #466)
The 'Stop spawning onto an exhausted account' entry described a flat
cooldown, which stopped being true with fleetd #466. Updated rather than
duplicated: the cooldown now doubles on each consecutive exhaustion of the
same credential, capped at 12x the base, so a chronically exhausted
credential is retried about a dozen times a week instead of ~336.
quarantineCooldownSeconds is now the BASE of the backoff, and the entry
says so where an operator reads what the key does.
Four gotchas added, all of them things an operator can get wrong:
- The reset is a TIME PROXY. Nothing reports a successful spawn back to the
tracker, so what clears the streak is a base cooldown of quiet, not
evidence the account recovered. Read it as 'we have not been told it is
still broken'.
- The multiplier and ceiling are constants, not YAML. No new knob, no new
hot/cold question.
- The ceiling is deliberate: an unbounded backoff is a permanent outage
only a restart clears, which is worse than the flat retrying it replaced.
- Cooling-off is a SEPARATE mechanism and is not escalated, because
escalating a 5xx storm would turn it into a multi-hour outage.
Entry count unchanged at 157 - this updates an existing entry.
Features: a usage limit names the model to turn off; the detector is hot
fleetd #446, merged as 1fb6176. One new entry:
- exhaustedPattern is now a hot config key, so detection can be armed without a
daemon restart. It used to be compiled once into a startup map.
- the quarantine warning names the fix: which model that profile runs, and the
exact edit to make.
- fleet_profiles reports model and reason on a quarantined row.
Also corrects the last gotcha on the fleetd #439 entry. It said the compat
overloads default to showing the coordinator row and pointed at #463 as open.
#463 shipped: the default is now false, so a forgotten argument means a missing
row rather than a silent disclosure. The count is corrected too -- listFleet has
seven declarations and six inherit the default; the old text said six overloads
without saying one of them writes it.
Both entries name what is still unproven: nothing tests that main() wires the
quarantine sink at all (fleetd #460).
Features: the coordinator row is lead-only; add the CB-548 quorum design page
Features entry for fleetd #439, merged as 92a96fc. A worker or architect
calling fleet_list now gets no coordinator key at all, and the gate runs
before the row is assembled so the key is absent rather than empty. Notes
the compat overloads that still default to showing it (fleetd #463).
Also commits CB-548-Lead-Quorum-Design.md, which was sitting untracked in
this working tree. 418 lines answering the operator's question about a
deterministic lead-plus-quorum decision procedure. It self-labels as a
proposal, not as built, and it is linked from the sidebar under a new
"Design proposals" heading rather than being numbered as a chapter.
One correction to that page before committing it. It reported a live defect:
that fleet.charters.architect told an architect to call bridge_send, a tool
CB-634 renamed away. That does not reproduce. Measured today in
fleetd/fleetd.yaml: bridge_send 0 times, any bridge_[a-z] name 0 times,
fleet_send twice, with "architects" twice as the control that the grep read
the right file. The page now carries that measurement, its re-measure
command, and the part of the argument that is still true: nothing checks
charter text against the registered tool surface, because
FleetConfig.validateCharters() reads only the key and the blankness. Filed
as fleetd #464.
The three mermaid diagrams were checked against the authoring rules by hand
(no hardcoded fills, quoted labels with parentheses, self-closing <br/>).
No mermaid renderer is installed on this host, so they were not rendered.
Features: fleet_ack does not destroy held peer mail, and heldDurable is derived
Two corrections to the "A lead can read its own held peer mail" entry.
1. It said fleet_ack "would have destroyed the message before anyone read
it". Measured false: fleet_ack calls MessageService.ackReply, which goes
to the worker reply inbox, never to the coordination channel. It removes
nothing and still reports "acknowledged <msgId>" (fleetd #437). The old
sentence made a no-op sound like a data-loss risk, which is the opposite
of the real defect.
2. heldDurable was a hardcoded true when this entry was written. Since
fleetd #440 it is derived from the queue durability and the ack mode, so
it can report false. Noted that one boolean over two inputs cannot say
which of the two changed.
Features: correct the stale owner of the unpinned startup-report call site
The 'usage-limit detection' entry said the six FleetConfig.validateXxx()
startup calls had the same unpinned shape and that fleetd #398 owned closing
it. Both halves were wrong.
- #398 is the closed models allow-list PR, not a ticket for this.
- The validateXxx half is FIXED: the six calls were collapsed into one
cfg.validateAll(), pinned by FleetdStartupValidationTest calling the real
Fleetd.main and asserting it refuses a bad config.
Also widened the gap statement. It is not only reportExhaustedPatternGap:
all four startup reports are referenced by exactly one test file each, and
each of those tests calls the method directly, so no test proves main still
calls any of them. fleetd #442 now owns it.
fleetd #421: a lead can read its own held peer mail
Features entry for fleet_poll{coordId} — primary-only, non-destructive,
full bodies. Records why READ was not reused (the roster-carries-no-secrets
grant does not cover a lead-to-lead body) and the mailbox.pending: 0 trap.
Also re-syncs the portable CLAUDE.md block template with the repo's copy:
the #421 work added an intent->tool row for reading held peer mail, and a
worker cannot commit this submodule.
Features: fleet_profiles now reports whether the model gate is armed
fleetd #434 (follow-up to #422). The off set alone could not tell 'no models:
block on this host' from 'a block that is armed and reporting zero off' — both
came back empty. modelGateArmed separates them and is reported even when false.
Documents the two fields, the three startup log lines, and why both the log and
the MCP field read the one PeerLauncher.modelGateState() accessor the spawn gate
itself reads.
Features: runtime model on/off (#422) and real architect-slot revocation (#424)
Two new entries, plus a correction to the #398 allow-list entry: its gotcha
said `models:` is deferred and needs a restart. Since #422 the block is hot,
so that line was telling operators the opposite of what the code does.
The #422 entry carries the spawn-path diagram, because the same trap has now
been hit twice: an explicit profile is REFUSED on a bad model while a blank one
is ROUTED AROUND it, so resolving a profile name before calling spawn silently
moves the caller from one path to the other. That is the open regression on
PR #430.
Also records that fleet01 has no `models:` block, so the gate is inert there,
and that automatic limit detection is not built.
Features: the credential scrub aborted, and a missing report means more than one thing
Two updates to the memberCredentials entry, both from fleetd #388 and #394.
#388: the scrub now runs from the generated `.zshenv` as well as `.zshrc` and
`.zlogin`. `.zshenv` is the only file zsh reads unconditionally, so the control
no longer depends on enumerating which kinds of shell exist.
#394: a missing `scrub-report.txt` had been read as evidence about which shell
ran. It is not. The report block is the last statement in the script, so its
absence proves only that the script did not reach the end. The real cause was
`export UID=`, a fatal zsh parameter error that terminates the whole sourced
file, with the message swallowed by the loop's `2>/dev/null`.
The severity was the selection, not the count: `env` lists inherited names
first and a startup file's own exports last, so the loop blanked the harmless
half and died immediately before the operator's own exports.
Also added, because both cost real time:
- do not infer shell kind from argv[0], and never from a probe run inside an
agent — opencode forks a fresh `zsh -c` per command, so the probe answers
about itself
- plant a decoy when auditing this: allow-listed names come back blank whether
or not the scrub ran, so their blankness proves nothing
- `.zcompdump` present plus `scrub-report.txt` absent is only explained by
"sourced, then aborted partway"
Root cause found by the fleet01 lead; the inverted-selection wording is theirs.
fleetd #365: a reply reports whether anything was waiting for it
The REST reply endpoint's response shape changed: 'delivered' is no longer
always true, and an 'outcome' field names which of three things happened.
15-REST-API-Reference.md still documented the old constant.
Adds the Features entry the charter requires for a visible behaviour change,
covering both doors and the nudge-counter rename (delivered -> sent).
Features: plugin renamed fleet@fleetd, mount name fleet, ${FLEETD_MCP_URL}
Also records why this entry alone did not prevent a rebuild: wiki/ is a
submodule whose pointer is never advanced, so no session reads it by default.
A capability an agent must not re-derive needs a line in CLAUDE.md too.
Refs fleetd #362