Features: the coordinator row is lead-only; add the CB-548 quorum design page
Features entry for fleetd #439, merged as 92a96fc. A worker or architect
calling fleet_list now gets no coordinator key at all, and the gate runs
before the row is assembled so the key is absent rather than empty. Notes
the compat overloads that still default to showing it (fleetd #463).
Also commits CB-548-Lead-Quorum-Design.md, which was sitting untracked in
this working tree. 418 lines answering the operator's question about a
deterministic lead-plus-quorum decision procedure. It self-labels as a
proposal, not as built, and it is linked from the sidebar under a new
"Design proposals" heading rather than being numbered as a chapter.
One correction to that page before committing it. It reported a live defect:
that fleet.charters.architect told an architect to call bridge_send, a tool
CB-634 renamed away. That does not reproduce. Measured today in
fleetd/fleetd.yaml: bridge_send 0 times, any bridge_[a-z] name 0 times,
fleet_send twice, with "architects" twice as the control that the grep read
the right file. The page now carries that measurement, its re-measure
command, and the part of the argument that is still true: nothing checks
charter text against the registered tool surface, because
FleetConfig.validateCharters() reads only the key and the blankness. Filed
as fleetd #464.
The three mermaid diagrams were checked against the authoring rules by hand
(no hardcoded fills, quoted labels with parentheses, self-closing <br/>).
No mermaid renderer is installed on this host, so they were not rendered.
Three pages the audit asked for that did not exist.
14 Fleet Manager — fleet-manager had zero coverage in this wiki, so an
operator had no way to learn it exists. Written from its own source: the
fleets.json shape from the parser rather than the example file, what each
command does, and how the probe decides working vs stalled by comparing
worktree modification times twice. It also records the two limits that come
from fleetd rather than from the tool: two daemons sharing one herdr session
tear down each other's members, and a session inside a pane is resolved as a
worker, which is why the manager sits outside and speaks REST.
15 REST API Reference — the 14 routes were documented nowhere as a set. The
page leads with why that matters: MCP and REST are sibling adapters over one
shared MessageService, so REST behaviour cannot be inferred from the MCP
contract. fleet_ack and fleet_whoami have no route at all.
It also records a trap found while writing it: GET /members returns its
rows under a "workers" key (FleetApp.java:322). The route was renamed from
/workers in CB-557 and the body key was left behind, so a caller reading
body["members"] sees an empty fleet instead of an error.
16 Security & Trust Boundary — the guard, the role table, the member
credential scrub and token scope were spread across three pages. Collected,
with the limits stated rather than glossed: the effective allow-list is a
union and so a strict superset of what an operator writes under allow:; the
ZDOTDIR scrub is zsh-only; blocking SSH_AUTH_SOCK does not stop a member
reaching a passphrase-free key file; and argv is world-readable, which
bypasses every environment control described on the page.
Home and _Sidebar link all three.
#168: revise the front matter, Approaches, and the OpenCode port
- Home and _Sidebar: renamed the product to fleet / fleetd. Dropped the claim
that AgentAPI is a "swappable fallback injector" — it was never built, and no
AgentAPI code exists under fleetd/src/main/java.
- Home also lost two claims that chapters 1 and 2 removed today: there is no SSE
route, and the page no longer names a gateway host as if it were fixed. The
allowed hosts are a config value and differ per deployment.
- Home no longer carries a release number, a ticket count or a test total. Those
go stale in days and then read as current facts; 8-Roadmap holds the record.
- 3-Approaches keeps AgentAPI as discarded research, which is that page's job,
but never in the present tense. Evidence: a search for agentapi under
fleetd/src/main/java finds nothing, and both Profile.kind values run through
HerdrPeerLauncher.
- 12-Claude-to-OpenCode: the sample mount name was `fleetd`, which teaches a
second product name. Every launcher writes the same constant,
PeerLauncher.MCP_MOUNT_NAME = "fleet". A hand-written config using another key
mounts under a name the role-detection ladder never checks.
Match the code cutover: daemon name, config (fleetd.yaml), scripts, launchd/
systemd units, module dir, and MCP tool prefix bridge_* -> fleet_*. Kept:
the BRIDGED_MEMBER security marker, mcp__bridge__ (historical mount name), and
the .bridged-worktrees on-disk path. The portable CLAUDE.md block stays
byte-identical with the repo's CLAUDE.md.
docs: add chapter 13, the operator user guide, for release 1.1
The wiki had twelve chapters and none of them told an operator how to run
the thing. Chapters 1-3 explain why the design is what it is, 9-12 explain
how the code is put together, 11 lists capabilities. The two pages that were
meant to cover bring-up and day-2 - 4-Setup and 5-Operations - were never
written past their scope note, and every technical detail in them had gone
wrong: a decommissioned model host, port 8080, herdr protocol 14, a systemd
unit that does not exist, Redis Streams and NATS that were never built,
"no per-session authz yet" after Authz shipped, and recycle() events that
have no code behind them.
So this adds 13-User-Guide.md, written against the running system on
2026-08-17, and points the two stubs at it rather than leaving wrong claims
in place.
The guide covers:
1. what this is and, more usefully, the five things it is NOT, each with
the reason it is not that;
2. install - herdr (check the PROTOCOL number, not the version), the
daemon, the login-shell rule for secrets.sh, and the lead's tab label;
3. configure - the four knobs that cost money, the live profile table with
who pays for each, the gateway paths, and memberCredentials' two halves;
4. run - the redeploy script, and the four checks that go beyond /healthz,
because health is green while every spawn fails;
5. delegate - the eleven tools, the spawn-all-then-send-all rule, the ~60s
client cap on a blocking send, and the authz table;
6. when it breaks - twelve traps hit for real this year, grouped by
bring-up, losing a member's work, and merging a member's work;
7. where to look next.
Home.md is corrected too: it claimed members launch against ollama.ltms.dev,
a host that no longer exists (the gateway is llm.ltms.dev), it framed the
system as Claude-only with one worker, it listed 8 of the 13 pages, and its
status still said "Design".
OpenCode reads CLAUDE.md natively (and ~/.claude/CLAUDE.md globally), so
the translate-then-hand-fix cycle that made the Codex port fragile does
not exist. The chapter is now one JSON mapping plus the rules for what
must not cross.
Keeps four verified Codex findings in an appendix — no CLAUDE.md
fallback, no config interpolation, a scrubbed MCP child environment, and
trust-gated project config — because they are what the comparison rests
on and were established by hand.
wiki: chapter 12 — porting a Claude Code workspace to Codex
Written as a procedure so it converts into a plugin skill without rewriting.
Every claim is from a hand-verified run against codex-cli 0.147.0 and
ai-config-sync-manager 0.1.10, not from documentation.
The two steps most likely to be skipped are the ones that cost us: filtering the
MCP sync so machine-local IDE servers never reach a project-root file that every
worktree inherits, and diffing the generated AGENTS.md every time — the
terminology map rewrites file paths into sentences that are wrong rather than
awkward, including turning "never commit .mcp.json" into a rule about a path in
~/ that cannot be committed at all.
Also records the three Codex seams a launcher cannot guess: --approve-for-me is
mandatory for unattended MCP calls, a fresh CODEX_HOME has no credentials and
fails as an opaque mid-run 401, and AGENTS.md is the only way to deliver a
standing instruction since Codex has no --append-system-prompt.
wiki: add chapter 11 Features, and stop claiming Stage 5 is finished
Roadmap said Stage 5 was "✅ All landed" and listed CB-501–505. Twenty tickets
shipped after that line was written (CB-506…CB-525) and none of them appeared
anywhere in the wiki — so the page was not merely incomplete, it asserted
something false.
Two fixes, because there were two problems:
- The stage row is corrected and a "Stage 5 continued (as-built)" section records
all twenty, split by who needs them: capabilities, contract changes, and the
quality infrastructure that makes the rest trustworthy. It also carries the
CB-521 version-coupling warning — /healthz can be green while every spawn
fails, because health pings herdr without checking that the adapter and the
binary agree on a protocol.
- Chapter 11 is new, and it is the structural fix. Chapters 1–10 are organised by
design topic and every one answers "how is this built / why this way". None
answers "what can it do and how do I turn it on", so a shipped knob like
`placement: weighted` had no page that wanted it and landed nowhere. Eleven
entries, each: what it does · the knob · why it exists · the gotcha. The `why`
line is the load-bearing one — it is what stops a decision being re-litigated
a month later.
Written from verified behaviour only; six entries that still need their config
surface confirmed against the code are listed as an explicit backfill list rather
than guessed at.
Mermaid block rendered with mermaid-cli to confirm it parses.
Fan-out code audit (one worker per layer, off-subscription) synthesized into
a verified as-built map: component layers, per-package class reference, REST
route table, bootstrap wiring, forward + reverse rendezvous sequence flows,
and two state machines (AgentStatus; Injector per-target turn lifecycle with
the real grace constants 8/120/240). Enums, routes, and the guard allowlist
verified against source at aa0cf81. All 6 mermaid diagrams mmdc-validated and
theme-safe.
wiki: consistency pass vs rewritten Architecture (5-agent review)
Independent cold reads of every page against 1-Architecture found no invariant
violations; fixed the drift the rewrite introduced plus one real contradiction:
- terminology: Channel 1/2 -> Mode 1/2, 'two-channel' -> 'two invariants / two
modes' (Approaches, Team, Home, Sidebar, Message-Server); north/south face ->
SERVER/CLIENT face (Message-Server, 6 spots)
- contradiction reconciled: Architecture now acknowledges a non-MCP *worker*
Stop-hook (POSTs reply to bridged) as well as the split-host-primary hook -
both target bridged, never a broker; Message-Server tier table split into
Unified / Hooked / Unmodified to match
- Approaches: footnote credits bridge_reply (Stop-hook = fallback); §4 subtitle
reframed; <payload> mermaid label de-angled (parse-safe)
- Team: SERVER 'role router' -> 'policy brain'; fan-out sequence quoted; inference edges labeled
- Home/README: CLIENT-face node regains 'status-gated injector'
- Operations: 'broker' -> 'internal broker/queue'; Stop-hook framed as split-host exception
- async ticket/bridge_poll reframed as injection-first (push), poll = non-pane fallback
All 16 mermaid blocks validated with mmdc.
wiki: number page filenames (1-..5-) so Gitea Pages list sorts
- git mv content pages to N-Name.md (history preserved); Home + _Sidebar kept
- convert [[wiki-links]] to [display](numbered-slug) markdown links so
resolution is unambiguous and prose display stays clean