4
13 User Guide
Dai Ha edited this page 2026-08-31 10:38:42 +07:00
This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

13 — User Guide

Status: 🟢 Written against the running system, 2026-08-17 (release 1.1). Every command here was run on this host. Where a claim is not measured, it says so.

This is the operator page. Chapters 1–3 say why the system is built this way, 9–12 say how the code is put together, 11 lists what it can do. This page is the one you read when you have to bring it up, run it, or fix it. It assumes you have done this before and forgotten the details.


1. What this is, and what it is not

fleetd is a message bus between AI agent sessions. One session orchestrates (the lead), and it delegates work to members running in other terminal panes, often on other models and other vendors. All traffic goes through fleetd's MCP tools. No session talks to another session, to a broker, or to the network directly.

flowchart LR
    LEAD["lead session<br/>(Claude Code, subscription)"]
    subgraph BD["fleetd — a plain Java daemon"]
        SRV["MCP + REST<br/>policy, authz, lifecycle"]
        INJ["injector<br/>status-gated"]
        SRV --> INJ
    end
    HERDR["herdr<br/>owns panes and PTYs"]
    M1["member pane<br/>claude-code"]
    M2["member pane<br/>opencode"]
    GW["llm.ltms.dev<br/>the one gateway"]

    LEAD -->|"fleet_send"| SRV
    M1 -.->|"fleet_reply"| SRV
    M2 -.->|"fleet_reply"| SRV
    INJ -->|"unix socket"| HERDR
    HERDR --> M1
    HERDR --> M2
    M1 --> GW
    M2 --> GW

    classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
    classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
    class LEAD,GW ext
    class SRV,INJ,HERDR core

Figure 1 — who talks to whom. The lead never touches herdr; the members never touch each other.

What it is not, and each "not" is a decision, not a gap:

Not Why
Not a build system It does not compile, lint, or test your code. Members run whatever the repo gives them. The lead verifies.
Not an environment manager It does not install toolchains or set up a member's language runtime. It only sets a small, named set of environment variables at spawn.
Not a Claude-only tool kind: claude-code and kind: opencode are both first-class. A member can be GPT, DeepSeek, or Claude.
Not a scheduler with a budget There is no cost metering and no refusal on spend. maxLoad per profile is the only throttle you have. See §3.
Not a security boundary against your own members A member runs as you, on your machine. memberCredentials limits which secrets it sees. It does not sandbox anything else.

The two invariants. Break either and nothing else matters.

  1. The lead never sets ANTHROPIC_BASE_URL or ANTHROPIC_AUTH_TOKEN. It stays on the subscription. Only the daemon moves a member off it, at spawn.
  2. The fleet is the only channel. Text printed in a pane reaches nobody. An answer that is not in a fleet_* call is discarded silently.

2. Install

Four things must be true before anything works. In order, because each one depends on the last.

2.1 herdr

fleetd does not own terminals. herdr does. fleetd drives it over a Unix socket.

herdr --version          # 0.8.0 on this host
curl -s http://127.0.0.1:8765/healthz
# {"status":"ok","herdr":{"version":"0.8.0","protocol":19}}

The protocol number is the thing to check, not the version. fleetd's adapter speaks one herdr wire protocol. If herdr is upgraded and the protocol moves, /healthz still says ok — because the socket connects — and every spawn fails. Green health with a broken fleet is the normal way this breaks. See trap 2 in §6.

2.2 The daemon

Build and run from this repo. There is no POM at the repo root:

mvn -f fleetd/pom.xml clean install      # never pipe this — a pipe hides BUILD FAILURE
java -jar fleetd/target/fleetd.jar fleetd/fleetd.yaml

In practice you never run those two by hand. Use the script (§4).

2.3 Secrets, and the login shell

All credentials live in one file: ${SHARED_ENV}/tools/secrets.sh. It is sourced only by a login shell. This one fact causes more lost hours than anything else in the system.

  • The daemon inherits AI_GATEWAY_TOKEN and WORKER_GITEA_TOKEN from the shell that started it.
  • Start it from a non-login shell and both are empty. The daemon boots, /healthz is green, and nothing complains.
  • The failure surfaces hours later: a gateway profile gets a 401, or a member cannot open a pull request.

launchd does not run a login shell either. That is the only reason scripts/fleetd-launchd-wrapper.sh exists — it execs zsh -lc so the store gets sourced, while keeping one process so launchd's PID tracking still works. Read its header; it explains the trap better than this paragraph.

On this host the daemon is not under launchd. It runs as a plain java -jar started by the redeploy script from a login shell. launchctl list | grep fleetd returns nothing.

The daemon logs which required secret names resolved at startup (Fleetd.reportRequiredSecrets). Read those lines. But note the gap: it skips profiles marked subscription: true, on purpose, because they need no token. So a green secret report says nothing about your subscription profiles.

2.4 The lead's tab

The daemon finds the lead by its herdr tab label, and by nothing else:

fleet:
  leaders:
    opus:
      tab: "lead: opus"        # must match the real tab label

Matching is case-insensitive and both ends are stripped. Internal spacing is not forgiving: "lead: opus" and "lead:opus" are different. A session whose tab does not match is resolved as a worker, and every orchestration call it makes is refused.

The old terminal: key is gone. A daemon at CB-579 or later refuses to start if terminal: is still in the config, and also if tab: is missing. That refusal is deliberate — a silently demoted lead was worse than a daemon that will not boot.

2.5 A second herdr daemon for members (optional)

§2.1 described one herdr. There is an optional top-level key, memberHerdrSocket:, that adds a second one (FleetConfig.java:84, fleetd/fleetd.example.yaml:117-118):

memberHerdrSocket: /run/fleet/herdr-members.sock    # optional — omit for one daemon

Set it to a second herdr API socket path and every member spawns on that daemon, while everything the lead does stays on the first one. Absent — the default — both names are the same client, so nothing changes. The daemon builds the two clients from the config on boot, and a HerdrRouter owns the split: each consumer gets the lead client, the member client, or routing by target id (Fleetd.java:152-158, herdr/HerdrRouter.java:6-24).

This key tells you what the feature will do, not how to turn it on today. The feature is not ready to switch on yet. The running blocker list is issue #185, and 11 Features → memberHerdrSocket explains why the split exists in the first place: herdr forks every pane as its own OS user, and has no user parameter in its socket API, so a member under a different user needs its own herdr. Two blockers are worth knowing up front:

  • The socket permissions are a race. herdr creates its API socket with mode 0600, and umask does not change that, so cross-user access still needs a chmod g+rw after the socket appears — a race on every start. That is a herdr-side measurement recorded in #185, not something the fleetd code can show.
  • Pane ids are still not daemon-qualified. herdr's ids are per-daemon counters, so two daemons can both hold w1:p1, pointing at different panes. fleet_stop{paneId} takes that bare id, and when two daemons really do claim one, the stop is refused rather than closing a pane on an arbitrary daemon (CompositePeerLauncher.java:49-60) — a safety net, not a design that removes the collision.

It is a boot-time-only key. The daemon reads memberHerdrSocket: once, at startup, and builds the member herdr client from it; nothing re-reads the key after that (Fleetd.java:153-155). A change therefore needs a daemon restart, not a config reload. In the reload machinery it sits in the set that cannot change under a running daemon, the same set as herdrSocket:; with configReload: enabled, a reload that touches it is refused on that account (ConfigRef.java:72-74, 190-195).

/healthz checks both daemons when the second is configured. It pings both, and returns 200 only when both answer; a down member daemon now shows 503 degraded instead of hiding behind a healthy lead, which is the old shape of trap 2 (FleetApp.java:229-255 and the comment at 210-216). The 200 body gains a member key with that daemon's version and protocol, and protocolMismatch: true when the two protocol numbers differ (FleetApp.java:256-264). The herdr key still carries the lead daemon's values, deliberately: scripts/redeploy-fleetd.sh and scripts/rename-checkout.sh already read this endpoint, and folding two daemons into one key would hide a mismatch from whichever reader looks only there (FleetApp.java:218-227). With one daemon the body is byte-identical to before — one ping, no new keys. The point of the change: it is the member daemon's protocol that decides whether a spawn works, and before, a bad protocol on the member daemon left /healthz green with every spawn failing.


3. Configure

fleetd/fleetd.yaml is the live config. It is gitignored. fleetd.example.yaml is the tracked, documented copy. Two consequences you will meet:

  • Members cannot see the live config. They work in worktrees of the tracked repo. So a change to what a config key means can pass review, merge green, and break the running fleet — because nobody who reviewed it could see the file it breaks. Check and fix the live config yourself at merge time.
  • The example file is the documentation. A key that is read by code but missing from the example is a real defect (that is how subscription: stayed undocumented for months — CB-610).

The knobs that cost money

This is the section to reread before you change anything.

Knob What it does The cost
subscription: true The member runs on your own Claude plan. No ANTHROPIC_BASE_URL, no token — it inherits your Claude Code auth. Every spawn bills your plan and eats your usage limit. Off-subscription is the whole point of the daemon, so treat true as a deliberate exception. Mutually exclusive with baseUrl; setting both is refused at load.
maxLoad Members allowed at once on this profile. The only throttle that exists. There is no metering, no budget, no refusal on spend. When maxLoad is reached, a spawn is refused with no capacity … no fallback — it does not fall back to a cheaper profile.
weight Share of automatic placement. weighted placement is not cheapest-first. It spreads by ratio across every profile that has a free slot, so paid spawns happen while the free local profile is idle. The workaround here is local.weight: 100; the real fix is open as CB-589.
fleet.reviewers Which profiles an unqualified reviewer spawn may land on. Review is the fan-out step — several members per pull request. Leaving the free profile out of this pool was the single largest avoidable cost in the fleet.

Live profiles on this host:

Profile Kind Model maxLoad Pays
local claude-code deepseek-v4-flash via llm.ltms.dev/anthropic 2 gateway
local-direct claude-code deepseek-v4-flash direct to gx00.gw:8000 2 nothing, LAN only, weight: 0
gx opencode gx/deepseek-v4-flash via llm.ltms.dev/v1 2 gateway
opus claude-code claude-opus-5 1 your Claude plan
sonnet claude-code claude-sonnet-5 3 your Claude plan
sol opencode openai/gpt-5.6-sol 1 OpenAI
terra opencode openai/gpt-5.6-terra 2 OpenAI

sol and terra share one OpenAI credential. When it hits its limit, both die at once, and an opencode member dying looks like a short clipped reply rather than an error.

The gateway

llm.ltms.dev is the one front door, and AI_GATEWAY_TOKEN is the single key.

  • kind: claude-code → https://llm.ltms.dev/anthropic. Never /v1 — on /v1 the model's thinking output silently disappears.
  • kind: opencode → https://llm.ltms.dev/v1. The path is taken as-is.

Both known gateway ceilings were fixed on 2026-08-15: the body limit went 32 KiB → 32 MiB, and the route timeout 60 s → 86400 s. The remaining risk is an Envoy HTTP/1.1 bug that can truncate a stream with a 200 and no terminator. Our members cannot detect that; it looks like a short answer.

Credentials a member can see

memberCredentials is deny-by-default. You name what is allowed; everything else known is replaced with a sentinel string. Current state: 34 known, 7 allowed, 29 blocked, measured inside a live member pane on 2026-08-17.

It has two halves, and only both together work:

  1. The launcher writes the member's environment at spawn.
  2. A BRIDGED_MEMBER-guarded block at the end of ${SHARED_ENV}/tools/secrets.sh re-applies it.

Half 2 is not optional. The member's pane runs a login shell, which re-sources the whole secret store and overwrites whatever the launcher set. The guarded block must be the last thing in the file, or the store overwrites it in turn.

The policy is re-read on every spawn, so an allow-list edit needs no restart. That is measured, not assumed.

Rule in force: only the lead and architects may use GITEA_ACCESS_TOKEN. Members get WORKER_GITEA_TOKEN, which is a minimal write:repository token, never the admin one.


4. Run, and prove it runs

Start and restart

Use the script. Do not hand-roll the steps.

scripts/redeploy-fleetd.sh --check   # read-only: reports state, changes nothing
scripts/redeploy-fleetd.sh           # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes     # skip the drain prompt (fleet already checked)
scripts/redeploy-fleetd.sh --no-build # restart the jar you already have

Run --check first, always. It is the only thing that reports whether the forge token resolves, and it reports that without printing the value.

The script builds before it stops anything, so a failed build never leaves the fleet down. It waits for the old process to exit instead of assuming. It polls /healthz. And it anchors its log checks to a marker taken before the restart, so old errors cannot be misread as new ones.

A merge is not a deployment. The running daemon holds the jar it started with. Code merged to main does nothing until you rebuild and restart. Saying "shipped" about code the live daemon has never loaded is a false report.

Drain first. fleet_list, collect anything you still want with fleet_poll, then fleet_stop each member. A restart drops in-flight tickets and rendezvous. A member's report is not recoverable once its ticket is gone.

Verify — /healthz is not enough

/healthz proves the socket connects. It does not prove a spawn works, that identity resolves, or that the new jar is the one running. Four checks, in order:

  1. A fresh boot line. Confirm a new fleetd listening line at the end of fleetd/fleetd.out, dated after the restart. An old daemon that never died looks identical from outside.
  2. Deferred keys. The startup log names which config keys it accepted and which it deferred. A deferred key needing a restart is usually the whole reason you restarted. Read those lines rather than assuming.
  3. Identity. fleet_whoami must still answer primary. If the tab label changed, the lead is now a worker and every orchestration call is refused.
  4. A real spawn. Spawn one cheap member and stop it. This is the only check that catches a herdr protocol mismatch.

Restarting the daemon cuts your own MCP mount, and it does not reconnect. So you cannot run fleet_whoami from the session that restarted it. Ask the operator to run /mcp to reconnect. This is why the restart is done from the lead but verified after a reconnect.

The REST surface

When the MCP mount is down — which is exactly when you need it most — this is the whole API. It is loopback only.

Route Does
GET /healthz 200 with herdr.protocol in the body; 503 degraded when herdr does not answer ping
GET /metrics Prometheus. Only present when metrics are enabled
GET /sessions every session, each row carrying its live agentStatus
GET /agents agents as herdr sees them
GET /members · POST /members · DELETE /members/{paneId} list, spawn, tear down
GET /profiles the backends configured
GET /sessions/{id}/status one session — the same view as fleet_status
POST /sessions/{id}/message · /reply · /ask the three message kinds
GET /sessions/{id}/replies drain the reply inbox — destructive, see trap 6
GET /tasks/{ticket} poll a detached ticket

Where it runs, and where the logs are

  • Log file: fleetd/fleetd.out, in both supervised and unsupervised modes.
  • Audit log: fleetd/logs/audit.log, rotated daily, 30 days kept.
  • Service units ship in deploy/: fleetd.service for Linux systemd (ordered After=herdr.service) and dev.ltms.fleetd.plist for macOS launchd. deploy/lavinmq holds the optional broker.
  • On this host neither is loaded. The daemon runs as a plain java -jar started by scripts/redeploy-fleetd.sh from a login shell. Verified with launchctl list | grep fleetd, which returns nothing. If you expected launchd here, that expectation is the bug.

5. Delegate

Eleven tools. This is the whole surface.

Intent Tool
Confirm your own role fleet_whoami
See backends available fleet_profiles
Start a member fleet_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}
See the fleet fleet_list → leads + members
One member's state fleet_status{sessionId}
Delegate, blocking fleet_send{sessionId, content}
Delegate, long task fleet_send{sessionId, content, wait:false} → ticket
Answer a member's fleet_ask fleet_send{turnId, content} — not sessionId
Collect a reply fleet_poll{ticket} or fleet_poll{target}
Clear a reply from the inbox fleet_ack{target, msgId} — target, not ticket
Member ends its turn fleet_reply{content}
Member asks the lead fleet_ask{question}
Tear down fleet_stop{paneId}

The loop

sequenceDiagram
    participant L as lead
    participant B as fleetd
    participant M as member
    L->>B: fleet_spawn — all units first
    L->>B: fleet_send with wait false — then all sends
    B-->>L: ticket
    B->>M: injected when the member is idle
    M->>B: fleet_reply
    B-->>L: nudge into the lead's own pane
    L->>B: fleet_poll by ticket
    L->>B: fleet_ack by target and msgId
    L->>B: fleet_stop by paneId

Figure 2 — the delegation loop. Spawn is separate from send on purpose.

Spawn every unit first, then send them all. Spawning and sending in one loop is how parallel work silently becomes serial. It is the most common way this whole layer is wasted.

Prefer wait:false. A blocking fleet_send is capped by your own MCP client timeout, about 60 seconds — far below any real task's runtime. The cap is in your client, not in the daemon, so no server setting fixes it.

Line 1 of every brief names a skill. Load the implementer skill. or Load the reviewer skill. Those skills are opt-in, and that line is what makes them reliable. The brief must be self-contained: the member sees your message and the repo, and nothing of your context, your plan, or your screen.

Since CB-588, a wait:false ticket that finishes nudges your pane by itself. You do not have to poll on a timer. It needs an injectable pane and is capped at 5 pending nudges, so still watch anything you cannot afford to lose.

Who may call what

Call Allowed
spawn / stop / drain lead only
send lead or architect
reply / ask any member, only as itself
read / status any authenticated caller

Identity comes from the connection, never from an argument. You cannot act as another session. A call outside your role is refused, not queued.

The lead's own job

Delegation does not delegate responsibility. These stay with you and are not delegable: the conversation with the operator, decomposition, the final judgment, verification, and the merge. Merging on a reviewer's word is delegating the merge by proxy.

A member cannot run your IDE tooling. Any forge tools it appears to have hold a deliberately blocked credential and fail every call. And a piped command hides failures behind a zero exit code. So never promote a member's "clean" to a fact — re-run the build yourself.


6. When it breaks

Twelve traps, all hit for real this year. Grouped by where they bite.

Bring-up

1. The daemon started from a non-login shell. Symptom: everything green for hours, then a member cannot open a pull request, or a gateway profile gets a 401. Nothing logs it at the time. Fix: scripts/redeploy-fleetd.sh --check is the only thing that reports it. Restart from a login shell, or via the launchd wrapper.

2. /healthz is green and every spawn fails. Cause: herdr was upgraded and its wire protocol moved past what the adapter speaks. The socket still connects, so health is ok. Fix: check protocol in the /healthz body, and verify any restart with one real spawn — not with health.

3. The lead is resolved as a worker. Cause: the herdr tab label no longer matches fleet.leaders.*.tab. Every orchestration call is refused. Fix: match the tab exactly — internal spacing counts. Note the old terminal: key is removed; a current daemon refuses to start if it is still present.

4. A merge is not a deployment. The running daemon holds its original jar. Rebuild and restart, then prove it with a fresh fleetd listening line. Also: the restart cuts your own MCP mount and it never reconnects, so you cannot verify identity from that session — ask for /mcp.

Losing a member's work

5. The ticket expired. A wait:false ticket is kept for about 10 minutes, counted from when the turn finishes — not from when you sent it (MessageService.pruneTerminalTickets). So a task that runs for an hour still hands you its report, as long as you collect within about 10 minutes of it finishing.

This used to be measured from creation, which meant any task longer than the TTL lost its report the moment it arrived. That was fixed in #197; a daemon started before that fix still has the old behaviour, so check what the running jar is before you trust a long ticket.

When a ticket has expired, fleet_poll{ticket} reports it as timed out. The member is usually fine and its answer went to the member's inbox instead. Drain it with fleet_poll{target}, then fleet_ack{target, msgId}. For anything long, still have the member write its report into a file or its pull request, so a lost ticket is never a lost report.

6. Reading the reply inbox is destructive. GET /sessions/{id}/replies drains on first read. If you curl it through head or a parser that dies, the reply is gone. Always redirect it to a file first.

7. The idle reaper ate the report. A member that is done holds its report only until lifecycle.idleTtlSeconds (1800 s here). Then the ticket, the pane, and the report are all gone, with no scrape fallback. Poll before the reaper.

8. A second send to a busy member never lands. timed_out_queued means never delivered — the member is fine. Worse, the stale brief stays queued and restarts the member when it next goes idle. Back up the member's commit first: git fetch <worktree> <branch>:refs/backup/x.

9. liveStatus: working is not progress. A member can report busy for hours doing nothing. Check file modification times in its worktree and snapshot git diff before and after. Do not read its terminal — use fleet_status.

10. Never brief a member to "ask me". fleet_ask blocks for about 55 seconds and no nudge extends it. It is also invisible to fleet_poll. If you are running async, you will not see the question in time. Decide before you delegate, or give the member an explicit default.

Merging a member's work

11. The reported branch is not the branch it committed to. fleet_list reports the branch spawn provisioned, not the one the member actually used. Check git -C <worktree> branch --show-current, or you push an empty ref and the merge says "Already up to date".

12. Members see a months-old wiki, and cannot see the live config. wiki/ is a submodule whose pointer is never advanced, so a member's checkout is a stale snapshot. Never brief "read wiki/…" — paste the text, and commit the member's wiki entry yourself. Same shape for fleetd.yaml: it is gitignored, so a member cannot see the file its change may break.

The general shape behind several of these

A checker that is narrower than what everyone believes it checks, with nothing saying so. The credential probe hardcoded 31 names against a policy of 29 and exited 0. A config test checked only column-0 keys, so a nested key that bills your Claude plan stayed undocumented. A scheduled sweep never ran once because its gate used a sentinel in a numeric field.

Two habits fix most of it: make a checker print its own denominator ("checked 26 of 29", not "26 blocked"), and when a green check is the evidence for a claim, ask what it does not cover before believing it.


7. Where to look next

You want Go to
Why the design is herdr-centric at all 1 Architecture, 3 Approaches
The daemon's own design 2 Message Server
The as-built code map, classes, state machines 9 Implementation
What a capability does, the knob, why it exists, the gotcha 11 Features
Orchestrating a mixed-vendor fleet 6 Team, 7 Use Cases
Adding a second host or a non-Claude peer 10 Cross-Host Messaging, 12 Claude → OpenCode
What is left to build 8 Roadmap
The rules every session must follow CLAUDE.md in the repo — it loads into every session
Rendezvous, fleet_ask, detached delivery, turn-done fallback docs/MCP-Contract.md §6 only — the rest of that page is a pre-build design doc and its tool names never caught up with the code
Every config key, documented fleetd/fleetd.example.yaml

The one thing that is easy to forget. This repo is the fleet, so the charter block in CLAUDE.md is not documentation about someone else's system — it is the instruction surface this codebase ships. A code change that quietly makes it untrue is an incomplete change. The block in CLAUDE.md and the template in 7 Use Cases must stay byte-identical; there is a check script in CLAUDE.md that proves it.