ecc590f344
Part of #145 (CB-632). Documentation only, plus one internal literal. Unit 1 renamed the package and classes, which left every doc describing classes that no longer exist. This fixes the prose across README.md, docs/ and bridged/docs/ -- 18 files. Renamed: dev.ltms.bridged -> dev.ltms.fleet, the five class names, and "bridged" where it names the daemon as a product rather than a path. Also renamed two literals, because a doc that disagrees with the code is worse than one that is out of date: - bridged-local-noauth -> fleetd-local-noauth. A placeholder apiKey OpenCodeLauncher sends when a profile resolves no token, to a local endpoint that does not check it. No test asserts the old string. - the vnd.ltms.bridged.* media type in the M4 design doc. It appears in no Java file, so nothing implements it yet. Deliberately NOT renamed, because each is still literally true today and changes only at the cutover: - paths: bridged/, bridged.yaml, bridged.example.yaml, bridged.jar, .bridged-worktrees, deploy/dev.ltms.bridged.plist, scripts/redeploy-bridged.sh, bridged-launchd-wrapper.sh - bridged_* metric names -- renaming these after the monitoring is wired would break dashboard continuity, so they move before it is - bridge_* MCP tool names, which answer alongside fleet_* on purpose - BRIDGED_* environment variables, read by a file outside this repo Method note: perl, not sed. BSD sed has no \b and no lookaround, and a word-boundary expression there fails silently. The prose replace uses (?<![\w./-])bridged(?![\w./-]) so it cannot touch a path or an identifier, then every remaining hit was read by hand. Verified: mvn clean install green, 51 classes, 878 tests, 0 failures.
120 lines
6.9 KiB
Markdown
120 lines
6.9 KiB
Markdown
# v1.0.0 — One leader, one host, complete
|
|
|
|
This is the first release of **`fleetd`**.
|
|
|
|
`fleetd` lets one main Claude Code session (the **leader**, on your Pro/Max subscription)
|
|
run a team of **workers** — extra Claude Code sessions on a cheaper or local model, and
|
|
non-Claude agents too. The leader's own session is never touched: it stays on subscription,
|
|
with a clean environment.
|
|
|
|
**This release finishes a full scope — it does not stop halfway.** The scope is *one leader
|
|
on one machine*, running many workers. Everything that setup needs is now built, tested, and
|
|
used daily: who-is-who, the subscription line, messages in both directions, worker start and
|
|
stop, security, and monitoring. Nothing on the single-machine path is left as a known gap.
|
|
|
|
This is also how the code is shaped: `PrimaryRegistry` holds exactly **one** leader. Running
|
|
across many machines is the next big step (see *What comes next* below) — not a missing piece
|
|
of this one.
|
|
|
|
## One gateway for all messages
|
|
|
|
- **Everyone talks through the same door.** The leader and every worker connect to the same
|
|
MCP server and use only its tools: `fleet_whoami` · `fleet_profiles` · `fleet_spawn` ·
|
|
`fleet_list` · `fleet_status` · `fleet_send` · `fleet_reply` · `fleet_ask` ·
|
|
`fleet_poll` · `fleet_ack` · `fleet_stop`.
|
|
- **You are who your connection says you are.** The bridge finds out who is calling from the
|
|
connection itself, never from a name the caller sends. So a worker cannot pretend to be
|
|
someone else, and `fleet_whoami` tells each agent its own role — no guessing.
|
|
- **The subscription line cannot be crossed.** Only a spawned worker gets
|
|
`ANTHROPIC_BASE_URL`; the leader never does. Each worker profile has a list of allowed
|
|
model hosts, checked before anything starts.
|
|
- **Messages wait for the right moment.** The bridge sends one message per turn, only when
|
|
the other side is ready — no hammering a busy agent.
|
|
|
|
## Worker lifecycle
|
|
|
|
- **Start → work → stop.** Each worker gets its own git worktree (its own copy of the repo)
|
|
with the same setup as the leader — `CLAUDE.md`, skills, hooks — so it commits on its own
|
|
branch and opens its own PR. Ready-made playbooks ship in the repo:
|
|
`.claude/skills/implementer` and `.claude/skills/reviewer`.
|
|
- **Fast failure, not a silent hang.** Starting a worker waits until it is really connected.
|
|
If it never connects, you get a clear error (`PeerUnreachableException`) instead of a stuck
|
|
send.
|
|
- **Workers don't live forever.** Idle workers are cleaned up (`idle_ttl`), long sessions have
|
|
a turn limit (`context_cap`), shutdown drains work first, and workers left behind by an old
|
|
daemon are found and removed at startup.
|
|
- **Predictable placement.** Each worker gets its own tab in a shared worker space, in the
|
|
same order every time.
|
|
|
|
## No reply gets lost
|
|
|
|
MCP only lets the client call the server, so the bridge could push to a worker but the leader
|
|
had to ask for its replies. That gap is now closed on a single machine:
|
|
|
|
- **Replies are kept, never dropped.** If a reply arrives and nobody is waiting, the bridge
|
|
holds it until the leader picks it up.
|
|
- **Replies can survive a restart.** With a broker (LavinMQ or RabbitMQ) set up, held replies
|
|
live on the broker, so a daemon restart does not lose them — they come back, and repeats are
|
|
filtered out by `msgId`. No `broker:` in the config → replies are held in memory instead.
|
|
- **The leader gets a tap on the shoulder.** When a reply lands, the bridge nudges the
|
|
leader's own pane — only when the leader is free, and only a few times. If the leader is on
|
|
another machine, this quietly falls back to pick-up mode; the reply still waits.
|
|
- **Workers can ask questions.** With `fleet_ask`, a worker can pause mid-task, ask the
|
|
leader something, and continue the *same* task with the answer.
|
|
|
|
## More than one kind of worker
|
|
|
|
Workers are started through a small plug-in interface (`PeerLauncher`). Two plug-ins ship:
|
|
one for Claude Code and one for **opencode** (tested live against opencode 1.18.5). The
|
|
opencode one proves the interface is neutral — it uses nothing Claude-specific. Each worker
|
|
only sees the tools its own launcher gives it.
|
|
|
|
## Security & operations
|
|
|
|
- **Auth.** Default is `loopback-trust`: only same-machine callers are trusted. Or set a
|
|
bearer `token`. Unknown callers count as `ANONYMOUS` — nobody is trusted by accident. If
|
|
the config would expose the daemon to the network without a token, it **refuses to start**.
|
|
Workers never need the token, so turning auth on cannot lock them out.
|
|
- **Rules + audit log.** The role rules are checked on both doors (REST and MCP). A worker
|
|
may only act as itself. The audit log is JSON and never contains message text.
|
|
- **Monitoring.** `/healthz` for liveness, `/metrics` for Prometheus — no extra libraries.
|
|
- **Runs as a service.** launchd (macOS) and systemd (Linux) files are included. If herdr
|
|
isn't up yet at boot, the daemon waits up to 30 seconds and then runs in a reduced mode
|
|
instead of crash-looping.
|
|
- **CI.** Every push builds and tests on the Gitea runner, including the broker test against
|
|
a real broker. Only the live-herdr test stays local (`-Pcontract`).
|
|
|
|
## What you need
|
|
|
|
Java 25 · Maven · **herdr 0.8.0 (protocol 19)** · optionally LavinMQ or RabbitMQ for
|
|
restart-proof replies · macOS (launchd) or Linux (systemd). For TLS, put a reverse proxy in
|
|
front — the daemon does not do TLS itself, by design.
|
|
|
|
## Tested
|
|
|
|
`mvn clean install` is green at `84081b2`: **399 tests**, coverage **75.6%** of instructions /
|
|
**64.5%** of branches. Live end-to-end runs under `e2e/`: a worker asking the leader a
|
|
question, one leader running several workers at once on a bug hunt, and a 30-turn
|
|
back-and-forth conversation. The bridge is used on itself — worker-run code reviews have led
|
|
to real committed fixes in this repo.
|
|
|
|
## What comes next
|
|
|
|
The single-machine story is done. The next big step stretches the same rules across machines:
|
|
|
|
- **Many machines (CB-308).** A leader on machine A with workers on machines B and C — built
|
|
on this release's broker layer, with one gateway per machine. The design is written
|
|
(`docs/CB-308-Multi-Host-Federation.md`); the open question is trust between machines.
|
|
- **Many leaders.** Today the bridge holds exactly one leader. The next idea is a small
|
|
council — for example one Claude and one Codex — that can discuss a problem together,
|
|
compare answers, and agree on a decision before the work is handed to workers. This needs
|
|
leader-to-leader messages and a simple way to settle disagreement, neither of which exists
|
|
yet.
|
|
- **Third-party launcher plug-ins.** Loading launcher plug-ins from outside the project needs
|
|
a trust model first, because a launcher runs with daemon rights and can hand secrets to
|
|
workers.
|
|
|
|
One thing we chose **not** to build, so it isn't read as a gap: the originally planned heavy
|
|
message envelope (CB-201). Connection identity already routes every reply to the right place,
|
|
so only a small `QUESTION` message kind and a `turn_id` were added.
|