Files
fleetd/plugin/skills/setup/SKILL.md
Dai Ha 2e138a199b CB-634: one shared "fleet" workspace + rename bridged -> fleetd cutover
Two changes ship together here.

1. One shared herdr workspace. The lead and every worker now live in one
   workspace called "fleet", so the operator sees one "session" with many
   windows, not two. Before, the lead sat in a "leads" workspace and workers
   in "bridged-workers", which read as two sessions. The lead is still told
   apart from workers by its exact tab label ("lead: <name>"), so putting them
   in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
   for split layouts; Fleetd now passes an empty exclude set.

2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
   launchd/systemd units, module dir, and MCP mount).
   - Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
   - Log line, comments, docs, and CLAUDE.md updated to say fleetd.
   - Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
     bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
   - Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
     bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
   - Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
     still read as a fallback, and still gitignored.
   - MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
     server name is "fleet". The mount name in the local .mcp.json becomes
     "fleet" (gitignored, not in this commit).
   - Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
     BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.

Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.

Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.

The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).

949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
2026-08-25 04:01:08 +02:00

228 lines
9.7 KiB
Markdown

---
name: setup
description: Make the current project bridge-ready — check the prerequisites, mount the fleetd MCP gateway into the project's .mcp.json, apply standard Claude Code settings, and verify this session resolves as the primary. Writes no credentials. Load this when asked to set up, install, configure, or onboard a project onto claude-bridge, or when fleet_* tools are expected but absent.
---
# Bridge setup — make this project bridge-ready
This skill configures **the project you are currently in** so that this Claude Code session can
orchestrate a fleet of delegated workers through `fleetd`.
**It writes no credentials, ever.** Every secret is referenced by environment-variable *name*, and
the user exports the value themselves. Nothing this skill creates is unsafe to commit. If you are
ever about to write a token, key, or password into a file, you have misread this skill — stop.
Work through the steps in order. Each one has a check; **report what actually happened**, including
failures. A setup that half-worked and was reported as done is worse than one that failed loudly.
## 0. Establish where you are
```bash
pwd
git rev-parse --show-toplevel 2>/dev/null || echo "(not a git repo)"
ls -a | head -30
```
Everything below is written into **this** project root. If the user meant a different directory,
confirm before writing anything.
## 1. Preflight — what must already exist
The bridge is three moving parts, and the plugin is only one of them. Check all of it before
changing any file, so you can tell the user the whole story at once instead of failing one step at
a time.
```bash
command -v herdr && herdr --version 2>&1 | head -1 || echo "MISSING: herdr"
command -v ccs && ccs version 2>&1 | head -1 || echo "MISSING: ccs (needed for worker profiles)"
command -v codex && codex --version 2>&1 | head -1 || echo "absent: codex (optional)"
curl -s -m 5 http://127.0.0.1:8765/healthz || echo "MISSING: fleetd daemon is not reachable"
```
A healthy daemon answers with its status **and the herdr protocol it negotiated**:
```json
{"status":"ok","herdr":{"version":"0.8.0","protocol":19}}
```
| Missing | What to tell the user |
|---|---|
| `herdr` | The PTY multiplexer that owns worker terminals. Install it first; nothing else works without it. |
| `fleetd` | The daemon. It is a separate service, not part of this plugin — the plugin only *mounts* it. Point the user at the project's own install instructions. |
| `ccs` | Only needed to launch worker profiles. The bridge itself will still start. |
| `codex` | Optional. Only needed if this fleet will run Codex peers. |
**Do not attempt to install these yourself.** They are system services with their own lifecycles;
guessing at an install is how you end up with two daemons on one socket. Report what is missing and
let the user install it.
## 2. Mount the bridge MCP — merge, never overwrite
The project's `.mcp.json` may already declare servers. **Read it first and merge**; clobbering
someone's existing MCP config is not a recoverable mistake.
```bash
cat .mcp.json 2>/dev/null || echo "(no .mcp.json yet)"
```
The entry to add, exactly:
```json
{
"mcpServers": {
"fleetd": {
"type": "http",
"url": "http://127.0.0.1:8765/mcp"
}
}
}
```
If `.mcp.json` already exists, add only the `fleetd` key and leave every other server untouched.
If a `fleetd` entry is already there with a different URL, **ask** rather than assuming yours is
right — a non-default port usually means a deliberate second daemon.
> **If this plugin is installed, you can skip this step entirely.** The plugin ships its own
> `.mcp.json`, so `fleetd` is already mounted for any session with the plugin enabled. Write the
> project-level file only when the user wants the mount to work *without* the plugin — for
> teammates who have not installed it, or for CI.
**Before writing it, settle whether `.mcp.json` is committed here:**
```bash
git ls-files --error-unmatch .mcp.json 2>/dev/null && echo "TRACKED" || echo "untracked"
```
A tracked `.mcp.json` is inherited by every checkout of this repo — including git worktrees the
bridge provisions for workers. Servers bound to *your* machine (an IDE index, a local language
server) will then be mounted by workers too, and every path they return points into **your**
checkout rather than the worker's. That failure is silent and expensive: it has produced a worker
that made all of its edits in the wrong tree while its builds passed, because it was building the
tree it was not editing. Keep machine-local servers out of a tracked `.mcp.json`, or keep the file
untracked.
## 3. Standard project settings
Create or merge `.claude/settings.json`. These are defaults, not requirements — keep anything the
project already set.
```json
{
"$schema": "https://json.schemastore.org/claude-code-settings.json",
"permissions": {
"allow": [
"mcp__fleetd__fleet_whoami",
"mcp__fleetd__fleet_list",
"mcp__fleetd__fleet_status",
"mcp__fleetd__fleet_profiles",
"mcp__fleetd__fleet_poll"
]
}
}
```
Only the **read-only** bridge verbs are pre-allowed. `fleet_spawn`, `fleet_send`, and
`fleet_stop` start processes, deliver work, and tear down terminals — those stay behind a prompt
on purpose. Do not "helpfully" add them.
Never write `settings.local.json` on the user's behalf; that file is personal and usually
gitignored.
## 4. Credentials — by reference only
The bridge takes every secret from the **environment**, and the daemon's config names the variable
rather than holding the value. Your job is to tell the user which variables to export, not to
collect or store them.
| Variable | Needed for | Notes |
|---|---|---|
| `FLEETD_WORKER_TOKEN` | authenticating a worker to the daemon | only when the daemon is configured with `tokenEnv` |
| `GITEA_TOKEN` / equivalent | letting a worker open its own PR | **minimal `write:repository` scope** — see below |
| `GITEA_HOST` | the forge base URL | no secret; safe anywhere |
Two rules to state plainly to the user:
- **The PR token must not be able to merge.** A worker opens a PR; the primary is the gate. A token
that can merge makes the gate decorative. Mint a narrow, repo-scoped token — never reuse a
personal admin token.
- **Never set `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN`** in this project, this shell, or any
settings file. The primary stays on subscription; only the daemon moves a *worker* off it, at
spawn. Mounting the bridge must never move a session across that boundary — if setup appears to
need this, something is wrong and you should stop and say so.
Write none of these into any file. Show the user the `export` lines to run themselves.
## 5. Verify — and do not trust a green health check
Reconnect MCP if needed (`/mcp`), then confirm the tools are live and this session is the primary:
```
fleet_whoami
```
- `{"role":"primary"}` — correct, you are done with this step.
- `{"role":"worker", …}` — **this is the trap.** If the primary runs inside a herdr pane, the
daemon resolves it to a terminal and classifies it as a worker, refusing `spawn`/`send`/`stop`:
every verb an orchestrator exists to call. It is **self-locking**, because the daemon can only
*learn* the primary's terminal from those same refused calls. The only way out is an
operator-set pin in the daemon's config:
```yaml
primary:
terminal: term_xxxxxxxxxxxx # the terminalId fleet_whoami just reported
```
The daemon reads this **at boot**, so it needs a restart. Re-pin whenever the primary moves
panes — a stale pin fails exactly as silently as no pin.
Then prove the fleet actually works, with a real spawn:
```
fleet_profiles → the configured backends
fleet_spawn{profile: "<one of them>"} → must reach state "ready"
fleet_stop{paneId: "<from spawn>"}
```
**`/healthz` reporting `ok` is not evidence that spawning works.** It reports that the daemon can
reach herdr — nothing more. A version mismatch between the daemon's adapter and the herdr binary
leaves health green while every single spawn fails. Only a real spawn proves the fleet. Do this
even when everything above looked fine.
## 6. Optional — Codex parity
Only if the user wants Codex and Claude Code to share instructions, skills, and MCP config:
```bash
npm install -g ai-config-sync-manager
ai-config-sync connect
ai-config-sync status # compare both hosts
ai-config-sync sync --dry-run # preview — always look before applying
ai-config-sync sync --apply
```
It maps `~/.claude/CLAUDE.md` ↔ `~/.codex/AGENTS.md`, `~/.claude/skills/` ↔ `~/.codex/skills/`, and
Claude's MCP servers ↔ `[mcp_servers.*]` in `~/.codex/config.toml`.
**Raise the boundary before running it.** That sync is *user-level* and bidirectional, while the
bridge deliberately isolates each worker's tool surface (§2). Syncing your MCP servers into
`~/.codex/config.toml` gives every Codex session your machine-local servers — the same
wrong-tree failure as §2, in a different runtime. Use the sync for the two CLIs *you* drive
interactively; leave anything the bridge spawns isolated. Always `--dry-run` first.
## 7. Report
State plainly:
```
prereqs: herdr <version> · fleetd <protocol> · ccs <version> · codex <version|absent>
written: <files created or merged, or "none">
role: <fleet_whoami result — and the pin, if one was needed>
spawn: <real spawn result: profile, state reached, torn down>
env: <variables the USER still needs to export — names only, never values>
skipped: <anything not done, and why>
```
Never report a step as done that you did not verify. If the daemon was unreachable, say so and stop
— the remaining steps cannot be checked, and guessing at them is how a broken setup gets called
finished.