Files
Dai Ha 457458437f
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Successful in 1m31s
#362: make the plugin visible, and fix the drift that made it unusable
CB-527 shipped a Claude Code plugin and a marketplace in this repo. Nothing in
CLAUDE.md or docs/ ever named it, so a later session planned the same feature
from scratch. The wiki Features entry existed and was correct, but wiki/ is a
submodule whose pointer is never advanced, so no session reads it.

Visibility:
- CLAUDE.md addendum now names plugin/ and both structural limits, so every
  session sees it. This is the change that stops the rebuild happening again.
- wiki/11-Features.md records the rename and why the entry alone was not enough.

Drift (each measured against the code, not assumed):
- mount name fleetd -> fleet, matching PeerLauncher.MCP_MOUNT_NAME. The old name
  gave a lead with both a project .mcp.json and the plugin two mounts of one
  daemon and a duplicated fleet_* tool set.
- url is now ${FLEETD_MCP_URL} instead of a hardcoded address, so one plugin can
  serve hosts running the daemon on different ports. Plain ${VAR}, the form
  kb-alms proves works here; ${VAR:-default} is untested and not used.
- plugin claude-bridge -> fleet, marketplace claude-bridge -> fleetd, version
  0.2.0. Breaking for a 0.1.0 install: mcp__fleetd__* becomes mcp__fleet__*.
- README install path ltms/claude-bridge -> the fleet/fleetd remote.
- the setup skill's §5 told operators to pin primary.terminal:. CB-579 replaced
  that with fleet.leaders.*.tab. Replaced, with the duplicate-tab warning (#359).

Scope: the plugin is lead-side only, and cannot be otherwise. The launcher adds
--agent only when <worktree>/.claude/agents/<role>.md exists in the member's own
tree (ClaudeCodeLauncher.java:371,391), and a member's CLAUDE_CONFIG_DIR points
at its profile's config dir (ClaudeCodeLauncher.java:285), so a member never
reads the operator's plugin store. On this Mac all four Claude profiles set
configDir, and the four ccs instances hold four separate copies of the plugin
store -- same md5, different inodes. Seeding member skills through the worktree
is #362 scope item 3, implemented separately.

Note for anyone verifying a plugin: `claude plugin validate` does NOT read
.mcp.json. Replacing it with `{ this is not json at all` still passes, exit 0.

Refs #362, #359
2026-09-05 12:42:20 +07:00

253 lines
11 KiB
Markdown

---
name: setup
description: Make the current project bridge-ready — check the prerequisites, mount the fleetd MCP gateway into the project's .mcp.json, apply standard Claude Code settings, and verify this session resolves as the primary. Writes no credentials. Load this when asked to set up, install, configure, or onboard a project onto claude-bridge, or when fleet_* tools are expected but absent.
---
# Bridge setup — make this project bridge-ready
This skill configures **the project you are currently in** so that this Claude Code session can
orchestrate a fleet of delegated workers through `fleetd`.
**It writes no credentials, ever.** Every secret is referenced by environment-variable *name*, and
the user exports the value themselves. Nothing this skill creates is unsafe to commit. If you are
ever about to write a token, key, or password into a file, you have misread this skill — stop.
Work through the steps in order. Each one has a check; **report what actually happened**, including
failures. A setup that half-worked and was reported as done is worse than one that failed loudly.
## 0. Establish where you are
```bash
pwd
git rev-parse --show-toplevel 2>/dev/null || echo "(not a git repo)"
ls -a | head -30
```
Everything below is written into **this** project root. If the user meant a different directory,
confirm before writing anything.
## 1. Preflight — what must already exist
The bridge is three moving parts, and the plugin is only one of them. Check all of it before
changing any file, so you can tell the user the whole story at once instead of failing one step at
a time.
```bash
command -v herdr && herdr --version 2>&1 | head -1 || echo "MISSING: herdr"
command -v ccs && ccs version 2>&1 | head -1 || echo "MISSING: ccs (needed for worker profiles)"
command -v codex && codex --version 2>&1 | head -1 || echo "absent: codex (optional)"
curl -s -m 5 "${FLEETD_MCP_URL%/mcp}/healthz" 2>/dev/null \
|| curl -s -m 5 http://127.0.0.1:8765/healthz \
|| echo "MISSING: fleetd daemon is not reachable"
[ -n "$FLEETD_MCP_URL" ] && echo "FLEETD_MCP_URL is set" || echo "MISSING: FLEETD_MCP_URL"
```
**`FLEETD_MCP_URL` is required.** The plugin's own `.mcp.json` mounts `${FLEETD_MCP_URL}` rather
than a hardcoded address, so one plugin can serve hosts that run the daemon on different ports. If
it is unset the mount does not resolve. The usual value is `http://127.0.0.1:8765/mcp`; tell the
user to export it, do not write it into a file for them.
A healthy daemon answers with its status **and the herdr protocol it negotiated**:
```json
{"status":"ok","herdr":{"version":"0.8.0","protocol":19}}
```
| Missing | What to tell the user |
|---|---|
| `herdr` | The PTY multiplexer that owns worker terminals. Install it first; nothing else works without it. |
| `fleetd` | The daemon. It is a separate service, not part of this plugin — the plugin only *mounts* it. Point the user at the project's own install instructions. |
| `ccs` | Only needed to launch worker profiles. The bridge itself will still start. |
| `codex` | Optional. Only needed if this fleet will run Codex peers. |
**Do not attempt to install these yourself.** They are system services with their own lifecycles;
guessing at an install is how you end up with two daemons on one socket. Report what is missing and
let the user install it.
## 2. Mount the bridge MCP — merge, never overwrite
The project's `.mcp.json` may already declare servers. **Read it first and merge**; clobbering
someone's existing MCP config is not a recoverable mistake.
```bash
cat .mcp.json 2>/dev/null || echo "(no .mcp.json yet)"
```
The entry to add, exactly:
```json
{
"mcpServers": {
"fleet": {
"type": "http",
"url": "http://127.0.0.1:8765/mcp"
}
}
}
```
**The server must be named `fleet`.** That is `PeerLauncher.MCP_MOUNT_NAME` in the daemon, the name
a spawned member's mount carries, and the name the `mcp__fleet__*` role heuristic in `CLAUDE.md`
keys on. An earlier version of this plugin named it `fleetd`, which gave a lead with both a project
file and the plugin **two mounts of the same daemon** and a duplicated `fleet_*` tool set.
If `.mcp.json` already exists, add only the `fleet` key and leave every other server untouched.
If a `fleet` entry is already there with a different URL, **ask** rather than assuming yours is
right — a non-default port usually means a deliberate second daemon.
> **If this plugin is installed, you can skip this step entirely.** The plugin ships its own
> `.mcp.json`, so `fleetd` is already mounted for any session with the plugin enabled. Write the
> project-level file only when the user wants the mount to work *without* the plugin — for
> teammates who have not installed it, or for CI.
**Before writing it, settle whether `.mcp.json` is committed here:**
```bash
git ls-files --error-unmatch .mcp.json 2>/dev/null && echo "TRACKED" || echo "untracked"
```
A tracked `.mcp.json` is inherited by every checkout of this repo — including git worktrees the
bridge provisions for workers. Servers bound to *your* machine (an IDE index, a local language
server) will then be mounted by workers too, and every path they return points into **your**
checkout rather than the worker's. That failure is silent and expensive: it has produced a worker
that made all of its edits in the wrong tree while its builds passed, because it was building the
tree it was not editing. Keep machine-local servers out of a tracked `.mcp.json`, or keep the file
untracked.
## 3. Standard project settings
Create or merge `.claude/settings.json`. These are defaults, not requirements — keep anything the
project already set.
```json
{
"$schema": "https://json.schemastore.org/claude-code-settings.json",
"permissions": {
"allow": [
"mcp__fleet__fleet_whoami",
"mcp__fleet__fleet_list",
"mcp__fleet__fleet_status",
"mcp__fleet__fleet_profiles",
"mcp__fleet__fleet_poll"
]
}
}
```
Only the **read-only** bridge verbs are pre-allowed. `fleet_spawn`, `fleet_send`, and
`fleet_stop` start processes, deliver work, and tear down terminals — those stay behind a prompt
on purpose. Do not "helpfully" add them.
Never write `settings.local.json` on the user's behalf; that file is personal and usually
gitignored.
## 4. Credentials — by reference only
The bridge takes every secret from the **environment**, and the daemon's config names the variable
rather than holding the value. Your job is to tell the user which variables to export, not to
collect or store them.
| Variable | Needed for | Notes |
|---|---|---|
| `FLEETD_WORKER_TOKEN` | authenticating a worker to the daemon | only when the daemon is configured with `tokenEnv` |
| `GITEA_TOKEN` / equivalent | letting a worker open its own PR | **minimal `write:repository` scope** — see below |
| `GITEA_HOST` | the forge base URL | no secret; safe anywhere |
Two rules to state plainly to the user:
- **The PR token must not be able to merge.** A worker opens a PR; the primary is the gate. A token
that can merge makes the gate decorative. Mint a narrow, repo-scoped token — never reuse a
personal admin token.
- **Never set `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN`** in this project, this shell, or any
settings file. The primary stays on subscription; only the daemon moves a *worker* off it, at
spawn. Mounting the bridge must never move a session across that boundary — if setup appears to
need this, something is wrong and you should stop and say so.
Write none of these into any file. Show the user the `export` lines to run themselves.
## 5. Verify — and do not trust a green health check
Reconnect MCP if needed (`/mcp`), then confirm the tools are live and this session is the primary:
```
fleet_whoami
```
- `{"role":"primary"}` — correct, you are done with this step.
- `{"role":"worker", …}` — **this is the trap.** If the lead runs inside a herdr pane whose tab the
daemon does not recognise, it is classified as a worker and refused on `spawn`/`send`/`stop`:
every verb an orchestrator exists to call. It is **self-locking**, because those are the same
calls that would tell the daemon who you are.
Identity is the **tab label**, matched exactly and case-insensitively:
```yaml
fleet:
leaders:
kb: # name it after coordinator.selfId if this host uses lead-to-lead
profile: opus
tab: "lead: kb" # the exact label of the tab this lead sits in
cwd: /path/to/the/project
```
A tab label is stable across restarts of the agent inside it, which is why CB-579 replaced the
older `primary.terminal:` pin — a herdr `terminal_id` changed on every restart and cost a config
edit each time. `primary.terminal:` still parses, but it is no longer the mechanism; do not
reach for it.
The daemon reads `leaders:` **at boot**, so a new entry needs a restart. Two things to check
afterwards: that `fleet_whoami` now answers `primary`, and that no *stale* tab carries the same
label — duplicate lead tabs are their own failure (#359), and they stall lead-to-lead delivery
until one lead is named after `coordinator.selfId`.
Then prove the fleet actually works, with a real spawn:
```
fleet_profiles → the configured backends
fleet_spawn{profile: "<one of them>"} → must reach state "ready"
fleet_stop{paneId: "<from spawn>"}
```
**`/healthz` reporting `ok` is not evidence that spawning works.** It reports that the daemon can
reach herdr — nothing more. A version mismatch between the daemon's adapter and the herdr binary
leaves health green while every single spawn fails. Only a real spawn proves the fleet. Do this
even when everything above looked fine.
## 6. Optional — Codex parity
Only if the user wants Codex and Claude Code to share instructions, skills, and MCP config:
```bash
npm install -g ai-config-sync-manager
ai-config-sync connect
ai-config-sync status # compare both hosts
ai-config-sync sync --dry-run # preview — always look before applying
ai-config-sync sync --apply
```
It maps `~/.claude/CLAUDE.md` ↔ `~/.codex/AGENTS.md`, `~/.claude/skills/` ↔ `~/.codex/skills/`, and
Claude's MCP servers ↔ `[mcp_servers.*]` in `~/.codex/config.toml`.
**Raise the boundary before running it.** That sync is *user-level* and bidirectional, while the
bridge deliberately isolates each worker's tool surface (§2). Syncing your MCP servers into
`~/.codex/config.toml` gives every Codex session your machine-local servers — the same
wrong-tree failure as §2, in a different runtime. Use the sync for the two CLIs *you* drive
interactively; leave anything the bridge spawns isolated. Always `--dry-run` first.
## 7. Report
State plainly:
```
prereqs: herdr <version> · fleetd <protocol> · ccs <version> · codex <version|absent>
written: <files created or merged, or "none">
role: <fleet_whoami result — and the pin, if one was needed>
spawn: <real spawn result: profile, state reached, torn down>
env: <variables the USER still needs to export — names only, never values>
skipped: <anything not done, and why>
```
Never report a step as done that you did not verify. If the daemon was unreachable, say so and stop
— the remaining steps cannot be checked, and guessing at them is how a broken setup gets called
finished.