#362: make the plugin visible, and fix the drift that made it unusable
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Successful in 1m31s

CB-527 shipped a Claude Code plugin and a marketplace in this repo. Nothing in
CLAUDE.md or docs/ ever named it, so a later session planned the same feature
from scratch. The wiki Features entry existed and was correct, but wiki/ is a
submodule whose pointer is never advanced, so no session reads it.

Visibility:
- CLAUDE.md addendum now names plugin/ and both structural limits, so every
  session sees it. This is the change that stops the rebuild happening again.
- wiki/11-Features.md records the rename and why the entry alone was not enough.

Drift (each measured against the code, not assumed):
- mount name fleetd -> fleet, matching PeerLauncher.MCP_MOUNT_NAME. The old name
  gave a lead with both a project .mcp.json and the plugin two mounts of one
  daemon and a duplicated fleet_* tool set.
- url is now ${FLEETD_MCP_URL} instead of a hardcoded address, so one plugin can
  serve hosts running the daemon on different ports. Plain ${VAR}, the form
  kb-alms proves works here; ${VAR:-default} is untested and not used.
- plugin claude-bridge -> fleet, marketplace claude-bridge -> fleetd, version
  0.2.0. Breaking for a 0.1.0 install: mcp__fleetd__* becomes mcp__fleet__*.
- README install path ltms/claude-bridge -> the fleet/fleetd remote.
- the setup skill's §5 told operators to pin primary.terminal:. CB-579 replaced
  that with fleet.leaders.*.tab. Replaced, with the duplicate-tab warning (#359).

Scope: the plugin is lead-side only, and cannot be otherwise. The launcher adds
--agent only when <worktree>/.claude/agents/<role>.md exists in the member's own
tree (ClaudeCodeLauncher.java:371,391), and a member's CLAUDE_CONFIG_DIR points
at its profile's config dir (ClaudeCodeLauncher.java:285), so a member never
reads the operator's plugin store. On this Mac all four Claude profiles set
configDir, and the four ccs instances hold four separate copies of the plugin
store -- same md5, different inodes. Seeding member skills through the worktree
is #362 scope item 3, implemented separately.

Note for anyone verifying a plugin: `claude plugin validate` does NOT read
.mcp.json. Replacing it with `{ this is not json at all` still passes, exit 0.

Refs #362, #359
This commit is contained in:
Dai Ha
2026-09-05 12:42:20 +07:00
parent 3759c41f99
commit 457458437f
7 changed files with 433 additions and 34 deletions
+337
View File
@@ -0,0 +1,337 @@
# Fleet as a Claude Code plugin — plan
Status: draft for architect review. Not implemented.
Author: primary (lead `opus`, Mac fleet). Date: 2026-09-05.
## 0. Correction — this already exists, and that changes the plan
I wrote sections below as if the plugin were new work. It is not. **This repo is already a Claude
Code marketplace and already ships a plugin**, added in `ef1e014` (CB-527) and last touched in
`2e138a1` (CB-634):
```
.claude-plugin/marketplace.json -> name "claude-bridge", plugins: [ ./plugin ]
plugin/.claude-plugin/plugin.json -> name "claude-bridge", version 0.1.0
plugin/.mcp.json -> mounts "fleetd" at http://127.0.0.1:8765/mcp
plugin/skills/setup/SKILL.md -> a full onboarding skill
plugin/README.md
```
The `setup` skill is good and covers most of what section 5 proposes: preflight, merge-not-clobber
into `.mcp.json`, read-only permissions only, credentials by env-var name, and a verify step that
insists on a **real spawn** because a green `/healthz` proves nothing.
So the operator's question — "can we pack things into plugins?" — is already answered *yes, and it
was built*. The real question is why it did nothing for the kb session. The answer is drift plus
invisibility.
### The drift, measured
| # | Finding | Evidence |
|---|---|---|
| 1 | **Mount name mismatch.** The plugin mounts the server as `fleetd`; the daemon's own constant is `fleet` | `plugin/.mcp.json` vs `PeerLauncher.java:34` `String MCP_MOUNT_NAME = "fleet"` |
| 2 | **URL hardcoded**, no env indirection, so one plugin cannot serve two hosts or ports | `plugin/.mcp.json` |
| 3 | **Ships no worker skills and no agents** | `plugin/` has 1 skill (`setup`); `.claude/skills/` has 5 and `.claude/agents/` has 3, none of them in `plugin/` |
| 4 | **Stale identity advice.** `setup` §5 tells the operator to pin `primary.terminal:` | CB-579 replaced that with `fleet.leaders.*.tab`. `record Primary` still exists (`FleetConfig.java:954`), so the advice is not dead — but it is no longer the mechanism |
| 5 | **Stale install path.** README says `/plugin marketplace add ltms/claude-bridge` | the repo is `fleet/fleetd` since CB-623 |
| 6 | **Stale names.** Plugin and marketplace are both `claude-bridge` | the project renamed to `fleetd` in CB-634 |
| 7 | **Nothing references it.** `grep -rn "plugin/" CLAUDE.md docs/*.md` returns nothing | so no session is ever told the plugin exists — which is exactly why I planned it from scratch |
Finding 7 is the root cause of the other six. A shipped capability that no instruction file
mentions gets no maintenance, and the next person rebuilds it. That is the same failure the
`CLAUDE.md` "Features" rule was written to stop.
Finding 3 is the one that matters most for the operator's actual problem. A worker spawned into a
**kb** worktree has no `implementer` skill, because only `claude-bridge` carries one in
`.claude/skills/`. Every brief that says "Load the implementer skill" is a no-op outside this repo.
The plugin is the right home for those skills and does not carry them yet.
## 1. The problem, restated
A project is "fleet-enabled" today by hand-edits nobody wrote down in one place:
- a `fleet.leaders.<name>` entry in a host's gitignored `fleetd.yaml`;
- the repo must carry `.claude/skills/*` for a worker to load `implementer` or `reviewer`;
- the repo must carry the canonical bridge block in its `CLAUDE.md`;
- the MCP mount arrives only because `LeadLauncher` and `ClaudeCodeLauncher` add `--mcp-config`
to the argv they build.
Shown live on 2026-09-05: an operator opened `claude` by hand in `/home/ltms/LTMS/kb` on fleet01
and the session had **no `fleet_*` tools at all**, because a hand-started agent never gets the
launcher's `--mcp-config`. The plugin would have fixed that — if it had been installed, and if it
had been mentioned anywhere.
## 2. What was verified, and how
| Claim | Evidence |
|---|---|
| A plugin can install globally | `~/.claude/plugins/installed_plugins.json` — scopes in use are `project` (5), `local` (4), **`user` (1)** |
| A plugin can mount an MCP server | `~/.claude/plugins/marketplaces/kb-alms/.mcp.json` mounts `memory` at `"url": "${KB_MEMORY_URL}"` |
| A plugin can carry skills, agents, commands, hooks | `kb-alms` ships `skills/` + `hooks/hooks.json`; `umputun-cc-thingz/plugins/planning` ships `agents/` + `skills/` |
| Env vars interpolate in a plugin's `.mcp.json` | same `kb-alms` file: `${KB_MEMORY_URL}`, `${MEMORY_MCP_TOKEN}` |
| A marketplace can be a plain git repo | `known_marketplaces.json` — `mgnl-code-review` has `"source": "git", "url": "https://..."` |
| **This repo is already such a marketplace** | `.claude-plugin/marketplace.json`, committed in `ef1e014` |
| A lead already binds to a project directory | `FleetConfig.java:1017` `record Leader(..., String workspace, String cwd)`; used at `LeadLauncher.java:193-199` |
| The lead's mount comes from argv, not config | `LeadLauncher.java:253` adds `--mcp-config` |
## 2b. Architect review + one measurement changed the design
The architect verified the plan against the code and returned **build it with these changes**. Two
of its findings are load-bearing. I checked both myself.
### A plugin cannot carry the agent definitions — confirmed
`ClaudeCodeLauncher.java:371` calls
`agentDefinitionFile(spec.cwd(), spec.role(), ".claude", "agents")`, and
`HerdrPeerLauncher.java:359-365` returns a path only when
`<cwd>/.claude/agents/<role>.md` `Files.isRegularFile`. `ClaudeCodeLauncher.java:391-393` adds
`--agent` **only** when that returns non-null. `OpenCodeLauncher` does the same for
`.opencode/agent`.
So the file must exist **in the member's worktree**. Moving `.claude/agents/*.md` into the plugin
would silently stop every member getting `--agent`. **The agents stay in the repo.** My plan had
this wrong.
### The plugin does not reach members at all — confirmed, and worse than the architect could see
`ClaudeCodeLauncher.java:285` does `putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir())`.
A member with `configDir` set reads that directory, not the operator's `~/.claude`.
The architect could not check how far that goes, because `fleetd.yaml` is gitignored. I measured it:
- **All four Claude profiles on the live fleet set `configDir`** — `local`, `local-direct`, `opus`,
`sonnet` (`grep -c configDir fleetd/fleetd.yaml` = 4).
- Each `ccs` instance has its **own** `plugins/` directory: `gx10` (8 entries), `ltms` (11),
`ollama` (9), `work` (9).
- Those directories are **four separate real directories with four separate inodes**, and
`installed_plugins.json` in each is a **separate inode with an identical md5**
(`51c6e1c853e32e656b817e123fbbfcc5`). They are *copies made once*, not links.
So a plugin installed at user scope lands in exactly one instance's store. It would have to be
installed once per `CLAUDE_CONFIG_DIR`, and each copy would then drift. **The plugin is not a
delivery mechanism for member-facing assets on this host.**
A side effect worth recording: my own session's `CLAUDE_CONFIG_DIR` is set, so the
`~/.claude/plugins/*` evidence in section 2 is not even this session's store. The claims about what
a plugin *can* do still hold — they were read from real manifests — but the directory I read them
from is the wrong one for any conclusion about *this* session.
### The design that follows
Split by audience, not by mechanism:
| Audience | Delivered by | Carries |
|---|---|---|
| operator / lead (a human opening any project) | **the plugin**, per config dir | the MCP mount, `setup`, the bridge charter |
| member (a worker in a provisioned worktree) | **worktree provisioning** | `.claude/skills/*`, `.claude/agents/*` |
The second row is not a new idea — it is what the code already does for agents, and it is why
`agentDefinitionFile` looks in the worktree. Extending worktree provisioning to seed
`.claude/skills/` from a fleetd-owned source is the consistent move, and it is what actually fixes
"a worker in kb has no `implementer` skill". The plugin never could.
### Other review findings I accepted
- **Mount-name collision is real.** `PeerLauncher.java:34` is `fleet`; the plugin mounts `fleetd`.
A lead with both gets two mounts of one daemon and duplicate `fleet_*` tools. Rename the
plugin's server to `fleet`.
- **Keep `--mcp-config` in `LeadLauncher`.** It is config→argv from `profile.mcpUrl()`, not drift,
and it is the only path that works for a member with its own `configDir`.
- **I overstated #359.** `LeadCoordLoop.java:174-197` returns null and logs a warning that names the
fix; `tick()` leaves the message unacked, so the broker holds it and delivers once a lead is
named. It **stalls loudly and recovers** — it is not silent, and it is not data loss. My wording
in issue #361 needs the same correction.
- **Stage 1 ends with tools that mostly cannot be used until stage 3**, because authority still
comes from the tab. Section 3 already said this; the stage table did not.
## 3. What a plugin can and cannot do
This is the part that decides the design, so it is stated before the design.
**A plugin gives tools. It does not give authority.**
`fleet_whoami` resolves a caller's role from the connection, not from what is mounted. A
hand-started session in kb that mounts `fleet_*` through a plugin will be resolved as a **worker**
and refused on every orchestration call, because its pane is not in a tab matching
`fleet.leaders.*.tab`.
So the plugin alone does not make a project fleet-enabled. It makes it *tool*-enabled. Registering
the lead stays fleetd's job. Any plan that forgets this ships a plugin that looks installed and
does nothing.
```mermaid
flowchart TB
P["fleet plugin<br/>(user scope, every session)"] --> T["fleet_* tools mounted"]
D["fleetd.yaml<br/>leaders.kb {tab, cwd}"] --> A["role = primary"]
T --> W["can call fleet_*"]
A --> W2["calls are authorized"]
W --> OK["working lead"]
W2 --> OK
T --> NO["tools mounted, every call refused"]
classDef good fill:#2f855a,stroke:#22543d,color:#ffffff;
classDef bad fill:#9b2c2c,stroke:#63171b,color:#ffffff;
class OK good
class NO bad
```
*Both halves are needed. The plugin is the left half only.*
## 4. Proposed architecture — three layers
### Layer 1: the plugin — lead-side only, fix the one that exists
Keep it at `plugin/`, keep the marketplace at `.claude-plugin/marketplace.json`. Do not create a
second one, and do not put member-facing assets in it (see 2b).
```
.claude-plugin/marketplace.json -> rename to "fleetd"; keep source ./plugin
plugin/
.claude-plugin/plugin.json -> rename to "fleet"; bump version
.mcp.json -> mount name "fleet" (match PeerLauncher.MCP_MOUNT_NAME),
url "${FLEETD_MCP_URL}", 8765 default documented
skills/setup/SKILL.md -> EXISTS. fix the stale primary.terminal advice (§5)
skills/bridge-charter/SKILL.md -> NEW: the canonical CLAUDE.md block
README.md -> fix the install path (fleet/fleetd, not ltms/claude-bridge)
```
**Not in the plugin:** `agents/*.md` (the launcher requires them in the member's worktree —
`ClaudeCodeLauncher.java:371,391`) and the three worker skills (a member with `configDir` set never
reads the operator's plugin store — measured in 2b). Those belong to layer 1b.
**A rename is a breaking change for anyone who installed 0.1.0.** The mount name goes `fleetd` ->
`fleet`, so a project whose `.claude/settings.json` pre-allows `mcp__fleetd__fleet_whoami` stops
matching. Only this fleet has it installed today, so the cost is small now and grows. Decide once.
**This solves propagation of the charter.** `CLAUDE.md` says the bridge block "must stay
byte-identical with the template in the wiki" and that "other projects carrying the block need the
same edit" — a hand-copy the file itself admits is fragile, with a python snippet to check it. A
plugin skill turns that into a version bump.
### Layer 1b: worker skills reach members through the worktree, not the plugin
This is the change that actually fixes "a worker in kb cannot load `implementer`".
Worktree provisioning already writes into the member's tree — the parity overlay, the neutralised
`.mcp.json`, the IDE overlay. Add one more: seed `<worktree>/.claude/skills/` from a fleetd-owned
source directory, so every member gets `implementer`, `reviewer` and `hunter` whatever repo it is
working in. `.claude/agents/` is already required there by the launcher, so this follows the
grain of the design rather than cutting across it.
Open question for implementation: copy or symlink, and where the source lives (a config key such
as `memberSkills:`, or the plugin's own directory read by the daemon). A symlink is one source of
truth but breaks if the member's tree is archived; a copy drifts but is self-contained.
### Layer 2: host-global fleet settings
`~/.fleet/fleetd.yaml` — the things that are true for the **machine**, not the project:
- `broker:` and `coordinator:` (URIs come from env, no secrets in the file)
- `profiles:` — backends, models, credentials, weights
- `memberCredentials:` policy
- `worktreeRoot`, `worktreeGroup`
### Layer 3: per-project settings, committable
`<project>/.fleet/project.yaml` — the things that are true for the **repo**:
```yaml
lead:
tab: "lead: kb"
profile: opus
ide:
projectDir: "" # kb is a Python repo, everything at the root
worktree: true
```
**This split fixes a contradiction that exists today.** `ideProjectDir` is a property of a *repo*
(fleetd's Maven module is a subdirectory; kb's code is at the root) but the config key is
per-*profile*. One profile therefore cannot serve both repos — measured on fleet01 on 2026-09-05,
where the key had to be commented out to make kb work. Moving it to a project file removes the
contradiction rather than working around it.
It is also committable, because it holds no secrets. A project that has been fleet-enabled once
stays fleet-enabled for everyone who clones it.
## 5. The `fleet-setup` skill
What the operator actually asked for: one command that makes any project fleet-compatible.
```mermaid
sequenceDiagram
participant Op as Operator
participant Sk as fleet-setup skill
participant Fs as project files
participant Fd as fleetd
Op->>Sk: /fleet-setup (in any project)
Sk->>Fs: write .fleet/project.yaml
Sk->>Fs: add the bridge block to CLAUDE.md (if absent)
Sk->>Fd: register the lead (tab + cwd)
Fd-->>Sk: tab created, lead launched
Sk-->>Op: report what changed, and what is still manual
```
*The skill writes the project half and asks the daemon for the host half.*
The registration step needs something that does not exist yet: an MCP tool such as
`fleet_workspace_add{path, tab, profile}`, or a `fleetd` config include so a project file is picked
up without hand-editing the host file. **This is the one genuinely new piece of daemon work.**
## 6. Does this reduce fleetd's complexity?
Honestly: **partly**. Claiming more than this would be wrong.
**Yes, in three places.**
1. Skill and agent delivery leaves the daemon and the repos entirely.
2. Worktree config neutralisation (fleetd #134) gets safer. It blanks `.mcp.json` so the primary's
IDE and forge servers do not leak into a worker. Today the fleet mount survives only because
the launcher re-adds it by argv. With a user-scope plugin the fleet mount is outside the file
being neutralised, so the two concerns stop fighting.
3. The `ideProjectDir` per-profile/per-repo contradiction disappears.
**No, in the places that matter most.** fleetd still owns spawn, authorization, worktrees, herdr,
the broker, tickets, and identity. A plugin cannot do any of those. The plugin is a **distribution**
mechanism, not a replacement for the daemon.
**And it adds one new risk.** The mount URL becomes a second source of truth. `fleetd.yaml` has the
port; the plugin has the URL. Mitigation: the plugin reads `${FLEETD_MCP_URL}` only, and the host
env is the single place it is set.
## 7. Rollout stages
| Stage | Content | Ends with |
|---|---|---|
| 0 | **Make it visible.** One `CLAUDE.md` line and one Features entry saying the plugin exists and where | nobody re-plans it a third time |
| 1 | Fix the plugin's drift: mount name `fleet`, `${FLEETD_MCP_URL}`, names, README, stale `primary.terminal` advice | a lead in any project can install one plugin and get the mount |
| 1b | Seed `.claude/skills/` into provisioned worktrees | **a worker in *kb* can load `implementer`** |
| 2 | `.fleet/project.yaml` schema + `FleetConfig` reads it; `ideProjectDir` moves there | kb and fleetd both work off one profile |
| 3 | Fix #359, then config-include for lead registration, wired into the existing `setup` skill | `/fleet:setup` in a fresh project produces a working lead |
| 4 | Roll out to fleet01; retire the hand-copied CLAUDE.md block in favour of the skill | one `git pull` propagates the charter |
Stage 0 is minutes of work and is the one that stops this happening again, so it goes first.
**Stage 1b carries most of the value** and is independent of the plugin — it could ship first if the
plugin rename needs more thought. Stage 1 alone ends with tools a hand-started session mostly
cannot use, because authority still comes from the tab; that is fixed in stage 3, not stage 1.
#359 moves ahead of stage 3 on the architect's advice, because stage 3 is what creates the second
lead.
## 8. Questions for the architect
1. **Is the layer-2 / layer-3 split right?** Specifically: should `profiles:` stay host-global, or
should a project be able to pin which profiles it uses? Cost of getting this wrong is a config
that has to be re-split later.
2. **Config include, or a new MCP tool, for registering a project's lead?** An include is passive
and survives a restart; a tool is live but writes to a gitignored file the daemon owns.
3. **What happens when the plugin is absent?** Should `LeadLauncher` keep its `--mcp-config`
belt-and-braces, or is that the drift risk we should remove? Note opencode members cannot use
Claude plugins at all, so `OpenCodeLauncher` keeps its ephemeral config either way.
4. **Does a user-scope plugin mount leak into members in a way we do not want?** Members already
inherit user-scope MCP servers (`--mcp-config` adds, it does not replace). A worker getting
`fleet_*` is correct and already happens. Confirm nothing else in the plugin should be
worker-invisible.
5. **Two leads on one host both hold a subscription seat.** Is per-project leads the right unit, or
should one lead serve several projects by changing cwd?
6. **Blocking defect to fix first or alongside:** `LeadCoordLoop.resolveLocalLead()` (lines
174-190) routes a peer message to "the sole lead" when no lead is *named* after
`coordinator.selfId`. The moment a host has two leads — exactly what this plan encourages —
cross-host coordination silently stops. Tracked as #359.