Files
Dai Ha 457458437f
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Successful in 1m31s
#362: make the plugin visible, and fix the drift that made it unusable
CB-527 shipped a Claude Code plugin and a marketplace in this repo. Nothing in
CLAUDE.md or docs/ ever named it, so a later session planned the same feature
from scratch. The wiki Features entry existed and was correct, but wiki/ is a
submodule whose pointer is never advanced, so no session reads it.

Visibility:
- CLAUDE.md addendum now names plugin/ and both structural limits, so every
  session sees it. This is the change that stops the rebuild happening again.
- wiki/11-Features.md records the rename and why the entry alone was not enough.

Drift (each measured against the code, not assumed):
- mount name fleetd -> fleet, matching PeerLauncher.MCP_MOUNT_NAME. The old name
  gave a lead with both a project .mcp.json and the plugin two mounts of one
  daemon and a duplicated fleet_* tool set.
- url is now ${FLEETD_MCP_URL} instead of a hardcoded address, so one plugin can
  serve hosts running the daemon on different ports. Plain ${VAR}, the form
  kb-alms proves works here; ${VAR:-default} is untested and not used.
- plugin claude-bridge -> fleet, marketplace claude-bridge -> fleetd, version
  0.2.0. Breaking for a 0.1.0 install: mcp__fleetd__* becomes mcp__fleet__*.
- README install path ltms/claude-bridge -> the fleet/fleetd remote.
- the setup skill's §5 told operators to pin primary.terminal:. CB-579 replaced
  that with fleet.leaders.*.tab. Replaced, with the duplicate-tab warning (#359).

Scope: the plugin is lead-side only, and cannot be otherwise. The launcher adds
--agent only when <worktree>/.claude/agents/<role>.md exists in the member's own
tree (ClaudeCodeLauncher.java:371,391), and a member's CLAUDE_CONFIG_DIR points
at its profile's config dir (ClaudeCodeLauncher.java:285), so a member never
reads the operator's plugin store. On this Mac all four Claude profiles set
configDir, and the four ccs instances hold four separate copies of the plugin
store -- same md5, different inodes. Seeding member skills through the worktree
is #362 scope item 3, implemented separately.

Note for anyone verifying a plugin: `claude plugin validate` does NOT read
.mcp.json. Replacing it with `{ this is not json at all` still passes, exit 0.

Refs #362, #359
2026-09-05 12:42:20 +07:00

19 KiB

Fleet as a Claude Code plugin — plan

Status: draft for architect review. Not implemented. Author: primary (lead opus, Mac fleet). Date: 2026-09-05.

0. Correction — this already exists, and that changes the plan

I wrote sections below as if the plugin were new work. It is not. This repo is already a Claude Code marketplace and already ships a plugin, added in ef1e014 (CB-527) and last touched in 2e138a1 (CB-634):

.claude-plugin/marketplace.json          -> name "claude-bridge", plugins: [ ./plugin ]
plugin/.claude-plugin/plugin.json        -> name "claude-bridge", version 0.1.0
plugin/.mcp.json                         -> mounts "fleetd" at http://127.0.0.1:8765/mcp
plugin/skills/setup/SKILL.md             -> a full onboarding skill
plugin/README.md

The setup skill is good and covers most of what section 5 proposes: preflight, merge-not-clobber into .mcp.json, read-only permissions only, credentials by env-var name, and a verify step that insists on a real spawn because a green /healthz proves nothing.

So the operator's question — "can we pack things into plugins?" — is already answered yes, and it was built. The real question is why it did nothing for the kb session. The answer is drift plus invisibility.

The drift, measured

# Finding Evidence
1 Mount name mismatch. The plugin mounts the server as fleetd; the daemon's own constant is fleet plugin/.mcp.json vs PeerLauncher.java:34 String MCP_MOUNT_NAME = "fleet"
2 URL hardcoded, no env indirection, so one plugin cannot serve two hosts or ports plugin/.mcp.json
3 Ships no worker skills and no agents plugin/ has 1 skill (setup); .claude/skills/ has 5 and .claude/agents/ has 3, none of them in plugin/
4 Stale identity advice. setup §5 tells the operator to pin primary.terminal: CB-579 replaced that with fleet.leaders.*.tab. record Primary still exists (FleetConfig.java:954), so the advice is not dead — but it is no longer the mechanism
5 Stale install path. README says /plugin marketplace add ltms/claude-bridge the repo is fleet/fleetd since CB-623
6 Stale names. Plugin and marketplace are both claude-bridge the project renamed to fleetd in CB-634
7 Nothing references it. grep -rn "plugin/" CLAUDE.md docs/*.md returns nothing so no session is ever told the plugin exists — which is exactly why I planned it from scratch

Finding 7 is the root cause of the other six. A shipped capability that no instruction file mentions gets no maintenance, and the next person rebuilds it. That is the same failure the CLAUDE.md "Features" rule was written to stop.

Finding 3 is the one that matters most for the operator's actual problem. A worker spawned into a kb worktree has no implementer skill, because only claude-bridge carries one in .claude/skills/. Every brief that says "Load the implementer skill" is a no-op outside this repo. The plugin is the right home for those skills and does not carry them yet.

1. The problem, restated

A project is "fleet-enabled" today by hand-edits nobody wrote down in one place:

  • a fleet.leaders.<name> entry in a host's gitignored fleetd.yaml;
  • the repo must carry .claude/skills/* for a worker to load implementer or reviewer;
  • the repo must carry the canonical bridge block in its CLAUDE.md;
  • the MCP mount arrives only because LeadLauncher and ClaudeCodeLauncher add --mcp-config to the argv they build.

Shown live on 2026-09-05: an operator opened claude by hand in /home/ltms/LTMS/kb on fleet01 and the session had no fleet_* tools at all, because a hand-started agent never gets the launcher's --mcp-config. The plugin would have fixed that — if it had been installed, and if it had been mentioned anywhere.

2. What was verified, and how

Claim Evidence
A plugin can install globally ~/.claude/plugins/installed_plugins.json — scopes in use are project (5), local (4), user (1)
A plugin can mount an MCP server ~/.claude/plugins/marketplaces/kb-alms/.mcp.json mounts memory at "url": "${KB_MEMORY_URL}"
A plugin can carry skills, agents, commands, hooks kb-alms ships skills/ + hooks/hooks.json; umputun-cc-thingz/plugins/planning ships agents/ + skills/
Env vars interpolate in a plugin's .mcp.json same kb-alms file: ${KB_MEMORY_URL}, ${MEMORY_MCP_TOKEN}
A marketplace can be a plain git repo known_marketplaces.json — mgnl-code-review has "source": "git", "url": "https://..."
This repo is already such a marketplace .claude-plugin/marketplace.json, committed in ef1e014
A lead already binds to a project directory FleetConfig.java:1017 record Leader(..., String workspace, String cwd); used at LeadLauncher.java:193-199
The lead's mount comes from argv, not config LeadLauncher.java:253 adds --mcp-config

2b. Architect review + one measurement changed the design

The architect verified the plan against the code and returned build it with these changes. Two of its findings are load-bearing. I checked both myself.

A plugin cannot carry the agent definitions — confirmed

ClaudeCodeLauncher.java:371 calls agentDefinitionFile(spec.cwd(), spec.role(), ".claude", "agents"), and HerdrPeerLauncher.java:359-365 returns a path only when <cwd>/.claude/agents/<role>.md Files.isRegularFile. ClaudeCodeLauncher.java:391-393 adds --agent only when that returns non-null. OpenCodeLauncher does the same for .opencode/agent.

So the file must exist in the member's worktree. Moving .claude/agents/*.md into the plugin would silently stop every member getting --agent. The agents stay in the repo. My plan had this wrong.

The plugin does not reach members at all — confirmed, and worse than the architect could see

ClaudeCodeLauncher.java:285 does putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir()). A member with configDir set reads that directory, not the operator's ~/.claude.

The architect could not check how far that goes, because fleetd.yaml is gitignored. I measured it:

  • All four Claude profiles on the live fleet set configDir — local, local-direct, opus, sonnet (grep -c configDir fleetd/fleetd.yaml = 4).
  • Each ccs instance has its own plugins/ directory: gx10 (8 entries), ltms (11), ollama (9), work (9).
  • Those directories are four separate real directories with four separate inodes, and installed_plugins.json in each is a separate inode with an identical md5 (51c6e1c853e32e656b817e123fbbfcc5). They are copies made once, not links.

So a plugin installed at user scope lands in exactly one instance's store. It would have to be installed once per CLAUDE_CONFIG_DIR, and each copy would then drift. The plugin is not a delivery mechanism for member-facing assets on this host.

A side effect worth recording: my own session's CLAUDE_CONFIG_DIR is set, so the ~/.claude/plugins/* evidence in section 2 is not even this session's store. The claims about what a plugin can do still hold — they were read from real manifests — but the directory I read them from is the wrong one for any conclusion about this session.

The design that follows

Split by audience, not by mechanism:

Audience Delivered by Carries
operator / lead (a human opening any project) the plugin, per config dir the MCP mount, setup, the bridge charter
member (a worker in a provisioned worktree) worktree provisioning .claude/skills/*, .claude/agents/*

The second row is not a new idea — it is what the code already does for agents, and it is why agentDefinitionFile looks in the worktree. Extending worktree provisioning to seed .claude/skills/ from a fleetd-owned source is the consistent move, and it is what actually fixes "a worker in kb has no implementer skill". The plugin never could.

Other review findings I accepted

  • Mount-name collision is real. PeerLauncher.java:34 is fleet; the plugin mounts fleetd. A lead with both gets two mounts of one daemon and duplicate fleet_* tools. Rename the plugin's server to fleet.
  • Keep --mcp-config in LeadLauncher. It is config→argv from profile.mcpUrl(), not drift, and it is the only path that works for a member with its own configDir.
  • I overstated #359. LeadCoordLoop.java:174-197 returns null and logs a warning that names the fix; tick() leaves the message unacked, so the broker holds it and delivers once a lead is named. It stalls loudly and recovers — it is not silent, and it is not data loss. My wording in issue #361 needs the same correction.
  • Stage 1 ends with tools that mostly cannot be used until stage 3, because authority still comes from the tab. Section 3 already said this; the stage table did not.

3. What a plugin can and cannot do

This is the part that decides the design, so it is stated before the design.

A plugin gives tools. It does not give authority.

fleet_whoami resolves a caller's role from the connection, not from what is mounted. A hand-started session in kb that mounts fleet_* through a plugin will be resolved as a worker and refused on every orchestration call, because its pane is not in a tab matching fleet.leaders.*.tab.

So the plugin alone does not make a project fleet-enabled. It makes it tool-enabled. Registering the lead stays fleetd's job. Any plan that forgets this ships a plugin that looks installed and does nothing.

flowchart TB
    P["fleet plugin<br/>(user scope, every session)"] --> T["fleet_* tools mounted"]
    D["fleetd.yaml<br/>leaders.kb {tab, cwd}"] --> A["role = primary"]
    T --> W["can call fleet_*"]
    A --> W2["calls are authorized"]
    W --> OK["working lead"]
    W2 --> OK
    T --> NO["tools mounted, every call refused"]
    classDef good fill:#2f855a,stroke:#22543d,color:#ffffff;
    classDef bad fill:#9b2c2c,stroke:#63171b,color:#ffffff;
    class OK good
    class NO bad

Both halves are needed. The plugin is the left half only.

4. Proposed architecture — three layers

Layer 1: the plugin — lead-side only, fix the one that exists

Keep it at plugin/, keep the marketplace at .claude-plugin/marketplace.json. Do not create a second one, and do not put member-facing assets in it (see 2b).

.claude-plugin/marketplace.json   -> rename to "fleetd"; keep source ./plugin
plugin/
  .claude-plugin/plugin.json      -> rename to "fleet"; bump version
  .mcp.json                       -> mount name "fleet" (match PeerLauncher.MCP_MOUNT_NAME),
                                     url "${FLEETD_MCP_URL}", 8765 default documented
  skills/setup/SKILL.md           -> EXISTS. fix the stale primary.terminal advice (§5)
  skills/bridge-charter/SKILL.md  -> NEW: the canonical CLAUDE.md block
  README.md                       -> fix the install path (fleet/fleetd, not ltms/claude-bridge)

Not in the plugin: agents/*.md (the launcher requires them in the member's worktree — ClaudeCodeLauncher.java:371,391) and the three worker skills (a member with configDir set never reads the operator's plugin store — measured in 2b). Those belong to layer 1b.

A rename is a breaking change for anyone who installed 0.1.0. The mount name goes fleetd -> fleet, so a project whose .claude/settings.json pre-allows mcp__fleetd__fleet_whoami stops matching. Only this fleet has it installed today, so the cost is small now and grows. Decide once.

This solves propagation of the charter. CLAUDE.md says the bridge block "must stay byte-identical with the template in the wiki" and that "other projects carrying the block need the same edit" — a hand-copy the file itself admits is fragile, with a python snippet to check it. A plugin skill turns that into a version bump.

Layer 1b: worker skills reach members through the worktree, not the plugin

This is the change that actually fixes "a worker in kb cannot load implementer".

Worktree provisioning already writes into the member's tree — the parity overlay, the neutralised .mcp.json, the IDE overlay. Add one more: seed <worktree>/.claude/skills/ from a fleetd-owned source directory, so every member gets implementer, reviewer and hunter whatever repo it is working in. .claude/agents/ is already required there by the launcher, so this follows the grain of the design rather than cutting across it.

Open question for implementation: copy or symlink, and where the source lives (a config key such as memberSkills:, or the plugin's own directory read by the daemon). A symlink is one source of truth but breaks if the member's tree is archived; a copy drifts but is self-contained.

Layer 2: host-global fleet settings

~/.fleet/fleetd.yaml — the things that are true for the machine, not the project:

  • broker: and coordinator: (URIs come from env, no secrets in the file)
  • profiles: — backends, models, credentials, weights
  • memberCredentials: policy
  • worktreeRoot, worktreeGroup

Layer 3: per-project settings, committable

<project>/.fleet/project.yaml — the things that are true for the repo:

lead:
  tab: "lead: kb"
  profile: opus
ide:
  projectDir: ""          # kb is a Python repo, everything at the root
worktree: true

This split fixes a contradiction that exists today. ideProjectDir is a property of a repo (fleetd's Maven module is a subdirectory; kb's code is at the root) but the config key is per-profile. One profile therefore cannot serve both repos — measured on fleet01 on 2026-09-05, where the key had to be commented out to make kb work. Moving it to a project file removes the contradiction rather than working around it.

It is also committable, because it holds no secrets. A project that has been fleet-enabled once stays fleet-enabled for everyone who clones it.

5. The fleet-setup skill

What the operator actually asked for: one command that makes any project fleet-compatible.

sequenceDiagram
    participant Op as Operator
    participant Sk as fleet-setup skill
    participant Fs as project files
    participant Fd as fleetd
    Op->>Sk: /fleet-setup (in any project)
    Sk->>Fs: write .fleet/project.yaml
    Sk->>Fs: add the bridge block to CLAUDE.md (if absent)
    Sk->>Fd: register the lead (tab + cwd)
    Fd-->>Sk: tab created, lead launched
    Sk-->>Op: report what changed, and what is still manual

The skill writes the project half and asks the daemon for the host half.

The registration step needs something that does not exist yet: an MCP tool such as fleet_workspace_add{path, tab, profile}, or a fleetd config include so a project file is picked up without hand-editing the host file. This is the one genuinely new piece of daemon work.

6. Does this reduce fleetd's complexity?

Honestly: partly. Claiming more than this would be wrong.

Yes, in three places.

  1. Skill and agent delivery leaves the daemon and the repos entirely.
  2. Worktree config neutralisation (fleetd #134) gets safer. It blanks .mcp.json so the primary's IDE and forge servers do not leak into a worker. Today the fleet mount survives only because the launcher re-adds it by argv. With a user-scope plugin the fleet mount is outside the file being neutralised, so the two concerns stop fighting.
  3. The ideProjectDir per-profile/per-repo contradiction disappears.

No, in the places that matter most. fleetd still owns spawn, authorization, worktrees, herdr, the broker, tickets, and identity. A plugin cannot do any of those. The plugin is a distribution mechanism, not a replacement for the daemon.

And it adds one new risk. The mount URL becomes a second source of truth. fleetd.yaml has the port; the plugin has the URL. Mitigation: the plugin reads ${FLEETD_MCP_URL} only, and the host env is the single place it is set.

7. Rollout stages

Stage Content Ends with
0 Make it visible. One CLAUDE.md line and one Features entry saying the plugin exists and where nobody re-plans it a third time
1 Fix the plugin's drift: mount name fleet, ${FLEETD_MCP_URL}, names, README, stale primary.terminal advice a lead in any project can install one plugin and get the mount
1b Seed .claude/skills/ into provisioned worktrees a worker in kb can load implementer
2 .fleet/project.yaml schema + FleetConfig reads it; ideProjectDir moves there kb and fleetd both work off one profile
3 Fix #359, then config-include for lead registration, wired into the existing setup skill /fleet:setup in a fresh project produces a working lead
4 Roll out to fleet01; retire the hand-copied CLAUDE.md block in favour of the skill one git pull propagates the charter

Stage 0 is minutes of work and is the one that stops this happening again, so it goes first. Stage 1b carries most of the value and is independent of the plugin — it could ship first if the plugin rename needs more thought. Stage 1 alone ends with tools a hand-started session mostly cannot use, because authority still comes from the tab; that is fixed in stage 3, not stage 1. #359 moves ahead of stage 3 on the architect's advice, because stage 3 is what creates the second lead.

8. Questions for the architect

  1. Is the layer-2 / layer-3 split right? Specifically: should profiles: stay host-global, or should a project be able to pin which profiles it uses? Cost of getting this wrong is a config that has to be re-split later.
  2. Config include, or a new MCP tool, for registering a project's lead? An include is passive and survives a restart; a tool is live but writes to a gitignored file the daemon owns.
  3. What happens when the plugin is absent? Should LeadLauncher keep its --mcp-config belt-and-braces, or is that the drift risk we should remove? Note opencode members cannot use Claude plugins at all, so OpenCodeLauncher keeps its ephemeral config either way.
  4. Does a user-scope plugin mount leak into members in a way we do not want? Members already inherit user-scope MCP servers (--mcp-config adds, it does not replace). A worker getting fleet_* is correct and already happens. Confirm nothing else in the plugin should be worker-invisible.
  5. Two leads on one host both hold a subscription seat. Is per-project leads the right unit, or should one lead serve several projects by changing cwd?
  6. Blocking defect to fix first or alongside: LeadCoordLoop.resolveLocalLead() (lines 174-190) routes a peer message to "the sole lead" when no lead is named after coordinator.selfId. The moment a host has two leads — exactly what this plan encourages — cross-host coordination silently stops. Tracked as #359.