fleet manager: run and drive several fleets on ONE host from a single session #160

Open
opened 2026-08-23 14:26:11 +02:00 by ltms · 1 comment
Owner

Design ticket. This is not #6. #6 (CB-308) is multi-host federation: one fleet spread over many machines. This is the opposite and much nearer: many fleets on one machine, driven from one Claude or OpenCode session, behind a fleet-manager MCP.

The motivating case is already real. One fleet works on fleetd; another should work on kb. A lead must be started in the project directory its fleet works on, so one fleet per project is the natural unit. Today an operator can only talk to one.

What a "fleet" is here

One fleetd process, plus the herdr session it drives, plus the project it works in. So a second fleet on this host means a second fleetd with its own port, its own herdr session, its own worktree root and its own broker vhost.

The good news: the config is already almost entirely per-instance

I read FleetConfig looking for things that would stop two daemons sharing a host. There are very few.

concern status
listen port bind.port, configurable. 8765 appears once in the codebase, as the default
herdr socket herdrSocket, configurable. ~/.config/herdr/herdr.sock is only the fallback
worktree root worktreeRoot, configurable
broker broker.uriEnv names an env var, so each daemon can point at its own vhost
member MCP mount Profile.mcpUrl, per profile, so a member can be pointed at its own daemon
pid/lock files none. Nothing in the daemon assumes it is the only one

fleet01 already proves the herdr half: it runs two herdr sessions side by side, and fleetd is pinned to the named one.

So "run two fleetds on one host" is mostly a packaging and convention problem, not a code problem. That is the cheap part.

Blocker 1: a fleet has no name

There is no fleetName, no instance id, nothing. I grepped for every spelling.

A manager cannot address what has no name. Every tool the manager offers needs a fleet argument, every roster row needs to say which fleet it came from, and every log line an operator reads needs to say which daemon wrote it. This is the first unit of work and everything else depends on it.

Blocker 2: the authz model breaks when one session talks to several fleets

This is the real problem, and it is not obvious.

CallerResolver.resolve decides who you are from the connection: remote address and port → peer pid → herdr terminal. Then:

  • if that terminal is named in this daemon's fleet.leaders.*.tab, you are a lead;
  • else, if the caller is in a herdr pane at all, you are a worker;
  • only a caller in no pane falls through to loopback-trust and becomes primary.

An operator's session normally runs inside a herdr pane — that is why fleet.leaders.*.tab exists at all. So if that session mounts a second daemon whose config does not name its tab, it is not a lead there. It falls to the worker branch, and every orchestration call is refused.

This means the zero-code option does not work. "Just put two fleetd servers in .mcp.json" looks like it should be enough — Claude Code mounts many MCP servers happily, and each would namespace its tools. But the second daemon would resolve the operator as a worker and refuse to spawn, send or stop.

There are three ways out, and choosing between them is the main design decision here:

  1. Name the operator's tab as a lead in every fleet's config. No code at all. It does not scale, it duplicates config, and it means every fleet trusts the same pane — but it would work today and it is worth trying first, because it tells us whether anything else breaks.
  2. A proxy manager MCP. One MCP server the session mounts; it holds a registry of fleets and forwards each call to the right daemon, adding a fleet argument. Because the proxy is not in a herdr pane, each daemon resolves it as primary, so the demotion problem disappears by construction.
  3. Teach authz about a manager principal. Cleanest in the long run, most work, and it widens the trust surface of the one part of this system that is deliberately unforgeable.

I lean to 2, with 1 run first as a one-hour experiment to find the unknown unknowns.

If we build the proxy, one rule is not negotiable

Members must never mount the manager. A member's identity works precisely because its connection comes from its own pane. If a member called through the proxy, the proxy's connection would be resolved instead, and every member would be seen as primary. That is a full privilege escalation.

So: members keep pointing at their own daemon via Profile.mcpUrl; the manager is an operator-side front door only. Whatever is built has to make that hard to get wrong, not merely documented.

Blocker 3: nothing starts, stops or supervises a fleet

Today a fleet is started by hand — the Mac has a launchd plist, fleet01 has a shell script and no supervisor at all (#103, and #156 §8). A manager that can list fleets but cannot start or restart one is a viewer, not a manager.

This is where this ticket and the ops gaps meet. "Start fleet kb" has to mean something.

What is worth deciding before any code

  • Does the manager own daemon lifecycle, or only conversation routing? A router is much smaller and useful on day one. A lifecycle manager is what the name promises.
  • Where does the fleet registry live? A config file the manager reads, or discovery by probing ports, or each daemon announcing itself on the shared broker. The broker option is interesting because the fabric already exists and already isolates by vhost — but a fleet that is down cannot announce itself, and a manager most needs to talk about fleets that are down.
  • How are ports allocated? By hand in each config is fine for two fleets and bad for ten.
  • One vhost per fleet, or one per host? Queue names are agent.<target>.inbox with no fleet namespace, so two fleets sharing a vhost would collide. One vhost per fleet avoids it and matches what we just did for the two hosts.
  • Tool surface. N fleets behind one mount must not mean N× the tools. A fleet argument on the existing tools is the obvious answer, but it changes every call site and every brief.

Suggested first units

  1. Give a fleet a name. name: in config, in /healthz, in the startup log, and in every roster row. Small, unblocks everything, useful even with one fleet.
  2. Run two fleetds on this host by hand, using option 1 above (name the operator's tab as a lead in both). Write down what breaks. This is a measurement, not a feature — but every later decision depends on it, and I would rather find the surprises now.
  3. Then choose the manager shape, informed by 2.

Related

  • #156 — the as-built survey of fleet01, including the config and lifecycle facts this ticket relies on.
  • #103 — no supervision on Linux; blocker 3's prerequisite.
  • #118 (CB-615) — there is no way to see the fleet. A manager and a dashboard want the same roster; they should not invent two.
  • #6 (CB-308) — multi-host federation. Different problem, but a fleet name is needed by both, so unit 1 should be designed so it does not have to be redone there.

Not verified

I reasoned blocker 2 from CallerResolver.resolve and the fleet.leaders.*.tab mechanism. I did not stand up a second daemon and watch a session get demoted to worker. That is exactly what suggested unit 2 would prove, and it should be proven before any manager design is committed to.

Design ticket. **This is not #6.** #6 (CB-308) is multi-*host* federation: one fleet spread over many machines. This is the opposite and much nearer: **many fleets on one machine, driven from one Claude or OpenCode session**, behind a fleet-manager MCP. The motivating case is already real. One fleet works on `fleetd`; another should work on `kb`. A lead must be started in the project directory its fleet works on, so one fleet per project is the natural unit. Today an operator can only talk to one. ## What a "fleet" is here One `fleetd` process, plus the herdr session it drives, plus the project it works in. So a second fleet on this host means a second `fleetd` with its own port, its own herdr session, its own worktree root and its own broker vhost. ## The good news: the config is already almost entirely per-instance I read `FleetConfig` looking for things that would stop two daemons sharing a host. There are very few. | concern | status | |---|---| | listen port | `bind.port`, configurable. `8765` appears **once** in the codebase, as the default | | herdr socket | `herdrSocket`, configurable. `~/.config/herdr/herdr.sock` is only the fallback | | worktree root | `worktreeRoot`, configurable | | broker | `broker.uriEnv` names an env var, so each daemon can point at its own vhost | | member MCP mount | `Profile.mcpUrl`, per profile, so a member can be pointed at its own daemon | | pid/lock files | none. Nothing in the daemon assumes it is the only one | `fleet01` already proves the herdr half: it runs **two** herdr sessions side by side, and `fleetd` is pinned to the named one. So "run two fleetds on one host" is mostly a packaging and convention problem, not a code problem. That is the cheap part. ## Blocker 1: a fleet has no name There is no `fleetName`, no instance id, nothing. I grepped for every spelling. A manager cannot address what has no name. Every tool the manager offers needs a `fleet` argument, every roster row needs to say which fleet it came from, and every log line an operator reads needs to say which daemon wrote it. This is the first unit of work and everything else depends on it. ## Blocker 2: the authz model breaks when one session talks to several fleets This is the real problem, and it is not obvious. `CallerResolver.resolve` decides who you are from the **connection**: remote address and port → peer pid → herdr terminal. Then: - if that terminal is named in **this daemon's** `fleet.leaders.*.tab`, you are a lead; - else, if the caller is in a herdr pane at all, you are a **worker**; - only a caller in no pane falls through to loopback-trust and becomes `primary`. An operator's session normally runs **inside a herdr pane** — that is why `fleet.leaders.*.tab` exists at all. So if that session mounts a second daemon whose config does not name its tab, it is not a lead there. It falls to the worker branch, and every orchestration call is refused. **This means the zero-code option does not work.** "Just put two fleetd servers in `.mcp.json`" looks like it should be enough — Claude Code mounts many MCP servers happily, and each would namespace its tools. But the second daemon would resolve the operator as a worker and refuse to spawn, send or stop. There are three ways out, and choosing between them is the main design decision here: 1. **Name the operator's tab as a lead in every fleet's config.** No code at all. It does not scale, it duplicates config, and it means every fleet trusts the same pane — but it would work today and it is worth trying first, because it tells us whether anything *else* breaks. 2. **A proxy manager MCP.** One MCP server the session mounts; it holds a registry of fleets and forwards each call to the right daemon, adding a `fleet` argument. Because the proxy is not in a herdr pane, each daemon resolves *it* as `primary`, so the demotion problem disappears by construction. 3. **Teach authz about a manager principal.** Cleanest in the long run, most work, and it widens the trust surface of the one part of this system that is deliberately unforgeable. I lean to **2**, with **1** run first as a one-hour experiment to find the unknown unknowns. ### If we build the proxy, one rule is not negotiable **Members must never mount the manager.** A member's identity works precisely because its connection comes from its own pane. If a member called through the proxy, the proxy's connection would be resolved instead, and every member would be seen as `primary`. That is a full privilege escalation. So: members keep pointing at their own daemon via `Profile.mcpUrl`; the manager is an operator-side front door only. Whatever is built has to make that hard to get wrong, not merely documented. ## Blocker 3: nothing starts, stops or supervises a fleet Today a fleet is started by hand — the Mac has a launchd plist, `fleet01` has a shell script and no supervisor at all (#103, and #156 §8). A manager that can list fleets but cannot start or restart one is a viewer, not a manager. This is where this ticket and the ops gaps meet. "Start fleet `kb`" has to mean something. ## What is worth deciding before any code - **Does the manager own daemon lifecycle, or only conversation routing?** A router is much smaller and useful on day one. A lifecycle manager is what the name promises. - **Where does the fleet registry live?** A config file the manager reads, or discovery by probing ports, or each daemon announcing itself on the shared broker. The broker option is interesting because the fabric already exists and already isolates by vhost — but a fleet that is down cannot announce itself, and a manager most needs to talk about fleets that are down. - **How are ports allocated?** By hand in each config is fine for two fleets and bad for ten. - **One vhost per fleet, or one per host?** Queue names are `agent.<target>.inbox` with no fleet namespace, so two fleets sharing a vhost would collide. One vhost per fleet avoids it and matches what we just did for the two hosts. - **Tool surface.** N fleets behind one mount must not mean N× the tools. A `fleet` argument on the existing tools is the obvious answer, but it changes every call site and every brief. ## Suggested first units 1. **Give a fleet a name.** `name:` in config, in `/healthz`, in the startup log, and in every roster row. Small, unblocks everything, useful even with one fleet. 2. **Run two fleetds on this host by hand**, using option 1 above (name the operator's tab as a lead in both). Write down what breaks. This is a measurement, not a feature — but every later decision depends on it, and I would rather find the surprises now. 3. Then choose the manager shape, informed by 2. ## Related - **#156** — the as-built survey of `fleet01`, including the config and lifecycle facts this ticket relies on. - **#103** — no supervision on Linux; blocker 3's prerequisite. - **#118** (CB-615) — there is no way to see the fleet. A manager and a dashboard want the same roster; they should not invent two. - **#6** (CB-308) — multi-*host* federation. Different problem, but a fleet name is needed by both, so unit 1 should be designed so it does not have to be redone there. ## Not verified I reasoned blocker 2 from `CallerResolver.resolve` and the `fleet.leaders.*.tab` mechanism. I did **not** stand up a second daemon and watch a session get demoted to worker. That is exactly what suggested unit 2 would prove, and it should be proven before any manager design is committed to.
Author
Owner

Experiment run — the "not verified" parts are now measured

I ran the two-daemon experiment from the bottom of this ticket. Two fleetds on one host, daemon 2 on port 8766, sharing the one herdr session, with a config that does not name my pane as a lead. Everything below is measured, not reasoned.

1. The proxy mechanism works (confirmed)

A process that is not in a herdr pane is resolved as primary by every fleetd on the host. I called fleet_whoami from a plain Python process against both daemons — both answered primary. And a lead-only fleet_stop from that non-pane process succeeded against both the experimental daemon and the live one. So a proxy manager MCP, which is not in a pane, is seen as primary by each fleet it fronts. Blocker 2's proposed fix (option 2) holds up.

2. Two daemons must NOT share one herdr session — they kill each other (new, important)

The moment daemon 2 booted, it ran its startup orphan-reap and stopped daemon 1's live member:

reaped orphan opencode opencode-xf-f649c6-1 (pane=w3:p61 tab=w3:t5V) left by a prior daemon
orphan-peer reap complete — 1 stale opencode peer(s) removed at startup

The reap keys on a random per-process nameNonce (HerdrPeerLauncher), so every daemon treats every other daemon's members as its own leaked orphans and tears them down. And daemon 1 never noticed: fleet_list still showed the killed member as state: ready with its capacity slot counted; only liveStatus flipped to unknown, and fleet_status on it returned agent_not_found.

Consequence for the design: each fleet needs its own herdr session (fleet01 already runs two side by side, so the mechanism exists). Co-tenancy on one herdr session is not just untidy, it is destructive. This is a hard requirement, not a preference.

3. Starting a second daemon silently spawns a second subscription lead (new)

Daemon 2's LeadLauncher found no pane matching its tab and launched one — a fresh opus (subscription) Claude Code lead. So naively starting a second fleet burns a second subscription lead with no prompt. A manager that starts fleets must own lead launch, or make it opt-in.

4. Identity is per-connecting-process, which is a single-fleet escalation risk (new, NOT fully verified)

Identity comes from the pid that opens the connection, matched against herdr's shell_pid and foreground pids only (PaneLocator.paneOwnsPid). A grandchild process of a pane (e.g. python3 a worker shells out to) matches neither, so it falls through to loopback-trust and is resolved as primary. My non-pane probe demonstrates the mechanism.

I tried to confirm it from inside a worker pane — I asked a real worker to run the lead-only probe — and could not: the worker's own local safety classifier blocked the script before it ran. So the in-pane escalation is plausible from the code and unproven in practice. It should be pinned down separately, because if it holds, a worker can escalate to primary on its own fleet just by shelling out, independent of any manager. Filing that as its own security ticket.

Direction from the operator (2026-08-23)

  • The fleet manager is a separate module or git repo — not mixed into the current modules. Recorded.
  • The first feature is a fleets-status tool: across fleets, what each is doing, which is working, which is stalled.
  • A dashboard is the intended future, built on the same status data.

Status tool — working prototype

I built it as a standalone probe that reads each fleet's REST API and touches no fleetd Java — already honoring "don't mix with current modules". It reads a small fleets.json registry (name → REST url → project), and for each fleet reports health, roster, and a per-member verdict.

The valuable part is telling working from stalled, because liveStatus alone lies — a member reads "working" for hours doing nothing. So the tool does not trust the flag: for any member claiming to be busy it samples the worktree's newest file mtime twice, a few seconds apart. Files changed ⇒ WORKING; unchanged ⇒ STALLED, with the age of the newest file.

Proven against the live fleet, all three states real, not mocked:

  • a member given real file work → WORKING (files changed)
  • the same member after it finished → idle, correctly not called working
  • fleet01 (other host) → remote (not probed), stated honestly rather than guessed

Known limit, written into the tool: a member that is legitimately thinking without writing files reads as stalled. And it sees only this host's fleets, because fleetd's REST binds to loopback — cross-host status needs the broker or ssh, which is the CB-308 boundary, not this one.

I have the prototype and its output ready to show. The one open decision is packaging: a separate git repo vs a top-level non-Maven directory in this repo. I lean to a separate repo, matching how fleet-agent-kit was planned, so the manager and the coming dashboard get their own lifecycle and CI and never entangle with the Java build.

## Experiment run — the "not verified" parts are now measured I ran the two-daemon experiment from the bottom of this ticket. Two fleetds on one host, daemon 2 on port 8766, sharing the one herdr session, with a config that does **not** name my pane as a lead. Everything below is measured, not reasoned. ### 1. The proxy mechanism works (confirmed) A process that is **not** in a herdr pane is resolved as `primary` by every fleetd on the host. I called `fleet_whoami` from a plain Python process against both daemons — both answered `primary`. And a lead-only `fleet_stop` from that non-pane process **succeeded** against both the experimental daemon *and* the live one. So a proxy manager MCP, which is not in a pane, is seen as `primary` by each fleet it fronts. Blocker 2's proposed fix (option 2) holds up. ### 2. Two daemons must NOT share one herdr session — they kill each other (new, important) The moment daemon 2 booted, it ran its startup orphan-reap and **stopped daemon 1's live member**: ``` reaped orphan opencode opencode-xf-f649c6-1 (pane=w3:p61 tab=w3:t5V) left by a prior daemon orphan-peer reap complete — 1 stale opencode peer(s) removed at startup ``` The reap keys on a random per-process `nameNonce` (`HerdrPeerLauncher`), so every daemon treats every *other* daemon's members as its own leaked orphans and tears them down. And **daemon 1 never noticed**: `fleet_list` still showed the killed member as `state: ready` with its capacity slot counted; only `liveStatus` flipped to `unknown`, and `fleet_status` on it returned `agent_not_found`. Consequence for the design: **each fleet needs its own herdr session** (fleet01 already runs two side by side, so the mechanism exists). Co-tenancy on one herdr session is not just untidy, it is destructive. This is a hard requirement, not a preference. ### 3. Starting a second daemon silently spawns a second subscription lead (new) Daemon 2's `LeadLauncher` found no pane matching its tab and **launched one** — a fresh `opus` (subscription) Claude Code lead. So naively starting a second fleet burns a second subscription lead with no prompt. A manager that starts fleets must own lead launch, or make it opt-in. ### 4. Identity is per-connecting-process, which is a single-fleet escalation risk (new, NOT fully verified) Identity comes from the pid that opens the connection, matched against herdr's `shell_pid` and foreground pids only (`PaneLocator.paneOwnsPid`). A **grandchild** process of a pane (e.g. `python3` a worker shells out to) matches neither, so it falls through to loopback-trust and is resolved as `primary`. My non-pane probe demonstrates the mechanism. I tried to confirm it from *inside* a worker pane — I asked a real worker to run the lead-only probe — and **could not**: the worker's own local safety classifier blocked the script before it ran. So the in-pane escalation is **plausible from the code and unproven in practice**. It should be pinned down separately, because if it holds, a worker can escalate to primary on its *own* fleet just by shelling out, independent of any manager. Filing that as its own security ticket. ## Direction from the operator (2026-08-23) - The fleet manager is a **separate module or git repo — not mixed into the current modules.** Recorded. - The **first feature is a fleets-status tool**: across fleets, what each is doing, which is working, which is stalled. - A **dashboard is the intended future**, built on the same status data. ## Status tool — working prototype I built it as a standalone probe that reads each fleet's REST API and touches **no** fleetd Java — already honoring "don't mix with current modules". It reads a small `fleets.json` registry (name → REST url → project), and for each fleet reports health, roster, and a per-member verdict. The valuable part is telling **working** from **stalled**, because `liveStatus` alone lies — a member reads "working" for hours doing nothing. So the tool does not trust the flag: for any member claiming to be busy it samples the worktree's newest file mtime **twice**, a few seconds apart. Files changed ⇒ WORKING; unchanged ⇒ STALLED, with the age of the newest file. Proven against the live fleet, all three states real, not mocked: - a member given real file work → `WORKING (files changed)` - the same member after it finished → `idle`, correctly not called working - fleet01 (other host) → `remote (not probed)`, stated honestly rather than guessed Known limit, written into the tool: a member that is legitimately thinking without writing files reads as stalled. And it sees only this host's fleets, because fleetd's REST binds to loopback — cross-host status needs the broker or ssh, which is the CB-308 boundary, not this one. I have the prototype and its output ready to show. The one open decision is packaging: a **separate git repo** vs a **top-level non-Maven directory in this repo**. I lean to a separate repo, matching how `fleet-agent-kit` was planned, so the manager and the coming dashboard get their own lifecycle and CI and never entangle with the Java build.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#160