CB-621 (epic): rename the product to "fleet" and move it to its own Gitea org #125

Open
opened 2026-08-22 21:34:02 +02:00 by ltms · 1 comment
Owner

Epic. The product outgrew both its name and its home. This ticket holds the decision and the order of work; each unit is its own ticket.

Decisions taken

1. The product is called fleet. The codebase already speaks this word: 258 uses across code, config and CLAUDE.md, including the fleet: config block, FleetHealth, FleetHealthMonitor, and fleet.leaders.*.tab (the lead-identity pin). The alternative, crew, appears 0 times. Picking fleet renames the artefacts around a word the code already uses. Picking anything else would have reassigned a word that already means something here, which is the more expensive kind of rename.

The daemon becomes fleetd. lead and member do not change.

2. The MCP tool namespace renames too: bridge_* -> fleet_*. Consistency was the deciding factor. A product called fleet whose tools are called bridge_* makes the code and the docs disagree permanently.

3. The lms org is NOT renamed. Measured: lms holds alms (OpenOLAT fork), alms-fe, alms-memory and claude-bridge. Three of four are the LMS product. This repo was parked there. So create a new org and transfer the repo out, leaving lms untouched.

Withdrawn after review

An independent review checked the first draft against the import graph and found two parts that had to go.

The 5-module Maven split is withdrawn. It was proposed as a layout choice. It is really a refactor, and its boundaries contradict the actual dependencies. Three cycles, all verified in the code:

cycle evidence
auth <-> mcp auth/CallerResolver.java:3 imports mcp.ConnectionIdentity; mcp/BridgeMcp.java has 5 auth imports
msg <-> mcp msg/LeadHeartbeatLoop.java:5 and msg/ReplyPushLoop.java:5 import mcp.PrimaryRegistry; BridgeMcp imports msg
core <-> peer-herdr 20 files across ~8 packages import herdr, while member imports config, guard, herdr, peer, placement

The draft also named only 8 packages. Package sizes total 16,795 LOC, so 7,307 LOC had no module at all - including msg at 2,665, the largest package in the project, and config at 2,225.

Doing that refactor during an org transfer risks both. The repo stays one Maven module. Boundaries get enforced by an ArchUnit test instead (CB-627). Revisit only when a real second consumer exists.

The fleet-registry repo is withdrawn for now. It has zero code behind it, and multi-host is a next-version feature. A fleetd-registry-client module with nothing in it is the "a silent default disables the feature" shape that has already cost this project several sessions. docs/CB-308-Multi-Host-Federation.md is the design home until the stage is scheduled.

Target layout

git.ltms.dev/fleet/                NEW org.  lms keeps alms, alms-fe, alms-memory.
  |
  +-- fleetd/         PRIVATE   transferred from lms/claude-bridge.
  |                             stays ONE Maven module.
  |                             docs/, e2e/, deploy/, scripts/ stay inside.
  |
  \-- fleet-agent-kit/ PUBLIC   role bodies + skills written once, generated for
                                both backends, with a drift check fleetd CI runs.

Two repos, not the four first proposed. No per-backend peer repos: HerdrPeerLauncher is 1025 lines and both adapters extend it (ClaudeCodeLauncher 427, OpenCodeLauncher 512), so the shared base is bigger than both adapters together. Splitting them would mean releasing and version-bumping that base for every change. No Codex adapter: dropped, with the work parked on salvage/cb-528a-codex-launcher and salvage/cb-528b-codex-home.

Risks measured before starting

  • Redirects hold. A probe on a throwaway repo showed that after a rename the web and API both 301 and git ls-remote on the old SSH URL still works; after a further cross-org transfer, redirects chain from the original org and name.
  • CI survives the transfer. Measured: lms has 0 org-level runners and the repo has 0 repo-level runners, yet CI has 384 runs with a green one today. So the runners are instance-level and serve every org. .gitea/workflows/ also references no secrets.* at all.
  • The wiki submodule does not survive by itself. .gitmodules pins lms/claude-bridge.wiki.git. The probe tested the repo redirect, not the .wiki.git companion. Handled in CB-623.
  • The blocking PRs are gone. #121 and #122 were both closed as already-landed; Gitea reports changed_files: 0 for each.

Order of work - this is the part that matters

  1. CB-622 - bridge_* -> fleet_*, in place, before any move. Least disruptive while the repo is still at its old path.
  2. CB-623 - create the fleet org, transfer the repo, rename it to fleetd. Then git remote set-url origin in the existing clone. Never reclone - a fresh clone loses the --skip-worktree flag on .mcp.json (verified: git ls-files -v returns S), and the local file holding the IDE and forge MCP servers gets clobbered on the next checkout.
  3. CB-624 - rename the local checkout directory. Alone, in its own step, together with the 5 hardcoded paths in deploy/dev.ltms.bridged.plist, the live worktree at /Users/dai.ha/LTMS/.bridged-worktrees/cb401-manual, and CLAUDE.md:315.
  4. CB-625 - extract fleet-agent-kit.
  5. CB-626 - config key fleet: -> roles: with a read-both shim.
  6. CB-627 - ArchUnit boundary test.

Steps 2 and 3 are separate on purpose. If issue links or the wiki remote break, we need to know which step did it.

Writing rule that comes with the name

JetBrains ships an IDE called Fleet, and our docs already discuss JetBrains tooling. In any document that also says fleet, always write "JetBrains Fleet" in full and never bare "Fleet".

Epic. The product outgrew both its name and its home. This ticket holds the decision and the order of work; each unit is its own ticket. ## Decisions taken **1. The product is called `fleet`.** The codebase already speaks this word: 258 uses across code, config and `CLAUDE.md`, including the `fleet:` config block, `FleetHealth`, `FleetHealthMonitor`, and `fleet.leaders.*.tab` (the lead-identity pin). The alternative, `crew`, appears 0 times. Picking `fleet` renames the artefacts around a word the code already uses. Picking anything else would have reassigned a word that already means something here, which is the more expensive kind of rename. The daemon becomes `fleetd`. `lead` and `member` do not change. **2. The MCP tool namespace renames too: `bridge_*` -> `fleet_*`.** Consistency was the deciding factor. A product called fleet whose tools are called `bridge_*` makes the code and the docs disagree permanently. **3. The `lms` org is NOT renamed.** Measured: `lms` holds `alms` (OpenOLAT fork), `alms-fe`, `alms-memory` and `claude-bridge`. Three of four are the LMS product. This repo was parked there. So create a new org and transfer the repo out, leaving `lms` untouched. ## Withdrawn after review An independent review checked the first draft against the import graph and found two parts that had to go. **The 5-module Maven split is withdrawn.** It was proposed as a layout choice. It is really a refactor, and its boundaries contradict the actual dependencies. Three cycles, all verified in the code: | cycle | evidence | |---|---| | `auth` <-> `mcp` | `auth/CallerResolver.java:3` imports `mcp.ConnectionIdentity`; `mcp/BridgeMcp.java` has 5 `auth` imports | | `msg` <-> `mcp` | `msg/LeadHeartbeatLoop.java:5` and `msg/ReplyPushLoop.java:5` import `mcp.PrimaryRegistry`; `BridgeMcp` imports `msg` | | core <-> peer-herdr | 20 files across ~8 packages import `herdr`, while `member` imports `config`, `guard`, `herdr`, `peer`, `placement` | The draft also named only 8 packages. Package sizes total 16,795 LOC, so **7,307 LOC had no module at all** - including `msg` at 2,665, the largest package in the project, and `config` at 2,225. Doing that refactor during an org transfer risks both. The repo stays **one Maven module**. Boundaries get enforced by an ArchUnit test instead (CB-627). Revisit only when a real second consumer exists. **The `fleet-registry` repo is withdrawn for now.** It has zero code behind it, and multi-host is a next-version feature. A `fleetd-registry-client` module with nothing in it is the "a silent default disables the feature" shape that has already cost this project several sessions. `docs/CB-308-Multi-Host-Federation.md` is the design home until the stage is scheduled. ## Target layout ``` git.ltms.dev/fleet/ NEW org. lms keeps alms, alms-fe, alms-memory. | +-- fleetd/ PRIVATE transferred from lms/claude-bridge. | stays ONE Maven module. | docs/, e2e/, deploy/, scripts/ stay inside. | \-- fleet-agent-kit/ PUBLIC role bodies + skills written once, generated for both backends, with a drift check fleetd CI runs. ``` Two repos, not the four first proposed. No per-backend peer repos: `HerdrPeerLauncher` is 1025 lines and both adapters extend it (`ClaudeCodeLauncher` 427, `OpenCodeLauncher` 512), so the shared base is bigger than both adapters together. Splitting them would mean releasing and version-bumping that base for every change. No Codex adapter: dropped, with the work parked on `salvage/cb-528a-codex-launcher` and `salvage/cb-528b-codex-home`. ## Risks measured before starting - **Redirects hold.** A probe on a throwaway repo showed that after a rename the web and API both 301 and `git ls-remote` on the old SSH URL still works; after a further cross-org transfer, redirects chain from the original org and name. - **CI survives the transfer.** Measured: `lms` has 0 org-level runners and the repo has 0 repo-level runners, yet CI has 384 runs with a green one today. So the runners are instance-level and serve every org. `.gitea/workflows/` also references no `secrets.*` at all. - **The wiki submodule does not survive by itself.** `.gitmodules` pins `lms/claude-bridge.wiki.git`. The probe tested the repo redirect, not the `.wiki.git` companion. Handled in CB-623. - **The blocking PRs are gone.** #121 and #122 were both closed as already-landed; Gitea reports `changed_files: 0` for each. ## Order of work - this is the part that matters 1. **CB-622** - `bridge_*` -> `fleet_*`, in place, before any move. Least disruptive while the repo is still at its old path. 2. **CB-623** - create the `fleet` org, transfer the repo, rename it to `fleetd`. Then `git remote set-url origin` in the existing clone. **Never reclone** - a fresh clone loses the `--skip-worktree` flag on `.mcp.json` (verified: `git ls-files -v` returns `S`), and the local file holding the IDE and forge MCP servers gets clobbered on the next checkout. 3. **CB-624** - rename the local checkout directory. **Alone, in its own step**, together with the 5 hardcoded paths in `deploy/dev.ltms.bridged.plist`, the live worktree at `/Users/dai.ha/LTMS/.bridged-worktrees/cb401-manual`, and `CLAUDE.md:315`. 4. **CB-625** - extract `fleet-agent-kit`. 5. **CB-626** - config key `fleet:` -> `roles:` with a read-both shim. 6. **CB-627** - ArchUnit boundary test. Steps 2 and 3 are separate on purpose. If issue links or the wiki remote break, we need to know which step did it. ## Writing rule that comes with the name JetBrains ships an IDE called Fleet, and our docs already discuss JetBrains tooling. In any document that also says fleet, **always write "JetBrains Fleet" in full and never bare "Fleet"**.
ltms added this to the 2.0 — one operation centre, many hosts milestone 2026-08-22 21:34:02 +02:00
Author
Owner

How the operator reaches a remote fleet: herdr --remote, not a bridged transport

Recorded 2026-08-23 after checking a proposal to containerise bridged and have it drive the host's herdr. That shape does not work, and the reason points at the right design for the vhost stage.

What herdr --remote actually is

--remote <target>              Attach through SSH to a remote Herdr server
--remote-keybindings <local|server>   Keybindings for --remote app attach
--handoff                      Opt into live handoff for update or remote attach

It sits beside keybindings and live handoff, so it is for putting a human's terminal on a remote herdr. Every control verb in the herdr CLI is described as "over the socket API", and bridged speaks only that:

// Bridged.java:119 — unix socket. No TCP, no SSH.
Path socket = Path.of(cfg.herdrSocket());

Forwarding the socket into a container over SSH is possible but is a workaround, and it was not tested.

Why "bridged in a container, herdr on the host" fails

The blocker is the filesystem, not the transport. bridged writes to local disk, then hands absolute path strings to herdr, which uses them on the host. Four seams, all verified in the code:

bridged writes (container) member reads (host)
GitWorktrees shells git -C <repoRoot> worktree add locally the member's cwd must exist on the host
Files.writeString(charter) --append-system-prompt-file <path>
OpenCodeLauncher.writeConfig() into a temp dir OPENCODE_CONFIG=<path>
overlay stubs written into the worktree .mcp.json / opencode.json neutralisation

WorkspaceControl.createTab(workspaceId, cwd, env) transmits the path string only. So the worktree lands in the container while the PTY opens on the host, and the member looks for a directory that is not there. Making it work would need the repo, the worktree root and the temp config root bind-mounted at byte-identical absolute paths in both places — and git worktrees additionally store absolute gitdir pointers, which is already half of CB-624.

It would also not deprecate the daemon. bridged stays a long-running process; Docker's restart policy would merely replace launchd as its supervisor. The members would still execute on the host, so the component that runs untrusted work gains no isolation at all — only bridged does, and bridged is the least risky part.

The design this implies for the vhost stage

One vhost holds bridged, herdr, the members and the project together. One filesystem, so no path contract exists to break, and the isolation is real because the members are inside it.

herdr --remote <vhost> is then the operator's attach path, not bridged's transport. The fleet runs on the vhost; you attach your local terminal to it over SSH. That is exactly what the flag was built for.

   LAPTOP                          VHOST
   ──────                          ─────
   herdr --remote <vhost>  ──SSH──►  herdr server
                                       │ unix socket
                                       ▼
                                     fleetd  ──► members, worktrees, project
                                     (all one filesystem)

Consequences for this epic

  • No containerisation work belongs in CB-621..CB-627. The vhost stage supersedes it.
  • When the vhost provisioning ticket is written, it must place herdr and fleetd on the same host as the worktrees, and name herdr --remote as the documented operator access path.
  • If a container is ever wanted, the only coherent version puts herdr inside it too. Half-containerising fleetd adds a bind-mount contract at four seams and buys a restart policy that launchd already provides (installed and verified 2026-08-22: kill -9 → restarted in 2s).
## How the operator reaches a remote fleet: `herdr --remote`, not a bridged transport Recorded 2026-08-23 after checking a proposal to containerise bridged and have it drive the **host's** herdr. That shape does not work, and the reason points at the right design for the vhost stage. ### What `herdr --remote` actually is ``` --remote <target> Attach through SSH to a remote Herdr server --remote-keybindings <local|server> Keybindings for --remote app attach --handoff Opt into live handoff for update or remote attach ``` It sits beside *keybindings* and *live handoff*, so it is for putting **a human's terminal** on a remote herdr. Every control verb in the herdr CLI is described as "over the socket API", and bridged speaks only that: ```java // Bridged.java:119 — unix socket. No TCP, no SSH. Path socket = Path.of(cfg.herdrSocket()); ``` Forwarding the socket into a container over SSH is possible but is a workaround, and it was not tested. ### Why "bridged in a container, herdr on the host" fails The blocker is the filesystem, not the transport. bridged writes to local disk, then hands **absolute path strings** to herdr, which uses them on the host. Four seams, all verified in the code: | bridged writes (container) | member reads (host) | |---|---| | `GitWorktrees` shells `git -C <repoRoot> worktree add` locally | the member's `cwd` must exist on the host | | `Files.writeString(charter)` | `--append-system-prompt-file <path>` | | `OpenCodeLauncher.writeConfig()` into a temp dir | `OPENCODE_CONFIG=<path>` | | overlay stubs written into the worktree | `.mcp.json` / `opencode.json` neutralisation | `WorkspaceControl.createTab(workspaceId, cwd, env)` transmits the path **string** only. So the worktree lands in the container while the PTY opens on the host, and the member looks for a directory that is not there. Making it work would need the repo, the worktree root and the temp config root bind-mounted at byte-identical absolute paths in both places — and git worktrees additionally store absolute `gitdir` pointers, which is already half of CB-624. It would also not deprecate the daemon. bridged stays a long-running process; Docker's restart policy would merely replace launchd as its supervisor. The members would still execute on the host, so the component that runs untrusted work gains no isolation at all — only bridged does, and bridged is the least risky part. ### The design this implies for the vhost stage **One vhost holds bridged, herdr, the members and the project together.** One filesystem, so no path contract exists to break, and the isolation is real because the members are inside it. **`herdr --remote <vhost>` is then the operator's attach path**, not bridged's transport. The fleet runs on the vhost; you attach your local terminal to it over SSH. That is exactly what the flag was built for. ``` LAPTOP VHOST ────── ───── herdr --remote <vhost> ──SSH──► herdr server │ unix socket ▼ fleetd ──► members, worktrees, project (all one filesystem) ``` ### Consequences for this epic - No containerisation work belongs in CB-621..CB-627. The vhost stage supersedes it. - When the vhost provisioning ticket is written, it must place **herdr and fleetd on the same host as the worktrees**, and name `herdr --remote` as the documented operator access path. - If a container is ever wanted, the only coherent version puts **herdr inside it too**. Half-containerising fleetd adds a bind-mount contract at four seams and buys a restart policy that launchd already provides (installed and verified 2026-08-22: `kill -9` → restarted in 2s).
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#125