diff --git a/plans/fleet01-standup/plan.md b/plans/fleet01-standup/plan.md new file mode 100644 index 0000000..6f5c3d2 --- /dev/null +++ b/plans/fleet01-standup/plan.md @@ -0,0 +1,369 @@ +# Plan — move the fleet back onto fleet01 + +**Goal.** Stop the fleet depending on a laptop that sleeps. + +**Written** 2026-09-05. **Rewritten the same day** after the operator pointed out that fleet01 is a VM +and already has the repo. They were right, and my first draft was wrong in an important way: this is +not a stand-up. **The whole fleet already ran on fleet01 in August.** It was abandoned, not attempted. + +Every fact below was measured on 2026-09-05. Where I did not measure something, the text says so. + +--- + +## Status, re-measured 2026-09-10 — phases 1 to 3 are DONE, and §9 is out of date + +This plan is prior art now, not a to-do list. Every number in this section came from a read-only +survey over `ssh fleet01` on 2026-09-10. **Delete this section and the plan once fleet01 is the +fleet's only daemon** — at that point the plan has been executed and stops being useful. + +| Plan item | State on 2026-09-10 | Command that re-measures it | +|---|---|---| +| Phase 1, refresh + build | **done, then drifted.** Checkout is on `main` at `4887731`, **77 commits behind** `origin/main`, 0 ahead. `fleetd/target/fleetd.jar` exists, built 2026-09-10 02:10 UTC. So it was rebuilt, and main has moved since. | `git -C ~/LTMS/fleetd rev-list --count HEAD..origin/main` | +| Phase 2, port the config | **done.** `fleetd/fleetd.yaml` exists on the host, with `bind`, `profiles`, `configReload`, `fleet`, `health`, `lifecycle`, `guard`, `memberCredentials`, `broker` and `coordinator` all present. | `test -f ~/LTMS/fleetd/fleetd/fleetd.yaml` | +| Phase 3, supervision | **done.** `~/.config/systemd/user/fleetd.service` and `herdr.service` both exist, both `active` and `enabled`. `loginctl show-user ltms -p Linger` prints `yes`. Exactly 1 java process, so the double-daemon problem in the old `restart.sh` is not present. | `systemctl --user is-active fleetd herdr` | +| Phase 4-5, reachable and a member proven | **partly.** `/healthz` answers 200 on loopback and reports `{"protocol":19,"version":"0.8.0"}`, matching the herdr pin. 1 herdr socket present. I did not spawn a member from here, so "a member on fleet01 opens a PR" is still unproven by me. | `curl -s http://127.0.0.1:8765/healthz` on the host | +| Phase 6-7, cutover and reboot proof | **not done.** The Mac still runs its own daemon and is still this fleet's lead. Uptime on fleet01 is 2 weeks 2 days, so no reboot proof has been taken since the units were installed. | `uptime -p` on the host | +| §9 "not moving the lead yet" | **out of date.** `claude` is installed at `/home/ltms/.local/bin/claude` and a fleet01 lead is live — it reaches this session over the coordination channel. So the headless-login blocker named in §9 is solved. | `ssh fleet01 'command -v claude'` | + +### The profiles fleet01 actually offers a member + +Measured from the live `fleetd.yaml` on the host, with `placement: weighted`: + +| profile | kind | model | weight | maxLoad | +|---|---|---|---|---| +| `gx` | opencode | `gx/deepseek-v4-flash` | 100 | 2 | +| `xf` | opencode | `opencode/mimo-v2.5-free` | 80 | 5 | +| `local` | claude-code | `deepseek-v4-flash` | 10 | 2 | +| `opus` | claude-code | `claude-opus-5` | 0 | 1 | + +`weighted` spreads by ratio across every profile with a free slot, so an **unqualified** spawn on +fleet01 lands on `gx` or `xf` almost every time. Pass `profile` explicitly there, as the canonical +block already says. + +### The PATH split on fleet01, which is not the one I expected + +I went looking for a defect and found the opposite, so this is written down to stop the next +session repeating the search. + +`opencode` is installed at `/home/ltms/.opencode/bin/opencode`. Whether a shell can see it depends +on which kind of shell it is: + +``` +zsh -ic 'command -v opencode' -> /home/ltms/.opencode/bin/opencode (1 PATH entry) +zsh -lc 'command -v opencode' -> nothing (0 PATH entries) +``` + +So it is on the **interactive** PATH (`.zshrc`), not the login one. Two consequences, and they +point in opposite directions: + +- A herdr pane on Linux is a plain non-login interactive zsh, so a pane **can** launch `opencode`. + fleetd types the launch command into the pane rather than exec'ing it, so `gx` and `xf` are not + broken by this. I have not spawned one to confirm, so that is inference from the shell + measurement plus the typing behaviour, not an end-to-end result. +- The daemon itself runs `ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` — + a **login** shell, on purpose, because credentials live in `.zprofile`. Its own PATH has 10 + entries and none contains `opencode`. + +**The general shape: the login shell and the interactive shell see different PATHs, and which one +matters depends on whether fleetd types a command or execs it.** Credentials live on the login +side; `~/.opencode/bin` lives on the interactive side. Anything fleetd must exec itself is +invisible to it if it lives only on the interactive PATH — and that is the mirror of the trap the +login-shell `ExecStart` was added to fix. + +--- + +--- + +## 1. My first draft was wrong — read this before the rest + +I wrote "fleetd is not deployed on fleet01". That was wrong, and I got there by looking in one place +and concluding about all of them. Three times: + +| I checked | I concluded | What is actually true | +|---|---|---| +| `~/LTMS/claude-bridge` | "no repo checkout" | the repo is at `~/LTMS/fleetd` — the directory follows the **renamed** repo, and my own notes record that rename | +| `~/.config/herdr/herdr.sock` | "herdr never ran" | the socket lives at `~/.config/herdr/sessions/fleet01/herdr.sock` — a **named session**, and its server log runs to Aug 28 | +| both of the above | "fleetd is not deployed" | `~/LTMS/fleetd/fleetd-run/` holds `restart.sh`, `start-herdr.sh` and a `fleetd.out` from **Aug 24** | + +The lesson is the one already written down here: enumerate one channel, conclude about all of them. +A single-path check is not a survey. + +--- + +## 2. What already worked on fleet01, proven from its own log + +`~/LTMS/fleetd/fleetd-run/fleetd.out` covers 06:07 to 16:48 on 2026-08-24. It shows: + +``` +4 distinct panes pane=c2504101-... w1:pC w1:pE +4 worktrees created and removed /home/ltms/LTMS/.fleet-worktrees/{05f2a7-4,dd9f51-1,f498dd-2,fdb522-3} +4 allow-list decisions "memberCredentials allow-list: pane w1:pC allowed 16 of 32 environment variables" +AMQP on 127.0.0.1:5672 "AMQP connection recovered; cleared held replies for fresh redelivery" +the full turn machinery SessionManager transitions, ReplyPushLoop, CompletionResolver fallback +``` + +So on fleet01, already: fleetd listened, herdr made panes, members spawned into git worktrees, the +credential allow-list fed them their environment, and the broker link was **loopback**. + +That last point is the whole reason for this move. The AMQP resets in section 3 cannot happen to a +loopback connection. + +### The two traps I was going to design around are already solved there + +My first draft named these as the biggest risks. Both were already handled in August: + +**Trap A — systemd sources no login shell.** `restart.sh` already starts through one, and says why in +its own comment: + +```sh +# 1. Start java from a LOGIN shell (zsh -lc). ~/.zprofile is where the credentials live, and a +# non-login shell starts the daemon fine with an empty AI_GATEWAY_TOKEN -- a failure that +# stays invisible until a member actually needs it. +setsid zsh -lc "exec java -jar target/bridged.jar fleetd.yaml" < /dev/null > "$RUN/fleetd.out" 2>&1 & +``` + +**Trap B — a Linux herdr pane is a plain zsh, so members get no credentials.** Not a problem, and not +for the reason I assumed. Members do not inherit from the pane's shell profile — fleetd hands them an +allow-list. The log proves it ran: `allowed 16 of 32 environment variables`, four times. The old +`fleetd.yaml` has a `memberCredentials:` block that configures it. + +**The headless pty trap is solved too.** `start-herdr.sh` carries the fix and the explanation: + +```sh +# Why the size matters: herdr creates each pane sized to the attached client's view. Started +# under a pty with no winsize, the client reports 0x0, and every pane.split / workspace.create +# then fails with "ghostty error -2" -- libghostty refusing a 0x0 surface. +cat > /tmp/herdr-inner.sh <<'INNER' +stty rows 50 cols 200 2>/dev/null || true +exec herdr --session fleet01 +INNER +setsid script -qfec /tmp/herdr-inner.sh /dev/null < /dev/null > /dev/null 2>&1 & +``` + +--- + +## 3. Why we are moving + +The daemon runs on a Mac laptop. On battery it idle-sleeps after **one minute**: + +``` +pmset -g custom -> Battery Power: sleep 1 + AC Power: sleep 0 +``` + +Since the last restart: 16 `Connection reset` events and 6 ERROR lines, all on the AMQP link. I +matched every one of the 16 to the nearest sleep or wake event in `pmset -g log`. Largest gap **55 +seconds**; most under 20. Not one reset lacked a nearby sleep or wake. The broker never sees a network +fault — it sees the client stop sending heartbeats, then closes. + +Every reconnect worked, so no message was lost. **The lost messages are not the problem.** The problem +is that a member mid-turn freezes with the host, and a long worker turn with nobody typing is exactly +the case that goes idle. + +--- + +## 4. How the two hosts relate today + +This answers "why is fleet01 related to this Mac at all?" + +``` + Mac laptop fleet01 (KVM/QEMU guest) + ┌────────────────────────────┐ ┌──────────────────────┐ + │ Claude Code lead │ │ LavinMQ :5672 │ + │ fleetd 127.0.0.1:8765 │ AMQP over │ vhost /mac │ + │ herdr │──Tailscale────>│ vhost /fleet01 │ + │ members + worktrees │ utun4, 1280 │ │ + └────────────────────────────┘ │ (fleetd idle since │ + │ Aug 24) │ + └──────────────────────┘ +``` + +**Today fleet01 runs only the broker.** Everything else — daemon, herdr, members — is on the Mac. The +single link between them is the Mac's fleetd opening AMQP to `10.10.20.13:5672` across Tailscale. + +So the errors I reported were **the Mac's client dying when the Mac slept**, not fleet01 failing. +fleet01 was healthy throughout: the container is up 11 days and its log shows a clean heartbeat +timeout each time, which is what a broker sees when a client vanishes. + +After the move that arrow becomes loopback and the whole class of problem is gone. + +--- + +## 5. The shape we are restoring + +```mermaid +flowchart LR + subgraph MAC["Mac laptop (free to sleep)"] + LEAD["Claude Code lead"] + TUN["ssh -N -L"] + end + subgraph F01["fleet01 (KVM guest, always on)"] + FD["fleetd
127.0.0.1:8765"] + HD["herdr --session fleet01"] + WT["members
~/LTMS/.fleet-worktrees"] + MQ["LavinMQ
127.0.0.1:5672"] + end + LEAD --> TUN + TUN -->|"ssh over Tailscale"| FD + FD -->|"unix socket"| HD + HD --> WT + FD -->|"AMQP, loopback"| MQ + WT -->|"AMQP, loopback"| MQ +``` + +*The lead stays on the Mac. Everything that must survive a sleep is already able to run on fleet01.* + +Two constraints fix this shape: + +1. **fleetd, herdr and the worktrees must share one filesystem.** The herdr link is a Unix socket plus + absolute path strings. `herdr --remote` is terminal attach, not a transport. +2. **fleetd fails fast on a non-loopback bind without token auth**, by its own design. So we do not + expose `:8765` to the `10.10.20.0/24` LAN. An SSH tunnel keeps the bind on loopback and needs no + new secret. + +**Why the lead stays on the Mac for now.** A session not in a herdr pane resolves as `primary`, so a +Mac-side session over the tunnel works with nothing new. What we give up: async ticket nudges type +into the lead's pane, and fleetd cannot type into a pane on another host — so `wait:false` tickets +stop nudging and I poll instead. Moving the lead as well is section 9; its hard part is +authenticating Claude Code on a headless box, which has nothing to do with sleep and must not block +this. + +--- + +## 6. The actual gap + +Everything below is what stands between "it ran in August" and "it runs supervised today". + +| # | Gap | Measured state | +|---|---|---| +| 1 | **Checkout is stale** | branch `cb-634-ide-mcp`, HEAD `7655f1b` (2026-08-24), **290 commits behind** `origin/main`, 0 ahead | +| 2 | **Never rebuilt after the rename** | no `target/` anywhere; the scripts still say `bridged.jar` and `cd bridged`, but the tree is now `fleetd/` and `fleetd.jar` | +| 3 | **No live `fleetd.yaml`** | gitignored, so not in git. Three backups exist under `bridged/` — the newest is `fleetd.yaml.bak-cb634-pin`, 5512 bytes, and it is **clean of inline secrets** (0 inline passwords, 6 uses of `uriEnv`/`tokenEnv`) | +| 4 | **No supervision** | `Linger=no`; **zero** systemd user unit files. Only the hand-rolled `restart.sh` / `start-herdr.sh` | +| 5 | **Nothing running now** | no herdr process, no fleetd, no answer on `:8765/healthz` | + +Untracked files in the checkout: `.idea/`, `fleetd-run/`, `docs/CB-634-Worker-IDE-Worktree.md`, and the +three yaml backups. All are **untracked, none modified**, and `7655f1b` is already an ancestor of +`origin/main` — so nothing is lost by updating the branch. Keep `fleetd-run/` and the backups; they +are the prior art this plan is built on. + +### The old config's keys, which tell us what to port + +``` +bind: herdrSocket: /home/ltms/.config/herdr/sessions/fleet01/herdr.sock +profiles: placement: weighted +configReload: fleet: health: +lifecycle: guard: worktreeRoot: /home/ltms/LTMS/.fleet-worktrees +memberCredentials: broker: +``` + +`herdrSocket` already points at the **named-session** path, and `memberCredentials` is already +configured. Those two are what made members work. + +### What the repo already has for this + +`deploy/fleetd.service` exists and is written for Linux. Three lines need fleet01's real paths: +`ExecStart` names `/usr/lib/jvm/temurin-25-jdk/bin/java` (fleet01 has `/usr/bin/java`), the `PATH` +names `/usr/share/maven/bin` (fleet01 has `/usr/bin/mvn`), and `WorkingDirectory` assumes +`%h/src/claude-bridge`. It also declares `After=herdr.service` — **and no `herdr.service` exists in +`deploy/`**. Writing that unit, from `start-herdr.sh`, is the one genuinely new piece of code here. + +--- + +## 7. Phases + +```mermaid +flowchart TD + P1["1. Refresh
update + build"] + P2["2. Config
port fleetd.yaml"] + P3["3. Supervise
linger + 2 units"] + P4["4. Reachability
tunnel, primary"] + P5["5. Prove a member"] + P6["6. Cutover"] + P7["7. Reboot proof"] + P1 --> P2 --> P3 --> P4 --> P5 --> P6 --> P7 +``` + +**Phase 1 — refresh the checkout.** Update to `origin/main` (290 commits). Keep the untracked +`fleetd-run/` and the yaml backups. Then `mvn clean install`, **unpiped** — a pipe hides a failure +behind a zero exit. *Check:* `fleetd/target/fleetd.jar` exists and the suite is green. + +**Phase 2 — port the config.** Write `fleetd/fleetd.yaml` from `bridged/fleetd.yaml.bak-cb634-pin`, +updating it for the rename and 290 commits of config changes. Diff its keys against +`fleetd/fleetd.example.yaml` on current main, key by key, and say what changed. *Check:* the daemon +starts and `journalctl ... | grep 'startup secret'` reports **no** `MISSING`. + +**Phase 3 — supervision.** This is the part that never existed. `loginctl enable-linger ltms`; write +`deploy/herdr.service` from `start-herdr.sh`, keeping the `stty` sizing; fix the three path lines in +`deploy/fleetd.service` and keep the login-shell `ExecStart`; install both under +`~/.config/systemd/user/`. Secrets go in a `systemctl --user edit` drop-in or a 0600 +`EnvironmentFile`, never in the committed unit. *Check:* log out of every ssh session, log back in, +and confirm the socket and the daemon are still there. That is what lingering is for, and it is the +check people skip. + +**Phase 4 — reachability.** `ssh -N -L :127.0.0.1:8765 fleet01` from the Mac. Use a **different +local port** for the first test so the Mac's own daemon on `:8765` is untouched and the whole test is +reversible. Point the lead's `.mcp.json` at it — that file is `--skip-worktree` and must never be +committed. Wrap the tunnel in `autossh` or a launchd `KeepAlive`, because the Mac still sleeps. +*Check:* `fleet_whoami` answers `primary`. If it answers `worker`, the `fleet.leaders.*.tab` pin does +not match — a known demotion, not a network fault. + +**Phase 5 — prove a member.** `healthz` can be green while every spawn fails, so only a real spawn +proves the herdr link. Spawn one member, then give it a real unit ending in a pushed PR — the push is +what proves `WORKER_GITEA_TOKEN` resolved. Confirm the allow-list line still appears +(`allowed N of M environment variables`) and that N is what you expect. *Check:* a member on fleet01 +opens a PR. + +**Phase 6 — cutover.** Drain the Mac fleet properly first: `fleet_list`, `fleet_poll` anything still +wanted, then `fleet_stop` each member — a restart drops in-flight tickets and a member's report is +gone with its ticket. Then stop the Mac's launchd agent. This is the migration itself, not a change to +the Mac's settings, and it is reversible in one command. + +**Phase 7 — reboot proof.** Reboot fleet01. Without touching anything: socket present, healthz +answering, `fleet_whoami` still `primary`, one spawn works. Until this passes, "supervised" is a claim. + +--- + +## 8. Risks + +| Risk | Why it bites | What this plan does | +|---|---|---| +| **290 commits of config drift** | the old yaml predates the rename and much else; a silently defaulted key turns a feature off with no error | phase 2 diffs key-by-key against current `fleetd.example.yaml` | +| **A new config key gets silently dropped** | `FleetConfig`'s back-compat constructor ladder can absorb an arity change, so a new key compiles and is defaulted away | separate ticket already in flight; matters most here because fleet01 gets a hand-edited yaml | +| **Nothing supervises herdr** | `deploy/fleetd.service` depends on a unit that does not exist | phase 3 writes it from the working script; phase 7 proves it | +| **Linger left off** | everything dies at logout and looks fine until then | phase 3, checked by logging out | +| **Wrong JDK/Maven path in the unit** | fleetd propagates its PATH to every member, so a bad PATH means no member can build | three lines fixed in phase 3, proven by phase 5 | +| **Headless pty with no winsize** | `ghostty error -2`, reported three steps later as a spawn failure | the `stty` fix is carried into `herdr.service` | +| **Port 8765 collides during the test** | both daemons want the same local port | phase 4 uses a different local port first | +| **Tunnel dies when the Mac sleeps** | same sleep, far smaller blast radius — it interrupts my session, not members | `autossh`/launchd `KeepAlive` | +| **Lead demoted to worker** | the tab pin no longer matches | phase 4's check is `fleet_whoami` | +| **Upgrading herdr** | 0.8.0 is **protocol 19**, pinned on purpose — 0.8.2 is protocol 20 and fleetd has **no version handshake** | do not upgrade herdr during this work; both hosts measured at 0.8.0 today | +| **`placement: tab` headless** | fails for the same 0x0 reason as `ghostty error -2` | fleet01 must keep `placement: pane`, as its August config did | +| **Profile launch settings are deferred** | editing `placement:` and waiting for the 10s config watch does nothing — the launcher holds a startup snapshot | restart the daemon after those keys, do not wait for the reload | + +--- + +## 9. Deliberately not doing + +- **No changes to the Mac's power settings or host config.** The point is to stop depending on it. +- **Not touching the leftover `bridged-lavinmq` container** on the Mac. It is unused and harmless. +- **Not moving the broker.** It is already on fleet01 and already the durable one. After the move its + connection becomes loopback, which is the fix. +- **Not building a second fleet.** This is a move. Two daemons on one herdr session kill each other's + members. +- **Not moving the lead yet.** That needs Claude Code authenticated on a headless Ubuntu box and a + `fleet.leaders.*.tab` pin on its pane. It buys back pane nudges. It has nothing to do with sleep, so + it must not hold up phases 1–7. Note that fleet01 already carries an `opus` profile defined purely so the lead slot resolves; it cannot spawn until someone runs `claude` and completes `/login` on the host. + +--- + +## 10. Open questions for the operator + +1. **vhost** — keep `/mac`, or rename now the fleet is not on the Mac? Renaming loses the existing + queues. (The old fleet01 config used its own; phase 2 must settle which this fleet owns.) +2. **Fallback week** after cutover, or stop the Mac daemon for good? +3. **Delete or keep the stale `cb-634-ide-mcp` branch** on fleet01 once the checkout is updated? Its + tip is already in main, so nothing is lost either way. + +The repo path question from the first draft is answered: **`/home/ltms/LTMS/fleetd`**, which already +exists. `deploy/fleetd.service` should be pointed there rather than the reverse.