plans: commit the fleet01 move plan, with today's state measured on the host
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m58s

This plan was written 2026-09-05 and has been sitting untracked in the
working tree since, so nobody but this machine could read it and nothing
recorded that it existed. plans/ is a tracked directory here.

Committed with a status section measured today over a read-only ssh survey,
because a five-day-old plan committed as-is would read as current:

- Phases 1 to 3 are done. Both systemd user units exist, are active and
  enabled, linger is on, one java process (so the old restart.sh
  double-daemon problem is gone), the live fleetd.yaml is on the host, and
  the jar was rebuilt 2026-09-10 02:10 UTC.
- The checkout has drifted again: main at 4887731, 77 commits behind
  origin/main, 0 ahead. So fleet01's daemon runs code from before this
  week's merges.
- Phases 6 and 7 are not done. The Mac still runs the daemon this fleet
  uses, and fleet01 has 2 weeks of uptime, so no reboot proof exists.
- Section 9 is out of date: claude is installed on fleet01 and a fleet01
  lead is live on the coordination channel, so the headless-login blocker
  it names is solved.

Also records fleet01's four profiles with their weights, because
placement: weighted plus gx at 100 and xf at 80 means an unqualified spawn
there almost never lands on a claude-code profile.

And one search that found the opposite of what I expected, written down so
the next session does not repeat it: opencode IS installed on fleet01, at
~/.opencode/bin/opencode, but only on the INTERACTIVE PATH. zsh -ic finds
it, zsh -lc does not. A herdr pane is a non-login interactive zsh and
fleetd types the launch command into the pane, so gx and xf are not broken
by this. The daemon's own login-shell ExecStart cannot see it, which is the
mirror of the trap that ExecStart exists to fix: credentials live on the
login side, ~/.opencode/bin on the interactive side.
This commit is contained in:
Dai Ha
2026-09-10 18:53:11 +07:00
parent 92a96fcbd8
commit f5e02fedd6
+369
View File
@@ -0,0 +1,369 @@
# Plan — move the fleet back onto fleet01
**Goal.** Stop the fleet depending on a laptop that sleeps.
**Written** 2026-09-05. **Rewritten the same day** after the operator pointed out that fleet01 is a VM
and already has the repo. They were right, and my first draft was wrong in an important way: this is
not a stand-up. **The whole fleet already ran on fleet01 in August.** It was abandoned, not attempted.
Every fact below was measured on 2026-09-05. Where I did not measure something, the text says so.
---
## Status, re-measured 2026-09-10 — phases 1 to 3 are DONE, and §9 is out of date
This plan is prior art now, not a to-do list. Every number in this section came from a read-only
survey over `ssh fleet01` on 2026-09-10. **Delete this section and the plan once fleet01 is the
fleet's only daemon** — at that point the plan has been executed and stops being useful.
| Plan item | State on 2026-09-10 | Command that re-measures it |
|---|---|---|
| Phase 1, refresh + build | **done, then drifted.** Checkout is on `main` at `4887731`, **77 commits behind** `origin/main`, 0 ahead. `fleetd/target/fleetd.jar` exists, built 2026-09-10 02:10 UTC. So it was rebuilt, and main has moved since. | `git -C ~/LTMS/fleetd rev-list --count HEAD..origin/main` |
| Phase 2, port the config | **done.** `fleetd/fleetd.yaml` exists on the host, with `bind`, `profiles`, `configReload`, `fleet`, `health`, `lifecycle`, `guard`, `memberCredentials`, `broker` and `coordinator` all present. | `test -f ~/LTMS/fleetd/fleetd/fleetd.yaml` |
| Phase 3, supervision | **done.** `~/.config/systemd/user/fleetd.service` and `herdr.service` both exist, both `active` and `enabled`. `loginctl show-user ltms -p Linger` prints `yes`. Exactly 1 java process, so the double-daemon problem in the old `restart.sh` is not present. | `systemctl --user is-active fleetd herdr` |
| Phase 4-5, reachable and a member proven | **partly.** `/healthz` answers 200 on loopback and reports `{"protocol":19,"version":"0.8.0"}`, matching the herdr pin. 1 herdr socket present. I did not spawn a member from here, so "a member on fleet01 opens a PR" is still unproven by me. | `curl -s http://127.0.0.1:8765/healthz` on the host |
| Phase 6-7, cutover and reboot proof | **not done.** The Mac still runs its own daemon and is still this fleet's lead. Uptime on fleet01 is 2 weeks 2 days, so no reboot proof has been taken since the units were installed. | `uptime -p` on the host |
| §9 "not moving the lead yet" | **out of date.** `claude` is installed at `/home/ltms/.local/bin/claude` and a fleet01 lead is live — it reaches this session over the coordination channel. So the headless-login blocker named in §9 is solved. | `ssh fleet01 'command -v claude'` |
### The profiles fleet01 actually offers a member
Measured from the live `fleetd.yaml` on the host, with `placement: weighted`:
| profile | kind | model | weight | maxLoad |
|---|---|---|---|---|
| `gx` | opencode | `gx/deepseek-v4-flash` | 100 | 2 |
| `xf` | opencode | `opencode/mimo-v2.5-free` | 80 | 5 |
| `local` | claude-code | `deepseek-v4-flash` | 10 | 2 |
| `opus` | claude-code | `claude-opus-5` | 0 | 1 |
`weighted` spreads by ratio across every profile with a free slot, so an **unqualified** spawn on
fleet01 lands on `gx` or `xf` almost every time. Pass `profile` explicitly there, as the canonical
block already says.
### The PATH split on fleet01, which is not the one I expected
I went looking for a defect and found the opposite, so this is written down to stop the next
session repeating the search.
`opencode` is installed at `/home/ltms/.opencode/bin/opencode`. Whether a shell can see it depends
on which kind of shell it is:
```
zsh -ic 'command -v opencode' -> /home/ltms/.opencode/bin/opencode (1 PATH entry)
zsh -lc 'command -v opencode' -> nothing (0 PATH entries)
```
So it is on the **interactive** PATH (`.zshrc`), not the login one. Two consequences, and they
point in opposite directions:
- A herdr pane on Linux is a plain non-login interactive zsh, so a pane **can** launch `opencode`.
fleetd types the launch command into the pane rather than exec'ing it, so `gx` and `xf` are not
broken by this. I have not spawned one to confirm, so that is inference from the shell
measurement plus the typing behaviour, not an end-to-end result.
- The daemon itself runs `ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` —
a **login** shell, on purpose, because credentials live in `.zprofile`. Its own PATH has 10
entries and none contains `opencode`.
**The general shape: the login shell and the interactive shell see different PATHs, and which one
matters depends on whether fleetd types a command or execs it.** Credentials live on the login
side; `~/.opencode/bin` lives on the interactive side. Anything fleetd must exec itself is
invisible to it if it lives only on the interactive PATH — and that is the mirror of the trap the
login-shell `ExecStart` was added to fix.
---
---
## 1. My first draft was wrong — read this before the rest
I wrote "fleetd is not deployed on fleet01". That was wrong, and I got there by looking in one place
and concluding about all of them. Three times:
| I checked | I concluded | What is actually true |
|---|---|---|
| `~/LTMS/claude-bridge` | "no repo checkout" | the repo is at `~/LTMS/fleetd` — the directory follows the **renamed** repo, and my own notes record that rename |
| `~/.config/herdr/herdr.sock` | "herdr never ran" | the socket lives at `~/.config/herdr/sessions/fleet01/herdr.sock` — a **named session**, and its server log runs to Aug 28 |
| both of the above | "fleetd is not deployed" | `~/LTMS/fleetd/fleetd-run/` holds `restart.sh`, `start-herdr.sh` and a `fleetd.out` from **Aug 24** |
The lesson is the one already written down here: enumerate one channel, conclude about all of them.
A single-path check is not a survey.
---
## 2. What already worked on fleet01, proven from its own log
`~/LTMS/fleetd/fleetd-run/fleetd.out` covers 06:07 to 16:48 on 2026-08-24. It shows:
```
4 distinct panes pane=c2504101-... w1:pC w1:pE
4 worktrees created and removed /home/ltms/LTMS/.fleet-worktrees/{05f2a7-4,dd9f51-1,f498dd-2,fdb522-3}
4 allow-list decisions "memberCredentials allow-list: pane w1:pC allowed 16 of 32 environment variables"
AMQP on 127.0.0.1:5672 "AMQP connection recovered; cleared held replies for fresh redelivery"
the full turn machinery SessionManager transitions, ReplyPushLoop, CompletionResolver fallback
```
So on fleet01, already: fleetd listened, herdr made panes, members spawned into git worktrees, the
credential allow-list fed them their environment, and the broker link was **loopback**.
That last point is the whole reason for this move. The AMQP resets in section 3 cannot happen to a
loopback connection.
### The two traps I was going to design around are already solved there
My first draft named these as the biggest risks. Both were already handled in August:
**Trap A — systemd sources no login shell.** `restart.sh` already starts through one, and says why in
its own comment:
```sh
# 1. Start java from a LOGIN shell (zsh -lc). ~/.zprofile is where the credentials live, and a
# non-login shell starts the daemon fine with an empty AI_GATEWAY_TOKEN -- a failure that
# stays invisible until a member actually needs it.
setsid zsh -lc "exec java -jar target/bridged.jar fleetd.yaml" < /dev/null > "$RUN/fleetd.out" 2>&1 &
```
**Trap B — a Linux herdr pane is a plain zsh, so members get no credentials.** Not a problem, and not
for the reason I assumed. Members do not inherit from the pane's shell profile — fleetd hands them an
allow-list. The log proves it ran: `allowed 16 of 32 environment variables`, four times. The old
`fleetd.yaml` has a `memberCredentials:` block that configures it.
**The headless pty trap is solved too.** `start-herdr.sh` carries the fix and the explanation:
```sh
# Why the size matters: herdr creates each pane sized to the attached client's view. Started
# under a pty with no winsize, the client reports 0x0, and every pane.split / workspace.create
# then fails with "ghostty error -2" -- libghostty refusing a 0x0 surface.
cat > /tmp/herdr-inner.sh <<'INNER'
stty rows 50 cols 200 2>/dev/null || true
exec herdr --session fleet01
INNER
setsid script -qfec /tmp/herdr-inner.sh /dev/null < /dev/null > /dev/null 2>&1 &
```
---
## 3. Why we are moving
The daemon runs on a Mac laptop. On battery it idle-sleeps after **one minute**:
```
pmset -g custom -> Battery Power: sleep 1
AC Power: sleep 0
```
Since the last restart: 16 `Connection reset` events and 6 ERROR lines, all on the AMQP link. I
matched every one of the 16 to the nearest sleep or wake event in `pmset -g log`. Largest gap **55
seconds**; most under 20. Not one reset lacked a nearby sleep or wake. The broker never sees a network
fault — it sees the client stop sending heartbeats, then closes.
Every reconnect worked, so no message was lost. **The lost messages are not the problem.** The problem
is that a member mid-turn freezes with the host, and a long worker turn with nobody typing is exactly
the case that goes idle.
---
## 4. How the two hosts relate today
This answers "why is fleet01 related to this Mac at all?"
```
Mac laptop fleet01 (KVM/QEMU guest)
┌────────────────────────────┐ ┌──────────────────────┐
│ Claude Code lead │ │ LavinMQ :5672 │
│ fleetd 127.0.0.1:8765 │ AMQP over │ vhost /mac │
│ herdr │──Tailscale────>│ vhost /fleet01 │
│ members + worktrees │ utun4, 1280 │ │
└────────────────────────────┘ │ (fleetd idle since │
│ Aug 24) │
└──────────────────────┘
```
**Today fleet01 runs only the broker.** Everything else — daemon, herdr, members — is on the Mac. The
single link between them is the Mac's fleetd opening AMQP to `10.10.20.13:5672` across Tailscale.
So the errors I reported were **the Mac's client dying when the Mac slept**, not fleet01 failing.
fleet01 was healthy throughout: the container is up 11 days and its log shows a clean heartbeat
timeout each time, which is what a broker sees when a client vanishes.
After the move that arrow becomes loopback and the whole class of problem is gone.
---
## 5. The shape we are restoring
```mermaid
flowchart LR
subgraph MAC["Mac laptop (free to sleep)"]
LEAD["Claude Code lead"]
TUN["ssh -N -L"]
end
subgraph F01["fleet01 (KVM guest, always on)"]
FD["fleetd<br/>127.0.0.1:8765"]
HD["herdr --session fleet01"]
WT["members<br/>~/LTMS/.fleet-worktrees"]
MQ["LavinMQ<br/>127.0.0.1:5672"]
end
LEAD --> TUN
TUN -->|"ssh over Tailscale"| FD
FD -->|"unix socket"| HD
HD --> WT
FD -->|"AMQP, loopback"| MQ
WT -->|"AMQP, loopback"| MQ
```
*The lead stays on the Mac. Everything that must survive a sleep is already able to run on fleet01.*
Two constraints fix this shape:
1. **fleetd, herdr and the worktrees must share one filesystem.** The herdr link is a Unix socket plus
absolute path strings. `herdr --remote` is terminal attach, not a transport.
2. **fleetd fails fast on a non-loopback bind without token auth**, by its own design. So we do not
expose `:8765` to the `10.10.20.0/24` LAN. An SSH tunnel keeps the bind on loopback and needs no
new secret.
**Why the lead stays on the Mac for now.** A session not in a herdr pane resolves as `primary`, so a
Mac-side session over the tunnel works with nothing new. What we give up: async ticket nudges type
into the lead's pane, and fleetd cannot type into a pane on another host — so `wait:false` tickets
stop nudging and I poll instead. Moving the lead as well is section 9; its hard part is
authenticating Claude Code on a headless box, which has nothing to do with sleep and must not block
this.
---
## 6. The actual gap
Everything below is what stands between "it ran in August" and "it runs supervised today".
| # | Gap | Measured state |
|---|---|---|
| 1 | **Checkout is stale** | branch `cb-634-ide-mcp`, HEAD `7655f1b` (2026-08-24), **290 commits behind** `origin/main`, 0 ahead |
| 2 | **Never rebuilt after the rename** | no `target/` anywhere; the scripts still say `bridged.jar` and `cd bridged`, but the tree is now `fleetd/` and `fleetd.jar` |
| 3 | **No live `fleetd.yaml`** | gitignored, so not in git. Three backups exist under `bridged/` — the newest is `fleetd.yaml.bak-cb634-pin`, 5512 bytes, and it is **clean of inline secrets** (0 inline passwords, 6 uses of `uriEnv`/`tokenEnv`) |
| 4 | **No supervision** | `Linger=no`; **zero** systemd user unit files. Only the hand-rolled `restart.sh` / `start-herdr.sh` |
| 5 | **Nothing running now** | no herdr process, no fleetd, no answer on `:8765/healthz` |
Untracked files in the checkout: `.idea/`, `fleetd-run/`, `docs/CB-634-Worker-IDE-Worktree.md`, and the
three yaml backups. All are **untracked, none modified**, and `7655f1b` is already an ancestor of
`origin/main` — so nothing is lost by updating the branch. Keep `fleetd-run/` and the backups; they
are the prior art this plan is built on.
### The old config's keys, which tell us what to port
```
bind: herdrSocket: /home/ltms/.config/herdr/sessions/fleet01/herdr.sock
profiles: placement: weighted
configReload: fleet: health:
lifecycle: guard: worktreeRoot: /home/ltms/LTMS/.fleet-worktrees
memberCredentials: broker:
```
`herdrSocket` already points at the **named-session** path, and `memberCredentials` is already
configured. Those two are what made members work.
### What the repo already has for this
`deploy/fleetd.service` exists and is written for Linux. Three lines need fleet01's real paths:
`ExecStart` names `/usr/lib/jvm/temurin-25-jdk/bin/java` (fleet01 has `/usr/bin/java`), the `PATH`
names `/usr/share/maven/bin` (fleet01 has `/usr/bin/mvn`), and `WorkingDirectory` assumes
`%h/src/claude-bridge`. It also declares `After=herdr.service` — **and no `herdr.service` exists in
`deploy/`**. Writing that unit, from `start-herdr.sh`, is the one genuinely new piece of code here.
---
## 7. Phases
```mermaid
flowchart TD
P1["1. Refresh<br/>update + build"]
P2["2. Config<br/>port fleetd.yaml"]
P3["3. Supervise<br/>linger + 2 units"]
P4["4. Reachability<br/>tunnel, primary"]
P5["5. Prove a member"]
P6["6. Cutover"]
P7["7. Reboot proof"]
P1 --> P2 --> P3 --> P4 --> P5 --> P6 --> P7
```
**Phase 1 — refresh the checkout.** Update to `origin/main` (290 commits). Keep the untracked
`fleetd-run/` and the yaml backups. Then `mvn clean install`, **unpiped** — a pipe hides a failure
behind a zero exit. *Check:* `fleetd/target/fleetd.jar` exists and the suite is green.
**Phase 2 — port the config.** Write `fleetd/fleetd.yaml` from `bridged/fleetd.yaml.bak-cb634-pin`,
updating it for the rename and 290 commits of config changes. Diff its keys against
`fleetd/fleetd.example.yaml` on current main, key by key, and say what changed. *Check:* the daemon
starts and `journalctl ... | grep 'startup secret'` reports **no** `MISSING`.
**Phase 3 — supervision.** This is the part that never existed. `loginctl enable-linger ltms`; write
`deploy/herdr.service` from `start-herdr.sh`, keeping the `stty` sizing; fix the three path lines in
`deploy/fleetd.service` and keep the login-shell `ExecStart`; install both under
`~/.config/systemd/user/`. Secrets go in a `systemctl --user edit` drop-in or a 0600
`EnvironmentFile`, never in the committed unit. *Check:* log out of every ssh session, log back in,
and confirm the socket and the daemon are still there. That is what lingering is for, and it is the
check people skip.
**Phase 4 — reachability.** `ssh -N -L <port>:127.0.0.1:8765 fleet01` from the Mac. Use a **different
local port** for the first test so the Mac's own daemon on `:8765` is untouched and the whole test is
reversible. Point the lead's `.mcp.json` at it — that file is `--skip-worktree` and must never be
committed. Wrap the tunnel in `autossh` or a launchd `KeepAlive`, because the Mac still sleeps.
*Check:* `fleet_whoami` answers `primary`. If it answers `worker`, the `fleet.leaders.*.tab` pin does
not match — a known demotion, not a network fault.
**Phase 5 — prove a member.** `healthz` can be green while every spawn fails, so only a real spawn
proves the herdr link. Spawn one member, then give it a real unit ending in a pushed PR — the push is
what proves `WORKER_GITEA_TOKEN` resolved. Confirm the allow-list line still appears
(`allowed N of M environment variables`) and that N is what you expect. *Check:* a member on fleet01
opens a PR.
**Phase 6 — cutover.** Drain the Mac fleet properly first: `fleet_list`, `fleet_poll` anything still
wanted, then `fleet_stop` each member — a restart drops in-flight tickets and a member's report is
gone with its ticket. Then stop the Mac's launchd agent. This is the migration itself, not a change to
the Mac's settings, and it is reversible in one command.
**Phase 7 — reboot proof.** Reboot fleet01. Without touching anything: socket present, healthz
answering, `fleet_whoami` still `primary`, one spawn works. Until this passes, "supervised" is a claim.
---
## 8. Risks
| Risk | Why it bites | What this plan does |
|---|---|---|
| **290 commits of config drift** | the old yaml predates the rename and much else; a silently defaulted key turns a feature off with no error | phase 2 diffs key-by-key against current `fleetd.example.yaml` |
| **A new config key gets silently dropped** | `FleetConfig`'s back-compat constructor ladder can absorb an arity change, so a new key compiles and is defaulted away | separate ticket already in flight; matters most here because fleet01 gets a hand-edited yaml |
| **Nothing supervises herdr** | `deploy/fleetd.service` depends on a unit that does not exist | phase 3 writes it from the working script; phase 7 proves it |
| **Linger left off** | everything dies at logout and looks fine until then | phase 3, checked by logging out |
| **Wrong JDK/Maven path in the unit** | fleetd propagates its PATH to every member, so a bad PATH means no member can build | three lines fixed in phase 3, proven by phase 5 |
| **Headless pty with no winsize** | `ghostty error -2`, reported three steps later as a spawn failure | the `stty` fix is carried into `herdr.service` |
| **Port 8765 collides during the test** | both daemons want the same local port | phase 4 uses a different local port first |
| **Tunnel dies when the Mac sleeps** | same sleep, far smaller blast radius — it interrupts my session, not members | `autossh`/launchd `KeepAlive` |
| **Lead demoted to worker** | the tab pin no longer matches | phase 4's check is `fleet_whoami` |
| **Upgrading herdr** | 0.8.0 is **protocol 19**, pinned on purpose — 0.8.2 is protocol 20 and fleetd has **no version handshake** | do not upgrade herdr during this work; both hosts measured at 0.8.0 today |
| **`placement: tab` headless** | fails for the same 0x0 reason as `ghostty error -2` | fleet01 must keep `placement: pane`, as its August config did |
| **Profile launch settings are deferred** | editing `placement:` and waiting for the 10s config watch does nothing — the launcher holds a startup snapshot | restart the daemon after those keys, do not wait for the reload |
---
## 9. Deliberately not doing
- **No changes to the Mac's power settings or host config.** The point is to stop depending on it.
- **Not touching the leftover `bridged-lavinmq` container** on the Mac. It is unused and harmless.
- **Not moving the broker.** It is already on fleet01 and already the durable one. After the move its
connection becomes loopback, which is the fix.
- **Not building a second fleet.** This is a move. Two daemons on one herdr session kill each other's
members.
- **Not moving the lead yet.** That needs Claude Code authenticated on a headless Ubuntu box and a
`fleet.leaders.*.tab` pin on its pane. It buys back pane nudges. It has nothing to do with sleep, so
it must not hold up phases 1–7. Note that fleet01 already carries an `opus` profile defined purely so the lead slot resolves; it cannot spawn until someone runs `claude` and completes `/login` on the host.
---
## 10. Open questions for the operator
1. **vhost** — keep `/mac`, or rename now the fleet is not on the Mac? Renaming loses the existing
queues. (The old fleet01 config used its own; phase 2 must settle which this fleet owns.)
2. **Fallback week** after cutover, or stop the Mac daemon for good?
3. **Delete or keep the stale `cb-634-ide-mcp` branch** on fleet01 once the checkout is updated? Its
tip is already in main, so nothing is lost either way.
The repo path question from the first draft is answered: **`/home/ltms/LTMS/fleetd`**, which already
exists. `deploy/fleetd.service` should be pointed there rather than the reverse.