Files
Dai Ha f5e02fedd6
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m58s
plans: commit the fleet01 move plan, with today's state measured on the host
This plan was written 2026-09-05 and has been sitting untracked in the
working tree since, so nobody but this machine could read it and nothing
recorded that it existed. plans/ is a tracked directory here.

Committed with a status section measured today over a read-only ssh survey,
because a five-day-old plan committed as-is would read as current:

- Phases 1 to 3 are done. Both systemd user units exist, are active and
  enabled, linger is on, one java process (so the old restart.sh
  double-daemon problem is gone), the live fleetd.yaml is on the host, and
  the jar was rebuilt 2026-09-10 02:10 UTC.
- The checkout has drifted again: main at 4887731, 77 commits behind
  origin/main, 0 ahead. So fleet01's daemon runs code from before this
  week's merges.
- Phases 6 and 7 are not done. The Mac still runs the daemon this fleet
  uses, and fleet01 has 2 weeks of uptime, so no reboot proof exists.
- Section 9 is out of date: claude is installed on fleet01 and a fleet01
  lead is live on the coordination channel, so the headless-login blocker
  it names is solved.

Also records fleet01's four profiles with their weights, because
placement: weighted plus gx at 100 and xf at 80 means an unqualified spawn
there almost never lands on a claude-code profile.

And one search that found the opposite of what I expected, written down so
the next session does not repeat it: opencode IS installed on fleet01, at
~/.opencode/bin/opencode, but only on the INTERACTIVE PATH. zsh -ic finds
it, zsh -lc does not. A herdr pane is a non-login interactive zsh and
fleetd types the launch command into the pane, so gx and xf are not broken
by this. The daemon's own login-shell ExecStart cannot see it, which is the
mirror of the trap that ExecStart exists to fix: credentials live on the
login side, ~/.opencode/bin on the interactive side.
2026-09-10 18:53:11 +07:00

21 KiB
Raw Permalink Blame History

Plan — move the fleet back onto fleet01

Goal. Stop the fleet depending on a laptop that sleeps.

Written 2026-09-05. Rewritten the same day after the operator pointed out that fleet01 is a VM and already has the repo. They were right, and my first draft was wrong in an important way: this is not a stand-up. The whole fleet already ran on fleet01 in August. It was abandoned, not attempted.

Every fact below was measured on 2026-09-05. Where I did not measure something, the text says so.


Status, re-measured 2026-09-10 — phases 1 to 3 are DONE, and §9 is out of date

This plan is prior art now, not a to-do list. Every number in this section came from a read-only survey over ssh fleet01 on 2026-09-10. Delete this section and the plan once fleet01 is the fleet's only daemon — at that point the plan has been executed and stops being useful.

Plan item State on 2026-09-10 Command that re-measures it
Phase 1, refresh + build done, then drifted. Checkout is on main at 4887731, 77 commits behind origin/main, 0 ahead. fleetd/target/fleetd.jar exists, built 2026-09-10 02:10 UTC. So it was rebuilt, and main has moved since. git -C ~/LTMS/fleetd rev-list --count HEAD..origin/main
Phase 2, port the config done. fleetd/fleetd.yaml exists on the host, with bind, profiles, configReload, fleet, health, lifecycle, guard, memberCredentials, broker and coordinator all present. test -f ~/LTMS/fleetd/fleetd/fleetd.yaml
Phase 3, supervision done. ~/.config/systemd/user/fleetd.service and herdr.service both exist, both active and enabled. loginctl show-user ltms -p Linger prints yes. Exactly 1 java process, so the double-daemon problem in the old restart.sh is not present. systemctl --user is-active fleetd herdr
Phase 4-5, reachable and a member proven partly. /healthz answers 200 on loopback and reports {"protocol":19,"version":"0.8.0"}, matching the herdr pin. 1 herdr socket present. I did not spawn a member from here, so "a member on fleet01 opens a PR" is still unproven by me. curl -s http://127.0.0.1:8765/healthz on the host
Phase 6-7, cutover and reboot proof not done. The Mac still runs its own daemon and is still this fleet's lead. Uptime on fleet01 is 2 weeks 2 days, so no reboot proof has been taken since the units were installed. uptime -p on the host
§9 "not moving the lead yet" out of date. claude is installed at /home/ltms/.local/bin/claude and a fleet01 lead is live — it reaches this session over the coordination channel. So the headless-login blocker named in §9 is solved. ssh fleet01 'command -v claude'

The profiles fleet01 actually offers a member

Measured from the live fleetd.yaml on the host, with placement: weighted:

profile kind model weight maxLoad
gx opencode gx/deepseek-v4-flash 100 2
xf opencode opencode/mimo-v2.5-free 80 5
local claude-code deepseek-v4-flash 10 2
opus claude-code claude-opus-5 0 1

weighted spreads by ratio across every profile with a free slot, so an unqualified spawn on fleet01 lands on gx or xf almost every time. Pass profile explicitly there, as the canonical block already says.

The PATH split on fleet01, which is not the one I expected

I went looking for a defect and found the opposite, so this is written down to stop the next session repeating the search.

opencode is installed at /home/ltms/.opencode/bin/opencode. Whether a shell can see it depends on which kind of shell it is:

zsh -ic 'command -v opencode'   ->  /home/ltms/.opencode/bin/opencode   (1 PATH entry)
zsh -lc 'command -v opencode'   ->  nothing                             (0 PATH entries)

So it is on the interactive PATH (.zshrc), not the login one. Two consequences, and they point in opposite directions:

  • A herdr pane on Linux is a plain non-login interactive zsh, so a pane can launch opencode. fleetd types the launch command into the pane rather than exec'ing it, so gx and xf are not broken by this. I have not spawned one to confirm, so that is inference from the shell measurement plus the typing behaviour, not an end-to-end result.
  • The daemon itself runs ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml" — a login shell, on purpose, because credentials live in .zprofile. Its own PATH has 10 entries and none contains opencode.

The general shape: the login shell and the interactive shell see different PATHs, and which one matters depends on whether fleetd types a command or execs it. Credentials live on the login side; ~/.opencode/bin lives on the interactive side. Anything fleetd must exec itself is invisible to it if it lives only on the interactive PATH — and that is the mirror of the trap the login-shell ExecStart was added to fix.



1. My first draft was wrong — read this before the rest

I wrote "fleetd is not deployed on fleet01". That was wrong, and I got there by looking in one place and concluding about all of them. Three times:

I checked I concluded What is actually true
~/LTMS/claude-bridge "no repo checkout" the repo is at ~/LTMS/fleetd — the directory follows the renamed repo, and my own notes record that rename
~/.config/herdr/herdr.sock "herdr never ran" the socket lives at ~/.config/herdr/sessions/fleet01/herdr.sock — a named session, and its server log runs to Aug 28
both of the above "fleetd is not deployed" ~/LTMS/fleetd/fleetd-run/ holds restart.sh, start-herdr.sh and a fleetd.out from Aug 24

The lesson is the one already written down here: enumerate one channel, conclude about all of them. A single-path check is not a survey.


2. What already worked on fleet01, proven from its own log

~/LTMS/fleetd/fleetd-run/fleetd.out covers 06:07 to 16:48 on 2026-08-24. It shows:

4 distinct panes                     pane=c2504101-...  w1:pC  w1:pE
4 worktrees created and removed      /home/ltms/LTMS/.fleet-worktrees/{05f2a7-4,dd9f51-1,f498dd-2,fdb522-3}
4 allow-list decisions               "memberCredentials allow-list: pane w1:pC allowed 16 of 32 environment variables"
AMQP on 127.0.0.1:5672               "AMQP connection recovered; cleared held replies for fresh redelivery"
the full turn machinery              SessionManager transitions, ReplyPushLoop, CompletionResolver fallback

So on fleet01, already: fleetd listened, herdr made panes, members spawned into git worktrees, the credential allow-list fed them their environment, and the broker link was loopback.

That last point is the whole reason for this move. The AMQP resets in section 3 cannot happen to a loopback connection.

The two traps I was going to design around are already solved there

My first draft named these as the biggest risks. Both were already handled in August:

Trap A — systemd sources no login shell. restart.sh already starts through one, and says why in its own comment:

#   1. Start java from a LOGIN shell (zsh -lc). ~/.zprofile is where the credentials live, and a
#      non-login shell starts the daemon fine with an empty AI_GATEWAY_TOKEN -- a failure that
#      stays invisible until a member actually needs it.
setsid zsh -lc "exec java -jar target/bridged.jar fleetd.yaml" < /dev/null > "$RUN/fleetd.out" 2>&1 &

Trap B — a Linux herdr pane is a plain zsh, so members get no credentials. Not a problem, and not for the reason I assumed. Members do not inherit from the pane's shell profile — fleetd hands them an allow-list. The log proves it ran: allowed 16 of 32 environment variables, four times. The old fleetd.yaml has a memberCredentials: block that configures it.

The headless pty trap is solved too. start-herdr.sh carries the fix and the explanation:

# Why the size matters: herdr creates each pane sized to the attached client's view. Started
# under a pty with no winsize, the client reports 0x0, and every pane.split / workspace.create
# then fails with "ghostty error -2" -- libghostty refusing a 0x0 surface.
cat > /tmp/herdr-inner.sh <<'INNER'
stty rows 50 cols 200 2>/dev/null || true
exec herdr --session fleet01
INNER
setsid script -qfec /tmp/herdr-inner.sh /dev/null < /dev/null > /dev/null 2>&1 &

3. Why we are moving

The daemon runs on a Mac laptop. On battery it idle-sleeps after one minute:

pmset -g custom   ->  Battery Power: sleep 1
                      AC Power:      sleep 0

Since the last restart: 16 Connection reset events and 6 ERROR lines, all on the AMQP link. I matched every one of the 16 to the nearest sleep or wake event in pmset -g log. Largest gap 55 seconds; most under 20. Not one reset lacked a nearby sleep or wake. The broker never sees a network fault — it sees the client stop sending heartbeats, then closes.

Every reconnect worked, so no message was lost. The lost messages are not the problem. The problem is that a member mid-turn freezes with the host, and a long worker turn with nobody typing is exactly the case that goes idle.


4. How the two hosts relate today

This answers "why is fleet01 related to this Mac at all?"

   Mac laptop                                    fleet01  (KVM/QEMU guest)
   ┌────────────────────────────┐                ┌──────────────────────┐
   │ Claude Code lead           │                │ LavinMQ  :5672       │
   │ fleetd      127.0.0.1:8765 │  AMQP over     │   vhost /mac         │
   │ herdr                      │──Tailscale────>│   vhost /fleet01     │
   │ members + worktrees        │  utun4, 1280   │                      │
   └────────────────────────────┘                │ (fleetd idle since   │
                                                 │  Aug 24)             │
                                                 └──────────────────────┘

Today fleet01 runs only the broker. Everything else — daemon, herdr, members — is on the Mac. The single link between them is the Mac's fleetd opening AMQP to 10.10.20.13:5672 across Tailscale.

So the errors I reported were the Mac's client dying when the Mac slept, not fleet01 failing. fleet01 was healthy throughout: the container is up 11 days and its log shows a clean heartbeat timeout each time, which is what a broker sees when a client vanishes.

After the move that arrow becomes loopback and the whole class of problem is gone.


5. The shape we are restoring

flowchart LR
    subgraph MAC["Mac laptop (free to sleep)"]
        LEAD["Claude Code lead"]
        TUN["ssh -N -L"]
    end
    subgraph F01["fleet01 (KVM guest, always on)"]
        FD["fleetd<br/>127.0.0.1:8765"]
        HD["herdr --session fleet01"]
        WT["members<br/>~/LTMS/.fleet-worktrees"]
        MQ["LavinMQ<br/>127.0.0.1:5672"]
    end
    LEAD --> TUN
    TUN -->|"ssh over Tailscale"| FD
    FD -->|"unix socket"| HD
    HD --> WT
    FD -->|"AMQP, loopback"| MQ
    WT -->|"AMQP, loopback"| MQ

The lead stays on the Mac. Everything that must survive a sleep is already able to run on fleet01.

Two constraints fix this shape:

  1. fleetd, herdr and the worktrees must share one filesystem. The herdr link is a Unix socket plus absolute path strings. herdr --remote is terminal attach, not a transport.
  2. fleetd fails fast on a non-loopback bind without token auth, by its own design. So we do not expose :8765 to the 10.10.20.0/24 LAN. An SSH tunnel keeps the bind on loopback and needs no new secret.

Why the lead stays on the Mac for now. A session not in a herdr pane resolves as primary, so a Mac-side session over the tunnel works with nothing new. What we give up: async ticket nudges type into the lead's pane, and fleetd cannot type into a pane on another host — so wait:false tickets stop nudging and I poll instead. Moving the lead as well is section 9; its hard part is authenticating Claude Code on a headless box, which has nothing to do with sleep and must not block this.


6. The actual gap

Everything below is what stands between "it ran in August" and "it runs supervised today".

# Gap Measured state
1 Checkout is stale branch cb-634-ide-mcp, HEAD 7655f1b (2026-08-24), 290 commits behind origin/main, 0 ahead
2 Never rebuilt after the rename no target/ anywhere; the scripts still say bridged.jar and cd bridged, but the tree is now fleetd/ and fleetd.jar
3 No live fleetd.yaml gitignored, so not in git. Three backups exist under bridged/ — the newest is fleetd.yaml.bak-cb634-pin, 5512 bytes, and it is clean of inline secrets (0 inline passwords, 6 uses of uriEnv/tokenEnv)
4 No supervision Linger=no; zero systemd user unit files. Only the hand-rolled restart.sh / start-herdr.sh
5 Nothing running now no herdr process, no fleetd, no answer on :8765/healthz

Untracked files in the checkout: .idea/, fleetd-run/, docs/CB-634-Worker-IDE-Worktree.md, and the three yaml backups. All are untracked, none modified, and 7655f1b is already an ancestor of origin/main — so nothing is lost by updating the branch. Keep fleetd-run/ and the backups; they are the prior art this plan is built on.

The old config's keys, which tell us what to port

bind:            herdrSocket: /home/ltms/.config/herdr/sessions/fleet01/herdr.sock
profiles:        placement: weighted
configReload:    fleet:              health:
lifecycle:       guard:              worktreeRoot: /home/ltms/LTMS/.fleet-worktrees
memberCredentials:                   broker:

herdrSocket already points at the named-session path, and memberCredentials is already configured. Those two are what made members work.

What the repo already has for this

deploy/fleetd.service exists and is written for Linux. Three lines need fleet01's real paths: ExecStart names /usr/lib/jvm/temurin-25-jdk/bin/java (fleet01 has /usr/bin/java), the PATH names /usr/share/maven/bin (fleet01 has /usr/bin/mvn), and WorkingDirectory assumes %h/src/claude-bridge. It also declares After=herdr.service — and no herdr.service exists in deploy/. Writing that unit, from start-herdr.sh, is the one genuinely new piece of code here.


7. Phases

flowchart TD
    P1["1. Refresh<br/>update + build"]
    P2["2. Config<br/>port fleetd.yaml"]
    P3["3. Supervise<br/>linger + 2 units"]
    P4["4. Reachability<br/>tunnel, primary"]
    P5["5. Prove a member"]
    P6["6. Cutover"]
    P7["7. Reboot proof"]
    P1 --> P2 --> P3 --> P4 --> P5 --> P6 --> P7

Phase 1 — refresh the checkout. Update to origin/main (290 commits). Keep the untracked fleetd-run/ and the yaml backups. Then mvn clean install, unpiped — a pipe hides a failure behind a zero exit. Check: fleetd/target/fleetd.jar exists and the suite is green.

Phase 2 — port the config. Write fleetd/fleetd.yaml from bridged/fleetd.yaml.bak-cb634-pin, updating it for the rename and 290 commits of config changes. Diff its keys against fleetd/fleetd.example.yaml on current main, key by key, and say what changed. Check: the daemon starts and journalctl ... | grep 'startup secret' reports no MISSING.

Phase 3 — supervision. This is the part that never existed. loginctl enable-linger ltms; write deploy/herdr.service from start-herdr.sh, keeping the stty sizing; fix the three path lines in deploy/fleetd.service and keep the login-shell ExecStart; install both under ~/.config/systemd/user/. Secrets go in a systemctl --user edit drop-in or a 0600 EnvironmentFile, never in the committed unit. Check: log out of every ssh session, log back in, and confirm the socket and the daemon are still there. That is what lingering is for, and it is the check people skip.

Phase 4 — reachability. ssh -N -L <port>:127.0.0.1:8765 fleet01 from the Mac. Use a different local port for the first test so the Mac's own daemon on :8765 is untouched and the whole test is reversible. Point the lead's .mcp.json at it — that file is --skip-worktree and must never be committed. Wrap the tunnel in autossh or a launchd KeepAlive, because the Mac still sleeps. Check: fleet_whoami answers primary. If it answers worker, the fleet.leaders.*.tab pin does not match — a known demotion, not a network fault.

Phase 5 — prove a member. healthz can be green while every spawn fails, so only a real spawn proves the herdr link. Spawn one member, then give it a real unit ending in a pushed PR — the push is what proves WORKER_GITEA_TOKEN resolved. Confirm the allow-list line still appears (allowed N of M environment variables) and that N is what you expect. Check: a member on fleet01 opens a PR.

Phase 6 — cutover. Drain the Mac fleet properly first: fleet_list, fleet_poll anything still wanted, then fleet_stop each member — a restart drops in-flight tickets and a member's report is gone with its ticket. Then stop the Mac's launchd agent. This is the migration itself, not a change to the Mac's settings, and it is reversible in one command.

Phase 7 — reboot proof. Reboot fleet01. Without touching anything: socket present, healthz answering, fleet_whoami still primary, one spawn works. Until this passes, "supervised" is a claim.


8. Risks

Risk Why it bites What this plan does
290 commits of config drift the old yaml predates the rename and much else; a silently defaulted key turns a feature off with no error phase 2 diffs key-by-key against current fleetd.example.yaml
A new config key gets silently dropped FleetConfig's back-compat constructor ladder can absorb an arity change, so a new key compiles and is defaulted away separate ticket already in flight; matters most here because fleet01 gets a hand-edited yaml
Nothing supervises herdr deploy/fleetd.service depends on a unit that does not exist phase 3 writes it from the working script; phase 7 proves it
Linger left off everything dies at logout and looks fine until then phase 3, checked by logging out
Wrong JDK/Maven path in the unit fleetd propagates its PATH to every member, so a bad PATH means no member can build three lines fixed in phase 3, proven by phase 5
Headless pty with no winsize ghostty error -2, reported three steps later as a spawn failure the stty fix is carried into herdr.service
Port 8765 collides during the test both daemons want the same local port phase 4 uses a different local port first
Tunnel dies when the Mac sleeps same sleep, far smaller blast radius — it interrupts my session, not members autossh/launchd KeepAlive
Lead demoted to worker the tab pin no longer matches phase 4's check is fleet_whoami
Upgrading herdr 0.8.0 is protocol 19, pinned on purpose — 0.8.2 is protocol 20 and fleetd has no version handshake do not upgrade herdr during this work; both hosts measured at 0.8.0 today
placement: tab headless fails for the same 0x0 reason as ghostty error -2 fleet01 must keep placement: pane, as its August config did
Profile launch settings are deferred editing placement: and waiting for the 10s config watch does nothing — the launcher holds a startup snapshot restart the daemon after those keys, do not wait for the reload

9. Deliberately not doing

  • No changes to the Mac's power settings or host config. The point is to stop depending on it.
  • Not touching the leftover bridged-lavinmq container on the Mac. It is unused and harmless.
  • Not moving the broker. It is already on fleet01 and already the durable one. After the move its connection becomes loopback, which is the fix.
  • Not building a second fleet. This is a move. Two daemons on one herdr session kill each other's members.
  • Not moving the lead yet. That needs Claude Code authenticated on a headless Ubuntu box and a fleet.leaders.*.tab pin on its pane. It buys back pane nudges. It has nothing to do with sleep, so it must not hold up phases 1–7. Note that fleet01 already carries an opus profile defined purely so the lead slot resolves; it cannot spawn until someone runs claude and completes /login on the host.

10. Open questions for the operator

  1. vhost — keep /mac, or rename now the fleet is not on the Mac? Renaming loses the existing queues. (The old fleet01 config used its own; phase 2 must settle which this fleet owns.)
  2. Fallback week after cutover, or stop the Mac daemon for good?
  3. Delete or keep the stale cb-634-ide-mcp branch on fleet01 once the checkout is updated? Its tip is already in main, so nothing is lost either way.

The repo path question from the first draft is answered: /home/ltms/LTMS/fleetd, which already exists. deploy/fleetd.service should be pointed there rather than the reverse.