This plan was written 2026-09-05 and has been sitting untracked in the
working tree since, so nobody but this machine could read it and nothing
recorded that it existed. plans/ is a tracked directory here.
Committed with a status section measured today over a read-only ssh survey,
because a five-day-old plan committed as-is would read as current:
- Phases 1 to 3 are done. Both systemd user units exist, are active and
enabled, linger is on, one java process (so the old restart.sh
double-daemon problem is gone), the live fleetd.yaml is on the host, and
the jar was rebuilt 2026-09-10 02:10 UTC.
- The checkout has drifted again: main at 4887731, 77 commits behind
origin/main, 0 ahead. So fleet01's daemon runs code from before this
week's merges.
- Phases 6 and 7 are not done. The Mac still runs the daemon this fleet
uses, and fleet01 has 2 weeks of uptime, so no reboot proof exists.
- Section 9 is out of date: claude is installed on fleet01 and a fleet01
lead is live on the coordination channel, so the headless-login blocker
it names is solved.
Also records fleet01's four profiles with their weights, because
placement: weighted plus gx at 100 and xf at 80 means an unqualified spawn
there almost never lands on a claude-code profile.
And one search that found the opposite of what I expected, written down so
the next session does not repeat it: opencode IS installed on fleet01, at
~/.opencode/bin/opencode, but only on the INTERACTIVE PATH. zsh -ic finds
it, zsh -lc does not. A herdr pane is a non-login interactive zsh and
fleetd types the launch command into the pane, so gx and xf are not broken
by this. The daemon's own login-shell ExecStart cannot see it, which is the
mirror of the trap that ExecStart exists to fix: credentials live on the
login side, ~/.opencode/bin on the interactive side.
21 KiB
Plan — move the fleet back onto fleet01
Goal. Stop the fleet depending on a laptop that sleeps.
Written 2026-09-05. Rewritten the same day after the operator pointed out that fleet01 is a VM and already has the repo. They were right, and my first draft was wrong in an important way: this is not a stand-up. The whole fleet already ran on fleet01 in August. It was abandoned, not attempted.
Every fact below was measured on 2026-09-05. Where I did not measure something, the text says so.
Status, re-measured 2026-09-10 — phases 1 to 3 are DONE, and §9 is out of date
This plan is prior art now, not a to-do list. Every number in this section came from a read-only
survey over ssh fleet01 on 2026-09-10. Delete this section and the plan once fleet01 is the
fleet's only daemon — at that point the plan has been executed and stops being useful.
| Plan item | State on 2026-09-10 | Command that re-measures it |
|---|---|---|
| Phase 1, refresh + build | done, then drifted. Checkout is on main at 4887731, 77 commits behind origin/main, 0 ahead. fleetd/target/fleetd.jar exists, built 2026-09-10 02:10 UTC. So it was rebuilt, and main has moved since. |
git -C ~/LTMS/fleetd rev-list --count HEAD..origin/main |
| Phase 2, port the config | done. fleetd/fleetd.yaml exists on the host, with bind, profiles, configReload, fleet, health, lifecycle, guard, memberCredentials, broker and coordinator all present. |
test -f ~/LTMS/fleetd/fleetd/fleetd.yaml |
| Phase 3, supervision | done. ~/.config/systemd/user/fleetd.service and herdr.service both exist, both active and enabled. loginctl show-user ltms -p Linger prints yes. Exactly 1 java process, so the double-daemon problem in the old restart.sh is not present. |
systemctl --user is-active fleetd herdr |
| Phase 4-5, reachable and a member proven | partly. /healthz answers 200 on loopback and reports {"protocol":19,"version":"0.8.0"}, matching the herdr pin. 1 herdr socket present. I did not spawn a member from here, so "a member on fleet01 opens a PR" is still unproven by me. |
curl -s http://127.0.0.1:8765/healthz on the host |
| Phase 6-7, cutover and reboot proof | not done. The Mac still runs its own daemon and is still this fleet's lead. Uptime on fleet01 is 2 weeks 2 days, so no reboot proof has been taken since the units were installed. | uptime -p on the host |
| §9 "not moving the lead yet" | out of date. claude is installed at /home/ltms/.local/bin/claude and a fleet01 lead is live — it reaches this session over the coordination channel. So the headless-login blocker named in §9 is solved. |
ssh fleet01 'command -v claude' |
The profiles fleet01 actually offers a member
Measured from the live fleetd.yaml on the host, with placement: weighted:
| profile | kind | model | weight | maxLoad |
|---|---|---|---|---|
gx |
opencode | gx/deepseek-v4-flash |
100 | 2 |
xf |
opencode | opencode/mimo-v2.5-free |
80 | 5 |
local |
claude-code | deepseek-v4-flash |
10 | 2 |
opus |
claude-code | claude-opus-5 |
0 | 1 |
weighted spreads by ratio across every profile with a free slot, so an unqualified spawn on
fleet01 lands on gx or xf almost every time. Pass profile explicitly there, as the canonical
block already says.
The PATH split on fleet01, which is not the one I expected
I went looking for a defect and found the opposite, so this is written down to stop the next session repeating the search.
opencode is installed at /home/ltms/.opencode/bin/opencode. Whether a shell can see it depends
on which kind of shell it is:
zsh -ic 'command -v opencode' -> /home/ltms/.opencode/bin/opencode (1 PATH entry)
zsh -lc 'command -v opencode' -> nothing (0 PATH entries)
So it is on the interactive PATH (.zshrc), not the login one. Two consequences, and they
point in opposite directions:
- A herdr pane on Linux is a plain non-login interactive zsh, so a pane can launch
opencode. fleetd types the launch command into the pane rather than exec'ing it, sogxandxfare not broken by this. I have not spawned one to confirm, so that is inference from the shell measurement plus the typing behaviour, not an end-to-end result. - The daemon itself runs
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"— a login shell, on purpose, because credentials live in.zprofile. Its own PATH has 10 entries and none containsopencode.
The general shape: the login shell and the interactive shell see different PATHs, and which one
matters depends on whether fleetd types a command or execs it. Credentials live on the login
side; ~/.opencode/bin lives on the interactive side. Anything fleetd must exec itself is
invisible to it if it lives only on the interactive PATH — and that is the mirror of the trap the
login-shell ExecStart was added to fix.
1. My first draft was wrong — read this before the rest
I wrote "fleetd is not deployed on fleet01". That was wrong, and I got there by looking in one place and concluding about all of them. Three times:
| I checked | I concluded | What is actually true |
|---|---|---|
~/LTMS/claude-bridge |
"no repo checkout" | the repo is at ~/LTMS/fleetd — the directory follows the renamed repo, and my own notes record that rename |
~/.config/herdr/herdr.sock |
"herdr never ran" | the socket lives at ~/.config/herdr/sessions/fleet01/herdr.sock — a named session, and its server log runs to Aug 28 |
| both of the above | "fleetd is not deployed" | ~/LTMS/fleetd/fleetd-run/ holds restart.sh, start-herdr.sh and a fleetd.out from Aug 24 |
The lesson is the one already written down here: enumerate one channel, conclude about all of them. A single-path check is not a survey.
2. What already worked on fleet01, proven from its own log
~/LTMS/fleetd/fleetd-run/fleetd.out covers 06:07 to 16:48 on 2026-08-24. It shows:
4 distinct panes pane=c2504101-... w1:pC w1:pE
4 worktrees created and removed /home/ltms/LTMS/.fleet-worktrees/{05f2a7-4,dd9f51-1,f498dd-2,fdb522-3}
4 allow-list decisions "memberCredentials allow-list: pane w1:pC allowed 16 of 32 environment variables"
AMQP on 127.0.0.1:5672 "AMQP connection recovered; cleared held replies for fresh redelivery"
the full turn machinery SessionManager transitions, ReplyPushLoop, CompletionResolver fallback
So on fleet01, already: fleetd listened, herdr made panes, members spawned into git worktrees, the credential allow-list fed them their environment, and the broker link was loopback.
That last point is the whole reason for this move. The AMQP resets in section 3 cannot happen to a loopback connection.
The two traps I was going to design around are already solved there
My first draft named these as the biggest risks. Both were already handled in August:
Trap A — systemd sources no login shell. restart.sh already starts through one, and says why in
its own comment:
# 1. Start java from a LOGIN shell (zsh -lc). ~/.zprofile is where the credentials live, and a
# non-login shell starts the daemon fine with an empty AI_GATEWAY_TOKEN -- a failure that
# stays invisible until a member actually needs it.
setsid zsh -lc "exec java -jar target/bridged.jar fleetd.yaml" < /dev/null > "$RUN/fleetd.out" 2>&1 &
Trap B — a Linux herdr pane is a plain zsh, so members get no credentials. Not a problem, and not
for the reason I assumed. Members do not inherit from the pane's shell profile — fleetd hands them an
allow-list. The log proves it ran: allowed 16 of 32 environment variables, four times. The old
fleetd.yaml has a memberCredentials: block that configures it.
The headless pty trap is solved too. start-herdr.sh carries the fix and the explanation:
# Why the size matters: herdr creates each pane sized to the attached client's view. Started
# under a pty with no winsize, the client reports 0x0, and every pane.split / workspace.create
# then fails with "ghostty error -2" -- libghostty refusing a 0x0 surface.
cat > /tmp/herdr-inner.sh <<'INNER'
stty rows 50 cols 200 2>/dev/null || true
exec herdr --session fleet01
INNER
setsid script -qfec /tmp/herdr-inner.sh /dev/null < /dev/null > /dev/null 2>&1 &
3. Why we are moving
The daemon runs on a Mac laptop. On battery it idle-sleeps after one minute:
pmset -g custom -> Battery Power: sleep 1
AC Power: sleep 0
Since the last restart: 16 Connection reset events and 6 ERROR lines, all on the AMQP link. I
matched every one of the 16 to the nearest sleep or wake event in pmset -g log. Largest gap 55
seconds; most under 20. Not one reset lacked a nearby sleep or wake. The broker never sees a network
fault — it sees the client stop sending heartbeats, then closes.
Every reconnect worked, so no message was lost. The lost messages are not the problem. The problem is that a member mid-turn freezes with the host, and a long worker turn with nobody typing is exactly the case that goes idle.
4. How the two hosts relate today
This answers "why is fleet01 related to this Mac at all?"
Mac laptop fleet01 (KVM/QEMU guest)
┌────────────────────────────┐ ┌──────────────────────┐
│ Claude Code lead │ │ LavinMQ :5672 │
│ fleetd 127.0.0.1:8765 │ AMQP over │ vhost /mac │
│ herdr │──Tailscale────>│ vhost /fleet01 │
│ members + worktrees │ utun4, 1280 │ │
└────────────────────────────┘ │ (fleetd idle since │
│ Aug 24) │
└──────────────────────┘
Today fleet01 runs only the broker. Everything else — daemon, herdr, members — is on the Mac. The
single link between them is the Mac's fleetd opening AMQP to 10.10.20.13:5672 across Tailscale.
So the errors I reported were the Mac's client dying when the Mac slept, not fleet01 failing. fleet01 was healthy throughout: the container is up 11 days and its log shows a clean heartbeat timeout each time, which is what a broker sees when a client vanishes.
After the move that arrow becomes loopback and the whole class of problem is gone.
5. The shape we are restoring
flowchart LR
subgraph MAC["Mac laptop (free to sleep)"]
LEAD["Claude Code lead"]
TUN["ssh -N -L"]
end
subgraph F01["fleet01 (KVM guest, always on)"]
FD["fleetd<br/>127.0.0.1:8765"]
HD["herdr --session fleet01"]
WT["members<br/>~/LTMS/.fleet-worktrees"]
MQ["LavinMQ<br/>127.0.0.1:5672"]
end
LEAD --> TUN
TUN -->|"ssh over Tailscale"| FD
FD -->|"unix socket"| HD
HD --> WT
FD -->|"AMQP, loopback"| MQ
WT -->|"AMQP, loopback"| MQ
The lead stays on the Mac. Everything that must survive a sleep is already able to run on fleet01.
Two constraints fix this shape:
- fleetd, herdr and the worktrees must share one filesystem. The herdr link is a Unix socket plus
absolute path strings.
herdr --remoteis terminal attach, not a transport. - fleetd fails fast on a non-loopback bind without token auth, by its own design. So we do not
expose
:8765to the10.10.20.0/24LAN. An SSH tunnel keeps the bind on loopback and needs no new secret.
Why the lead stays on the Mac for now. A session not in a herdr pane resolves as primary, so a
Mac-side session over the tunnel works with nothing new. What we give up: async ticket nudges type
into the lead's pane, and fleetd cannot type into a pane on another host — so wait:false tickets
stop nudging and I poll instead. Moving the lead as well is section 9; its hard part is
authenticating Claude Code on a headless box, which has nothing to do with sleep and must not block
this.
6. The actual gap
Everything below is what stands between "it ran in August" and "it runs supervised today".
| # | Gap | Measured state |
|---|---|---|
| 1 | Checkout is stale | branch cb-634-ide-mcp, HEAD 7655f1b (2026-08-24), 290 commits behind origin/main, 0 ahead |
| 2 | Never rebuilt after the rename | no target/ anywhere; the scripts still say bridged.jar and cd bridged, but the tree is now fleetd/ and fleetd.jar |
| 3 | No live fleetd.yaml |
gitignored, so not in git. Three backups exist under bridged/ — the newest is fleetd.yaml.bak-cb634-pin, 5512 bytes, and it is clean of inline secrets (0 inline passwords, 6 uses of uriEnv/tokenEnv) |
| 4 | No supervision | Linger=no; zero systemd user unit files. Only the hand-rolled restart.sh / start-herdr.sh |
| 5 | Nothing running now | no herdr process, no fleetd, no answer on :8765/healthz |
Untracked files in the checkout: .idea/, fleetd-run/, docs/CB-634-Worker-IDE-Worktree.md, and the
three yaml backups. All are untracked, none modified, and 7655f1b is already an ancestor of
origin/main — so nothing is lost by updating the branch. Keep fleetd-run/ and the backups; they
are the prior art this plan is built on.
The old config's keys, which tell us what to port
bind: herdrSocket: /home/ltms/.config/herdr/sessions/fleet01/herdr.sock
profiles: placement: weighted
configReload: fleet: health:
lifecycle: guard: worktreeRoot: /home/ltms/LTMS/.fleet-worktrees
memberCredentials: broker:
herdrSocket already points at the named-session path, and memberCredentials is already
configured. Those two are what made members work.
What the repo already has for this
deploy/fleetd.service exists and is written for Linux. Three lines need fleet01's real paths:
ExecStart names /usr/lib/jvm/temurin-25-jdk/bin/java (fleet01 has /usr/bin/java), the PATH
names /usr/share/maven/bin (fleet01 has /usr/bin/mvn), and WorkingDirectory assumes
%h/src/claude-bridge. It also declares After=herdr.service — and no herdr.service exists in
deploy/. Writing that unit, from start-herdr.sh, is the one genuinely new piece of code here.
7. Phases
flowchart TD
P1["1. Refresh<br/>update + build"]
P2["2. Config<br/>port fleetd.yaml"]
P3["3. Supervise<br/>linger + 2 units"]
P4["4. Reachability<br/>tunnel, primary"]
P5["5. Prove a member"]
P6["6. Cutover"]
P7["7. Reboot proof"]
P1 --> P2 --> P3 --> P4 --> P5 --> P6 --> P7
Phase 1 — refresh the checkout. Update to origin/main (290 commits). Keep the untracked
fleetd-run/ and the yaml backups. Then mvn clean install, unpiped — a pipe hides a failure
behind a zero exit. Check: fleetd/target/fleetd.jar exists and the suite is green.
Phase 2 — port the config. Write fleetd/fleetd.yaml from bridged/fleetd.yaml.bak-cb634-pin,
updating it for the rename and 290 commits of config changes. Diff its keys against
fleetd/fleetd.example.yaml on current main, key by key, and say what changed. Check: the daemon
starts and journalctl ... | grep 'startup secret' reports no MISSING.
Phase 3 — supervision. This is the part that never existed. loginctl enable-linger ltms; write
deploy/herdr.service from start-herdr.sh, keeping the stty sizing; fix the three path lines in
deploy/fleetd.service and keep the login-shell ExecStart; install both under
~/.config/systemd/user/. Secrets go in a systemctl --user edit drop-in or a 0600
EnvironmentFile, never in the committed unit. Check: log out of every ssh session, log back in,
and confirm the socket and the daemon are still there. That is what lingering is for, and it is the
check people skip.
Phase 4 — reachability. ssh -N -L <port>:127.0.0.1:8765 fleet01 from the Mac. Use a different
local port for the first test so the Mac's own daemon on :8765 is untouched and the whole test is
reversible. Point the lead's .mcp.json at it — that file is --skip-worktree and must never be
committed. Wrap the tunnel in autossh or a launchd KeepAlive, because the Mac still sleeps.
Check: fleet_whoami answers primary. If it answers worker, the fleet.leaders.*.tab pin does
not match — a known demotion, not a network fault.
Phase 5 — prove a member. healthz can be green while every spawn fails, so only a real spawn
proves the herdr link. Spawn one member, then give it a real unit ending in a pushed PR — the push is
what proves WORKER_GITEA_TOKEN resolved. Confirm the allow-list line still appears
(allowed N of M environment variables) and that N is what you expect. Check: a member on fleet01
opens a PR.
Phase 6 — cutover. Drain the Mac fleet properly first: fleet_list, fleet_poll anything still
wanted, then fleet_stop each member — a restart drops in-flight tickets and a member's report is
gone with its ticket. Then stop the Mac's launchd agent. This is the migration itself, not a change to
the Mac's settings, and it is reversible in one command.
Phase 7 — reboot proof. Reboot fleet01. Without touching anything: socket present, healthz
answering, fleet_whoami still primary, one spawn works. Until this passes, "supervised" is a claim.
8. Risks
| Risk | Why it bites | What this plan does |
|---|---|---|
| 290 commits of config drift | the old yaml predates the rename and much else; a silently defaulted key turns a feature off with no error | phase 2 diffs key-by-key against current fleetd.example.yaml |
| A new config key gets silently dropped | FleetConfig's back-compat constructor ladder can absorb an arity change, so a new key compiles and is defaulted away |
separate ticket already in flight; matters most here because fleet01 gets a hand-edited yaml |
| Nothing supervises herdr | deploy/fleetd.service depends on a unit that does not exist |
phase 3 writes it from the working script; phase 7 proves it |
| Linger left off | everything dies at logout and looks fine until then | phase 3, checked by logging out |
| Wrong JDK/Maven path in the unit | fleetd propagates its PATH to every member, so a bad PATH means no member can build | three lines fixed in phase 3, proven by phase 5 |
| Headless pty with no winsize | ghostty error -2, reported three steps later as a spawn failure |
the stty fix is carried into herdr.service |
| Port 8765 collides during the test | both daemons want the same local port | phase 4 uses a different local port first |
| Tunnel dies when the Mac sleeps | same sleep, far smaller blast radius — it interrupts my session, not members | autossh/launchd KeepAlive |
| Lead demoted to worker | the tab pin no longer matches | phase 4's check is fleet_whoami |
| Upgrading herdr | 0.8.0 is protocol 19, pinned on purpose — 0.8.2 is protocol 20 and fleetd has no version handshake | do not upgrade herdr during this work; both hosts measured at 0.8.0 today |
placement: tab headless |
fails for the same 0x0 reason as ghostty error -2 |
fleet01 must keep placement: pane, as its August config did |
| Profile launch settings are deferred | editing placement: and waiting for the 10s config watch does nothing — the launcher holds a startup snapshot |
restart the daemon after those keys, do not wait for the reload |
9. Deliberately not doing
- No changes to the Mac's power settings or host config. The point is to stop depending on it.
- Not touching the leftover
bridged-lavinmqcontainer on the Mac. It is unused and harmless. - Not moving the broker. It is already on fleet01 and already the durable one. After the move its connection becomes loopback, which is the fix.
- Not building a second fleet. This is a move. Two daemons on one herdr session kill each other's members.
- Not moving the lead yet. That needs Claude Code authenticated on a headless Ubuntu box and a
fleet.leaders.*.tabpin on its pane. It buys back pane nudges. It has nothing to do with sleep, so it must not hold up phases 1–7. Note that fleet01 already carries anopusprofile defined purely so the lead slot resolves; it cannot spawn until someone runsclaudeand completes/loginon the host.
10. Open questions for the operator
- vhost — keep
/mac, or rename now the fleet is not on the Mac? Renaming loses the existing queues. (The old fleet01 config used its own; phase 2 must settle which this fleet owns.) - Fallback week after cutover, or stop the Mac daemon for good?
- Delete or keep the stale
cb-634-ide-mcpbranch on fleet01 once the checkout is updated? Its tip is already in main, so nothing is lost either way.
The repo path question from the first draft is answered: /home/ltms/LTMS/fleetd, which already
exists. deploy/fleetd.service should be pointed there rather than the reverse.